hypernetwork
超网络元学习、贝叶斯适应与任务参数关系建模
这些研究聚焦超网络作为元学习器、动态权重生成器或任务上下文编码器的基础机制,通过任务嵌入、层次化上下文、贝叶斯估计和参数关系建模实现少样本适应、跨任务迁移及快速优化;其中神经影像归一化工作体现了该机制在跨设备领域适应中的应用。
- Bayesian estimation and model averaging of convolutional neural networks by hypernetwork(Kenya Ukai, Takashi Matsubara, Kuniaki Uehara, 2019, Nonlinear Theory and Its Applications IEICE)
- Meta-Learning via Hypernetworks(Dominic Zhao, Johannes von Oswald, Seijin Kobayashi, J. Sacramento, B. Grewe, 2020, No journal)
- Hypernetwork Approach to Bayesian MAML (Student Abstract)(P. Borycki, Piotr Kubacki, Marcin Przewiezlikowski, Tomasz Kusmierczyk, Jacek Tabor, P. Spurek, 2025, AAAI Conference on Artificial Intelligence)
- Hypernetworks for Dynamic Weight Generation in Few-Shot Learning(T. Ramaprabha, 2026, Eduschool International Journal of Data Science and Machine learning (EIJDSML))
- Hierarchical Context-Based Meta-Learning for Stochastic LPV System Identification(Tianjun Su, Yawen Mao, Hanyan Huang, Chen Xu, 2026, 2026 IEEE 15th Data Driven Control and Learning Systems (DDCLS))
- Lifelong Learning of Task-Parameter Relationships for Knowledge Transfer(S. Srivastava, Mohammad Yaqub, K. Nandakumar, 2023, 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW))
- Spectral-conditioned hypernetworks for meta-learning neuroimaging normalization with concept-guided transparency(Hashim Ali, 2026, Advanced Neurology)
持续学习、参数高效适配与模型遗忘
这些文献围绕模型生命周期管理展开,利用超网络生成稀疏参数、前缀、任务特定权重或适配模块,以支持持续学习、参数高效微调、跨域知识迁移、少样本增量学习和无数据遗忘,并缓解灾难性遗忘。
- EvoCL: Continual Learning over Evolving Domains(Vishnuprasadh Kumaravelu, P. Srijith, Sunil Gupta, 2025, IEEE Workshop/Winter Conference on Applications of Computer Vision)
- Bridging pre-trained models to continual learning: A hypernetwork based framework with parameter-efficient fine-tuning techniques(Fengqian Ding, Chen Xu, Han Liu, Bin Zhou, Hong-Chao Zhou, 2024, Information Sciences)
- Task-Adaptive FFN Editing for Continual Blind Image Quality Assessment(Satish Maurya, Parimala Kancharla, 2026, International Conference on Signal Processing and Communications)
- Hy2CRE: Hypernetworks with Hybrid Data Augmentation for Continual Relation Extraction(Yang Zhang, Yidong Chen, 2025, IEEE International Joint Conference on Neural Network)
- Hypernetwork-Assisted Parameter-Efficient Fine-Tuning with Meta-Knowledge Distillation for Domain Knowledge Disentanglement(Changqun Li, Linlin Wang, X. Lin, Shi-Zhou Huang, Liang He, 2024, No journal)
- PerFRDiff: Personalised Weight Editing for Multiple Appropriate Facial Reaction Generation(Hengde Zhu, Xiangyu Kong, Weicheng Xie, Xin Huang, Linlin Shen, Lu Liu, Hatice Gunes, Siyang Song, 2024, ACM Multimedia)
- Self-Referential Meta-Learning for Continual Few-Shot Learning(Hossein Jamali, Sergiu M. Dascalu, Frederick C. Harris, 2026, International Conference on Software Engineering Research and Applications)
- A Hypernetwork Framework for Data-Free Unlearning and Continual Learning(Sayanta Adhikari, Vishnuprasadh Kumaravelu, P. Srijith, 2026, Machine-mediated learning)
- Growing a Brain with Sparsity-Inducing Generation for Continual Learning(Hyundong Jin, Gyeong-Hyeon Kim, Chanho Ahn, Eunwoo Kim, 2023, IEEE International Conference on Computer Vision)
- Supplementary Material of AdaPrefix++: Integrating Adapters, Prefixes and Hypernetwork for Continual Learning(Sayanta Adhikari, Dupati Srikar Chandra, P. Srijith, P. Wasnik, Naoyuki Oneo, No journal)
- AdaPrefix++: Integrating Adapters, Prefixes and Hypernetwork for Continual Learning(Sayanta Adhikari, Dupati Srikar Chandra, P. Srijith, P. Wasnik, Naoyuki Oneo, 2025, IEEE Workshop/Winter Conference on Applications of Computer Vision)
- A federated continuous learning framework based on hypernetworks and dual-domain distillation(Qi Chang, Yanyan Zhang, Nong Si, Zimeng Wang, Zikai Liu, Xiaoai Gu, 2025, International Conference on Conceptual Structures)
个性化联邦学习与异质客户端参数生成
这些研究均面向联邦学习中的客户端异质性、动态参与、个性化建模和通信约束,通过超网络生成客户端特定参数、联邦元模型或自适应量化模块,减少聚合造成的信息损失并提升跨客户端泛化与通信效率。
- An Adaptive Federated Service Evolution Framework for Industrial Operations Analytics Under Dynamic Participation(Dao-Bin Luo, Yu-Ru Liu, Weishan Zhang, Yuan-Ge Liu, Qiao Qiao, Ling-Zhao Meng, Lianyong Qi, 2026, 2026 IEEE International Conference on Web Services (ICWS))
- Heterogeneous Federated Dynamic Graph HyperNetwork for Image Classification(Liu Yang, Kegen Chen, Qilong Wang, Zheng Xu, Shiqiao Gu, Qinghua Hu, 2026, IEEE Transactions on Image Processing)
- FedQCab: A Common Plug-and-Play Layer-wise Quantization Tool for Communication-Efficient Federated Learning(Haizhou Du, Tonghao Chen, 2026, International Workshop on Quality of Service)
- Personalized Federated Learning With Adaptive Transformer Pruning and Hypernetwork-Driven Personalization in Wireless Networks(Moqbel Hamood, Abdullatif Albaseer, H. El-Sallabi, Mohamed M. Abdallah, Ala I. Al-Fuqaha, Bechir Hamdaoui, 2026, IEEE Transactions on Machine Learning in Communications and Networking)
- Toward Personalized Federated Meta-Learning With Constrained Hypernetwork on Non-IID Data(Lizhao Wu, Xiaoding Wang, Hui Lin, Xu Yang, Jiwu Shu, Xun Yi, Ibrahim Khalil, Albert Y. Zomaya, 2026, IEEE transactions on computers)
- Heterogeneous Federated Learning Based on Graph Hypernetwork(Zheng-Ning Xu, Liu Yang, Shiqiao Gu, 2023, International Conference on Artificial Neural Networks)
动态环境强化学习、调度控制与多智能体决策
这些研究将超网络嵌入强化学习、元强化学习、博弈决策和调度控制框架,根据环境状态、历史轨迹、市场或用户上下文以及智能体规模动态生成策略参数,重点解决非平稳环境适应、跨场景迁移和多智能体可扩展性问题。
- A meta-learning framework for cross-domain order scheduling in intelligent warehouses(Honglin Zhang, Jianling Chen, Jianing Ren, Yingjie Zhao, Yanyan Wang, Jing Chen, Bindong Gao, 2025, International Journal of Production Research)
- Scalable Multi-Agent Reinforcement Learning With Permutation-Free Networks(Hyunwoo Park, Baekryun Seong, Sang-Ki Ko, 2026, IEEE Access)
- Meta-Reinforcement Learning with Hypernetworks for Variable Speed Limit Control under Adverse Weather and Work Zones(Xiaojun Zhao, Gao-Qiang Zhang, Tao Wen, Bingshuo Chen, Xun-Abulikemu TuEr, Xiaodong Li, Zhao-Qing Li, 2026, Transportation Research Record)
- Population-Invariant MADRL for AoI-Aware UAV Trajectory Design and Communication Scheduling in Wireless Sensor Networks(Xuanhan Zhou, Jun Xiong, Hai-Tao Zhao, Chao Yan, Haijun Wang, Ji-Bo Wei, 2025, IEEE Internet of Things Journal)
- Teacher-apprentices RL (TARL): leveraging complex policy distribution through generative adversarial hypernetwork in reinforcement learning(S. Tang, Athirai Aravazhi Irissappane, F. Oliehoek, Jie Zhang, 2023, Autonomous Agents and Multi-Agent Systems)
- Context-Aware Meta-Reinforcement Learning for Intelligent Diverse Indoor HVAC Control(Song-Ling Liu, Jing Li, T. Buganza, 2025, IEEE International Joint Conference on Neural Network)
- Adaptive transparent cloaking tunnel enabled by Meta-Reinforcement-Learning Metasurfaces(Ji-Wei Zhao, Pei-Xuan Zhu, Zhi-Bin Wen, Fan Tang, Bin Zheng, Rongrong Zhu, Haoliang Qian, Chao Qian, Huan Lu, Hongsheng Chen, 2026, PhotoniX)
- GTH-Net: A Dynamic Game-Theoretic HyperNetwork for Non-Stationary Financial Time Series Forecasting(Fujie Chen, Chen Ding, 2026, Applied Sciences)
社会关系、图超图表示与高阶信息扩散
这些文献把超网络或超图作为高阶关系、群体协作和动态扩散的结构表示,研究社会关系、科学合作、交通行为、知识传播、图表示学习以及时空级联预测中的结构演化、链接预测和信息传播机制。
- A Knowledge Generation Model via the Hypernetwork(Jian-Guo Liu, Guang-Yong Yang, Zhao-Long Hu, 2014, PLoS ONE)
- Dynamical evolution behavior of scientific collaboration hypernetwork(Xiang-Bo Li, Gangjin Wang, Dai-Jun Wei, 2022, AIP Advances)
- Bayesian hypernetwork collaborates with time-difference evolutional network for temporal knowledge prediction(Pengpeng Shao, Yan Wen, Jianhua Tao, 2024, Neural Networks)
- Predicting hyperlinks via weighted hypernetwork loop structure(Hao Peng, Shuzhe Li, Dandan Zhao, Ming-Hong Zhong, Cheng-Chun Qian, Wei Wang, 2024, The European Physical Journal Special Topics)
- Information dissemination in dynamic hypernetwork(Xingpeng Jiang, Zhiping Wang, Wei Liu, 2019, Physica A: Statistical Mechanics and its Applications)
- Ex Post Path Choice Estimation for Urban Rail Systems Using Smart Card Data: An Aggregated Time-Space Hypernetwork Approach(Baichuan Mo, Zhenliang Ma, H. Koutsopoulos, Jinhua Zhao, 2022, Transportation Science)
- Transmission Mechanism and Influencing Factors of Green Behavior in Dynamic Multiplex Networks(Haofei Yin, Zhiping Wang, Zhaohui Xu, 2021, IEEE Access)
- Hypernetwork Representation Learning With the Transformation Strategy(Y. Zhu, Haixing Zhao, Jianqiang Huang, Xiaoying Wang, 2022, Proceedings of the 6th International Conference on High Performance Compilation, Computing and Communications)
- CasDacGCN: A Dynamic Attention-Calibrated Graph Convolutional Network for Information Popularity Prediction(Bo-Feng Zhang, Yan-Ling Zhu, Zhirong Zhang, Kaili Liao, Sen Niu, Bingchun Li, Haiyan Li, 2025, Entropy)
视觉生成、三维表示与多模态自适应融合
这些研究面向图像生成、超分辨率、神经辐射场、三维场景、材质外观、人脸生成及多模态视觉理解,通过内容、文本提示或场景上下文动态生成网络权重或融合参数,以提高生成质量、分辨率独立性和样本级适应能力。
- Enhancing Imaging Generation through Implicit Neural Representations and HyperNetwork for Spatial Variability(Jaehoon Cha, Siu-Lun Yeung, S. Dhanpal, J. Thiyagalingam, 2025, IEEE International Conference on Acoustics, Speech, and Signal Processing)
- Hyper-SNBRDF: Hypernetwork for Neural BRDF Using Sinusoidal Activation(Zhiqiang Li, Xu-Kun Shen, Xueyang Zhou, Yong Hu, Bowen Li, 2024, International Conference on 3D Vision)
- HyperNeRFGAN: Camera-Free 3D Scene Generation via Hypernetwork-Driven Neural Radiance Fields(Adam Kania, A. Kasymov, Jakub Kościukiewicz, Artur Górak, M. Mazur, Maciej Zięba, P. Spurek, 2025, International Conference on Data Science and Advanced Analytics)
- Learning Resolution-independent Image Representations(J. Hammer, Michael S. Gashler, 2018, IEEE International Conference on Cognitive Informatics and Cognitive Computing)
- Toward Open-World Text-Driven Face Generation and Manipulation via StyleGAN3(Zonglin Li, Zhaoxing Zhang, Pei-Qiang Liu, Qinglin Liu, Xin Sun, 2024, IEEE transactions on circuits and systems for video technology (Print))
- HyperBuff: Branched Per-Frame Neural Radiance Fields using HyperNetwork(Chengkun Lu, 2023, No journal)
- HyperSOR: Context-Aware Graph Hypernetwork for Salient Object Ranking(Minglang Qiao, Mai Xu, Lai Jiang, Peng Lei, Shijie Wen, Yunjin Chen, Leonid Sigal, 2024, IEEE Transactions on Pattern Analysis and Machine Intelligence)
- HyperSR: A Hypernetwork Based Framework for Image Super-Resolution(Divya Mishra, Ofek Finkelstein, O. Hadar, 2024, IEEE International Joint Conference on Neural Network)
- A Study on the Application of Using Hypernetwork and Low Rank Adaptation for Text-to-Image Generation Based on Diffusion Models(A. O. Levin, Yuri S. Belov, 2024, 2024 6th International Youth Conference on Radio Electronics, Electrical and Power Engineering (REEPE))
- HERA: Fusing Deliberative Reasoning and Adaptive Hypernetworks for Knowledge-Enhanced Multimodal Sentiment Analysis(Qingguang Li, Kai Zhao, Linlin Zhang, 2026, International Conference on Computer Supported Cooperative Work in Design)
动态视觉感知、三维配准与场景检测
这些工作关注输入相关的视觉感知参数生成,分别服务于无人机语义分割、三维点云配准和动态场景目标检测,通过空间位置、局部几何或场景条件调制网络权重,提升复杂环境中的识别与配准鲁棒性。
- Research on Deep Learning-based Semantic Segmentation Algorithm for UAV Images(Qiang Yan, G. Cheng, 2023, Proceedings of the 2023 7th International Conference on Electronic Information Technology and Computer Engineering)
- Adaptive point cloud registration method based on iterative weighted SVD with dynamic multiscale attention mechanism(Yuanhao Deng, Qiang Liu, Guorui Zhao, Yunpeng Wu, Rongliang Zhu, 2025, Measurement science and technology)
- Learning Dynamic Scene-Conditioned 3D Object Detectors(Yu Zheng, Yueqi Duan, Zongtai Li, Jie Zhou, Jiwen Lu, 2023, IEEE Transactions on Pattern Analysis and Machine Intelligence)
医学影像、脑功能超网络与临床智能分析
这些研究将超网络用于医学影像、脑功能连接、脑疾病分析、MRI运动校正和脑电特征选择,重点处理医学数据的设备差异、动态运动、个体异质性及高维特征选择问题。
- Spectral-conditioned hypernetworks for meta-learning neuroimaging normalization with concept-guided transparency(Hashim Ali, 2026, Advanced Neurology)
- Construction and Multiple Feature Classification Based on a High-Order Functional Hypernetwork on fMRI Data(Yao Li, Qifan Li, Tao Li, Zijing Zhou, Yong-Qing Xu, Yanli Yang, Jun-Jie Chen, Hao Guo, 2022, Frontiers in Neuroscience)
- Deep convolutional GAN and hypernet-based neural architecture search for brain tumor diagnosis detection and classification(S. Swathi, M. Rajalakshmi, 2026, PLoS ONE)
- Motion-Aware Neural Networks Improve Rigid Motion Correction of Accelerated Segmented Multislice MRI(Nalini M. Singh, Malte Hoffmann, E. Adalsteinsson, Bruce Fischl, Polina Golland, A. Dalca, Robert Frost*, 2023, Proceedings of the International Society for Magnetic Resonance in Medicine ... Scientific Meeting and Exhibition. International Society for Magnetic Resonance in Medicine. Scientific Meeting and Exhibition)
- A Hypernetwork-Based Feature Selection Method for EEG Classification(Dunmin Chen, Yu Tang, Zhiping Tan, 2025, International Journal of Cognitive Informatics and Natural Intelligence)
无线通信安全、信道处理与边缘推理自适应
这些文献面向无线通信、无线安全和边缘推理任务,根据无线信道、编码结构、设备位置、攻击状态或通信资源动态生成模型参数,重点提升信道估计、译码、定位、自干扰消除、对抗防御和边缘功率控制的自适应能力。
- HyperAdv: Dynamic Defense Against Adversarial Radio Frequency Machine Learning Systems(Milin Zhang, Michael DeLucia, A. Swami, Jonathan D. Ashdown, K. Turck, F. Restuccia, 2024, IEEE Military Communications Conference)
- Low-Complexity Compressive Channel Estimation for IRS-Aided mmWave Systems With Hypernetwork-Assisted LAMP Network(Wen-Chiao Tsai, Chi-Wei Chen, Chieh-Fang Teng, A. Wu, 2022, IEEE Communications Letters)
- Deep Hypernetwork-based Robust Localization in Millimeter-Wave Networks(Roman Klus, Jukka Talvitie, Benjamin W. Domae, D. Cabric, Mikko Valkama, 2024, IEEE International Symposium on Personal, Indoor and Mobile Radio Communications)
- Hypernetwork-Aided Channel Estimation for Integrated Data and Energy Transfer(Yushi Lei, Yusha Liu, Jie Hu, Kun Yang, 2025, IEEE Transactions on Vehicular Technology)
- Convolutional Autoencoders Coupled With Hypernetworks for Recognizing Attacks in 5G Networks and Beyond(Loukas Ilias, Stefanos Palmos, George Doukas, Afroditi Blika, G. Kiokes, Christos Ntanos, Dimitris Askounis, 2025, IEEE Open Journal of the Communications Society)
- Hypernetwork Based Model-Driven Channel Neural Decoding(Yuanhui Liang, C. Lam, Qingle Wu, B. Ng, Sio-Kei Im, 2024, IEEE Access)
- Hypernetwork-Based Adaptive Self-Interference Cancellation for Full-Duplex Wireless Communication Systems(S. Islam, Xin Ma, Chunxiao Chigan, 2024, 2024 IEEE International Conference on Communications Workshops (ICC Workshops))
- Power Control for Edge ML Inference With Hypernetwork Meta-Parameters(Jiaying Zhang, Qiushuo Hou, Guan-Ding Yu, 2025, IEEE Communications Letters)
物理信息建模、设备健康预测与工程逆问题
这些研究将超网络与物理信息神经网络、设备状态建模和工程逆问题结合,根据物理参数、运行工况或设备老化状态生成求解器和预测模型参数,从而提高跨工况泛化、少样本预测和物理一致性。
- Strategies for multi-case physics-informed neural networks for tube flows: a study using 2D flow scenarios(H. Wong, W. Chan, Binghuan Li, C. Yap, 2024, Scientific Reports)
- Few-shot RUL Prediction with A Hypernetwork Structure Incorporating Uncertainty Quantification and Calibration(Ying Wang, Fangyu Li, Di Wang, 2024, 2024 IEEE 20th International Conference on Automation Science and Engineering (CASE))
- Aging-Aware Hypernetwork LSTM With Multi-Constraint Physics-Informed Neural Network (AAH-LSTM-PINN) for Lithium-Ion Battery SOC Estimation(A. A. Isa, Sheik Mohammed Sulthan, M. Dani, Tan Soon Jiann, 2026, IEEE Access)
- Physics-Informed Neural Networks for Inverse Electromagnetic Problems(M. Baldan, P. di Barba, D. Lowther, 2023, IEEE transactions on magnetics)
跨工况故障诊断与物理耦合动态适配
这组研究专门针对复杂工况和跨域分布偏移下的故障诊断,通过域条件编码、频域动态适配和物理耦合特征调制生成条件化参数,以提高未见工况中的故障识别稳健性。
- Domain-conditioned feature modulation for cross-condition fault diagnosis in rotating machinery(Jie Zhang, Quan Zhou, Weigen Chen, Wenjie Zhou, 2026, Measurement science and technology)
- Physics-Coupled Frequency Dynamic Adaptation Network for Domain Generalized Underwater Object Detection(Linxuan Luo, Pan Mu, Cong Bai, 2025, ACM Multimedia)
神经架构搜索、硬件部署与计算性能建模
这些工作关注超网络在模型结构层面的作用,包括神经架构搜索、硬件感知部署、忆阻器并行计算、联合任务分类结构设计以及异构设备性能测量。共同目标是在大规模结构空间中共享或生成候选架构,并兼顾精度、资源和设备性能约束。
- Bayesian Differentiable Architecture Search for Efficient Domain Matching Fault Diagnosis(Zheng Zhou, Tianfu Li, Zilong Zhang, Zhibin Zhao, Chuang Sun, Ruqiang Yan, Xuefeng Chen, 2021, IEEE Transactions on Instrumentation and Measurement)
- Hardware-Aware Iterative One-Shot Neural Architecture Search With Adaptable Knowledge Distillation for Efficient Edge Computing(O. Chen, Yu-Xuan Chang, Chih-Yu Chung, Yanfu Cheng, Manh-Hung Ha, 2025, IEEE Access)
- HyperNetwork Designs for Improved Classification and Robust Meta-Learning(Sudarshan Babu, Pedro H. P. Savarese, Michael Maire, No journal)
- D-GHNAS for Joint Intent Classification and Slot Filling(Yanxi Tang, Jianzong Wang, Xiaoyang Qu, Nan Zhang, Jing Xiao, 2020, No journal)
- AdaptPerf: On Adaptive and Scalable Computing Power Measurement for Heterogeneous Devices(Zhuo Li, Chengxu Han, 2025, IEEE Internet of Things Journal)
- AdaptPerf: Adaptive Measurement for Computing Power of Heterogeneous Devices Based on NAS(Chengxu Han, Zhuo Li, 2024, IEEE International Conference on High Performance Computing and Communications)
语言模型、知识图谱与关系特定权重生成
这些研究面向自然语言处理和知识表示任务,通过语言、任务或关系类型生成动态权重,用于文本扩散模型、多任务语言适配和知识图谱关系特定参数建模,体现超网络对任务结构和语义关系的细粒度适配能力。
- On feedback connections in text-diffusion LLMs(A. Sandhu, Edward Kim, No journal)
- Language Adaptive Weight Generation for Multi-task Visual Grounding(2023, No journal)
- Generating relation-specific weights for ConvKB using a HyperNetwork architecture(Thanh-Huong Le, Duy Nguyen, Bac Le, 2023, Applied Intelligence)
个性化预测、药物协同与不确定性后验建模
这些研究将超网络用于推荐、药物协同预测和隐式后验估计等专门任务,根据用户、药物或模型不确定性信息生成个性化预测参数或后验表示,重点体现超网络在数据稀疏、个体差异和概率建模场景中的应用价值。
- HyperRS: Hypernetwork-Based Recommender System for the User Cold-Start Problem(Yu-Xun Lu, Kosuke Nakamura, R. Ichise, 2023, IEEE Access)
- Few-Shot Drug Synergy Prediction With a Prior-Guided Hypernetwork Architecture(Qing-Qing Zhang, Shao-Wu Zhang, Yue-Hua Feng, Jianyu Shi, 2023, IEEE Transactions on Pattern Analysis and Machine Intelligence)
- Hypernetwork-based Implicit Posterior Estimation and Model Averaging of Convolutional Neural Networks Kenya(Ukai, 2018, No journal)
合并后的文献可归纳为十四个相互并列的方向。整体来看,hypernetwork的核心作用是依据任务、样本、域、客户端、环境状态、物理条件或网络关系动态生成目标模型的权重、结构或适配模块。研究一方面集中于元学习、持续学习、联邦个性化和参数高效适配,另一方面扩展至视觉生成与感知、医学影像、无线通信、物理工程、故障诊断、神经架构搜索及社会关系超网络建模。由此形成从高阶关系建模到深度模型参数生成、从算法适应到硬件与工程部署的完整研究谱系。
总计 87 篇相关文献
: Neural networks have a rich ability to learn complex representations and have achieved remarkable results in various tasks. However, they are prone to overfitting owing to the limited number of training samples and regularizing the learning process of neural networks is essential. In this paper, we propose a regularization method that estimates the parameters of a large convolutional neural network as probabilistic distributions using a hypernetwork, which generates the parameters of another network. Additionally, we perform model averaging to improve the network performance. Then, we apply the proposed method to a large model such as wide residual networks. The experimental results demonstrate that our method and its model averaging outperform the commonly used maximum a posteriori estimation with L2 regularization.
Fluid dynamics computations for tube-like geometries are crucial in biomedical evaluations of vascular and airways fluid dynamics. Physics-Informed Neural Networks (PINNs) have emerged as a promising alternative to traditional computational fluid dynamics (CFD) methods. However, vanilla PINNs often demand longer training times than conventional CFD methods for each specific flow scenario, limiting their widespread use. To address this, multi-case PINN approach has been proposed, where varied geometry cases are parameterized and pre-trained on the PINN. This allows for quick generation of flow results in unseen geometries. In this study, we compare three network architectures to optimize the multi-case PINN through experiments on a series of idealized 2D stenotic tube flows. The evaluated architectures include the ‘Mixed Network’, treating case parameters as additional dimensions in the vanilla PINN architecture; the “Hypernetwork”, incorporating case parameters into a side network that computes weights in the main PINN network; and the “Modes” network, where case parameters input into a side network contribute to the final output via an inner product, similar to DeepONet. Results confirm the viability of the multi-case parametric PINN approach, with the Modes network exhibiting superior performance in terms of accuracy, convergence efficiency, and computational speed. To further enhance the multi-case PINN, we explored two strategies. First, incorporating coordinate parameters relevant to tube geometry, such as distance to wall and centerline distance, as inputs to PINN, significantly enhanced accuracy and reduced computational burden. Second, the addition of extra loss terms, enforcing zero derivatives of existing physics constraints in the PINN (similar to gPINN), improved the performance of the Mixed Network and Hypernetwork, but not that of the Modes network. In conclusion, our work identified strategies crucial for future scaling up to 3D, wider geometry ranges, and additional flow conditions, ultimately aiming towards clinical utility.
Densely captured real-world materials require effective compression for rendering, material generation and reconstruction. Neural networks with high compression rates and the ability to fit complex functions can encode each BRDF into the corresponding network. However, current works that take advantage of single implicit neural representations are incapable of effectively modeling the high-frequency details of the highlight region. In this paper, we propose an improved compact neural network representation of BRDF data based on the sinusoidal activation. The lightweight network and the periodic activation function improve the fidelity of the reproduction material appearance under the condition of a high compression rate. Furthermore, the method of building a unified model using neural networks can decode all materials from latent space. However, the deep structure of the network model increases memory consumption. To overcome this challenge, we propose a hypernetwork framework that compresses measured BRDFs to latent space and generates weights for the neural network-based representation of materials. The lightweight implicit representation of BRDF generated by training directly from original materials shows the characteristics of a low memory footprint and high-precision reproduction of appearance. Additionally, we apply the hypernetwork to reconstruct materials from a single image. Thanks to implicit representation of BRDF that can reproduce the appearance with high fidelity, the reflectance properties can be accurately recovered.
No abstract available
Physics-informed neural networks (PINNs) have been successfully applied in electromagnetism (EM) for the solution of direct problems. However, since PINNs typically do not take system parameters (like geometry or material properties) as input, when embedded in inverse problems or adopted for parametrical studies, to output the solution of the governing equations, they require additional training for each new system parameter set. To overcome this issue, we propose a hypernetwork (HNN) that receives system parameters and outputs the network weights of a PINN, which in turn provides the solution of the direct problem. Therefore, once trained, the HNN acts as a parametrized real-time field solver that allows the fast solution of inverse problems, in which the objective(s) are defined a posteriori (i.e., after HNN’s training). This method is adopted for a coil optimal design task in magnetostatics.
No abstract available
Motion-Aware Neural Networks Improve Rigid Motion Correction of Accelerated Segmented Multislice MRI
Synopsis We demonstrate a deep learning approach for fast retrospective intraslice rigid motion correction in segmented multislice MRI. A hypernetwork uses auxiliary rigid motion parameter estimates to produce a reconstruction network based on the motion parameters that are specific to the input image. This strategy produces higher quality reconstructions than those produced by model-based techniques or by networks that do not use motion estimates. Further, this approach mitigates sensitivity to misestimation of the motion parameters.
Integrated data and energy transfer (IDET) is an important component of future communication networks due to its characteristic to provide continuous and reliable wireless power to battery-limited devices. In order to get full advantages of IDET technology in future communication networks and realize efficient beamforming design, the base station (BS) must obtain accurate downlink channel state information (CSI). In this work, we propose an energy harvesting (EH) aided channel estimation scheme using a novel deep learning (DL) architecture in an end-to-end training mode. This architecture simultaneously obtains CSI from both an energy harvester and a data decoder at an IDET receiver. By designing a hypernetwork-aided deep neural network (DNN), we achieve more accurate channel estimation, while effectively reducing the pilot overhead in channel estimation. Simulation results show that compared to the state of the art, our EH aided channel estimation scheme attains a lower normalized mean square error (NMSE).
No abstract available
In recent years, there has been a significant advancement in memristor-based neural networks, positioning them as a pivotal processing-in-memory deployment architecture for a wide array of deep learning applications. Within this realm of progress, the emerging parallel analog memristive platforms are prominent for their ability to generate multiple feature maps in a single processing cycle. However, a notable limitation is that they are specifically tailored for neural networks with fixed structures. As an orthogonal direction, recent research reveals that neural architecture should be specialized for tasks and deployment platforms. Building upon this, the neural architecture search (NAS) methods effectively explore promising architectures in a large design space. However, these NAS-based architectures are generally heterogeneous and diversified, making it challenging for deployment on current single-prototype, customized, parallel analog memristive hardware circuits. Therefore, investigating memristive analog deployment that overrides the full search space is a promising and challenging problem. Inspired by this, and beginning with the DARTS search space, we study the memristive hardware design of primitive operations and propose the memristive all-inclusive hypernetwork that covers 2×1025 network architectures. Our computational simulation results on 3 representative architectures (DARTS-V1, DARTS-V2, PDARTS) show that our memristive all-inclusive hypernetwork achieves promising results on the CIFAR10 dataset (89.2% of PDARTS with 8-bit quantization precision), and is compatible with all architectures in the DARTS full-space. The hardware performance simulation indicates that the memristive all-inclusive hypernetwork costs slightly more resource consumption (nearly the same in power, 22%∼25% increase in Latency, 1.5× in Area) relative to the individual deployment, which is reasonable and may reach a tolerable trade-off deployment scheme for industrial scenarios.
Recent advances in the field of image generation have attracted attention due to the growing number of diverse data sources and test samples. A primary driver of this evolution is the application of neural networks, particularly for generating high-quality images from textual prompts. Despite the potential of diffusion models in this sector, they typically face computational challenges associated with vast datasets. This paper describes the research on two existing solutions: Hypernetworks and Low-Rank Adaptation (LoRA), both aiming to streamline and optimize the image generation process. While hypernetworks dynamically adjust model parameters based on the input text, increasing flexibility and performance, LoRA efficiently adapts the primary model style without requiring it to be trained from scratch. Using the Stable Diffusion 1.5 model as a benchmark, this research evaluates the influence of hypernetwork and LoRA modifications. The results indicate that both approaches provide efficient and highly accurate image generation, confirming their efficacy in contemporary image generation tasks.
Full-duplex (FD) systems have emerged as a promising avenue to optimize temporal and spectral resource utilization by enabling simultaneous data transmission and reception on a single frequency. Nonetheless, the presence of robust self-interference (SI) signals at the receiver presents a critical challenge that necessitates effective SI cancellation strategies. Recent advancements have introduced neural networks (NN) to address this challenge, offering computational advantages over conventional polynomial models. This paper delves into the realm of leveraging Hyper Neural Networks (hyperNet) for SI cancellation (SIC), thereby exploring their unique attributes to enhance the efficiency of FD systems. The proposed method utilizes a dynamic hyperNet model to learn the non-linear characteristics of the SI channel and cancel it from the received signal. The efficacy of the proposed SIC technique is assessed using datasets derived from a MATLAB simulation platform that accurately replicates the transmission process observed in real-world scenarios. The simulation results demonstrate the superior performance of the proposed hyperNet in comparison to other state-of-art NN based SIC methods. Our proposed hyperNet-based SIC technique demonstrates its ability to autonomously adapt to more complex characteristics of the varying self-interference channel.
Image super-resolution models often require a large number of parameters to capture the complex mapping between low-resolution and high-resolution images. Hypernetworks allow for efficient parameterization by generating the weights of the target network dynamically based on the input low-resolution image. This enables the model to have a smaller set of fixed parameters while still being able to adapt and generate high-resolution images. Hypernetworks are meta-learning neural networks that generate the weights or parameters of another neural network, known as the target network, based on the input data. With only a 0.12% increase in computation parameters complexity for SRCNN as the target network, the proposed framework, Hyper-SR: a Hyper-network-based framework for single-image Super-Resolution, outperforms the target network in terms of perceptual image quality at higher scaling factors and faster convergence time at fewer epochs. We demonstrate results and ablation experiments using existing SRCNN as the target network and reported an average gain on SET5 dataset of +0.83 db for PSNR and +0.0208 for SSIM, on SET14 dataset, a gain of +0.62 db for PSNR and +0.0109 for SSIM at scaling factor of 4. However, our methodology may be used for any existing super-resolution network as the target network to obtain marginally improved resolution without necessitating a large number of computational parameters.
No abstract available
No abstract available
No abstract available
Unmanned aerial vehicles (UAVs) are recognized as effective data collectors for wireless sensor networks. The Age of Information (AoI), a metric indicating data freshness, is crucial for decision making in time-sensitive applications. It can be significantly reduced by jointly optimizing UAV trajectories and communication scheduling of sensor nodes (SNs). However, rapid changes in the environment make it challenging to predesign UAV trajectories and communication scheduling decisions using traditional methods, especially when central controllers are absent and the numbers of UAVs and SNs vary. In this article, we propose hypernetwork-based QMIX (HyperQMIX), a population-invariant multiagent deep reinforcement learning (MADRL) algorithm capable of transferring policies across tasks with varying population sizes. First, we design neural network modules adaptable to varying input and output dimensions, facilitated by parameter generation through a hypernetwork. Then, HyperQMIX leverages these modules to process fluctuations in state and action dimensions. This approach ensures that the network structure remains consistent regardless of population sizes, thereby enhancing the algorithm’s scalability. Extensive simulations demonstrate that HyperQMIX significantly outperforms state-of-the-art algorithms in terms of learning efficiency and converged performance. Moreover, agents pretrained with HyperQMIX perform well in tasks of different population sizes without additional training. Fine-tuning these models achieves performance comparable to training from scratch.
The evolution of Beyond 5G (B5G) and 6G networks requires the design of powerful deep neural networks to effectively address the emerging threats and ensure predictive, adaptive security measures that can safeguard the performance and integrity of these advanced systems. Existing studies rely on fully supervised learning settings, lack mechanisms for dynamic adaptation, and perform their experiments on datasets, which do not represent the characteristics of existing 5G scenarios. To address these limitations, we present the first study incorporating hypernetworks into convolutional autoencoders for identifying attacks in 5G datasets collected under realistic conditions. Specifically, the input data is reshaped into a 2D matrix and fed into a convolutional autoencoder, consisting of an encoder and a decoder. At the same time, raw input data is passed through a hypernetwork, which is trained to generate weights for the target network. The target network receives as input the latent representation vector and gives as output the final prediction. Experiments are performed on two real 5G network datasets, namely NANCY and 5G-NIDD. Results show that the proposed deep neural network achieves an Accuracy of 98.51% and 99.91% on the NANCY and 5G-NIDD datasets respectively. Findings of an ablation study demonstrate the effectiveness of the proposed method.
Resting-state functional connectivity hypernetworks, in which multiple nodes can be connected, are an effective technique for diagnosing brain disease and performing classification research. Conventional functional hypernetworks can characterize the complex interactions within the human brain in a static form. However, an increasing body of evidence demonstrates that even in a resting state, neural activity in the brain still exhibits transient and subtle dynamics. These dynamic changes are essential for understanding the basic characteristics underlying brain organization and may correlate significantly with the pathological mechanisms of brain diseases. Therefore, considering the dynamic changes of functional connections in the resting state, we proposed methodology to construct resting state high-order functional hyper-networks (rs-HOFHNs) for patients with depression and normal subjects. Meanwhile, we also introduce a novel property (the shortest path) to extract local features with traditional local properties (cluster coefficients). A subgraph feature-based method was introduced to characterize information relating to global topology. Two features, local features and subgraph features that showed significant differences after feature selection were subjected to multi-kernel learning for feature fusion and classification. Compared with conventional hyper network models, the high-order hyper network obtained the best classification performance, 92.18%, which indicated that better classification performance can be achieved if we needed to consider multivariate interactions and the time-varying characteristics of neural interaction simultaneously when constructing a network.
Hypernetworks—neural networks that generate the weights of a target network—offer a compelling paradigm for rapid task adaptation by producing task-specific parameters in a single forward pass, bypassing the need for iterative gradient-based fine-tuning. This paper presents a comprehensive examination of hypernetwork architectures for few-shot learning, spanning theoretical foundations, architectural design principles, and empirical performance across standard benchmarks. We trace the evolution from foundational hypernetwork formulations through task-conditioned and attention-based variants to modern approaches integrating hypernetworks with pretrained foundation models via prompt generation and low-rank adaptation. We provide detailed comparisons with optimization-based meta-learning (MAML), metric learning (Prototypical Networks), and amortized inference approaches. Our analysis covers both classification and regression settings, addressing challenges including weight space dimensionality, generalization bounds, and computational efficiency. We identify key open problems including scaling hypernetworks to generate weights for billion-parameter models and establishing tighter theoretical guarantees for hypernetwork-generated parameters.
Wireless localization and sensing are increasingly important capabilities when the networks are evolving towards the $6^{t h}$ generation era. While the physics-inspired geometrical models are known to perform well in line-of-sight (LoS) dominant scenarios, harnessing the power of artificial intelligence (AI) to improve robustness, efficiency, and performance in more complex propagation scenarios is an intriguing prospect. To this end, the hypernetwork (HN) is an emerging neural network (NN) architecture, where one model is used to parameterize the weights of the other, promising dynamic weight adaptation among other performance improvements. In this work, we propose the concept of Hypernetwork Localization (HypLoc) - a hybrid HN-based architecture for localization in beamforming millimeter-wave (mmWave) networks, while combining angle-of-arrival (AoA), time-of-flight (ToF), and received power (RP) as representative measurements. Considering a realistic urban vehicular environment, we first demonstrate the baseline effectiveness of HypLoc with a fixed and known gNodeB (gNB) deployment scenario. We then also study a scenario where the factory pre-training covers multiple different gNB deployment constellations and show that the proposed HypLoc clearly outperforms the traditional NNs. Finally, we also show that the HypLoc adapts faster and requires less training data when adapting to a previously unseen deployment scenario. Overall, the proposed approach facilitates efficient factory pre-training when operating under multiple different gNB deployment options.
Most existing text-driven face image generation and manipulation methods are based on StyleGAN2, which is inherently limited to aligned faces and therefore makes these methods fail to preserve the highly variable face placement. Additionally, these methods directly leverage a pairwise loss to learn the correspondence between the image and text, which can not handle complex text descriptions, e.g., the text with multiple captions describes multiple facial attributes. To address these issues, we explore the feasibility of applying the more advanced StyleGAN3 to generate and manipulate the face images in an Open-World setup, e.g., the target face image is not required to be aligned and the text description contains multiple captions. To this end, we first design an improved iterative refinement strategy that adaptively predicts the generator weight offsets rather than residuals for the inverted latent code via a hypernetwork, which efficiently finds a desired generator with no image-specific optimization. We further analyze the disentanglement of different StyleGAN3 latent spaces and demonstrate that the ${\mathcal {S}}$ space learns a more semantically-disentangled representation. To enable complex edits mentioned by the multi-caption text, we propose a cross-modal feature filtration module with a probability adaptation strategy to capture the image-text correspondences. Finally, we incorporate a channel-wise attention mechanism to obtain a global latent manipulation direction, which learns to assign importance weights to different channels. Extensive experiments demonstrate the superior performance of our proposed method compared against the state-of-the-art methods.
Federated learning, as an emerging paradigm in the field of distributed machine learning, can handle both "data silos" and privacy protection issues by jointly training global models with scattered data parties while ensuring that data from all parties remains locally at all times. However, it faces two significant challenges in Non-IID data scenarios: performance degradation due to client-side data heterogeneity and catastrophic forgetting caused by the continuous evolution of local tasks. To address these issues, this paper presents a new federated continuous learning framework that combines Hypernetwork Federated dynamic weight generation with Dual-Domain knowledge distillation (HFDD). Specifically, deploy the hypernetwork on the server side, generate personalized aggregated weights based on client features, and construct an initialization model that considers both global generalization and local characteristics. To address the problem of catastrophic forgetting, client-local training introduces a dual-domain knowledge distillation loss, providing information across and within domains without compromising leakage. Experimental verification shows higher training accuracy on the Fashion-MNIST and CIFAR-10 datasets compared to FedAvg and FCCL, effectively mitigating the effects of data heterogeneity and catastrophic forgetting.
Most of the current image semantic segmentation algorithms on UAV vision segment remote sensing images, which cannot represent ground detail information, resulting in obstacles to real-time autonomous environment perception of UAVs in low altitude flight missions. To address this problem, this paper proposes a real-time image semantic segmentation method for low altitude UAVs. The network designs a novel hypernetwork structure that incorporates a context header weight generation module at the last layer of the encoder, the weights of each block in the decoder are generated before the end of encoding in the encoder to reduce the number of parameters and computation of the model to achieve real-time segmentation. In the decoder, a dynamic patch-wise convolutional algorithm is designed using the locally connection layer mechanism to take full account of the contextual semantic information when targeting large segmented objects that are in more than one piece, so that the decoder's weights change with the spatial location of the input feature map, and at the same time, the dynamic weights are used to target the segmentation of different objects, to maximise the network's adaptive nature. In order to verify the effectiveness of the method, this experiment uses the transfer learning technique to carry out pre-training on Cityscape data set, and uses the UAVid data set to validate the method of this paper, and the experimental results show that the mean intersection over union of this method is 66.3% for the images of the categories of buildings, roads, and static cars, and the prediction speed reaches 37.9 FPS, which significantly improves segmentation accuracy under the condition of guaranteeing the real-time performance.
Recent advancements in Large Language Models (LLMs) have focused on scaling architectures and developing efficient fine-tuning strategies like Low-Rank Adaptation (LoRA). While traditionally applied to autoregressive models, these methods are increasingly being adapted for text-based diffusion models. Standard fine-tuning, however, applies a static set of weight modifications across all samples and generation steps. In this work, we propose a dynamic, per-step adaptation method. We introduce a hypernetwork, termed the feedbackward (FB) model, which generates unique LoRA weights for a base feedforward (FF) model at each step of the diffusion process. Inspired by feedback connections in the brain, the FB model processes a masked input to generate LoRA weights for the query, key, value, and output projections within each of the FF model's attention blocks. Notably, these weights are generated in reverse sequence, from the final block to the first, allowing high-level contextual adjustments to inform lower-level feature representations. This dynamic weight generation is hypothesized to better guide the sampling trajectory for each input, offering more nuanced control than a static set of weights. The FF model, augmented with these bespoke LoRA weights, then processes the same input to produce the output for that generation step.
No abstract available
Collecting data from diverse perspectives is essential across many fields to achieve high-resolution imaging, from Synthetic Aperture Imaging (SAI) to advanced microscopy techniques. However, due to persistent challenges, fully capturing variability in position, angle, and scale remains difficult, driving the need for advanced data generation techniques. In this paper, we utilise a combination of Implicit Neural Representations and HyperNetwork framework to generate datasets enriched with spatial variability, including variations in position, angle, and scale. This method enhances the diversity of generated data, enabling more flexible and robust image datasets applicable to various imaging tasks. Our approach is evaluated on the MNIST and cryo-EM datasets, where we demonstrate the generation of spatially diverse images that maintain high-quality attributes despite changes in spatial configuration. This work highlights the potential to improve data generation processes, leading to more versatile and comprehensive datasets for scientific use.
Training 3D generative models often faces bottle-necks due to dependencies on precise camera pose estimation, particularly when using Neural Radiance Fields (N eRFs) for photorealistic novel-view synthesis. We introduce HyperNeR-FGAN, a generative framework that eliminates camera pose requirements by integrating a hypernetwork with a Generative Adversarial Network (GAN). This architecture maps Gaussian noise directly to the weights of a NeRF model, bypassing viewing direction inputs during training. Our experiments demonstrate that HyperNeRFGAN achieves state-of-the-art performance on datasets where camera position estimation is impractical - notably in medical imaging scenarios with limited or ambiguous viewpoint metadata. Despite its architectural simplicity compared to existing methods, the model produces high-fidelity 3D reconstructions across diverse modalities, including MRI and X-ray-derived 2D scans. The framework's efficiency and robustness suggest broad applicability in domains requiring 3D generation from unstructured or poorly annotated 2D data. Key revisions emphasize the camera-pose independence, clinical relevance, and architectural efficiency while maintaining technical nrecision.
Human facial reactions play crucial roles in dyadic human-human interactions, where individuals (i.e., listeners) with varying cognitive process styles may display different but appropriate facial reactions in response to an identical behaviour expressed by their conversational partners. While several existing facial reaction generation approaches are capable of generating multiple appropriate facial reactions (AFRs) in response to each given human behaviour, they fail to take human's personalised cognitive process in AFRs generation. In this paper, we propose the first online personalised multiple appropriate facial reaction generation (MAFRG) approach which learns a unique personalised cognitive style from the target human listener's previous facial behaviours and represents it as a set of network weight shifts. These personalised weight shifts are then applied to edit the weights of a pre-trained generic MAFRG model, allowing the obtained personalised model to naturally mimic the target human listener's cognitive process in its reasoning for multiple AFRs generations. Experimental results show that our approach not only largely outperformed all existing approaches in generating more appropriate and diverse generic AFRs, but also serves as the first reliable personalised MAFRG solution. Our code is made available at https://github.com/xk0720/PerFRDiff.
Abstract Hunting for candidate compounds with favorable pharmacological, toxicological, and pharmacokinetic properties in drug discovery is essentially a low-data problem, as data acquisition is both challenging and costly. This inherent data limitation clashes with the requirements of many powerful deep learning models, which typically require large datasets. Here, we present Meta-Mol, a novel few-shot learning framework based on Bayesian Model-Agnostic Meta-Learning. Meta-Mol introduces a novel atom-bond graph isomorphism encoder that captures molecular structure information at the atomic and bond levels. This representation is further enhanced by a Bayesian meta-learning strategy, allowing for task-specific parameter adaptation and reducing overfitting risks. Additionally, a hypernetwork is employed to dynamically adjust weight updates across tasks, facilitating more complex posterior estimation. Our results demonstrate that Meta-Mol significantly outperforms existing models on several benchmarks, providing a robust solution to address data scarcity in drug discovery.
Salient object ranking (SOR) aims to segment salient objects in an image and simultaneously predict their saliency rankings, according to the shifted human attention over different objects. The existing SOR approaches mainly focus on object-based attention, e.g., the semantic and appearance of object. However, we find that the scene context plays a vital role in SOR, in which the saliency ranking of the same object varies a lot at different scenes. In this paper, we thus make the first attempt towards explicitly learning scene context for SOR. Specifically, we establish a large-scale SOR dataset of 24,373 images with rich context annotations, i.e., scene graphs, segmentation, and saliency rankings. Inspired by the data analysis on our dataset, we propose a novel graph hypernetwork, named HyperSOR, for context-aware SOR. In HyperSOR, an initial graph module is developed to segment objects and construct an initial graph by considering both geometry and semantic information. Then, a scene graph generation module with multi-path graph attention mechanism is designed to learn semantic relationships among objects based on the initial graph. Finally, a saliency ranking prediction module dynamically adopts the learned scene context through a novel graph hypernetwork, for inferring the saliency rankings. Experimental results show that our HyperSOR can significantly improve the performance of SOR.
No abstract available
Financial time series forecasting remains a challenging task due to the high non-stationarity and concept drift inherent to market data. Existing deep learning models, such as LSTMs and transformers, typically employ static weights after training, limiting their ability to adapt to rapid market regime shifts (e.g., from trends to reversals). To bridge this gap between static parameters and dynamic environments, we propose a novel framework named Game-Theoretic HyperNetwork (GTH-Net), which introduces a context-aware meta-learning mechanism to achieve adaptive forecasting. Specifically, we first introduce an Evolutionary Game-Theoretic Correction Module (E-GTCM) to explicitly extract latent buying and selling pressure based on market microstructure priors through an iterative gated evolution process. Subsequently, we propose a HyperNetwork-based fusion mechanism that treats the extracted game state as a meta-context to dynamically generate the weights of the forecasting head. This allows the model to automatically switch its prediction rules in response to shifting market regimes. Extensive experiments on real-world stock datasets demonstrate that GTH-Net significantly outperforms baselines in terms of machine learning predictive accuracy and simulated financial profitability. Furthermore, ablation studies and parameter analysis confirm that the dynamic weight generation mechanism effectively captures market reversals caused by overcrowded trades.
Federated learning (FL) enables privacy-preserving collaboration among distributed clients, but practical deployments often face heterogeneous models and non-IID data, leading to degraded communication and personalization. In addition, real-world FL systems frequently encounter newly joined clients that require rapid adaptation and abnormal clients that may upload corrupted updates, further exacerbating instability and hindering global convergence. To address these challenges in image classification, we propose HFedDGHN, a Heterogeneous Federated Dynamic Graph HyperNetwork that jointly models inter-client relations and personalized parameter generation. Specifically, a graph structure learner adaptively captures client correlations to construct a dynamic collaboration graph, while a graph-convolutional hypernetwork generates model parameters for heterogeneous architectures, enabling implicit knowledge transfer without sharing local data or weights. Moreover, the framework naturally supports meta-learning-based generalization, allowing efficient adaptation to newly joined clients. Furthermore, the dynamic graph enhances robustness by isolating abnormal clients, as they tend to be excluded from most neighborhoods during adaptive graph construction. Extensive experiments across multiple benchmarks demonstrate that HFedDGHN achieves superior accuracy compared to state-of-the-art personalized and heterogeneous FL methods, while naturally improving robustness and scalability in real-world deployments.
Abstract Recently, the information diffusion models based on static hypernetworks have been proposed, in which the nodes and hyperedges represent the individuals and the social groups consisting of several individuals, respectively. However, the social networks represented by hypernetwork should be a dynamic network because of social activities. Hence, we establish an SIS model to present the dynamics of information spreading in a dynamic social hypernetworks with three dynamic processes. Then the theoretical analysis and simulation of the model are carried out, and theoretical results are accordant with the simulation results. In addition, the effects of different parameters on the information dissemination dynamics are analyzed in RP and CP strategy respectively. Most noteworthy, reorganization of social relations leads to a decrease in the proportion of informed individuals in stable state. By analyzing the effect of hyperedge reorganization on hyperdegree distribution, we discover the reorganization result in the increasing number of isolated nodes. These isolated nodes is hard to be joined in new social relationships, and cannot get information from the social network.
Channel decoding algorithms based on model-driven deep learning, also known as channel neural decoding algorithms, have received a lot of attention in recent years. However, the internal parameters and number of layers of the current channel neural decoding algorithm cannot be changed after training. Once changed, retraining of the channel neural decoding network is required. Hypernetwork is a neural network that can generate internal parameters for the main neural network to reduce the training cost of the main neural network and improve the flexibility of the main neural network. In this study, a novel hypernetwork based channel neural decoder is proposed for neural belief propagation algorithms (NBP), including the neural normalized min-sum (NNMS) and neural offset min-sum (NOMS) algorithms. According to the type of information interaction between the hypernetwork and the main decoding network, hypernetwork-based channel neural decoders can be divided into two types: static and dynamic. The internal parameters of the static hypernetwork-based channel neural decoder can be updated as needed without retraining of the main network. In addition to this benefit, the number of layers of the dynamic hypernetwork-based channel neural decoder can also be adjusted. Experimental results show that, compared with the existing NNMS decoding algorithms, the proposed hypernetwork-based NNMS decoding algorithms can achieve better performance on both low-density parity-check (LDPC) and Bose-Chaudhuri-Hocquenghem (BCH) codes.
Deploying transformer models in Personalized Federated Learning (PFL) at the wireless edge faces critical challenges, including high communication overhead, latency, and energy consumption. Existing compression methods, such as pruning and sparsification, typically degrade performance due to the sensitivity of self-attention layers (SALs) to parameter reduction. Also, standard federated averaging (FedAvg) often diminishes personalization by blending crucial client-specific parameters. To overcome these issues, we propose PFL-TPP (Personalized Federated Learning with Transformer Pruning and Personalization). This dual-strategy framework effectively reduces computational and communication burdens while maintaining high model accuracy and personalization. Our approach employs dynamic, learnable threshold pruning on feed-forward layers (FFLs) to eliminate redundant computations. For SALs, we introduce a novel server-side hypernetwork that generates personalized attention parameters from client-specific embeddings, significantly cutting communication overhead without sacrificing personalization. Extensive experiments demonstrate that PFL-TPP achieves up to 82.73% energy savings, 86% reduction in training time, and improved model accuracy compared to standard baselines. These results demonstrate the effectiveness of our proposed approach in enabling scalable, communication-efficient deployment of transformers in real-world PFL scenarios.
Despite technological progress in underwater object detection, there are still problems such as the domain shift stemming from diverse environmental conditions and severe image degradation in underwater environments. To address these challenges, we propose a hypernetwork-powered domain generalization framework that synergizes physical priors with multi-frequency feature learning, called Hy-UOD. To combat domain shift challenges, we devise a meta-learning empowered hypernetwork architecture that synthesizes domain-generalization parameters through environment-specific physical descriptor encoding for cross-domains. To further mitigate the impact of complex degradation on object detection performance, we designed a Multi-frequency Feature Dynamic Adaptation (i.e., MFDA) module based on hypernetwork features and domain-specific information. This module implements a systematic compensation for degraded features through a multi-level dynamic adaptation mechanism: ''low-frequency correction, high-frequency refinement, and mid-frequency reconstruction''. Experiments on multiple underwater datasets demonstrate the robust detection performance and strong cross-domain generalization capability of our method. The source code will be available at https://github.com/White-cat-ed/HyUOD.
This paper proposes an ex post path choice estimation framework for urban rail systems using an aggregated time-space hypernetwork approach. We aim to infer the actual passenger flow distribution in an urban rail system for any historical day using the observed automated fare collection (AFC) data. By incorporating a schedule-based dynamic transit network loading (SDTNL) model, the framework captures the crowding correlation among stations and the interaction between the path choice and passenger left behind, which is important for the path choice estimation in a “near-capacity” operated urban rail system. The path choice estimation is formulated as an optimization problem, which aims to minimize the difference between the model-derived and observed information with path choice parameters as decision variables. The original problem is intractable because of nonlinear (logit model) and nonanalytical (SDTNL) constraints. A solution procedure is proposed to decompose the original problem into three tractable subproblems, which can be solved efficiently. Solving the decomposed problem is equivalent to finding a fixed point. We prove that the solution to the original problem is the same as the decomposed problem (i.e., the fixed point) when passenger path choices follow the predefined behavior model. If this condition does not hold, the solution of the original problem is proved to be an “almost fixed point” for the decomposed problem. The model is validated using both synthetic and real-world AFC data from a major urban railway system. The analysis with synthetic data validates the model’s effectiveness in estimating path choice parameters and left behind probabilities, which outperforms state-of-art simulation-based optimization methods and probabilistic models in both accuracy and efficiency. The analysis using actual data shows that the estimated path shares are more reasonable than the baseline uniform path shares and survey-derived path shares. The model estimation is robust to different initial parameter values and AFC data from various dates.
Aiming at the problems of insufficient local geometric representation, low efficiency of multi-scale feature fusion and noise sensitivity of point cloud registration in complex scenes, this paper proposes an adaptive deep learning framework based on iterative weighted SVD with dynamic multi-scale attention mechanism. A dynamic convolutional kernel with adaptive inputs is generated through a hypernetwork, and combined with multi-scale channel attention to achieve hierarchical fusion of local-global features, which enhances the ability to characterize complex geometric structures; a coordinate-feature joint attention model is designed to filter high-confidence neighborhoods using a hybrid distance metric to improve the consistency of the local structure; a multitasking supervisory strategy is designed to filter the reliable corresponding pairs of points dynamically and optimize the transformation parameters to reduce the noise interference through an iteratively weighted SVD. Experiments on the ModelNet40 dataset show that, when a random noise test set (SNR = 10 dB) is added, the proposed method’s registration accuracy (RMSE(R) = 0.078, RMSE(t) = 0.0018) is significantly better than the current best method. Furthermore, a group of experiments prove that the method can achieve complete and accurate 3D point cloud registration for high dynamic range surfaces with any complex structure compared to other methods. Ablation experiments validate the effectiveness of the module.
Intelligent reflecting surface (IRS) can enhance the wireless communication environment by smartly reflecting the incident signal toward desired directions. However, the acquisition of channel state information (CSI) is challenging since IRS usually consists of a massive number of passive elements that have no capabilities of sensing and processing the pilot signals. Although by exploiting the sparsity of the angular domain channel, the huge pilot overhead can be reduced with conventional compressive sensing algorithms, such as approximate message passing (AMP), these algorithms cannot achieve satisfactory channel estimation performance. Based on the learned AMP (LAMP) network, we propose a hypernetwork-assisted LAMP (HN-LAMP) network with dynamic shrinkage parameters to improve the channel estimation accuracy. Furthermore, a recurrent architecture is adopted to reduce the large memory overhead arising from the LAMP network. Simulation results show that the proposed HN-LAMP network can improve the channel estimation accuracy or reduce the computational complexity under satisfactory estimation performance. Moreover, the proposed hypernetwork-assisted recurrent LAMP (HNR-LAMP) architecture can effectively reduce 50% memory overhead by sharing learnable weights.
Scientific collaboration has a complex hypernetwork structure. How to construct scientific collaboration in a complex system is an open issue. In this paper, a non-uniform dynamic collaborative evolution model is proposed. In the proposed method, each scholar is viewed as a node, and each cooperation relationship is regarded as a hyperedge. This model includes three processes: adding hyperedges, entering nodes, and forming hyperedges by new nodes. It is theoretically proved that the hyperdegree distribution of nodes follows the power law distribution. Furthermore, the effects of different parameters on the proposed model are numerically simulated in this paper. The experimental results are consistent with the theoretical ones. In addition, experiments show that the influence of nodes and hyperedges will affect the selection of old nodes when new nodes enter the network. This paper not only considers the construction of hyperedges with old nodes but also considers the possibility that new nodes construct new hyperedges among themselves. This model provides a reference for the research of the evolution process of scientific collaboration hypernetworks.
Radio Frequency Machine Learning Systems (RFMLS) have attracted increasing interest over the past few years. However, it has been demonstrated that RFMLS are vulnerable to Adversarial Machine Learning (AML). While AML has been extensively investigated in traditional domains, current state of the art often compromises the performance on benign data or introduces excessive computational overhead. As such, it cannot meet the strict requirements of tactical RFMLS. In this paper, we propose a novel defense approach based on dynamic adaptation of Deep Neural Network (DNN). Specifically, we leverage a hypernetwork to dynamically generate diverse parameters for a target DNN during inference. In addition, an ensemble learning and multi-stage training framework is proposed to train such a hypernetwork. Experimental results show that the proposed defense can increase the accuracy on adversarial examples by 48% and 16% in comparison to naturally trained DNN and defensive training strategies, respectively.
Industrial Operations Analytics (IOA) services increasingly rely on federated learning to enable privacy-preserving collaborative model training across distributed participants. However, real-world service deployments face a critical challenge: service participants join dynamically throughout the service lifecycle, introducing distribution shifts that compromise service stability. Existing federated learning frameworks assume static participant membership, leaving this dynamic service evolution problem largely unaddressed. We propose FedJoin, a serviceaware federated learning framework that maintains service quality during dynamic participant joining. FedJoin employs a dualembedding conditioned hypernetwork that jointly reasons about service status (encoded in the current model) and participant data distributions, enabling effective adaptation to dynamically joining participants. For resource-constrained deployments, we introduce FedJoin-Lite, a lightweight variant using low-rank factorization that reduces parameters by 95.89% while retaining $\mathbf{9 9. 1 0 {\%}}$ performance.
In this paper, we propose a dynamic 3D object detector named HyperDet3D, which is adaptively adjusted based on the hyper scene-level knowledge on the fly. Existing methods strive for object-level representations of local elements and their relations without scene-level priors, which suffer from ambiguity between similarly-structured objects only based on the understanding of individual points and object candidates. Instead, we design scene-conditioned hypernetworks to simultaneously learn scene-agnostic embeddings to exploit sharable abstracts from various 3D scenes, and scene-specific knowledge which adapts the 3D detector to the given scene at test time. As a result, the lower-level ambiguity in object representations can be addressed by hierarchical context in scene priors. However, since the upstream hypernetwork in HyperDet3D takes raw scenes as input which contain noises and redundancy, it leads to sub-optimal parameters produced for the 3D detector simply under the constraint of downstream detection losses. Based on the fact that the downstream 3D detection task can be factorized into object-level semantic classification and bounding box regression, we furtherly propose HyperFormer3D by correspondingly designing their scene-level prior tasks in upstream hypernetworks, namely Semantic Occurrence and Objectness Localization. To this end, we design a transformer-based hypernetwork that translates the task-oriented scene priors into parameters of the downstream detector, which refrains from noises and redundancy of the scenes. Extensive experimental results on the ScanNet, SUN RGB-D and MatterPort3D datasets demonstrate the effectiveness of the proposed methods.
The influence of the statistical properties of the network on the knowledge diffusion has been extensively studied. However, the structure evolution and the knowledge generation processes are always integrated simultaneously. By introducing the Cobb-Douglas production function and treating the knowledge growth as a cooperative production of knowledge, in this paper, we present two knowledge-generation dynamic evolving models based on different evolving mechanisms. The first model, named “HDPH model,” adopts the hyperedge growth and the hyperdegree preferential attachment mechanisms. The second model, named “KSPH model,” adopts the hyperedge growth and the knowledge stock preferential attachment mechanisms. We investigate the effect of the parameters on the total knowledge stock of the two models. The hyperdegree distribution of the HDPH model can be theoretically analyzed by the mean-field theory. The analytic result indicates that the hyperdegree distribution of the HDPH model obeys the power-law distribution and the exponent is . Furthermore, we present the distributions of the knowledge stock for different parameters . The findings indicate that our proposed models could be helpful for deeply understanding the scientific research cooperation.
With the global warming, soil erosion and a series of environmental problems are worsening. Green sustainable development has increasingly become a major issue of human concern. This study established a dynamic interaction model to explore the influence of green information and the social relations. Due to the continuous evolution of the network and the heterogeneity of individuals, the transition probability of each node is combined with time-varying parameters that reflect its current state. Then, guided by hypernetwork theory and microscopic Markov chain approach, this model is analyzed. Furthermore, the effects of various adjustment parameters on the green behaviors are compared and tested. The results show that behavior diffusion is not only affected by the diffusion of green information, but also closely related to the change of social relations. Finally, combined with the diffusion of green information, a propagation mechanism that simulates damped harmonic motion is proposed to maximize the green behavior. The results show that the final practice fraction of green behavior has been significantly improved.
Rotating machinery operating under varying speeds and loads exhibits strong nonstationarity, causing distribution shifts between training and deployment conditions that hinder reliable cross-condition fault diagnosis. Most domain generalization (DG) methods attempt to address this issue by enforcing domain-invariant representations; however, in rotating machinery, operating-condition variations are often physically coupled with fault-related features, making strict invariance assumptions less suitable—particularly when source domains are limited or discrepancies are large. To overcome these limitations, we propose a domain-conditioned dynamic DG framework for fault diagnosis under unseen operating conditions. The method introduces an explicit domain embedding branch and a lightweight hypernetwork with feature-wise linear modulation to generate channel-wise modulation parameters. By conditioning task features on operating information, the model adaptively adjusts its decision behavior across domains without adversarial training or explicit distribution alignment. Experiments on four public datasets (CWRU bearings, HUST bearings, HUST gearboxes, and PHM2009 gearboxes) demonstrate that the proposed framework achieves competitive or superior multi-source generalization performance and exhibits consistently improved training stability across diverse tasks and hyperparameter settings.
Multimodal sentiment analysis is often challenged by inter-modality conflicts and a lack of external context, which are difficult to resolve using conventional fusion strategies that typically rely on static mechanisms with limited sample-level flexibility. To address these issues, this paper introduces HERA (Hypernetwork-Enhanced Reasoning and Adaptation), a framework that integrates deliberative reasoning with a dynamic fusion module. HERA first grounds its analysis by retrieving external knowledge through a hybrid retriever, followed by an Iterative Cross-modal Deliberative Prompting process where a Large Language Model synthesizes raw inputs and context through a cycle of critique and refinement. To achieve precise integration, an Adaptive Hypernetwork Fusion module dynamically generates sample-specific parameters for a fusion operator. By conditioning on multimodal features and label embeddings, the hypernetwork synthesizes low-rank matrices that define a unique fusion tensor for each instance, allowing for a highly context-sensitive integration of text, augmented text, and visual cues. Extensive experiments on multiple benchmark datasets demonstrate that HERA significantly outperforms baseline methods, proving its efficacy in handling complex multimodal sentiment dynamics.
In recent years, machine learning (ML) tasks have been widely deployed at the edge of wireless networks, e.g., autonomous cars and tactile robots. However, the impairments of wireless channels between devices, such as fading and noise, deteriorate the effectiveness of ML inference tasks. In this work, we propose an efficient framework with an adaptive power control mechanism, which considers the constraint of the limited energy budget of edge devices. To guarantee the inference performance of ML tasks that are transmitted through wireless channels, we design hypernetworks with meta-parameters. The hypernetwork takes the context, such as the network condition, as the input and outputs the parameters of the power control network and artificial intelligence (AI) model. The training loss is designed by minimizing the trade-off between inference performance and energy consumption. Simulation results verify the effectiveness of the proposed adaptive inference framework on energy saving while ensuring the accuracy of inferring ML tasks.
Personalized Federated Learning (pFL) tailors models to each client’s local data distribution in heterogeneous federated learning settings. Federated Meta-Learning (FML) is a branch of pFL that uses meta-learning to achieve fast adaptation, where clients start with a meta-model and personalize it by fine-tuning it with local data. Since a single global meta-model has limitations when the data distribution of clients varies significantly, meta-model personalization should be considered in FML. However, most benchmark pFL methods lack meta-model personalization, and usually lack meta-learning or relying on a single global meta-model. Besides, these methods can neither provide meta-model personalization nor guarantee generalization and convergence, due to the challenges in measuring the distance between the meta-model and the client model in FML. To address these issues, we combine FML with hypernetwork and propose a constrained hypernetwork-based FML framework called FMLH, which innovatively utilizes hypernetwork to capture the differences in fine-tuned models, thereby providing personalized meta-models for each client. We provide rigorous mathematical proofs illustrating how the hypernetwork affects the convergence and generalization bounds of FMLH. Experimental results demonstrate that FMLH significantly improves the generalization of the model in cross-client shifts, with the lowest decile accuracy improved by up to 18.71%. FMLH also outperforms representative pFL algorithms by up to 5.6% in terms of maximum accuracy improvement.
Dynamic order scheduling in intelligent warehouse systems faces critical challenges within the domain of production research, including heterogeneous robot allocation, load balancing, and cross-scenario generalisation. This study proposes a meta-learning-based warehouse order scheduling framework (Meta-Scheduler) that adapts optimisation paradigms from cloud computing task scheduling to achieve efficient, adaptive scheduling in dynamic environments. The framework comprises three innovative components: (1) a hierarchical meta-feature extractor that encodes warehouse topology and order patterns via heterogeneous graph neural networks, (2) a dynamic hypernetwork policy generator enabling cross-scenario knowledge transfer through a differentiable neural architecture search, and (3) a robust optimal transport allocator that balances efficiency and load balancing via the entropy-regularized Sinkhorn algorithm. Experiments demonstrate that compared with traditional genetic algorithms and deep reinforcement learning, Meta-Scheduler decreases the order completion time and load Gini coefficient by 40.6% and 55.8%, respectively, in cold-start scenarios while achieving 81% faster convergence in multi-warehouse transfer tasks. Furthermore, the framework maintains real-time subsecond decision capability at the 200-robot scale, meeting industrial deployment requirements. This research establishes a scalable theoretical framework for dynamic warehouse scheduling and provides new insights for the production research community into the migration of the cross-domain task scheduling algorithm.
Domain adaptation from labeled source domains to the target domain is important in practical summarization scenarios. However, the key challenge is domain knowledge dis-entanglement. In this work, we explore how to disentangle domain-invariant knowledge from source domains while learning specific knowledge of the target domain. Specifically, we propose a hypernetwork-assisted encoder-decoder architecture with parameter-efficient fine-tuning. It leverages a hypernetwork instruction learning module to generate domain-specific parameters from the encoded inputs accompanied by task-related instruction. Further, to better disentangle and transfer knowledge from source domains to the target domain, we introduce a meta-knowledge distillation strategy to build a meta-teacher model that captures domain-invariant knowledge across multiple domains and use it to transfer knowledge to students. Experiments on three dialogue summarization datasets show the effectiveness of the proposed model. Human evaluations also show the superiority of our model with regard to the summary generation quality.
Multi-site neuroimaging studies are essential for developing clinically deployable artificial intelligence (AI) systems, yet deep learning models remain highly sensitive to scanner-induced domain shift. Variations in scanner manufacturer, field strength, acquisition protocol, reconstruction pipeline, and site-specific preprocessing can alter image intensity, texture, and spatial-frequency characteristics, thus degrading model performance when deployed on unseen scanners. This paper presents a spectral-conditioned hypernetwork framework for meta-learning neuroimaging normalization with concept-guided transparency. The proposed framework integrates a spectral-conditioned hypernetwork into the batch normalization layers of a three-dimensional convolutional encoder. For each input magnetic resonance imaging volume, a compact spectral embedding is extracted from the three-dimensional Fourier magnitude spectrum and used to dynamically generate the channel-wise normalization scale and shift parameters. This enables input-adaptive normalization without requiring explicit site labels or scanner metadata. To improve generalization to unseen imaging centers, this study formulates a bi-level meta-learning protocol in which small site-specific support sets adapt only the hypernetwork parameters, while the outer loop jointly optimizes the encoder and hypernetwork across multiple source sites. To enhance clinical interpretability, the proposed framework introduces a concept-guided regularization term that aligns gradient-weighted class activation mapping attribution maps with predefined anatomical regions of interest, including the hippocampus, ventricles, temporal pole, and entorhinal cortex. The framework is evaluated on a multi-site Alzheimer’s disease classification benchmark using leave-one-site-out validation. The proposed method improves diagnostic performance, reduces cross-site variability in the area under the receiver operating characteristic curve, and produces stronger alignment of attribution with clinically meaningful anatomical concepts compared to statistical harmonization, adversarial domain adaptation, meta-learning domain generalization, site-conditioned normalization, and concept-based interpretability baselines. These findings suggest that coupling adaptive normalization with few-shot meta-learning and anatomical concept guidance provides a practical pathway toward robust and transparent AI-assisted neuroimaging diagnosis.
Meta-learning enables rapid adaptation to new tasks, while continual learning addresses sequential task acquisition without catastrophic forgetting. However, existing approaches assume either fixed learning algorithms or slow gradient-based adaptation, failing to handle realistic scenarios where agents encounter novel task types in rapid succession. We propose SR-MCL (Self-Referential Meta-Learning for Continual Few-Shot Learning), a framework that learns to improve its own adaptation procedure through experience. SR-MCL employs a hypernetwork conditioned on compressed task history to generate meta-parameter updates, a probe network for interference prediction, and dynamic anchors with decay-aware freezing for stability. We provide theoretical analysis characterizing when improvement occurs: sublinear regret for known task types and amortized penalty for novel types. Experiments on Omniglot, MiniImageNet, and Meta-Dataset demonstrate 8.3% improvement in accuracy and 75% reduction in catastrophic forgetting compared to MAML, with bounded computational overhead. Our probe network achieves correlation $\rho=0.86$ between predicted and measured interference, validating the embedding-based proxy. Analysis confirms the hypernetwork learns task-conditioned update rules that diverge systematically from standard meta-gradients.
Multitask identification of linear parameter-varying (LPV) systems under process and measurement noise remains an open problem, as existing methods lack principled handling of scheduling-dependent stochastic state distributions, task-specific adaptation within a shared framework, and tractable optimization. We propose SMS-HCL, a data-driven framework that unifies scheduled meta-state space representation with hierarchical context learning. Under a bijective mapping assumption, scheduling-dependent state distributions are compressed into finite-dimensional meta-states, yielding deterministic transition dynamics and GMM-based output distributions. The task context is split into state-transition and observation sub-vectors, each feeding a dedicated hypernetwork, so that new tasks can be adapted by updating only the relevant context. Bilevel optimization is stabilized via an exponential moving average on shared parameters, avoiding second-order derivatives. Experiments on a four-task discrete-time LPV benchmark show that SMS-HCL outperforms CNN, BiLSTM, meta-learning, and a regularized LPV-ARX baseline on the held-out task; selective adaptation of the observation context alone attains comparable performance to full adaptation, supporting the hierarchical design.
Ensuring high indoor air quality (IAQ) while minimizing energy consumption and preserving occupant comfort is a central challenge in building management systems. Traditional rule-based or single-objective controls often neglect dynamic fluctuations in pollution levels and occupant behavior, leading to suboptimal trade-offs. In this paper, we propose a context-aware Meta-Reinforcement Learning (Meta-RL) framework that simultaneously addresses multiple objectives—IAQ, energy efficiency, and comfort—under a variety of building configurations and disturbances (e.g., wildfires, equipment faults, occupancy surges). Our approach integrates a Transformer-based encoder for latent context extraction, a Meta-Pareto hypernetwork that generates diverse policies for user-driven preferences, and safety-constrained adaptation to maintain strict pollutant thresholds. Through extensive simulations using EnergyPlus and real-world calibration data, the proposed framework demonstrates (1) significantly lower IAQ violations compared to standard RL and rule-based baselines, (2) reduced energy usage while maintaining comfortable thermal conditions, and (3) rapid transfer to new buildings via few-shot meta-training. These findings underscore the potential of Meta-RL to deliver robust, flexible HVAC control solutions in complex, real-world indoor environments.
No abstract available
Conventional electromagnetic cloaking paradigms predominantly necessitate the encasing of static objects within predefined topological enclosures, fundamentally restricting invisibility to fixed, closed geometries. Realizing dynamic, adaptive concealment for arbitrary moving targets within an open, boundary-free aperture remains a formidable challenge. Here, we report a meta-reinforcement-learning metasurface (Meta2Surface) that enables the first experimental demonstration of a "transparent cloaking tunnel" (TCT)—an open corridor permitting the undetected passage of diverse objects. Distinguished from traditional adaptive cloak, the Meta2Surface employs a sensor-in-the-loop meta-policy governed by a task-adaptive hypernetwork. This architecture fuses real-time sensing with historical interaction trajectories to instantly synthesize impedance strategies that actively nullify object-dependent scattering with millisecond-scale latency. Comprehensive full-wave simulations and microwave experiments confirm robust, high-fidelity cloaking of diverse dynamic targets—varying in shape, size, material, and trajectory—even under abrupt object substitution. By transitioning invisibility from static encapsulation to a dynamic, open architecture, this work establishes a new paradigm for fusing artificial intelligence with reconfigurable metasurfaces to achieve cognitive, large-scale electromagnetic wave control.
No abstract available
Variable speed limits (VSL) have been widely implemented to alleviate highway congestion and enhance operational efficiency. However, most existing studies focus on fixed traffic scenarios, making them inadequate when addressing uncertainties such as fluctuating traffic flows, extreme weather conditions, and construction-induced closures. Consequently, traditional VSL control strategies exhibit limited adaptability and generalization capability in unfamiliar scenarios. To overcome these limitations, this paper proposes a VSL control strategy based on Meta-Reinforcement Learning (Meta-RL) and Multi-Agent Proximal Policy Optimization (MAPPO) (Meta-MAPPO). This method leverages the meta-learning mechanism of Meta-RL and integrates a Hypernetwork module to dynamically adjust the network parameters of the control policy. By doing so, it adapts to diverse traffic scenarios and environmental disturbances, facilitating rapid policy transfer across scenarios and enhancing control performance. The training results demonstrate that Meta-MAPPO achieves faster convergence and superior model performance than MAPPO and Meta Multi-Agent Soft Actor-Critic (Meta-MASAC). Simulation experiments reveal that, compared with traditional feedback control methods and conventional multi-agent RL approaches, Meta-MAPPO exhibits significant advantages in unseen scenarios: it effectively mitigates traffic congestion and substantially reduces total travel time. The findings provide a more applicable solution for the practical implementation of VSL and offer valuable insights for further exploration of multi-agent methodologies in intelligent transportation systems.
The main goal of Few-Shot learning algorithms is to enable learning from small amounts of data. One of the most popular and elegant Few-Shot learning approaches is Model-Agnostic Meta-Learning (MAML). In this paper, we propose a novel framework for Bayesian MAML called BH-MAML, which employs Hypernetworks for weight updates. It learns the universal weights point-wise, but a probabilistic structure is added when adapted for specific tasks. In such a framework, we can use simple Gaussian distributions or more complicated posteriors induced by Continuous Normalizing Flows.
Deep learning (DL) has emerged as a powerful tool for predicting the remaining useful lifetime (RUL) of components and systems. However, there are two challenges. First, DL-based methods require a sufficient number of labelled samples, while the number of run-to-failure units in practical is often small due to the high cost and time-consuming life test. Second, it is desirable to provide the prediction uncertainty of RUL for maintenance decision-making. To tackle above issues, we propose a novel few-shot RUL prediction model with a hypernetwork structure incorporating uncertainty quantification and calibration (Hyper-UC). The proposed Hyper-UC uses a shared feature embedding network and a unit-specific weight to model the mapping from sensor signals to RUL, and a generative network is designed to learn the meta-knowledge of generating distributions over the unit-specific weights. To accurately provide the predictive uncertainty, the Hyper-UC systematically models two types of uncertainty: epistemic uncertainty and aleatoric uncertainty. In the online phase, the unit-specific weights of in-service units are obtained through calibration, and subsequently the predictive distribution of the RUL can be obtained. The proposed method is evaluated to have superior model performance than the benchmark methods in a case study using the C-MAPSS dataset.
Predicting drug synergy is critical to tailoring feasible drug combination treatment regimens for cancer patients. However, most of the existing computational methods only focus on data-rich cell lines, and hardly work on data-poor cell lines. To this end, here we proposed a novel few-shot drug synergy prediction method (called HyperSynergy) for data-poor cell lines by designing a prior-guided Hypernetwork architecture, in which the meta-generative network based on the task embedding of each cell line generates cell line dependent parameters for the drug synergy prediction network. In HyperSynergy model, we designed a deep Bayesian variational inference model to infer the prior distribution over the task embedding to quickly update the task embedding with a few labeled drug synergy samples, and presented a three-stage learning strategy to train HyperSynergy for quickly updating the prior distribution by a few labeled drug synergy samples of each data-poor cell line. Moreover, we proved theoretically that HyperSynergy aims to maximize the lower bound of log-likelihood of the marginal distribution over each data-poor cell line. The experimental results show that our HyperSynergy outperforms other state-of-the-art methods not only on data-poor cell lines with a few samples (e.g., 10, 5, 0), but also on data-rich cell lines.
Meta-learning has been proven to be effective for the cold-start problem of recommender systems. Many meta-learning recommender systems that are designed for the user cold-start problem are gradient-based. They use a global parameter learned from existing users to initialize the recommender system parameter for new users that provides a personalized recommendation with limited user-item interactions. Many systems require users’ demographic information to learn the global parameter. This requirement raises privacy concerns, and demographic information is not always available. In addition, some of the gradient-based systems need dozens or even hundreds of user-item interactions from existing users to learn the global parameter. This requirement is difficult to satisfy in specific scenarios in which active users and their item interactions are rare. Moreover, gradient-based meta-learning systems rarely capture user preferences over different item attributes, as updating the corresponding weights requires many optimization steps. We proposed HyperRS, a Hypernetwork-based Recommender System, for the user cold-start problem. Our system does not rely on demographic information to provide personalized recommendations. In our system, a hypernetwork generates all weights in the underlying recommender system. The hypernetwork enables the weights to adapt quickly to capture user interest in both item attributes and the contents of the item attributes. The experimental results show that our method outperforms several state-of-the-art meta-learning recommender systems for the user cold-start problem.
Continual learning allows systems to continuously learn and adapt to the tasks in an evolving real-world environment without forgetting previous tasks. Developing deep learning models that can continually learn over a sequence of tasks is challenging. We propose a novel method, AdaPrefix, which addresses this and empowers continual learning capability in pretrained large models (PLMs). AdaPrefix provide a continual learning method for transformer-based deep learning models by appropriately integrating the parameter-efficient methods, adapters and prefixes. AdaPrefix is an effective approach for smaller PLMs and achieves better results than state-of-the-art approaches. We further improve upon AdaPrefix by proposing AdaPrefix++, enabling knowledge transfer across the tasks. It leverages hypernetworks to generate prefixes and continually learns the hypernetwork parameters to facilitate knowledge transfer. AdaPrefix++ has a smaller parameter growth compared to AdaPrefix and is more effective and valuable for continual learning in PLMs. We performed several experiments on various benchmark datasets to demonstrate the performance of our approach for different PLMs and continual learning scenarios. Code is available on Github Link
No abstract available
No abstract available
Continual Learning aspires to build models capable of learning new tasks, without forgetting previously learnt tasks. In real-world settings, the distributions underlying the tasks are prone to shift. This necessitates a model capable of observing how the task distributions drift with time and adapt proactively. We present a novel framework of continual learning under evolving domains. Our approach employs a hypernetwork with separate embeddings conditioned on both domain and task to address this problem. The hypernetwork generates customised classifier weights corresponding to any domain-task pair. We employ a separate network that is trained end to end along with the hypernetwork to predict the next domain embedding, which in turn helps to generate classifier parameters corresponding to the next future domain in the evolution. We conduct extensive experiments on various datasets with a wide variety of distribution shifts to demonstrate the efficacy of our model in generalizing to future domains across all the tasks.
Deep neural networks suffer from catastrophic forgetting in continual learning, where they tend to lose information about previously learned tasks when optimizing a new incoming task. Recent strategies isolate the important parameters for previous tasks to retain old knowledge while learning the new task. However, using the fixed old knowledge might act as an obstacle to capturing novel representations. To overcome this limitation, we propose a framework that evolves the previously allocated parameters by absorbing the knowledge of the new task. The approach performs under two different networks. The base network learns knowledge of sequential tasks, and the sparsity-inducing hyper-network generates parameters for each time step for evolving old knowledge. The generated parameters transform old parameters of the base network to reflect the new knowledge. We design the hypernetwork to generate sparse parameters conditional to the task-specific information and the structural information of the base network. We evaluate the proposed approach on class-incremental and task-incremental learning scenarios for image classification and video action recognition tasks. Experimental results show that the proposed method consistently outperforms a large variety of continual learning approaches for those scenarios by evolving old knowledge.
No abstract available
Continual Relation Extraction (CRE) aims to learn constantly emerging tasks while avoiding forgetting the learned tasks. Previous studies suggest that interference among similar relations is the primary factor contributing to the forgetting of learned tasks. To address this issue, robust training strategies have been employed to enhance the model’s ability for distinguishing between similar relations. However, these methods use the same model parameters for learning all tasks, which increases the risk of conflicts between similar relations within the unified feature space. To address this issue, we propose a model that utilizes a Hypernetwork with Hybrid data augmentation for CRE (Hy2CRE). Specifically, Hy2CRE employs a hypernetwork-based network generator to generate task-specific projection heads for each task during the continual learning process, thereby mitigating the risk of conflicts emerging between similar relations within the model’s feature space. Meanwhile, we introduce a hybrid data augmentation method to further enhance the model’s robustness to similar relations, which integrates data augmentation in both the discrete text space and the continuous semantic space. Experimental results on two benchmark datasets prove the effectiveness of our model.
Blind Image Quality Assessment (BIQA) models trained on one distortion distribution often degrade when exposed to new ones, making sequential adaptation without forgetting a fundamental challenge. While continual learning offers a natural solution, existing methods typically retrain the entire backbone per task, limiting scalability and parameter efficiency. We propose ContEditIQA, a parameter-efficient framework for continual BIQA that selectively edits a pre-trained Vision Transformer (ViT) rather than retraining it. Following a locate-then-edit strategy, a lightweight attention-guided hypernetwork identifies distortion-sensitive Feed-Forward Network (FFN) parameters for each incoming task and restricts updates to those regions, while attention layers remain frozen to preserve globally shared representations. This targeted editing enables robust sequential adaptation without model expansion or memory replay. Experiments across six BIQA benchmarks demonstrate superior knowledge retention and cross-dataset generalization while modifying fewer than 30% of backbone parameters, establishing selective model editing as an effective and scalable paradigm for continual BIQA.
The ability to acquire new skills and knowledge continually is one of the defining qualities of the human brain, which is critically missing in most modern machine vision systems. In this work, we focus on knowledge transfer in the lifelong learning setting. We propose a lifelong learner that models the similarities between the optimal weight spaces of tasks and exploits this in order to enable knowledge transfer across tasks in a continual learning setting. To characterize the "task-parameter relationships", we propose a metric called adaptation rate integral (ARI), which measures the expected rate of adaptation over a finite number of steps for a (task, parameter) pair. These task-parameter relationships are learned using an auxiliary network trained on guided explorations of parameter space. The learned auxiliary network is then used to heuristically select the best parameter sets on seen tasks, which are consolidated using a hypernetwork. Given a new (unseen) task, knowledge transfer occurs through the selection of the most suitable parameter set from the hypernetwork that can be rapidly finetuned. We show that the proposed approach can improve knowledge transfer between tasks across standard benchmarks without any increase in overall model capacity, while naturally mitigating catastrophic forgetting.
Brain tumors are complex and life-threatening conditions that require accurate and efficient diagnostic approaches. However, existing approaches often face limitations in precision and computational efficiency, mainly due to the heterogeneous and limited nature of medical imaging datasets. Recent advancements in deep learning, mostly Neural Architecture Search (NAS) and Generative Adversarial Networks (GANs), have show significant potential for enhancing diagnostic performance. In this study, a novel framework integrating HyperNet-based Neural Architecture Search (HN-NAS) with Deep Convolutional Generative Adversarial Networks (DCGANs) is proposed for brain tumor detection and classification. The DCGAN model is employed to generate high-quality synthetic MRI images of brain lesions, thereby developing dataset diversity and mitigating the issue of limited training data. Meanwhile, HN-NAS is utilized to efficiently recognize optimal neural network architectures for accurate tumor diagnosis. The use of a HyperNetwork allows the generation of weights for multiple candidate architectures, provocatively decreasing the computational cost of architecture search and facilitating scalable model exploration. Experimental results establish that the proposed technique developments both segmentation and classification performance while maintaining computational efficiency. The findings designate that a reliable and scalable solution for real-time clinical applications can be accomplished by combining advanced NAS methods with generative models. Overall, this study establishes that integrating data augmentation with architecture optimization can suggestively improve medical imaging diagnosis.
Intelligent fault diagnosis is essential for downtime reduction in modern industries. However, domain (i.e., working condition) variants are unavoidable for fault diagnosis under changing environment. Though one can obtain a customized deep learning model for a given domain, it is not practical to design deep neural networks manually for multiple domains. Differentiable architecture search (DARTS), as an automatic machine learning technique, can automate the network design process for a specific domain efficiently by using hypernetwork and differentiable search strategy. Nevertheless, the representation of hypernetwork is restricted by several unfair factors, and the searched architecture of DARTS is overconfident. To address these issues, a one-shot neural architecture search approach, which involves two-stage learning, is proposed for efficient domain matching fault diagnosis. In the first training stage, the warmup and path-dropout strategies are taken to enhance the competitiveness of the parametric operators and alleviate the coadaptation problem to obtain an intuitively fair hypernetwork. In the second matching stage, variational inference is introduced into a differentiable search strategy to estimate the uncertainty of model matching, and a scale mixture prior is used to softly constrain the matching stage. Then, a candidate architectures set can be sampled and ordered from the posterior. Multidomain experiment is implemented by adding noise to the raw signal, and the proposed method outperforms four commonly used deep neural networks for aeroengine bevel gear fault diagnosis.
Accurate State-of-Charge (SOC) estimation of batteries remains challenging across aging states, temperature variations, and operating conditions. This paper introduces a novel Aging-Aware Hypernetwork Long Short-Term Memory with Multi-Constraint Physics-Informed Neural Network (AAH-LSTM-PINN) aimed at addressing some of the most pertinent challenges of estimating battery SOC through three key contributions. First, the introduction of a continuous four-component age parameter to quantify the temporal progression, the effects of capacity fade, resistance growth, and thermal stress. Second, the ability of a hypernetwork architecture to generate age-specific adaptation weights for explicit feature transformation. Third, the incorporation of physics-informed constraints based on Coulomb Counting (CC), Equivalent Circuit Model (ECM), and Arrhenius temperature with weights optimized using a three-phase systematic grid search. The effectiveness of the AAH-LSTM-PINN was evaluated on the Lithium-ion Battery (LIB) dataset under three extreme scenarios, which are temperature interpolation with a 0°C holdout, temperature extrapolation with a −20°C holdout, and drive cycle generalization. AAH-LSTM-PINN achieved 0.516% average Root Mean Square Error (RMSE) which gives a 42.8% improvement over age-agnostic LSTM. Age awareness, hypernetwork adaptation, and physics-informed constraints implementation contributed to 27.8%, 15.3%, and 6.5% improvement respectively. The AAH-LSTM-PINN framework has a real-time inference capability of $11.94~\mu $ s, which is appropriate for embedded Battery Management Systems (BMS) while also offering robust generalization to unseen situations. This work advanced the battery SOC estimation by showing that accurate state estimation across different aging states and operational conditions is possible when explicit age-aware mechanisms, coupled with deep learning and physics-informed regularization are employed.
Electroencephalography (EEG) is an important method of detecting brain activity. Various techniques are used to classify EEG for tasks such as motor imagery, emotion recognition, and medical diagnosis. However, many of these methods solely focus on the temporal information in EEG and ignore the spatial information provided by the positions of the electrodes. To address this gap, a combine technique of hypernetwork and recurrent neural network (RNN), so-called HyperRNN, is designed to utilize both spatial and temporal information in EEG and select better features subsets for EEG classification. In more details, hypernetwork is introduced to incorporates spatial information in EEG and search valid feature subsets. The RNN is employed to evaluate feature subsets, which consider temporal information with its memory architecture. Experiments on public EEG datasets demonstrate the proposed algorithm outperforms classical feature selection methods and improve EEG classification accuracy.
Computing power measurement faces the challenges of heterogeneity of devices, poor scalability with workloads, and discrepancies between measurement methods and models. In this article, we propose a computing power measurement framework called AdaptPerf to address these issues. AdaptPerf uses the technique of neural architecture search to build a hypernetwork of benchmark models for heterogeneous devices and adaptively selects the ones suitable for them. It achieves equivalency in computing power measurement through kernel-level latency analysis. We further prove that the adaptive benchmark measurement model selection problem (ABMSP) is an NP-hard problem and introduced a deep reinforcement learning approach to solve it. The theoretical analyzes show that the proposed method can guarantee that the obtained solution approximates the Pareto frontier. We implement AdaptPerf in a real-world environment, constructing a hypernetwork of over $4.2 \times 10^{20}$ models for heterogeneous devices to autonomously select the benchmark model adapted to the devices. We use different tools, such as Nsight Compute, Nsight System, to analyze and evaluate the feasibility and performance of AdaptPerf. It is found that the models generated by AdaptPerf increase computational throughput by 46.63% and improve data reuse rate by 93.7%, compared with more than 60 traditional models. AdaptPerf also exhibits superior load adaptability and model complexity compared with existing methods, such as MLPerf.
Computing power measurement is critical for effective resource utilization in computing power networks. However, it faces the challenges of heterogeneity of devices, poor scalability with workloads, and discrepancies between measurement methods and models. In this paper, we propose a computing power measurement framework called AdaptPerf to address these issues. AdaptPerf uses the technique of Neural Architecture Search to build a hypernetwork of benchmark models for heterogeneous devices and adaptively selects the ones suitable for them. It eliminates the impact of memory access factors among heterogeneous models by conducting kernel-level latency analysis, achieving equivalency in computing power measurement. We implement AdaptPerf in a real-world environment, constructing a hypernetwork of over 4.2 × 1020 models for heterogeneous devices to autonomously select the benchmark model adaptive to the devices. We use different tools, such as Nsight Compute, Nsight System, and etc., to analyze and evaluate the feasibility and performance of AdaptPerf. It is found that the models generated by AdaptPerf increase parameter volume by 5.19 times and improve by 4.99% in data reuse rate, compared with more than 50 conventional models. AdaptPerf also exhibits superior load adaptability and model complexity compared with existing methods such as MLPerf.
No abstract available
The growing demand for edge applications calls for efficient and optimized deep neural network models. Neural Architecture Search (NAS) is instrumental in designing such models, but achieving optimal architectures quickly remains a key challenge. To address this, we propose Hardware-aware Iterative One-shot NAS (HIO-NAS), a highly efficient approach that iteratively explores architectures across predefined search spaces for depth, filter, and width, all while respecting hardware constraints. HIO-NAS operates in four main steps: full training, random search, hardware verification, and retraining with adaptable knowledge distillation, repeated for each search space. The hardware-aware mechanism incorporates a lookup table to evaluate and filter subnetworks sampled during random search, ensuring only the most promising candidates proceed to hardware verification. The top-performing subnetwork(s) are then deployed on the target hardware for further validation. The adaptable knowledge distillation technique dynamically adjusts the teacher model’s influence based on deviations in the cost function during training. By progressively refining search spaces, HIO-NAS reduces computational overhead and avoids convergence issues. Its hypernetwork design emphasizes chain-like structures, where layers maintain connectivity. Applied to MobileNet v2 on CIFAR-100 and YOLOv7 on BDD100K, HIO-NAS delivered significant improvements: a 2.1% accuracy gain and 20.8% reduction in latency for MobileNet v2, and better performance over YOLOv7 and its variants. Interestingly, the findings highlight that unit-sharing topologies excel in depth searches, whereas unit-non-sharing topologies perform better in filter and width searches. Overall, HIO-NAS showcases a robust capability for efficiently discovering high-performance, hardware-optimized architectures, making it ideal for edge applications.
Information popularity prediction is a critical problem in social network analysis. With the increasing prevalence of social platforms, accurate prediction of the diffusion process has become increasingly important. Existing methods mainly rely on graph neural networks to model structural relationships, but they are often insufficient in capturing the complex interplay between temporal evolution and local cascade structures, especially in real-world scenarios involving sparse or rapidly changing cascades. To address this issue, we propose the Cascading Dynamic attention-calibrated Graph Convolutional Network, named CasDacGCN. It enhances prediction performance through spatiotemporal feature fusion and adaptive representation learning. The model integrates snapshot-level local encoding, global temporal modeling, cross-attention mechanisms, and a hypernetwork-based sample-wise calibration strategy, enabling flexible modeling of multi-scale diffusion patterns. Results from experiments demonstrate that the proposed model consistently surpasses existing approaches on two real-world datasets, validating its effectiveness in popularity prediction tasks.
In cooperative multi-agent reinforcement learning (MARL), the permutation problem—where the state space grows exponentially with the number of agents—reduces sample efficiency. Additionally, many existing architectures struggle with scalability, as they rely on a fixed structure tied to a specific number of agents, limiting their applicability to environments with variable entity counts. While approaches such as graph neural networks (GNNs) and self-attention mechanisms have made progress in addressing these challenges, they have significant limitations: dense GNNs and self-attention mechanisms incur high computational costs. To overcome these limitations, we propose an architecture composed of a novel agent network and a scalable non-linear mixing network, both of which ensure permutation-freeness and scalability, enabling generalization to environments with variable numbers of agents. Our agent network significantly reduces the computational complexity of attention from $O(n^{2}d)$ to $O(nd)$ , and our scalable hypernetwork enables efficient weight generation for non-linear mixing. Additionally, we introduce curriculum learning to improve training efficiency. Experiments on StarCraft Multi-Agent Challenge v2 (SMACv2) and Google Research Football (GRF) show that our proposed method generally outperforms representative MARL baselines, including HPN, UPDeT, VDN, and QMIX-based variants, particularly on SMACv2 scenarios with varying numbers and types of agents.
In real life, there are many cases that cannot be described by the network abstracted as the graph, but can be described perfectly by the hypernetwork abstracted as the hypergraph. Different from the network, the hypernetwork structure is more complex and poses a great challenge to the existing network representation learning methods. Therefore, in order to overcome the challenge of the hypernetwork structure, a hypernetwork representation learning method with the transformation strategy is proposed. Firstly, as three types of transformation strategies from the hypergraph to the graph, line graph, incidence graph and 2-section graph are combined into three types of integral graphs with the hyperedge information, namely incidence graph + 2-section graph, line graph + incidence graph and line graph + incidence graph + 2-section graph. Secondly, a shallow neural network algorithm is trained respectively on five types of networks abstracted as incidence graph, 2-section graph, incidence graph + 2-section graph, line graph + incidence graph and line graph +incidence graph + 2-section graph to obtain node representation vectors. Finally, the evaluation experiment is conducted on four different types of hypernetwork datasets. The experimental results demonstrate that the node classification performance of 2-section graph is better than that of other graphs, and the link prediction performance of incidence graph + 2-section graph is better than that of other graphs.
Federated Learning (FL) enables collaborative model training across distributed edge devices while preserving data privacy. However, severe communication overhead remains a primary bottleneck, especially in resource-constrained and heterogeneous environments. Existing quantization methods typically rely on uniform compression or rigid heuristics, which often overlook the varying learning behaviors across network layers and the dynamic bandwidth budget, inevitably sacrificing model accuracy for communication efficiency. To resolve these challenges, we propose FedQCab, a generic and plug-and-play layer-wise quantization tool designed to minimize communication overhead without compromising model performance. FedQCab introduces a lightweight hypernetwork to formulate an adaptive quantization policy. By leveraging a Differentiable Quantization Policy Generation mechanism via Gumbel-Softmax relaxation, we transform discrete bit-width selection into a differentiable training process to support hypernetwork training. Furthermore, FedQCab employs a bandwidth-constrained and task-aware optimization that directly minimizes the actual task loss while strictly adhering to real-time client bandwidth budgets. This feedback mechanism allows the system to intelligently allocate limited bit resources to sensitive layers while aggressively compressing robust layers. Extensive experiments across diverse datasets and architectures demonstrate that FedQCab significantly outperforms baselines, achieving up to 31.8% accuracy improvement and reducing communication overhead by up to 9.59× under dynamic bandwidth constraints.
Humans are well-known to be highly effective at comprehending continuous patterns within digital images. We present a collection of methods that enable analogous capabilities in deep neural networks. These methods train neural networks to represent images with continuous resolution-independent representations. They utilize an MCMC algorithm that directs attention during the learning phase to regions of the image that deviate from the current model. An encoding hypernetwork learns to generalize from a collection of images, such that it can effectively compute resolution-independent representations in constant time. These methods have immediate applications in super-resolution scaling of images, image compression, and secure image processing, and additionally suggest improved capabilities for image processing with neural networks in several future applications.
合并后的文献可归纳为十四个相互并列的方向。整体来看,hypernetwork的核心作用是依据任务、样本、域、客户端、环境状态、物理条件或网络关系动态生成目标模型的权重、结构或适配模块。研究一方面集中于元学习、持续学习、联邦个性化和参数高效适配,另一方面扩展至视觉生成与感知、医学影像、无线通信、物理工程、故障诊断、神经架构搜索及社会关系超网络建模。由此形成从高阶关系建模到深度模型参数生成、从算法适应到硬件与工程部署的完整研究谱系。