深度强化学习共享储能
多智能体深度强化学习在能源系统协调与共享中的应用
该组文献共同关注多智能体系统(MAS)在能源管理、P2P能源交易及多区域协调调度中的应用,利用多智能体深度强化学习(MADRL)解决多个独立参与者之间的博弈与合作问题。
- Integrated three-stage decentralized scheduling for virtual power plants: A model-assisted multi-agent reinforcement learning method(Biao Xu, Wenpeng Luan, Jing Yang, Bochao Zhao, Chao Long, Qian Ai, Jiani Xiang, 2024, Applied Energy)
- Real-Time Multi-Home Energy Management with EV Charging Scheduling Using Multi-Agent Deep Reinforcement Learning Optimization(Niphon Kaewdornhan, Chitchai Srithapon, Rittichai Liemthong, R. Chatthaworn, 2023, Energies)
- Demand Side Management and Peer-to-Peer Energy Trading for Industrial Users Using Two-Level Multi-Agent Reinforcement Learning(Jiatong Wang, D. Mishra, Li Li, Jiangfeng Zhang, 2023, IEEE Transactions on Energy Markets, Policy and Regulation)
- 基于多智能体强化学习的综合能源分布式优化(陶彩霞, 陈乃焜, 高锋阳, 张建刚, 2026, 系统仿真学报)
- Finding individual strategies for storage units in electricity market models using deep reinforcement learning(Nick Harder, A. Weidlich, Philipp Staudt, 2023, Energy Informatics)
- Multi-Agent Reinforcement Learning for Coordinated Smart Grid and Building Energy Management Across Urban Communities(Lei Qiu, 2025, Computer Life)
- Joint Energy and Carbon Trading for Multi-Microgrid System Based on Multi-Agent Deep Reinforcement Learning(Yanting Zhou, Zhongjing Ma, Tianyu Wang, Jinhui Zhang, Xingyu Shi, Suli Zou, 2024, IEEE Transactions on Power Systems)
- Multi-agent reinforcement learning for energy management in microgrids with shared hydrogen storage(David Toquica, K. Agbossou, N. Henao, 2025, International Journal of Hydrogen Energy)
- Energy Optimization for Building Energy Management with Thermal Storage: A Multi-Agent Deep Reinforcement Learning Approach(Shifeng Zhu, Wei Lu, Yi Feng, Changhai Sun, 2024, 2024 43rd Chinese Control Conference (CCC))
混合智能算法在复杂储能与微电网调度中的优化研究
该组文献集中研究深度强化学习与启发式算法(如PSO)或分层控制框架的结合,旨在解决高维、非线性及复杂约束条件下的电网优化及储能调度问题。
- Revolutionising Battery Energy Storage Systems Energy Management: Dynamic‐Aware Solutions With Integrated Quantum Particle Swarm Optimisation and Deep Reinforcement Learning(Yousef Asadi, Mohsen Eskandari, Milad Mansouri, 2026, IET Energy Systems Integration)
- 基于深度强化学习算法的跨区域多微电网系统扩展规划研究(周剑, 聂孝婷, 庞可欣, 王小越, 马义中, 2025, 中国管理科学)
- A swarm intelligence and deep learning strategy for wind power and energy storage scheduling in smart grid(Lin Geng, Lei Zhang, Fangming Niu, Yang Li, Feng Liu, 2024, International Journal of Intelligent Networks)
- 基于深度强化学习的分布式能源高渗透场景主动配电网动态重构(张学岩)
- 基于SAC-PSO分层协同强化学习的微网群优化调度策略(王一全)
- Optimal Scheduling Strategy of Electricity and Thermal Energy Storage Based on Soft Actor-Critic Reinforcement Learning Approach(Yingying Zheng, Hui Wang, Kailei Guo, Yuanrui Sang, 2023, Journal of energy storage)
面向碳中和与多能互补的深度强化学习调度策略
该组文献侧重于环境效益(碳减排)与多种能源载体(如水、光、氢、热、储)的互补优化,通过深度强化学习实现可持续、低碳的能源系统运行。
- Integration of deep reinforcement learning and parametric rule-based control for thermal storage management of district heating systems under spot price variations(Han Du, Xinlei Zhou, N. Nord, Y. Carden, Ping Cui, Zhenjun Ma, 2026, Energy)
- Mixture-of-experts based multi-critic deep reinforcement learning for sustainable management of data center microgrids(Qiong Liu, Zhanhua Pan, Ye Guo, Yue Chen, Goran Strbac, D. Qiu, 2026, Applied Energy)
- 深度强化学习驱动的水光储互补系统优化调度(向聪, 黄显峰, 李俊臣, 周士浩, 方国华, 周论, 2026, 水力发电学报)
- 基于深度强化学习的含氢综合能源系统低碳调度策略(李占凯, 曹明昊, 程学文, 2025, 现代电力)
- Efficient Low-Carbon Construction Pathways for Energy Ecosystems in Sustainable Cities through Deep Reinforcement Learning Management with Green Hydrogen Diversified Utilization under Trust Safe Peer-to-Peer Trading and Social Welfare(Pang Zhang, 2026, Energy)
储能系统运行中的鲁棒性、安全与前沿综合综述
该组文献主要关注储能系统在各种复杂环境下的韧性、安全性以及针对其在不同场景下的应用现状、模型方法与未来趋势的系统性研究。
- Multilevel Deep Reinforcement Learning for Secure Reservation-Based Electric Vehicle Charging via Differential Privacy and Energy Storage System(Sangyoon Lee, Dae-Hyun Choi, 2024, IEEE Transactions on Vehicular Technology)
- Energy management approach for wayside energy storage system in urban rail transit considering real-observable characteristics: A deep reinforcement learning method based on fuzzy logic guided(Yan Li, Fei Lin, Zhongping Yang, Xiaochun Fang, Xudong Lu, 2025, Journal of Energy Storage)
- Optimal Self-Consumption Scheduling of Highway Electric Vehicle Charging Station Based on Multi-Agent Deep Reinforcement Learning(Jianshu Zhou, Yue Xiang, Xin Zhang, Zhou Sun, Xuefei Liu, Junyong Liu, 2024, Renewable Energy)
- Resilient coordination of energy storage and renewable generation in residential systems using deep reinforcement learning considering extreme weather conditions(M. Salehpour, M.J. Hossain, M. Shafie-khah, 2026, Journal of Energy Storage)
- Powering Future Advancements and Applications of Battery Energy Storage Systems Across Different Scales(Zhaoyang Dong, Yuechuan Tao, Shuying Lai, Tianjin Wang, Zhijun Zhang, 2025, Energy Storage and Applications)
- Energy management of a microgrid considering nonlinear losses in batteries through Deep Reinforcement Learning(David Domínguez-Barbero, Javier García-González, M. A. Sanz-Bobi, A. García-Cerrada, 2024, Applied Energy)
现有文献围绕深度强化学习在共享储能及能源管理领域的研究呈现出高度的专业化分工:一是利用多智能体框架处理分布式能源交易与协作;二是采用混合算法和分层结构应对复杂电网系统的非线性调度;三是重点考虑低碳排放与多能互补的长期运行优化;四是针对系统韧性、安全性及前沿综述的研究,为储能作为能源体系核心的未来应用提供了理论与实证支持。
总计26篇相关文献
微电网在提升电力供应韧性和减少温室气体排放等方面展现出了巨大潜力,孤岛型微电网通过互联成为跨区域的多微电网系统,有利于实现微电网的经济性和供电韧性。针对跨区域多微电网系统的扩展规划问题,考虑相邻微电网间的能量互济,将供电韧性和环境效益作为约束,提出以最小化多微电网系统总成本为目标的长期扩展规划框架。基于深度强化学习算法,对此动态、随机决策优化问题给出了求解方法,结合真实数据构造了包含三个区域的多微电网系统,并以此作为算例验证模型的有效性。算例仿真结果表明,针对跨区域多微电网系统的规划框架不仅可提升微电网的供电韧性,而且能够考虑跨区域的微电网结构的影响,适时调整投资规划,选取电力依赖性更高、用途更广泛的区域进行微电网设施的投资,有效解决了跨区域多微电网系统的规划问题。
水、光、储电站独立运行的模式受输电通道容量约束,易出现弃水、弃光,制约电网系统的清洁能源消纳能力。为解决此问题,本文提出基于异步优势动作评价算法的水光储互补系统优化调度方法,适用于大规模水光储协同运行场景。首先,搭建水光储电站运行场景,以短期—中长期互补引导机制为基础构造优化调度模型;其次,将水光储互补系统的优化调度问题转化为马尔科夫决策过程,通过深度强化学习算法实现策略的高效探索与学习;最后以新疆叶尔羌河流域水光储互补系统为实例进行验证。结果表明,异步优势动作评价算法能稳定收敛到高奖励值,系统消纳电量显著提升,且计算时间显著低于其他算法,具有良好的工程应用价值。
为了提高含氢综合能源系统(hydrogen-integrated energy system,H-IES)的新能源利用率、降低碳排放水平,提出一种基于深度强化学习的H-IES调度策略。首先,构建包含氢能相关设备的H-IES模型,并引入奖惩阶梯式碳交易机制,利用阶梯碳价激励减排行为。然后,采用马尔可夫决策过程将调度问题转化为序列决策任务,综合考虑总成本、风光消纳率等多个目标,设计自适应奖励函数,动态调整各目标的权重。最后,在实测数据基础上,利用双延迟深度确定性策略梯度算法训练智能体。案例分析结果表明,所提方法能够挖掘H-IES的碳减排潜力,促进新能源消纳,实现经济性与环保性的双重优化。
基于预设规则开展工作,往往仅能达成单一目标的局部优化,导致主动配电网重构效率低下。针对上述问题,提出基于深度强化学习的分布式能源高渗透场景主动配电网动态重构。明确问题的抽象框架,并围绕该框架定义状态、动作空间以及优化目标与约束条件,进而构建配电网动态重构问题模型。运用深度Q网络,构建图结构并开展图卷积操作,同时结合损失函数,基于深度强化学习实现状态与动作的映射。将配电网划分为多个控制区域并部署独立智能体,通过信息交互与动作协调,实现主动配电网动态重构。实验显示,研究方法实现了主动配电网动态重构多目标的有效平衡,相较于对比方法具有更优的收敛速度与最终收敛值,提升了主动配电网重构效率。
针对分布式综合能源系统协调优化面临的能量管理和隐私保护问题,提出基于多智能体近端策略优化算法的分布式协调优化策略。在MDP框架下构建能源管理模型; 考虑电热异能特性,构建多区域双层交互机制; 在集中训练-分散执行的框架下, 利用同态加密避免协调过程中的隐私泄露问题 , 同时精准量化个体贡献 ,缓解多智能体策略评估方差激增问题;在系统小时级调度中以日成本最低为目标函数寻找最优策略。仿真结果表明:该算法可以基于大量历史数据自适应训练,完成最优策略的推导,在满足工程约束条件下同步缩减各区域的运营成本。
针对微网群优化调度中全局协调困难、局部约束复杂及风光出力不确定性强等问题,提出一种基于改进软演员—评论家(SAC)与改进粒子群优化(PSO)的分层协同调度方法。构建由微网群聚合商和多个子微网组成的双层优化模型,其中上层负责微网群整体功率协调,下层负责子微网内部自治优化。针对上层高维连续决策问题,引入变分自编码器(VAE)进行状态压缩,并结合时间衰减优先经验回放机制改进SAC算法,以提升学习效率与环境适应能力;针对下层多约束、非线性优化问题,设计包含自适应惯性权重、精英领导者选择及约束修复机制的改进PSO算法,以增强复杂可行域下的求解能力。
… for community-scale microgrids with hybrid energy storage. The method employed is the … thermal energy storage units with time-of-use electricity prices and stochastic renewable energy …
Battery Energy Storage Systems (BESSs) are critical in modernizing energy systems, addressing key challenges associated with the variability in renewable energy sources, and enhancing grid stability and resilience. This review explores the diverse applications of BESSs across different scales, from micro-scale appliance-level uses to large-scale utility and grid services, highlighting their adaptability and transformative potential. This study also includes advanced applications such as mobile energy storage, second-life battery utilization, and innovative models like Energy Storage as a Service (ESaaS) and energy storage sharing. Additionally, it discusses the integration of machine learning (ML) and large language models (LLMs), including advanced reinforcement learning (RL) algorithms, to optimize BESS operations and ensure safety through dynamic and data-driven decision-making. By examining current technologies, modeling methods, and future trends, this review provides a comprehensive overview of BESSs as a cornerstone technology for sustainable and efficient energy management, leading to a resilient energy future.
… This technique is what permits DRL to reach selections rapidly in convoluted, ever-… the dynamic features of Wind Energy (WE) and storage management are incomparable to traditional …
… integrated energy ecosystems. This paper proposes a knowledge-network-assisted deep reinforcement learning (DRL) framework for coordinated management of multi-energy clusters, …
… reinforcement learning (DRL), which learns robust coordination strategies for residential energy resources directly from stochastic environments. The framework is designed to optimize …
… a DRL-based ESS scheduling algorithm in which an ESS agent minimizes the operational energy cost of smart EVCS and conceals the energy … proposed three-level DRL algorithm. The …
Modeling energy storage units realistically is challenging as their decision-making is not governed by a marginal cost pricing strategy but relies on expected electricity prices. Existing electricity market models often use centralized rule-based bidding or global optimization approaches, which may not accurately capture the competitive behavior of market participants. To address this issue, we present a novel method using multi-agent deep reinforcement learning to model individual strategies in electricity market models. We demonstrate the practical applicability of our approach using a detailed model of the German wholesale electricity market with a complete fleet of pumped hydro energy storage units represented as learning agents. We compare the results to widely used modeling approaches and demonstrate that the proposed method performs well and can accurately represent the competitive behavior of market participants. To understand the benefits of using reinforcement learning, we analyze overall profits, aggregated dispatch, and individual behavior of energy storage units. The proposed method can improve the accuracy and realism of electricity market modeling and help policymakers make informed decisions for future market designs and policies.
Building Energy Management (BEM) with Thermal Energy Storage (TES) poses significant challenges due to the intricate coordination required among components such as Power-to-Heat (P2H) converters, TES units, and zone temperature controllers. In this paper, we propose a novel multi-agent Deep Reinforcement Learning (DRL) method for BEM, capable of optimizing the energy efficiency and flexibility of Building Energy Systems (BESs) integrated with TES without requiring manual efforts to set up a control-oriented model. Our method-distributed optimized NNPC via DRL-consists of two parts: i) a Neural Network Predictive Control (NNPC) for zone temperature regulation, leveraging an attention-based artificial neural network for forecasting operational conditions and a Dynamic Programming (DP) algorithm to jointly optimize both comfort levels and energy efficiency; and ii) a distributed optimization strategy realized through a coordinating policy that facilitates the collaboration between conventional on/off and NNPC agents/controllers. The efficacy of our method is confirmed through simulations established based on a real-world BES that integrates a TES unit. We compare our method with three alternatives: a conventional on/off control, an optimized on/off control, and an optimized Naïve NNPC, with the latter two also incorporating our distributed optimization strategy. Specifically, our method realizes approximately 7% reduction in energy consumption compared to the conventional on-off method while maintaining equivalent comfort levels. Furthermore, it attains parity with the performance of the optimized Naïve NNPC without requiring manual tuning, thereby illustrating its adaptability and ease of implementation. Collectively, these findings underscore our method’s superiority across metrics of energy efficiency, comfort and applicability, demonstrating the profound potential of DRL in BEM, particularly when TES is involved.
… of PCM storage. Unlike most existing studies, this study first trains the DRL agent using the … thermal energy storage, bridging simulation-based research and real-world application. …
The increasing integration of inverter‐based resources (IBRs) into microgrids (MGs) poses considerable challenges for dynamic stability and energy management, primarily due to their variable and uncertain inertia and damping characteristics. Unlike conventional synchronous generators (SGs), IBRs exhibit complex nonlinear behaviour, complicating both mathematical modelling and real‐time control. Battery energy storage systems (BESS) are essential for ensuring grid stability, operational efficiency and flexibility. Nevertheless, dynamic‐aware energy management of BESS remains insufficiently explored, with current approaches often lacking adaptability to uncertainty and real‐time requirements. This paper proposes a two‐step hybrid framework combining quantum particle swarm optimisation (Q‐PSO) with deep reinforcement learning (DRL). In the first step, Q‐PSO efficiently generates an initial solution, significantly reducing computational demands. Subsequently, DRL dynamically refines this solution, effectively managing real‐time uncertainties linked to inertia and damping variations. The proposed method addresses the non‐Markovian nature of MG dynamics by constraining the DRL action space using the Q‐PSO‐derived solution, thereby alleviating the curse of dimensionality and enhancing training stability. Furthermore, dynamic constraints on frequency deviations and the rate of change of frequency (RoCoF) are incorporated to maintain robust grid stability during transients. Extensive simulations demonstrate that the proposed dynamic‐aware energy management system (EMS) achieves economic efficiency improvements of 52.2% compared to Q‐PSO alone and 22.7% compared to DQN alone. Additionally, BESS charging efficiency improves by 39.6% and 22.7%, whereas discharging efficiency increases by 38.3% and 28.25%, respectively, against the same benchmarks.
… proposes a DRL approach … DRL methods applied to ESS energy management in urban rail systems, the fuzzy logic guided agent demonstrates a substantial enhancement in energy-…
… DRL) algorithms have been employed to solve the latter, gaining popularity in recent years. Since DRL … – a state-of-the-art method in DRL – for the EMS of a microgrid that includes …
… In DCs integrated with aquifer thermal energy storage, a context-aware RL approach has … shared MoE structure within the multi-critic DRL framework. Each critic head leverages shared …
Energy management for multi-home installation of solar PhotoVoltaics (solar PVs) combined with Electric Vehicles’ (EVs) charging scheduling has a rich complexity due to the uncertainties of solar PV generation and EV usage. Changing clients from multi-consumers to multi-prosumers with real-time energy trading supervised by the aggregator is an efficient way to solve undesired demand problems due to disorderly EV scheduling. Therefore, this paper proposes real-time multi-home energy management with EV charging scheduling using multi-agent deep reinforcement learning optimization. The aggregator and prosumers are developed as smart agents to interact with each other to find the best decision. This paper aims to reduce the electricity expense of prosumers through EV battery scheduling. The aggregator calculates the revenue from energy trading with multi-prosumers by using a real-time pricing concept which can facilitate the proper behavior of prosumers. Simulation results show that the proposed method can reduce mean power consumption by 9.04% and 39.57% compared with consumption using the system without EV usage and the system that applies the conventional energy price, respectively. Also, it can decrease the costs of the prosumer by between 1.67% and 24.57%, and the aggregator can generate revenue by 0.065 USD per day, which is higher than that generated when employing conventional energy prices.
… reinforcement learning algorithm to address such a management problem. Each customer learns an energy … for the customers to tune the energy management strategies and outperform …
Carbon trading has emerged as an effective way to promote the renewable generation and sustainable energy development. Since carbon emissions are closely coupled to energy system, it is a challenge to design a market mechanism for joint energy and carbon trading to achieve better strategies. In this study, the energy management problem with a specific focus on joint trading in multi-microgrid system is investigated by utilizing a multi-agent deep reinforcement learning approach. Initially, a joint energy and carbon trading market is established and the dispatch optimization problem is formulated as a Markov decision process without modeling uncertainties accurately. This mechanism enables direct one-to-one energy transactions among all areas, avoiding the market clearing in traditional multi-party local energy trading markets. To enhance the learning efficiency and maintain agent privacy, an enhanced multi-agent proximal policy optimization (MAPPO) algorithm that incorporates a parameter sharing mechanism is introduced. Moreover, the recurrent neural networks (RNN) structure is leveraged to perform feature encoding for individual agents, which improves the overall feature extraction capability. Through comprehensive experiments involving various algorithms, the proposed approach reduce operating costs 14.86 $\%$ and carbon emissions 19.04 $\%$ compared with traditional MAPPO, which validates the effectiveness and performance benefits.
… scheduling based on the MADRL is proposed. Regarding the multi-objective scheduling of the … and non-convex problem is transformed into multi-agent Markov game model. Finally, the …
Over the past few years, with more concerns on energy equity, energy security and environmental sustainability, peer-to-peer (P2P) energy trading market has increasingly attracted attention to energy users. With the rising electricity price and the growing public doubt to unsustainable development, more efforts are given to the integration between P2P energy trading and industrial practice. To date, the current research work has emphasized electricity cost reduction using demand-side management (DSM) and P2P energy trading market to promote energy sharing and create a win–win situation for all industrial market participants. However, less attention has been given to evaluating industrial production. With this attention, this paper formulates a cost minimization problem for multiple discrete manufacturers, where DSM and P2P energy trading market are integrated to optimize manufacturing productivity and energy consumption. Moreover, this paper proposes a two-level reinforcement learning (RL) method to obtain the optimal DSM and energy trading strategies for discrete manufacturers. Unlike the existing RL approaches, the proposed method depicts a two-level learning approach, where the upper-level program suggests a production plan for the manufacturer to improve the cost return, while the lower-level program optimizes the net benefits. Further, to verify the performance of the proposed method, simulations are conducted based on four different case studies, including different weather and load conditions. Numerical results demonstrate that the proposed algorithm provides a better production plan with enhanced learning speed and economic benefits compared to the existing RL approaches.
The increasing complexity of urban energy systems necessitates sophisticated coordination mechanisms between smart grids and building energy management systems to achieve optimal energy efficiency and sustainability goals. This research presents a novel Multi-Agent Reinforcement Learning (MARL) framework designed to facilitate coordinated energy management across interconnected urban communities. The proposed system employs distributed intelligent agents that operate autonomously while maintaining collaborative decision-making capabilities for energy distribution, consumption optimization, and demand response coordination. Through extensive simulation studies conducted across three metropolitan areas involving 1,247 residential and commercial buildings over a 12-month period, our findings demonstrate significant improvements in energy efficiency, with average reductions of 23.4% in peak demand loads and 18.7% in overall energy consumption costs. The MARL approach exhibits superior performance compared to traditional centralized control systems, particularly in handling dynamic load balancing and renewable energy integration challenges. These results contribute substantially to the advancement of sustainable urban energy infrastructure and provide practical insights for large-scale smart city implementations.
… -stage scheduling framework for VPPs and develops a model-assisted multi-agent reinforcement learning (… In the proposed framework, the VPP scheduling problem is formulated as a …
现有文献围绕深度强化学习在共享储能及能源管理领域的研究呈现出高度的专业化分工:一是利用多智能体框架处理分布式能源交易与协作;二是采用混合算法和分层结构应对复杂电网系统的非线性调度;三是重点考虑低碳排放与多能互补的长期运行优化;四是针对系统韧性、安全性及前沿综述的研究,为储能作为能源体系核心的未来应用提供了理论与实证支持。