CHEN Guili, CHEN Danhong, WEN Hongwu, ZHENG Yongwei, MAI Liang, FENG Xiashan, ZHONG Qishen, HU Yiming, ZHENG Jiehui
[Objective] To address the variations in resource allocation and energy consumption behavior among prosumers in peer-to-peer (P2P) community energy trading scenarios, as well as the limited adaptability of traditional model-based methods in uncertain environments, this paper proposes a multi-agent reinforcement learning method that features both scalability and privacy protection capabilities. [Methods] First, three representative types of heterogeneous prosumer models are constructed. Second, a community energy trading model based on the mid-market rate pricing mechanism is established, and a flexibility incentive mechanism is introduced. Finally, the energy trading decision-making problem of prosumers is formulated as a partially observable Markov decision process, and a soft actor-critic algorithm based on dynamic mean-field (DMF-SAC) approximation is proposed to solve the energy management strategies of prosumers. [Results] Simulation results demonstrate that the proposed method outperforms baseline methods in terms of convergence performance, computational overhead, and operating costs. It also effectively improves the local consumption of distributed energy and enhances peak-shaving and valley-filling capabilities. [Conclusions] The proposed method effectively improves the efficiency and economic benefits of collaborative optimization for heterogeneous prosumers while balancing privacy protection and system scalability, which holds significant value for energy trading and management in community-based markets.