xiaoqiang wen, Zhongan Duan, Wang jianguo, Qiao Hong
Abstract The uncertainties associated with renewable energy generation and the challenges in the precise modeling of microgrids present significant difficulties for conventional planning algorithms. Owing to its model-free nature, strong hyperparameter robustness, and broad applicability, the Proximal Policy Optimization (PPO) algorithm, which is based on the Actor-Critic architecture, shows considerable promise as an ideal solution for the optimal scheduling of microgrids. However, in the standard PPO training process, the update of the value network lags behind that of the policy network, which compromises training efficiency. Moreover, the inherent randomness in agent exploration can lead to instability in policy updates. To address these limitations, this study proposes an enhanced PPO algorithm that incorporates a policy feedback mechanism and employs a segmented clipping mechanism. These improvements significantly accelerate the overall training convergence speed and effectively mitigate severe fluctuations during training, thereby enhancing the stability of the network updates. Experimental results in a microgrid energy management scenario demonstrate that the proposed algorithm achieves faster convergence and greater stability during the training phase. During the testing phase, it is capable of generating high-quality intra-day scheduling strategies based on real-time electricity consumption conditions. Over a testing period of ten days, the proposed algorithm achieved a 19.87% reduction in cumulative electricity costs and a 20.25% decrease in cumulative power imbalance compared to the standard PPO. The proposed algorithm outperforms the existing PPO method, offering an effective solution for microgrid energy management and demonstrating substantial application value.