Xinfu Pang, Bin Wang, Yefeng Liu, Wei Liu, Zedong Zheng
Against the backdrop of China’s carbon peak and carbon neutrality goals, the energy system is undergoing rapid transformation. As a vital component, virtual power plants (VPPs) have become a key pathway for integrating distributed energy resources (DERs) into power systems and electricity markets. However, VPP operations still face challenges, such as fluctuating renewable energy output, uncertainty in demand response capabilities, and strategic behavior. To address this, a bi-level optimization scheduling method for VPPs is proposed via Stackelberg game and Q-learning-based differential evolution (QDE). First, a two-stage Stackelberg dynamic game model is established for VPP operators and the user side. VPP operators optimize electricity prices and controllable unit outputs, while industrial users and electric vehicles adjust their loads according to electricity prices to maximize flexibility. Secondly, a solution method combining Q-learning-based differential evolution with quadratic programming is proposed. The upper layer employs QDE optimization, while the lower layer is solved via quadratic programming. Q-learning adaptively adjusts crossover and mutation rates to enhance search capability; parameters are determined through orthogonal experiments, with multi-factor variance analysis validating its effectiveness. Thirdly, sensitivity analysis reveals the impact of carbon trading interval lengths on emissions, identifies the optimal benchmark electricity price, and demonstrates the model’s effectiveness under renewable generation uncertainty. Finally, simulation results demonstrate that this approach reduces carbon emissions and EV charging costs by 20.31% and 8.67%, respectively.