Bin Huang, Tianqiao Zhao, Meng Yue, Xiaodong Zheng, Jianhui Wang
The emergence of quantum computing offers a transformative approach to energy arbitrage (EA) in energy storage systems (ESSs) by enabling efficient, data-driven solutions with enhanced learning capabilities. This study introduces a Quantum Policy Learning (QPL) algorithm designed to optimize online EA processes in energy markets. Unlike traditional reinforcement learning, which relies on neural networks for policy approximation, QPL utilizes a Variational Quantum Circuit (VQC) as the parameterized policy model. The VQC is tailored for the Noisy Intermediate-Scale Quantum era, featuring a low-depth and low-width design with alternating variational entanglement and embedding sub-layers. The loss function and gradient calculations are derived using action re-parameterization. Case studies indicate that QPL achieves faster convergence and competitive optimization performance compared to classical algorithms while using only 1.6% of the number of trainable parameters. To highlight the expressivity advantage, VQCs are further compared with classical neural networks through a harmonic-regression benchmark and the downstream ESS arbitrage task. Experiments on IBM quantum hardware confirm QPL’s resilience to gate errors and qubit decoherence, and additional analyses demonstrate its superior training stability and efficiency.