Duo Yang, Haoran Lv, Yu Yan, Mince Li, Rui Pan, Jing Liang
Hydrogen fuel cell vehicles typically integrate energy storage components like lithium batteries to enhance dynamic response and support regenerative braking. The energy management strategy (EMS) plays a crucial role in optimizing power distribution among power components. Although deep reinforcement learning (DRL) offers strong adaptability and optimization capabilities for EMS, its training process often demands extensive data, computational resources, and careful hyperparameter tuning, posing challenges in stability and safety. To address these limitations, this paper proposes a hierarchically structured knowledge-transfer framework that synergistically fuses offline expert guidance with online reinforcement autonomy. First, a supervised behavior cloning procedure extracts optimal power distribution policies from dynamic programming trajectories, achieving hydrogen consumption minimization through direct policy imitation. A deep neural network (DNN) is then trained to replicate these rules, providing a pre-trained model for online EMSs. Subsequently, a twin-delayed deep deterministic policy gradient (TD3)-based EMS is developed for real-time power allocation, incorporating hydrogen consumption, fuel cell degradation, and battery state-of-charge fluctuations into the reward function. Finally, the pre-trained DNN initializes the TD3 actor network, effectively improving training efficiency and strategy security. The results demonstrate that the proposed framework accelerates convergence, improves system performance, and ensures safer control strategies compared to conventional DRL-based methods.