Zhige Yuan, Yu Zeng, Amer M. Y. M. Ghias, Pengfeng Lin, Josep Pou
Conventional deep reinforcement learning (DRL) frameworks for power electronics face several critical limitations, including nonreal-time training environments, poor generalization due to sim-to-real mismatches, inadequate thermal modeling, cumbersome controller deployment, and prolonged training durations. To address these challenges, this paper proposes a real-time augmented DRL (RTA-DRL) framework that integrates the PLECS environment to enhance training fidelity and ensure real-time compatibility. Within this framework, the policy network is periodically deployed into PLECS hardwarein- the-loop (HIL), where real-time data are incorporated into the training process to strengthen policy robustness. The HIL stage further serves as an intermediate evaluation step to verify network execution and deployment feasibility on the target platform. An automated workflow based on a JSON-RPC proxy enables seamless communication between MATLAB and PLECS. Additionally, the framework incorporates the power switches' thermal model and parameter mismatch considerations in HIL evaluations to enhance generalization and bridge the sim-toreal gap. The proposed RTA-DRL is validated on a dualactive- bridge converter with triple-phase-shift control, targeting efficiency optimization under varying battery voltage conditions. Hardware experimentations demonstrate reduced training time, improved energy efficiency, and enhanced policy generalization of the proposed RTA-DRL framework.