S. E. Lee, Ngan-Khanh Chau, Sanghun Choi
This study presents a deep reinforcement learning (DRL)-based temperature control strategy for systems utilizing thermoelectric modules. Although traditional proportional–integral–derivative (PID) controllers are widely adopted, they frequently encounter issues such as overshooting and slow convergence, particularly in nonlinear and time-delay systems. To address these limitations, we developed a simulation-based environment powered by an artificial neural network trained on experimental data. Within this environment, a reinforcement learning agent was trained using the proximal policy optimization (PPO) algorithm, enabling it to learn optimal control actions directly, without relying on predefined PID structures. The RL controller was evaluated within an ANN-based simulation environment trained on experimental data, serving as a virtual platform for validation prior to hardware implementation. The results indicate that the RL-based controller significantly outperforms conventional PID control, effectively reducing overshoot and maintaining the target temperature within a ±1 K range. Moreover, compared to the PID controller, the RL-based approach reduced overshoot by 78.6% and shortened the settling time by 51.1%. The proposed framework was validated across various target scenarios, demonstrating strong generalization capabilities. While currently evaluated in a computational setting, the method demonstrates considerable promise for real-world deployment and can be extended to control applications demanding precision, robustness, and adaptability.