Elham Yazdani Bejarbaneh, Haiping Du, Jun Shen
Cooperative driving of connected and automated vehicles (CAVs) is envisioned as a promising approach to improving fuel efficiency, safety, and traffic flow. However, achieving robust and efficient control in heterogeneous CAV platoons remains challenging, especially under uncertain dynamics and external disturbances. Most existing reinforcement learning (RL)-based platoon controllers often ignore model uncertainties or rely on centralized training, limiting their scalability and robustness in real-world applicability. To address these limitations, this study proposes a fully distributed, model-free RL framework integrated with a robust compensator for optimal platoon control of heterogeneous vehicles with unknown dynamics. The RL agent simultaneously learns the optimal control policy and estimates control-relevant dynamic parameters using only local input-output data, without requiring explicit vehicle models. These estimates are then used in real time to construct a disturbance-rejection input that ensures robust trajectory tracking. A distributed observer based on consensus theory is embedded within the hybrid controller to estimate leader-relative reference trajectories using only local neighbor information, eliminating the need for global communication. Theoretical analysis guarantees policy convergence and bounded tracking errors under dynamic uncertainties. The proposed method is experimentally validated using the high-fidelity Mixed Traffic Simulation (MiTaS) platform, combining the SUMO microscopic traffic simulator with MATLAB, demonstrating improved tracking, damping of traffic oscillations, and up to 15.8% fuel savings compared to recent RL-based methods.