Huang Chen, MingJun Dai
The centralized training with decentralized execution (CTDE) paradigm has achieved strong performance in multi-agent reinforcement learning (MARL), yet on-policy actor–critic methods such as MAPPO can still exhibit training instability and late-stage performance collapse on coordination-intensive tasks. We address this issue with Self-Distillation MAPPO (SD-MAPPO), which maintains an Exponential Moving Average (EMA) teacher policy and regularizes the current actor with a KL-based temporal consistency term. Rather than correcting the underlying critic estimates, the proposed mechanism provides a slowly evolving policy reference. From a local analytical perspective, the EMA teacher bounds teacher–student drift over time, while the KL penalty locally damps abrupt policy fluctuations induced by noisy learning signals. Empirically, we evaluate on challenging SMAClite micromanagement tasks and complement win rate with test return, final-stage statistics, and late-stage update diagnostics. Across the primary four-task suite, SD-MAPPO improves final-stage robustness on the more collapse-prone maps while remaining broadly competitive on the easier tasks, with especially clear gains on 3s5z_vs_3s6z and 10m_vs_11m . Diagnostic analyses further show that the method consistently reduces late-stage policy-shift indicators such as old–new KL and clipping activation, even when critic-side changes are only mild. Overall, the results support SD-MAPPO as a lightweight actor-side stabilizer for instability-prone on-policy CTDE training rather than a universal improvement module. The method is plug-and-play, preserves decentralized execution, and adds only a small constant-factor training-time overhead. At deployment, action selection remains in the same actor-only regime as MAPPO because the EMA teacher is not used.