Tianheng Xu, Yang Gu, Wen Du, Kai Ying, Qingqing Wu, Pei Peng, Dusit Niyato
This letter proposes a distributed path selection algorithm based on the meta multi-agent proximal policy optimization (Meta-MAPPO). The algorithm leverages transferable knowledge to achieve faster and more stable policy optimization in dynamic satellite networks. We integrate meta-learning into the MAPPO framework, equipping agents with rapid adaptation capabilities and enhancing convergence efficiency through experience sharing. Simulation results on a 96-satellite Walker–Delta constellation demonstrate that the proposed framework achieves at least a 5% reduction in average end-to-end delay, maintains zero packet loss, and converges faster, demonstrating its efficiency and robustness in dynamic satellite network environments.