Yongjia Nian, Hao Liu, Renwen Chen, Xintong Hou, Aocheng He
This paper investigates the challenge of topology optimization for UAV swarms in dynamic environments and proposes a reinforcement learning–driven distributed framework. Under the centralized training and decentralized execution (CTDE) paradigm, a MADDPG-based topology reconfiguration algorithm is developed that integrates partial observability with a bi-directional interest game, enabling nodes to achieve distributed Nash equilibrium decisions under local information constraints. At the communication layer, a channel model, topology maintenance scheme, and CSDMA-based distributed slot allocation process are introduced to ensure reliable connectivity in the presence of interference and dynamic node access. Simulation results show that the proposed method attains faster convergence, greater robustness, lower communication latency, and higher path efficiency than benchmark approaches such as MST and PSO, with reconfiguration completed within milliseconds. These results highlight both the effectiveness and scalability of the framework for large-scale swarm networking. Beyond its theoretical contributions, the approach holds practical promise for deployment in critical scenarios such as emergency communications, disaster relief, and mission-critical operations, offering a viable pathway toward intelligent UAV swarm networks.