Zhewen Cui, Wei Guan, Xianku Zhang, C. Guedes Soares
This study addresses a significant gap in the application of multi-agent reinforcement learning (MARL) to unmanned aerial vehicle and unmanned surface vehicle (UAV-USV) heterogeneous collaboration collision avoidance decision-making (HCCAD). The MARL has demonstrated considerable promise in homogeneous multi-agent systems; however, its extension to heterogeneous platforms with disparate dynamics, sensing modalities, and action spaces remains underexplored. To bridge this gap, an optimal baseline multi-agent proximal policy optimisation (OB-MAPPO) framework tailored for HCCAD problem is proposed. The innovations of the proposed method are: (1) The framework incorporates the advantage decomposition mechanism and optimal baseline to reduce gradient variance and mitigate redundant parameter sharing caused by heterogeneous action spaces. (2) A multimodal fusion network based on the Swin Transformer is introduced to effectively integrate visual data from the UAV and LiDAR observations from the USV, resolving spatial discrepancies in cross-domain perception. (3) The reward function is designed to incorporate collaborative error, ship domain, and immediate danger situations. Simulations show our method achieves superior convergence, success rate, and generalization in complex dynamic environments compared to existing MARL and conventional methods. Our work presents a novel, robust, and scalable paradigm for collision avoidance in cross-domain operations involving both aerial and maritime vehicles.