Hao Peng, Ke Jiang, Deli Yuan, Ziqing Zeng, Zihan Wu
To overcome the critical energy weakest-link effect in Autonomous Underwater Vehicle (AUV) swarms, this study introduces a wild goose formation-inspired hierarchical energy-equilibrium framework. The proposed approach employs a Long Short-Term Memory-Enhanced Multi-Agent Proximal Policy Optimization with energy penalty reward function (EP-LSTM-MAPPO) to enable dynamic energy-role matching, where high-SOC AUVs autonomously lead long-range tasks while low-SOC units follow short-range roles. A three-tier architecture ensures multi-agent collaboration, path feasibility, and precise motion control. Experimental evaluations against Frontier-Based Exploration (FBE) and LSTM-MAPPO without energy penalty (NE-LSTM-MAPPO) algorithms demonstrate the framework's significant superiority: it reduces average State-of-Charge range (ΔSOC) by 55.9 %, increases coverage completeness by 128.9 %, and lowers average energy consumption per cycle by 52.0 %, while achieving an 82.0 % improvement in the Comprehensive Efficiency Index. This framework maintains robust performance under hydrodynamic disturbances and scales effectively in complex environments, establishing a sustainable paradigm for long-duration AUV swarm operations with direct applications in large-area underwater mapping, environmental monitoring, and persistent surveillance.