Ye Wang, Jingjing Wang, Jianrui Chen, Xiangwang Hou, Ziyang Wang, Chunxiao Jiang
The evolution of the Internet of vehicles (IoV) has introduced computation-intensive and latency-sensitive applications that challenge traditional cloud architectures. Although drone-aided IoV offers a flexible solution, it presents a complex optimization problem. The core challenge lies in balancing task offloading efficiency with crucial operational safety constraints, such as collision avoidance and battery management, a gap often overlooked in existing research. This paper addresses this problem by first modeling the drone-aided task offloading system as a constrained multi-agent Markov decision process. Based on this framework, we propose a novel safe multi-agent reinforcement learning algorithm (MARL) named Lagrangian-constrained multi-agent policy optimization (LC-MAPO). The LC-MAPO integrates safety constraints into the twin delayed deep deterministic policy gradient (TD3) actor-critic framework using Lagrangian duality theory. The algorithm's effectiveness was validated in three distinct simulation scenarios and compared against an unconstrained multi-agent deep deterministic policy gradient (MADDPG) algorithm and a greedy algorithm. Experimental results demonstrate that LC-MAPO achieves superior performance in both safety adherence and task processing efficiency.