Chen Dai, Yize Wang, Tianyu Wu, Xiaolin Gu, H. Cao, Guozi Sun
Integrated sensing and communication (ISAC) is envisioned as a key enabler of 6 G networks, allowing communication and sensing to share spectrum and hardware resources in a unified manner. Unmanned aerial vehicles (UAVs) can further enhance ISAC systems by providing flexible deployment and reliable line-of-sight (LoS) links. However, the openness of UAV communication channels, coupled with their dynamic mobility and limited energy, makes them highly vulnerable to eavesdropping and exacerbates physical layer security (PLS) concerns. To address these challenges, we propose a reconfigurable intelligent surface (RIS)-assisted framework for enhancing PLS in dynamic and uncertain UAV-enabled ISAC networks. We formulate a secrecy-oriented joint optimization problem that integrates UAV trajectory design, RIS phase shift configuration, beamforming, and decoupled uplink–downlink (DUDe) user association. To render this problem tractable, we adopt a two-stage approach. First, the beamforming subproblem is solved via fractional programming and successive convex approximation to obtain closed-form solutions. Subsequently, with beamforming fixed, we reformulate the remaining coupled design of trajectories, associations, and RIS phases as a Partially Observable Markov Decision Process (POMDP). To address the challenges of high-dimensional sequential decision-making and the need for scalable, real-time coordination inherent in multi-UAV systems, we develop a robust multi-agent deep reinforcement learning (MADRL) solution, effectively circumventing the prohibitive communication overhead of centralized control and the intractability of single-agent formulations suffering from the curse of dimensionality. Specifically, we design a stochastic environment-aware Proximal Policy Optimization (PPO) algorithm that incorporates random environment distribution to capture environmental randomness, intrinsic rewards to guide UAV exploration of secure regions, and uncertainty estimation to refine policy updates. Simulation results verify that the proposed framework significantly outperforms benchmark methods in terms of secrecy rate and robustness, especially under dynamic scenarios with varying numbers of users and eavesdroppers.