Hao Jiang, Xuting Pan, Wangqi Shi, Linzhou Zeng, Zhen Chen, Feng Shu, Jiangzhou Wang
This paper proposes a geometry-based air-to-ground (A2G) channel model that captures time-varying velocities of unmanned aerial vehicles (UAVs), mobile receiver (MR), and scatterers in realistic propagation environments. To enhance the UAV energy efficiency during flight, the model integrates deep reinforcement learning (DRL) to enable real-time decision-making within each discrete time interval. Specifically, the approach utilizes an enhanced twin-delayed deep deterministic policy gradient (TD3) algorithm with a dual-layer actor network and twin critic networks, which further improves its ability to efficiently handle complex decision-making tasks in dynamic environments. Additionally, key statistical characteristics are derived rigorously and confirmed through numerical experiments, including spatial, temporal, and frequency domain correlations. The numerical results demonstrate that the proposed approach achieves superior energy efficiency and stability over the soft actor-critic (SAC), proximal policy optimization (PPO), and deep deterministic policy gradient (DDPG) methods, ensuring optimal communication performance and UAV trajectories. The study underscores the effectiveness of DRL in optimizing UAV trajectories, offering valuable insights for the design of advanced A2G wireless communication systems.