科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ IEEE Open Journal of Vehicular Technology2025-12-10· Reinforcement learning

A UAV Path Planning Method Based on Deep Reinforcement Learning With Dense Rewards

Jianhong Zhou, Yong Wang, Qian Xie, Zixia Shang, Yinliang Jiang, Qiuyu DU

原始摘要(英文原文)· Original abstract
Most state-of-the-art (SOTA) unmanned aerial vehicle (UAV) path planning approaches depend on global environmental knowledge. While algorithms like adaptive soft actor-critic (ASAC) have improved training efficiency, their obstacle avoidance in partially observable environments remains limited. To address this, we propose a depth-based collision risk prediction (DCRP) algorithm that integrates into the ASAC framework. DCRP processes depth images alongside UAV pose and velocity to calculate a dense collision risk signal, enriching the reward function for more effective avoidance learning. Furthermore, we enhance the policy network with a novel skip connection that directly injects critical state information into the final action output. This innovation mitigates gradient vanishing and accelerates policy learning. Additionally, a generalized transfer learning (GTL) strategy accelerates convergence in complex environments by leveraging policies pre-trained in simpler ones. Extensive evaluation in high-fidelity AirSim environments demonstrates the superiority of our method. It outperforms several SOTA baselines, achieving an approximately 20% higher task success rate and 39% faster training efficiency on average, while maintaining a real-time inference time of around 15 ms.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

A UAV Path Planning Method Based on Deep Reinforcement Learning With Dense Rewards — 科研速览 Science Skim