Nader Alotaibi, Wojdan BinSaeedan
Cooperative multi-UAV path planning under dynamic and adversarial conditions demands simultaneous satisfaction of safety, efficiency, and coordination constraints, yet existing swarm-intelligence and RL–swarm hybrids rely on deterministic switching rules, tabular states, and ad hoc training schedules. This paper proposes RL-JSO, a hybrid framework in which a dueling double deep Q-network with prioritized experience replay adaptively selects among the drift, passive, and active phases of a jellyfish search optimizer, replacing the deterministic time-control rule with a learned policy. The framework integrates a five-layer hierarchical safety control mechanism, a mastery-gated nine-stage curriculum, and a shared reward module that architecturally enforces fairness between RL-JSO and a paired RL-PSO counterpart. Evaluation across four progressive campaigns with 160 independent runs per algorithm shows that, within the evaluated JSO/PSO family, RL-JSO is the only method that sustains a 100% collision-free rate across all four progressive difficulty campaigns, its Cliff’s delta over standard JSO grows monotonically with difficulty from medium to large, and under a composite cooperation metric its coordination score remains nearly invariant while comparators degrade by 17–23%. A paired inference-time ablation on the trained checkpoint provides controlled inference-time evidence that adaptive phase switching is a principal contributor to the observed test-time performance within the trained system, rather than the heuristic fallback layers.