Qi Liu, Daqiao Zhang, Shaopeng Li
In multi-UAV mountain search and rescue scenarios, the perception system of multi-UAV suffers from low utilization of noise resources, poor collaboration of multi-modal data, and a persistent imbalance between speed and detection accuracy. The paper proposes a federated multi-modal perception method based on terrain-adaptive variational positive-incentive noise (FedVPN). The framework transforms complex mountain interference into task-related beneficial noise, constructs a privacy-preserving federated multi-modal collaborative architecture for distributed feature fusion, and adopts a two-stage training pipeline. Under three typical scenarios, FedVPN outperforms all five baseline methods. In the basic scenario, it achieves an F1-score of 89.23% with a noise gain rate of 7.86%. Under dynamic interference conditions and large-scale heterogeneous environments, the performance decay is only 3.59% and the rescue response time is reduced to 48.60 s. The method significantly improves the accuracy, robustness, and efficiency of the perception module for autonomous rescue systems.