Z. -Zhao Xiao, Shenghai Yuan, Guili Xu, Xianglong Zeng, Huanran Hu, Junwei He, Yizhuo Yang
The rapid proliferation of compact Unmanned Aerial Vehicles (UAVs) poses serious challenges to low-altitude airspace security. However, existing audio-based and audio-visual-based 3D UAV detection models rely on Convolutional Neural Networks (CNNs) for feature extraction. The translation-invariance of CNNs does not align with the physical properties of audio mel-spectrograms, thereby impairing trajectory estimation and classification. Furthermore, the audio-visual models lack adaptive mechanisms for evaluating modality-specific contributions and reliability, which leads to significant performance degradation when any single modality becomes unstable. To address these limitations, we propose an annotation-free Audio-Visual fusion framework for Drone Trajectory Estimation and Classification (AV-DTEC). AV-DTEC introduces three key innovations. First, we design the Audio mamba (Adm) to address the translation-invariance of CNNs, which is integrated with Vim-Tiny for visual encoding to construct a dual-stream state-space model capable of effectively capturing fine-grained spatiotemporal patterns across modalities. Second, we develop a Primary-Auxiliary Feature Fusion Module (PAFFM) equipped with an adaptive adjustment mechanism to evaluate modality reliability, whichdynamicallyintegrates auxiliary features into primary features. Third, we leverage an unsupervised LiDAR-based pipeline and a Vision-Language Model (VLM) to generate 3D trajectory and classification pseudo-labels, enabling self-supervised training without the reliance on costly manual annotations. Extensive experiments on the real-world MMAUD dataset demonstrate that AV-DTEC achieves superior trajectory estimation and classification performance compared to single modality models and multimodal fusion models. Code and models are available at: https://github.com/AmazingDay1/AV-DETC.