Xiaona Song, Haozhe Zhang, Zhengyi Huang, Jianlin Zhao, Lijun Wang
This paper proposes an improved multimodal fusion framework for 3D object detection, termed BEV-Nexus, which aims to address the issues of inaccurate depth estimation and inefficient fusion paradigms in existing image-point cloud fusion methods. We introduce a Point-Cloud-Guided Depth Prediction Network (PCGD-Net), which enhances the image branch's depth prediction capability by embedding point cloud spatial prior, ground-truth loss constraint, and projected point cloud depth filling. Additionally, we design a Dynamic Self-adaptive Feature Fusion Module (DSF-Module), which computes multimodal feature similarity using window attention and performs weighted fusion based on self-adaptive weights, resolving alignment deviations in BEV features. Finally, we propose a Dilated Attention Enhancement Block (DAEB), which expands the receptive field through dilated convolution and integrates parameter-free attention mechanism (SimAM) for feature enhancement, ensuring efficiency while improving overall feature representation. Experimental results on nuScenes validation set show that BEV-Nexus outperforms it baseline (BEVFusion) by 1.8% mAP and 1.5% NDS. On the test set, BEV-Nexus improves mAP and NDS by 1.6% and 1.4%, respectively. Furthermore, the detection FPS remains nearly unchanged, demonstrating significant lightweight advantages.