Weiming Jing, Xiyuan Chen, Shuhan Nie, Zhiyuan Jiao, Jianghui Ma
In autonomous driving, bird’s-eye view (BEV) representations have emerged as the dominant approach for 3-D object detection. However, projecting 3-D objects into BEV space can lead to distant and nearby objects appearing similar in size, making it challenging to discern depth relationships between foreground and background objects. Furthermore, inadequate modeling of intermodal discrepancies and correlations hampers effective contextual integration in cross-modal fusion. To address these limitations, we propose hierarchical dynamic fusion (HD-Fusion), a novel end-to-end multimodal fusion framework consisting of a scene-level fusion (SLF) module and a contextual-level fusion (CLF) module. The SLF module fuses depth details from point cloud pillars with image features, generating BEV image representations enhanced with depth cues. The CLF module further enhances the features of LiDAR and cameras with a bidirectional cross-modal attention (BCMA) block and a discrete wavelet transform (DWT) encoder. The BCMA captures long-range interactions between LiDAR and image tokens, while the DWT separates multiscale frequency components to suppress noise and artifacts. Extensive experiments on the nuScenes benchmark show that HD-Fusion achieves 70.5% mAP and 72.9% NDS, improving over the baseline by 12.1 and 6.6 points, respectively. Additional evaluations on rainy/night subsets, simulated camera/LiDAR failures, and cross-dataset transfer from nuScenes to Lyft further demonstrate that HD-Fusion maintains superior performance on small and distant objects and exhibits strong robustness and generalization in challenging autonomous-driving scenarios.