Changjun Gu, Like Wang, Jiaxu Leng, Xinbo Gao
Visual-Inertial Odometry (VIO) is a fundamental and challenging task in computer vision and robotics for estimating trajectory. Learning-based VIO methods have gained significant attention for enhancing robustness. However, existing learning-based methods rarely focus on effectively exploring their complementary benefits and selecting the most representative features to improve accuracy. To address these limitations, we propose a novel Multi-modal Interaction and Selection (MIS) approach for visual-inertial localization that enhances robustness while improving localization accuracy. Specifically, we first introduce a global-local cross-attention module that enables effective knowledge exchange between visual and inertial feature representations, enhancing feature representation while preserving their respective strengths. Then, we present a dual-path dynamic fusion module that adaptively identifies and selects the most informative features in complex environments, seamlessly fusing information from visual and inertial modalities. Finally, we evaluate the localization accuracy using benchmark datasets, such as the KITTI Odometry dataset and others. Extensive experimental results show that the proposed MIS-VIO method achieves superior performance compared to existing approaches.