Shunxing Chen, Xiangyang Miao, Yunmei Zhao, Bin Yang, Guobao Xiao
In this paper, we propose a Spatial-Motional Transformer Network (SMFormer) for two-view correspondence pruning. Compared to most existing methods that mainly focus on spatial coordinates for correspondence learning, SMFormer jointly considers correspondence reliability and motion consistency by integrating feature screening and sparse motion representations. Specifically, we develop a robust feature learning framework that combines feature screening and spatial-motional interaction to enhance correspondence representation under heavy outlier contamination. We present a Feature Screening Transformer Block (FSTB) to selectively suppress low-confidence features and enhance feature reliability under heavy outlier contamination. In addition, we design a Spatial-Motional Interaction Module (SMIM) to perform cross-stream interaction and fusion between correspondence features and motion features. This module employs graph-based cross-attention aggregation over sparse spatial and motion graphs to further improve contextual feature representation. The proposed FSTB and SMIM collaboratively improve correspondence pruning by enhancing feature reliability and exploiting complementary spatial-motional cues. In consequence, our proposed SMFormer outperforms state-of-the-art methods by a significant margin, achieving a remarkable 9.75% improvement under an error threshold of 5∘ over the second-best method for the relative pose estimation task on an unknown outdoor dataset. Our source code will be available at: https://github.com/CSX777/SMFormer.