Wenjun Xie, Kejun Chen, Dong Wang, Xiaoping Liu
Recently, Mamba has gained widespread attention due to its ability to model long-range dependencies with linear computational complexity. To explore the application of Mamba in 2D human pose estimation, we propose MatPose, a Mamba-Transformer hybrid model specifically designed for efficient 2D human pose estimation. The model aims to combine Mamba’s efficient modeling of long-range dependencies with the powerful global context modeling capabilities of the Transformer to effectively extract human pose keypoints. Initially, to address the lack of local features when Mamba is applied to computer vision tasks, we design a Cross-Stage Multi-Scale Convolution (CSMSC) module by integrating multi-scale convolution, cross-stage feature fusion, and spatial attention mechanisms to effectively extract local features. Then, to mitigate the long-range forgetting issue inherent in Mamba, we shorten the sequence length using the Conv-Reduce operation. In addition, we design a Channel Selection Attention (CSA) mechanism to compensate for the feature loss caused by the Conv-Reduce operation. Finally, to explore a suitable integration method for the Mamba-Transformer hybrid model in 2D human pose estimation, we conduct a comprehensive ablation study on the feasibility of integrating Mamba and Transformer models. Experimental results show that the proposed method, compared to the baseline model, improves performance while reducing computational overhead. On the COCO val2017 dataset, MatPose achieves an AP of 74.6 with only 5.18 GFLOPs, outperforming most existing human pose estimation models.