Xiandai Cui, Li Zhang
Precise analysis of multisource remote sensing imagery is essential for producing accurate land-cover classifications. Although attention-based architectures excel at modeling long-range dependencies, their computational complexity scales quadratically with input size, posing challenges for fusing high-resolution hyperspectral, light detection and ranging (LiDAR), and synthetic aperture radar (SAR) data. Although the Mamba architecture enables linear-time sequence modeling, its inherent reliance on unidirectional scanning limits its capacity to fully capture and integrate bidirectional spatial contextual information. To overcome this limitation, we introduce a architecture based on Mamba, enhanced with attention mechanisms and U-Net-inspired design elements for multimodal remote sensing image classification. Specifically, we introduce two key modules: (1) a dual-scanning Mamba (DSM) module that models bidirectional context (forward-backward) with linear complexity, eliminating unidirectional bias while enabling efficient long-range dependency capture, and (2) a crossing-attention module that jointly attends to horizontal and vertical spatial orientations, extending receptive fields and enabling adaptive feature refinement crucial for complex boundaries. The proposed architecture, named CADSM (crossing-attention dual-scanning Mamba), was evaluated on the Muufl, Houston University, and Augsburg datasets, achieving state-of-the-art (SOTA) accuracies of 96.10%, 99.84%, and 97.24%, respectively. In addition, tests on the Indian Pines hyperspectral dataset and the San Francisco PolSAR (polarimetric synthetic aperture radar) dataset demonstrated that CADSM significantly outperforms baseline models, highlighting its remarkable generalization capability. Codes are available at https://github.com/cuixiandai/CADSM.