Nan Xiao, Xiyou Fu, Qi Ren, Wangquan He, Siqi Wei, Sen Jia
Existing transformer-based hyperspectral image (HSI) and multispectral image (MSI) fusion methods often suffer from limited cross-modal interaction and insufficient capability to model heterogeneous spatial and spectral patterns across different regions, which restricts their ability to fully exploit modality complementarity. To address these issues, we propose a Region-Aware Mixture of Experts network for HSI and MSI fusion. The proposed framework consists of three key components. First, an attention sharing module (ASM) is designed to explicitly propagate attention information between modalities, enabling more balanced interaction between spatial and spectral representations. Second, the region-aware MoE (RAMoE) is introduced, where multiple region-aware experts (RAEs) capture region-specific feature variations, while an invariant-weight expert (IWE) provides global contextual guidance, improving robustness and generalization. Third, a spatial and spectral reconstruction module (SSRM) is employed to jointly recover spatial details and spectral information through residual learning. Extensive experiments on benchmark datasets demonstrate that the proposed method consistently outperforms state-of-the-art fusion methods in terms of both quantitative metrics and visual quality. Furthermore, experiments on real-world datasets verify the practical effectiveness of RAMoE in recovering fine spatial details while preserving spectral fidelity. The code will be released at https://github.com/XXn25/RAMoE.