Jie He, Jiahui Qu, Xiaohan Zhang, Wenqian Dong, Yunsong Li
Fusion of hyperspectral and multispectral images has become a mainstream technique for obtaining high spatial resolution hyperspectral images. However, in practical applications, it is challenging to obtain spatially well-matched pairs of low-resolution hyperspectral images (LRHSI) and high-resolution multispectral images (HRMSI). To tackle this challenge, we propose a Matching-Fusion Mutual Enhancement Diffusion Network (MFME-DiffNet), which leverages the unpaired multi-scale matched images to facilitate the generation of high-quality HRHSI through iterative matching-fusion mechanism. Specifically, MFME-DiffNet comprises two core modules: a Gradient-Aligned Multi-Scale Matching Module (GAMM) and a Similarity-Weighted Attention Fusion Module (SWAF). GAMM is designed to generate multi-scale guidance images using dynamic sliding windows and to perform gradient-based feature alignment, resulting in high-quality auxiliary images that are consistent with the target HSI. SWAF is proposed to adaptively integrate multi-scale matched images based on their texture similarity with the HSI through a similarity-weighted fusion strategy, and reweight the fused features using the overall similarity of the matched images, thereby significantly enhancing the texture details of HSI. GAMM and SWAF are iteratively executed during the reverse process of the diffusion model, forming a dynamic feedback mechanism for matching and fusion. Extensive experimental results demonstrate that the proposed MFME-DiffNet can effectively achieve the HSI-MSI fusion with incomplete spatial correspondence, and is significantly superior to existing advanced methods in terms of visual effects and quantitative evaluation indicators. The code is available at https://github.com/Jiahuiqu/MFME-DiffNet.