Xintong Cao, Wenqian Dong, Jiahui Qu, Yunsong Li
Global surface changes are increasingly monitored using multi-temporal remote sensing technologies. As an emerging technology, change captioning can organically integrate the location information of changed regions with the semantic analysis of regional attributes to generate natural language descriptions of changes, providing critical support for the intuitive interpretation of remote sensing monitoring results. However, existing methods have two key limitations: first, single-stream CNNs fail to extract spatiotemporal information sufficiently which leads to the easy omission of subtle changes, while single-stream Transformers suffer from relatively high parameter counts; second, using pixel-level difference information directly causes the model to focus on pseudo-changes induced by illumination or noise, reducing the accuracy of real change characterization and subsequent description. To address these issues, this paper proposes RMNet which is a dual-dimensional difference recalibration-guided CNN-VMamba synergistic network, with two core innovations: 1) a dual-stream architecture adopted that combines the strong local feature extraction capability of CNNs with the powerful global context modeling ability of VMamba, further enhanced by a channel-wise spatial window attention mechanism; 2) a dual-dimensional difference recalibration mechanism that optimizes features by highlighting core changes and suppressing pseudo-changes through dimension-specific enhancement strategies. Extensive experiments on the LEVIR-CC dataset demonstrate significant performance improvements across all evaluation metrics, validating the effectiveness of RMNet in overcoming current limitations in remote sensing change captioning tasks. The code is available at https://github.com/Jiahuiqu/RMNet.