Xiang Li, Min Zhang, Yuanhao Jin, Lixiang Xu, Bowen Wang, Jing Yang
Remote Sensing Image Super-Resolution (RSISR) is a core task in geospatial image analysis. Convolutional neural networks (CNNs) have achieved significant breakthroughs in RSISR tasks by extracting local features. However, CNN-based methods struggle to capture long-range dependencies, thereby limiting SR performance. Recently, Transformer-based methods have demonstrated remarkable performance in capturing global information. Nevertheless, they remain inadequate for exploring high-frequency details and local features. To overcome these limitations, this work introduces a novel dual-path collaborative architecture, named DCTNet, which combines Transformer-based global modeling with convolution-driven local feature extraction. DCTNet is a hybrid network composed of a CNN-Transformer Residual Hybrid Group (CTHG). This group consists of two core components: the Dual-domain Fusion Window Attention Block (DFWAB) and the Stepwise Dilated Convolution (SDC). Specifically, the DFWAB incorporates channel and frequency attention mechanisms following the standard Transformer block to recover high-frequency details. Furthermore, by integrating stepwise dilated convolutions into the conventional Transformer architecture, the CTHG effectively captures both multi-scale local and global features. Additionally, we employ dense connections among the DFWAB modules to facilitate feature reuse across layers. Experimental results on the AID and UCMerced datasets demonstrate that DCTNet achieves competitive reconstruction performance across different scale factors, with statistically significant improvements observed in specific settings.