Min Zhao, Jiajun Cai, Man Zhou, Bo Huang
Optical remote sensing images are often degraded by cloud contamination, leading to significant information loss. Moreover, due to the inherent trade-off between temporal and spatial resolutions, it is challenging to acquire images with both high temporal frequency and fine spatial detail. Existing works have achieved strong performance in cloud removal and spatiotemporal fusion (STF) when treating them as separate tasks. However, processing them independently can introduce cumulative errors that degrade the reliability of downstream applications. To address these challenges, we propose SuperSTF, an all-in-one framework that simultaneously reconstructs cloud-free, fine resolution image series over time from coarse resolution inputs and cloud contaminated fine resolution observations. By jointly modeling these tasks, SuperSTF adaptively exploits their intrinsic correlations, thereby enhancing both cloud removal and STF performance. Specifically, we design an efficient latent diffusion model for image generation, where a Swin Transformer-based network serves as the pixel space autoencoder. Cross-attention modules are incorporated into the diffusion network to facilitate multi-source feature fusion. Furthermore, cloud location encoding and acquisition date modulation are integrated into the framework to further improve reconstruction quality. Experiment results demonstrate the superiority of our proposed method in fusing multi-temporal and multi-source data to generate image series with fine details.