Yawen Bai, Weisheng Li, Yidong Peng, Yusha Liu
Existing remote sensing spatiotemporal fusion methods often suffer from excessive smoothing at object boundaries, insufficient modeling of global contextual information, and underutilization of temporal change features. To improve these issues, a triple-branch convolutional neural network is proposed, consisting of a dual-stream spatial network and a temporal difference network.Within the spatial branches, a spatiotemporal adaptive modulation (STAM) module is integrated, in which spatial attention is combined with squeeze-and-excitation (SE) channel attention to enhance feature discrimination. In the temporal branch, lightweight depthwise separable convolutions are employed to efficiently refine temporal difference features. Furthermore, a composite loss function is designed to simultaneously optimize pixel-level accuracy, structural fidelity, and edge detail restoration. Experiments on the CIA, LGC, and AHB datasets demonstrate that the proposed method achieves marked improvements over current models in reconstructing high-frequency edges and low-frequency structures, maintaining global semantic consistency, and suppressing dynamic change noise, with an average increase of 1.46 dB in PSNR and an average decrease of 0.15 in ERGAS.