Zhihao Xi, Yu Meng, Yupeng Deng, Yuman Feng, Diyou Liu, Jingbo Chen
In conventional Unsupervised domain adaptation (UDA), model knowledge is transferred to the target domain with access to annotated source data, which is an unsuitable strategy for cross-domain scenarios involving data privacy and confidentiality. In this paper, we focus on source-free domain adaptation (SFDA) for semantic segmentation tasks, which adapts source-trained models to unannotated target domains without relying on source data. Self-training paradigms dominate the existing approaches, but often suffer from significant performance degradation due to unreliable pseudolabels and knowledge forgetting. To address these challenges, we propose innovative self-supervised spatial–temporal consistency learning for source-free domain adaptive segmentation (S3T-SFDA). Specifically, to address pseudolabel ambiguity in various spatial contexts, a spatial multiview consistency (SMVC) mechanism is proposed to constrain the semantic consistency across different spatial views of the same image. To mitigate the knowledge forgetting problem caused by a lack of supervisory information, a temporal dynamic consistency (TDC) mechanism, which harnesses historical knowledge consistency to regularize the current model evolution direction, is proposed. Furthermore, to improve the category-discriminative representation capabilities under domain shifts, a spatial-temporal contrastive (STC) strategy is designed to promote the intrinsic semantic association of features belonging to different categories in the target domain. Extensive experiments conducted on two domain-adaptive remote sensing (RS) segmentation benchmarks. The results demonstrate the flexibility of the proposed method when integrated into various advanced segmentation architectures, as well as its excellent generalization performance across different cross-domain RS scenarios.