科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ IEEE Transactions on Geoscience and Remote Sensing2026-01-01· Remote sensing

Diffusion Feature Completion-Driven Textual Dual Distillation for Multimodal Remote Sensing Classification With Missing Modality

Dan Xu, Wenqian Dong, Song Xiao, Jiahui Qu, Y F Li

原始摘要(英文原文)· Original abstract
In recent years, research on joint classification of multi-modal remote sensing data has achieved remarkable progress. However, constrained by imaging conditions or sensor issues, missing modality frequently occurs in practice. Knowledge distillation, as one of the mainstream strategies to deal with missing modality, is widely used due to its excellent transfer ability. However, most existing knowledge distillation methods focus on directly transferring available modal information through feature alignment or predictive constraints, without explicitly modeling or refining the latent semantic structure of missing modality. Furthermore, the heterogeneity between different modalities makes effective alignment difficult through direct transfer, thus limiting the utilization of cross-modal complementary information and consequently hindering model performance improvement in remote sensing scenarios with missing modality. Therefore, in this paper, we propose a diffusion feature completion-driven textual dual distillation model (DFC-TD2) for multi-modal remote sensing image classification with missing modality. This model explicitly models the missing modality, introduces textual semantic information to guide cross-modal feature transfer, and achieves multi-constraint collaborative optimization under the supervision of classification labels, thereby effectively improving the overall performance of multi-modal classification tasks with missing modality. Specifically, a shared-constrained diffusion feature completion network (SCFC) is designed, which uses shared features extracted and frozen from multi-modal data as constraints to guide the diffusion model to effectively model and complete missing modal features, thereby generating semantically consistent and more discriminative modal representations. Furthermore, a text-guided dual distillation network (TGD2) is designed to guide cross-modal feature alignment using text semantics and combine it with classification label supervision to construct a feature-label dual constraint, which effectively improves feature discriminativeness and supervision efficiency, thereby mitigating cross-modal representation heterogeneity and improving the stability and effectiveness of knowledge transfer. Experimental results on three public datasets demonstrate the effectiveness of this method in dealing with missing modality.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Diffusion Feature Completion-Driven Textual Dual Distillation for Multimodal Remote Sensing Classification With Missing Modality — 科研速览 Science Skim