科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ IEEE transactions on pattern analysis and machine intelligence2026-09-22

Alignment-Consistent Multimodal Learning under Uncertain Correspondence.

Qian Li, Yaheng Wang, Qian Huang, Cheng Ji, Shangguang Wang

原始摘要(英文原文)· Original abstract
Multimodal learning aims to integrate heterogeneous observations such as images, text, depth, and radar to improve perception and reasoning. However, most existing multimodal models implicitly assume that cross-modal observations are well aligned, an assumption that rarely holds in real-world scenarios due to viewpoint variation, sensor noise, temporal asynchrony, and incomplete observations. Such misalignment introduces substantial uncertainty in cross-modal correspondence, often leading to unreliable feature matching and unstable multimodal representations. In this paper, we revisit multimodal representation learning from the perspective of uncertain cross modal correspondence and introduce the principle of alignment consistency, which states that tokens describing the same semantic entity across modalities should preserve consistent high-level representations even when their low-level correspondence is ambiguous. Based on this principle, we propose an Alignment-Consistent Multimodal Learning (ACML) framework that explicitly models probabilistic token correspondence and enforces alignment-consistent representations through uncertainty-aware matching and globally coherent alignment. Specifically, ACML integrates probabilistic correspondence estimation with uncertainty-aware optimal transport to capture ambiguous cross-modal relations while suppressing unreliable matches, and learns modality-invariant representations via alignment-consistent representation learning. Extensive experiments on diverse multimodal bench-marks, including vision-language understanding, RGB-depth scene understanding, and multimodal remote sensing analysis, demonstrate that ACML consistently improves performance and robustness under cross-modal misalignment, missing modalities, and noisy sensing conditions. The evaluation includes matched-mechanism controls, five-seed statistics, direct SAR-optical tie-point registration, remove-one ablations, calibration diagnostics, sensitivity sweeps, and explicit failure cases. The resulting claim is deliberately bounded: ACML improves uncertainty-aware correspondence and representation stability when the modalities retain appreciable shared semantic support. These results highlight that explicitly modeling correspondence uncertainty provides a principled foundation for robust multimodal representation learning.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Alignment-Consistent Multimodal Learning under Uncertain Correspondence. — 科研速览 Science Skim