科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ IEEE Transactions on Affective Computing2026-01-01· Modalities

TEMPO: Training-time Equilibration of Modalities for Per-sample Optimization in Multimodal Sentiment

Yi Zhao, Erik Cambria, E Xiaosong, Xianxun Zhu

原始摘要(英文原文)· Original abstract
Multimodal sentiment models often become over-reliant on the “easiest” modality (typically text), leading to three coupled sub-problems: (i)representation-level dominance, where weaker modalities contribute little to the fused representation; (ii)optimization-level dominance, where the strongest modality drives most gradient updates and suppresses learning in others; and (iii)robustness degradation, where audio or vision fail under noise or missing inputs at test time. We present TEMPO, a plug-and-play training framework that mitigates these issues by rebalancing learning pressure across modalities while leaving inference unchanged. For each mini-batch, TEMPO estimates relative modality strength and applies two synchronized, training-only controls: selective forward attenuation and backward gradient equilibration. On IEMOCAP and MELD, TEMPO improves accuracy and weighted F1 over strong multimodal baselines, increases the standalone usefulness of weaker modalities, and offers higher robustness under missing- or corrupted-modality stress. Across three benchmarks, TEMPO improves accuracy by 3.2–6.7% and reduces calibration error by 18–34%, demonstrating consistent gains in both performance and reliability with negligible computational overhead.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

TEMPO: Training-time Equilibration of Modalities for Per-sample Optimization in Multimodal Sentiment — 科研速览 Science Skim