科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ ACM Transactions on Multimedia Computing Communications and Applications2026-08-04· Computer science

Multimodal Infusion Tuning for Large Models

Hao Sun, Yu Song, Shiyu Teng, Xinyao Yu, Ziwei Niu, Yen‐Wei Chen

原始摘要(英文原文)· Original abstract
Recent advancements in large-scale models have showcased remarkable generalization capabilities in various tasks. However, integrating multimodal processing into these models presents a significant challenge, as it often comes with a high computational burden. To address this challenge, we introduce a new parameter-efficient multimodal tuning strategy for large models in this paper, referred to as Multimodal Infusion Tuning (MiT). MiT leverages decoupled self-attention mechanisms within large language models to effectively integrate information from diverse modalities such as images and acoustics. In MiT, we also design a novel adaptive rescaling strategy at the attention head level, which optimizes the representation of infused multimodal features. Notably, all foundation models are kept frozen during the tuning process to reduce the computational burden and only 2.5% parameters are tunable. We conduct experiments across a range of multimodal tasks, including image-related tasks like referring segmentation and non-image tasks such as sentiment analysis. Our results showcase that MiT achieves highly competitive performance among parameter-efficient multimodal tuning methods, while significantly reducing computational overhead (to 10% of previous adapter-based methods). Moreover, our tuned model exhibits robust reasoning abilities even in complex scenarios.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Multimodal Infusion Tuning for Large Models — 科研速览 Science Skim