科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Computation2026-05-29· Computer science

A Comprehensive Survey and Guide to Multimodal Large Language Models in Vision–Language Tasks

Chia Xin Liang, Pu Tian, Caitlyn Heqi Yin, Yao Yua, Wei An-Hou, Li Ming, Xinyuan Song, Tianyang Wang, Ziqian Bi, Ming Liu, Riyang Bao, Pengbin Feng

原始摘要(英文原文)· Original abstract
This survey provides a comprehensive guide to Multimodal Large Language Models (MLLMs) with a focus on vision–language tasks, including image captioning, visual question answering, cross-modal retrieval, visual grounding, multi-image reasoning, long-video understanding, and embodied AI. We examine architectures, training pipelines, and practical applications, covering visual encoders, language model backbones, connector modules, contrastive pre-training, instruction tuning, and preference alignment. We also foreground first-principles constraints—information bottlenecks, data-processing limits, and statistical co-occurrence bias—that shape architecture, robustness, and evaluation. This survey centers on vision–language systems and does not cover audio-only models or code-generation tools without visual inputs. Through task-level analysis and system-level case studies, we examine prominent MLLM implementations while addressing key challenges in scalability, memory, energy use, inference cost, robustness, and cross-modal learning. We present a unified taxonomy of the MLLM design space, a comparative overview of representative models and evaluation benchmarks, and a discussion of open problems. Concluding with ethical considerations and responsible AI development, this survey offers theoretical frameworks and practical insights for researchers, practitioners, and students working at the intersection of natural language processing and computer vision.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

A Comprehensive Survey and Guide to Multimodal Large Language Models in Vision–Language Tasks — 科研速览 Science Skim