科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ IEEE Access2026-01-01· Interpretability

Multimodal Vision–Language Models in Medical Imaging: A Survey of Retrieval, Interpretability, and Trust

Muhammad Imran, Yugyung Lee

原始摘要(英文原文)· Original abstract
Medical diagnosis requires synthesizing radiological images, laboratory results, clinical notes, and patient history—capabilities increasingly supported by multimodal vision–language models (VLMs). Yet despite rapid advances, no general-purpose medical VLM has achieved routine clinical deployment. Existing surveys typically examine retrieval-augmented generation (RAG), interpretability, and trustworthiness in isolation, leaving unclear how these components jointly determine real-world clinical readiness. This survey analyzes 183 studies and evaluates three pillars essential for deployment: (1) factual accuracy enabled by RAG, which grounds model outputs in clinically verified evidence; (2) interpretability aligned with clinician reasoning, ensuring explanations that support diagnostic decision-making; and (3) trustworthy operation encompassing safety, fairness, uncertainty handling, workflow integration, and regulatory alignment. We show that benchmark gains alone are insufficient, as current VLMs still face persistent challenges in multimodal alignment, temporal reasoning, retrieval latency and scalability, explanation fidelity, and regulatory and cost-effectiveness constraints. Our contributions include a deployment-oriented taxonomy of medical VLM architectures; a comparative analysis of RAG designs and their clinical implications; an evaluation of interpretability methods grounded in clinical workflows; and an integrated framework for trust, uncertainty reporting, and regulatory pathways. We also curate models, datasets, and benchmarks that support reproducible evaluation and outline a roadmap toward causal, temporally aware, and computationally efficient VLMs suitable for real-world clinical use. To the best of available literature, no prior survey has jointly analyzed retrieval, interpretability, and trust within a unified framework for evaluating the clinical viability of medical VLMs.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Multimodal Vision–Language Models in Medical Imaging: A Survey of Retrieval, Interpretability, and Trust — 科研速览 Science Skim