科研速览 · Science Skim继续刷下去 · Keep skimming →
◇ medRxiv2026-09-11· radiology and imaging

Accuracy Overstates Evidence Grounding and Abstention Reliability in Mammography Vision-Language Models

B. Qu, w. Liu, M. Murrow, M. Burger, X. Guo, M. S. Vaidya, S. L. Rose, M. Kantarcioglu, B. A. Malin, Z. Yin

原始摘要(英文原文)· Original abstract
Answer accuracy alone cannot determine whether a Vision-Language Model (VLM) relies on clinically relevant mammographic evidence or recognizes when that evidence is unavailable. We introduce an evidence-grounded selective evaluation benchmark that evaluates Pathology classification and Abnormality identification together with label-aware lesion localization. Each original image--task instance is paired with a lesion-removal counterfactual and, when feasible, a size-matched non-lesion random-removal control. Lesion evidence is removed by local tissue reconstruction, while the random-removal view provides a task-irrelevant regional perturbation baseline of comparable spatial extent. We evaluate 16 general-purpose and medically specialized VLMs, including open-source and proprietary models, under a common protocol with the same prompt, image conditions, label spaces, and structured-output requirements. The results show substantial gaps between classification performance and evidence-grounded reliability. InternVL3.5-8B achieved the highest Pathology Macro-F1 at 53.75\%, Huatuo Vision-34B achieved the highest Abnormality Exact Match at 52.44\%, and GPT-5.6 Sol achieved the highest Grounded Answer Success at 21.99\%. Lingshu-7B achieved the highest Counterfactual Specificity at 35.19\%, whereas GPT-5.6 Sol achieved the highest Joint Reliability at 6.47\%. These findings suggest that answer accuracy and abstention behavior can substantially overstate the reliability of current VLMs when predictions are not verified against the visual evidence on which they should depend.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Accuracy Overstates Evidence Grounding and Abstention Reliability in Mammography Vision-Language Models — 科研速览 Science Skim