科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Frontiers in medicine2026-01-01

Comparative performance of contemporary multimodal large language models in retinal imaging question answering.

Xiaochen Gu, Yizhou Yang, Xuanqiao Lin, Xu Chen

一句话结论 · In one sentence

Contemporary multimodal LLMs showed promising but uneven performance on OCTCases retinal questions. Their performance was consistently better on nonimage-based than image-based questions, indicating that retinal image interpretation remains a major limitation. Voting-based group answers may provide a useful reliability signal when model agreement is present, whereas model disagreement may help identify difficult or visually ambiguous questions requiring specialist review. These findings support continued benchmarking and cautious, specialist-supervised use of multimodal LLMs in retinal education and structured imaging question answering.

原始摘要(英文原文)· Original abstract
BACKGROUND: Contemporary multimodal large language models (LLMs) can process both clinical text and medical images, but their reliability in retinal imaging-based question answering remains uncertain. This study evaluated the performance of six contemporary multimodal LLMs on retinal multiple-choice questions derived from OCTCases. METHODS: This cross-sectional benchmark study included 226 multiple-choice questions from 78 OCTCases retinal cases, comprising 151 image-based and 75 nonimage-based questions. Six multimodal LLMs were evaluated: ChatGPT-5.5 Instant, ChatGPT-5.5 Thinking, Grok 4, Gemini 3, DeepSeek V4, and Kimi 2.5. Model-selected answers were compared with the OCTCases answer key, which was reviewed by three retina specialists. Accuracy was assessed overall and by question subset. Pairwise model comparisons, inter-model agreement, voting-based group-answer performance, item difficulty distribution, and interface latency were analyzed. RESULTS: Overall accuracy ranged from 59.3 to 77.4%, with the highest accuracy observed for ChatGPT-5.5 Thinking, followed by Gemini 3 and ChatGPT-5.5 Instant. DeepSeek V4 showed the lowest overall accuracy. All models performed better on nonimage-based questions than on image-based questions. In the image-based subset, Gemini 3 achieved the highest accuracy, whereas DeepSeek V4 showed the lowest accuracy. Under the majority-vote rule, group answers were generated for 182 of 226 questions and achieved an accuracy of 89.6% among questions with a determinate group answer. The plurality-vote rule generated group answers for more questions but with lower accuracy. Inter-model agreement was generally higher for nonimage-based than image-based questions, and all questions answered incorrectly by all six models were image-based. Interface latency varied across models and showed no consistent association with response correctness. CONCLUSION: Contemporary multimodal LLMs showed promising but uneven performance on OCTCases retinal questions. Their performance was consistently better on nonimage-based than image-based questions, indicating that retinal image interpretation remains a major limitation. Voting-based group answers may provide a useful reliability signal when model agreement is present, whereas model disagreement may help identify difficult or visually ambiguous questions requiring specialist review. These findings support continued benchmarking and cautious, specialist-supervised use of multimodal LLMs in retinal education and structured imaging question answering.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Comparative performance of contemporary multimodal large language models in retinal imaging question answering. — 科研速览 Science Skim