Xiaochen Gu, Yizhou Yang, Xuanqiao Lin, Xu Chen
Contemporary multimodal LLMs showed promising but uneven performance on OCTCases retinal questions. Their performance was consistently better on nonimage-based than image-based questions, indicating that retinal image interpretation remains a major limitation. Voting-based group answers may provide a useful reliability signal when model agreement is present, whereas model disagreement may help identify difficult or visually ambiguous questions requiring specialist review. These findings support continued benchmarking and cautious, specialist-supervised use of multimodal LLMs in retinal education and structured imaging question answering.
BACKGROUND: Contemporary multimodal large language models (LLMs) can process both clinical text and medical images, but their reliability in retinal imaging-based question answering remains uncertain. This study evaluated the performance of six contemporary multimodal LLMs on retinal multiple-choice questions derived from OCTCases.
METHODS: This cross-sectional benchmark study included 226 multiple-choice questions from 78 OCTCases retinal cases, comprising 151 image-based and 75 nonimage-based questions. Six multimodal LLMs were evaluated: ChatGPT-5.5 Instant, ChatGPT-5.5 Thinking, Grok 4, Gemini 3, DeepSeek V4, and Kimi 2.5. Model-selected answers were compared with the OCTCases answer key, which was reviewed by three retina specialists. Accuracy was assessed overall and by question subset. Pairwise model comparisons, inter-model agreement, voting-based group-answer performance, item difficulty distribution, and interface latency were analyzed.
RESULTS: Overall accuracy ranged from 59.3 to 77.4%, with the highest accuracy observed for ChatGPT-5.5 Thinking, followed by Gemini 3 and ChatGPT-5.5 Instant. DeepSeek V4 showed the lowest overall accuracy. All models performed better on nonimage-based questions than on image-based questions. In the image-based subset, Gemini 3 achieved the highest accuracy, whereas DeepSeek V4 showed the lowest accuracy. Under the majority-vote rule, group answers were generated for 182 of 226 questions and achieved an accuracy of 89.6% among questions with a determinate group answer. The plurality-vote rule generated group answers for more questions but with lower accuracy. Inter-model agreement was generally higher for nonimage-based than image-based questions, and all questions answered incorrectly by all six models were image-based. Interface latency varied across models and showed no consistent association with response correctness.
CONCLUSION: Contemporary multimodal LLMs showed promising but uneven performance on OCTCases retinal questions. Their performance was consistently better on nonimage-based than image-based questions, indicating that retinal image interpretation remains a major limitation. Voting-based group answers may provide a useful reliability signal when model agreement is present, whereas model disagreement may help identify difficult or visually ambiguous questions requiring specialist review. These findings support continued benchmarking and cautious, specialist-supervised use of multimodal LLMs in retinal education and structured imaging question answering.