科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Advances in ophthalmology practice and research2026-01-01

Evaluating large language model clinical reasoning in glaucoma using retrieval-augmented generation.

Houfa Yin, Qi Miao, Wanshu Zhou, Chenyang Hu, Yongwei Guo, Lixia Shen, Haiyan Cai, Andrzej Grzybowski, Kai Jin

一句话结论 · In one sentence

Retrieval augmentation was associated with more accurate, more complete, and more safety-aware glaucoma reasoning under this scenario-based evaluation. These findings support the potential value of RAG-enhanced LLMs as supervised clinical decision-support tools, but they do not establish standalone clinical use. Safety was evaluated as a separate audit layer rather than as part of the weighted composite score.

原始摘要(英文原文)· Original abstract
BACKGROUND: Large language models (LLMs) demonstrate strong performance in knowledge-based medical tasks, yet their clinical reasoning capabilities in complex ophthalmic decision-making, particularly glaucoma management, remain insufficiently characterized. Retrieval-augmented generation (RAG) has been proposed as a strategy to improve factual grounding and safety, but its clinical value requires systematic evaluation. METHODS: We conducted a retrospective, scenario-based comparative evaluation using 40 real-world glaucoma cases spanning primary, secondary, postoperative, and end-stage disease. Six model conditions (GPT, Gemini, and Grok, each with and without RAG) were assessed using a guideline-grounded framework covering medical accuracy, key point coverage, logical completeness, and a separate qualitative safety audit. Model outputs were compared with written responses from four practicing ophthalmologists. All responses were independently scored by two masked glaucoma specialists using a prespecified ordinal rubric, and formal inter-rater reliability and paired sensitivity analyses were performed. RESULTS: RAG-enhanced models consistently outperformed their matched non-RAG counterparts across evaluation domains. In human-rating sensitivity analyses, the composite RAG advantage remained significant for GPT (mean difference=0.119, 95% CI: 0.068-0.173), Gemini (mean difference=0.095, 95% CI: 0.030-0.161), and Grok (mean difference=0.110, 95% CI: 0.033-0.194). Inter-rater agreement of the ordinal rubric was limited, supporting consensus adjudication and cautious interpretation of artificial intelligence-human comparisons. CONCLUSIONS: Retrieval augmentation was associated with more accurate, more complete, and more safety-aware glaucoma reasoning under this scenario-based evaluation. These findings support the potential value of RAG-enhanced LLMs as supervised clinical decision-support tools, but they do not establish standalone clinical use. Safety was evaluated as a separate audit layer rather than as part of the weighted composite score.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Evaluating large language model clinical reasoning in glaucoma using retrieval-augmented generation. — 科研速览 Science Skim