科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Graefe's archive for clinical and experimental ophthalmology = Albrecht von Graefes Archiv fur klinische und experimentelle Ophthalmologie2026-08-06

Comparative evaluation of large language models and clinicians in real-world glaucoma clinical reasoning.

Houfa Yin, Lixia Shen, Haiyan Cai, Wei Wu, Andrzej Grzybowski, Kai Jin

一句话结论 · In one sentence

In this limited 34-case evaluation, large language model-based AI systems produced structured glaucoma-related reasoning with performance that overlapped with attending ophthalmologists but did not establish clinical equivalence. These systems require specialist oversight and further validation before clinical use, but may have potential as supervised decision-support and educational tools.

原始摘要(英文原文)· Original abstract
PURPOSE: Clinical decision-making in glaucoma is complex and requires integration of heterogeneous information, including patient history, examination findings, and risk stratification. While artificial intelligence (AI) has shown strong performance in image-based ophthalmic tasks, its capability in specialty-specific clinical reasoning remains insufficiently explored. METHODS: Performance was evaluated by glaucoma specialists using a predefined rubric across three clinically oriented domains: medical accuracy (40%), key-point recall (30%), and logical completeness (30%). The weighted composite score was used as a descriptive summary of case-based reasoning quality. RESULTS: AI models showed structured clinical reasoning performance in this case-based dataset, with weighted mean scores overlapping with those of attending ophthalmologists and exceeding those of some lower-performing trainees. These findings should be interpreted as exploratory performance patterns rather than evidence of equivalence. Inter-individual variability was substantial among human clinicians, particularly residents. AI systems often included safety-critical diagnostic and management elements, while the best-performing human clinician achieved the highest individual score overall. CONCLUSION: In this limited 34-case evaluation, large language model-based AI systems produced structured glaucoma-related reasoning with performance that overlapped with attending ophthalmologists but did not establish clinical equivalence. These systems require specialist oversight and further validation before clinical use, but may have potential as supervised decision-support and educational tools.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Comparative evaluation of large language models and clinicians in real-world glaucoma clinical reasoning. — 科研速览 Science Skim