科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Seminars in ophthalmology2026-08-29

Evaluation of Large Language Model-Generated Recommendations in Glaucoma Surgical Decision-Making.

Müge Toprak, Büşra Yılmaz Tuğan, Nurşen Yüksel

一句话结论 · In one sentence

LLM-based chatbots can provide acceptable surgical guidance for straightforward primary surgical cases, but their utility is limited in high-risk or complex clinical settings. The observed deficiencies in rationale and increased safety risks in complex cases suggest that LLMs should be integrated as auxiliary decision-support tools under expert supervision rather than used as autonomous decision-makers in glaucoma surgery planning.

原始摘要(英文原文)· Original abstract
OBJECTIVE: To evaluate the utility, rationality, and safety of glaucoma surgery recommendations generated by three prominent large language models (LLMs) - ChatGPT, Microsoft Copilot, and Google Gemini - when applied to real-world clinical scenarios. METHODS: Retrospective records from a tertiary hospital were converted into standardized scenarios and stratified into "primary" and "complex" glaucoma groups. Each LLM was prompted to suggest a single surgical approach and provide a rationale. A blinded team of glaucoma specialists evaluated the outputs based on six criteria: appropriateness, rationale quality, specificity, adherence to guidelines, feasibility, and safety risk, using a normalized 0-100 scale. RESULTS: Median overall quality scores across all cases were 80.7 for ChatGPT, 80.7 for Copilot, and 82.7 for Gemini, showing no statistically significant difference in general performance (p = .367). However, case complexity significantly affected performance. For Gemini, appropriateness and rationale quality scores dropped significantly in complex cases and were accompanied by a statistically significant increase in safety risk (p = .009). Although ChatGPT and Copilot demonstrated more stability across groups, their rationale quality was significantly lower in complex scenarios than in primary ones (p = .014 and <0.001, respectively). Pairwise analyses revealed that ChatGPT offered superior rationale quality compared to Copilot, while Gemini exhibited higher specificity. CONCLUSIONS: LLM-based chatbots can provide acceptable surgical guidance for straightforward primary surgical cases, but their utility is limited in high-risk or complex clinical settings. The observed deficiencies in rationale and increased safety risks in complex cases suggest that LLMs should be integrated as auxiliary decision-support tools under expert supervision rather than used as autonomous decision-makers in glaucoma surgery planning.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Evaluation of Large Language Model-Generated Recommendations in Glaucoma Surgical Decision-Making. — 科研速览 Science Skim