科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Journal of ophthalmology2026-01-01

Assessing Guideline Knowledge Alignment of Large Language Models in Ophthalmology: A Preclinical Benchmarking Study on KLEx Evidence-Based Guidelines.

Hongxia Lu, Yan Huo, Ruisi Xie, Zhengyuan Qu, Yutong Li, Shengjin Wang, Haohan Zou, Yan Wang

一句话结论 · In one sentence

Both models aligned closely with the guidelines, with significant concordance in GRADE-based recommendation strength. They may serve as preclinical reference tools in refractive surgery but not as a substitute for specialist judgment.

原始摘要(英文原文)· Original abstract
PURPOSE: To evaluate the guideline knowledge alignment of two large language models (LLMs), GPT-5.5 Instant and DeepSeek-V4, and to determine their preclinical reliability as reference tools in refractive surgery. METHODS: Using the 38 evidence-based recommendations of the international keratorefractive lenticule extraction (KLEx) guidelines as the gold standard, both LLMs were evaluated in their default configurations. Two ophthalmologists independently assessed the clinical safety and medical accuracy of the model responses using a 5-point Likert scale (1-5 points). Agreement between each model's recommendation strength and the guideline was quantified by intraclass correlation coefficient (ICC), structural reliability by the DISCERN scale, and readability by the Flesch Reading Ease (FRE) and Flesch-Kincaid Grade Level (FKGL) indices. RESULTS: Likert ratings did not differ between GPT-5.5 Instant (4.96 ± 0.206) and DeepSeek-V4 (4.89 ± 0.385; p = 0.134). Across 114 independent generations, the ICC for agreement with the guideline was 0.884 (95% CI, 0.836-0.918) for GPT-5.5 Instant and 0.739 (0.640-0.813) for DeepSeek-V4 (both p < 0.001). DISCERN scores were 70.18 ± 5.16 and 68.05 ± 5.41 (p = 0.085); FRE, 9.93 ± 8.33 and 3.74 ± 5.24 (p < 0.001); and FKGL, 16.07 ± 2.21 and 18.95 ± 1.96 (p < 0.001). CONCLUSION: Both models aligned closely with the guidelines, with significant concordance in GRADE-based recommendation strength. They may serve as preclinical reference tools in refractive surgery but not as a substitute for specialist judgment.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Assessing Guideline Knowledge Alignment of Large Language Models in Ophthalmology: A Preclinical Benchmarking Study on KLEx Evidence-Based Guidelines. — 科研速览 Science Skim