科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ International ophthalmology2026-09-07

Generative artificial intelligence to augment ethical problem solving in ophthalmology: GPT-5.1 versus a human ethicist.

Daniel C Kelly, Ishan Chillikatil, Jonathon M Monroe, Rebika Khanal, Andrew Trippiedi, Chelsea-Jane Arcalas, Matthew R Claxton, Tochukwu Ndukwe, Elizabeth Pogrebniak, Joshua Barnett, Jacquelyn O'Banion, Jeremy K Jones, Rebecca F Neustein

一句话结论 · In one sentence

Compared to human ethicist responses, GPT-5.1 responses demonstrated lower readability; however, GPT-5.1 responses received significantly higher ratings for both outcomes of likelihood-of-use and perceived patient impact, respectively. These results advocate for further explorations of GPT-5.1 and other LLMs as potentially useful tools for evaluating common bioethical challenges in ophthalmology.

原始摘要(英文原文)· Original abstract
PURPOSE: While many large language models (LLMs) have been extensively investigated for their clinical decision-making capabilities, few studies have characterized their abilities to reason through complex, open-ended ethics cases. This study compared GPT-5.1 and expert human ethicist responses to real-world ethical scenarios specific to ophthalmology. METHODS: Ten ethical scenarios from the American Academy of Ophthalmology's Ask the Ethicist website were randomly selected and presented to GPT-5.1 using ChatGPT. AI-generated responses were subsequently compared to those of the expert human ethicist using conventional readability metrics. A panel of 10 physicians independently rated all responses via 6-point and 5-point Likert scales for both outcomes of likelihood-of-use in their own careers and perceived patient impact, respectively. RESULTS: GPT-5.1 versus human ethicist responses differed significantly on Flesch Reading Ease (10.9 ± 9.6 vs. 25.4 ± 10.9, p = 0.002) and Gunning Fog Index (21.3 ± 1.9 vs. 19.5 ± 2.9, p = 0.037). For likelihood-of-use, median ratings were 4.1 [3.8-4.3] for human ethicist versus 5.0 [4.7-5.2] for GPT-5.1 responses (p = 0.013). Median ratings for perceived patient impact of human ethicist versus GPT-5.1 responses were 3.6 [3.3-3.7] versus 3.9 [3.8-4.4], p = 0.008. Inter-rater reliability was moderate for both outcomes and response sources (ICC [2, 10]: 0.56-0.69). CONCLUSIONS: Compared to human ethicist responses, GPT-5.1 responses demonstrated lower readability; however, GPT-5.1 responses received significantly higher ratings for both outcomes of likelihood-of-use and perceived patient impact, respectively. These results advocate for further explorations of GPT-5.1 and other LLMs as potentially useful tools for evaluating common bioethical challenges in ophthalmology.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Generative artificial intelligence to augment ethical problem solving in ophthalmology: GPT-5.1 versus a human ethicist. — 科研速览 Science Skim