科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Cureus2026-07-01

Anesthesiologists' Assessment of the Correctness, Completeness, and Coherence of ChatGPT Answers to Patient Preoperative Questions.

Daniel Werry, Garrett Barry, Vishal Uppal, Jonathan G Bailey

原始摘要(英文原文)· Original abstract
Introduction Patients are increasingly turning to large language models (LLMs) for answers. However, LLMs like ChatGPT are relatively new technology and are untested for this purpose. Inconsistency and "hallucinations" are commonly described problems with LLMs. This study tested the consistency, correctness, completeness, and coherence of answers to frequently asked patient questions, as scored by expert anesthesiologists. Methods We gathered commonly asked patient questions from professional association websites. Patient partners were consulted to confirm relevance of the questions. A subset of questions was tested for consistency using intraclass correlation ICC (3,k) across multiple leading prompts. Then anesthesiologists and trainees scored the responses on a Likert scale for 10 anesthesia-related patient questions, with and without a leading prompt. Interrater agreement was calculated, and Likert scores were described. Results In terms of consistency, all ICC (3,k) values were below the prespecified threshold of 0.75, although confidence intervals were wide. Overall median ratings of the correctness, completeness, and coherence of each question ranged from 3.4 to 3.9, corresponding to average ratings between "good" and "very good." Ratings for all responses were similar regardless of whether the question was submitted with or without the prompt. Excellent scores were rarer, receiving 81 (16%) for correctness, 66 (13%) for completeness, and 121 (23%) for coherence. Discussion ChatGPT offers appropriate and moderately consistent answers to anesthesia-related patient questions. Anesthesiologists should be aware that while ChatGPT responses are generally correct, complete, and coherent, this does not translate to patient comprehension, and the answers may not be appropriate for the particular patient and institution.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Anesthesiologists' Assessment of the Correctness, Completeness, and Coherence of ChatGPT Answers to Patient Preoperative Questions. — 科研速览 Science Skim