科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Hospital pediatrics2026-09-03

A Real-World Evaluation of Large Language Model-Generated Hospital Courses in Pediatrics.

Jasmine E Kim, Jonathan D Hron, Daniel J Kats, Kara Wong, Lily Kutz, Chase R Parsons

一句话结论 · In one sentence

In this pediatric evaluation of LLM-generated hospital courses reviewed by frontline clinicians, errors were common, but perceived potential harm was low, even assuming use without clinician correction. These findings support the use of LLM-generated hospital courses as starting drafts when paired with clinician review and institutional safeguards.

原始摘要(英文原文)· Original abstract
BACKGROUND: Large language model (LLM)-generated hospital courses are increasingly integrated into electronic health records (EHRs), yet their accuracy and safety in pediatric populations remain poorly characterized. OBJECTIVE: To evaluate the accuracy, text quality, and perceived potential harm of EHR-integrated and LLM-generated hospital courses in pediatric inpatient care during early clinical implementation. METHODS: We conducted a descriptive evaluation from June 10 to August 8, 2025, at an academic freestanding children's hospital using an Epic EHR with an integrated LLM tool (GPT-4o and GPT-4.1). Clinicians across multiple roles, including attending physicians, residents, and advanced practice providers, reviewed LLM-generated hospital courses for their own patients. Clinicians identified and categorized errors (hallucinations, inaccuracies, or omissions). They also rated text quality (comprehensiveness, conciseness, coherence) on a 5-point scale and perceived harm on an 8-point scale. RESULTS: A total of 129 LLM-generated hospital courses were reviewed (median length of stay, 3 days; IQR, 2-7) by 50 involved clinicians. Hallucinations occurred in 21% (95% CI, 14%-29%) of the hospital courses, inaccuracies in 41% (53/129; 95% CI, 33%-50%), and omissions in 24% (31/129; 95% CI, 17%-32%). Overall, perceived harm ratings were low (median, 0; IQR, 0-1). Text quality ratings were high (median [IQR]: comprehensiveness, 4 [3-5]; conciseness, 4 [4-5]; coherence, 4 [4-5]) and comparable with prior literature. CONCLUSION: In this pediatric evaluation of LLM-generated hospital courses reviewed by frontline clinicians, errors were common, but perceived potential harm was low, even assuming use without clinician correction. These findings support the use of LLM-generated hospital courses as starting drafts when paired with clinician review and institutional safeguards.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

A Real-World Evaluation of Large Language Model-Generated Hospital Courses in Pediatrics. — 科研速览 Science Skim