科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ European archives of oto-rhino-laryngology : official journal of the European Federation of Oto-Rhino-Laryngological Societies (EUFOS) : affiliated with the German Society for Oto-Rhino-Laryngology - Head and Neck Surgery2026-09-16

Large language models for tympanostomy patient education: readability and guideline adherence.

Shreeya Bahethi, Hetal Lad, Shrey Shah, Sudeepti Vedula, Brian Manzi

一句话结论 · In one sentence

Google Search AI demonstrated the highest guideline concordance, though all models produced material too complex for typical patient comprehension. Structured prompting enhances clinical accuracy, but readability remains a key barrier to accessibility.

原始摘要(英文原文)· Original abstract
PURPOSE: To evaluate the accuracy and readability of large language model (LLM)-generated patient education materials regarding tympanostomy tube placement. METHODS: Over a two-month period, ChatGPT 4o, Gemini 2.5 Flash, and Google Search AI were prompted daily using long-form and layered prompt formats covering typical concerns regarding tympanostomy. Responses were scored on a 12-point rubric adapted from the AAO-HNS Clinical Practice Guidelines (CPG), assessing diagnostic accuracy, procedural clarity, and postoperative care. Readability was evaluated using Flesch Reading Ease and Flesch-Kincaid Grade Level. For each model and prompt type, average scores and variability were analyzed with 95% confidence intervals. Between-model differences were tested with Welch's t-tests and Cohen's d; temporal trends were analyzed via linear regression. RESULTS: For long-form outputs, Google Search AI and Gemini 2.5 Flash demonstrated the highest mean CPG adherence (96.4% and 96.6%, respectively), both significantly exceeding ChatGPT 4o (84.3%; both P < .001; Cohen's d ≈ 2.05). For layered prompt sessions, Google Search AI again demonstrated the highest adherence (91.7%), followed by Gemini 2.5 Flash (88.1%) and ChatGPT 4o (76.2%). Guideline adherence was temporally stable across all models (P > .05 for most). All outputs exceeded the recommended sixth-grade reading threshold (mean FKGL, 9.4 for long-form; 10.6 for layered), with no statistically significant readability differences between prompt strategies. CONCLUSIONS: Google Search AI demonstrated the highest guideline concordance, though all models produced material too complex for typical patient comprehension. Structured prompting enhances clinical accuracy, but readability remains a key barrier to accessibility.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Large language models for tympanostomy patient education: readability and guideline adherence. — 科研速览 Science Skim