科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ World journal of otorhinolaryngology - head and neck surgery2026-08-27

Generative AI and Clinicians Show Comparable Prognostic Reasoning From Clinical Narratives in Biologic-Treated CRSwNP.

Sholem Hack, Chase Kahn, Ainhoa Garcia-Lliberos, Cristina Rodriguez-Prado, Ameen Biadsee, Maxime Fieux, Mark Liu, Eran Glikson, Masayoshi Takashima, Omar G Ahmed, Raj Sindwani, Dennis M Tang, Christopher R Roxbury

一句话结论 · In one sentence

When restricted to short pretreatment text, LLMs produced prognostic judgments within the range of clinician variability for severe, biologic-treated CRSwNP outcomes. These findings suggest that in information-limited settings, LLMs may provide an auxiliary reasoning signal rather than a standalone decision tool.

原始摘要(英文原文)· Original abstract
OBJECTIVE: To compare the ability of large language models (LLMs) and otolaryngologists to identify prognostic signals from brief clinical vignettes in CRSwNP. METHODS: In this blinded study, 68 adults initiating biologic therapy for CRSwNP (≥ 36 months follow-up) were represented by standardized vignettes derived from documentation immediately before biologic initiation. Vignettes summarized symptoms, prior surgery, comorbidities, and medications. Biomarkers, imaging scores, smell testing, and follow-up data were excluded to isolate text-based prognostic reasoning under identical constraints. Five attending otolaryngologists, one rhinology fellow, and two residents independently predicted four 5-year outcomes: subsequent endoscopic sinus surgery, recurrent systemic steroid bursts, biologic switch, and a composite endpoint. Multiple LLMs were queried in identical zero-shot format across three sessions to assess stability. Predictions were compared with verified outcomes using accuracy, sensitivity, specificity, F1 score, Cohen's κ, and Matthews correlation coefficient. RESULTS: Five-year outcome prevalences were 38.2% (26 of 68) for subsequent sinus surgery, 26.5% (18 of 68) for steroid bursts, 14.7% (10 of 68) for biologic switch, and 58.8% (40 of 68) for composite failure. Macro-averaged LLM accuracies ranged from 65.0% to 76.1%, compared with 65.7% ± 7.4% for human raters. The best-performing LLM achieved accuracy comparable to the top attending. Both groups showed higher negative than positive predictive values, indicating better discrimination for patients without adverse events. Agreement between LLMs and clinicians was moderate and within the range of inter-clinician variability. CONCLUSION: When restricted to short pretreatment text, LLMs produced prognostic judgments within the range of clinician variability for severe, biologic-treated CRSwNP outcomes. These findings suggest that in information-limited settings, LLMs may provide an auxiliary reasoning signal rather than a standalone decision tool.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Generative AI and Clinicians Show Comparable Prognostic Reasoning From Clinical Narratives in Biologic-Treated CRSwNP. — 科研速览 Science Skim