科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Revista da Associacao Medica Brasileira (1992)2026-01-01

ChatGPT-4o in delirium recognition: a pilot study.

Kubra Cingar Alpay, Suna Avci, Aysegul Gunduz, Duygu Ozata, Gokalp Kurthan Avlagı, Seyda Bilgin, Ummugulsum Durak, Alper Doventas, Deniz Suna Erdincler

一句话结论 · In one sentence

ChatGPT-4o showed promising but context-sensitive delirium recognition in geriatric vignettes. Large language models may serve as vigilance-enhancing decision-support tools under expert supervision, pending prospective validation.

原始摘要(英文原文)· Original abstract
OBJECTIVE: Delirium is common yet frequently under-recognized in older adults and is associated with increased morbidity and mortality. Large language models may support clinical reasoning, but their performance and consistency in complex geriatric syndromes remain unclear. METHODS: We analyzed 20 PubMed-indexed geriatric case reports converted into structured vignettes with explicit references to delirium removed. ChatGPT-4o was tested in three conditions (two independent sessions and one consecutive-session run) to assess delirium identification and temporal consistency; DeepSeek-V2 served as a comparator. Two clinicians (geriatrician and neurologist) independently rated outputs for clinical plausibility, concordance with the reference diagnosis, and decision-support value. RESULTS: ChatGPT-4o identified delirium in 70% of cases in both independent sessions and 90% in the consecutive-session run, whereas DeepSeek-V2 identified 50%. Differences across ChatGPT-4o conditions were not significant (p=0.135). Agreement between independent ChatGPT-4o sessions was substantial (κ=0.76), while agreement involving the consecutive-session run was low, suggesting contextual priming. Expert ratings were higher in successfully identified cases across domains (all p<0.05). Inter-rater reliability was highest for concordance (Intraclass Correlation Coefficient=0.86). CONCLUSION: ChatGPT-4o showed promising but context-sensitive delirium recognition in geriatric vignettes. Large language models may serve as vigilance-enhancing decision-support tools under expert supervision, pending prospective validation.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

ChatGPT-4o in delirium recognition: a pilot study. — 科研速览 Science Skim