科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Frontiers in Pediatrics2026-02-17· Medicine

Should we leave paediatric emergency triage to artificial intelligence? A comparison of ChatGPT 4o and Grok 3

Emre Aygun, Aysenur Imdat, Nazan Karabulut Dalgiç

原始摘要(英文原文)· Original abstract
Background The growing number of patients in paediatric emergency departments requires fast and precise triage assessments. The implementation of large language models faces obstacles due to their limited interpretability. We aimed to compare the performance of ChatGPT 4o and Grok 3 with that of nurses and physicians in paediatric emergency triage. Methods This prospective observational study evaluated paediatric emergency patients presenting to our paediatric emergency department between March and April 2025. Demographic data, chronic disease status, presenting complaints, and vital signs were documented. Patients were triaged according to ESI criteria by nurses, paediatric specialists (gold standard), ChatGPT 4o, and Grok 3. Inter-rater agreement was analysed using Cohen's kappa**. Cochran's Q and McNemar's tests were used for paired comparisons.** Results A total of 1,505 paediatric emergency patients were included in the analysis. No ESI-1 cases were observed; therefore, critical patients were defined as ESI-2. Nurses achieved 53.1% (95% CI: 50.6–55.6) accuracy in triage assessments, while ChatGPT 4o achieved 76.1% (95% CI: 73.9–78.2) and Grok 3 achieved 47.0% (95% CI: 44.5–49.6) accuracy (Cochran's Q = 275.68, p < 0.001). ChatGPT 4o showed good agreement with physicians ( κ = 0.69). For critical patient identification, sensitivity was 37.2% for nurses, 82.9% for ChatGPT 4o, and 97.7% for Grok 3; however, Grok 3 demonstrated substantial over-triage (36.3%) and low positive predictive value (37.2%). ChatGPT 4o achieved the lowest mean absolute ESI error (0.25 ± 0.45). Nurses' critical patient recognition improved from 28.3% to 59.5% ( p < 0.01) for children with chronic illnesses. Conclusion ChatGPT 4o achieved the most favourable balance of sensitivity and specificity. The superior performance of nurses in recognising critically ill patients with chronic diseases suggests that AI systems should augment nursing expertise rather than replace it.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Should we leave paediatric emergency triage to artificial intelligence? A comparison of ChatGPT 4o and Grok 3 — 科研速览 Science Skim