科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Neurosurgical review2026-09-16

Inter-Model Variability of ChatGPT, Gemini and Claude in Treatment Recommendations for Unruptured Intracranial Aneurysms.

Anish Narayan, Frederick Mariajoseph, Malik Farooq, Adrian Praeger, Ronil Chandra, Idrees Sher, Lee-Anne Slater, Calvin Gan, Andrew Gauden, Hamed Asadi, Justin Moore

原始摘要(英文原文)· Original abstract
Patients increasingly consult large language models (LLMs) before specialist review, yet whether frontier models agree with one another in nuanced clinical domains such as unruptured intracranial aneurysm (UIA) management remains uncharacterised. We therefore quantified inter-model variability across ChatGPT, Gemini and Claude, anchored against neurovascular multidisciplinary team (MDT) consensus and the Unruptured Intracranial Aneurysm Treatment Score (UIATS). Sixty-seven UIA cases referred to our neurovascular service (January-December 2025) were retrospectively analysed. De-identified clinical vignettes were submitted to Claude Opus 4.6, ChatGPT-5.4 and Gemini 3 Pro Thinking, each run five times. Within-model reproducibility was assessed using Fleiss' κ; inter-model agreement via Cohen's κ and McNemar's test; and each LLM's majority-vote anchored against MDT and UIATS using the same methods. Within-model reproducibility was almost perfect (Fleiss' κ 0.837-0.860). Pairwise inter-model agreement was asymmetric: ChatGPT-Gemini behaved near-identically (Cohen's κ = 0.850, 95% CI 0.71-0.97), whereas Claude diverged from both (κ = 0.688 and 0.667). Recommendations were non-unanimous in 13/67 cases (19.4%); Claude was the sole outlier in 8/13 (conservative in 7). Gemini was significantly more pro-treatment than Claude (McNemar p = 0.0117). Against MDT, Gemini and ChatGPT showed significant pro-treatment propensity (p = 0.0022, p = 0.0153). Claude was significantly more conservative than UIATS (p = 0.0162). Frontier LLMs are highly reproducible internally but diverge from one another asymmetrically. ChatGPT and Gemini behave near-identically while Claude diverges conservatively. Clinicians should anticipate AI-driven treatment expectations and future work should explore prompting, patient sentiment, and multicentre moderators.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Inter-Model Variability of ChatGPT, Gemini and Claude in Treatment Recommendations for Unruptured Intracranial Aneurysms. — 科研速览 Science Skim