科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ JCO clinical cancer informatics2026-01-01

Accuracy and Safety of Large Language Models in Endometrial Cancer Decision Making: A Case-Based In Silico Benchmarking Study.

Emanuele Perrone, Giuseppe Parisi, Maria Consiglia Giuliano, Ilaria Capasso, Nicola Macellari, Luca Russo, Angela Santoro, Francesco Fanfani

一句话结论 · In one sentence

Mainstream consumer large language models showed substantially different performance in postoperative endometrial cancer decision making. Although Gemini achieved higher concordance and fewer major safety issues than ChatGPT and Claude, no model demonstrated performance sufficient to support autonomous clinical use in multidisciplinary management.

原始摘要(英文原文)· Original abstract
PURPOSE: To compare the concordance of ChatGPT, Gemini, and Claude with a prespecified expert guideline-based reference standard in fabricated endometrial cancer clinical vignettes under standardized prompting. METHODS: We conducted a case-based in silico benchmarking study using 35 fabricated postoperative endometrial cancer vignettes representing a broad spectrum of ESGO-ESTRO-ESP 2025 management scenarios. Each vignette was submitted to ChatGPT, Gemini, and Claude in independent chat sessions using the same standardized prompt. The primary end point was concordance with the prespecified expert reference standard, scored as 0 (discordant), 1 (partially concordant), or 2 (fully concordant). Secondary end points were major safety issues and recognition of missing critical information. A post hoc subgroup analysis evaluated the effect of guideline-informed prompting in 12 cases. RESULTS: Concordance differed significantly across models (P < .001). Gemini achieved the highest performance, with a median concordance score of 2 (IQR 1-2), compared with 1 (IQR 0-1) for ChatGPT and 0 (IQR 0-1) for Claude. Fully concordant recommendations were generated in 65.7% of cases by Gemini, 2.9% by ChatGPT, and 8.6% by Claude. Major safety issues also differed across models (P = .007), occurring in 22.9% of Gemini responses, 48.6% of ChatGPT responses, and 54.3% of Claude responses. In the subset of vignettes with intentionally missing decisive information, Gemini identified the need for additional data in 83.3% of cases, compared with 50.0% for both ChatGPT and Claude. In the post hoc subgroup analysis, guideline-informed prompting significantly improved concordance for Gemini and Claude. CONCLUSION: Mainstream consumer large language models showed substantially different performance in postoperative endometrial cancer decision making. Although Gemini achieved higher concordance and fewer major safety issues than ChatGPT and Claude, no model demonstrated performance sufficient to support autonomous clinical use in multidisciplinary management.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Accuracy and Safety of Large Language Models in Endometrial Cancer Decision Making: A Case-Based In Silico Benchmarking Study. — 科研速览 Science Skim