科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Updates in surgery2026-08-26

Assessing the performance of three large language models in thyroid cancer tumour board decision-making.

Angus M H White, Kerry A Leyton, Neil Patel, David M Scott-Coombes, Richard J Egan, Michael J Stechman

原始摘要(英文原文)· Original abstract
Large language models (LLMs) have potential to support clinical decision-making, but their role in thyroid cancer multidisciplinary team (MDT) meetings remains uncertain. This study evaluated the performance of ChatGPT GPT-5.5, Muse Spark 1.0 (Meta AI) and DeepSeek-V4-Flash in reproducing the management decisions of a regional thyroid cancer MDT. One hundred thyroid cancer cases discussed by a regional MDT were submitted to each LLM using an identical standardised prompt referencing ATA, BTA and UICC guidelines. Recommendations were independently assessed by four consultant endocrine surgeons using a four-point concordance scale (0-3). Consensus scores were determined using the median. Inter-rater agreement was assessed using Fleiss' Kappa. Differences between models were analysed using Friedman and post hoc Wilcoxon signed-rank tests. Three hundred LLM recommendations were assessed. Inter-rater agreement was moderate (Fleiss' Kappa = 0.422, 95% CI 0.369-0.474). ChatGPT achieved the highest proportion of complete concordance with MDT recommendations (69%), followed by Meta AI (60%) and DeepSeek (53%). Clinically acceptable recommendations (scores 2-3) were produced in 93%, 88% and 89% of cases, respectively. Overall concordance differed significantly between models (Friedman χ2(2) = 10.03, p = 0.0066). ChatGPT significantly outperformed Meta AI (p = 0.015) and DeepSeek (p < 0.001), while no difference was observed between Meta AI and DeepSeek (p = 0.384). All three LLMs demonstrated high concordance with consultant-led thyroid cancer MDT decisions. ChatGPT achieved the highest overall concordance, although all models generated clinically acceptable recommendations in most cases. LLMs show promise as adjunctive decision-support tools but require continued clinician oversight and robust governance before routine clinical implementation.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Assessing the performance of three large language models in thyroid cancer tumour board decision-making. — 科研速览 Science Skim