科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Cancer medicine2026-09-01

Agreement in Staging and Treatment Recommendations Among Clinicians, Text-Only Clinicians, DeepSeek-V3, and ChatGPT-4o for Nasopharyngeal Carcinoma Patients.

Shenglan Lin, Cai Zhang, Feifei Zhong, Xiyi Liao, Liuyun Gong, Dunhuang Wang

一句话结论 · In one sentence

The AI tools demonstrated moderate agreement with clinicians in overall clinical staging for NPC, with heterogeneous performance across T, N, and M categories (almost perfect for M stage, moderate for T and N stages), whereas their preferred treatment recommendations showed only slight agreement with clinical decision-making.

原始摘要(英文原文)· Original abstract
PURPOSE: With AI tools being increasingly utilized for medical inquiries, this study evaluated the agreement between clinicians, text-only clinicians, DeepSeek-V3, and ChatGPT-4o in TNM staging and preferred treatment recommendations for newly diagnosed nasopharyngeal carcinoma (NPC) patients. METHODS: A retrospective study analyzed 322 consecutive NPC patients treated at our institution from January 2023 to February 2025. TNM staging (AJCC 8th edition) and preferred treatment recommendations were independently assessed by text-only clinicians, DeepSeek-V3, and ChatGPT-4o. Interrater agreement was quantified using Cohen's kappa coefficient and Fleiss' kappa analysis (κ), with the range of κ = 0.81-1.00 considered almost perfect agreement. RESULTS: The cohort comprised 244 males and 78 females (median age, 52 years, range, 18-77 years). Cohen's kappa analysis showed that for clinical staging, moderate agreement was exhibited in clinician-AI comparisons, while substantial agreement was observed between clinicians and text-only clinicians. For preferred treatment recommendations, slight agreement was noted in clinician-AI comparisons, while moderate agreement was found between clinicians and text-only clinicians. Fleiss' kappa analysis demonstrated moderate agreement among the 4 raters for T stage, N stage, and clinical staging, while M stage achieved almost perfect agreement. However, overall agreement for treatment recommendations was fair. CONCLUSIONS: The AI tools demonstrated moderate agreement with clinicians in overall clinical staging for NPC, with heterogeneous performance across T, N, and M categories (almost perfect for M stage, moderate for T and N stages), whereas their preferred treatment recommendations showed only slight agreement with clinical decision-making.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Agreement in Staging and Treatment Recommendations Among Clinicians, Text-Only Clinicians, DeepSeek-V3, and ChatGPT-4o for Nasopharyngeal Carcinoma Patients. — 科研速览 Science Skim