科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Insights into imaging2026-09-21

Commercial large language models for oral cavity cancer staging using descriptive pre-treatment MRI reports: ready for standalone use in clinical practice?

Qi Yong Hemis Ai, Hoi Ming Kwok, Ming-Yi Lu, Tracy T S Lau, Kuo Feng Hung, Lun M Wong, Tiffany Y So, Ann D King, Ho Sang Leung

一句话结论 · In one sentence

Variable tested LLM performance in staging OCC from descriptive MRI reports suggested that they are not suitable as standalone staging tools. Additionally, Low MDT staging accuracy highlighted the need for structured reports that include clinical cancer staging to facilitate MDT assessment, thus potentially ensuring optimised disease management.

原始摘要(英文原文)· Original abstract
OBJECTIVES: MRI is used for staging head and neck cancer (HNC), but assigning T- and N-category criteria requires specialised expertise, and so many institutions offer only descriptive MRI reports. This study assessed potentials of commercial large language models (LLMs) to stage oral cavity cancer (OCC) using descriptive MRI reports and compared their accuracy with human experts. MATERIALS AND METHODS: 104 eligible MRI reports were processed by five commercial LLMs (ChatGPT5.4, ChatGPT5.0, ChatGPT4.1, Gemini3.1 and DeepSeekV3.2). T- and N-categories, and overall stage were extracted from outputs of the LLMs. MRI reports were also staged by two multidisciplinary team (MDT) members. Accuracies of the LLMs and MDT members for cancer staging were assessed against the reference staging standard and compared using the McNemar test. RESULTS: The LLMs showed accuracy of 54.8-76.0% for T-categorisation, 63.5-88.5% for N-categorisation, and 53.8-80.8% for overall stage. Compared with DeepSeekV3.2, ChatGPT and Gemini3.1 showed significantly higher accuracy (p ≤ 0.001), except for Gemini3.1 for T-categorisation (p = 0.07). No differences in accuracy for staging between ChatGPT versions and between them and Gemini3.1 (p = 0.08 to > 0.99). MDT members showed accuracy of 74.0-76.9% for T-categorisation, 76.9-77.9% for N-categorisation and 68.3-71.2% for overall stage. The tested LLMs did not consistently outperform MDT members for staging. CONCLUSION: Variable tested LLM performance in staging OCC from descriptive MRI reports suggested that they are not suitable as standalone staging tools. Additionally, Low MDT staging accuracy highlighted the need for structured reports that include clinical cancer staging to facilitate MDT assessment, thus potentially ensuring optimised disease management. KEY POINTS: Question Can commercial LLMs accurately assign cancer stage based on descriptive MRI reports to overcome the lack of specialised expertise in clinical practice? Findings Tested LLMs demonstrated variable performances, ranging 54.8-76.0% for T-categorisation, 63.5-88.5% for N-categorisation and 53.8-80.8% for overall stage, and failed to consistently outperform MDT members. Critical relevance statement Variable tested LLM performance in staging OCC from descriptive MRI reports suggested that they are not suitable as standalone staging tools. MDT members' low performances highlight that structured MRI reports with clinical cancer staging are needed to improve multidisciplinary assessment.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Commercial large language models for oral cavity cancer staging using descriptive pre-treatment MRI reports: ready for standalone use in clinical practice? — 科研速览 Science Skim