科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ JCO clinical cancer informatics2026-01-01

Data Extraction From Oncology Imaging Reports by Large Language Models: A Comparative Accuracy Study.

Lea P Passweg, Johannes M Schwenke, Christof M Schönenberger, Flavio Locher, Julia Picker, Manuel Dieterle, Benjamin Thiele, Dimitri Hasler, Alessia Danelli, Andreas M Schmitt, Tobias Heye, Thomas Stojanov, Matthias Briel, Benjamin Kasenda

一句话结论 · In one sentence

In this study, LLMs were noninferior to human accuracy for classification of metastasis status but were inferior for response to treatment assessment.

原始摘要(英文原文)· Original abstract
PURPOSE: Manual data extraction from clinical text is resource-intensive. Locally hosted large language models (LLMs) may offer a privacy-preserving solution, but their performance on non-English data remains unclear. We investigated whether the accuracy of locally hosted LLMs is noninferior to human accuracy when determining metastasis status and treatment response from German radiology reports. METHODS: In this retrospective comparative accuracy study, five locally hosted LLMs (llama3.3:70b, mistral-small:24b, qwq:32b, qwen3:32b, and gpt-oss:120b) were compared against humans. A ground truth was established via duplicate human extraction and adjudication of discrepancies by a senior oncologist. The study was conducted at a tertiary referral hospital in Switzerland. We randomly sampled 400 radiology reports from adult patients with cancer (computed tomography, magnetic resonance imaging, positron emission tomography) generated between January 2023 and May 2025 and split them into a prompt optimization set (n = 100) and test set (n = 300). Primary outcomes were noninferiority (5 percentage points [pp] margin) of LLM classification accuracy compared with human accuracy for metastasis status (presence/absence by anatomic site) and treatment response categories. Secondary outcomes included accuracy for primary tumor diagnosis and radiologic absence of tumor. RESULTS: The analysis included 400 reports from 317 patients. In the test set (n = 300), the human accuracy for metastasis status was 98.4% (95% CI, 98.0 to 98.8). All LLMs were noninferior; gpt-oss:120b performed best (97.6% accuracy; difference, -0.8 pp [90% CI, -1.3 to -0.3 pp]). For response to treatment, the human accuracy was 86.0% (95% CI, 83.2 to 88.8). All LLMs were inferior; the most accurate model, gpt-oss:120b, achieved 78.3% (difference, -7.7 pp [90% CI, -11.6 to -3.8 pp]). CONCLUSION: In this study, LLMs were noninferior to human accuracy for classification of metastasis status but were inferior for response to treatment assessment.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Data Extraction From Oncology Imaging Reports by Large Language Models: A Comparative Accuracy Study. — 科研速览 Science Skim