科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Frontiers in oncology2026-01-01

Comparative assessment of DeepSeek-R1 and GPT-4 for structured ultrasound reporting of adnexal masses.

Mengjuan Zhang, Bin Wang, Ran Li, Xinwu Cui, Xiaofeng Zhang, Jingzhe Wang, Lan Jia, Lijuan Guo, Yan Kong

一句话结论 · In one sentence

Both DeepSeek-R1 and GPT-4 demonstrate potential for automating structured ultrasound reporting of AMs, with DeepSeek-R1 showing superior performance in classification and recommendations. However, their diagnostic performance in discriminating benign from malignant masses, while promising, requires further refinement before independent clinical application. These findings support their role as assistive tools, pending integration with image data and prospective validation.

原始摘要(英文原文)· Original abstract
OBJECTIVE: This study primarily evaluated the ability of two large language models (DeepSeek-R1 and GPT-4) to generate structured ultrasound reports from free-text adnexal mass reports. Secondarily, we assessed their accuracy in O-RADS classification and management recommendations, with an exploratory analysis of their performance in benign versus malignant discrimination. METHODS: This study included 215 free-text ultrasound reports of adnexal masses (AMs) between July 2024 and March 2025. Each report was processed three times per model; majority voting was used to determine the final output. Three senior radiologists, blinded to model identity, evaluated structured reports against a predefined template, O-RADS categories, and management guidelines. Histopathology served as the reference standard for benign/malignant discrimination; while expert consensus served as the reference for the other endpoints. RESULTS: GPT-4 showed numerically better performance than DeepSeek-R1 in generating structured reports (99.5% vs. 96.7%, p = 0.07).DeepSeek-R1 outperformed GPT-4 in O-RADS accuracy (64.4% vs. 52.9%, p < 0.001) and appropriate management recommendations (66.4% vs. 58.0%, p = 0.007). Both models demonstrated good discrimination ability for benign versus malignant classification (AUC > 0.80 for both). CONCLUSION: Both DeepSeek-R1 and GPT-4 demonstrate potential for automating structured ultrasound reporting of AMs, with DeepSeek-R1 showing superior performance in classification and recommendations. However, their diagnostic performance in discriminating benign from malignant masses, while promising, requires further refinement before independent clinical application. These findings support their role as assistive tools, pending integration with image data and prospective validation.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Comparative assessment of DeepSeek-R1 and GPT-4 for structured ultrasound reporting of adnexal masses. — 科研速览 Science Skim