科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Digital health2026-01-01

Open-source versus proprietary large language models for academic writing in medicine: A controlled evaluation of DeepSeek-R1 and ChatGPT-o1 in neuroradiology.

Semil Eminovic, Bogdan Levita, Robin Schmidt, Maximilian Lindholz, Dmitriy Desser, Tobias Penzkofer, Mike P Wattjes, Jawed Nawabi

一句话结论 · In one sentence

DeepSeek-R1 outperformed ChatGPT-o1 in generating academic content for neuroradiology, highlighting the potential of transparent, locally deployable models. However, moderate quality, citation errors and hallucinations indicate that LLMs are not yet sufficient for unsupervised academic writing and demand rigorous human review.

原始摘要(英文原文)· Original abstract
OBJECTIVES: Large language models (LLMs) are being used to facilitate academic writing. We aimed (1) to assess the quality, integrity, and factual reliability of LLM-assisted academic writing in medicine, and (2) to compare the performance of open-source versus proprietary LLMs, reflecting differences in model transparency. METHODS: In this prospective, controlled evaluation study, an open-source model (DeepSeek-R1) and proprietary model (ChatGPT-o1) completed ten academic tasks: generating five scientific essays and five evidence-based question-answers covering clinical topics in neuroradiology. Two radiologists rated outputs using 14 Likert and 7 binary criteria across the domains of academic quality, linguistic expression, factual reliability. Paired Wilcoxon/McNemar (Holm) assessed differences; ICC/Cohen's κ assessed reliability. RESULTS: DeepSeek-R1 achieved higher overall mean Likert scores than ChatGPT-o1 (mean ratings ± standard deviation: 3.23 ± 0.44 vs 3.02 ± 0.30, p = 0.021). It significantly outperformed in reasoning depth (p = 0.015), contextual coherence (p = 0.043), subtlety (p = 0.037), and evidence integration (p = 0.027). Although both models cited sources in every output, citation reliability was poor. DeepSeek-R1 cited more real publications (55.1% vs. 33.3%, p=0.053) but showed a higher confirmed fabrication rate (36.7% vs. 2.6%, p<0.001). DeepSeek-R1 respected word limits in 90%, ChatGPT-o1 only in 50%. Neither reviewer reliably noticed AI-generated text features. Inter-rater reliability was poor for Likert criteria (ICC = 0.466) and substantial for binary items (κ = 0.746). CONCLUSION: DeepSeek-R1 outperformed ChatGPT-o1 in generating academic content for neuroradiology, highlighting the potential of transparent, locally deployable models. However, moderate quality, citation errors and hallucinations indicate that LLMs are not yet sufficient for unsupervised academic writing and demand rigorous human review.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Open-source versus proprietary large language models for academic writing in medicine: A controlled evaluation of DeepSeek-R1 and ChatGPT-o1 in neuroradiology. — 科研速览 Science Skim