Semil Eminovic, Bogdan Levita, Robin Schmidt, Maximilian Lindholz, Dmitriy Desser, Tobias Penzkofer, Mike P Wattjes, Jawed Nawabi
DeepSeek-R1 outperformed ChatGPT-o1 in generating academic content for neuroradiology, highlighting the potential of transparent, locally deployable models. However, moderate quality, citation errors and hallucinations indicate that LLMs are not yet sufficient for unsupervised academic writing and demand rigorous human review.
OBJECTIVES: Large language models (LLMs) are being used to facilitate academic writing. We aimed (1) to assess the quality, integrity, and factual reliability of LLM-assisted academic writing in medicine, and (2) to compare the performance of open-source versus proprietary LLMs, reflecting differences in model transparency.
METHODS: In this prospective, controlled evaluation study, an open-source model (DeepSeek-R1) and proprietary model (ChatGPT-o1) completed ten academic tasks: generating five scientific essays and five evidence-based question-answers covering clinical topics in neuroradiology. Two radiologists rated outputs using 14 Likert and 7 binary criteria across the domains of academic quality, linguistic expression, factual reliability. Paired Wilcoxon/McNemar (Holm) assessed differences; ICC/Cohen's κ assessed reliability.
RESULTS: DeepSeek-R1 achieved higher overall mean Likert scores than ChatGPT-o1 (mean ratings ± standard deviation: 3.23 ± 0.44 vs 3.02 ± 0.30, p = 0.021). It significantly outperformed in reasoning depth (p = 0.015), contextual coherence (p = 0.043), subtlety (p = 0.037), and evidence integration (p = 0.027). Although both models cited sources in every output, citation reliability was poor. DeepSeek-R1 cited more real publications (55.1% vs. 33.3%, p=0.053) but showed a higher confirmed fabrication rate (36.7% vs. 2.6%, p<0.001). DeepSeek-R1 respected word limits in 90%, ChatGPT-o1 only in 50%. Neither reviewer reliably noticed AI-generated text features. Inter-rater reliability was poor for Likert criteria (ICC = 0.466) and substantial for binary items (κ = 0.746).
CONCLUSION: DeepSeek-R1 outperformed ChatGPT-o1 in generating academic content for neuroradiology, highlighting the potential of transparent, locally deployable models. However, moderate quality, citation errors and hallucinations indicate that LLMs are not yet sufficient for unsupervised academic writing and demand rigorous human review.