科研速览 · Science Skim继续刷下去 · Keep skimming →
◇ medRxiv2026-09-04· health informatics

Citation reliability of frontier large language models in medical writing and its automated verification

R. Shin, J.-M. Lee, J. Park, J.-S. Kwun, H.-W. Cho, S.-H. Kang, K.-H. Jeon

原始摘要(英文原文)· Original abstract
Large language models (LLMs) are increasingly used to draft medical manuscripts, yet their citations are unreliable and clinicians lack a validated way to verify them. We evaluated three frontier LLMs, Claude Opus 4.8, GPT-5.5, and Gemini 3.5 Flash, generating 270 cardiology narrative reviews with web search enabled, and verified all 8,050 references against PubMed. Problematic references accounted for 11.5% of GPT-5.5 output, 29.2% of Claude output, and 29.6% of Gemini output (P < 0.001), with no significant gradient across topics of differing publication volume (P = 0.052). Misattribution, a valid PubMed identifier that resolves to a different article, made up 77% of errors, whereas fabrication was rare (0.6%). Against an expert-adjudicated set of 270 references, an LLM-based Chain-of-Verification (CoVe) detected 60 of 62 problematic references (sensitivity 96.8%, specificity 98.6%), including every misattribution and fabrication. LLM-generated citations require identifier-level verification, and CoVe provides it at expert-level accuracy.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Citation reliability of frontier large language models in medical writing and its automated verification — 科研速览 Science Skim