科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Journal of evaluation in clinical practice2026-09-01

Performance and Consistency of Large Language Models in Key Labor-Intensive Tasks of Systematic Reviews.

Yi-Ran Liu, Xi-Ling Wang, Zi-Xuan Zhou, Yuan-Ji He, Yu-Zhang Li, Chuan Liu, Dong-Dong Zhou

一句话结论 · In one sentence

We provide a suite of automated tools for key SR tasks. By leveraging these tools to validate model outputs rather than starting manually, reviewers can significantly improve workflow efficiency while maintaining methodological rigour.

原始摘要(英文原文)· Original abstract
OBJECTIVE: To evaluate the performance and consistency of Large Language Models (LLMs) in core systematic review (SR) tasks and to introduce open-source tools for automated batch processing that provide decision rationales. METHODS: We assessed GPT-4o, Kimi-K2, DeepSeek-V3, and DeepSeek-R1 on five SR tasks: title/abstract screening (3550 records), full-text screening (233 texts), data extraction (112 RCTs), Risk of Bias (ROB) assessment (112 RCTs), and AMSTAR-2 assessment (20 SRs). Each model was evaluated twice to measure consistency. All outputs required supporting rationales and verbatim evidence. RESULTS: LLMs demonstrated proficiency across tasks, with generally high intra-model but lower inter-model consistency. In screening, models showed lower precision (0.27-0.40) but high recall (0.83-0.91) and specificity (0.83-0.91). DeepSeek-R1 and DeepSeek-V3 excelled in title/abstract and full-text screening, respectively. Data extraction accuracy was similar across models (0.78-0.82). Kimi-K2 achieved the highest ROB F1 score (0.71). AMSTAR-2 assessments were generally acceptable. DISCUSSION: While effective, LLMs showed variable performance across SR tasks. The mandatory output of rationales and evidence enhances transparency and allows for human verification of AI decisions. CONCLUSION: We provide a suite of automated tools for key SR tasks. By leveraging these tools to validate model outputs rather than starting manually, reviewers can significantly improve workflow efficiency while maintaining methodological rigour.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Performance and Consistency of Large Language Models in Key Labor-Intensive Tasks of Systematic Reviews. — 科研速览 Science Skim