S. L. Kampman, K. P. Braun, W. M. Otte
Background: Risk-of-bias (ROB) assessments represent an integral component of systematic reviews. However, this task is often highly repetitive, time-consuming, and may lack inter-rater consistency. Large language models (LLMs) offer opportunities for automation in systematic reviews, which may expedite and enhance the quality and consistency of research synthesis. Methods: Using zero-shot prompting, we designed an LLM-based pipeline as a virtual mimic of a human reviewer for the Quality in Prognosis Studies (QUIPS) framework. Then, focusing on prognostic research in a single discipline (neurology), we applied this pipeline to articles included in previously published systematic reviews. We studied inter-rater agreement between both (1) the LLM and the original human ROB assessments and (2) between original human ROB assessments. Results: 298 articles from 15 reviews across three domains (epilepsy, traumatic brain injury, stroke) were included. We demonstrate the feasibility of a tailored, prompt-engineered LLM pipeline for automating ROB assessments with the QUIPS tool. While LLM-human agreement was limited (Cohen's weighted kappa; = 0.22, 95% CI, 0.12 - 0.33), our data tentatively suggest, based on a small sample (n=5), that it may not be inferior to human-human agreement (Cohen's weighted kappa; = -0.25, 95% CI, -1.04 - 0.54). Wilcoxon signed-rank tests were statistically significant (p < 0.05) across four bias domains and for the overall risk scores, and rank-biserial correlations demonstrated human raters' tendency to assign higher risk scores than LLM counterparts. Conclusions: With targeted methodological refinements - including standardization of QUIPS implementation and validation against expert ratings - automated ROB assessments may meaningfully reduce time and cost of systematic reviews of prognosis studies in neurology and beyond.