Barbara Nussbaumer-Streit, Andreea Dobrescu, Amin Sharifan, Irma Klerings, Camilla Neubauer-Bruckner, Claus Nowak, Gerald Gartlehner
LLM-based classifiers excluded by automated screening 5.8-46.6% of 1,129 records identified by the original search and 7.7-34.3% of 1,003 records based on the SKP search. More specific classifiers excluded more records than broader ones. Of the 29 combinations applied to the original search results, 10 incorrectly excluded the same two relevant studies. Based on the SKP review, four combinations missed one study, and 10 combinations missed two studies. However, exclusion of these studies resulted in minimal changes to NMA-effect estimates (risk ratio changes of 0.0-0.04 in both scenarios) and did not alter certainty-of-evidence ratings.
BACKGROUND: The Epistemonikos Sustainable Knowledge Platform (SKP) integrates a suite of machine learning tools designed to make evidence synthesis more efficient, including access to the Epistemonikos Database of Trials (ED-Trials). One feature is the use of classifiers combined with large language models (LLMs) to support abstract screening during study selection. LLM-based classifiers can be applied to populations, interventions, and study design features. Their primary aim is to reduce the proportion of records needed to be screenedby automatically excluding records with a high probability of being ineligible. We aimed to evaluate SKP's LLM-based classifiers on abstract screenings of records retrieved via: (scenario 1) a traditional systematic literature search (original ), and (scenario 2) a search conducted within ED-Trials and subsequently forwarded to SKP.
METHODS: We used data from a completed but at the time unpublished systematic review and network meta-analysis (NMA)as reference standards. We evaluated 29 LLM-based classifier combinations in both scenarios and assessed the proportion of records excluded by automated screening as well as whether any relevant study was wrongly excluded and therefore missed. To determine the impact of missed studies, we re-ran the network meta-analyses of the original review, excluding missed studies from the model and we evaluated how missed studies impact the certainty-of-evidence ratings.
RESULTS: LLM-based classifiers excluded by automated screening 5.8-46.6% of 1,129 records identified by the original search and 7.7-34.3% of 1,003 records based on the SKP search. More specific classifiers excluded more records than broader ones. Of the 29 combinations applied to the original search results, 10 incorrectly excluded the same two relevant studies. Based on the SKP review, four combinations missed one study, and 10 combinations missed two studies. However, exclusion of these studies resulted in minimal changes to NMA-effect estimates (risk ratio changes of 0.0-0.04 in both scenarios) and did not alter certainty-of-evidence ratings.
DISCUSSION: LLM-based classifiers are a promising strategy for reducing the amount of records to screen while keeping the risk of missing relevant studies low. However, further evaluations using multiple-use cases across different medical topics are necessary to learn more about the generalizability of our results.