Ronald Luo, Abu Ilius Faisal, Ziya Sastimoglu, Jamal Deen
Background: Systematic reviews are crucial for evidence-based practice but are often hindered by labor-intensive methodologies. Recent advances in Large Language Models (LLMs) may help improve review efficiency and replicability.Objective: Building on prior research, this study evaluates the impact of prompt engineering on LLM performance in automating systematic reviews, specifically examining the "zero-shot" capability for predicting article inclusion based on titles and abstracts alone. Methods: We reanalyzed 24,534 studies across three systematic reviews using two sets of prompts with GPT-3.5 Turbo. One set rated the suitability of a study on a scale of one to ten, while the other elicited a binary response. Both ordinal-response and binary-response prompts were assessed for accuracy, sensitivity, specificity, predictive values, F1 score, and Matthews Correlation Coefficient (MCC). We also evaluated Receiver Operating Characteristic (ROC) curves for ordinal-response prompts and statistically validated the results using K-fold permutations and one-way analysis of variance.Results: Ordinal prompting allowed for fine-tuning of sensitivity and specificity, proving to be a viable zero-shot inclusion prompt with ROC Area Under Curve (AUC) scores between 80-90%. Including studies rated seven or higher achieved sensitivity and specificity of ~80%, while using the mean score increased sensitivity to nearly 100%, a task not possible with earlier strategies.Conclusion: Prompt engineering offers a scalable solution that allows researchers to target highly relevant articles or maximize sensitivity to include nearly all relevant studies if more time is available.