Fedan Avrumova, Michael D D Cesar, Giuseppe Loggia, Klemenz Strebel, Franziska C S Altorfer, Rafael I Castro, Jiaqi Zhu, Celeste Abjornson, Christopher M Bono, Darren R Lebl
OE demonstrated superior CSCG concordance and substantially lower susceptibility to hallucinations than ChatGPT. Nonetheless, physician oversight remains essential for safe AI integration into clinical practice.
BACKGROUND CONTEXT: Generative artificial intelligence (AI) is increasingly used in spine care; however, concerns remain regarding citation hallucinations and reliability. ChatGPT may generate inaccurate or fabricated references, whereas OpenEvidence (OE) prioritizes verified, peer-reviewed literature. This is the first study comparing OE and ChatGPT using cervical spine clinical guideline (CSCG) queries.
PURPOSE: To compare guideline alignment, citation validity, sourcing, and prompt-engineering effects between OE and ChatGPT using CSCGs.
STUDY DESIGN/SETTING: Cross-sectional comparative analysis.
PATIENT SAMPLE: No patient population was included.
OUTCOME MEASURES: Primary outcomes were guideline alignment score and citation validity (fully correct, partially hallucinated, or fully hallucinated). Secondary outcomes included source type, publication year, proportion published after CSCG release, and prompt-engineering effects.
METHODS: A total of 110 evidence-based clinical questions derived from 10 CSCGs authored by 6 academic societies were submitted to OE (v2.0) and ChatGPT-4o from June 1 to July15, 2025. A subset of prompts was repeated to evaluate prompt-engineering effects.
RESULTS: OE generated 999 citations with 100% accuracy, whereas only 184/393 (46.8%) ChatGPT citations met accuracy criteria (p<0.001). OE demonstrated higher guideline alignment than ChatGPT (4.6 ± 0.8 vs 4.1 ± 0.7; p = 0.03), with almost perfect interrater agreement (weighted Cohen's κ = 0.88; 95% CI, 0.80-0.97). . ChatGPT produced 88 partially hallucinated citations (22.4%), most commonly due to incorrect hyperlinks, author names, or publication years. OE cited more peer-reviewed literature (79.8% vs 60.1%; p<0.001) and more recent studies (2018±5.6 vs 2012±7.4; p<0.001). Prompt-engineering analysis showed OE maintained higher citation validity and fewer hallucinations, although accuracy declined when outputs were reformatted through ChatGPT, suggesting cross-model contamination.
CONCLUSION: OE demonstrated superior CSCG concordance and substantially lower susceptibility to hallucinations than ChatGPT. Nonetheless, physician oversight remains essential for safe AI integration into clinical practice.