Jack Lott, Nicholas Dietrich
LLMs demonstrated high reliability in identifying absolute contraindications but were less consistent for relative contraindications. RAG significantly improved classification accuracy across all models, supporting further evaluation of guideline-augmented LLMs for pre-procedural safety screening.
BACKGROUND: Large language models (LLMs) are increasingly used for clinical decision support in spine care, but their ability to classify contraindications to interventional procedures according to safety guidelines is unknown.
OBJECTIVE: This study evaluated LLM accuracy using the International Pain and Spine Intervention Society (IPSIS) Safety Practices and assessed whether retrieval-augmented generation (RAG) improves performance.
METHODS: A database of 318 clinical scenarios was curated from IPSIS Safety Practice guidelines spanning 12 procedure categories (102 absolute contraindications, 57 relative contraindications, 159 matched controls). Three LLMs (GPT-5.4, Gemini 3.1 Pro, Claude Sonnet 4.6) classified each case as no contraindication, relative, or absolute under baseline and RAG conditions. McNemar tests with Bonferroni correction compared conditions within models. Between-model comparisons used Cochran's Q test with post-hoc pairwise analysis.
RESULTS: Baseline accuracy ranged from 85.8% to 87.1% across models and improved to 92.5% to 98.4% with RAG (all p < 0.05). Weighted kappa improved from 0.885-0.913 to 0.949-0.990. Safety catch rates for absolute contraindications were 99.0% to 100.0% across all conditions. Relative contraindications showed the largest improvements with RAG, increasing from 42.1%-56.1% to 59.6%-93.0%.
CONCLUSION: LLMs demonstrated high reliability in identifying absolute contraindications but were less consistent for relative contraindications. RAG significantly improved classification accuracy across all models, supporting further evaluation of guideline-augmented LLMs for pre-procedural safety screening.