Erika Yamazawa, Ziad Rifi, John K Houten, James D Lin, Konstantinos Margetis, Jeremy Steinberger
LLMs provide generally accurate and understandable responses to common ACDF preoperative questions but lack nuanced risk stratification and individualized clinical guidance. These findings support their use as adjuncts, rather than replacements, for physician-led counseling.
BACKGROUND: Patients increasingly use large language models (LLMs) to obtain medical information, including preoperative guidance. However, the quality of LLM-generated responses to questions regarding anterior cervical discectomy and fusion (ACDF) remains unclear. This study evaluated the performance of GPT-5 and GROK 4 in answering common preoperative ACDF questions. We hypothesized that both models would provide generally accurate and comprehensible responses while demonstrating limitations in completeness, clinical nuance, and personalization.
METHODS: As an expert-rated pilot study, eighteen frequently asked ACDF preoperative questions were independently entered into GPT-5 and GROK 4 without additional prompting. Responses were evaluated by three attending spine neurosurgeons and one neurosurgery resident using 5-point Likert scales for accuracy, completeness, and comprehensibility. Inter-rater reliability was assessed using intraclass correlation coefficients (ICCs) and ordinal Krippendorff's alpha. Readability was measured using the Flesch-Kincaid Grade Level (FKGL).
RESULTS: Mean accuracy scores were identical for GPT-5 and GROK 4 (4.52). GROK 4 demonstrated higher completeness (4.61 vs. 4.26), whereas GPT-5 showed higher comprehensibility (4.74 vs. 4.41). Inter-attending ICC values were low (0.12-0.21), while Krippendorff's alpha indicated very high ordinal consistency despite restricted score variance. Agreement between the attending mean score and resident ratings was moderate (ICC ≈0.51-0.52). FKGL scores were approximately 10 for both models.
CONCLUSIONS: LLMs provide generally accurate and understandable responses to common ACDF preoperative questions but lack nuanced risk stratification and individualized clinical guidance. These findings support their use as adjuncts, rather than replacements, for physician-led counseling.