Jaden Wise, Isabella Merem, Justin Gumbs, Zachary Comella, Angela Rodio, Victoria Lin, Victoria Verrengia, Rudy Paul, Min Shi, Maohua Lin, Frank D Vrionis
AI reliability in spine surgery depends on parameter-level development choices rather than algorithm selection alone. More standardized reporting, calibration assessment, prospective validation, and lifecycle monitoring are needed before AI tools can be integrated safely into spine surgery workflows.
BACKGROUND AND OBJECTIVE: Artificial intelligence (AI) is increasingly applied in spine surgery for diagnosis, operative planning, intraoperative guidance, and outcome prediction. However, reported technical performance does not always translate into clinical reliability. This narrative review evaluates how parameter-level decisions influence the reliability and clinical translation of AI models in spine surgery.
METHODS: A structured literature search of PubMed, Scopus, and Google Scholar was performed for English-language studies published from January 2015 through March 2025. Search terms combined spine surgery concepts with AI, machine learning, deep learning, hyperparameters, learning rate, feature selection, regularization, model validation, and training strategy. Studies were synthesized narratively according to recurring parameter domains and clinical implementation context.
KEY CONTENT AND FINDINGS: Across preoperative, intraoperative, and postoperative applications, learning rate, feature selection, regularization, model architecture, optimization strategy, and evaluation metrics shaped convergence behavior, overfitting risk, interpretability, calibration, and external validity. Reported performance varied substantially by task, dataset, outcome definition, and validation strategy. Common limitations included retrospective single-center datasets, inconsistent parameter reporting, limited external validation, and challenges related to bias, interpretability, workflow integration, and model drift.
CONCLUSIONS: AI reliability in spine surgery depends on parameter-level development choices rather than algorithm selection alone. More standardized reporting, calibration assessment, prospective validation, and lifecycle monitoring are needed before AI tools can be integrated safely into spine surgery workflows.