Tongxin Li, Xiaoqin Zhang, Jincheng Chen, Zhili Liu, Shanshan Xu, Lie Jiang, Ziyan Wei, Wei Guo, Yan He, Wei Wu, Liu Yang, Yi Wu
An interpretable NLR-centered XGBoost model built on routine clinical and hematological variables provided moderate-accuracy prediction of ICI response in ESCC, with an optimal NLR cutoff of 3.40 for initial clinical screening, pending external validation.
BACKGROUND: Immune checkpoint inhibitors (ICIs) have transformed the treatment of esophageal cancer, yet only a subset of patients derive durable clinical benefit. The neutrophil-to-lymphocyte ratio (NLR) is a readily accessible inflammatory biomarker derived from routine blood tests; however, no interpretable machine learning model tailored to esophageal squamous cell carcinoma (ESCC) currently integrates NLR with multidimensional clinical features for response prediction. We developed an NLR-centered model to predict ICI response in patients with ESCC.
METHODS: This was a multicenter retrospective cohort study of 419 ESCC patients treated with ICIs at four medical centers. Eighteen pretreatment clinical and hematological features were used as input features; the outcome was ICI response. Six machine learning algorithms spanning the principal tabular-learning families were compared. Random Forest hyperparameters were selected by RandomizedSearchCV; the remaining models used pre-specified regularized hyperparameters to control overfitting. Internal validation used 5-fold stratified cross-validation (CV). Interpretability used SHapley Additive exPlanations (SHAP), and the optimal NLR cutoff was determined by the Youden index of the receiver operating characteristic (ROC) curve.
RESULTS: Under 5-fold stratified CV, Gradient Boosting Machine (GBM) achieved the highest area under the curve (AUC) (0.783); eXtreme Gradient Boosting (XGBoost) and Random Forest tied second (both 0.771), followed by Logistic L1 (0.747), Multilayer Perceptron (0.700), and K-Nearest Neighbors (0.694). XGBoost was selected as the primary interpretable model based on the overall balance of discrimination, calibration, clinical utility, and SHAP interpretability rather than on AUC alone. At the Youden-optimal threshold (0.612), XGBoost reached sensitivity 0.777, specificity 0.680, accuracy 74.2%, F1-score 0.795, positive predictive value 0.813, negative predictive value 0.630, and Brier 0.178. SHAP ranked NLR first by both mean |SHAP| (0.633) and gain (0.185), followed by Eastern Cooperative Oncology Group performance status and neutrophil count. NLR was lower in responders than non-responders (P<0.001; optimal cutoff 3.40).
CONCLUSIONS: An interpretable NLR-centered XGBoost model built on routine clinical and hematological variables provided moderate-accuracy prediction of ICI response in ESCC, with an optimal NLR cutoff of 3.40 for initial clinical screening, pending external validation.