Shuting Yang, Haosen Jiao, Xuan Zheng, Chengbin Duan, Weiyu Hu, Dimin Zhu, Jinping Chen, Zongming Wang, Wenli Chen
Under model-specific tuning strategies, the clinically informed strategy showed little change in discrimination compared with the reference strategy (area under the receiver operating characteristic curve, 0.8054 vs. 0.8022; area under the precision-recall curve, 0.4603 vs. 0.4534), but had a lower Brier score (0.1100 vs. 0.1480) and calibration closer to ideal (intercept, 0.0256 vs. -0.3778; slope, 1.0209 vs. 2.5233). It also showed greater model-based net benefit across an exploratory threshold range of 0.05-0.50 and separated cohort-derived tertiles with observed non-GTR rates of 4.2%, 8.5%, and 35.6%. Postoperative pathological variables provided little incremental value.
INTRODUCTION: Preoperative estimation of non-gross-total resection (non-GTR), as defined on early postoperative magnetic resonance imaging (MRI), may support patient counseling and postoperative surveillance planning after surgery for nonfunctioning pituitary neuroendocrine tumors (NF-PitNETs). We evaluated whether a clinically informed, interpretable preoperative modeling strategy improved risk estimation compared with a reference strategy.
METHODS: We retrospectively analyzed 354 first-surgery patients from a single center with a documented early postoperative MRI-based resection-status endpoint; 57 (16.1%) had non-GTR. A prespecified clinically informed logistic strategy extended a reference model based on routine preoperative clinical and MRI variables by adding transformed terms representing nonlinear tumor burden and invasion severity. Models underwent stratified repeated nested cross-validation with five outer folds repeated 20 times and four-fold inner tuning. Performance was evaluated using patient-level averaged held-out predictions.
RESULTS: Under model-specific tuning strategies, the clinically informed strategy showed little change in discrimination compared with the reference strategy (area under the receiver operating characteristic curve, 0.8054 vs. 0.8022; area under the precision-recall curve, 0.4603 vs. 0.4534), but had a lower Brier score (0.1100 vs. 0.1480) and calibration closer to ideal (intercept, 0.0256 vs. -0.3778; slope, 1.0209 vs. 2.5233). It also showed greater model-based net benefit across an exploratory threshold range of 0.05-0.50 and separated cohort-derived tertiles with observed non-GTR rates of 4.2%, 8.5%, and 35.6%. Postoperative pathological variables provided little incremental value.
DISCUSSION: These findings represent single-center internal validation of complete modeling strategies. Because the strategies differed in both predictor representation and tuning objective, the calibration difference cannot be attributed to the transformed predictors alone. Independent external validation using standardized postoperative MRI, together with uniform-tuning and reduced-model sensitivity analyses, is required before application to counseling or surveillance planning.