İbrahim Doğan, Serdar Çoban, Ezgi Sönmez, Veysi Gökyer, Serpil Sevimli Deniz
In this contemporary low-mortality surgical cohort, the Boey and PULP scores retained useful discriminatory ability but showed calibration mismatch when applied to outcomes beyond their original development setting. Local recalibration improved agreement between predicted and observed risks, illustrating the importance of model updating when established prediction tools are transported across populations with different baseline risks. These recalibrated models should be considered proof-of-concept updates requiring independent validation rather than immediately deployable clinical tools.
BACKGROUND: The Boey and Peptic Ulcer Perforation (PULP) scores were developed to predict mortality following perforated peptic ulcer (PPU), but their performance may vary when applied to contemporary populations with different case mix and baseline risk. Their performance for clinically relevant postoperative outcomes and the potential role of model updating remain uncertain. This study externally validated established prognostic scores and evaluated their discrimination, calibration, decision-analytic net benefit, and the potential effect of score-based recalibration.
METHODS: This retrospective single-center cohort study included 286 consecutive patients who underwent emergency surgery for PPU. The primary outcome was postoperative complications according to the Clavien-Dindo classification, while secondary outcomes included intensive care unit (ICU) admission and 30-day mortality. The discriminatory performance of the Boey, PULP, American Society of Anesthesiologists (ASA), and Charlson Comorbidity Index (CCI) scores was assessed using receiver operating characteristic (ROC) analysis. Calibration, decision curve analysis, score-based logistic recalibration, and the incremental prognostic value of the neutrophil-to-lymphocyte ratio (NLR) were also evaluated.
RESULTS: The median age was 34 years, 88.1% of patients were male, and the 30-day mortality rate was 2.8%. Postoperative complications occurred in 25.2% of patients, and 11.5% required ICU admission. For postoperative complications, the PULP score demonstrated the highest discrimination (AUC, 0.713; 95% CI, 0.646-0.778), followed by the Boey score (AUC, 0.668; 95% CI, 0.600-0.734). For ICU admission, the Boey score showed the highest discrimination (AUC = 0.840), followed by the PULP score (AUC = 0.827). The ASA classification and CCI showed AUC values of 0.801 and 0.736, respectively. Direct comparison of the PULP and Boey scores showed no statistically significant difference in discrimination for either postoperative complications (P = 0.081) or ICU admission (P = 0.549). Calibration mismatch was observed in the original score-based models, whereas logistic recalibration improved agreement between predicted and observed risks. Decision curve analysis showed greater net benefit for the disease-specific scores across selected threshold probabilities, while addition of NLR did not significantly improve predictive performance.
CONCLUSIONS: In this contemporary low-mortality surgical cohort, the Boey and PULP scores retained useful discriminatory ability but showed calibration mismatch when applied to outcomes beyond their original development setting. Local recalibration improved agreement between predicted and observed risks, illustrating the importance of model updating when established prediction tools are transported across populations with different baseline risks. These recalibrated models should be considered proof-of-concept updates requiring independent validation rather than immediately deployable clinical tools.