María Lourdes Del Río-Solá, Clara de la Torre-Casaseca, Marina Jiménez-Caja, Carlos García-Padrón, Eva Álvarez-García, José Antonio Brizuela-Sanz
Using a uniform internal and temporal validation framework, gradient boosting achieved the highest numerical performance but did not demonstrate a conclusive advantage over penalised logistic regression. The MLP provided no incremental predictive benefit. These findings indicate that increasing algorithmic complexity does not necessarily improve mortality prediction in modest-sized emergency vascular datasets and support interpretable regression as an essential benchmark. External multicentre validation and more complete prospective data collection are required before clinical implementation.
BACKGROUND: Ruptured abdominal aortic aneurysm (rAAA) remains associated with substantial in-hospital mortality. Although machine-learning methods can model complex nonlinear relationships between admission characteristics and outcome, their incremental value over conventional regression remains uncertain. We compared three prediction models using identical admission variables and a uniform validation framework.
METHODS: This retrospective single-centre cohort included 196 unique patients with rAAA managed between 2006 and 2024. Five prespecified admission predictors were evaluated: age, hypovolaemic shock, maximal aneurysm diameter, haemoglobin and systolic blood pressure. Missing data were imputed independently within each training partition. Penalised logistic regression, gradient boosting and a multilayer perceptron (MLP) were compared using nested five-fold cross-validation repeated five times. Secondary temporal validation was performed in the 181 patients with a recoverable treatment year: models were developed in the 2006-2018 cohort (n=110) and evaluated, without refitting or recalibration, in the 2019-2024 cohort (n=71).
RESULTS: In-hospital mortality occurred in 99/196 patients (50.5%). Gradient boosting and penalised logistic regression achieved comparable moderate discrimination, with AUCs of 0.739 (95% CI 0.666-0.805) and 0.729 (95% CI 0.654-0.796), respectively. Their paired AUC difference was 0.010 (95% CI -0.024 to 0.045). The MLP showed lower discrimination (AUC 0.643, 95% CI 0.563-0.719), with a paired difference versus penalised logistic regression of -0.086 (95% CI -0.160 to -0.013). In temporal validation, AUCs were 0.800 (95% CI 0.686-0.904) for gradient boosting, 0.739 (95% CI 0.607-0.858) for penalised logistic regression and 0.632 (95% CI 0.496-0.762) for the MLP. Temporal calibration demonstrated overprediction of absolute mortality risk, with observed mortality of 40.8% compared with mean predicted mortality of 49.9%, 50.0%, and 63.4%, respectively.
CONCLUSIONS: Using a uniform internal and temporal validation framework, gradient boosting achieved the highest numerical performance but did not demonstrate a conclusive advantage over penalised logistic regression. The MLP provided no incremental predictive benefit. These findings indicate that increasing algorithmic complexity does not necessarily improve mortality prediction in modest-sized emergency vascular datasets and support interpretable regression as an essential benchmark. External multicentre validation and more complete prospective data collection are required before clinical implementation.