Rachel A Hadler, Sahithi Krishnaveni Lakamana, Tristan Moorman, Selen Bozkurt
SafeNET's performance was reproducible in an external health system, supporting its portability. However, local models trained on system-specific electronic health record patterns achieved markedly higher accuracy and calibration. Real-world replication revealed major data-availability barriers and highlighted the tension between acceptable recall and suboptimal precision in critical-care prognostic applications.
BACKGROUND: Mortality risk is often uncertain at the time of interhospital transfer (IHT).
OBJECTIVES: We examined what happens when a published machine learning model for post-transfer mortality, SafeNET, is deployed in an independent health system with substantial data fragmentation and missingness and assessed the extent to which local model development can compensate for cross-system data limitations.
DERIVATION COHORT: SafeNET was evaluated from published specifications using its original 14-variable feature schema and applied without modification. Locally trained gradient-boosting models (Categorical Boosting [CatBoost], Light Gradient Boosting, Extreme Gradient Boosting) were developed using a stratified 70/30 train-test split with five-fold cross-validation, class weighting, and random under-sampling to address class imbalance (~1:27).
VALIDATION COHORT: We included 14,728 adult patients (≥ 18 yr) undergoing IHT at a multihospital academic health system from January 1, 2023, to December 31, 2023. Functional status variables were missing in 85% of encounters, limiting direct implementation of the original feature set.
PREDICTION MODEL: Discrimination, classification metrics (precision, recall, F1 score), and calibration (Brier score) were compared across SafeNET and locally trained models on the held-out test set.
RESULTS: SafeNET achieved performance that closely approximated its development-site results (area under the receiver operating characteristic curve [AUC-ROC], 0.87; F1, 0.73; 95% CI, 0.69-0.77), indicating strong external consistency. Locally trained CatBoost models outperformed SafeNET (AUC-ROC, 0.92; F1, 0.74; 95% CI, 0.72-0.78) and demonstrated improved calibration (Brier 0.04 vs. 0.055). Key variables, including functional status, were missing in 85% of encounters, limiting direct implementation of the original feature set. Although recall exceeded 0.70 across most models, precision remained below 080 for several configurations, indicating important trade-offs relevant to clinical deployment.
CONCLUSIONS: SafeNET's performance was reproducible in an external health system, supporting its portability. However, local models trained on system-specific electronic health record patterns achieved markedly higher accuracy and calibration. Real-world replication revealed major data-availability barriers and highlighted the tension between acceptable recall and suboptimal precision in critical-care prognostic applications.