А. А. Ivshin, Yu. S. Boldina, N. A. Malyshev
Aim : to systematically compare the discrimination and calibration of eight machine learning (ML) algorithms with the Fetal Medicine Foundation (FMF) competing risks algorithm for first-trimester prediction of early-onset (< 34 weeks), late-onset ( > 34 weeks), severe, and preterm (< 37 weeks) preeclampsia (PE) by assessing the incremental value of markers with a nested model architecture (M0 [model includes only clinical and anamnestic data] → M2 [model supplemented with biophysical markers] → M4 [model supplemented with biochemical markers]) and external validation in an independent dataset (n = 7,581). Materials and Methods . A retrospective cohort study (TRIPOD 2b + 3) was conducted. Development: a development cohort (n = 7,581) with 10-fold stratified cross-validation. External validation: an independent cohort (n = 7,581). A total of 96 benchmark configurations were developed (8 algorithms × 3 predictor levels × 4 outcomes): M0 – 15 clinical and history-based factors; M2 – adding mean arterial pressure (MAP) and uterine artery pulsatility index (UtAPI); M4 – adding placental growth factor (PlGF) and pregnancy-associated plasma protein-A (PAPP-A). Evaluation metrics: AUC-ROC (Area Under the Receiver Operating Characteristic curve), calibration (O:E ratio (observed-to-expected ratio), calibration slope, Brier score), net reclassification improvement (NRI). integrated discrimination improvement (IDI), decision curve analysis (DCA), and SHapley Additive exPlanations (SHAP). Results . At the M4 level, logistic regression demonstrated the highest discrimination: early-onset PE – AUC = 0.976 (an optimistic estimate with EPV (events per variable) = 3.7; detection rate 92.9 % at 10 % FPR (false positive rate); late-onset PE – 0.837; severe PE – 0.885, preterm PE – 0.908; with the most robust calibration among all algorithms (O:E = 0.99–1.00), confirmed in the external cohort for M0 and M2 levels. Nonlinear algorithms and stacking did not achieve a significant advantage. For preterm PE (the target outcome of the FMF algorithm), logistic regression and FMF posterior were comparable in discrimination (ΔAUC = +0.002; p > 0.05) level, with the former showing superior calibration in the Russian population. For late-onset PE (the dominant disease phenotype not modeled by FMF), logistic regression provided significantly better risk stratification (AUC = 0.837 vs. 0.807; p < 0.05). The FMF algorithm exhibited miscalibration: risk overestimation for its target outcomes (O:E = 0.55 for early-onset PE) and structural underestimation for non-target outcomes (O:E = 2.84 for late-onset PE). Biophysical markers were critical for early-onset PE, whereas biochemical markers were critical for late-onset PE; this pattern was consistent across algorithms. External validation confirmed robust discrimination at the M0 level (AUC = 0.790–0.823; n = 7,581) and M2 level (AUC = 0.801–0.931; n = 4,080); results at the M4 level (n = 1,010; 11–40 events) are preliminary. Conclusion . Nonlinear ML algorithms do not outperform logistic regression, which demonstrated the most robust calibration confirmed by external validation at the M0 and M2 levels. Logistic regression with an M0 → M2 → M4 architecture is recommended for clinical decision support systems in first-trimester PE screening. The FMF algorithm requires population-specific recalibration for its target outcomes and is structurally not designed to predict late-onset PE.