科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ PloS one2026-01-01

Reproducible benchmarking of four machine-learning classifiers on a large synthetic heart-disease competition dataset.

Raed M Ennab

一句话结论 · In one sentence

All four models discriminated strongly and showed close internal calibration, while gradient-boosting models had only marginally higher AUROC than logistic regression. The principal contribution is a fully rerunnable, controlled benchmark rather than a novel algorithm or a clinically validated prediction model. External performance in real clinical cohorts remains unknown.

原始摘要(英文原文)· Original abstract
BACKGROUND: Public synthetic datasets enable transparent model comparison but cannot establish performance in clinical populations. We asked how four fixed classifiers compare under one leakage-controlled validation framework on a large synthetic heart-disease competition dataset. METHODS: The official training set contained 630,000 observations (44.83% outcome prevalence) and 13 predictors; the 270,000-row test set had no outcome labels. L2-regularized logistic regression, HistGradientBoosting, CatBoost, and LightGBM were evaluated using identical stratified five-fold outer splits. Logistic-regression preprocessing was fitted within outer-training data, and early stopping used outer-training data only. The primary measure was pooled out-of-fold (OOF) AUROC; secondary measures were average precision, Brier score, calibration intercept and slope, fold variability, and train-validation optimism. Duplicate/overlap, identifier-only, and permuted-outcome controls assessed leakage. RESULTS: Pooled OOF AUROC was 0.955443 for CatBoost, 0.955324 for LightGBM, 0.955066 for HistGradientBoosting, and 0.952874 for logistic regression. Corresponding average precision ranged from 0.945912 to 0.948828 and Brier scores from 0.081116 to 0.083518. Calibration intercepts ranged from -0.001299 to 0.002153 and slopes from 0.998896 to 1.004388. Mean optimism was 0.000017-0.002686. No duplicated identifiers, duplicated feature profiles, or train-test feature-profile overlap were found; identifier-only AUROC was 0.500121 and mean permuted-outcome AUROC was 0.500615. CONCLUSIONS: All four models discriminated strongly and showed close internal calibration, while gradient-boosting models had only marginally higher AUROC than logistic regression. The principal contribution is a fully rerunnable, controlled benchmark rather than a novel algorithm or a clinically validated prediction model. External performance in real clinical cohorts remains unknown.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Reproducible benchmarking of four machine-learning classifiers on a large synthetic heart-disease competition dataset. — 科研速览 Science Skim