科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Journal of electrocardiology2026-09-16

Why blinding sex is not enough: equitable discrimination, unequal operation in machine-learning detection of myocardial infarction from the ECG.

Guilherme Cesar Fernandes de Oliveira da Costa, Jose Gregorio Valero Rodriguez, Giulia Fernandes de Oliveira da Costa, Erivelton Alessandro do Nascimento

一句话结论 · In one sentence

Discrimination and calibration showed no detectable between-sex difference, yet the 0.5 threshold produced lower recall in women, most markedly in younger women, a descriptive and exploratory finding tied to the operating point. Fairness is best verified per subgroup and in external validation.

原始摘要(英文原文)· Original abstract
BACKGROUND AND PURPOSE: Automated electrocardiogram (ECG) interpretation is proposed to standardise myocardial infarction (MI) detection, but whether such models reproduce under-diagnosis of MI in women is unclear. METHODS: Sex- and age-related equity in MI-versus-normal detection was audited on PTB-XL as a benchmark classification audit (14,437 ECGs, 12,940 patients; MI-labelled record proportion 0.296 in women, 0.455 in men), across four models sharing a patient-grouped, hash-frozen 70/15/15 split: interpretable gradient-boosted trees with and without sex and age, and a raw-signal convolutional network under two normalisation schemes. Discrimination was separated from the operating point; bootstrap intervals were resampled by patient. RESULTS: Discrimination by sex and by age showed no statistically detectable between-sex difference in any model (area under the ROC curve, AUC, 0.96 to 0.98; DeLong non-significant), and no male-minus-female contrast in calibration intercept or slope survived correction. At the default 0.5 threshold recall was lower in women, the disparity declined continuously with age in exploratory analysis, reaching 0.176 and 0.240 below 60 years. Female median scores were lower than male ones within both true classes (Holm-adjusted p 0.002 to 0.038), a calibrated shift tracking the smaller female class proportion. Amplitude features predicted sex collectively (AUC 0.845). Sex-specific thresholding narrowed the gap but gave no stable younger-female cut. CONCLUSIONS: Discrimination and calibration showed no detectable between-sex difference, yet the 0.5 threshold produced lower recall in women, most markedly in younger women, a descriptive and exploratory finding tied to the operating point. Fairness is best verified per subgroup and in external validation.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Why blinding sex is not enough: equitable discrimination, unequal operation in machine-learning detection of myocardial infarction from the ECG. — 科研速览 Science Skim