Maria Rosario Borroek, Bambang Tutuko, Ermatita, Jasmir
Diabetes mellitus (DM) and pre-diabetes (PD) represent a growing global health burden, yet classification from population-based survey data remains challenging because of class imbalance, overlapping risk profiles, and evaluation protocols that may overestimate predictive performance.This study presents a transparent, leakage-free classification framework based on duplicate removal, body mass index (BMI) outlier filtering, composite cardiovascular risk-burden stratification, and analysis of variance (ANOVA) feature selection.Twelve machinelearning classifiers were evaluated on two complementary cohorts: the complete deduplicated cohort (primary evaluation) and a risk-burden-stratified sub-population (sensitivity analysis), across four diabetes classification scenarios using macro-averaged metrics and one-versus-rest (OvR) receiver operating characteristic area under the curve (ROC-AUC).On the complete deduplicated cohort, which serves as the primary evaluation dataset in this study, classification accuracy was substantially lower than that observed on the risk-burden-stratified cohort, indicating that realistic population-wide diabetes classification remains considerably more challenging than suggested by studies using cohort-selection strategies.Results obtained from the risk-burden-stratified cohort are therefore presented only as a sensitivity analysis to quantify the influence of cohort construction.Under identical experimental settings, the stratified cohort yielded higher accuracy, with Extra Trees (ET) achieving 0.9905 for ND vs. PD and 0.9725 for ND vs. DM, compared with 0.6285 and 0.6975 on the complete deduplicated cohort.Re-implementation of representative literature baselines under identical preprocessing, train-test partitions, thresholds, and evaluation metrics demonstrated practically equivalent classification accuracy, indicating that previously reported improvements are largely attributable to differences in cohort construction rather than classifier design.Robustness was assessed using repeated stratified cross-validation (CV), bootstrap confidence intervals (CIs), Friedman and Nemenyi statistical tests, calibration analysis, decision-curve analysis (DCA), and external validation on the National Health and Nutrition Examination Survey (NHANES) 2013-2014 dataset.External validation suggested substantially lower transfer performance than internal testing, reinforcing the importance of cohort definition when interpreting reported classification accuracy.These findings position the proposed framework as a retrospective methodological investigation for quantifying cohort effects rather than as evidence of immediate clinical screening readiness.