科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Computers in biology and medicine2026-09-26

Duplicate leakage and evaluation integrity on a public synthetic sleep-health tabular dataset: A methodological case study.

Pramod K B Rangaiah, B P Pradeep Kumar, Robin Augustine

原始摘要(英文原文)· Original abstract
Public tabular datasets are widely reused to benchmark machine learning for health-related screening, but their evaluation validity is rarely audited. Using the widely cited, synthetic Sleep Health and Lifestyle Dataset (374 records, three classes) as a case study, we show that the very high accuracies reported for it reflect evaluation optimism rather than a property of the prediction task: several published studies report accuracies approaching 98%, whereas our reproduction under conventional cross-validation reaches approximately 91%, and confining duplicate records to a single fold reduces this further. Our central finding is duplicate-record leakage: only 132 of the 374 rows are distinct, 64.7% of rows are exact copies of another row, and the conventional cross-validation used in prior work allows identical rows to fall on both sides of the train/test split. When duplicate records are confined to a single fold using group-aware cross-validation, a gradient-boosted model drops from 91.2% to 72.6%, and on fully de-duplicated data to 61.7%: standard cross-validation overestimates performance by 18.6 percentage points relative to duplicate-group-aware evaluation, and performance falls by a further 10.9 percentage points after full de-duplication, reflecting the reduced effective sample size and conflicting-label structure. Under the corrected protocol we benchmark seven classifiers with correlation-corrected (Nadeau-Bengio) confidence intervals, nested cross-validation for hyperparameter selection, and effect-size-oriented significance testing. Simple models (logistic regression, support vector machine, multilayer perceptron) and CatBoost remain robust (≈87-91%), whereas the untuned random forest, XGBoost, and LightGBM show substantial variability, with very wide intervals; after nested tuning they partially recover, underscoring that model rankings on this dataset are not reliable. As a secondary, deliberately negative result, a feature-graph attention network we designed for the task (with a learnable sparse graph and a residual gate) does not outperform a plain multilayer perceptron (full model 88.2%±4.2 vs 89.3%±3.4; paired Wilcoxon p=0.03), indicating that graph inductive biases do not provide measurable benefit for this small static tabular dataset. We also report a near-deterministic association between the Occupation field and the label (Cramér's V=0.75; an 85% no-learning lookup), which we treat as suggestive of a data-construction artefact rather than as proven leakage. We conclude, in line with recent leakage-aware work in medical imaging, that this dataset is unsuitable for strong claims about sleep-disorder screening unless duplicate structure and synthetic-construction artefacts are handled explicitly; the contribution is a reusable, duplicate-aware benchmarking protocol rather than a clinical result.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Duplicate leakage and evaluation integrity on a public synthetic sleep-health tabular dataset: A methodological case study. — 科研速览 Science Skim