科研速览 · Science Skim继续刷下去 · Keep skimming →
◇ bioRxiv2026-09-10· bioinformatics

Spurious model comparisons are widespread in biomedical artificial intelligence

T. Zeng, H. Li, S. Zhang, Y. Q. Tan, F. Tian, C. Orban, L. An, W. Che, J. Cheng, J. S. X. Chong, N. Dehestani, Z. Dong, X. Li, Z. Li, M. J. R. Lim, Y. Lin, Q. Ling, Z. Ling, X. Z. Low, S. Mansour L., K. K. Ng, T. T. Nguyen, L. Q. R. Ooi, S. Pande, X. Qian, J. Ruan, Z. Wang, Y. Xie, C. Zhang, Y. Zhang, K. Patil, L. Parkes, E. Dhamala, S. Chopra, A. Zalesky, A. J. Holmes, S. Eickhoff, J. H. Zhou, O. Renaud, N. Dosenbach, K. P. Kording, D. Bzdok, T. Nichols, B. T. T. Yeo

原始摘要(英文原文)· Original abstract
Cross-validation is routinely used to compare performance in biomedical artificial intelligence. Standard tests ignore correlation across cross-validation folds, inflating false-positive rates. In a PRISMA-guided meta-analysis of 184 studies (impact factor [≥] 15) across 30 biomedical fields, 97% use invalid tests. Among studies with abstract-level claims supported by invalid tests, 59% rely on a spurious comparison: one that loses significance after correcting for this correlation. On average, regaining significance requires a 43% larger effect. Extending these findings to the broader literature, we estimate that spurious comparisons occur in nearly two in three studies and support abstract-level claims in nearly one in three studies. Simulations confirm false-positive rates paradoxically approach 100% when cross-validation is repeated to improve stability. We introduce SHARP, a redesign of cross-validation, which best balances false-positive control and power among 13 benchmarked tests. These results reveal widespread fragility in biomedical artificial intelligence and offer a practical route to valid comparisons.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Spurious model comparisons are widespread in biomedical artificial intelligence — 科研速览 Science Skim