科研速览 · Science Skim继续刷下去 · Keep skimming →
◇ bioRxiv2026-08-17· physiology

Towards a Physiological Scaling Law: Model Quality vs. Cohort Size for Stochastic Sequence Data

G. Sunil, B. Ravikumar, B. Ramsundar, S. Subramanian

原始摘要(英文原文)· Original abstract
Scaling laws help determine the optimal data size for training large models but are established in domains where the target is deterministic. Physiological signals are different: heartbeat sequences are stochastic, so part of the error is irreducible even with large amounts of data. Metrics such as MAE do not account for non-deterministic behavior, and therefore assessing scaling requires evaluating distributional calibration (measuring how well predicted probability densities capture true conditional characteristics). We formulate a scaling law metric(n) = E + A*n^- and evaluate it with five metrics: accuracy (MAE, RMSE), distributional calibration (KS distance, goodness-of-fit), and training objective (negative log loss) using a neural temporal point process trained on a cohort of four-ECG datasets. The law fits all five metrics. While point accuracy is near saturation at n=183, KS distance and goodness-of-fit improve by 6% and 12% respectively when extrapolated to 10,000 subjects, showing that scaling decisions in stochastic domains must be guided by distributional calibration rather than point accuracy.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Towards a Physiological Scaling Law: Model Quality vs. Cohort Size for Stochastic Sequence Data — 科研速览 Science Skim