科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Journal of Sports Analytics2026-08-01· Software deployment

Synthetic data generation in sport: A framework for design, privacy, utility, fidelity, and deployment

Ruiz Pinto Paul, Paul Pao‐Yen Wu, John Warmenhoven, Kerrie Mengersen, Divya Mehta

原始摘要(英文原文)· Original abstract
The increasing use of synthetic data in sports science reflects persistent limitations in real-world sport datasets, including small sample sizes, class imbalance, privacy restrictions, and the absence of ground truth. While synthetic data generation has been applied across a wide range of sporting contexts, its use remains heterogeneous, with limited consistency in design, evaluation, and reporting practices. This study synthesises evidence from twelve sport studies to address a central gap in the literature: the lack of an established, domain-specific framework to guide the systematic generation, evaluation, and deployment of synthetic data in sport. Rather than proposing a new generative model, this work introduces a decision-oriented framework that structures synthetic data projects across six interrelated dimensions: objective of use, data structure, generation strategy, domain constraints, utility and fidelity evaluation, and deployment risk. Analysis of the reviewed studies shows that synthetic data are most effective when generation strategies are aligned with sport-specific domain knowledge and clearly defined operational goals. The framework highlights recurring challenges related to domain shift, constrained realism, and class imbalance, and demonstrates how these issues can be addressed through explicit design and evaluation choices.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Synthetic data generation in sport: A framework for design, privacy, utility, fidelity, and deployment — 科研速览 Science Skim