科研速览 · Science Skim继续刷下去 · Keep skimming →
◇ medRxiv2026-09-14· health informatics

AI-Based Synthetic Data in Biomedicine: A Decade of Growth and a Persistent Translation Gap

N. Asgari, I. F. Perez, G. Epelde, L. Zhang, L. Horesh, C. Saab, M. Muszkat, M. Rosen-Zvi

原始摘要(英文原文)· Original abstract
AI-generated synthetic data are increasingly used to address data scarcity, privacy constraints and experimental limitations in biomedicine, but how far these methods have translated into practice remains unclear. We conducted a systematic mapping and bibliometric analysis of 4,143 publications spanning 2015-2025, combining expert annotation with LLM-assisted classification across data modality, medical domain, paper type, deployment status and research stance. Publication volume grew continuously; 77.8% of papers were strongly supportive while critical work remained below 1%. Medical imaging dominated the corpus, consistent with well-characterized transformation-group invariances supporting data augmentation and generative modeling. Highly cited primary research concentrated disproportionately in molecular and pharmaceutical applications, where SE(3)-equivariant architectures and structure-prediction models accelerated generative approaches. Only 27 publications reported operational use; omics and tabular clinical data, lacking well-characterized invariance structures, remained underrepresented. These findings reveal a gap between methodological growth and deployment, motivating investment in evaluation standards, deployment reporting and encoding domain-relevant invariances.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

AI-Based Synthetic Data in Biomedicine: A Decade of Growth and a Persistent Translation Gap — 科研速览 Science Skim