科研速览 · Science Skim继续刷下去 · Keep skimming →
◇ bioRxiv2026-09-11· microbiology

Corpusome, a cross-body-site human microbiome corpus for representation learning

H. Xuan, Y. Huang, J. Bian

原始摘要(英文原文)· Original abstract
Machine-learning models of the human microbiome are trained mostly on stool samples from single cohorts, limiting cross-body-site representation and cross-study generalization. Progress is constrained less by algorithms than by the absence of a harmonized multi-body-site corpus carrying the technical metadata needed to model, rather than ignore, batch structure. Here we release Corpusome, a harmonized two-tier cross-body-site human microbiome corpus for representation learning: a harmonized corpus of 187,546 human microbiome samples integrating standardized profiles from curatedMetagenomicData, the American Gut Project, and the EBI MGnify platform. Corpusome follows a two-tier design preserving both functional depth and cross-body-site breadth: a shotgun tier (22,588 samples, 93 studies) with species- and pathway-level profiles, and a 16S tier (164,958 samples, from a full pull of 708 MGnify studies) with genus-level profiles extending coverage to oral, skin, respiratory, and urogenital sites. It spans six body sites and two modalities, with harmonized metadata for batch-aware modelling. Body-site signal exceeds technical/source variance in the 16S tier by approximately 2.4-fold.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Corpusome, a cross-body-site human microbiome corpus for representation learning — 科研速览 Science Skim