科研速览 · Science Skim继续刷下去 · Keep skimming →
◇ bioRxiv2026-09-10· microbiology

The StrainDiscoveryDatabase: an open framework for standardized microbial strain data

J. F. Witte, A. Lissin, I. Schober, C. Ebeling, H. Lueken, J. Koblitz, A. Yurkov, B. Bunk, J. M. Lopez-Coronado, A. Vaello, A. Zuzuarregui, R. Aznar, A. Garcia-Donate, A. M. P. Melo, G. Verkleij, R. P. de Vries, V. Robert, W. Meyer, J. Overmann, L. C. Reimer

原始摘要(英文原文)· Original abstract
The vast amount of existing data on microbial strains holds immense potential to revolutionize bioindustry through the application of Artificial Intelligence (AI). However, the training of robust predictive AI models requires large-scale, unified, and non-redundant microbial datasets, which is currently severely hindered by the deep fragmentation of the data and the existence of synonymous strain identifiers in different culture collections. To overcome these infrastructural bottlenecks, we have established the StrainDiscoveryDatabase (SDD), a comprehensive, machine-readable dataset encompassing over 6.2 million harmonized data points for 256,889 microbial strains. The SDD does not rely on its own data repository, but rather on existing data that is retrieved on the fly from highly curated databases. Through an automated pipeline phenotypic, genotypic, and contextual data are systematically retrieved via the Application Programming Interfaces (APIs) of the Bacterial Diversity database (BacDive), the Microbial Resource Research Infrastructure Information System (MIRRI-IS) and the catalogue of the DSMZ. In order to reliably resolve synonymous strain identifiers, the StrainInfo database and its identification tools are employed, enabling the accurate deduplication and unification of records from disparate sources. The resulting aggregated, globally unique dataset is provided in a highly standardized JSON format in strict adherence to the FAIR data principles. By bridging isolated database silos and linking distributed knowledge to discrete biological entities, the SDD provides a high-quality, foundational resource designed to accelerate trait-based strain discovery, large-scale comparative analysis, and machine learning applications, thereby supporting the translation of the extensive existing knowledge on microbial traits to bioindustrial applications.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

The StrainDiscoveryDatabase: an open framework for standardized microbial strain data — 科研速览 Science Skim