科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Journal of integrative bioinformatics2026-08-17

The DSMZ Digital Diversity Annotation Hub: a pipeline for database expansion via text mining and human curation.

Emanuel Quadros, Lorenz C Reimer, Julia Koblitz

原始摘要(英文原文)· Original abstract
Maintaining scientific databases that depend on continuous curation of research literature often requires labor-intensive, slow, and error-prone annotation processes. To address these challenges, we present a pipeline that integrates text mining with expert supervision to support database expansion. Using the BRENDA enzyme database as a case study, we compiled a relation extraction dataset by aligning document-level annotations with literature references through distant supervision. We then developed a neural model that performs entity recognition and relation classification, enabling the extraction of enzyme-strain associations from full-text articles. To close the loop between machine learning and expert curation, we designed a web-based interface that allows annotators to review and refine predicted relations. While preliminary, our initial experiments show the potential of combining weak supervision and human-in-the-loop validation to accelerate the integration of literature-derived information into knowledge bases.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

The DSMZ Digital Diversity Annotation Hub: a pipeline for database expansion via text mining and human curation. — 科研速览 Science Skim