科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Nucleic acids research2026-09-22

Targeted ortholog search in unannotated genome assemblies with fDOG-Assembly.

Hannah Muelbaier, Freya Arthen, Vinh Tran, Ina Schaefer, Miklós Bálint, Ingo Ebersberger

原始摘要(英文原文)· Original abstract
Whole genome shotgun sequencing and assembly is routine. However, identifying protein-coding genes in newly assembled genomes remains complex, time-consuming, and labour-intensive. Therefore, most eukaryotic genome assemblies in public databases lack gene annotations reducing their value for evolutionary and functional genomics. Here, we present fDOG-Assembly (fDA), a novel tool for targeted, feature architecture-aware ortholog searches directly in unannotated genome assemblies. Benchmarking shows that fDA performs similarly to BUSCO and Compleasm in ortholog identification while offering the advantage of not being restricted to universal single-copy genes. Applied to identify orthologs of 5000 human genes in rat and Nematostella vectensis, fDA approaches the performance of traditional ortholog search tools that rely on pre-annotated proteomes. Importantly, it can recover orthologs missed by conventional methods because of incomplete gene annotations, helping to fill gaps in phylogenetic profiles. As a case study, we screened 176 soil invertebrate genome assemblies for genes involved in antibacterial compound production. We found that orthologs of β-lactam biosynthesis genes are widespread in springtails, with individual species possessing nearly complete cephamycin biosynthetic gene sets, suggesting they may represent previously unrecognized natural producers of β-lactam antibiotics. Overall, fDA is a powerful resource for orthology-based analyses of the rapidly growing collection of unannotated genome assemblies.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Targeted ortholog search in unannotated genome assemblies with fDOG-Assembly. — 科研速览 Science Skim