科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ BMC Bioinformatics2026-06-16· De Bruijn sequence

GnnDebugger: GNN based error correction in De Bruijn Graphs

Marijo Simunovic, Lovro Vrček, Mile Šikić, Anton Bankevich

原始摘要(英文原文)· Original abstract
BACKGROUND: Modern sequencing technologies have enabled the reconstruction of complete mammalian genomes from telomere to telomere. However, scaling this achievement to thousands of species and population-level studies remains a challenge. Key bottlenecks include the low quality of the draft assemblies and the high coverage requirements. In particular, reconstructing complete and accurate sequences of both haplotypes in diploid genomes is especially difficult since the sequencing depth is not always sufficient to properly reconstruct diverged regions. We aim to explore the use of machine learning, specifically graph neural networks, for scalable error correction in De Bruijn Graphs, addressing the limitations of existing heuristic methods in genome assembly. RESULTS: Inspired by the success of neural networks in extracting patterns from the data on a massive scale, we introduce a method for correcting errors in De Bruijn Graphs using Graph Neural Networks. Our model provides a reliable classification of edges into correct and erroneous, especially for diploid genomes with coverage depth 35 and lower. We demonstrate that these predictions can guide the downstream read error correction algorithm and genome assembly, ultimately allowing for more accurate genome assembly. CONCLUSIONS: Machine learning methods have the potential to replace heuristic methods commonly used in genome assembly. Learning-based approaches can enhance the performance of existing assemblers in challenging scenarios and facilitate adaptation to newly sequenced species.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

GnnDebugger: GNN based error correction in De Bruijn Graphs — 科研速览 Science Skim