科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Peer Community Journal2026-06-17· Sequence (biology)

Accelerating k-mer-based sequence filtering

Igor Martayan, Léa Vandamme, Bede Constantinides, Bastien Cazaux, Charles Paperman, Antoine Limasset

原始摘要(英文原文)· Original abstract
Abstract Motivation The exponential growth of global sequencing data repositories presents both analytical challenges and opportunities. While k -mer-based indexing has improved scalability over traditional alignment for identifying relevant documents, pinpointing the exact sequences matching numerous queries remains a hurdle. In particular, searching for numerous k -mers with a single large query or multiple distinct queries strains existing exact matching tools, whose performance scales poorly with an increasing number of patterns. At the same time, indexing entire vast datasets for infrequent or ad-hoc searches is often resource-prohibitive. Designing fast methods for matching a large number of k -mers without exhaustive pre-indexing is therefore critical. Contributions We propose an efficient solution to the problem of k-mer-based sequence filtering : given a set of k -mers of interests and a threshold, quickly evaluate whether an arbitrary sequence has a number of k -mer matches above or below the threshold. Our approach demonstrates how minimizer-based based sketching, alongside SIMD acceleration, can enhance the performance of streaming searches, and is implemented as a Rust tool named K2Rmini . On a consumer laptop, K2Rmini is able to filter long reads at 2 Gbp/s. Availability https://github.com/Malfoy/K2Rmini .
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Accelerating k-mer-based sequence filtering — 科研速览 Science Skim