科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Nucleic Acids Research2026-03-19· Biology

A foundation model for nucleotide sequences

Xilin Shen, Jie Li, Meng Yang, Lei Shi, Kexin Chen, Xiangchun Li

原始摘要(英文原文)· Original abstract
Foundation models have demonstrated exceptional performance across diverse downstream tasks. However, in genomics and transcriptomics, the integration of nucleotide sequences with their rich annotations remains underexplored, potentially limiting model generalizability across species and biological contexts. Here, we introduce OmniNA (Omni-applicable foundation model for nucleic acid), a self-supervised generative foundation model trained on 91.7 million nucleotide sequences and their associated annotations, totaling 1076.2 billion bases and 197 million words spanning diverse species. Unlike most existing approaches that focus solely on functional genomic element recognition, OmniNA jointly learns from both sequences and annotations, leveraging their complementarity to enhance semantic understanding and representation learning. We demonstrate that OmniNA captures sequence grammar and annotation semantics, facilitating robust transfer across nucleotide-level tasks. OmniNA can be fine-tuned under natural language paradigms and achieves state-of-the-art or competitive performance in 23 benchmarks, including sequence detection and species classification. Its learned representations also help reveal mutation effects on DNA and RNA processing. We release the model publicly as a community resource for genomics and transcriptomics research. OmniNA represents a step toward advancing foundational modeling by integrating annotation-aware learning, offering a powerful tool for genomics and transcriptomics research.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

A foundation model for nucleotide sequences — 科研速览 Science Skim