科研速览 · Science Skim继续刷下去 · Keep skimming →
◇ bioRxiv2026-09-06· bioinformatics

ContrasTED: contrastive domain embeddings for scalable remote homology classification

D. M. Miller, N. Bordin, J. Jeyananthan, V. Waman, M. Heinzinger, C. Orengo

原始摘要(英文原文)· Original abstract
Protein structure prediction has expanded structural databases to hundreds of millions of domains. Classifying these domains into homologous superfamilies reveals evolutionary and functional relationships that can persist despite low sequence similarity. As the size of structural databases continues to grow, homology classification requires methods that combine scalability with accuracy. Here we present ContrasTED, which uses CATH-supervised center-contrastive learning to project structure-aware embeddings into a domain-level metric space for nearest-centroid superfamily assignment. On a sequence-filtered S20 benchmark (n = 1,028), superfamily assignment accuracy reached 92.9% (1-NN) and 91.4% (nearest centroid), exceeding sequence search, profile HMMs, Foldseek, and a classifier trained on embeddings. The learned latent space separates superfamilies while retaining structural information below 20% sequence identity, with the largest gains among sparsely represented superfamilies. ContrasTED produces 4.67 million new candidate assignments across 3,796 superfamilies in The Encyclopedia of Domains (TED), extending annotation coverage beyond previous structure-based methods.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

ContrasTED: contrastive domain embeddings for scalable remote homology classification — 科研速览 Science Skim