科研速览继续刷下去 →
◆ Nature Methods2026-03-30· Computer science

Compressing the collective knowledge of ESM into a single protein language model

Tuan Dinh, Seon-Kyeong Jang, Noah Zaitlen, Vasilis Ntranos

一句话结论

We allow individual PLMs to self-improve by distilling the most confident predictions from multiple models of the same family and demonstrate that co-distillation of ESM models suffices to achieve state-of-the-art performance across multiple VEP benchmarks.

原始摘要(原文)
Protein language models (PLMs) have recently emerged as a promising approach for next-generation variant-effect prediction (VEP). Most high-performing VEP methods currently utilize PLMs combined with additional information, such as homology, protein structure and population genetics data to improve prediction accuracy. This performance gain, however, comes with added complexity or limited applicability compared to pure PLMs trained only on raw, unaligned sequences, such as evolutionary scale modeling (ESM). Here we challenge the prevailing view that sequence-only PLMs are intrinsically limited and present an efficient co-distillation approach to adapt them for high-accuracy VEP without requiring additional information beyond evolutionary signals captured during pretraining. We allow individual PLMs to self-improve by distilling the most confident predictions from multiple models of the same family and demonstrate that co-distillation of ESM models suffices to achieve state-of-the-art performance across multiple VEP benchmarks. We further show that this performance increase enables accurate quantification of the severity of variant effects on continuous clinical phenotypes in biobank data.
读原文 ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文

Compressing the collective knowledge of ESM into a single protein language model — 科研速览 Science Skim