科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Bioinformatics advances2026-01-01

Evaluating transformer-based models for structural characterization of orphan proteins.

Ercan Seçkin, Dominique Colinet, Etienne G J Danchin, Edoardo Sarti

一句话结论 · In one sentence

We compared predictions from several widely used TBM architectures on an expert-curated set of orphan proteins from the Meloidogyne genus, comprising some of the most destructive plant-parasitic nematodes. Multiple sequence alignment-based approaches such as AlphaFold2 performed poorly on orphan proteins, as did single-sequence or embedding-based language models ESMFold, OmegaFold, and ProtT5. This limited performance cannot be fully attributed to intrinsic disorder, as confirmed by independent non-TBM disorder predictors. While accurate tertiary structure prediction remains out of reach, secondary structure is more reliably captured: predictors share about 70% of secondary structure elements, regardless of global fold similarity, and these elements are consistently identified by dedicated secondary structure tools.

原始摘要(英文原文)· Original abstract
MOTIVATION: Transformer-based models (TBMs) are state-of-the-art deep learning architectures that predict protein structural features with high accuracy. Despite methodological differences, they all rely on large datasets structured in families of homologous sequences. However, 5%-30% of eukaryotic proteomes consist of orphan proteins, which are sequences without detectable similarity to known families. Although they may share structural traits with characterized proteins, their lack of homology makes them an ideal dataset for evaluating TBM generalization beyond familiar sequence space. RESULTS: We compared predictions from several widely used TBM architectures on an expert-curated set of orphan proteins from the Meloidogyne genus, comprising some of the most destructive plant-parasitic nematodes. Multiple sequence alignment-based approaches such as AlphaFold2 performed poorly on orphan proteins, as did single-sequence or embedding-based language models ESMFold, OmegaFold, and ProtT5. This limited performance cannot be fully attributed to intrinsic disorder, as confirmed by independent non-TBM disorder predictors. While accurate tertiary structure prediction remains out of reach, secondary structure is more reliably captured: predictors share about 70% of secondary structure elements, regardless of global fold similarity, and these elements are consistently identified by dedicated secondary structure tools. AVAILABILITY: All data and analysis scripts are available at https://doi.org/10.5281/zenodo.18788931.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Evaluating transformer-based models for structural characterization of orphan proteins. — 科研速览 Science Skim