科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Patterns2025-10-15· Transformer

Unraveling learning characteristics of transformer models for molecular design

Jannik P. Roth, Jürgen Bajorath

原始摘要(英文原文)· Original abstract
In drug design, transformer networks adopted from natural language processing are applied in a variety of ways. We have used sequence-based generative compound design as a model system to explore the learning characteristics of transformers and determine if these models learned information relevant for protein-ligand interactions. The analysis reveals that sequence-based predictions of active compounds using transformer models required a proportion of at least ∼60% of the original test sequences. Moreover, predictions depended on sequence and compound similarity of training and test data and on compound memorization effects. The predictions were purely statistically driven by associating sequence patterns with molecular structures, thus rationalizing their strict dependence on detectable similarities. Moreover, the transformer models did not learn target sequence information relevant for ligand binding. While the results do not call sequence-based compound design approaches generally into question, they caution against over-interpretation of transformer models used for such applications.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Unraveling learning characteristics of transformer models for molecular design — 科研速览 Science Skim