科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ npj Drug Discovery.2026-08-04· Scaling

Diversity Beats Size Scaling for Chemical Language Models

Borja Medina, Alessandro Tibo, Jiazhen He, Jon Paul Janet, Nicklas Österbacka

原始摘要(英文原文)· Original abstract
Chemical language models, such as transformers trained on SMILES strings, are increasingly used in drug design and have seen rapid growth in both model capacity and training dataset size. The impact of this scaling on practical downstream performance remains unclear, however. We systematically evaluate how model size and dataset size affect encoder-decoder transformers trained on paired textual molecular representations. We find that, beyond a minimal threshold, further model scaling yields no gain in hit generation rate, while dataset scaling gives diminishing returns. We further introduce a dataset diversification strategy that substantially increases hit diversity. These results suggest that, for molecular hit discovery, data curation and diversity may be more impactful than continued scaling of model size or dataset volume, and they motivate a shift from scale-first to diversity-first training paradigms.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Diversity Beats Size Scaling for Chemical Language Models — 科研速览 Science Skim