科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ The journal of physical chemistry letters2026-08-27

Zipf-like Statistical Regularities in Molecular Sequence Representations for Chemical Language Models.

Anyu Liu, Chao Fang, Yuntao Li, Zongguo Wang, Tao Qi, Guoping Hu

原始摘要(英文原文)· Original abstract
Large language model (LLM)-based approaches increasingly use molecular strings such as SMILES and SELFIES for molecular generation and property prediction. However, the statistical properties of molecular token distributions have not been systematically characterized. Here, we analyzed rank-frequency distributions of tokens derived by byte-pair encoding (BPE) across large molecular databases. BPE-derived tokens showed approximate Zipf-like rank-frequency scaling across the examined representations and chemical spaces, with fitted slopes moderately steeper than the canonical value of -1, paralleling a statistical pattern widely observed in natural language. Moreover, when BERT models were pretrained using BPE vocabularies with different rank-frequency slopes, the closeness of these slopes to the ideal Zipf value of -1 strongly correlated with performance on molecular property prediction tasks (Pearson r = 0.91, p < 0.001) and remained associated after adjustment for vocabulary size (partial r = 0.86, p < 0.001). Together, these findings show that molecular BPE vocabularies exhibit an approximate Zipf-like rank-frequency regularity and that slope closeness provides an empirical diagnostic for comparing vocabulary sizes within the examined SMILES/SELFIES BPE framework.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Zipf-like Statistical Regularities in Molecular Sequence Representations for Chemical Language Models. — 科研速览 Science Skim