科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Ecotoxicology (London, England)2026-09-01

Transfer learning for honey bee toxicity prediction: MolFormer versus classical QSAR representations.

Alan Victor de Souza Pinho, Edilson Beserra de Alencar Filho, Rosalvo Ferreira de Oliveira Neto

原始摘要(英文原文)· Original abstract
Honey bee (Apis mellifera) toxicity assessment is essential for environmental risk evaluation of agrochemicals, yet experimental data are costly and often limited, constraining the development of robust predictive models. This study investigates whether transfer learning via pretrained chemical language models can provide competitive QSAR performance for bee toxicity classification. Using the ApisTox dataset (1,035 compounds; 296 toxic and 739 non-toxic for bees), molecular representations were derived from SMILES in three ways: (i) PaDEL molecular descriptors, (ii) RDKit Morgan fingerprints, and (iii) MolFormer embeddings extracted from a publicly available reduced-scale pretrained checkpoint (~ 100 M molecules; ~10% ZINC + ~ 10% PubChem), used as frozen features. Three classical classifiers: Random Forest, Support Vector Machine, and Multilayer Perceptron were trained and evaluated under 5-fold cross-validation. Fingerprints paired with RF achieved the best overall discrimination, as measured by the Area Under the Receiver Operating Characteristic (ROC-AUC = 0.866). Importantly, MolFormer embeddings combined with SVM reached near-parity (ROC-AUC = 0.859) and consistently outperformed PaDEL descriptors for all classifiers (ΔROC-AUC = + 0.005 to + 0.021). These results demonstrate that transfer-learned chemical embeddings can rival established QSAR baselines while simplifying feature engineering, supporting their practical adoption for ecotoxicological screening under limited labeled data.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Transfer learning for honey bee toxicity prediction: MolFormer versus classical QSAR representations. — 科研速览 Science Skim