科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Journal of chemical information and modeling2026-09-14

MolSelector: A Machine Learning Framework for Interpretable Subset Selection from Molecular Science Data Sets.

Caitlin Whitter, Aurora Evelyn Clark, Alex Pothen, Rajiv Khanna

原始摘要(英文原文)· Original abstract
We describe MolSelector, a machine learning framework for selecting representative subsets of molecular science data sets for efficient and accurate neural network training. As part of MolSelector, we introduce the interpretable Atypicality Score algorithm, which identifies molecules that are typical versus atypical of their data set and selects a representative sample of the data set on that basis. These subsets can be used for subsequent neural network training, leading to fast model training with reduced memory requirements while still maintaining low test set error across several molecular property prediction tasks. Our experiments demonstrated that the Atypicality Score subsets resulted in errors close to the errors obtained when the entire training set is used, while achieving a 3× or greater training time speedup compared to this baseline. Additionally, we analyzed the Atypicality Score algorithm's typical and atypical molecule assignments to gain insight into the molecular characteristics the algorithm determined most beneficial for subset selection.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

MolSelector: A Machine Learning Framework for Interpretable Subset Selection from Molecular Science Data Sets. — 科研速览 Science Skim