Joshua S Harris, Patricia A Vignaux, Thomas R Lane, Sean Ekins
Uncertainty quantification is crucially important for small-molecule machine learning (ML) models. For companies and regulatory bodies that rely on ML models to inform high-cost and high-risk decisions, accurate uncertainty quantification can reveal whether a given prediction is trustworthy. Existing methods for uncertainty quantification have been shown to be less reliable outside of a model's applicability domain. We have developed a novel approach to uncertainty quantification called domain classification (DC) modeling, which combines results from three independent binary classification models to capture the relationship between the subsets of active and inactive molecules in a training set, as well as a large and diverse "out of domain" set. Applying this modeling approach to acetylcholinesterase inhibition, we show that it robustly captures aleatoric and epistemic uncertainty even for molecules far outside the applicability domain. We show that filtering predictions with uncertainty-based thresholds leads to superior prediction performance even for external test sets with significant out-of-domain character, at the cost of discarding uncertain predictions. In particular, while screening out the top 50% of uncertain predictions from a binary random forest model, recall on external test sets improves from 0.56 to 0.80, precision improves from 0.82 to 0.87, and specificity stays about the same (0.92 to 0.91). While screening out 70% of predictions, recall and precision increase to 0.91 and 0.92, respectively, while specificity stays at 0.92. Improvements in recall and precision are statistically significant with a p-value of 0.05, whereas there is no significant change in specificity. We also use DC modeling to define uncertainty windows for prediction probabilities based on confidence levels and demonstrate that these confidence levels accurately represent the likelihood that ground truth activity probabilities fall within the uncertainty window.