Bailang Liu, Yong Liu, Hui Yang, Yang Li, Shanshan Yin, Junnan Wang, Zhikui Cheng, Xin Zhao, Xuejun Zhang, Shu Li, Lifei Liu
Tree-based models generally showed the strongest external performance. In particular, models using RDKit2D descriptors achieved Matthews correlation coefficients (MCC) above 0.65 and a ROC-AUC of 0.87, while several other tree-based model representation combinations achieved MCC values around 0.45-0.50. Boruta markedly reduced feature dimensionality, but its effect on predictive performance depended on both the molecular representation and learning algorithm; it was most beneficial for selected high-dimensional fingerprints and tree-based models, while offering limited or unfavorable effects in some support vector machine settings. Applicability-domain and SHAP analyses further revealed representation-dependent generalization patterns and highlighted contributions from lipophilicity, polarity, solubility, and molecular complexity.
INTRODUCTION: Bile salt export pump (BSEP) inhibition is an important mechanistic risk factor for cholestatic drug-induced liver injury. Quantitative structure-activity relationship (QSAR) models may support early safety screening, but practical BSEP datasets are often limited in size and imbalanced.
METHODS: We evaluated a machine-learning workflow combining molecular descriptors and fingerprints, Boruta feature selection, and class-imbalance training. Multiple algorithms were assessed using cross-validation, a held-out test set, and an independent external validation set composed of compounds synthesized in an industrial research setting.
RESULTS: Tree-based models generally showed the strongest external performance. In particular, models using RDKit2D descriptors achieved Matthews correlation coefficients (MCC) above 0.65 and a ROC-AUC of 0.87, while several other tree-based model representation combinations achieved MCC values around 0.45-0.50. Boruta markedly reduced feature dimensionality, but its effect on predictive performance depended on both the molecular representation and learning algorithm; it was most beneficial for selected high-dimensional fingerprints and tree-based models, while offering limited or unfavorable effects in some support vector machine settings. Applicability-domain and SHAP analyses further revealed representation-dependent generalization patterns and highlighted contributions from lipophilicity, polarity, solubility, and molecular complexity.
DISCUSSION: Overall, the workflow maintained useful predictive performance under small and imbalanced conditions and provides a practical framework for early-stage BSEP risk screening.