Udbhas Garai, Aditya S. Pal, Koyel Ghosh, Deepak B. Salunke, Utpal Garain
Potency (IC50) prediction of small molecules is pivotal for anticancer drug development. This study benchmarked five deep learning (DL) models for IC50 prediction—DeepCDR, DrugCell, PaccMann, Precily, and tCNN—against a simple mean-based Baseline using standardized GDSC datasets and recently published anticancer compounds. To ensure practicality, conventional error metrics were supplemented with percentage error, log error, three-sigma limit, and a newly proposed Experimental Variability-Aware Prediction Accuracy statistic. The models performed well on randomly split data and unseen cell lines but showed sharply reduced accuracy for unseen compounds. Though all DL models exhibited similar performance trends, DeepCDR, DrugCell, and tCNN held a slight edge in most testing scenarios. Interestingly, several DL algorithms could not significantly outperform the Baseline model in many tests. Assessing prediction error against physicochemical and biological properties of compounds and cell lines revealed weak correlation, highlighting an underexplored aspect of model performance. A user-friendly web server (https://nlplab1.isical.ac.in/ic50.php) was also developed for IC50 prediction of new compounds against cancer cell lines. Accurate IC50 prediction of small molecules is crucial for advancing anticancer drug development, however, the effectiveness of deep learning methods for predicting IC50 values remains underexplored. Here, the authors benchmark five deep learning models against a mean-based baseline, revealing that while DeepCDR, DrugCell, and tCNN slightly outperform others, all models struggle with unseen compounds, underscoring the need for improved predictive accuracy.