Kiran Naz, Hafiz Muhammad Bilal, Faraha Ashraf, Muhammad Kamran Siddiqui
The quantitative description of drug molecules through their molecular structures is central to computational studies of drugs. This paper presents a machine learning-driven quantitative structure-property relationship (QSPR) model linking graph-based structural descriptors to the physicochemical properties of anti-blood cancer drugs. A dataset of 13, anti-blood cancer drug molecules was characterized using six degree-based topological indices: the first Zagreb, second Zagreb, Sombor, reduced Sombor, harmonic, and Randić indices. These graph-derived descriptors were used as molecular features to train and compare six machine learning algorithms, namely linear regression, ridge regression, lasso regression, AdaBoost, XGBoost, and gradient boosting, for predicting nine physicochemical properties. Model performance was assessed using MSE,MAE,RMSE, and the coefficient of determination R2, together with a graphical comparison of observed and predicted values. The novelty of this work lies in the systematic, comparative evaluation of multiple graph-based descriptors across conventional, regularized, and nonlinear ensemble learning paradigms, rather than reliance on a single descriptor or algorithm, thereby revealing how effectively structural information embedded in topological indices can be transferred into property prediction. Gradient boosting emerged as the most efficient predictor, achieving R2=0.9996 to 0.9999,MSE=0.0027 to 33.8688,MAE=0.0425 to 3.9874, and RMSE=0.0519 to 5.8197, across all nine properties, confirming the strength of nonlinear ensemble learning for this task. These results demonstrate that combining degree-based topological descriptors with machine learning provides an effective and generalizable computational framework for QSPR modeling of anti-blood cancer drugs.