Jiang-Hua Tang, Sadia Noureen, Nazma Ashraf, Saood Azam, Adnan Aslam
High Resolution Image Download MS PowerPoint Slide This study addresses the challenge of accurately predicting the physicochemical properties of arthritis drugs using simple graph-theoretic descriptors and evaluates whether combining degree-based topological indices with modern machine learning can outperform traditional linear QSPR models. Nine classical degree-based topological indices (including reformulated Zagreb indices, augmented Zagreb index, hyper-Zagreb index, forgotten index, symmetric division degree index, and inverse sum Zagreb index) were computed for 50 structurally diverse arthritis drugs ranging from NSAIDs (ibuprofen and diclofenac) and COX-2 inhibitors (celecoxib) to disease-modifying agents (methotrexate) and immunomodulators with experimental property values (boiling point, melting point, critical pressure, molar refractivity, topological polar surface area, and calculated LogP) obtained from ChemDraw and PubChem. Using these indices as input features, linear regression, Random Forest, and XGBoost models were developed and compared. XGBoost consistently outperformed the other methods across all properties, achieving for molar refractivity a mean absolute error (MAE) of 0.80787, a root-mean-square error (RMSE) of 1.13352, and coefficient of determination ( r 2 = 0.9945); for boiling point, RMSE = 16.22 and r 2 = 0.9866; for melting point, RMSE = 25.86 and r 2 = 0.9450; for critical pressure, RMSE = 1.176 and r 2 = 0.9833; for topological polar surface area, RMSE = 4.38 and r 2 = 0.9762; and for CLogP, RMSE = 0.586 and r 2 = 0.8857. Random Forest also performed well but with slightly higher errors, while linear regression showed moderate to strong correlations for boiling point, melting point, molar refractivity, and tPSA but failed to capture nonlinear structure–property relationships. The novelty of this work lies in the first systematic integration of nine degree-based topological indices with advanced ensemble learning (XGBoost and Random Forest) for QSPR modeling of a large, diverse arthritis drug data set, demonstrating that simple connectivity measures─when paired with nonlinear machine learning─can achieve near-perfect predictive accuracy for key properties, offering a low-cost, computationally efficient framework for drug design and virtual screening.