Guorong Wu, Muhammad Waheed Rasheed, Samirah Alsulami
Quantitative Structure-Property Relationship (QSPR) modelling provides an efficient computational framework for predicting physicochemical properties of drug molecules when experimental data are limited. In this study, we investigate the predictive capability of degree-based topological indices (TIs) derived from SMILES (Simplified Molecular Input Line Entry System) representations for modelling physicochemical properties of anti-alkaptonuria drugs. Nine representative compounds, including Nitisinone, Ascorbic Acid, Ibuprofen, Naproxen, Paracetamol, Tramadol, Methotrexate, Sulfasalazine, and Glucosamine, were analysed using several molecular descriptors such as molecular weight, logP, hydrogen bond donors and acceptors, rotatable bonds, and polar surface area. A total of 58 regression models were developed using Linear Regression (LR) and two machine learning algorithms, Random Forest (RF) and Extreme Gradient Boosting (XGBoost, abbreviated XGB). Model performance was evaluated using Mean Absolute Error (MAE), Mean Squared Error (MSE), Root Mean Squared Error (RMSE), and the coefficient of determination R 2 . The results demonstrate that machine learning models significantly outperform classical regression, with XGB achieving the most accurate and stable predictions for the investigated physicochemical properties. This study introduces a machine learning-driven QSPR framework that integrates SMILES-derived degree-based topological indices with ensemble learning techniques for predicting physicochemical properties of anti-alkaptonuria drugs. The proposed approach demonstrates improved predictive performance on small datasets and highlights the effectiveness of combining graph-theoretic molecular descriptors with advanced machine learning methods.