Özge Çolakoğlu, Anthuvan Joseph Benjamin
The increasing demand for efficient and interpretable computational methods for molecular property prediction has stimulated the development of quantitative structure-property relationship (QSPR) models based on graph-theoretical molecular descriptors. Topological indices are numerical descriptors derived from graph-theoretical representations of chemical structures and are widely used in QSPR studies to predict physicochemical and biological properties. In this study, degree-based and neighborhood-degree-sum-based topological indices were computed for 18 drugs used to treat bipolar disorder and employed as descriptors for modeling physicochemical properties. linear regression, ridge regression, random forest, and XGBoost were systematically evaluated to investigate the suitability of classical and machine learning regression methods for QSPR modeling of a small molecular data set. The comparative analysis demonstrates that although ensemble machine-learning models (random forest and XGBoost) were capable of capturing nonlinear relationships, their predictive performance deteriorated under leave-one-out cross-validation because of the limited sample size. In contrast, ridge regression consistently produced the lowest prediction errors while preserving interpretability. Neighborhood-degree-sum-based descriptors produced lower prediction errors for seven of the nine investigated physicochemical properties, whereas degree-based descriptors remained stronger for boiling point and heavy atom count. Although meaningful relationships between molecular topology and physicochemical properties were identified, the findings are limited by the small data set and lack of external validation. Therefore, the proposed models should be considered as exploratory approximations. This study provides new insights into the structure-property relationships of bipolar disorder drugs and establishes a foundation for future studies using larger molecular data sets.