Ali Jafarian, Fatemeh Keshmiri Nasrabadi, Mehrshad Khosraviani
Logistic Regression provided the most robust and interpretable results. These findings support the utility of simple demographic-based models as noninvasive tools that may support early supportive screening of subjective cognitive difficulties in cancer survivorship care. However, external validation using independent datasets is required.
OBJECTIVES: Cancer-related cognitive impairment is a decline in cognitive functioning following cancer treatment that negatively affects survivors' quality of life. Self-report tools such as the Cognitive Failure Questionnaire (CFQ) offer a rapid and cost-effective method for early detection of subjective cognitive difficulties. This study aimed to develop and compare machine learning models using demographic variables to classify subjective cognitive impairment (SCI) severity into low, moderate, and high levels.
METHODS: Data from 437 cancer survivors were analyzed. SCI severity was defined based on self-reported CFQ scores. Demographic predictors included age, gender, marital status, education, and occupation. Four supervised machine learning algorithms, Logistic Regression (LR), Support Vector Machine (SVM), Artificial Neural Network (ANN), and k-Nearest Neighbors (KNN), were applied for multiclass classification. Model performance was evaluated using accuracy, sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV), F1-score and area under the ROC curve (AUC), with stratified 5-fold cross-validation.
RESULTS: LR model demonstrated the best performance (accuracy = 0.89, AUC = 0.95), followed by SVM (accuracy = 0.86, AUC = 0.95), ANN (accuracy = 0.85, AUC = 0.93), and KNN (accuracy = 0.82, AUC = 0.91).
CONCLUSION: Logistic Regression provided the most robust and interpretable results. These findings support the utility of simple demographic-based models as noninvasive tools that may support early supportive screening of subjective cognitive difficulties in cancer survivorship care. However, external validation using independent datasets is required.