Cosmin-Daniel Minciuna, Dorina Minciuna, Ingrid-Andrada Vasilache, Lucian Miron
Background/Objectives: Diffuse large B-cell lymphoma (DLBCL) remains clinically heterogeneous despite standard first-line immunochemotherapy. We evaluated whether machine learning (ML) models using routinely available diagnostic variables could predict primary refractory DLBCL. Materials and Methods: We performed a retrospective single-center study of 369 patients with DLBCL treated between 2015 and 2023. Primary refractory disease was defined as failure to achieve at least a partial response after first-line therapy. Logistic regression, random forest, gradient boosting, support vector machine with radial basis function kernel (SVM-RBF), and multilayer perceptron models were trained using baseline clinical and laboratory predictors. Performance was assessed using a stratified 80/20 holdout split, nested cross-validation, repeated random splits, and sensitivity analysis with complete hyperparameter retuning. Results: Primary refractory disease occurred in 111 patients (30.1%). In the holdout test set (n = 73), IPI achieved the highest discrimination (AUC 0.692; 95% CI 0.545-0.828). Among ML models, gradient boosting performed best (AUC 0.658; 95% CI 0.503-0.801), followed by SVM-RBF (AUC 0.619; 95% CI 0.471-0.760). DeLong comparisons showed no significant differences between ML models and IPI. In internal validation, gradient boosting showed the highest nested CV AUC (0.712 ± 0.049). Feature importance identified IPI as the dominant predictor. A weighted ML-derived clinical score identified a high-risk group with the highest refractory rate (52.9%), although the trend was not statistically significant (p = 0.063). Conclusions: ML models showed comparable but not superior performance to IPI for predicting primary refractory DLBCL. These findings support further external validation and development of multimodal clinically interpretable prediction frameworks for risk-adapted management.