Javier H Gil-Gómez, Carlos Carvajal-Fierro, Ricardo Bruges-Maya, Ixchel Rodríguez Parada, Andrés Mosquera-Zamudio, Julián Riaño-Moreno, José Fernando Polo, Marcela Gómez-Suárez, John Jaime Sprockel, Rafael Parra-Medina
In this cohort, machine learning models showed potential for overall survival prediction in lung adenocarcinoma. RSF achieved the highest observed discrimination, while the penalized Cox model showed competitive performance with lower complexity. These findings suggest that RSF may complement conventional survival models; however, the retrospective single-center design, moderate sample size, missing data, and lack of external validation limit their generalizability.
BACKGROUND: Lung adenocarcinoma is the most common subtype of non-small cell lung cancer and a leading cause of cancer-related mortality worldwide. Its marked clinical heterogeneity results in variable outcomes and poor prognosis for many patients. This study aimed to develop and compare statistical, machine learning, and deep learning models for predicting overall survival (OS) in patients with primary lung adenocarcinoma.
METHODS: This retrospective cohort study included patients with primary lung adenocarcinoma. Missing data were handled using fold-specific multiple imputation by chained equations within a nested cross-validation framework. Predictor selection and hyperparameter tuning were performed within 5 outer and 3 inner folds using univariable Cox analysis, Elastic Net, XGBoost-based importance, and clinical criteria. Overall survival was modeled using penalized Cox regression, Random Survival Forest, DeepSurv, and DeepHit. Model performance was evaluated from out-of-fold predictions using concordance indices, time-dependent AUC, calibration and the integrated Brier score. Sensitivity analyses included propensity score-matched survival comparisons by biomarker status among patients with stage IV disease.
RESULTS: The cohort included 254 patients, with a median overall survival of 12.8 months. Random Survival Forest showed the best performance (Harrell's C-index, 0.760; mean time-dependent AUC, 0.851; integrated Brier score, 0.160) and clearly separated high- and low-risk groups (p < 0.001). Clinical stage, ECOG status, systemic treatment, sex, and age were among the most influential predictors. Calibration was acceptable, with slight survival overestimation. Exploratory propensity score-matched analyses in stage IV subgroups suggested the highest apparent performance for RSF, although the small matched samples limited definitive conclusions.
CONCLUSIONS: In this cohort, machine learning models showed potential for overall survival prediction in lung adenocarcinoma. RSF achieved the highest observed discrimination, while the penalized Cox model showed competitive performance with lower complexity. These findings suggest that RSF may complement conventional survival models; however, the retrospective single-center design, moderate sample size, missing data, and lack of external validation limit their generalizability.