D. Garger, J. H. Burks, N. I. Paton, D. J. Sloan, G. Maher-Edwards, F. Theis, M. P. Menden, F. P. Casale
Relapse after apparently successful tuberculosis (TB) therapy remains difficult to predict, and how relapse risk evolves throughout treatment remains unclear. Using harmonised clinical data from two Phase 3 trials (2,918 participants), we performed time-resolved prediction of end-of-therapy (EOT) outcomes and post-treatment relapse by training models using tabular data at monthly intervals from baseline to therapy end. Prediction of EOT outcomes improved after month 3 (ROC-AUC up to 0.84), driven by sputum-smear and solid culture. In contrast, relapse prediction among participants with favourable EOT outcomes and completed follow-up improved only modestly through month 3 (ROC-AUC 0.58-0.63) before declining, with age, sex, clinical symptoms and bacterial burden contributing most strongly to prediction. Models trained on large language model-derived embeddings, a more flexible representation of the same variables, matched tabular relapse models in performance throughout therapy, and outperformed tabular EOT outcome models at months 3 and 4 ({Delta}ROC-AUC: 0.12 and 0.14), with longitudinal modelling and sparse variable inclusion only improving relapse prediction at months 4-6 ({Delta}ROC-AUC: 0.09, 0.19 and 0.12), however with reduced interpretability. Models incorporating data after baseline provided incremental improvements in post-treatment relapse risk stratification compared with baseline cavitation and sputum smear alone (4-month relapse-free survival: 94.8%/78.5% for model-derived low/high risk groups, vs. 91.4%/82.9% for baseline easy-/hard-to-treat groups). Overall, these findings suggest that while routine clinical data collected after baseline can improve post-treatment risk stratification, it offers limited predictive value for relapse, underscoring the need for relapse-specific biomarkers.