Lipika R Pal, E Michael Gertz, Nishanth Ulhas Nair, Sumit Mukherjee, Sumeet Patiyal, Thomas Cantore, Emma M Campagnolo, Tian-Gen Chang, Saugato Rahman Dhruba, Yewon Kim, Eldad D Shulman, Padma Sheila Rajagopal, Danh-Tai Hoang, Sridhar Hannenhalli, Alejandro A Schäffer, Eytan Ruppin
Precision oncology aims to guide treatment decisions using biomarkers. While DNA-based panels are increasingly applied, RNA transcriptomics remain underused due to limited datasets and the absence of robust models. To address this, we assembled the largest transcriptomic resource for drug response prediction to date, spanning 91 cohorts, 5,675 patients, nine cancer types, and six frontline therapies: anti-PD-1/PD-L1 immune-checkpoint inhibitors, trastuzumab, bevacizumab, BRAF inhibitors, paclitaxel, and FAC/FEC (fluorouracil-adriamycin-cyclophosphamide/fluorouracil-epirubicin-cyclophosphamide) chemotherapy. A supervised machine-learning framework, EXPRESSO (EXpression-Profile-RESponSe-Optimizer), was developed to predict treatment response from pre-treatment transcriptomes by integrating drug targets and context-specific biomarkers. EXPRESSO effectively predicted patient treatment responses across therapies, outperforming 20 published transcriptomic signatures and other machine learning methods. Prospective validation on 22 independent cohorts confirmed that performance generalized beyond cross-validation. The EXPRESSO signature additionally stratified progression-free survival in immune checkpoint blockade-treated cohorts, demonstrating prognostic value beyond binary response prediction. Robustness analysis revealed that predictive performance plateaued for some therapies with increasing training cohorts but continued to improve for others. These findings suggest inherent limits of supervised brute-force learning for certain treatments, but additional data and deeper mechanistic modeling may further enhance transcriptomics-based predictors.