Martin Culka, Nicolas W Lounsbury, William Thrift, Santrupti Nerli, Andrew Wallace, Gergő Nikolényi, Darya Orlova, Kiran Mukhyala, Mohammed AlQuraishi
Understanding T cell receptor (TCR) discrimination of MHC-presented epitope peptides (pMHCs) remains challenging. While machine-learning (ML)-based predictions of TCR specificity have gained attention, their capacity to generalize to unseen peptides is often misinterpreted. Using a proprietary cancer patient dataset, we show that ML methods succeed in predicting TCR specificity for known peptides but fail to generalize to novel peptides. Conversely, physics-based methods outperform ML methods on novel peptides but underperform on known peptides. In light of these observations, we develop a new ML method that leverages protein foundation models to achieve better or comparable performance than existing ML and biophysical methods on both in- and out-of-distribution TCR-pMHC specificity prediction. We furthermore characterize method performance as a function of distance of TCR sequence specificity between training and test sets. Our analysis elucidates the current limitations of modeling TCR-pMHC interactions and outlines new avenues for method development and data acquisition.