Carlos Victor Montefusco-Pereira
Sequence contains assay-associated ranking signal, but performance depends strongly on validation design and deteriorates temporally. Context improves retrospective prediction but may encode study structure and is not causal evidence. The framework supports benchmarking and candidate prioritization, not clinical or vaccine-efficacy prediction.
OBJECTIVE: To benchmark leakage-aware machine-learning methods for ranking viral T-cell peptides by qualitative interferon-gamma (IFN-gamma) assay outcome in public Immune Epitope Database (IEDB) records.
METHODS: Recovered curated IEDB records were evaluated with exact peptide-disjoint, edit-distance < =2 component-disjoint, and publication-year temporal tests. Character n-gram models were tuned separately within each design, and uncertainty was estimated by grouped bootstrap. Context models excluded outcome-derived evidence variables and were tested using grouped ablations and three definitions of discordant peptide-context labels. Calibration and a frozen 8-million-parameter ESM-2 baseline were supporting analyses.
RESULTS: The recovered curated dataset contained 31,502 assays from 17,336 peptides. Best sequence-only PR-AUC was 0.479 (95% CI 0.406-0.552) in the exact peptide-disjoint test, 0.404 (0.273-0.546) under component separation, and 0.151 (0.126-0.183) temporally. Leakage-screened context increased PR-AUC from 0.535 to 0.771 under exact separation and from 0.501 to 0.755 by component. After strict censoring of post-cutoff training evidence, temporal context PR-AUC increased from 0.324 to 0.363 (paired uplift 0.039; 95% CI 0.016-0.061); the direction persisted after excluding 662 tied outcomes and restricting to consistent labels. Frozen ESM-2 embeddings did not outperform the strongest classical model in any design.
CONCLUSION: Sequence contains assay-associated ranking signal, but performance depends strongly on validation design and deteriorates temporally. Context improves retrospective prediction but may encode study structure and is not causal evidence. The framework supports benchmarking and candidate prioritization, not clinical or vaccine-efficacy prediction.