科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Bioinformatics (Oxford, England)2026-08-01

Mitigating Goodhart's law in epitope-conditioned TCR generation using plug-and-play reward designs.

Pengfei Zhang, Xiaoyi He, Fredo Guan, Hao Mei, Gloria Grama, Seojin Bang, Heewook Lee

一句话结论 · In one sentence

We introduce a plug-and-play reward-design framework for RL-based TCR generation that combines heuristic biological priors, model ensembling, and binding-specificity objectives based on max-margin and contrastive formulations. These components suppress degenerate sequences, reduce model-specific biases, and discourage cross-epitope binding. During RL fine-tuning, the proposed rewards stabilize optimization, limit surrogate-reward inflation, preserve canonical CDR3β sequence patterns and repertoire diversity, and maintain closer alignment with experimentally validated TCR-binder distributions. In evaluations on unseen epitopes, the specificity-aware reward formulations provide the strongest overall performance, improving diversity and ground-truth distributional alignment while retaining biological authenticity and predicted target binding. These findings demonstrate that Goodhart-resistant reward design improves the reliability, controllability, and generalization of epitope-conditioned TCR generation.

原始摘要(英文原文)· Original abstract
MOTIVATION: Epitope-conditioned T cell receptor (TCR) generation extends protein language modeling to the design of therapeutically relevant receptors. Reinforcement learning (RL) post-training with surrogate binding predictors can improve generation controllability, but it is vulnerable to Goodhart's Law: optimizing an imperfect surrogate reward can lead to reward inflation, distributional drift, and biologically implausible or nonspecific sequences. We investigate whether reward hacking can be mitigated through improved reward formulations without modifying the generator architecture or training pipeline and without requiring additional training data. RESULTS: We introduce a plug-and-play reward-design framework for RL-based TCR generation that combines heuristic biological priors, model ensembling, and binding-specificity objectives based on max-margin and contrastive formulations. These components suppress degenerate sequences, reduce model-specific biases, and discourage cross-epitope binding. During RL fine-tuning, the proposed rewards stabilize optimization, limit surrogate-reward inflation, preserve canonical CDR3β sequence patterns and repertoire diversity, and maintain closer alignment with experimentally validated TCR-binder distributions. In evaluations on unseen epitopes, the specificity-aware reward formulations provide the strongest overall performance, improving diversity and ground-truth distributional alignment while retaining biological authenticity and predicted target binding. These findings demonstrate that Goodhart-resistant reward design improves the reliability, controllability, and generalization of epitope-conditioned TCR generation. AVAILABILITY AND IMPLEMENTATION: Code and models are available in a public repository (https://github.com/Lee-CBG/TCRRobustRewardDesign).
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Mitigating Goodhart's law in epitope-conditioned TCR generation using plug-and-play reward designs. — 科研速览 Science Skim