科研速览 · Science Skim继续刷下去 · Keep skimming →
◇ bioRxiv2026-08-27· biochemistry

Multi-Peptide Prompting Enables In-Context Learning in Protein Language

J. Almonte, M. T. Vu, A. Ahn, E. H. Thiede

原始摘要(英文原文)· Original abstract
Protein language models (PLMs) are trained primarily on individual protein sequences, yet many peptide-discovery problems require inference from only a small number of labeled examples. Here, we show that single-sequence PLMs can perform in-context peptide learning without gradient updates, task-specific retraining, or architectural modification. We introduce multi-peptide example prompts (MPEPs), in which demonstration peptides are concatenated with glycine spacers and used as context for scoring query peptides by their prompted probability. We evaluate this approach across a synthetic pattern-completion task, secondary-structure classification, and MHC-II binder prediction using both encoder-only ESM-2 models and decoder-only ProGen2 models. Across tasks, performance improves with the number of peptide examples and with model scale, indicating that PLMs can extract shared sequence-level properties from prompted examples. We further introduce a difference score that contrasts positive-example and negative-example MPEPs, reducing compositional biases in raw PLM probabilities and substantially improving classification. On MHC-II binder prediction, MPEP-based classification with larger ESM-2 models matches or exceeds low-data classifiers trained on frozen ESM-2 embeddings, while requiring no training. These results reveal an unexpected in-context inference capability in single-sequence PLMs and establish MPEP conditioning as a lightweight strategy for low-data peptide classification.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Multi-Peptide Prompting Enables In-Context Learning in Protein Language — 科研速览 Science Skim