科研速览 · Science Skim继续刷下去 · Keep skimming →
◇ bioRxiv2026-08-14· biophysics

PepXPro: a framework for curating, generating, and optimizing structure-affinity protein-peptide datasets

L. A. Chi, F. M. Ytreberg

原始摘要(英文原文)· Original abstract
Protein-peptide interactions are central to cellular signaling and to a growing class of peptide therapeutics, yet the datasets used to develop and benchmark computational methods for protein-peptide modeling remain poorly standardized. Available databases prioritize comprehensive coverage but require task-specific curation, while published benchmarks are typically distributed as static collections built with heterogeneous curation, quality-filtering, redundancy-reduction, and sampling strategies, limiting reproducibility and cross-study comparison. We present PepXPro, a modular framework that transforms publicly available protein-peptide structure-affinity resources into curated datasets and reproducible benchmark collections generated under user-defined criteria. PepXPro is organized into three components: Scrape, for deterministic curation of protein-peptide complex entries from public resources; GenSample, for constructing configurable subsets under explicit quality, redundancy, and sampling constraints; and Benchmark, for evaluating candidate subsets and selecting a non-redundant, representative, general-purpose benchmark for distribution. Starting from PDBbind and complementary resources, the curation pipeline yields a pool of protein-peptide complex entries that retains chemically complex cases, including disulfide-linked cyclic peptides, which are commonly excluded from existing benchmarks. We release PepXPro Benchmark v1, a benchmark comprising 70 non-redundant protein-peptide complexes with experimentally determined structures and binding affinities. The underlying framework provides an extensible foundation for reproducible protein-peptide benchmark construction.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

PepXPro: a framework for curating, generating, and optimizing structure-affinity protein-peptide datasets — 科研速览 Science Skim