Mark Fonteyne, Nada Badr, Peter J K Kuppen, Alexander L Vahrmeijer, Gerard J P van Westen, Willem Jespers
Benchmarked co-folding methods (AlphaFold-Multimer, AlphaFold3, Boltz-2) against peptide docking methods (AutoDock-CrankPep, HADDOCK) using the PepPro dataset containing conformationally flexible protein-peptide complexes. Co-folding methods outperformed docking approaches in pose prediction accuracy and ranking performance when related complexes were represented in the training set, while HADDOCK outperformed all evaluated co-folding methods in the absence of such representation. Boltz-2 achieved the highest accuracy on the PepPro dataset but showed reduced generalization on non-redundant complexes.
Synthetic peptides are increasingly important as therapeutic and diagnostic agents due to their high specificity, ease in synthesis, and suitability for targeting protein interfaces. However, accurate prediction of peptide binding poses remains a major challenge, particularly for conformationally flexible protein targets and de novo designed peptides. Deep learning-based co-folding methods, including AlphaFold-Multimer (AFM), AlphaFold3 (AF3), and Boltz-2, offer an alternative to conventional peptide docking by simultaneously predicting protein-peptide complex structures without requiring predefined binding conformations. In this study, we systematically benchmarked these three co-folding methods against two peptide-optimized molecular docking methods, AutoDock-CrankPep (ADCP) and HADDOCK, using the PepPro dataset. This established dataset contains conformationally flexible protein-peptide complexes that are represented in the training sets of all co-folding methods. To assess the generalization performance of the co-folding methods, an independent test set was constructed classified by its level of redundancy with the training sets. The benchmark results showed that co-folding methods outperformed docking approaches in pose prediction accuracy and ranking performance when related complexes were represented in the training set. In the absence of such representation, however, HADDOCK outperformed all evaluated co-folding methods. Boltz-2 achieved the highest accuracy on the PepPro dataset but showed reduced generalization on non-redundant complexes, indicating strong dependence on training-set coverage. In contrast, AFM and AF3 demonstrated more robust generalization and maintained consistent performance on independent protein-peptide complexes, highlighting improved suitability for de novo peptide binder design. Furthermore, extensive analysis of scoring metrics revealed that interface predicted template modelling (ipTM) provides strong early enrichment performance, while peptide-specific predicted local distance difference test (pLDDT) correlates closely with structural accuracy, suggesting complementary value for candidate prioritization. These findings demonstrate that robust peptide-binder prediction requires careful consideration of training-set redundancy, target conformational flexibility, and complementary confidence metrics when selecting computational methods for de novo candidate prioritization.