Parker Dryja, Morné Muller, Monique Horn, Ilya Balabin, Zackery W Dentmon, Prawin Rimal, Yuri K Peterson, Fourie Joubert, Thomas M Kaiser, Pieter B Burger
The rapid advancement of machine learning (ML)-based protein structure prediction, exemplified by AlphaFold2 and extended by newer models such as AlphaFold3 and Boltz-2, has generated significant optimism for structure-guided drug discovery. In particular, ligand-protein cofolding approaches offer the potential to overcome limitations in generating starting structures for physics-based free energy perturbation (FEP) calculations. However, the practical readiness of ML-predicted structures for FEP applications remains insufficiently evaluated. Here, we systematically assess experimentally determined crystal structures, a homology model, and ML-predicted protein structures as inputs for FEP using a well-characterized congeneric series targeting the tyrosine kinase cSrc. A data set of 133 compounds was evaluated through more than 1400 FEP calculations under minimal optimization to approximate "out-of-the-box" performance. By maintaining consistent preparation protocols, we isolate the impact of structural origin on predictive accuracy. Variable performance was observed across both experimental and ML-predicted structures, highlighting that even under this idealized benchmark scenario, significant challenges remain in reliably generating and refining predictive protein-ligand complexes. This study demonstrates that predictive variation in micro and macro conformational states─rather than the structural source─governs predictive reliability, underscoring the need for careful validation when integrating ML-derived structures into FEP workflows.