Eric Chen, J. Green, Yingkai Zhang
Allosteric kinase inhibitors are an important modality for overcoming resistance and achieving selectivity, yet most structure-based docking and deep generative models are trained predominantly on orthosteric protein-ligand complexes. As a result, current methods often misplace allosteric kinase ligands into the adenosine triphosphate (ATP)-binding site and fail to recover the correct binding mode. Here we curate AlloSet, a kinome-wide, time-split data set of kinase-ligand complexes annotated by binding mode, to systematically evaluate and fine-tune the diffusion-based docking model DiffDock-L for allosteric pose prediction. We explore several fine-tuning strategies, including increased dropout, freezing of torsion parameters with translation/rotation-only fine-tuning, and molecular dynamics-based supersampling of receptor conformations and ligand poses. The resulting DiffDock-L-Allo model is found to markedly improve pose-recovery metrics for Type III/IV allosteric binders while preserving the performance on ATP-site ligands. Binding-mode-resolved evaluations and comparisons with cofolding models such as AlphaFold3 and Boltz-2 highlight how targeted retraining reshapes the generative model's sampling distribution, offering practical guidance for adapting AI-driven docking to challenging, low-data binding modes in kinase structure-based drug design.