Huiling Zhang, Shuwen Wen, Ziqi Xu, Xiaochuan Chen, Shaozhen Cai, Guanshang Huang, Haowen Zhao, Ling Yin, Yanjie Wei
Allosteric regulation plays a crucial role in modulating protein function, and allosteric pockets represent promising targets for drug discovery. However, identifying allosteric pockets remains challenging due to their structural and evolutionary diversity. In this study, we propose AlloPoc, a novel framework that integrates multi-modal protein language model (PLM) representations with biochemical features for accurate allosteric pocket prediction. Unlike existing PLM-based methods that rely exclusively on sequence-level features, AlloPoc leverages complementary sequence- and structure-based information to capture a more comprehensive view of allosteric pockets. To our knowledge, this represents the first integration of dual-modality PLMs in this field. AlloPoc comprises three core modules: feature extraction, feature processing (including selection and enhancement), and ensemble prediction. By integrating sequence- and structure-based PLM embeddings with physicochemical properties extracted by Fpocket, AlloPoc leverages a hierarchical ensemble strategy to mitigate prediction biases inherent to individual classifiers. Evaluated on an independent test set, AlloPoc achieves state-of-the-art performance with an AUC of 0.927 and an MCC of 0.524, surpassing existing methods including AlloFusion, DeepAllo, and AlloPED. These results demonstrate the effectiveness of integrating multi-modal PLMs with ensemble learning for allosteric pocket identification.