Binyu Li, Zhihua Huang, Yijie Ding, Xiaoyi Guo
MDFA-MLP provides an effective and reliable framework for ACP prediction. By integrating handcrafted descriptors, protein language model representations, and ensemble learning, the proposed method can facilitate large-scale computational screening of candidate ACPs prior to experimental validation.
INTRODUCTION: Experimental identification of anticancer peptides (ACPs) is timeconsuming and costly, which limits large-scale ACP discovery and screening. To address this challenge, we developed MDFA-MLP, a novel computational framework for ACP prediction that integrates multi-scale feature learning and ensemble classification strategies.
METHODS: The proposed framework combines physicochemical descriptors, including amino acid composition (AAC), dipeptide composition (DPC), composition-transition-distribution (CTD), and pseudo-amino acid composition (PseAAC), with ProtBERT-derived embeddings. A Multiscale Dilated Fusion Attention (MDFA) module was designed to capture sequence patterns at different scales and enhance feature fusion. An ensemble classifier consisting of a multilayer perceptron (MLP), support vector machine (SVM), and histogram-based gradient boosting (HGB) was employed to improve prediction robustness and stability.
RESULTS: The proposed model was evaluated on the AntiCP 2.0 dataset under the different negative-sample settings. On Dataset A, MDFA-MLP achieved an accuracy of 93.9%, sensitivity of 92.2%, specificity of 95.8%, and an MCC of 0.89. On the more challenging Dataset B, the model achieved an accuracy of 78.2%, sensitivity of 76.5%, specificity of 82.6%, and an MCC of 0.65. Comparative experiments demonstrated that MDFA-MLP achieved competitive and balanced performance across multiple evaluation metrics. Although the improvement over existing methods was moderate in some cases, the model maintained stable predictive performance under different negative-sample settings, indicating good robustness and generalization ability.
DISCUSSION: The results indicate that traditional sequence descriptors and deep protein language model embeddings provide complementary biological information. The MDFA module effectively enhances feature representation by integrating multi-scale sequence characteristics, while the ensemble strategy improves model robustness and generalization under varying data distributions.
CONCLUSION: MDFA-MLP provides an effective and reliable framework for ACP prediction. By integrating handcrafted descriptors, protein language model representations, and ensemble learning, the proposed method can facilitate large-scale computational screening of candidate ACPs prior to experimental validation.