Bo Zhang, Ruifang Li, Kedong Yin, Yufeng Yang, Jinhua Zhang, Mengwan Jiang, Huijie Wang, Shiyu Li, Lujing Jia
Artificial intelligence accelerates anticancer peptides (ACPs) discovery. However, existing computational methods lack integration of identification with activity-based candidate prioritization. Here, we present DeepACPred, a three-stage pipeline encompassing ACP binary classification model, ACP multilabel classification model, and ACP IC50 prediction model, leveraging multimodal features from ESM2 protein language model embeddings, AAindex physicochemical descriptors, and sequence composition. On 5712 benchmark sequences, the binary classifier achieved 95.10% accuracy (AUC = 0.9913), with performance remaining stable under CD-HIT cluster-aware splitting at 40%-90% identity thresholds. Multilabel cancer-type prediction yielded macro-F1 = 0.9124 across seven cancer types, and log10(IC50) regression achieved Spearman ρ = 0.8602 under 5-fold cross-validation. Ablation experiments showed task-dependent feature contributions rather than uniformly additive multimodal effects. Applied to 260 000 motif-enriched 18-mer candidates, DeepACPred selected 12 peptides predicted to be active against breast cancer cells, all of which showed measurable in vitro cytotoxic activity against murine 4T1 cells in OD-derived dose-response assays (IC50: 0.88-36.83 μg/ml). Although prospective IC50 ranking showed limited fine-grained resolution, these results support the use of the regression module for coarse candidate enrichment. In conclusion, DeepACPred provides a systematic framework for ACP candidate enrichment and prioritization.