Yan Miao, Tingting Zou, Zhenyuan Sun, Yuming Zhao, Guohua Wang
We propose DeepGVS, a bimodal deep learning framework integrating CDS-derived features with protein sequence and structural representations for VF prediction. DeepGVS extracts multi-scale sequence-composition features from CDS. Concurrently, it employs a parallel graph attention network (GAT) and bidirectional Mamba (Bi-Mamba) architecture to process ESMFold-predicted structures and residue representations, capturing spatial and long-range dependencies. A neural additive model (NAM) functions as a meta-learner to integrate base-classifier predictions. DeepGVS was evaluated on the unchanged accession-level independent test set of Dataset_B and achieved an accuracy of 87.50%, corresponding to absolute improvements of 6.30, 2.60 and 1.40 percentage points over the published benchmark values of DeepVF, GTAE-VF and PLMVF, respectively. The incremental benefit of bimodal integration was dataset- and metric-dependent, and additional taxonomic analyses identified taxonomy as a potential confounding factor that does not fully reproduce the performance of the complete model.
MOTIVATION: Virulence factors (VFs) mediate host adhesion, invasion, immune evasion and toxin-mediated damage, making accurate VF prediction important for understanding bacterial pathogenesis and antimicrobial intervention. Existing predictors mainly use one-dimensional (1D) protein sequences, overlooking complementary coding DNA sequence (CDS)-level information and three-dimensional (3D) structural topology.
RESULTS: We propose DeepGVS, a bimodal deep learning framework integrating CDS-derived features with protein sequence and structural representations for VF prediction. DeepGVS extracts multi-scale sequence-composition features from CDS. Concurrently, it employs a parallel graph attention network (GAT) and bidirectional Mamba (Bi-Mamba) architecture to process ESMFold-predicted structures and residue representations, capturing spatial and long-range dependencies. A neural additive model (NAM) functions as a meta-learner to integrate base-classifier predictions. DeepGVS was evaluated on the unchanged accession-level independent test set of Dataset_B and achieved an accuracy of 87.50%, corresponding to absolute improvements of 6.30, 2.60 and 1.40 percentage points over the published benchmark values of DeepVF, GTAE-VF and PLMVF, respectively. The incremental benefit of bimodal integration was dataset- and metric-dependent, and additional taxonomic analyses identified taxonomy as a potential confounding factor that does not fully reproduce the performance of the complete model.
AVAILABILITY AND IMPLEMENTATION: Source code, datasets and pretrained models are available at https://github.com/guoguo26/DeepGVS. The archived code version used for the reported experiments is available at https://doi.org/10.5281/zenodo.21756890.
SUPPLEMENTARY INFORMATION: Supplementary data are available online.