Xinjing Gong, Ji Li, Mu Su, Peishen Yu, Te Ma, Ruiyang Zhai, Chenye Zhang, Mengyan Zhang, Yan Zhang
Prioritizing cancer driver genes amid passenger alterations remains challenging because protein-protein interaction (PPI) networks are heterophilic, multi-omics evidence is heterogeneous and can conflict, and driver annotations are sparse. DRIVE is a semi-supervised graph framework integrating mutation frequency, copy-number aberration, DNA methylation, and gene expression with biological networks. It separates PPI neighborhoods into tight and loose semantic views based on learnable representation consistency, reducing cross-class signal mixing. Multi-omics evidence is decomposed into omics-common and omics-specific components through contrastive mutual-information learning and the soft orthogonality constraint. Joint training combines self-supervised learning with focal and max-margin objectives to improve prioritization under sparse, imbalanced annotations. Across six benchmark PPI networks, DRIVE outperforms ten methods, achieving mean areas under the precision-recall curve (AUPRC) and receiver operating characteristic curve (AUROC) of 0.9204 and 0.9704, respectively. Ablation, representation, and masked-driver recovery analyses show that DRIVE captures complementary network and molecular signals and remains robust to incomplete annotations. DRIVE identifies 186 high-confidence candidate driver genes enriched near known drivers, 80.1% of which receive DepMap CRISPR dependency support. These candidates reveal underappreciated connections to tumor regulatory programs, particularly NF-κB-associated inflammation, T-cell activation, and immune checkpoint regulation. Pharmacogenomic associations further suggest therapeutic vulnerabilities.