Jothi Prakash V, Arul Antran Vijay S, Ganeshkumar Pugalendhi
Multiomics learning is central to computational biology, yet existing models often prioritise predictive accuracy over biological admissibility, yielding brittle representations and limited cross-cohort generalisation. We address the problem of learning transferable multiomics representations that remain consistent with known biological structure under missing modalities and distribution shifts. We propose Bio-Constrained Fusion Modeling (BCFM), a constraint-aware framework that integrates pathway coherence and interaction sparsity directly into representation learning via an augmented Lagrangian formulation. Modality-specific encoders project heterogeneous omics data into a shared latent space, fused through adaptive gating and aligned across modalities, while biological constraints are enforced as feasibility conditions rather than heuristic regularisers. Extensive experiments on TCGA pan-cancer data and external GEO cohorts show consistent gains over nine representative baselines: up to 3.4% higher classification accuracy, 3.1% higher macro-F1, and a 2.5% absolute gain in survival C-index, alongside substantially reduced biological constraint violations, mitigated degradation under missing modalities and cross-cancer transfer, and stable constraint convergence during training. Overall, embedding biological admissibility into optimisation yields robust, generalisable, and biologically plausible multiomics representations that transfer across cohorts and cancer types for bioinformatics applications in biomedical research.