Pu Wang, Weihao Liu, Bo Hang
Umami peptides are a class of functional biomolecules that significantly enhance food flavor, holding substantial application value in the food industry and nutritional science. Traditional prediction methods rely heavily on manual feature engineering, failing to adequately capture the complex semantic information and spatial structural characteristics of peptide sequences. This study presents UmamiFusion, a multimodal deep learning framework that integrates sequence local patterns with spatial structural features through attention-enhanced fusion to overcome the limitations of single-modal modeling approaches. Specifically, the protein language model ESM2 is employed to generate graph-structured representations of peptide molecules (with nodes representing amino acid residues and edges defining spatial contact relationships), combined with graph convolutional networks to extract high-order topological features. Simultaneously, one-dimensional convolutional neural networks (1D-CNNs) capture local contextual information from amino acid sequences. A bidirectional cross- attention mechanism enables adaptive interaction between graph nodes and sequence positions before the features are concatenated and fed into a multilayer perceptron for end-to-end classification. Experimental results on two meticulously constructed datasets, UMP442 and UMP614, demonstrate that UmamiFusion achieves accuracies of 96.27% and 98.24%, respectively, with MCC scores of 0.8214 and 0.8529, and AUC values of 0.9711 and 0.9879, significantly surpassing current state-of-the-art models. This research validates the critical role of attention-enhanced multimodal feature fusion in umami peptide prediction, providing an efficient computational tool for high-throughput screening of functional peptides and food flavor optimization.