Xinlong Luo, Gaoshi Li, Zhipeng Hu, Jingli Wu, Wei Peng, Jiafei Liu, Xiaoshu Zhu
Essential proteins are a crucial component of living organisms, and their absence will lead to cell death or reproductive arrest. Discovering these proteins can propel advancements in synthetic biology and facilitate the development of novel antibiotics and therapies for various diseases. However, current computational methods suffer from two major drawbacks that hinder their discovery rate: one is the significant noise in protein-protein interaction (PPI) data, and the other is the inadequate consideration of feature relationships. To enhance identification capabilities, this study proposes a novel essential protein prediction method, Feature Synergy Method (FSM), which leverages a features synergy model and GO pure centrality. The FSM is described as follows:Firstly, based on the principle of co-expression, gene expression data are integrated with the original PPI network to construct a pure PPI network (PPIN). Subsequently, GO annotation data are employed to calculate GO_sim weights for the interactions within the original PPI network, forming a GS_PIN. The PPIN and GS_PIN are then fused to establish the GS_PPIN, which helps mitigate the impact of noise in PPI data. Secondly, a new centrality measure, GO pure centrality (GPC), is designed based on this GO similarity-weighted pure PPI network. Thirdly, an evolutionary conservation score (ECS) is extracted from subcellular localization and orthologous proteins data. Fourthly, after analyzing the relationship between GPC and ECS, a novel fusion model, the features synergy model, is developed to integrate GPC and ECS, ultimately leading to the proposal of the new essential protein prediction method, FSM. To validate the performance of FSM, six computational methods (PeC, WDC, ION, NCCO, E_POC, and JDC) and six centrality measures (NC, IC, EC, SC, CC, and DC) were evaluated on three distinct yeast datasets. The results demonstrate that FSM achieves a higher essential protein identification rate. Similarly, GPC identifies more essential proteins compared to the six centrality-based approaches (NC, IC, EC, SC, CC, and DC).