Ilya Kandinov, Dmitry Trukhin, Dmitry Gryadunov, Elena Savvateeva
Hybrid insulin peptides (HIPs) are neoepitopes involved in type 1 diabetes (T1D), but their complete repertoire remains unknown. The vast combinatorial space makes experimental screening unfeasible and requires bioinformatics-based prioritization. We developed a multi-level machine learning pipeline for ranking HIP candidates. First, 36 physicochemical and junction-specific features were computed for a reference library of 240 HIPs with known enzyme-linked immunospot (ELISPOT) reactivity, and a baseline Ridge regression model was trained. Next, all possible HIP candidates with 7-9 amino acid residues per fragment were generated from eight pancreatic β-cell secretory granule source proteins, including insulin chains and C-peptide, islet amyloid polypeptide, chromogranin A, neuropeptide Y, and two secretogranins, yielding 1,057,374 candidates. For each source protein, a local weighted XGBoost (Extreme Gradient Boosting) model was trained using Ridge-score-derived pseudo-labels together with weighted ELISPOT-derived and literature-derived reference HIPs. Finally, anchor-calibrated re-ranking was performed in the global model using cosine similarity to positive anchors (n = 46) and negative anchors (n = 210). The Ridge model achieved 5-fold out-of-fold R² = 0.711 and an area under the receiver operating characteristic curve (AUC) of 0.967. The global model produced a prioritized list of 40 HIP candidates, five per source protein. The highest ranks were observed for candidates with right fragments from neuropeptide Y, secretogranins 1 and 2, islet amyloid polypeptide, and chromogranin A. Candidates carrying the insulin fragment on the right side were systematically down-ranked, suggesting asymmetry in HIP formation. The proposed pipeline reduces the HIP search space from more than one million sequences to a limited set of candidates for experimental validation and provides a framework adaptable to other chimeric neoepitopes in autoimmunity.