Maria Cairoli, Morten Nielsen, Catherine Betts, Olga Obrezanova, Leonardo De Maria
Peptide-level fingerprints showed poor performance due to loss of positional information, while residue-level fingerprints matched the performance of sequence-based encodings (BLOSUM62 and one-hot) for NAA-based peptides while accurately identifying binding cores and motifs. Considering citrullinated peptides as a case study, we observed similar linear correlation performance across encoding strategies, with residue-level fingerprints showing marginal improvements in quantitative prediction accuracy. Citrulline was rarely found at canonical anchor positions, suggesting that the advantage of explicit chemical encoding may be more pronounced for modifications occurring at positions critical for binding.
INTRODUCTION: Peptide therapeutics are increasingly explored to treat challenging diseases, but immunogenicity risks limit their clinical success. In silico tools enable immunogenicity screening through prediction of peptide-MHCII binding, yet current methods do not account for chemical properties of non-natural amino acids routinely incorporated to improve peptide properties.
METHODS: Here, we present a machine learning approach combining chemical fingerprints with sequence information to predict MHC class II binding for peptides including both natural (NAA) and non-natural (NNAA) amino acids. We evaluate peptide-level fingerprints against residue-level fingerprints (direct-encoding and similarity-based fingerprints) that preserve positional information while encoding chemical diversity, for a total of 31 different representations.
RESULTS: Peptide-level fingerprints showed poor performance due to loss of positional information, while residue-level fingerprints matched the performance of sequence-based encodings (BLOSUM62 and one-hot) for NAA-based peptides while accurately identifying binding cores and motifs. Considering citrullinated peptides as a case study, we observed similar linear correlation performance across encoding strategies, with residue-level fingerprints showing marginal improvements in quantitative prediction accuracy. Citrulline was rarely found at canonical anchor positions, suggesting that the advantage of explicit chemical encoding may be more pronounced for modifications occurring at positions critical for binding.
DISCUSSION: The proposed framework allows for inclusion of diverse modifications, including synthetic amino acids, and is compatible with existing pan-allele architectures. While the full advantage of explicit chemical encoding remains to be demonstrated, this framework is designed to capture these effects as experimental data becomes available, supporting immunogenicity risk assessment for emerging peptide therapeutics.