Irene L Hudson, Anders Yeo, Sean Andrew Hudson, David Akman
Use of two novel Machine Learning algorithms, Simultaneous perturbation Instance Selection (SpIS) and Simultaneous perturbation Feature and Instance Selection (SpFIS) enabled identification (and visualization) of unique Kinase Protein Interactions (KPIs), so-called instances and their features. The resultant minimal subset of both instances and features, based on the full data set of n = 188 KPIs and p = 83 features, successfully characterized the six human kinase groups: KG = 1 (TK), KG = 2 (Other), KG = 3 (CAMK), KG = 4 (AGC), KG = 5 (CMGC), KG = 6 (STE). A new partition of eight groups of KPIs based on deep Gaussian mixture clustering (deepGMM, M1DGMM), and a parsimonious reduced data set of 35 peptide-centric interaction features for the full 188 KPIs, also well characterized by a minimal subset of instances and features. Irrespective of the target groups, whether the six human kinase groups or the eight M1DGMM partition, the minimal subsets selection chose between 10 and 12 unique KPIs (with high fivefold cross validation accuracy). Also, minimal subset selection provided between 13 and 19 features (with high fivefold cross validation accuracy). Importantly, the minimal subsets of instances and features derived from SpFIS and SpIS differed according to whether the target group was Kinase group or M1DGMM. This highlights the value of our new classification based on DGMM. KPIs such as 3ALO, 1UKI, and 4UBX, so-called outliers, gave further insights into structural differences with potential to add to drug discovery innovation. There is potential for other peptide-protein sets (not just Kinases) to be similarly explored, with extensions to protein-protein (not just peptide) interactions.