Samuel S Agboola, Oluwaseun E Agboola, Oluwaseun Ruth Olasehinde, Temitope C Aribigbola, Anuoluwapo Bukola Shaleye, I Adefolake Alebiosu, Zainab A Ayinla, Babatunji E Oyinloye, Olutosin S Ilesanmi, Idowu O Omotuyi, Abel Kolawole Oyebamiji
Activity cliffs, that is pairs of closely related molecules with large potency differences, limit the reliability of structure-activity models, and it is commonly assumed that the chemistry responsible for a cliff carries information that generalizes to other targets. We tested this assumption directly across the protein kinase family using machine-learning models trained on measured bioactivity data. From 16,897 Ki and Kd measurements for 20 human kinases (8661 unique standardized structures) we generated 186,663 single-cut matched molecular pairs and labelled 116,845 as activity cliffs (|ΔpActivity| ≥ 2.0) or smooth pairs (≤0.5). Three findings emerged. First, cliff-forming transformations rarely recur across targets: 97.1% occurred on a single kinase, and agreement among those that did recur was indistinguishable from zero (median r = -0.058, 95% CI -0.13 to 0.41). Second, leave-one-kinase-out random forests exceeded both a prevalence baseline (average precision 0.111) and a transformation-frequency baseline (0.149), reaching 0.238. Third, the apparent advantage of within-target learning proved to be a validation artefact: under leakage-controlled grouped cross-validation, within-kinase performance fell from 0.821 to 0.309 (scaffold grouping) and 0.211 (molecular-component grouping) and became statistically indistinguishable from cross-kinase performance (p = 0.31 and p = 0.17). Activity cliffs are therefore difficult to predict in general rather than specifically difficult to transfer, and inferences about target specificity are highly sensitive to how within-target validation is grouped.