Ali Onur Kaya, Mert Can Emre
Background/Objectives: Public thrombin bioactivity records contain heterogeneous endpoints, replicate measurements, related chemical series, and potentially reactive compounds that may bias quantitative structure-activity relationship models. In this study, we developed an explainable, assay-aware, and leakage-safe machine learning framework for predicting thrombin-inhibitory activity. Methods: Exact Ki and IC50 records for human thrombin (CHEMBL204) were standardized, converted to pActivity, and aggregated using predefined criteria. The final dataset comprised 5189 unique compounds represented by development-filtered Mordred descriptors and Morgan fingerprint. The models were optimized using development-only out-of-fold validation and evaluated using random and scaffold-disjoint held-out tests. Results: In the random-split analysis, ConsensusAll achieved an out-of-fold R2 of 0.7588 and a held-out test R2 of 0.7480, with an RMSE of 0.7360, MAE of 0.5307, and concordance correlation coefficient of 0.8540. In the scaffold-disjoint locked test, ConsensusTop5 achieved R2 = 0.5831, RMSE = 0.9484, and MAE = 0.7290, respectively. The applicability domain covered 92.68% of the locked test compounds and yielded R2 = 0.6041. One hundred Y-randomization runs produced a mean R2 of -0.1233 (empirical p = 0.0099). Conclusions: The framework provides useful predictions within the represented chemical space and measurable generalization for unseen scaffolds. This supports compound prioritization, although prospective biochemical validation remains necessary.