Alaa M Elsayad, Khaled A Elsayad
Background/Objectives: Human carbonic anhydrase (CA) isoforms are clinically relevant zinc metalloenzymes, but their highly conserved catalytic Zn2+ site makes isoform-selective inhibitor design challenging. This study developed an explainable multi-isoform QSAR workflow for prioritizing selective inhibitors of human CA I, CA II, CA IX, and CA XII, integrating PubChem experimental concordance and applicability-domain (AD) analysis to distinguish model-supported candidates from exploratory extrapolative hypotheses. Methods: Curated ChEMBL Ki datasets were standardized to pKi and modeled in KNIME using 4185 two-dimensional descriptors, Random Forest-based feature selection, and H2O.ai AutoML stacked ensembles. A matched 3200-compound four-isoform matrix supported direct selectivity profiling. Potent ChEMBL inhibitors (Ki < 10 nM) seeded 90% PubChem similarity expansion; analogues were scored across all four isoform-specific models, while PubChem records were filtered to retain only human isoform-specific Ki evidence. AD was calibrated using Morgan fingerprint similarity, descriptor-space distance, and leverage analysis. Candidates were prioritized by predicted potency, selectivity, SAR plausibility, SwissADME profile, structural alerts, and AD membership. Results: Models achieved held-out R2 values of 0.727, 0.719, 0.652, and 0.607 for CA I, CA II, CA IX, and CA XII, respectively. PubChem concordance identified 1098 prioritized compounds with assay records, including 847 with direct Ki evidence and 428 with complete four-isoform coverage; same-target agreement showed Pearson r > 0.86 and linear-fit R2 > 0.74. Test-set error rose from very-high-AD to low/out-of-domain classes. Conclusions: The workflow integrates QSAR prediction, PubChem concordance, explainable SAR, ADME triage, and AD-based reliability assessment. CA II and CA XII candidates showed the strongest support, CA IX showed intermediate support, and CA I candidates require cautious interpretation as low-domain exploratory hypotheses.