Shiv Garg, Sara Garg
Quantum machine learning (QML) is widely proposed as a successor to classical models in cheminformatics, but rigorous benchmarks under realistic contemporary data constraints remain scarce. We compared two quantum classifiers (quantum support vector machine, variational quantum classifier) against gradient boosting, random forest, and neural network baselines across ten therapeutically diverse protein targets (45-18,286 compounds), using Bemis-Murcko scaffold-stratified cross-validation. Across 48 valid target-threshold tasks, no model class globally dominated (best-quantum versus best-classical Wilcoxon p = 0.189). Classical ensembles excelled on data-rich, structurally coherent targets (e.g., BACE1), while quantum kernels held a marginal edge on sparse, structurally bimodal targets (e.g., Mtb DprE1). A potential predictor of quantum advantage was the per-target best-quantum/best-classical model Cohen's κ gap (Spearman ρ = +0.685, p = 0.029), suggesting that decision calibration, not ranking, drives QML's edge. We propose the κ gap as a preliminary low-cost screening criterion for prioritizing QML in early stage discovery campaigns.