Shuyan Zhang, Bowen Shi, Jiongzhou Wu, Jinghao Fang, Linglong Yin, Jisheng Chen, Jin Yang
Artificial intelligence now guides choices across small-molecule discovery, yet benchmark gains rarely reveal whether a model changed the next compound, assay or project decision. This Perspective organizes public evidence around Design, Make, Test and Analyze (DMTA) handoffs. It distinguishes reproducible benchmark evidence (C1), use-relevant retrospective robustness (C2), executed Make or Test handoffs (C3), prospective workflow comparisons (C4-W) and traceable downstream trajectories (C4-D), with separate Evaluation, Deliverability and Governance profiles. Two authors independently screened 99 public sources and classified a reconciled inventory of 75 claim units. Consensus coding assigned 41 claims to C1, 6 to C2, 19 to C3 and 9 to Context; none met C4-W or C4-D. These counts characterize the purposive public-evidence map. Four pressure tests examine structural prediction, low-data active learning, ultra-large screening and reward hacking. Stronger claims depend on inspectable worklists, denominators, negative outcomes, decision rules and provenance. The proposed reporting framework separates workflow comparison from downstream traceability and supports tiered disclosure when compound details are commercially sensitive.