Ethan Waisberg, Joseph W Guarnieri
Clinical medicine evaluates decision-support technology using methods built around a single output presented for clinician review. Agentic artificial intelligence does not fit this design. Agent scaffolding, the software infrastructure surrounding a large language model that turns passive text generation into active, multi-step execution, allows the model to retrieve information, invoke external tools, and act on the results before a clinician sees any of it. The relevant unit of clinical risk therefore shifts from a single inference to a trajectory of actions. We describe four challenges this creates for responsible development: silent error propagation across multi-step tasks, oversight that reviews conclusions rather than processes, validation that does not survive changes to model or tooling, and scope that widens faster than the evidence supporting it. We argue that each is addressed by a clinical testing harness: a structured evaluation environment comprising scenario libraries built from clinical edge cases, full-trajectory observability, explicit escalation testing, and staged evidence thresholds tied to scope of practice. Medicine already possesses these tools in the form of simulation, credentialling, and morbidity and mortality review, and requires their adaptation rather than their invention.