Shiv Patil, Mert Karabacak, Matthew Southerby, Konstantinos Margetis
Large language models (LLMs) have moved rapidly from experimental demonstrations to early clinical, educational, and administrative use. The first phase of medical LLM discourse emphasized opportunities in transfer learning, domain adaptation, clinical decision support, education, privacy, fairness, and regulation. The next phase requires a more specific framework for clinical stewardship. Since 2023, medical LLMs have demonstrated substantial progress in medical question answering, diagnostic reasoning in interactive clinical dialogue, patient communication, documentation support, and retrieval-augmented knowledge synthesis. At the same time, prospective and randomized evaluations show that model performance alone does not guarantee improved physician reasoning, safer decisions, or better patient outcomes. The central challenge is therefore implementation: defining appropriate use cases, validating models in context, training clinicians, protecting patients, and monitoring deployed systems across their lifecycle. This Viewpoint proposes a practical roadmap for medical LLM adoption. We distinguish lower-risk administrative and communication tasks from higher-risk diagnostic, triage, and treatment tasks. We argue for tiered evidence standards, transparent reporting, local validation, human oversight, equity auditing, patient-centered consent, and continuous post-deployment surveillance. The regulatory environment has also matured, with risk-based frameworks and lifecycle-oriented oversight increasingly shaping how these systems are evaluated and maintained. LLMs should be treated as workflow-embedded sociotechnical systems, whose performance depends on the interaction among models, clinicians, patients, interfaces, local data, and governance. Responsible adoption will require collaboration among clinicians, clinical informaticians, health systems, patients, regulators, ethicists, and developers. Realizing the promise of medical LLMs depends on a shift from fascination with fluent outputs toward disciplined stewardship of clinical performance, accountability, equity, and trust.