Kathrin Cresswell, Robin Williams
Although artificial intelligence (AI) shows considerable promise in healthcare, its introduction and safe and effective exploitation involve several challenges. We argue that to maximise the benefits of AI and understand, track and mitigate potential harms there must be a shift in how AI is evaluated. Evaluation should be more closely linked to the local contexts in which AI is deployed and used and better integrated with clinical practice. To support this shift, we propose three interrelated levels of evaluation: evaluation of implementation, which examines transformations in contextually situated practices; evaluation of optimisation, which explores changes to clinical pathways over extended periods; and evaluation of scaling, which focuses on coordination through regional and national structures. This approach offers significant opportunities to advance both healthcare practice and the epistemology of medical knowledge.