U. Genske, A. Laudani, L. Yan, Y. Peng, G. Boening, S. T. Ulas, M. P. Wagner, T. Diekhoff, B. Hamm, P. Jahnke
Clinical deployment of medical imaging artificial intelligence (AI) requires objective and continuous quality assurance, yet standardised methods for this purpose have not been established. Here, we present a framework using physical phantoms for standardised on-site testing and monitoring of AI, demonstrated in CT-based liver lesion detection. We begin by designing phantoms tailored to the anatomical input domain expected by AI algorithms, and then systematically assess how AI performance is affected by variations in scanner technology and operation across two clinical CT systems. Next, we perform longitudinal monitoring, yielding consistent results over fifteen months on both systems. Finally, we validate clinical relevance by demonstrating that AI models trained on phantom data generalize effectively to patients and exhibit no evidence of phantom-specific adaptation. Our findings show that clinically realistic phantoms enable standardised, site-specific testing and monitoring of AI, providing a proactive method for local and cross-institutional quality assurance.