P. Bodanki, C. Capone, N. S. Hazare, S. Agaron, M. Vijayaraghavan, S. Japa, I. Zaretsky, J. Epstein, A. Sawant, A. Shaikh, E. Leibner, A. Sharma, S. Gangadharan, T. Ghani, H.-H. Tai, D. Ngai, S. El-Haj, P. Parchure, A. Itwaru, B. Kaplan, A. Vakil, R. Tamler, A. Klein, R. Freeman, P. Kovatch, B. Darrow, J. McGreevy, L. Stump, G. N. Nadkarni, P. Timsina, A. Sakhuja
Objective: To evaluate Epic IP Insights, an electronic health record integrated generative AI summarizer used across a seven-hospital health system, using a multidimensional assurance framework. Materials and Methods: We developed and validated an agentic hallucination detector that decomposed summaries into atomic content units (ACUs) and verified each against source documentation. We then applied it to remaining summaries to estimate the hallucination rate. We also measured source note utilization, textual and semantic similarity of regenerated summaries, and clinician perceptions through structured evaluations. Results: The EPIC IP Insights generated 706 summaries across 445 encounters and used a mean of 5.8% of available notes. The detector was developed on 30 summaries (2,012 ACUs) and validated on 5 held out summaries (295 ACUs). The detector agreed with physician adjudication on 96.95% ACUs in the validation set. Across 671 remaining summaries containing 40,452 ACUs, the hallucination rate was 11.79% (95% CI, 11.47% to 12.10%). Among 385 regenerated pairs of summaries, 32.7% were textually identical and 37.9% were semantically identical. Clinicians found the tool easy to use (97.8%), while responses were more mixed regarding reliance on the output with little verification (45.9%), expected efficiency gains (45.3%), and frequent use (46.1%). Discussion: Clinicians found the tool easy to use but believed that outputs required verification. Systematic evaluation identified the frequency of hallucinations and limited use of available notes which provide directions for future improvement. Conclusion: Multidimensional assurance frameworks are needed to evaluate the safety, reliability, and consistency of generative AI tools. Keywords generative artificial intelligence; electronic health records; clinical summarization; hallucination detection; AI assurance; large language