科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Journal of the American Medical Informatics Association : JAMIA2026-09-23

Comparative evaluation of ambient digital scribe systems in clinical documentation.

Everett Weiss, Balazs Zsenits, Farhad Nasar, Matt Phillips, Venugopal Mudgundi, Tamer Salhab, Andrew Muth, Kaiwen Zhu, Joanna Kolosko, Sean Hansen, Quang Neo Bui, Ahmad Yaseen, Mostafa Balboul, Jee Lee, Orlando Andres Nunez Isaac

一句话结论 · In one sentence

Pairing clinician Likert assessments with transcript-level error audits may be a useful adjunct in evaluating ADS and other generative AI clinical tools; confirmation in larger studies is needed before this framework drives procurement decisions.

原始摘要(英文原文)· Original abstract
OBJECTIVES: To demonstrate a replicable in-house framework for evaluating ambient digital scribe (ADS) systems and, applying it, to assess 3 commercial systems and examine alignment between clinician Likert ratings and a manual transcript-level error audit. MATERIALS AND METHODS: Two standardized clinical encounters-a brief problem-focused visit and a comprehensive annual visit-were simulated and audio-recorded. Three ADS vendors generated notes from identical recordings. A multidisciplinary panel of 56 physicians rated the notes on 7-point Likert scales across 4 domains: relevance, accuracy, quality, and provider confidence. In parallel, physician reviewers cross-referenced each note against the source recording, enumerating hallucinations, omissions, and inaccuracies using a pre-specified rubric. RESULTS: In this exploratory sample, Vendor A received the highest mean Likert ratings across all 4 domains (Friedman χ²(2) = 15.529, P < .001 short visit; χ²(2) = 10.753, P = .005 long visit). After Bonferroni correction, only Vendor A vs Vendor B on the short visit remained significant (P = .0098). The audit identified fewest errors in Vendor B's notes (n = 3), followed by Vendor A (n = 8) and Vendor C (n = 9). Physicians flagged hallucinations more often than the audit confirmed, a pattern we tentatively label "pseudo-hallucinations." DISCUSSION: In this sample, clinician preference did not track audit-identified error counts, raising the hypothesis that note-style features influence ratings alongside documentation correctness; larger samples are required. CONCLUSION: Pairing clinician Likert assessments with transcript-level error audits may be a useful adjunct in evaluating ADS and other generative AI clinical tools; confirmation in larger studies is needed before this framework drives procurement decisions.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Comparative evaluation of ambient digital scribe systems in clinical documentation. — 科研速览 Science Skim