科研速览 · Science Skim继续刷下去 · Keep skimming →
◇ medRxiv2026-09-14· health informatics

Beyond word error rate: clinical risk as the necessary standard for ambient AI scribe evaluation: evidence from 77 global languages

H. Bergman, V. Liu, B. Austin, R. Sanghera

原始摘要(英文原文)· Original abstract
Abstract Objective Ambient AI scribes evaluated using frequency-based metrics such as word error rate (WER), which do not represent clinical consequence. We tested whether variation in these metrics tracks consequential transcription errors. Methods We constructed a multilingual corpus from five clinical dictation scripts spanning a complexity gradient, translated into 99 languages, rendered to synthetic speech under three acoustic conditions, and transcribed by a production ambient scribe. Six frequency metrics were computed. Three independent large language model raters from external providers assessed clinically meaningful error patterns in context using a Severity x Likelihood framework informed by UK digital clinical-safety-risk-management principles. Results Across 59,819 genuine transcription-error occurrences, 58,329 (97.5%) were LOW risk and 251 (0.42%) CRITICAL or HIGH. None of six frequency metrics showed a statistically detectable association with serious clinical risk across languages; correlations were small (absolute Spearman {rho}<0.16). A Severity x Likelihood sum remained strongly correlated with WER ({rho}=0.80), showing that the aggregate remained dominated by benign errors. At complexity level 3, low-resource languages had worse WER than high-resource languages ({beta}=+0.078, 95% CI +0.045 to +0.111; p<0.0001), without a detectable difference in CRITICAL/HIGH risk (OR 1.21, 95% CI 0.43 to 3.43; p=0.72). Consultation complexity was the principal predictor of serious risk (OR 3.06 per level, p<0.0001). Conclusion Across this controlled multilingual corpus, aggregate transcription-frequency metrics did not reliably track the sparse severe tail of clinically consequential errors. WER remains appropriate for transcription quality, but these data do not support its use alone as a proxy for clinical safety. Context-aware assessment of error consequence provides complementary information that frequency measures can dilute.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Beyond word error rate: clinical risk as the necessary standard for ambient AI scribe evaluation: evidence from 77 global languages — 科研速览 Science Skim