H. Bergman, V. Liu, B. Austin, R. Sanghera
Abstract Objective Ambient AI scribes evaluated using frequency-based metrics such as word error rate (WER), which do not represent clinical consequence. We tested whether variation in these metrics tracks consequential transcription errors. Methods We constructed a multilingual corpus from five clinical dictation scripts spanning a complexity gradient, translated into 99 languages, rendered to synthetic speech under three acoustic conditions, and transcribed by a production ambient scribe. Six frequency metrics were computed. Three independent large language model raters from external providers assessed clinically meaningful error patterns in context using a Severity x Likelihood framework informed by UK digital clinical-safety-risk-management principles. Results Across 59,819 genuine transcription-error occurrences, 58,329 (97.5%) were LOW risk and 251 (0.42%) CRITICAL or HIGH. None of six frequency metrics showed a statistically detectable association with serious clinical risk across languages; correlations were small (absolute Spearman {rho}<0.16). A Severity x Likelihood sum remained strongly correlated with WER ({rho}=0.80), showing that the aggregate remained dominated by benign errors. At complexity level 3, low-resource languages had worse WER than high-resource languages ({beta}=+0.078, 95% CI +0.045 to +0.111; p<0.0001), without a detectable difference in CRITICAL/HIGH risk (OR 1.21, 95% CI 0.43 to 3.43; p=0.72). Consultation complexity was the principal predictor of serious risk (OR 3.06 per level, p<0.0001). Conclusion Across this controlled multilingual corpus, aggregate transcription-frequency metrics did not reliably track the sparse severe tail of clinically consequential errors. WER remains appropriate for transcription quality, but these data do not support its use alone as a proxy for clinical safety. Context-aware assessment of error consequence provides complementary information that frequency measures can dilute.