Andrea Urqueta Alfaro, Karina A Roundtree, Mark S Pfaff, Nathan Bos, John D Edell, Ronna Ten-Brink
Findings support the need to include diverse accents in ASR engine training and assessment as well as evaluation of assistive technologies that leverage ASR applications. Results also support the use of the SSM in future research with accented speech, including investigating which accuracy measures are more appropriate (i.e., better alignment with human judgement) for comparing performance amongst ASR engines.
PURPOSE: Automated Speech Recognition (ASR) is used in assistive technologies such as caption phone and video calls for people who are deaf or hard of hearing. Evidence indicates that ASR's poor caption accuracy with "accented" speech (not the "standard" United States Broadcasting Mid-Western accent) is a barrier to successful communication. Efforts to improve ASR shortcomings should evaluate how caption errors alter the meaning intended by the speaker. Prior studies measured caption accuracy only using Word Error Rate (WER) which does not account for impact on meaning. This study expands existing work by assessing three major ASR engines' (Microsoft Azure, IBM Watson, Google Speech to Text) accuracy with a wide range of diverse accents using a measure of semantic similarity (SSM) between what the speaker said and how it was captioned, in addition to WER.
RESULTS: ASR caption accuracy is worse for accented speech compared to the Broadcasting Mid-Western accent, as measured by SSM and WER (p < 0.001). However, results were mixed when investigating whether some ASR engines perform better than others with accented speech; Azure performed better based on WER (p = .026), but not based on SSM.
CONCLUSION: Findings support the need to include diverse accents in ASR engine training and assessment as well as evaluation of assistive technologies that leverage ASR applications. Results also support the use of the SSM in future research with accented speech, including investigating which accuracy measures are more appropriate (i.e., better alignment with human judgement) for comparing performance amongst ASR engines.