Helene Springer, Henrik Garde, Frida Splendido, Marianne Gullberg
Seeing a speaker's face can facilitate speech comprehension, but the magnitude of this effect depends on the visual salience of speech sounds. Visual salience has traditionally been defined through descriptions of speakers' articulatory gestures using discrete binary features. We examined which contrasts among Swedish long vowels are visually different and whether these align with traditional binary feature descriptions based on vowel openness and lip rounding. This proof-of-concept case study quantifies lip configurations in continuous speech to characterize the visual information available to listeners. Using a computer-vision approach, we extracted the area, height, and width of the mouth opening from video-recorded Swedish speech produced by a single speaker. Despite gradual variation within the vowel categories, the results reveal visually distinguishable groups of vowels. Crucially, they suggest that production-based binary feature matrices can misrepresent the visual information available to addressees, particularly in continuous speech. We discuss the advantages of quantitative visual parameters over traditional discrete features based on tongue position and lip rounding for describing visual speech. We also evaluate the use of computer vision for extracting facial landmarks and measuring lip configurations.