Chethana Saligram, Vishal Shrivastava, Marisha Speights
Co-articulation transition magnitudes varied systematically by age and clinical group. Attention improved temporal classification specifically in younger children, and PER dropped markedly between ages 3-4 and 4-5. On long-form naturalistic narratives, CTC training and inference remained numerically stable, but greedy decoding produced degenerate output dominated by blank and repetition tokens.
INTRODUCTION: Pediatric speech sound disorders (SSDs) affect many young children and are commonly assessed through auditory-perceptual judgments and IPA transcription, which are limited by listener bias, variable interrater reliability, and difficulty attributing deviations at the phoneme level. We present a developmentally informed framework for pediatric vocal biomarker foundations that prioritizes age-aware, phoneme-resolved interpretability over utterance-level accuracy alone.
METHODS: Using a CAAP-derived subset of the SEED corpus (27-94 months) with clinically specified phoneme targets and SSD labels, we construct age-stratified phoneme profiles and quantify error structure via phoneme error rate (PER) and age-banded confusion signatures. We operationalize co-articulation as transition dynamics from MFCC trajectories and their first and second derivatives, reflecting the velocity and acceleration of spectral change across adjacent segments. We learn compact acoustic representations with a variational autoencoder (VAE) and model temporal evolution with BiLSTMs, including attention, to characterize disorder-relevant instability in latent trajectories. For phoneme transcription, we train BiLSTM-CTC sequence models on clinically elicited speech and evaluate disorder classification and phoneme substitution patterns within age groups, then stress-test generalization on ECSC "Frog Story" narratives from TalkBank/CHILDES.
RESULTS: Co-articulation transition magnitudes varied systematically by age and clinical group. Attention improved temporal classification specifically in younger children, and PER dropped markedly between ages 3-4 and 4-5. On long-form naturalistic narratives, CTC training and inference remained numerically stable, but greedy decoding produced degenerate output dominated by blank and repetition tokens.
DISCUSSION: Together, these results support an age-aware approach to pediatric phoneme analytics, in which articulatory coordination, not just phoneme identity, carries developmentally and clinically relevant information; and phoneme-level performance should be interpreted relative to developmental stage rather than a single fixed benchmark. The decoding failures on naturalistic speech indicate a bottleneck specific to decoding rather than to the underlying representations, motivating constrained, hierarchical, or duration-aware decoding as a tractable next step. These findings position age-stratified, phoneme-resolved analysis as a foundation for interpretable and scalable pediatric screening and biomarker development.