科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Journal of speech, language, and hearing research : JSLHR2026-09-22

Automatic Prediction of Vocal Strain Scores in Singing Voice Using Audio and Electroglottographic Modalities.

Yuanyuan Liu, Okko Räsänen, Tero Ikävalko, Tua Hakanpää, Vesa Ronkainen, Anne-Maria Laukkanen

一句话结论 · In one sentence

Despite the higher numerical performance of audio-based models, EGG features demonstrated robust predictive potential and superior interpretability within the context of professional singing. The results confirm that combining multimodal data with rigorous feature selection provides a robust framework for the objective assessment of singing voice strain. Direct linkage analysis further verified the tight coupling between glottal and acoustic parameters, grounding the multimodal approach in proven vocal physics.

原始摘要(英文原文)· Original abstract
PURPOSE: This study developed machine learning models to predict perceptual strain scores in the singing voice using audio and electroglottographic (EGG) recordings. The study examined the predictive capability of distinct feature sets extracted from audio and EGG modalities and assessed the contributions of participant metadata (META) and feature selection to model performance. METHOD: Data were split into mutually exclusive train-validation (n = 16 singers, n ≈ 450 samples) and independent test (n = 11 singers, n ≈ 240 samples) sets to ensure singer-independent generalization. Extracted features included mel-frequency cepstral coefficients, extended Geneva minimalistic acoustic parameter set, wavelet scattering coefficients, and domain-specific descriptors (audio: amplitude modulation; EGG: glottal dynamics), along with singer META. Using leave-one-singer-out cross-validation, we trained regression models (support vector regressor, random forest, and ridge) to predict the expert strain ratings. Recursive feature elimination was employed to optimize feature subsets, and feature-level fusion was implemented to assess multimodal integration. RESULTS: During the training-validation phase, the highest Spearman correlations were ρ = .823 (p < .001) for the audio-only model and ρ = .777 (p < .001) for the EGG-only model. On the held-out test set, the audio-only model achieved ρ = .812 (p < .001), while the EGG-only model reached ρ = .759 (p < .001). The optimal overall performance on the test data set (ρ = .825, p < .001) was achieved by a ridge model integrating selected features from audio and META. CONCLUSIONS: Despite the higher numerical performance of audio-based models, EGG features demonstrated robust predictive potential and superior interpretability within the context of professional singing. The results confirm that combining multimodal data with rigorous feature selection provides a robust framework for the objective assessment of singing voice strain. Direct linkage analysis further verified the tight coupling between glottal and acoustic parameters, grounding the multimodal approach in proven vocal physics.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Automatic Prediction of Vocal Strain Scores in Singing Voice Using Audio and Electroglottographic Modalities. — 科研速览 Science Skim