科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Journal of voice : official journal of the Voice Foundation2026-08-15

Optimal Acoustic Parameter Subsets for Dimension-Specific Voice Quality Prediction.

Jorge C Lucero

一句话结论 · In one sentence

Different perceptual dimensions of voice quality are optimally predicted by distinct subsets of acoustic parameters. Log-transforming right-skewed perturbation measures improves linear prediction incrementally for overall severity and substantially for roughness. Using sustained vowels alone, data-driven parameter selection matches or exceeds the classification performance of established multiparametric indices scored on more speech material, and identifies candidate parameter sets for roughness and strain-dimensions lacking published composite indices. External validation and extension to connected speech remain necessary before clinical adoption.

原始摘要(英文原文)· Original abstract
OBJECTIVES: To determine the minimal set of acoustic parameters that optimally predicts each major perceptual voice quality dimension and to quantify the redundancy among commonly measured acoustic voice parameters. STUDY DESIGN: Secondary analysis of a publicly available voice database. METHODS: Fourteen acoustic parameters were extracted from sustained vowels of 284 speakers in the Perceptual Voice Qualities Database (PVQD). Four right-skewed parameters (jitter, shimmer, shimmer dB, and period standard deviation) were log-transformed to linearize their relationship with perceptual ratings. QR factorization with column pivoting characterized the redundancy structure among parameters. Orthogonal Matching Pursuit (OMP) ranked parameters by incremental predictive value for each perceptual dimension (CAPE-V Severity, Breathiness, Roughness, and Strain; GRBAS Grade, Breathiness, and Roughness). The optimal number of parameters was determined by convergence of adjusted R2, Bayesian Information Criterion, and cross-validated prediction error. Classification performance against dichotomized perceptual ratings was evaluated using ROC analysis with 10-fold stratified cross-validation. RESULTS: The 14-parameter acoustic space had an effective dimensionality of approximately 11, with slope contributing negligible independent information. Log transformation improved predictive accuracy for all dimensions, except breathiness, with gains that were modest for overall severity and substantial for roughness. Optimal parameter subsets varied across perceptual dimensions: overall severity required six parameters (log shimmer dB, GNE, Hno-6000, log PSD, F0, and HNR-D; rs = 0.77, AUC = 0.870), breathiness required three (CPPS, GNE, and Hno-6000; rs = 0.73, AUC = 0.865), roughness required 5 (log shimmer dB, H1-H2, GNE, log PSD, and log jitter; rs = 0.71, AUC = 0.830), and strain required six (HNR, F0, log PSD, tilt, H1-H2, and GNE; rs = 0.66, AUC = 0.790). The severity subset yielded a higher area under the curve (AUC) than the Acoustic Voice Quality Index (AVQI: six parameters, AUC = 0.827) and the Cepstral Spectral Index of Dysphonia (CSID: AUC = 0.787). This advantage was obtained from sustained vowels alone, whereas AVQI and CSID were scored on connected speech in addition to the vowel; the comparison is therefore task-asymmetric and favors the established indices in speech material. Notably, log shimmer dB displaced smoothed cepstral peak prominence (CPPS) as the most important predictor of overall severity, suggesting that the widely reported primacy of CPPS may partly reflect its distributional advantage in linear models. CONCLUSIONS: Different perceptual dimensions of voice quality are optimally predicted by distinct subsets of acoustic parameters. Log-transforming right-skewed perturbation measures improves linear prediction incrementally for overall severity and substantially for roughness. Using sustained vowels alone, data-driven parameter selection matches or exceeds the classification performance of established multiparametric indices scored on more speech material, and identifies candidate parameter sets for roughness and strain-dimensions lacking published composite indices. External validation and extension to connected speech remain necessary before clinical adoption.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Optimal Acoustic Parameter Subsets for Dimension-Specific Voice Quality Prediction. — 科研速览 Science Skim