科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ The Journal of the Acoustical Society of America2026-09-01

Spectral envelope estimation of high-pitched vowels using a shared excitation proxy and autoregressive modeling with exogenous input.

Zhao Zhang, Ju Zhang, Wenhuan Lu, Jianguo Wei

原始摘要(英文原文)· Original abstract
Characterizing the vocal tract spectral envelope in high-pitched vowels is challenging due to sparse harmonics and source-filter coupling. This study proposes an analytical framework using a shared glottal driving proxy to guide source-tract decoupling. Speech is decomposed into multi-band Hilbert envelopes; first-order differentiation and non-linear rectification are then applied to isolate transient energy increments during glottal closure. Non-negative matrix factorization captures the cross-band temporal dynamics of these increments, yielding a macroscopic temporal proxy of glottal excitation. This proxy acts as a physical constraint within an auto-regressive with exogenous input model to estimate the underlying vocal tract transfer function. To validate the estimated envelopes, formant estimation error serves as an objective metric on the OPENGLOT database. Particularly under high-pitched conditions where the fundamental frequency (F0) exceeds 300 Hz, the overall mean first formant (F1) estimation error across all evaluated vowels and parameter combinations is 11.43%, yielding an approximate 30% relative reduction from the traditional linear predictive coding baseline (16.51%). The results confirm that this approach provides a complementary pathway for vocal tract envelope modeling.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Spectral envelope estimation of high-pitched vowels using a shared excitation proxy and autoregressive modeling with exogenous input. — 科研速览 Science Skim