科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ PloS one2026-01-01

Zero-shot emotional speech synthesis based on feature decoupling and adaptive loss-threshold reweighting.

Yan Zhu, Yu Wang, Yijin Zhou, Bomin Liu, Rui Zhou

原始摘要(英文原文)· Original abstract
Deep learning-based zero-shot speech synthesis has achieved substantial progress in speaker generalization, but stable modeling remains challenging in fine-grained emotional scenarios. Existing systems often process textual and emotional conditions through shared or closely coupled pathways, which may introduce interference between semantic content and emotional expression. In addition, uniformly averaging per-sample Conditional Flow Matching (CFM) losses may provide insufficient optimization emphasis to high-loss emotional samples. This study proposes a CFM-based zero-shot emotional speech synthesis method. Emotion-Text Decoupling Attention (ETDA) processes semantic and emotional conditions through parallel cross-attention streams and combines them through adaptive gated fusion, allowing the two conditions to retain their respective information before fusion. A 16-class fine-grained emotion space is constructed through classifier filtering and K-Means clustering based on pitch and energy features, and the resulting labels are mapped to continuous representations using a trainable lookup embedding. During training, Adaptive Loss-Threshold Reweighting estimates a threshold from mini-batch loss statistics and assigns larger weights to samples whose individual CFM losses exceed that threshold. Under speaker-disjoint evaluation on ESD, the proposed method achieves a word error rate of 3.62±0.06%, a mel-cepstral distortion of 4.45±0.05 dB, and an emotional expressiveness mean opinion score of 4.64±0.04. Controlled text-length expansion and analyses on a fixed subset of high-loss emotional samples further indicate that the proposed method maintains linguistic content and emotional expression more consistently under the evaluated conditions. These results suggest that the proposed method can, to some extent, improve content preservation, emotional expression, and acoustic reconstruction in fine-grained zero-shot emotional speech synthesis.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Zero-shot emotional speech synthesis based on feature decoupling and adaptive loss-threshold reweighting. — 科研速览 Science Skim