科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ International journal of intelligent engineering and systems2026-08-29· Computer science

Lightweight Swin Transformer with Prosodic Tokens and Huber‑based HuBERT Distillation for Cross‑lingual Speech Emotion Recognition

Maather Alkhafaji, Amir Lakizadeh

原始摘要(英文原文)· Original abstract
Speech emotion recognition (SER) is essential for affect-aware human-computer interaction, yet real-world deployment demands models that are simultaneously accurate, compact, cross-lingually robust, and reproducible.In this paper, we present a lightweight SER framework based on a Swin Transformer that achieves all four goals through three synergistic innovations.First, a dual-branch input fuses high-resolution log-Mel spectrograms with a handful of lightweight prosodic tokens (derived from pitch, energy, and rate dynamics) at negligible overhead, enriching the representation with supra-segmental cues.Second, a stage-wise local-window attention schedule (3×3→5×5→7×7→7×7) progressively expands the receptive field in deeper layers, capturing fine time-frequency micro-structure in early stages and broader prosodic phrases in later ones.Third, we distill knowledge from a frozen HuBERT teacher using a feature-alignment loss; the Huber loss proves superior to MSE and L1, yielding consistent gains.Beyond within-language tests, we perform an initial cross-lingual generalization experiment -training on four English corpora (CREMA-D, RAVDESS, SAVEE, TESS) and testing zero-shot on unseen German EMO-DBachieving 63.44% weighted accuracy with our compact student.The full model attains state-of-the-art or competitive accuracy on five benchmarks (94.02% on CREMA-D, 96.32% on EMO-DB, 96.59% on RAVDESS, 98.08% on SAVEE, 99.89% on TESS) while requiring only 0.85 million parameters and 0.32 GFLOPs, making it approximately 33× lighter than a standard Swin-Tiny model.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Lightweight Swin Transformer with Prosodic Tokens and Huber‑based HuBERT Distillation for Cross‑lingual Speech Emotion Recognition — 科研速览 Science Skim