科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ PloS one2026-01-01

Hybrid CNN-embedding fusion with MFCC-SVM for speech emotion recognition: Random vs actor-wise evaluation on CREMA-D.

Parveen Kumari, Yogita Yashveer Raghav, Vimmi Kochher, Prashanth Kumar Katta, Abhilasha A, Prabhakar M, Jayanthiladevi A, Fikir Gizachew Belete

原始摘要(英文原文)· Original abstract
Speech Emotion Recognition (SER) is an important component of human-centered intelligent systems, yet robust performance remains challenging when speaker identities differ between training and testing. This study presents a protocol-aware and reproducible comparison on the CREMA-D corpus using three pipelines: (i) a classical MFCC-based Support Vector Machine (SVM), (ii) a log-mel Convolutional Neural Network (CNN), and (iii) a lightweight hybrid model that concatenates handcrafted acoustic descriptors with CNN-derived embeddings and uses an SVM classifier. The methodological contribution is not a new standalone classifier; it is the controlled integration of identical preprocessing, random and actor-wise evaluation, five-seed robustness reporting, class-wise error analysis, and CPU-oriented deployment within one experimental framework. The Hybrid approach achieves the best overall performance, obtaining 62.03% ± 0.94% Macro-F1 on the random split and 58.09% ± 1.36% on the actor-wise split, outperforming MFCC+SVM (55.72% ± 0.95% and 52.05% ± 1.95%) and the log-mel CNN (50.20% ± 1.51% and 42.68% ± 1.84%). A Streamlit interface supports WAV upload, live prediction, and export of per-seed confusion matrices and summary figures. The results show that a controlled lightweight fusion framework can improve robustness while making the performance gap between speaker-overlapping and speaker-independent evaluation explicit.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Hybrid CNN-embedding fusion with MFCC-SVM for speech emotion recognition: Random vs actor-wise evaluation on CREMA-D. — 科研速览 Science Skim