科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Journal Of Big Data2025-11-18· Computer science

A multimodal fusion model for real-time environment emotion recognition using audio-visual-textual features

Chhaya Gupta, Nasib Singh Gill, Preeti Gulia, Abhinav Kumar, Hanen Karamti, Demmelash Mollalign Moges, Imen Safra

原始摘要(英文原文)· Original abstract
Multimodal combines multiple modalities to create insightful conclusions or to make more precise predictions. Nowadays, the multimodal concept is used to identify human emotions precisely. This study proposes a three-stage novel M-fusHER (Multimodal fusion Human Emotion Recognition) multimodal model for human emotion recognition in real-time with the help of text, audio, and videos. In the first stage, features are extracted with the help of a convolutional neural network merged with multiplicative LSTM. In the second stage, video and audio data, text, and audio are fused in binary form. In the third stage, real-time object detection for human emotion recognition on real videos is implemented. The experimental results are obtained by fusing audio, text, and videos by considering the standard features. For object detection, a fine-tuned YOLOv6 model was used for detecting facial features and expressions from the video. The multiplicative LSTM is also used to extract and learn from the text features. Three datasets, i.e., IEMOCAP, MOSEI, and MELD are used for implementation, and the detection accuracy of the proposed model M- fus HER on IEMOCAP, MOSEI, and MELD datasets is 95.45%, 88.76%, and 95.41% approximately, which is quite encouraging.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

A multimodal fusion model for real-time environment emotion recognition using audio-visual-textual features — 科研速览 Science Skim