科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Journal of Intelligence2025-12-03· Computer science

Bridging Text and Speech for Emotion Understanding: An Explainable Multimodal Transformer Fusion Framework with Unified Audio–Text Attribution

Ashutosh Pandey, Jasmeet Singh, Maninder Kaur

原始摘要(英文原文)· Original abstract
Conversational interactions, rich in both linguistic and vocal cues, provide a natural context for studying these processes. In this work, we propose an explainable multimodal transformer framework that integrates textual semantics (via RoBERTa) and acoustic prosody (via WavLM) to advance emotion understanding. By projecting both modalities into a shared latent space, our model captures the complementary contributions of language and speech to affective communication, achieving an 0.83 accuracy value across five emotion categories. Crucially, we embed explainable AI (XAI) techniques including Integrated Gradients and Occlusion to attribute predictions to specific linguistic tokens and prosodic patterns, thereby aligning computational mechanisms with human cognitive processes of emotion perception. Beyond performance gains, this work demonstrates how multimodal AI systems can support transparent, human-centered emotion recognition.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Bridging Text and Speech for Emotion Understanding: An Explainable Multimodal Transformer Fusion Framework with Unified Audio–Text Attribution — 科研速览 Science Skim