Ashutosh Pandey, Jasmeet Singh, Maninder Kaur
Conversational interactions, rich in both linguistic and vocal cues, provide a natural context for studying these processes. In this work, we propose an explainable multimodal transformer framework that integrates textual semantics (via RoBERTa) and acoustic prosody (via WavLM) to advance emotion understanding. By projecting both modalities into a shared latent space, our model captures the complementary contributions of language and speech to affective communication, achieving an 0.83 accuracy value across five emotion categories. Crucially, we embed explainable AI (XAI) techniques including Integrated Gradients and Occlusion to attribute predictions to specific linguistic tokens and prosodic patterns, thereby aligning computational mechanisms with human cognitive processes of emotion perception. Beyond performance gains, this work demonstrates how multimodal AI systems can support transparent, human-centered emotion recognition.