Deepika Rajwade, Aabha Patel, Sayed Athar Ali Hashmi
The ability to accurately perceive human emotions is essential for fluid and natural Human-Robot Interaction (HRI).Although unimodal emotion recognition has advanced considerably, it remains fragile in unconstrained real-world settings due to noise, occlusions, and contextual ambiguity.This paper introduces CCAF-Net, a novel Cascaded Cross-Attention Fusion Network for robust multimodal emotion recognition.CCAF-Net employs a hierarchical architecture that first fuses acoustic and linguistic features from speech via cross-attention, generating a rich audio-linguistic representation.This representation then serves as a contextual query to attend to visual features from facial expressions in a second cross-attention stage.This cascaded design enables dynamic, context-aware weighting of each modality according to its relevance.We evaluate CCAF-Net on the CMU-MOSEI dataset, achieving a weighted F1-score of 88.7% and accuracy of 87.9% on seven-class emotion recognitionoutperforming several state-of-the-art baselines including TFN, LMF, MuIT, and MISA.To assess real-world efficacy, we integrated CCAF-Net into a Pepper robot's control pipeline and conducted a user study (N=30) in a simulated social HRI scenario.Participants rated the CCAF-Net-equipped robot significantly higher than a vision-only baseline across key metrics, including perceived empathy, intelligence, likeability, and interaction naturalness (all p < 0.01).This work highlights the importance of hierarchical, context-aware multimodal fusion for advancing the emotional and social intelligence of autonomous robots in dynamic interactions.