Yanan Zhang, Chenxu Guo, Kexin Zhu, Wenhui Hu, Bin Hu, Jian Shen
With the rapid advancement of emotion recognition technology, multimodal physiological signals have garnered increasing research attention due to their rich affective representations. However, the substantial heterogeneity across different physiological modalities poses a significant challenge for effective multimodal fusion, limiting the performance of current emotion recognition systems. Moreover, while demographic information inherently encodes valuable emotional cues, its systematic integration into emotion recognition remains underexplored. To address these challenges, we propose R2G$^{3}$Net, a novel hierarchical framework for multimodal emotion recognition. Our model leverages a three-tier architecture: 1) Regional-to-Global Brain Feature Extraction: A BiLSTM-GNN hybrid network hierarchically encodes EEG signals, capturing spatio-temporal patterns from local brain regions to global functional connectivity. 2) Regional-to-Global Cross-Modal Fusion: Peripheral nervous system (PNS) signals are extracted and fused with brain features to enhance physiological representation learning. 3) Regional-to-Global Social Context-Aware Modeling: A hypergraph neural network (HGNN) integrates demographic data to construct dynamic social networks, uncovering higher-order emotional interactions for improved interpretability. Extensive experiments on three benchmark datasets demonstrate R2G$^{3}$Net's superiority in joint spatio-temporal feature learning and social context-aware emotion recognition. Ablation studies and visual analytics further validate that our fused representations outperform state-of-the-art methods in both discriminative capability and model transparency.