科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Mathematics2026-05-19· Computer science

SGMT with S-PACE: A Framework for Temporal Alignment and Quality-Aware Multimodal Fusion in Emotion Recognition

Jun-Young Ahn, Sathiyamoorthi Arthanari, Sathishkumar Moorthy, Yeon-Kug Moon

原始摘要(英文原文)· Original abstract
Multimodal emotion recognition is challenging because behavioral signals and physiological responses evolve at different temporal rates. Facial expressions and speech often change rapidly after an emotional event, whereas peripheral biosignals such as electrodermal activity, blood volume pulse, and skin temperature exhibit delayed and smoother dynamics. This temporal inconsistency can degrade fusion performance, particularly in real-world recordings with noisy or missing modalities. To address this issue, this study proposes SGMT, an S-PACE Gated Multimodal Transformer for emotion recognition using speech, facial video, and physiological signals. The proposed SGMT introduces S-PACE, a physiology-guided cross-attention mechanism that aligns fast behavioral cues with slower biosignal representations without assuming a fixed temporal delay. A Quality-Aware Gate further improves robustness by adaptively weighting modalities according to signal reliability. The fused representations are processed using a Temporal Swin Transformer and a Perceiver Fusion module for arousal–valence prediction and emotion quadrant classification. Experiments are conducted on the Korean multimodal emotion datasets KEMDy20 and K-EmoCon under different modality settings. SGMT achieves arousal UARs of 68.4% on KEMDy20 and 62.9% on K-EmoCon, with quadrant accuracies of 44.7% and 62.5%, respectively. Ablation studies demonstrate that the proposed alignment and gating strategies provide more stable multimodal fusion than conventional feature concatenation. The results indicate that SGMT effectively adapts to varying modality availability and improves multimodal emotion recognition in naturalistic environments.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

SGMT with S-PACE: A Framework for Temporal Alignment and Quality-Aware Multimodal Fusion in Emotion Recognition — 科研速览 Science Skim