科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Alexandria Engineering Journal2026-05-22· Computer science

A mamba-driven transformer method for multimodal emotion recognition optimization

Dongwei Ren, Song Guo

原始摘要(英文原文)· Original abstract
Emotion recognition, a core technology of affective computing, directly impacts human–computer interaction. However, existing CNN and Transformer-based models have limitations: CNNs struggle with long sequence data due to local receptive fields, while Transformers face quadratic computational complexity with increasing sequence lengths, limiting complex emotion state recognition. To address these challenges, we propose the TriModalMam model, a multimodal emotion prediction framework. TriModalMam integrates the advantages of the Mamba architecture, optimizing emotion feature extraction and fusion by combining text, audio, and visual features. The model uses the Monomodal Sequential Mamba (MSM) module for deep feature extraction, projects modality features into a similarity subspace through shared encoders, and optimizes feature representations with Mamba-driven Multi-modal Extraction (MMFE). Cross-Modal Interaction (CMI) enhances information flow and interaction between modalities, and the fused features are processed by a Transformer encoder and MLP network for emotion prediction. Experiments on the CMU-MOSI and CMU-MOSEI datasets show that TriModalMam outperforms traditional CNN and Transformer models. On CMU-MOSI, the seven-class accuracy reaches 87.34, with a correlation of 0.855; on CMU-MOSEI, accuracy reaches 87.45, with correlation of 0.891. The model has only 110M parameters, balancing high performance and low computational complexity.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

A mamba-driven transformer method for multimodal emotion recognition optimization — 科研速览 Science Skim