科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Neural networks : the official journal of the International Neural Network Society2026-08-14

Multimodal sentiment analysis with multi-level representation learning and global tri-modality unified fusion.

Li Zelong, Lu Shuhua, Cui Fang, Sheng Chunlei, Liu Chengkai

原始摘要(英文原文)· Original abstract
Multimodal sentiment analysis (MSA) is a popular research topic particularly for predicting human emotional attitudes. However, most existing methods fail to learn the multi-level nonlinear information in MSA due to their usually adopting single-level representation learning, and alternatively suffer from insufficiently extracting the comprehensive correlation features among the triple modalities. Here, to address these issues, we propose a MSA network with Multi-Level representation learning and global Tri-Modality unified fusion, termed as MLTM. Specifically, we devise a multi-level encoding strategy with a hierarchical progressive encoder and multi-level perceptual attention to dynamically weight the information at each level, thereby enhancing the nonlinear representation ability. Furthermore, a dynamic representation optimization mechanism is developed to enhance the semantic relevance of shared features while preserving the uniqueness of private ones. Subsequently, we design a Global Tri-Modality Transformer (GTMT) that first performs parallel fusion of the three modalities and then conducts deep integration guided by the textual modality to achieve the cross-modal semantic alignment and correlation, significantly improving the global unified fusion effectiveness of the tri-modality information. Extensive experiments on three public MSA datasets demonstrate that MLTM outperforms various state-of-the-art methods by a wide margin across various evaluation metrics, indicating its effectiveness and robustness. Specifically, MLTM achieves a superior Acc7 of 55.10 and 48.98, an enhancement of 2.39% and 2.44% compared to the second-best baselines on CMU-MOSEI and CMU-MOSI. Moreover, it reduces MAE to 0.502 and 0.590, improving by 2.14% and 15.7%, on the above two datasets, respectively.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Multimodal sentiment analysis with multi-level representation learning and global tri-modality unified fusion. — 科研速览 Science Skim