Lisong Ou, Zhixin Li
The rapid development of mobile internet has turned multimodal sentiment analysis (MSA) into a prominent research focus. Despite the progress achieved by existing models, the heterogeneity problem resulting from modality differences presents significant challenges to MSA. Furthermore, while much focus has been placed on dynamic learning within a single sample, the learning of inter-sample and inter-class relationships has been largely overlooked. To tackle these problems, we propose a multi-scale text-guided progressive fusion network. Initially, a domain adaptation strategy is employed to fine-tune BERT model, enhancing its support for downstream tasks. Subsequently, we design a progressive fusion module that centers on text as the core modality, guiding the integration of audio and visual modalities to generate more enriched derivative feature representations. Within the cross-modal transformer, the integration of core and super-modal features captures fine-grained cross-modal dependencies. Furthermore, we introduce dual contrastive learning to facilitate the modeling of inter-sample and cross-category relationships, which further alleviates the impact of distributional discrepancies among modalities. Experimental results on the widely used CMU-MOSI, CMU-MOSEI, and CH-SIMS datasets demonstrate that the proposed method outperforms existing methods in multimodal sentiment classification tasks. The method achieved sentiment binary classification accuracies of 84.90%, 85.03%, and 81.76% on these datasets, and its effectiveness is further validated through multiple experiments.