YuGuo Wang, Honglin Wang, Ze Wang, Yan Wu
To address the challenges of inconsistent inter-modal contributions, high reliance on labeled data, and limited cross-modal synergy in multimodal emotion recognition, this paper proposes a dual-stream Siamese Self-supervised Emotion Recognition method (DSSER) based on EEG and fNIRS signals. During the pretraining phase, DSSER uses a dual-stream Siamese architecture to extract spatiotemporal features from electroencephalography (EEG) and functional near-infrared spectroscopy (fNIRS) signals. It also aligns their representation spaces through a co-training strategy. This approach eliminates the need for negative samples in conventional contrastive learning. In the fine-tuning stage, a Dynamic Gated Cross-Attention Fusion module is introduced. It assigns modality weights and integrates complementary cross-modal features. Experimental results on a ternary emotion dataset involving 16 participants show that DSSER achieves an accuracy of 98.03% under an intra-subject emotion-recognition protocol. This result reflects performance under the adopted intra-subject setting. It should not be interpreted as subject-independent generalization to unseen participants. The results support the effectiveness of DSSER in improving cross-modal interaction.