Jianbo Li, Zixuan Wang
Experiments on CMU-MOSEI, MELD, and CREMA-D show MetaEmo achieves F1-scores of 87.9%, 86.3%, and 84.7%, outperforming the best baseline by 2.9%-3.4%. It reduces cross-racial accuracy variance to 8.4% and cuts training time by 60% compared to traditional models, verifying efficiency and cultural robustness.
INTRODUCTION: Cross-cultural classroom emotional recognition is challenged by significant cultural disparities in emotional expression and insufficient labeled data across diverse cultural groups, leading to poor generalization in traditional models.
METHODS: This study proposes MetaEmo, a dynamic recognition framework with three core modules. The Transformer-CME module uses cross-modal Transformer architecture to fuse speech prosody, facial micro-expressions, and body language, constructing culture-agnostic emotional representations by filtering cultural noise. The Cultural Adaptive Augmentation Engine (CAAE) employs conditional generative adversarial networks to generate synthetic samples for underrepresented cultures, while a discriminator enhances sensitivity to cultural expression differences. The MAML module optimizes model initialization via meta-training, enabling rapid adaptation to new cultural scenarios with minimal target data.
RESULTS: Experiments on CMU-MOSEI, MELD, and CREMA-D show MetaEmo achieves F1-scores of 87.9%, 86.3%, and 84.7%, outperforming the best baseline by 2.9%-3.4%. It reduces cross-racial accuracy variance to 8.4% and cuts training time by 60% compared to traditional models, verifying efficiency and cultural robustness.
DISCUSSION: MetaEmo effectively addresses cultural differences and data scarcity but has limitations in minority culture coverage. Future work will expand dataset diversity, develop dynamic cultural weighting, and deploy the framework in real-time educational monitoring systems.