L. Chen, M. Zheng
Establishing reliable associations between visual appearance and dynamic behavior is essential for multimodal perception systems that require semantic consistency across heterogeneous data sources. This study proposes a cultural semantic-driven framework for modeling the relationship between ethnic dance costumes and movement expression through cross-modal representation learning. Clothing images are encoded using CLIP to extract texture, color, and structural features, while spatiotemporal characteristics of dance movements are captured by ST-GCN from threedimensional skeletal sequences. Cultural keywords are introduced as semantic anchors, and a cross-modal InfoNCE objective aligns visual and motion embeddings within a shared feature space under weak supervision. To improve interpretability, cross-modal attention and Grad-CAM are employed to generate association heatmaps that reveal the correspondence between costume elements and movement semantics. Experimental evaluations on representative ethnic dance datasets demonstrate superior semantic consistency, lower semantic alignment error, robust interpretability under increasing motion complexity, and improved agreement with human cognitive judgments compared with mainstream multimodal models. Educational intervention experiments further confirm significant gains in cultural knowledge acquisition and aesthetic judgment stability. By integrating visual feature extraction, semantic alignment, and interpretable multimodal reasoning into a unified architecture, the proposed framework provides a transferable computational strategy for image– motion information fusion and perception-oriented intelligent systems, offering methodological insights for multidisciplinary engineering applications involving multimodal sensing, feature transmission, and semantic association analysis.