Yuan Zhuang, Deqiang He, Zhenzhen Jin, Jian Miao, Juan Lu, Jinxin Wu, Baofu Qin
Fault diagnosis in industrial environments is often constrained by the scarcity of labeled data, variability of operating conditions and cross-device deployment. Moreover, relying solely on vibration signals results in limited diagnostic accuracy and robustness for motors under complex working conditions. To address these challenges, this paper proposes a novel self-supervised motor fault diagnosis framework based on vibration–current multimodal contrastive learning, in which current features are innovatively incorporated into the contrastive learning paradigm. Specifically, a vibration–current multimodal contrastive learning is first developed to enhance discriminative representation learning while improving cross-modal feature consistency, thereby enabling deeper semantic interaction and complementary information fusion between different modalities. Subsequently, to address the adaptive feature encoding of multimodal signals, a dynamically weighted multi-scale residual encoder is constructed, which effectively extracts heterogeneous modal features by adaptively adjusting the importance of multi-scale representations. Finally, the effectiveness of the proposed framework is validated through three case studies. Experimental results demonstrate that an average diagnostic accuracy of 97.75% was achieved across all cross-device tasks using only 5% labeled samples. In addition, the proposed model contains 11.89M parameters and requires only 2.58s for fine-tuning, highlighting its robustness, reliability, and strong application potential in label-limited industrial scenarios.