Kehui Zhu, Xinming Li, Z.-Y. Liu, Jinrui Zhang, Yuzhou Wang
Deep learning has achieved significant progress in industrial equipment fault diagnosis. However, large-scale models are challenging to deploy in industrial settings that require low power consumption and high portability. Moreover, the high latency introduced by model inference reduces the efficiency of real-time fault detection. In this study, we propose a lightweight fault diagnosis network and develop a convolutional neural network (CNN) acceleration for field-programmable gate array (FPGA) deployment. Firstly, a progressive knowledge distillation strategy is proposed to reduce the discrepancy between the teacher and student networks, thereby enhancing the overall distillation effectiveness. Experimental results demonstrate that the distilled student network, with only 1.32K parameters, achieves remarkable diagnostic performance and robustness against noise. Secondly, we designed an FPGA-based CNN accelerator. This design integrates the Winograd algorithm with a systolic array architecture to substantially accelerate convolution operations while reducing resource consumption. Moreover, parallel computation and module reuse are employed to further enhance inference throughput. Experimental results demonstrate that, compared to CPU and GPU implementations, the proposed accelerator achieves speedups of 142.8× and 1.40×, respectively, while reducing power consumption by 14.5× and 30.2×. Compared to state-of-the-art FPGA accelerators, the proposed design provides an approximate 16% improvement in inference speed. The deployed models achieve F1-scores exceeding 98% on two bearing datasets, validating the proposed method’s capability for real-time fault diagnosis in low-power industrial environments.