Xiao-qiang Zhao, Xiangde Qi, Jixiang Zhao
Abstract To address the challenges of underutilization of frequency domain attributes, limited noise resilience, and low interpretability in prevailing deep learning fault diagnosis techniques like CNN and Transformer, this study proposes a novel frequency-space dynamic collaborative residual network, which designs with a four-layer cascaded residual structure and incorporates a customizable Fourier feature decomposition module (FFC) in each layer. Specifically, the designed FFC module decomposes input signals into low-frequency and high-frequency components: Low-frequency branches leverage multi-scale convolution to enhance local feature scale invariance, while high-frequency branches incorporate dimensionality reduction in the spatial domain and depth-separable convolution in the frequency domain (FFT → Frequency-Domain Enhancement → IFFT) for parallel processing. In the frequency-domain processing of the high-frequency branch, frequency domain transformation via Fourier decomposition first identifies fault characteristic frequencies and their harmonic constituents. Subsequently, depth-separable convolution is employed to amplify the frequency band response of periodic impacts in the frequency domain. Finally, the time domain features are reconstructed via inverse Fourier transform to preserve periodic features that may be lost in conventional spatial models. Bidirectional cross-frequency feature enhancement is achieved through a bidirectional cross-frequency interaction mechanism, involving low-frequency feature injection and high-frequency attention compensation. A generalized second-order pool-directed channel attention module is developed to boost the response of fault-sensitive bands by jointly considering spatial covariance matrix and channel energy distribution. On two datasets, comparative experiments with other methods validate the proposed model’s favorable noise robustness and generalization. Meanwhile, interpretability analyses based on salience and rule extraction further prove that the proposed model is interpretable.