Nasrin Rezaei, Zahrasadat Mirtalebi
ABSTRACT The reliability of deep learning models for machinery fault diagnosis is frequently compromised by poor data quality in industrial environments, yet the literature remains dominated by architectural innovations validated on pristine laboratory datasets. This study addresses this gap through a rigorous controlled experiment quantifying the impact of data quality on diagnostic performance, training stability, and model calibration. We evaluate three distinct deep learning architectures, Long Short‐Term Memory (LSTM), ResNet1D, and Transformer, across two benchmark datasets: Case Western Reserve University (CWRU) and Machinery Failure Prevention Technology (MFPT). A systematic corruption pipeline was developed to inject realistic imperfections, including missing data segments and additive noise, followed by a targeted cleaning pipeline utilizing linear interpolation and Winsorization. Experimental results reveal that data corruption degrades performance nonuniformly, with LSTM models suffering catastrophic collapse (56% F1‐score reduction) while Transformers exhibit superior robustness due to attention‐based masking. Crucially, the study uncovers a dataset‐specific pathology in the CWRU dataset where all models failed to detect Outer Race faults; this was resolved solely through the data cleaning process, which restored the periodic fault signatures. Furthermore, in five out of six experimental conditions, the cleaned data yielded higher diagnostic accuracy than the original pristine baseline, suggesting that systematic cleaning acts as a form of implicit regularization. These findings argue for a paradigm shift from model‐centric to data‐centric diagnosis, demonstrating that simple, interpretable data cleaning strategies can yield performance gains comparable to complex architectural modifications.