Xianzhi Wang, Yang Zhang, Yu Wei, Chenyang Ma
Multi-scale entropy is a powerful analytical tool that extends single-scale entropy analysis into a multi-scale feature extraction approach, enabling more reliable fault diagnosis in complex machinery by revealing hidden dynamic information that single-scale metrics neglect. Driven by advances in symbolic dynamics, statistical processing, and robustness enhancement, multi-scale methods have evolved into more than 80 variants. Despite this diversity, a comprehensive and systematic evaluation of these methods is still lacking, particularly with regard to the trade-off between diagnostic accuracy and computational efficiency. This study provides a comprehensive benchmark of 85 multi-scale methods evaluated on the XJTU-Gearbox dataset to establish an evidence-based roadmap that balances diagnostic accuracy and computational cost. Through performance comparison and algorithmic-principle analysis, this study elucidates the mechanisms underlying the performance differences among these variants, highlighting how advanced multi-scale methods preserve richer fault information and enhance feature separability. Furthermore, this work identifies promising development trajectories and offers a basis for future integration of multi-scale methods with deep learning models. To avoid overgeneralization, the results are interpreted as a controlled single-dataset benchmark rather than a universal ranking across all datasets, classifiers, operating conditions, or fault severities.