Jian Zou, Yannick Düren, Xinyi Wang, Ying Xiang, Yunhui Qi, Miao Wang, Yilin Wu, Samuel Singer, Li-Xuan Qin
Reliable translation of microRNA (miRNA) sequencing data depends on effective harmonization to mitigate artifacts from variable experimental handling. Although many harmonization methods exist, prior evaluations have focused mainly on differential expression analysis, leaving the impact on subgroup discovery understudied. We present a framework for evaluating miRNA sequencing data harmonization in the context of sample clustering that integrates artificial intelligence (AI)-augmented datasets, statistical evaluation pipelines, and accessible software tools, enabling systematic comparisons across diverse signal-to-artifact ratios and cluster-composition settings. Using this framework, we show that harmonization can, often partially, mitigate artifact-associated losses in clustering accuracy, especially at moderate signal-to-artifact ratios, with the extent of mitigation depending on the specific harmonization method, the paired clustering technique, and the cluster-composition setting. We further confirm these findings by analyzing reconstructed cohorts from The Cancer Genome Atlas breast cancer miRNA sequencing data. Collectively, these results underscore the need for tailored harmonization to support reliable subgroup discovery and highlight the broader importance of context-specific workflows in translational genomics.