Tong Zhang, Zekun Yin, Xiaoming Xu, Lifeng Yan, Yang Yang, Yijie Gao, Xiaohui Duan, Bertil Schmidt, Weiguo Liu
Large-scale genome clustering supports genome database organization, pathogen surveillance, and downstream comparative analysis, yet rapidly growing genome databases demand methods that are efficient, scalable, and flexible across diverse analysis scenarios. This paper presents RabbitTClust2, a fast, scalable, and versatile genome clustering tool based on compact sketch representations. RabbitTClust2 supports both KSSD and MinHash sketches and introduces sketch-based inverted indexes with threshold-aware pruning to reduce unnecessary distance computations. It provides two complementary clustering strategies: a Greedy mode for efficient one-off clustering on massive datasets, and an MST mode that enables reusable clustering structures for fast incremental updates and multi-threshold analysis. To further improve scalability, RabbitTClust2 includes an MPI-based distributed MST implementation for large genome collections. It also supports practical output formats, including CD-HIT-style cluster files, Newick/PHYLIP files, and representative-sketch search tables for genome assignment. Experiments on large bacterial genome datasets show that RabbitTClust2 substantially reduces runtime and memory usage while maintaining high clustering quality. It clusters an 11 TB (FASTA format) GenBank bacterial dataset within one hour on a single server and completes MST-based clustering in less than 30 minutes on 16 nodes. RabbitTClust2 also accelerates incremental and multi-threshold clustering and helps reveal potential inconsistencies between genome similarity and taxonomic labels.