Cong-Cong Liu, Shan-Shan Dong, Jing Guo, Zhen Xu, Chen Wang, Yun-Xiao Li, Li-Li Meng, Xi-Cheng Yang, Meng Li, Kun Fu, Yan Guo, Tie-Lin Yang
Recovering high-quality microbial genomes from metagenomic sequencing data is essential for accurate profiling and understanding microbial variation. However, existing clustering methods often suffer from limited accuracy and scalability. Here we present MetaCAT (Metagenome Clustering and Association Tool), a framework that combines recovery of microbial genomes from metagenomic data and analysis of their associations with host traits. MetaCAT incorporates a Sparse Weighted Dirichlet Process Gaussian Mixture Model (SWDPGMM) to accurately and efficiently decompose complex datasets and combines k-mer frequency with read coverage to improve genome reconstruction. It also provides a dedicated workflow for microbial single-nucleotide polymorphism identification and metagenome-wide association studies with the host. MetaCAT outperforms existing methods in both clustering accuracy and computational efficiency across diverse datasets. Using metagenomic data from colorectal cancer cohorts, it revealed previously unrecognized marker species and microbial single-nucleotide polymorphisms associated with colorectal cancer. MetaCAT provides a scalable framework for microbial community profiling and advances our understanding of host-microbe interactions.