Youshuai Fang, Weiqin Ma, Wen-Tso Liu, Lin Ye
An integrative framework combining a large language model (LLM) with sequence-based machine learning was developed and applied to 16S rRNA gene-based microbial analyses in anaerobic digestion (AD) systems that constantly receive wasted biomass from upstream activated sludge (AS) processes. Using ChatGPT API, all genus-level microbes in the MiDAS 16S rRNA gene database for wastewater microbiome were annotated according to their environmental preferences (aerobic, anaerobic, or facultative) and categorized into four groups: AS-specific, AD-specific, both AS and AD, and unsure. The results were subsequently paired with full-length 16S rRNA gene sequences for training a random forest classifier based on DNA k-mer frequency patterns. The resulting model demonstrated strong predictive performance across multiple sequence types, achieving 98.76% accuracy on full-length taxonomically classified sequences, 78.87% accuracy on full-length taxonomically unclassified sequences, and over 94% accuracy on partial amplicon sequences. When applied to datasets from previous laboratory-scale and full-scale AD studies, the model successfully identified AS-specific sequences from AD samples, substantially reducing biological noise and revealing clearer beta-diversity clustering. Overall, this study establishes a scalable LLM-enhanced machine learning framework for microbial functional inference directly from 16S rRNA gene sequences, with broad potential applicability to microbial community analyses in a wide range of biological processes.