Alison M Cupples, James Richards, Maria Basaldua Del Cid
This study examined freely available whole genome sequencing (WGS) data for genes associated with contaminant biodegradation. Thirteen WGS datasets (>600 individual samples) involving more than 12,000 Gbases from multiple countries were examined. The Department of Energy Systems Biology Knowledgebase (KBase) was used to create metagenome-assembled genomes (MAGs) containing the operons of interest. The arrangement and length of subunits for each operon were compared. Phylogenetic trees were created for common biomarkers (tmoA, pmoA, prmA, mmoX, dmpN). MAGs were uploaded into publicly available KBase narratives. Thirty-four MAGs, within the phyla Actinomycetota, Pseudomonadota and Chloroflexota, were identified with the full propane monooxygenase operon (prmABCD). Twenty-six MAGs, within the classes Gammaproteobacteria and Alphaproteobacteria, contained the full operon for soluble methane monooxygenase (mmoXYBZDC). More than 100 MAGs contained the full operon for ammonia/particulate methane monooxygenase (pmoCAB) and were classified within the Gammaproteobacteria and Alphaproteobacteria groups, as well as other phyla. Fifty-five MAGs, within Burkholderiales (Gammaproteobacteria) and Alphaproteobacteria, contained the full operon for toluene-4-monooxygenase (tmoABCDEF). From the MAGs containing the full operon for toluene monooxygenase, thirty-three also contained the full operon for phenol monooxygenase (dmpKLMNOP). The MAGs generated and their associated functional gene sequences have the potential to improve molecular detection methods for site bioremediation.