T. Totu, G. Jaques, B. Heiniger, T. Segessemann, M. Schmid, M. Bourqui, A. Wicki, J. E. Frey, C. H. Ahrens
Microorganisms hold great promise for urgent global needs such as increasing sustainable agricultural production while reducing chemical fertilizer and pesticide use or providing novel classes of antimicrobials/therapeutics. Moving from analyzing microbiome composition to applying synthetic communities and studying their functions requires access to isolates and complete genome sequences. By spanning the frequent repeats, long-read sequencing can resolve complex prokaryotic genomes, yet error-prone short-read assemblies dominate. We here release the GenomeCompendium, a public database and interactive analysis tool for complete prokaryotic genomes (https://genome-compendium.com/). Using NCBI RefSeq (~47,000) and GenBank (~13,000) genomes, we integrated available metadata, GTDB taxonomy and computed features including repeat classes and gene content screening, intragenomic 16S rRNA sequence identity, and biosynthetic gene cluster co-occurrences. Evaluating repeat content and assembly complexity metrics, we identify taxonomic ranks dominated by difficult-to-assemble genomes and show that complex, repeat-rich genomes are more common than previously estimated. By mining metadata, our quality control flags 6.3% of RefSeq assemblies as potentially erroneous or incomplete. As valuable reference for data mining and to track taxonomic coverage, the GenomeCompendium links ~90 features across genomes, offers downloadable reports and -as unique features- pre-computed proteogenomics databases to improve genome annotations of RefSeq strains and the ability to analyze any uploaded prokaryotic genome.