A. Adamov, C. L. Mueller, N. Bokulich
Microbiome sequencing datasets are sparse, high-dimensional, compositional, and hierarchically structured, and predictive modeling from them typically relies on ad hoc feature representation choices that obscure their impact on performance and interpretation. We present ritme, an open-source Python package that jointly optimizes microbiome-specific feature representation and model selection - combined algorithm selection and hyperparameter optimization - tailored to these data. ritme systematically searches taxonomic aggregation, sparsity-aware selection, compositional transforms, and metadata enrichment together with model class and hyperparameters, using state-of-the-art optimizers that scale from a laptop to a compute cluster. Across three real-world use cases, ritme outperformed the original study pipelines by 7-29% on the primary task metric and surpassed three AutoML baselines in six of seven comparisons, while selecting substantially fewer features and exposing how feature and model choices drive performance. Open-source and modular, ritme supports reproducible, parsimonious predictive modeling, downstream biological investigation, and extension to other multi-omics modalities.