Haonan Zhu, Andre R Goncalves, Car Reen Kok, Camilo Valdes, Hiranmayi Ranganathan, Boya Zhang, Jose Manuel Martí, Monica K. Borucki, Nisha Mulakken, James B. Thissen, Crystal Jaing, A. Hero, Nicholas A Be
Abstract Background: Microbiome-based disease prediction across pooled studies is challenging because the data are high dimensional, heterogeneous, and often too limited to support reliable study-specific models. We propose a hierarchical sparse Bayesian multitask logistic regression model that encourages shared sparsity across related studies while retaining interpretability and uncertainty quantification. To make posterior inference scalable, we derive a variational approximation for the model parameters. Results: We evaluate the method on synthetic datasets and on a pooled metagenomic collection comprising 61 previously published microbiome studies spanning 19 disease conditions. In simulation, the proposed approach improves support recovery when regression coefficients share a common sparse structure across tasks. On the pooled microbiome application, the proposed method achieves competitive predictive performance while offering clear advantages in interpretability, consistently identifying sparse sets of informative taxa and providing calibrated probabilistic predictions with credible intervals, even under substantial cross-study heterogeneity. Conclusion: These results suggest that the method is particularly useful as an interpretable, uncertainty-aware framework for extracting shared microbial signals from heterogeneous microbiome studies.