Yiming Shi, Lili Liu, Jun Chen, Kristine M Wylie, Todd N Wylie, Sung Hee Park, Ruiwen Zhou, Yin Cao, Stephanie A Fritz, Molly J Stout, Maria Cristina Vazquez Guillamet, Lei Liu
Microbiome studies often seek to determine how the absolute abundances of individual taxa change across biological conditions, yet sequencing read counts are sample-specific scaled representations of those abundances. Because sampling depth can differ across samples, fold changes calculated directly from sequencing read counts do not generally represent absolute-abundance fold changes. Normalization methods attempt to account for these between-sample differences in sampling depth, but their accuracy depends on the reference used. In particular, total-sum scaling uses all taxa as the reference and can introduce compositional bias. Reference-based methods instead rely on taxa that are stable across conditions, but contamination of the reference set by differentially abundant (DA) taxa can distort sampling-depth estimation and downstream inference. Here, we present iterative reference selection (IRS), a robust normalization method that iteratively screens and refines a candidate reference set to exclude DA taxa. By deriving a clean reference set, IRS accurately captures between-sample differences in sampling depth and recovers absolute-abundance fold changes. Benchmarking using simulations and datasets with experimental absolute quantification shows that IRS outperforms standard scaling and existing reference-based methods in controlling false discovery rates while maintaining power.