Nishat Anjum Bristy, Russell Schwartz
We explore the theoretical and empirical basis for one strategy for managing these large data sets: subsampling mutations for the computationally challenging phylogeny problem followed by faster placement of mutations on a putatively known guide tree. We specifically focus on determining the number of mutations sufficient to recover the true phylogeny at some level of resolution with high probability. We theoretically analyze variants of several common models that underlie popular tools for building clonal lineage trees. We further test these bounds through simulations of these models, extensions of them, and real biological datasets. The results suggest that modest numbers of mutations suffice to reconstruct clonal tree topologies for typical numbers of clones, supporting subsampling as a general strategy for managing the challenges of ever-growing data.
MOTIVATION: Phylogenetics faces a growing challenge from increasingly large and complicated data sets enabled by ever-improving sequencing technologies. The issue is particularly acute for somatic evolution studies, such as cancer cell lineages, where single-cell data sets may include tens of thousands of mutations in hundreds of thousands of genetically distinct cells. Simultaneously, the biological complexity of somatic evolution has led to complex phylogeny methods that struggle to scale to even modest data sizes.
RESULTS: We explore the theoretical and empirical basis for one strategy for managing these large data sets: subsampling mutations for the computationally challenging phylogeny problem followed by faster placement of mutations on a putatively known guide tree. We specifically focus on determining the number of mutations sufficient to recover the true phylogeny at some level of resolution with high probability. We theoretically analyze variants of several common models that underlie popular tools for building clonal lineage trees. We further test these bounds through simulations of these models, extensions of them, and real biological datasets. The results suggest that modest numbers of mutations suffice to reconstruct clonal tree topologies for typical numbers of clones, supporting subsampling as a general strategy for managing the challenges of ever-growing data.
AVAILABILITY AND IMPLEMENTATION: All analysis code and scripts used for data simulation are implemented in Python 3 and available at https://github.com/CMUSchwartzLab/mutation-subsampling.git.