Marc B. Gros-La-Faige, Emmanuelle Génin, Anthony F. Herzig
Genetic polymorphisms are common in pharmacogenes, with sometimes important implications for drug metabolism. Assessing the correct enzyme phenotype from genetic data is thus a crucial step into the development of personalized medicine. Many bioinformatics star-allele callers have been developed for this purpose of identifying the correct star alleles and the associated phenotype, each of them having their specific method and limitations. Despite the important benchmarks that have been made so far, their performances have not yet been fully explored depending on various parameters, such as the type of genetic data provided as input or the individuals' ancestry. Hence, we provide a multi-gene, multi data-type comparison of the accuracy of four commonly used and open-access star-allele callers: PyPGx, ursaPGx, PharmCAT, and Aldy. We found that PyPGx and Aldy are overall more performant than the others, except for CYP2D6 where ursaPGx was the most accurate with its CYP2D6 dedicated caller that relies on the Cyrius software. Comparing to the commercial solution DRAGEN, PyPGx, and Aldy showed better results, except for CYP2D6 where DRAGEN performed best. When only SNP-chip or low-pass sequencing data is available, the use of imputation greatly improves the performance of star-allele callers, allowing performance comparable to that achieved with sequencing data. We also analyzed how concordance between star-allele callers varies depending on population ancestry. Our findings offer guidance on the choice of star-allele caller, depending on the pharmacogene being studied and the resolution of available genetic data.