Kaixin Wu, Mei Zheng, Peng He, Decheng Wang, Mi Zhao, Yanchun Feng, Guangqing Liang, Jun Xiong, Biqing Zhu, Guo-Wang Lin
Non-invasive prenatal testing (NIPT) generates vast amounts of low-depth sequencing data, offering a valuable resource for studying maternal genetic traits. However, standard NIPT genotype imputation workflows include time-consuming post-alignment processing steps from GATK, whose benefits for low-depth data remain uncertain. Additionally, merging imputation results from large, batch-processed cohorts presents a challenge, particularly for accurately combining imputation information scores (INFO). This study therefore aimed to develop an efficient imputation pipeline for NIPT data by evaluating the necessity of standard post-alignment steps and validating a batch-merging strategy, using maternal folate metabolism genotyping as a clinical application. The omission of GATK post-alignment steps, including duplicate marking and base quality score recalibration, did not compromise imputation accuracy across multiple simulated low depths but substantially reduced computational time. A sample-size weighted averaging method enabled accurate merging of imputation INFO scores from batch-processed data, yielding results nearly identical to single-cohort imputation for high-quality variants. Applying this optimized pipeline to 517 real-world NIPT samples demonstrated high genotype and allele concordance for the MTHFR rs1801131 and MTRR rs1801394 loci when compared to a sequencing capture method, with both metrics exceeding 96% at GP80. In conclusion, this study validates a simplified, computationally efficient imputation workflow for low-depth NIPT data. It enables accurate assessment of maternal folate metabolism genotypes, offering a cost-effective strategy for large-scale genetic screening of specific maternal traits without additional experimental burden, using existing clinical sequencing data.