Mitchell M Holland, Jennifer A McElhoe, Charity A Holland, Jasmeen K Khosa, Claudia Prieto Alcaide, Daniel Hannon, Erin D Brownfield, Jade T Korber, Miriam G Foster, Rachel T Cundey, Floor Claessens, Jennifer Higginbotham, Elise Anderson, Kimberly Sturk-Andreaggi, Charla Marshall, Walther Parson
Sequence analysis of the human mitochondrial genome (mitogenome) is of interest to the molecular anthropology, medical, and forensic communities. Quality mitogenome data is an essential component of haplotype search databases, serving as an important element of forensic investigations to ensure that weight estimates are reflective of accurate coincidental match probabilities. The European DNA Profiling group (EDNAP) Mitochondrial DNA Population database (EMPOP) is considered the gold standard for this purpose, serving as a reference database and a quality-control tool. The current study reports on the development of a sequencing pipeline for mitogenomes that is user friendly, robust, and cost effective for uploading mitogenome sequences to EMPOP. Whole blood or buffy coat samples were extracted using the Zymo Research Quick-DNA Miniprep Plus kit. Amplification of the mitogenome was performed using two overlapping long-range amplicons of approximately 8.5 kb. Batches of amplicons from 372 samples, plus eight DNA extraction reagent blanks and four amplification negative controls, were normalized and pooled using SequalPrep plates. A library of amplicons was prepared by ligation of SMRT bells (single molecule, real time adaptors), and prepared libraries were run on the PacBio Sequel IIe instrument using a high-fidelity (HiFi) approach. The total time for laboratory processing of 384 samples, prior to SMRT bell ligation, was up to 62 working hours. Total cost of reagents and supplies for all steps was approximately 20 U.S. dollars (USD) per sample. Including labor, the cost was approximately 30 USD. The success rate for 10,394 total samples tested was ∼98.2%, with only one of the two target amplicons failing to produce suitable sequence data. Therefore, on a per amplicon basis, the success rate was ∼99.2%. Concordance studies using two short-read sequencing methods confirmed the reliability of the long-read approach. The long-read pipeline can be easily adopted by laboratories and used in high-throughput studies involving quality biological samples to generate large mitogenome databases, including those for upload to EMPOP.