Karthikasri Karuppusamy, Shobana Sundar
Whole exome sequencing (WES) focuses on the protein-coding regions of the genome and it serves as a cost-effective technique for identifying disease-causing mutations. However, the analysis of WES data remains time-consuming and complicated due to the extensive amount of data generated and the numerous tools available to analyze the data. In this study, we have developed an integrated pipeline for detecting and annotating genetic variants in WES data. The developed pipeline helps in efficiently analyzing the large volumes of genomic information produced by WES. It streamlines the workflow by integrating several open-source bioinformatics tools within the Snakemake workflow management system (WMS), ensuring scalability, reproducibility, and ease of use. The developed Snakemake pipeline covers the entire WES analysis workflow, from initial quality control and pre-processing of raw sequencing data to final variant calling and annotation. It includes implementing robust quality control measures using tools like FastQC and Trimmomatic and developing efficient read mapping with Burrows-Wheeler Aligner-Maximum Exact Matches (BWA). It also focuses on creating accurate variant calling and filtration processes using GATK (Genome Analysis Toolkit). This work also focuses on building a comprehensive variant annotation approach. This process encompasses a fully integrated, end-to-end pipeline for WES analysis. The pipeline will significantly improve accuracy in identifying clinically relevant genetic variants. It provides a standardized and reproducible workflow for clinical research. Furthermore, its open-source nature will allow for community contributions and ongoing refinement of WES analysis methods, ensuring that the pipeline remains at the forefront of genomic research technologies for disease diagnosis.