Nicole Kuznetsov, Kensuke Daida, Mary B Makarious, Bashayer Al-Mubarak, Kajsa Atterling Brolin, Laksh Malik, Cedric Kouam, Breeana Baker, Raquel Real, Kathryn Step, Lara M Lange, Lesley Wu, Miriam Ostrozovicova, Katherine M Andersh, Pin-Jui Kung, Yasser Mecheri, Yi-Wen Tay, Behloul Soundous Malek, Nada Al Tassan, Maria Teresa Periñan, Samantha Hong, Mathew J Koretsky, Lana Sargeant, Kristin Levine, Cornelis Blauwendraat, Kimberley J Billingsley, Sara Bandres-Ciga, Hampton L Leonard, Soraya Bardien, Huw R Morris, Andrew B Singleton, Mike A Nalls, Dan Vitale
We present CNV-Finder, a deep learning pipeline employing Long Short-Term Memory (LSTM) networks for large-scale CNV identification within user-defined genomic regions. Trained on expert-annotated samples from the Global Parkinson's Genetics Program across four neurodegenerative disease-associated genes (PRKN, LINGO2, MAPT, SNCA), CNV-Finder integrates human feedback to iteratively improve performance. In benchmarking across 105 936 samples spanning 11 ancestries and nearly 150 cohorts, the model achieved 91% and 89% visual confirmation rates for PRKN deletions and duplications at high-confidence thresholds. In two validation cohorts, CNV-Finder nominated 83% fewer candidates than a popular Hidden Markov Model-based caller while maintaining higher confirmation rates. Validation through MLPA, short-read, and long-read sequencing demonstrated robust performance, generalizing to diverse signatures including homozygous deletions and SNCA triplications absent from training. Our findings highlight human expertise's value in complex loci like 17q21.31.
MOTIVATION: Copy Number Variations (CNVs) play pivotal roles in complex disease etiology, often requiring large sample sizes to analyze disease associations. While genotyping arrays offer a cost-effective approach for CNV detection using Log R Ratio (LRR) and B Allele Frequency (BAF) signals, existing independent array-based callers suffer from high false positive rates and noise susceptibility, burdening manual validation.
RESULTS: We present CNV-Finder, a deep learning pipeline employing Long Short-Term Memory (LSTM) networks for large-scale CNV identification within user-defined genomic regions. Trained on expert-annotated samples from the Global Parkinson's Genetics Program across four neurodegenerative disease-associated genes (PRKN, LINGO2, MAPT, SNCA), CNV-Finder integrates human feedback to iteratively improve performance. In benchmarking across 105 936 samples spanning 11 ancestries and nearly 150 cohorts, the model achieved 91% and 89% visual confirmation rates for PRKN deletions and duplications at high-confidence thresholds. In two validation cohorts, CNV-Finder nominated 83% fewer candidates than a popular Hidden Markov Model-based caller while maintaining higher confirmation rates. Validation through MLPA, short-read, and long-read sequencing demonstrated robust performance, generalizing to diverse signatures including homozygous deletions and SNCA triplications absent from training. Our findings highlight human expertise's value in complex loci like 17q21.31.
AVAILABILITY AND IMPLEMENTATION: CNV-Finder is freely available at https://github.com/nvk23/CNV-Finder.