Yuchen Zhang, Hu Chen, Xiaoqing Liu, Pingan He, Qi Dai
Long-read sequencing enables improved genome inference but remains challenged by high error rates that lead to excessive false-positive (FP) variant calls, particularly for small INDELs in complex genomic regions. We present VCboost, a deep learning-based post-calling framework designed to reduce FP in single-nucleotide polymorphism (SNP) and indel detection from long-read sequencing data. VCboost extracts discriminative features from pileup reads and consensus sequences and employs a dedicated filtering model integrating convolutional and recurrent neural networks with residual connections and multi-head attention. Evaluated on multiple human long-read datasets, VCboost consistently improved variant calling performance over Clair3, achieving a 5%-8% increase in precision and a 2%-5% gain in F1-score for INDELs, with minimal recall loss. For SNPs, VCboost improved precision by 4%-8% and F1-score by 2%-5%, with negligible impact on recall. Substantial performance gains were observed in difficult-to-map regions, including low-mappability and segmental-duplication regions, where SNP precision increased by up to 17.6%. Overall, VCboost effectively enhances the accuracy of long-read variant calling while preserving high sensitivity, offering a robust solution for variant detection in challenging genomic regions.