BalaSubramani Gattu Linga, Faisal E Ibrahim, Nader Al-Dewik
Background: Variants of uncertain significance (VUS) in the LDLR gene remain a major barrier to the molecular diagnosis of familial hypercholesterolemia. Although computational approaches offer scalable prioritization, their clinical utility is limited by predictor discordance, incomplete annotations, and inflated performance arising from variant-type imbalance. We aimed to develop a calibrated, uncertainty-aware machine learning framework for LDLR VUS triage. Methods: We developed an end-to-end pipeline integrating multi-source variant annotations and 45 engineered features (expanded to 68 encoded features) to train RandomForest, XGBoost, and support vector machine models. The framework incorporates predictor concordance analysis, splice-aware annotation, and explicit uncertainty modeling. Models were trained using nested cross-validation and evaluated on an independent ClinGen FH-VCEP expert-panel dataset. Results: The XGBoost model achieved high discriminative performance (ROC-AUC up to 0.99 for all variants and ~0.97 for missense variants). Among 1092 LDLR VUS, 69% were assigned a directional classification, including 485 (44.4%) pathogenic-leaning and 268 (24.5%) benign-leaning variants. The remaining 31% were explicitly categorized as predictor-discordant (n = 170, 15.6%) or unresolved (n = 169, 15.5%), preserving uncertainty. Predictor discordance was particularly enriched among missense variants, indicating that naive aggregation of in silico predictors may lead to overclassification. Conclusions: This framework provides a scalable, transparent, and clinically aligned approach for LDLR VUS prioritization, generating ACMG/AMP-compatible computational evidence to support expert curation and downstream functional validation.