Fayha Salah Ahmed, Mohammad Imteyaz Ahmad, Mohamedi Begum, Deepa Rao, Sheifa Joan
Machine learning-based LDL-C estimation shows improved agreement with routinely reported direct LDL-C values under specific triglyceride conditions. However, performance limitations in hypertriglyceridemia and the absence of external validation indicate that ML approaches should be considered complementary rather than replacement methods.
BACKGROUND: Accurate estimation of Low-Density Lipoprotein Cholesterol (LDL-C) is critical for cardiovascular risk assessment. Traditional equations like Friedewald often underperform, especially in patients with high triglyceride levels. This study compares the performance of Machine Learning (ML) algorithms with conventional equations for LDL-C estimation.
METHODS: A retrospective analysis was conducted using 96,492 lipid profiles from Dubai Health. LDL-C was directly measured for triglycerides > 400 mg/dL and estimated using Friedewald, Martin-Hopkins, and Sampson equations for TG < 400 mg/dL. ML models, Random Forest, XGBoost, Neural Network, k-Nearest Neighbors (k-NN), and Bayesian Ridge, were trained using total cholesterol, HDL-C, and triglycerides as predictors.
RESULTS: Machine learning models demonstrated higher statistical agreement with directly measured LDL-C than conventional equations in patients with triglyceride levels < 400 mg/dL. Model performance declined substantially in hypertriglyceridemia subgroups (> 400 mg/dL), with increased prediction error across all approaches.
CONCLUSION: Machine learning-based LDL-C estimation shows improved agreement with routinely reported direct LDL-C values under specific triglyceride conditions. However, performance limitations in hypertriglyceridemia and the absence of external validation indicate that ML approaches should be considered complementary rather than replacement methods.