科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ BMC Public Health2026-02-09· Interpretability

Geospatial and machine learning analyses of cardiovascular disease mortality across the continental United States: Identifying associated variables using Shapley values

Nima Kianfar, Mahdi Taghi, Shayan Dasdar, Abe Mollalo, Behzad Kiani

原始摘要(英文原文)· Original abstract
BACKGROUND: Cardiovascular diseases (CVDs) remain the leading cause of mortality in the U.S. and exhibit pronounced geographic variation. Although prior studies have documented regional disparities, fewer have combined spatial pattern detection with systematic comparisons of machine learning and deep learning approaches to identify county-level variables associated with CVD mortality at a national scale. METHODS: County-level CVD mortality data (2018–2021) were analyzed across the continental U.S. Spatial autocorrelation was examined using Global Moran’s I and Getis–Ord Gi* statistics to identify clustering patterns. Separately, predictive modeling was conducted using five machine learning algorithms: linear regression, decision tree, random forest, support vector machine, and extreme gradient boosting, and a deep learning artificial neural network (ANN), drawing on 40 demographic, clinical, socioeconomic, environmental, healthcare, and behavioral variables. Model interpretability was assessed using Shapley Additive Explanations (SHAP). RESULTS: Significant spatial clustering of CVD mortality was observed, with persistent high-mortality hotspots concentrated in the southeastern U.S., consistent with the “Stroke Belt.” Among the evaluated models, the ANN achieved the highest predictive performance (R² = 0.89), followed by XGBoost (R² = 0.82). SHAP analyses consistently identified hypertension prevalence, population aged 65 years and older, poverty, long-term PM2.5 exposure, and rural–urban status as the most influential contributors to CVD mortality predictions. CONCLUSIONS: These findings highlight the strong geographic clustering of CVD mortality in the U.S. and demonstrate the value of interpretable predictive modeling for clarifying how multiple, co-occurring population-level factors align with observed spatial disparities. Together, spatial analysis and explainable machine learning provide complementary insights into the distribution of CVD mortality, informing place-based public health assessment.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Geospatial and machine learning analyses of cardiovascular disease mortality across the continental United States: Identifying associated variables using Shapley values — 科研速览 Science Skim