Monica Isgut, Andrew Hornback, Han Bao, Yiting Sun, Katherine Choi, Blake Anderson, Shriprasad R. Deshpande, Anthony Chang, May D. Wang
Polygenic risk scores (PRSs) are increasingly being used to predict disease risk from genetic data. While promising in research, their clinical utility—especially when combined with non-genetic (NG) data such as lab results, physical measurements, and diagnostic history—remains uncertain. Myocardial infarction (MI), a leading cause of morbidity and mortality, is a key use case for assessing the incremental value of PRSs in risk models. Using UK Biobank data, we evaluated the added value of PRSs for 10-year MI risk prediction. We trained models with NG data alone and in combination with PRSs, varying model complexity and the NG feature space. Two modeling frameworks were used: logistic regression and a neural network. NG data was defined using two feature sets: NG1, which included established MI risk factors from structured fields; and NG2, a high-dimensional dataset derived from millions of diagnostic codes across five linked UK Biobank electronic health records (EHR) datasets combined with NG1 features. NG2 was generated using a deep representation learning approach that produced low-dimensional embeddings capturing latent medical concepts and disease co-occurrence patterns. Each model was trained with and without PRSs and evaluated using metrics such as the area under the ROC curve (AUC). PRSs add minimal predictive value when used alone. In contrast, diagnostic data from EHRs significantly improve performance. The best results are achieved using a multimodal neural network combining NG1, NG2, and PRSs. PRSs provide limited standalone utility for MI prediction compared to detailed diagnostic data. Their clinical value likely lies in integration with EHR-based models. Future work should focus on multi-modal approaches that contextualize PRS information within broader clinical data. Polygenic risk scores are a measure of how likely it is a person will get a particular disease based on the genes they inherited from their parents. We studied whether polygenic risk scores (PRSs) help predict heart attack occurrence when combined with clinical data such as information from blood tests and medical history. We tested two different computational models using two types of clinical data: one with common risk factors for heart attacks and another with broader data. We found that PRSs added little benefit by themselves, instead, detailed clinical information was much more useful. The best predictions came from combining all data types. This information and our models could be helpful to better identify people who will have heart attacks. Isgut et al. compare the effectiveness of polygenic risk score (PRS) versus electronic health records (EHR) in predicting 10-year myocardial infarction risk using a machine learning approach. While PRS adds minimal value compared to EHR data, integrating both could optimise risk stratification.