Hui-Qi Qu, Hakon Hakonarson
Polygenic risk scores (PRS) quantify the component of disease risk captured by measured additive common variants. Their use as a study design variable, rather than only as a predictive endpoint, broadens their scientific value in human genetics. PRS can define informative extremes of common-variant burden, identify discordance between phenotype and PRS-estimated risk, and facilitate the detection of subgroup-specific or residual mechanisms that may be obscured in conventional case-control analyses. This review outlines four major PRS-informed design strategies: tail sampling based on PRS extremes; discordance sampling based on mismatch between phenotype and PRS-implied risk; conditional and stratified genome-wide association analyses using PRS to adjust for / partition background risk; and residual phenotype analysis of the component of phenotype remaining after the PRS-associated component has been removed. It thus provides a practical framework for sample enrichment, subgroup definition, and calibrated epidemiologic comparison. These designs may be especially informative when integrated with sequencing, multi-omics, longitudinal cohorts, and translationally oriented intervention studies. Their application requires careful attention to data leakage, ancestry-related bias, collider structures, and the limits of interpretation.