科研速览 · Science Skim继续刷下去 · Keep skimming →
◇ bioRxiv2026-08-25· bioinformatics

From Data Curation to Risk Reporting: A Pipeline for Polygenic Risk Scores

P. V. Barbosa Araujo, T. d. S. Fiuza, J. E. Kroll, R. L. Andrade, D. H. F. Gomes, L. Varuzza, G. A. de Souza, S. J. de Souza

原始摘要(英文原文)· Original abstract
Polygenic risk scores (PRS) have emerged as a powerful tool for quantifying genetic susceptibility to complex traits and diseases. However, their calculation and interpretation require standardized data curation, robust statistical methods, and clear reporting strategies. In this work, we present an integrated pipeline designed to address these challenges. The pipeline begins with the construction of a curated genotype/phenotype database derived from public repositories, ensuring that only phenotypes with appropriate metadata, statistical distributions, and ethical suitability are retained. The final dataset comprises 2,346 phenotypes covering 38,256,468 unique SNPs. These phenotypes serve as the final analytical units for PRS calculation, risk stratification, and individual-level interpretation. The generated reports integrate sample-level results, phenotype categorization, risk classification, study references, and variant tables, providing a structured and interpretable output for end users. Together, the curated database and reporting framework establish a comprehensive toolbox for PRS analysis, enhancing reproducibility, transparency, and usability in both research and clinical contexts.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

From Data Curation to Risk Reporting: A Pipeline for Polygenic Risk Scores — 科研速览 Science Skim