科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ bioRxiv : the preprint server for biology2026-09-18· bioinformatics

Post-selection inference in testing for phenotypic differences with scRNA-Seq.

Nicolas Sanchez, Lucas Etourneau, Elizabeth Purdom

原始摘要(英文原文)· Original abstract
For the purpose of differential expression (DE) analysis in single-cell RNA-sequencing (scRNA-Seq), phenotype differences between samples are often tested within specific cell types. Cell types are regularly imputed by clustering the same gene expression data which is later used for phenotype testing. This creates the potential for a "double-dipping" or post-selection inference problem resulting in inflated rates of false discoveries. While this selection bias is known to inflate significance in cell-type marker identification, its effect on sample-level phenotype testing, e.g. in patient cohorts, has never been explored despite the growing preponderance of this type of analysis. To address this, we perform an extensive simulation study and demonstrate that naive clustering on uncorrected embeddings can severely inflate the False Discovery Rate (FDR) in the presence of strong phenotypic differences. However, we further show that applying batch-correction methods to remove phenotypic effects prior to clustering resolves the FDR inflation with no obvious loss of power. Finally, we provide measures of phenotypic imbalance that can be applied to real datasets which closely track the false discovery proportion and thus can be used to as part of data exploration to gauge the risk of post-selection inflation of p-values in a particular dataset.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Post-selection inference in testing for phenotypic differences with scRNA-Seq. — 科研速览 Science Skim