科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Communications for Statistical Applications and Methods2026-07-31· False discovery rate

Resampling-based mirror statistics to stabilize variable selection for high-dimensional data with binary outcomes

Sijeong Kim, Hokeun Sun

原始摘要(英文原文)· Original abstract
For variable selection problems of high-dimensional data with a binary outcome, regularization methods based on logistic likelihood are conventionally applied to identify relevant variables.Recently, data splitting methods using mirror statistics have been proposed to control the false discovery rate of selected variables.However, we found that the number of selected variables is unstable when we repeatedly apply the splitting methods to the same data.Moreover, the computational time is drastically increased as the number of variables increases.In this article, we propose new computational strategy using resampling-based mirror statistics for not only stabilizing the number of selected variables but also performing feasible computation for an extremely large number of variables.In our extensive simulation study, we demonstrated that the proposed strategy can significantly reduce the variance of the number of selected variables without loss of true positive rates, while it controls the false discovery rate at a designated level.We also applied it to high-dimensional gene expression data from a study of acute lymphoblastic leukemia where some cancer-related genes were consistently selected.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Resampling-based mirror statistics to stabilize variable selection for high-dimensional data with binary outcomes — 科研速览 Science Skim