Sijeong Kim, Hokeun Sun
For variable selection problems of high-dimensional data with a binary outcome, regularization methods based on logistic likelihood are conventionally applied to identify relevant variables.Recently, data splitting methods using mirror statistics have been proposed to control the false discovery rate of selected variables.However, we found that the number of selected variables is unstable when we repeatedly apply the splitting methods to the same data.Moreover, the computational time is drastically increased as the number of variables increases.In this article, we propose new computational strategy using resampling-based mirror statistics for not only stabilizing the number of selected variables but also performing feasible computation for an extremely large number of variables.In our extensive simulation study, we demonstrated that the proposed strategy can significantly reduce the variance of the number of selected variables without loss of true positive rates, while it controls the false discovery rate at a designated level.We also applied it to high-dimensional gene expression data from a study of acute lymphoblastic leukemia where some cancer-related genes were consistently selected.