Y-H Taguchi
Genomic research frequently encounters "large p small n" problems, where the number of variables significantly outweighs the number of samples. To address this challenge, this chapter presents tensor decomposition (TD)-based unsupervised feature extraction (FE) as a robust framework for feature selection in bioinformatics. By employing Tucker decomposition via higher-order singular value decomposition (HOSVD), multidimensional omics data are decomposed into low-dimensional latent spaces. Relevant features are identified using singular value vectors associated with specific biological conditions, with statistical significance assessed through an empirical Gaussian null hypothesis. A novel optimization method is introduced to refine the null hypothesis by maximizing the flatness of the P-value distribution. The utility of TD-based unsupervised FE is demonstrated across diverse scenarios, including multiomics integration, temporal data analysis-where it detects patterns like periodicity without prior knowledge-drug repositioning, and biomarker identification. To support practical implementation, two Bioconductor packages, TDbasedUFE and TDbasedUFEadv, are introduced. This methodology provides a versatile and effective approach for extracting meaningful biological insights from complex, high-dimensional datasets.