科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Biology2026-09-17

paraORA: GPU-Accelerated Exact Over-Representation Analysis for Repeated Queries of Large Gene-Set Libraries.

Zheng Wu, Zejun Zhang, Wenqing Feng, Jinlei Sun, Zhichun Liu, Guoqiang Wang, Yunqing Liu

原始摘要(英文原文)· Original abstract
Over-representation analysis (ORA) is widely used to interpret selected-gene lists, including differentially expressed genes and cell-type marker genes. Unlike ranked-list gene-set enrichment analysis (GSEA), ORA tests whether selected genes are over-represented in predefined gene sets under an explicit background. Repeated analysis of large gene-set libraries, however, can be computationally costly. GPU-accelerated methods such as rapidGSEA mainly target ranked-list GSEA and do not address repeated exact-ORA processing. We developed paraORA to accelerate repeated ORA by preparing a gene-set library once and reusing its compressed representation across queries. In the GPU-exact path, overlap counting, exact hypergeometric testing, Benjamini-Hochberg correction, and odds-ratio calculation are performed on the GPU. Across 425 numerical validation cases on three GPU models, paraORA agreed with CPU-reference results within the prespecified tolerance while preserving statistical decisions and rankings. On an RTX 4090, the CPU/GPU runtime ratio for the complete in-memory workflow increased from 1.89 for one query to 60.10 for 100 queries against 141,002 gene sets, including one-time preparation but excluding file writing. In a separate complete file-to-file benchmark, GPU-exact was 2.55-fold faster than CPU execution. An exploratory limma reanalysis showed direction-specific Hallmark enrichment. paraORA is available as a web service and an open-source package.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

paraORA: GPU-Accelerated Exact Over-Representation Analysis for Repeated Queries of Large Gene-Set Libraries. — 科研速览 Science Skim