科研速览 · Science Skim继续刷下去 · Keep skimming →
◇ bioRxiv2026-08-21· bioinformatics

Interpretable biomarker discovery from small-sample microarray datasets using XGBoost rank aggregation and SVM-RFECV

M. M. Prapty, M. S. Rahman

原始摘要(英文原文)· Original abstract
MotivationHigh-dimensional microarray datasets remain valuable for cancer biomarker discovery, but their small sample sizes make robust and interpretable feature selection challenging. Efficient workflows are needed to derive compact gene signatures while preserving biological interpretability. ResultsWe developed a two-stage biomarker-discovery workflow that combines cross-validated XGBoost rank aggregation with support vector machine recursive feature elimination and cross-validation (SVM-RFECV) to identify compact candidate biomarker panels. The workflow was evaluated on 21 public binary and multiclass microarray datasets using repeated stratified cross-validation for internal validation. Across the dataset collection, the selected panels demonstrated strong internal discriminative performance while remaining sufficiently compact for downstream biological interpretation. SHAP analysis identified dataset- and class-specific discriminative genes, and functional enrichment analysis supported the biological coherence of representative consensus signatures. The proposed workflow provides an interpretable and reproducible framework for candidate biomarker discovery from small-sample microarray datasets. AvailabilitySource code and processed outputs are freely available at https://github.com/mashiyat-mahjabin-prapty/microarray-feature-selection.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Interpretable biomarker discovery from small-sample microarray datasets using XGBoost rank aggregation and SVM-RFECV — 科研速览 Science Skim