Xiaoyue Zhu, Ruixin Zhang, Changhong Guo, Yongjun Shu
Soybean (Glycine max (L.) Merr.) is an important crop worldwide, and improving genomic prediction efficiency for complex traits is essential for soybean breeding. In this study, two public soybean datasets for drought tolerance (DT) and plant height (PH) were used to evaluate the predictive value of SNP, simple sequence repeat (SSR), and insertion-type structural variant (INS-SV) markers. Genome-wide markers were identified from whole-genome resequencing data and filtered for downstream analyses. GWAS-assisted marker prioritization was performed using five methods, including GLM, MLM, FarmCPU, fastGWA, and BOLT-LMM, and Top-K marker panels with different densities were evaluated using 12 genomic selection models. GWAS-prioritized marker panels generally outperformed random 5 K SNP panels and achieved comparable or improved prediction accuracy relative to the full-marker SNP GBLUP baseline within the internal validation framework. A marker density of 5 K provided a practical balance between prediction accuracy and marker number in the present datasets. The relative performance of combined-marker panels was trait- and GWAS-method-dependent, with small positive (Δr) values for some DT panels but negligible gains for PH. GBLUP and Bayesian regression models showed relatively stable performance. Overall, GWAS-assisted marker prioritization provides an efficient feature-selection strategy for soybean genomic prediction, whereas multi-type marker panels should be evaluated against a standard SNP reference on a trait- and method-specific basis.