科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Agronomy2025-10-28· Randomness

Randomness in Data Partitioning and Its Impact on Digital Soil Mapping Accuracy: A Comparison of Cross-Validation and Split-Sample Approaches

Dorijan Radočaj, Mladen Jurišić, Ivan Plašćak, Lucija Galić

原始摘要(英文原文)· Original abstract
Digital soil mapping has become increasingly important for large-scale soil organic carbon (SOC) assessments, yet the choice of accuracy assessment method significantly influences model performance interpretation. This study investigates the impact of cross-validation fold numbers on accuracy metrics and compares cross-validation with split-sample validation approaches in national-scale SOC mapping. Five machine learning algorithms (Random Forest, Cubist, Support Vector Regression, Bayesian Regularized Neural Networks, and ensemble modeling) were evaluated to predict SOC content across France (539,661 km2) and Czechia (78,873 km2) using 2731 and 445 soil samples, respectively. Environmental covariates included satellite imagery (Sentinel-1, Sentinel-2, and MODIS), climate data (CHELSA), and topographic variables. Four cross-validation approaches (k = 2, 4, 5, 10) were utilized with 100 repetitions each and the results were compared with the existing literature using both cross-validation and split-sample methods. Ensemble models consistently produced the highest prediction accuracy and lowest variance per fold across all validation approaches. Higher fold numbers (k = 10) also produced higher accuracy estimates compared to lower folds (k = 2) and had the greatest value ranges of accuracy assessment metrics. This confirmed the observations from previous studies, in which split-sample validation reported higher R2 values (0.10–0.90) compared to cross-validation studies (0.03–0.68), suggesting a strong effect of randomness in training and test data split in the split-sample approach. This suggests that k-fold cross-validation should preferably be used in reporting prediction accuracy in similar studies, with the split-sample approach being strongly affected by value properties from training and test data from particular splits used for validation.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Randomness in Data Partitioning and Its Impact on Digital Soil Mapping Accuracy: A Comparison of Cross-Validation and Split-Sample Approaches — 科研速览 Science Skim