Tong Li, Deeksha Beniwal, Ram C. Dalal, Sean Manning, Peng Fu, Anquan Xia, Jinran Wu, Qingyang Liu, Vilim Filipović, Georgios Tsiminis, Yash P. Dang
Mid-infrared (MIR) spectroscopy is increasingly used to estimate soil organic carbon (SOC), yet predictive performance depends strongly on spectral preprocessing. However, there is no consensus on best practice, especially for models that must generalize across regions. We benchmark five most common techniques identified in our meta-analysis presented in Part1—none, Standard Normal Variate (SNV), Multiplicative Scatter Correction (MSC), Savitzky–Golay (SG) smoothing, and SG-based second derivative (SGD)—using the ISRIC global soil IR library (Global, n = 3997) and the three largest national subsets (China, n = 262; Kenya, n = 245; Indonesia, n = 236). Parameters that are held constant include the MIR window (4000–600 cm⁻¹), the chemometric model, and the validation design. We created a stratified 80/20 validation splits (by SOC quartiles and soil type where available) with a fixed random seed. Replicates from the same sample/site were kept together to avoid leakage. We apply the same preprocessing to training and testing dataset. We trained the model with 80% data (cross-validation folds) and used the remainder 20% to test the performance of model in "real world conditions". Partial least squares regression (PLSR, mean-centered) was tuned by repeated 10-fold cross-validation (5 repeats) using the one-SE rule, refit on calibration, and evaluated on the validation 20%. We report R², Root Mean Squared Error (RMSE), Ratio of Performance to Interquartile Distance (RPIQ), and mean bias deviation (MBD). SGD delivered the best global validation performance (R² ≈ 0.79; RMSE ≈ 1.56%), whereas SNV generalized best for China and Kenya, and MSC for Indonesia, thus, demonstrating dataset-specific optimization. These results emphasize the importance of leakage-free external validation, multi-metric reporting, and region-tailored preprocessing.