Yapeng Kang, Jun Jiang, Lingfeng Meng, Yuhang Zhao, Yunfei Ma, Shengjiang Wu, Yingpeng Dai, Kesu Wei
Moisture content (MC) during tobacco curing is critical to final sensory quality, yet conventional methods cannot satisfy the need for rapid, non-destructive monitoring. This study constructed MC quantitative prediction models based on near-infrared absorbance spectra of tobacco leaves at different roasting stages, combined with various spectral preprocessing methods (SG, CT, SNV, DT, D1st), feature wavelength selection algorithms (SHAP, FI, CARS, SPA), and machine learning models (RF, GBDT, XGBoost, LightGBM, PLSR). To reduce the randomness of a single data split, the optimal XGBoost–SHAP model was evaluated under four splitting strategies: stratified sampling, random split, Kennard–Stone, and SPXY. Under full-spectrum, D1st–XGBoost achieved the best performance with stratified sampling ( R 2 p = 0.9121), and SHAP-selected 86 wavelengths further improved R 2 p to 0.9132. Across splitting strategies, SHAP yielded the highest mean R 2 p (0.8955) with the lowest standard deviation (0.0165), demonstrating superior robustness. SHAP-based interpretability indicated that first-derivative absorbance below 1450 nm contributed mainly positively to MC, whereas wavelengths above 1450 nm contributed primarily negatively. These results show that SHAP-assisted feature selection reduces redundant variables while improving prediction accuracy and stability, and provides a basis for spectral contribution analysis and MC quality control during curing.