Lanting Wang, Tengshuang Ma, Ling Deng, Likun Xu, Tian Yuan, Xuan Gao, Jing Fang
Predicting volatile fatty acids (VFAs) in Fe3O4-mediated high-load anaerobic fermentation (AF) is hindered by stochastic uncertainty and intrinsic data scarcity. To bridge the gap between the lab-scale datasets and industrial prediction demands, this study established the RSDS-ML framework integrating random standard deviation sampling (RSDS) with four machine learning (ML) models-K-Nearest Neighbors (KNN), Support Vector Regression (SVR), Random Forest (RF), and Extreme Gradient Boosting (XGBoost). The raw experimental data (30 points) was randomly partitioned into a training set (80%) and a testing set (20%). Subsequently, the training set was enriched by merging the 80% raw samples expanded via RSDS (expanded limited data to 500-15,000) with the supplementary data (18 points) via Piecewise Cubic Hermite Interpolating Polynomial, which was defined as 'name of ML-Data samples'. Results indicate that the primary KNN-500, SVR-10000 and XGBoost-1000 models are the optimal frameworks, achieving a testing set R2 more than 0.80. Crucially, a secondary fitting strategy was adopted by selecting easily monitored indicators (soluble carbohydrate, soluble proteins, pH, NH4+-N, and organic loading rates), which maintained high predictive accuracy (R2 = 0.8725) via SVR-10000-2 model, where the suffix '-2' explicitly designates the secondary 'lightweight' SVR model trained on the same data volume (10,000 samples) but using only the 5 selected easily monitored features. Summarily, RSDS-ML effectively captures nonlinear features from virtual data, extending the applicability of data-driven modeling to small-sample AF processes. This robust framework not only maintains prediction accuracy but also achieves significance for reducing both experimental workloads and temporal costs.