Hao Xu, Liting Liu, Xianting Zeng, Yanling Bin, Renman Wu, Chengrong Qin, Chen Liang, Shuangquan Yao, Wenxuan Mo, Huanfei Xu, Yingchao Wang, Baojie Liu
Understanding the factors influencing acid pretreatment efficiency is crucial for improving the separation efficiency of lignocellulosic biomass. This study adopted a data-driven method to investigate the separation behavior of cellulose, hemicellulose, and lignin under different acidic solvent conditions. The study constructed a dataset containing 17 variables, which originated from different pretreatment environments. Correlation analysis (PCA, PLS, Spearman) was applied to explore the relationships between variables, and machine learning models—Random Forest (RF), XGBoost, and TabPFN—were developed to predict the component separation efficiency. The results indicated that TabPFN performed best in predicting hemicellulose separation with a test set R² of 0.894 and RMSE of 10.12, while XGBoost showed higher accuracy in predicting lignin with a test set R² of 0.764 and RMSE of 14.21. In all models, key influencing factors included the number of oxygen-containing functional groups, solvent acidity, and reaction temperature. These variables were closely related to the improvement of lignin removal efficiency. This work is a multivariate prediction problem; its correlation prediction is intended to assist the credibility of model interpretation. All data was derived from literature extraction. The research results provide valuable guidance for optimizing pretreatment strategies and improving the selective valorization of lignocellulosic biomass.