Shijian Dong, Tianyu Yu, Jiahao Liu, Z. W. Wang, Xiaoqing Jiang
To address the challenge of small and sparse datasets, this study presents a first attempt to integrate slag mechanism knowledge with SHAP-guided stacking and data generation for accurately predicting end-point quality in the converter steelmaking process. The slag mechanism knowledge is adopted to increase key characteristic variables and interpretability of prediction results. The missing rate and abnormal values in the smelting data are hierarchically identified and processed by combining the interquartile range (IQR) method and smelting experience. The conditional tabular generative adversarial network (CTGAN) is utilized to expand the training dataset for alleviating the overfitting problem while avoiding data leaking and distribution shifts. The cumulative contribution rate of SHapley Additive exPlanations (SHAP) is utilized to screen out KNN, TabTransformer, RF, TabNet and ET as the base models (Layer-0). The Ridge is used as a meta-learner (Layer-1) to alleviate the multicollinearity through L2 regularization. By comparing the ablation experiments and existing networks, the superiority of the proposed model is verified by the production data from a true steel plant. The experimental results illustrate that the prediction R² of the proposed model in terms of phosphorus (P) content and end-point temperature are 0.932 and 0.948, respectively, and other evaluation indicators are also significantly better than comparing models. The proposed modeling technology lays a foundation for optimization of converter steelmaking process.