Kadeejathul Kubra, Munawar A Shaik
The upstream biopharmaceutical processes of monoclonal antibody (mAb) production generate high-dimensional and irregularly sampled time-series data with many missing values, which can limit the robustness and reusability of data-driven models. In this study, a data-driven workflow has been developed for a previously published industrial fed-batch dataset for the mAb production process. The workflow provides comprehensive data analysis from data cleaning and imputation through model benchmarking and interpretability to enable soft sensing and predictive modelling of mAb bioreactors. Initially, the missing value imputation and their evaluation were carried out using eight supervised machine learning (ML) regressors with temporally engineered time-dependent features. Light gradient boosting machine (LightGBM) provided the best imputed data among the eight ML models and was used as the imputation model for all the variables. Compared with a previously published dual-hybrid imputation method, the LightGBM-based imputation pipeline preserved all observed measurements exactly. It reproduces more accurately both marginal distributions and the multivariate geometry of the data. Using this imputed dataset, thirty-four single and stacked ML models were benchmarked for two industrially relevant tasks. These were used for (i) pointwise soft sensing of the mAb titer using concurrent process variables and (ii) early-stage prediction of final mAb titer using data from 1-7 days. Nonlinear tree-based ensembles achieved test R2 up to 0.98 for soft sensing and 0.93 for early prediction as compared to linear and kernel-based models. However, stacked ensembles provided modest but consistent gains in accuracy and robustness over the best single tree-based ensembles. Finally, permutation variable importance, Shapley additive explanations, and partial dependence analysis showed that elapsed culture time, total cell density, culture volume, glutamine, and glutamate are the most important variables of the predicted titer, whereas tightly controlled variables such as glucose and pH contribute little within the studied operating range. This is the first study that combined modern ML-based imputation, extensive model benchmarking, and explainable artificial intelligence on an industrial mAb dataset to develop both soft sensors and early prediction models. Overall, the workflow offers a simple, reusable, and explainable framework for handling missing data and building soft sensors in industrial mAb upstream development with the potential to reduce analytical burden, speed up feedback, and support data-driven decision-making in mAb upstream processes.