Yunfei Shen, Yuxiang Wan, Sijun Wu, Xu Yan, Hang Chen, Zhenhua Chen, Haibin Wang, Haibin Qu
Continuous monitoring of therapeutic cell cultures via Raman spectroscopy traditionally relies on product-specific calibration models, which often fail to generalize across diverse commercial processes when process parameter or cell lines change. To overcome the need for extensive recalibration, an integrated generic Raman modeling framework was developed by aggregating in-line spectra from four distinct 1500-L commercial Chinese hamster ovary fed-batch processes. Fifteen combinations of regression algorithms and feature selection strategies were first evaluated under sample-level splitting, and representative pipelines were subsequently assessed through repeated batch-level validation splits. A Support Vector Machine combined with the interval whale optimization algorithm demonstrated competitive performance on validation and test sets. The model achieved high prediction accuracy for glucose, lactate, and viable cell density, with R2 values exceeding 0.94 and low prediction errors observed on the independent test set. Shapley Additive Explanations and permutation importance analyses indicated that the retained intervals were consistent with chemically plausible vibrational regions. Transfer to a fifth unseen commercial process further demonstrated the applicability limits of the global model and the potential of local updating with limited labeled spectra. The integration of machine learning algorithms and feature selection provides a practical strategy for multiproduct Raman PAT development in commercial biopharmaceutical manufacturing.