Timothy Hermanto, Kenny W.L. Lam, Bhagya S. Yatipanthalawa, Gregory J.O. Martin, Yih Yean Lee, Sally L. Gras
Monoclonal antibodies are important for treating a wide range of diseases, making efficient industrial production essential to public health. Mathematical and machine learning models that can predict antibody production during cell culture can potentially reduce process cost and development time. While hybrid models combining both mechanistic and data-driven approaches can successfully predict several cell culture metrics, the effect of different training methods and algorithm types on the predictive performance of hybrid models is not well understood. This study aimed to compare serial hybrid models using four machine learning algorithms, trained with two prominent training strategies, the train-on-rates and train-on-concentrations approach, for both next-day and recursive predictions. Six hybrid machine learning models and four control data-driven models were trained on a narrow set of process data and then tested using four different process datasets to mimic a common scenario in the development of new antibody processes. The hybrid models could predict performance for the next day or across twelve working days. The two hybrid model training strategies and four machine learning algorithms were found to have similar predictive performance, although the hybrid models showed greater predictive capacity than data-driven control models for next-day predictions. A strong correlation was found between the similarity of train and test datasets, as assessed by the Euclidean distance and the error of recursive predictions for both hybrid and data-driven models. Overall, the ToR strategy was found to offer greater flexibility than ToC and the greater susceptibility to error was acceptable for the low level of noise in the data examined. There is also flexibility in choice of algorithm, with Random Forest and XGBoost algorithms performing better in some instances with lower estimates of uncertainty. This study enables a better understanding of the two training approaches and provides insights into future directions for hybrid modelling for process optimisation.