Md Nasre Alam, Sami Ullah Bhat, Hariprasad Kodamana, Anurag S. Rathore
Effective control of bioprocess systems requires a deep understanding of the fundamental equations governing them. However, in a data-driven framework, the discovery of such guiding rules is hampered by noisy data and intricate nonlinear dynamics. This paper presents a data-driven symbolic regression (SR) approach to enhance the modeling of microbial fermentation systems for bioprocessing. Unlike traditional machine learning (ML) models that fit data to predefined structures, SR discovers both the model structure and the parameters, offering interpretable mathematical expressions. To this extent, in this study, dissolved oxygen, aeration, agitation, temperature, and pH were utilized to develop a symbolic expression with the cell biomass and protein concentration. Results demonstrated that SR models achieved better performance with higher R -square and lower root-mean-square error and mean absolute error than various baseline ML models. A multiobjective optimization technique, namely, nondominated sorting genetic algorithm-II (NSGA-II), was employed over the developed model to optimize cell biomass, protein concentration, and batch time. Further, the optimal results obtained by NSGA-II were experimentally verified through three sequential runs, yielding a mean percentage deviation of 4.50% and 3.26% for cell biomass and protein concentration, respectively. Overall, the results indicate that SR modeling offers a powerful, interpretable, and data-driven method for optimizing complex bioprocesses, outperforming conventional ML models in both understanding and enhancing microbial fermentation systems.