Yi Zhang, Peng Quan, Fangyun Wang, Zhixing Pan, Xinzi Wang, Hengyu Lin, Zheng Hao Leong, Jiayue Zhang, Jianfeng Peng, Yeqing Li, Zongguo Wen
Traditional machine learning (ML) generalizes poorly in anaerobic digestion (AD) prediction, failing to interpret unstructured contextual data. To bridge theoretical mechanisms and industrial scaling, we propose a knowledge-driven, multimodal large language model (LLM) framework with a two-stage strategy: literature pretraining to establish biochemical domain knowledge, then scenario-adaptive fine-tuning on industrial plant data. Inputs integrate numerical parameters (pH, temperature, total solids, and organic loading rate) with textual contexts on system types, feedstocks, pretreatment, and reactor configurations. Ablation studies demonstrate that textual context serves as semantic anchors, enabling the model to internalize process logic rather than fitting numerical noise. The framework achieved an R2 of 0.94 on a heterogeneous test set, outperforming optimized traditional ML models by 13%. Knowledge transfer improved prediction accuracy by 53% and reduced training loss by 50% versus from-scratch training. This study provides a sustainable, data-efficient AI paradigm for autonomous decision-making in bioenergy plants.