Junjie Fu, Chunjun Qin, Rui Zhou, Zhaohong Deng, Jing Hu, Peter H. Seeberger, Biao Yu, Jian Yin
Glycosylation reactions are the cornerstone of carbohydrate synthesis, yet stereoselectivity has remained a long-standing challenge and a classic black box problem due to complex mechanistic ambiguities and a multitude of influencing factors. To move beyond laborious trial-and-error, early efforts established predictive frameworks based on empirical parameters and computational chemistry, laying a crucial foundation. This review examines the recent emergence of artificial intelligence (AI) as a data-driven paradigm to overcome this unpredictability. We highlight pioneering studies that employ diverse machine learning strategies, including applying transfer learning to address data scarcity, using random forest models with quantum chemical or empirical descriptors on curated datasets, and leveraging interpretable decision trees to predict outcomes from large-scale literature data. Beyond mere prediction, we showcase how AI is transitioning into a tool for discovery, identifying novel stereocontrol principles and guiding methodology development through active learning approaches like Bayesian optimization. Future progress, however, depends on overcoming key barriers in data standardization, advanced algorithms, and closed-loop automation. Addressing these has the potential to transform glycosylation from an intricate art into a predictive science, making the complexities of the glyco-world increasingly tractable.