Natalia Onishchenko, Eric S Larsen, Govind Paneru, Kisung Lee, Diana V Kolygina, Yankai Jia, Elizabeth Maria Clarissa, Wai-Shing Wong, Bartosz A Grzybowski
The utility of enzymes in organic synthesis is constrained by the limited ability to predict the scope of small-molecule substrates that a given enzyme can act upon. The problem has proven challenging to both classical computational-chemistry approaches and to modern Machine Learning (ML) methods, the latter suffering from the combination of data scarcity (including all-important negative examples) and the inaccuracy of chemoinformatic vectorization schemes. The current work addresses both problems, deploying affordable chemical robotics to curate a structurally diverse set of both active and inactive substrates, and then using this data to develop a physics-grounded ML model for substrate scope prediction. This model uses only a handful of features to learn the proper balance between steric and electronic factors even from small datasets. It maintains useful predictive power on external enzyme families and shows improved out-of-distribution performance compared with descriptor-heavy state-of-the-art ML algorithms.