科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Journal of the American Chemical Society2026-04-01· Scalability

Yield Prediction of Organic Reactions in Biased Data Sets via Positive-Unlabeled Learning

Florian Boser, Jan C. Spies, Frank Glorius

原始摘要(英文原文)· Original abstract
The vast reaction data within scientific literature represents a rich resource for training predictive machine learning models. However, this resource is fundamentally compromised by a pervasive selection and reporting bias, resulting in imbalanced data sets. In this work, we introduce "Positivity is All You Need" (PAYN), a machine learning framework that addresses this data-scarcity problem by learning directly from biased, positive-only data. PAYN leverages a spy-based positive-unlabeled (PU) learning strategy, treating reported high-yielding reactions as the "positive" class and the vast, unexplored chemical space as the "unlabeled" class. To validate our approach, we simulated literature bias on fully labeled high-throughput experimentation (HTE) data sets, including Ni-catalyzed borylations, Buchwald-Hartwig and Suzuki-Miyaura couplings. We demonstrated that PAYN significantly improves the performance of models trained on biased data by balancing the data with augmented negative data points. This work establishes a robust strategy for leveraging biased data, paving a path toward more scalable and accessible data-driven strategies for accelerating synthesis design, optimization, and chemical discovery.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Yield Prediction of Organic Reactions in Biased Data Sets via Positive-Unlabeled Learning — 科研速览 Science Skim