科研速览 · Science Skim继续刷下去 · Keep skimming →
◇ ChemRxiv2026-07-31· Computer science

Federated Learning for Atom-Wise Chemical Property Prediction: Effects of Dataset Size and Data Heterogeneity in NMR Shift Prediction

Stefan Kühn, Ismail Kara, Turgay Altındağ

原始摘要(英文原文)· Original abstract
Purpose: Federated learning (FL) is an established method to preserve privacy of training data in machine learning. Chemical data, often in a pharmaceutical context, have also been used for FL. So far, this has concentrated on per-molecule data. In this paper, we use per-atom data, specifically nuclear magnetic resonance (NMR) data, and examine the quality of the results obtained. Furthermore, we study how the number of overall data used for training influences the results. Methods: We use a deep-learning based neural network model to predict NMR shifts. We run this centralized and in a federated learning setting using FedN with random and skewed data distribution. For comparison, we also use a smaller neural network and HOSE codes. Results: We show that with low amounts of data, FL achieved better performance than centralized learning for the smallest datasets considered, and for large amounts of data, is almost as good as centralized learning. On the other 1 hand, we find constantly high variance for FL results, even with the full amount of data available. Conclusion: Federated learning is generally suitable for atom-wise data in chemistry. For low amounts of data, it can even significantly improve results. On the other hand, high variances in results show that there can be negative effects of the use of federated learning as well.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Federated Learning for Atom-Wise Chemical Property Prediction: Effects of Dataset Size and Data Heterogeneity in NMR Shift Prediction — 科研速览 Science Skim