Kati Asikainen, Matti Alatalo, Marko Huttula, Assa Aravindh Sasikala Devi
Machine-learning models are increasingly applied in materials science, yet their predictive power is often constrained by data scarcity. Here we show that accurate predictions can be achieved even with a limited number of training examples, provided that the dataset is compact and grounded in physically relevant quantities. By combining density functional theory calculations with a machine learning framework, we constructed accurate descriptor-based models to predict the formation energies of doped lepidocrocite TiO2 monolayers. The predictive accuracy of machine learning models was first evaluated for single-dopant Pt configurations, demonstrating that the selected structural and electronic descriptors reliably capture the key factors governing dopant stability. Chemical transferability was then examined by extending the dataset to include Ag-doped configurations. The predictive accuracy improved systematically as additional Ag-doped data points were included in the training set, while the performance of Pt remained robust. These results highlight the potential of small and well-curated datasets combined with physically informed descriptors to enable not only accurate but also chemically transferable machine-learning-driven screening in doped TiO2 monolayers.