Kisung Lee, Galymzhan Moldagulov, Bartosz A. Grzybowski
Machine learning (ML) has become the dominant approach for predicting biological activity from molecular structure, yet its true ability to generalize beyond known chemical space remains uncertain. Although numerous models have been proposed, recent work suggests that simple similarity-based methods can perform on par with far more sophisticated architectures, raising questions about whether current approaches genuinely learn transferable structure-activity relationships. Here, we systematically examine how different dataset-splitting strategies affect the performance of k-nearest neighbors (k-NN) and a representative set of modern ML models. Across all splitting regimes, k-NN performs comparably to state-of-the-art methods. More importantly, as dataset splits become increasingly stringent and test molecules move further out-of-distribution (OOD), the predictive accuracy of all models deteriorates sharply, approaching random guessing. This collapse persists even when standard fingerprints are augmented with 3D geometric descriptors, indicating that many previously reported accuracies-typically obtained under nonrigorous splitting strategies-are likely inflated. These results underscore the need for more rigorous evaluation standards and the development of representations capable of supporting true OOD generalization.