Eneko López, Giulia Gorla, Jaione Etxebarria‐Elezgarai, José Manuel Amigo, Andreas Seifert
Overfitting remains one of the most pervasive and deceptive pitfalls in predictive modeling. It leads to models that perform exceptionally well on training data but cannot be transferred nor generalized to real-world scenarios. Although overfitting is usually attributed to excessive model complexity, it is often the result of inadequate validation strategies, faulty data preprocessing and biased model selection, problems that can inflate apparent accuracy and compromise predictive reliability. In this second part of our series, we examine the most common yet overlooked practices that contribute to overfitting, ranging from data leakage in preprocessing to the pressures of scientific publishing that encourage result-driven overoptimization. By identifying these pitfalls and providing practical guidelines for performing robust validation protocols, this work serves as a blueprint for researchers to ensure their models are not only high-performing but also trustworthy, reproducible, and generalizable. • Comprehensive tutorial on external validation • Focus on overfitting as one of the most pervasive and deceptive pitfalls • An eight-step checklist for reliable chemometric modeling • Overfitting as a result of a chain of avoidable missteps