Agnieszka Mikołajczyk-Bareła, Michał Grochowski
Current research on bias in machine learning often focuses on fairness while overlooking its underlying causes. Bias was originally defined as a "systematic error," often caused by humans at different stages of the research process. This paper aims to bridge the gap between past and present literature on bias in research by providing a taxonomy of potential sources of bias and errors in data and models, with a special focus on bias in machine learning pipelines. The survey analyzes over forty potential sources of bias in the machine learning (ML) pipeline, providing clear examples for each. By understanding the sources and consequences of bias in machine learning, researchers can develop better methods for its detection and mitigation, leading to fairer, more transparent, and more accurate ML models.