Huda Alrashidi, Fadi N. Sibai, Abdullah A. Abonamah, M. S. Al-Rashidi, Ahmad Alsaber
Air pollution poses a significant threat to public health and the environment, particularly fine particulate matter (PM2.5). Machine learning (ML) models have proven their accuracy in classifying and predicting air pollution levels. This research trains and compares the performance of eight machine learning regression models on a time series air quality dataset containing data from 12 dispersed air quality stations in Kuwait, to predict the PM2.5 Air Quality Index (AQI). After cleaning then trimming the large dataset to about 13.4% of its original size, we performed thorough data visualization and analysis of the dataset to identify important patterns. Next, in a set of five experiments exploring feature pruning, the tree-based models, namely Gradient Boosting and AdaBoost, generated mean square errors below 1.5 and R2 numbers above 0.998, outperforming the other ML models. By integrating meteorological data, pollution source information, and geographical factors specific to Kuwait, these models provide a precise prediction of air quality levels. This research contributes to a deeper understanding and visualization of Kuwait’s air pollution challenges, and draws some public policy recommendations to mitigate environmental and health impacts.