Mirjana Perišić, Gordana Jovanović, Timea Bezdan, Svetlana Stanišić, Andreja Stojić
This study applies an explainable artificial intelligence framework to investigate PM10 variability using routine regulatory air-quality data from a single monitoring station, targeting data-limited conditions. A four-year dataset (2020-2023) of PM10, PM2.5, NO2, SO2, O3, and meteorological predictors was analyzed using ensemble machine-learning models with metaheuristic hyperparameter optimization. The best-performing model achieved high predictive performance (R2 = 0.913), supporting model-based interpretation. Clustering of SHAP-derived predictor-impact profiles identified ten recurrent environmental settings associated with PM10 enhancement, reduction, or transitional behavior. The strongest positive model-derived contributions were linked to cold-season accumulation settings: E0 was frequent and predominantly nocturnal, with a mean impact of 50.9 μg m-3, whereas E7 was less frequent but showed the largest mean impact of 82.9 μg m-3 and persistent daytime-nighttime occurrence. In contrast, warm-season and better-mixed settings showed negative model-derived impacts, with reductions of approximately 26-29 μg m-3 relative to the model-expected baseline. These results show that PM10 variability at the studied site is not fully described by concentration levels alone, because similar concentrations may correspond to different pollutant-meteorology configurations. The proposed setting-oriented ML-XAI framework provides a practical approach for extracting interpretable information from routine monitoring data where detailed chemical speciation is unavailable, while broader transferability requires validation across additional sites and environmental conditions.