Xinyan Lv, Zi Guo, Yanqi Li, Lei Zheng, Aijun Lin, Xiao Tan
Global soil microplastic (MP) pollution is an emerging environmental concern, yet its large-scale distribution remains difficult to characterize because observations are spatially clustered and environmental relationships are highly nonlinear. We compiled a global soil MP dataset and developed an interpretable ensemble machine learning framework using 500-km spatial blocking and nested spatial cross-validation. Among four tree-based algorithms, CatBoost showed the strongest spatial generalization (mean spatial-CV R2 = 0.587 ± 0.115; pooled out-of-fold R2 = 0.623). SHAP analysis identified soil organic carbon, land cover, PM2.5, and pH as the most influential predictors, while soil, meteorological, and socioeconomic variables contributed 38.2%, 34.0%, and 27.8% of total mean absolute SHAP attribution, respectively. Nonlinear SHAP dependence analysis revealed model-derived breakpoints for several predictors, indicating transitions in model contributions rather than universal environmental thresholds. Predictions for 2022, 2025, and 2030 showed pronounced spatial heterogeneity and region-specific temporal redistribution, with persistent and potentially developing high-abundance areas under the assumed predictor trajectories. Spatial block bootstrap analysis revealed geographic variation in prediction stability. EPI and ADD captured complementary patterns of population-weighted exposure priority and scenario-based potential intake, with the conditional 2030 projection indicating greater exposure priority in parts of Africa and South and Southeast Asia. Overall, the framework supports global soil MP mapping, uncertainty-informed monitoring, exposure-priority assessment, and region-specific management.