Yuhan Zhou, Haipeng Lu, Sicen Liu, Shuliang Zhang
With the intensification of global climate change and rapid urbanization, urban flooding poses an increasing threat to urban safety and sustainable development. Flood susceptibility mapping (FSM) serves as a practical approach for recognizing areas that may be vulnerable to flooding and is therefore essential for flood mitigation and urban planning. In this study, an interpretable ensemble machine-learning framework for urban FSM was developed using social media data. First, the spatial locations of flood events were extracted from social media posts and news reports to construct a flood inventory. Subsequently, a non-flood sample selection strategy, termed Similarity- and Diversity-Based Representative Sampling (SDRS), was proposed to ensure both sample similarity and diversity. Based on these samples, a heterogeneous bagging-based ensemble machine learning model was established for flood susceptibility assessment. To enhance model interpretability, the GeoShapley method was introduced to quantify the contributions of key conditioning factors and reveal their directional effects. The findings indicated that the proposed SDRS strategy delivered the best performance, yielding an AUC of 0.893 and a test-set precision of 0.859. The resulting susceptibility map exhibited a clear south-to-north decreasing gradient, with High- and Very-high-susceptibility zones accounting for approximately 26% of the study area (1897.23 km2). The interpretability analysis further indicated that the Nighttime Light Index (NLI), Impervious Surface Percentage (ISP), and population density were among the most strongly associated positive factors in the model, with a Global Spatial Share of 7.18%. These findings demonstrate that the proposed framework can reliably recognize areas vulnerable to flooding and offer a scientific basis for urban flood management in Guangzhou.