Ali Haghizadeh, Leila Ghasemi, Saeid Derikvand
Predicting flood-prone areas is often challenging due to data fragmentation and privacy constraints. This study developed a novel federated learning framework for flood hazard zoning in the Dorud–Borujerd Plain, incorporating eight influential parameters: slope, aspect, precipitation, geology, infiltration, soil hydrological groups, land use/land cover (LU/LC), and distance from the river. Six federated models were evaluated: Random Forest (RF), Deep Neural Network (DNN), an RF+DNN ensemble, Logistic Regression (LR), Extreme Gradient Boosting (XGBoost), and Light Gradient Boosting Machine (LGBM). The federated network consisted of four clients, representing different organizations (the Regional Water Company, Meteorological Office, Department of Natural Resources, and National Cartographic Centre). Each client held vertically partitioned data, with sample sizes distributed unevenly (ranging from 15% to 40% of the total samples per client). RF achieved the best performance (AUC = 0.92, 95% CI; precision = 0.89; recall = 0.87; F1 = 0.88; Cohen's kappa = 0.84). XGBoost followed (AUC = 0.86), then LGBM (0.78), LR (0.56), and DNN (0.51). A simulated centralized baseline confirmed that the federated RF model preserved high accuracy, with a marginal performance gap (ΔAUC = 0.02) compared to a centralized model trained on pooled data. Feature importance analysis identified distance from the river as the primary predictor. The framework converged within 40 communication rounds using secure aggregation via FedAvg, transferring approximately 12.5 MB per round. These findings demonstrate that, particularly with tree‑based ensemble models, federated learning can yield high‑accuracy flood susceptibility predictions while preserving data privacy and sovereignty across decentralized institutions.