Afshin Jahanshahi, Martijn J. Booij
Robust hydrological projections under climate change depend on model performance under conditions outside those used for calibration. This study presents a large-sample evaluation across 37 Iranian catchments to evaluate the transferability of six hydrological models using Differential Split-Sample Testing (DSST). Four ensemble-averaging approaches were also evaluated to determine whether under contrasting hydroclimatic conditions. DSST experiments were based on non-sequential 2–3-year periods selected according to (i) annual wettest and driest years and (ii) contrasting seasonal precipitation regimes. Model transferability varied substantially with the transfer setting, catchment attributes, and performance criterion. Ensemble averaging generally outperformed most individual models, although performance depended on the averaging method. Bayesian Model Averaging (BMA) and Granger-Ramanathan Averaging (GRA) consistently outperformed the Simple Arithmetic Mean (SAM) and Akaike Information Criterion Averaging (AICA). Based on the Nah-Sutcliffe Efficiency (NSE) metric, GRA exceeded the performance of the best individual model in 50–87% of transfer cases. This study provides a coordinated DSST-based evaluation of conceptual rainfall-runoff models and ensemble-averaging approaches across a large set of Iranian catchments. The results demonstrate how these established methods can be combined into a practical framework for climate-impact assessment in semi-arid, strongly seasonal regions. Based on this large-sample Iranian case study, we recommend DSST-based evaluation using climatically relevant analogues, multiple performance metrics, and advanced ensemble-averaging methods. Among the evaluated approaches, GRA emerged as the most practical compromise between predictive performance and computational efficiency. However, the relative ranking of models and averaging methods may differ under other climatic, hydrological, or data conditions