Danilo Couto de Souza, Luis Manoel Paiva Nunes, Ricardo de Camargo, Felipe M. Pimenta, Marcelo Andrioni, Eric O. Ribeiro
Reliable wind-resource assessment hinges on a robust statistical description of the wind-speed frequency distribution, often obtained via a two-parameter Weibull fit. International guidelines routinely require that such fits be accompanied by formal goodness-of-fit (GoF) tests, with acceptance based on p -value thresholds. However, for modern multi-year, high-frequency records, GoF tests become overly sensitive and p-values collapse toward zero, leading to rejections of otherwise satisfactory fits — even though larger samples should, in principle, yield more reliable parameter estimates. This apparent paradox is known as the large-sample-size effect and can bias practical decisions on model adequacy and complexity. Here, we propose an objective framework to mitigate this effect by combining Monte Carlo (MC) subsampling with temporal segmentation. The MC procedure draws repeated subsamples from the original record, fits the target parametric distribution, and evaluates each fit using the χ 2 and Kolmogorov–Smirnov (KS) tests and the root-mean-square error (RMSE). While GoF-test p-values systematically decrease with increasing sample size, RMSE stabilizes beyond a characteristic threshold. We detect this stabilization point to define an optimal effective sample size, from which statistically supported parameters are recommended for subsequent use. We apply the method to reanalysis winds and in situ meteorological observations (INMET). Across ERA5 and INMET records, full-sample Weibull fits frequently failed χ 2 /KS due to inflated test power, whereas temporal segmentation combined with RMSE-stability subsampling restored statistically interpretable p-values and increased “pass-both” rates from ∼ 25% to ∼ 76% on average across stations, without distorting the inferred diurnal and seasonal structure of ( k , c ) . Temporal segmentation further indicates that fitted parameters are not intrinsic site constants, but emergent summaries of superimposed weather regimes, with diurnal–seasonal variability in ( k , c ) reflecting shifts in the relative dominance of distinct atmospheric circulation patterns. Overall, the proposed framework mitigates large-sample-size artefacts while preserving physically meaningful variability, providing a practical, standards-compatible route for reporting parametric wind-speed statistics from modern high-frequency records. Although demonstrated with the Weibull law, the approach is readily transferable to any desired parametric distribution and helps distinguish genuinely non-conforming regimes from purely sample-size-driven rejections.