Jack C Kennedy, William Ferguson, Owen R. Jones, Steven Riley, Thomas Ward, Maria Tang, Jonathon Mellor
Nationally, there was a 47% improvement in Influenza pcWIS versus sub-ensembles. However, Influenza operational ensembles were on average 22% worse than sub-ensembles, when measured by RPS. For COVID-19, operational ensembles were 43% and 280% worse on average, than retrospective sub-ensembles by pcWIS and RPS, respectively. However, COVID-19 operational ensembles were on average 2% (pcWIS) and 13% (RPS) better than individual operational models. For influenza, operational ensembles were, on average, 58% (pcWIS) and 41% (RPS) better than individual models. The sub-ensemble simulation showed how individual models influenced the ensemble scores during different epidemic phases. The Pareto analysis demonstrated that there can be a trade-off between relative direction and absolute count score optimisation.
Abstract Background Epidemic forecasting research often assesses ensembles and their component models using probabilistic scoring rules. Quantifying how individual models affect ensemble performance is challenging, particularly across multiple targets and spatial scales. Methods We present Winter 2024-25 forecasts of Influenza and COVID-19 hospital admissions in England and conduct a retrospective simulation using the operational component models. Forecasts were scored using the per capita weighted interval score (pcWIS) for counts and the ranked probability score (RPS) for ordinal trend direction. We compared operational retrospective forecasts, used generalised additive models (GAMs) to estimate the expected change in score from the inclusion of a model in a sub-ensemble (an ensemble formed from a subset of available models), and used Pareto analysis to understand which sub-ensembles were Pareto-optimal across scoring rules. Results Nationally, there was a 47% improvement in Influenza pcWIS versus sub-ensembles. However, Influenza operational ensembles were on average 22% worse than sub-ensembles, when measured by RPS. For COVID-19, operational ensembles were 43% and 280% worse on average, than retrospective sub-ensembles by pcWIS and RPS, respectively. However, COVID-19 operational ensembles were on average 2% (pcWIS) and 13% (RPS) better than individual operational models. For influenza, operational ensembles were, on average, 58% (pcWIS) and 41% (RPS) better than individual models. The sub-ensemble simulation showed individual models influenced the ensembles during different epidemic phases. The Pareto analysis demonstrated that there can be a trade-off between relative direction and absolute count score optimisation. Interpretation Our analysis shows that UK Health Security Agency forecasts were well calibrated with observations and often had comparable performance to optimal ensembles. Our GAM and Pareto analyses inform model selection for future ensembles. Author Summary Forecasts of winter hospital pressures in England are an important tool for senior healthcare leaders. It is common practice to produce a forecasting ensemble, i.e. combine the predictions of multiple models to create a single, more accurate prediction. Forecasting teams should strive to produce the best forecast possible; one tool for this is retrospective evaluation over a forecasting season using proper scoring rules to assess performance. Our forecasts are constructed of two components, an epidemic trend direction estimate as well as forecast of hospital admission numbers. There are two main challenges we address. The first is understanding at which epidemic phase different ensemble contributions are most effective, the second is the joint optimisation of an ensemble for both trend direction and admission numbers forecast. We apply these methods to a variety of ensembles (sub-ensembles) based on our own modelling suite, and compare the sub-ensembles to our operational forecasts from the Winter 2024/25 season.