unseen-awg v1.0: spatio-temporal weather generation using analogs and unseen data
Abstract. Weather generators allow anticipating unseen weather and help prepare for possible weather-related hazards by providing long continuous time series representative of a given climate. Accurately representing dependencies between variables and locations within weather generators is challenging yet important – ignoring them can result in biased risk estimates. Daily analog weather generators trivially capture spatial and multivariate dependencies within each single time step. These generators resample a historical dataset while ensuring that successive sampled days have consistent large-scale atmospheric fields, thereby also ensuring temporally consistent local weather to some extent. Nevertheless, analog weather generators so far underestimate temporal correlations and are limited by the length of the available dataset they sample from, usually observations or reanalysis data. We propose unseen-awg, an analog weather generator based on data from weather forecasts initialized with historical conditions (reforecasts) and apply it to Europe in a case study. Combined with a novel tuning strategy and block sampling, this large, high-resolution dataset representative of present-day climate allows unseen-awg to simulate weather for the full annual cycle, improve on the temporal continuity of the generated time series, and generate unseen extremes at a daily timescale. We demonstrate that unseen-awg captures both the distributional properties of the individual variables and the dependence between summer temperature and precipitation at the grid-cell scale. We further highlight its ability to simulate droughts and heatwaves of unprecedented spatial extent. Combined with climate impact models, unseen-awg holds great potential for assessing weather-related risks across sectors such as water, agriculture, and forestry, domains that require simulating multiple variables and spatial dependencies across a large number of locations.
General Comments
The paper presents a novel tool, unseen-awg, which is an analog weather generator that resamples ECMWF reforecast data rather than reanalysis. The combination to the best of my knowledge is novel with a large reforecast pool as the sampling basis, with a forecast-skill criterion for tuning σ, and multi-day block sampling to preserve autocorrelation. The methodology, as well as results and discussions are well put forth. The limitations of the study are also openly acknowledged, and the code and data release allow for ease of replication. I view this as a valuable contribution and recommend publication following minor to moderate revision once the following comments have been addressed
Specific Comments
1. Lines 82-84: The manuscript notes that the dependencies between the ensemble members and between consecutive initializations reduces the effective sample size. Since the generation of unprecedented daily extremes is a central claim, a quantitative estimate of the unique/ independent states that exist (particularly in the tail) may strengthen the argument further.
2. In appendix B2, the authors point out that the largest connected heat event is drawn from the same reforecast state at different threshold levels, while this may indeed be a interesting result, it could also point to limited no of distinct candidate states present in the extreme tail.
3. Line 130-131: Unweighted Euclidean distance may give more importance or weightage to the high latitude cells. This may be addressed/ justified using area weighting.
4. The extreme events highlighted in Fig. B3 and Appendix B2 seem to come from fairly long lead times (26–28 days), where reforecasts have drifted towards the model's own climate. The additve bias correction applied per lead time takes care of the mean drift ,but not the variance? It may be useful to show that the t_lead distribution of the generated extremes aren't preferentially drawn from long leads with inflated spread?
Minor Technical corrections:
1. L231: you mention roughly 50 % increase; however 0.31 to 0.54 is more than 70 % ?
2. L236 : contiguous instead of coniguous?
3. Fig. B1 caption refers to panel (d); but the panels are labelled as (c1), (c2). Please correct
4. In Appendix B2, you report that the state selected at all three thresholds as t_init = 29 July 2012, t_lead = 28, m = 8, whereas the corresponding panel titles in Fig. B3 read as (2022-06-24, 26d, 6) ?