Isolating flow-dependent uncertainty in ensemble reanalysis data and its relation to Euro-Atlantic weather regimes and warm conveyor belts
Abstract. Data assimilation uncertainty varies greatly based on the quality and amount of ingested observations and the uncertainty in the background forecast. Ensemble data assimilation (EDA) schemes quantify this combined uncertainty. We investigate climatologically what governs this uncertainty on daily to synoptic time scales, but isolating the effects of changes to the observation system from the truly flow-dependent part of the uncertainty is not straightforward.
Drawing on the EDA produced for the ECMWF 5th Generation Reanalysis product (ERA5), we isolate the long- from the short-time-scale uncertainty components using grid-point-wise statistical models. We investigate the patterns and causes of the short time-scale component for a set of weather regimes and the occurrence of warm conveyor belts (WCBs) in the European North Atlantic sector. This provides coherent regions of higher and lower assimilation uncertainty in different weather regimes that can be explained by a combination of mathematical and physical arguments. Moisture presence is a key contributor for elevated assimilation uncertainty and our results indicate that this influence is mediated largely through state-dependent observation uncertainties in addition to model uncertainties usually considered to govern sub-daily error growth.
Consequently, we argue that what is commonly conceived as a model or weather system specific uncertainty (the "error-of-the-day") in ERA5 reanalysis is actually highly influenced by derived instrument uncertainties and assimilation challenges. Given ERA5's wide use as well as the role of EDA in providing initial ensemble forecast spread, considering the dynamics of assimilation uncertainty should be standard practice.
Competing interests: At least one of the (co-)authors is a member of the editorial board of Weather and Climate Dynamics.
Publisher's note: Copernicus Publications remains neutral with regard to jurisdictional claims made in the text, published maps, institutional affiliations, or any other geographical representation in this paper. While Copernicus Publications makes every effort to include appropriate place names, the final responsibility lies with the authors. Views expressed in the text are those of the authors and do not necessarily reflect the views of the publisher.
Review of manuscript titled “Isolating flow-dependent uncertainty in ensemble reanalysis data and its relation to Euro-Atlantic weather regimes and warm conveyor belts”, by Henry Schoeller and Stephan Pfahl.
Summary
It is well known that ensemble reanalysis uncertainty decreases with time. Since other factors are kept constant, this decrease must be due to improving observational information (quantity and/or quantity). From this baseline understanding, the current study aims to determine the flow-dependence of this uncertainty, and the processes behind this flow-dependence. The authors find that the relative impact of moisture / moist processes on reanalysis uncertainty has increased with time. They speculate that perturbations to moisture-sensitive observations (within the ECMWF Ensemble of Data Assimilations) might play a role here, in addition to model uncertainties (and inherent, aleatoric uncertainty more generally, I suspect?) which govern sub-daily error growth.
I found the manuscript engaging and thought-provoking, particularly because of the use of generalised least squares (GLS) regression applied to categorical input (relating to Euro-Atlantic weather regimes, WRs). However, I also found the manuscript quite long and lacking motivation at key points, particularly in the Methods section. I suspect that there could be other explanations for some of their results, which would be worth discussing. Overall, my feeling is that the conclusions do not currently justify the length and complexity of the manuscript. I do wonder, however, if it might be possible (from the existing framework) for the authors to draw out more understanding of the mechanisms that govern flow-dependent predictability?
Below, I elaborate on these points. Balancing my responsibility to the authors, readers, and the journal, I recommend that this manuscript could be suitable for publication subject to major changes.
Major general comments
1. I found the Methods section lacked some motivation and was unbalanced in the sense that some aspects are discussed in great detail (e.g. error covariance matrices), while there is little discussion of what the covariates might be other than those describing the annual cycle, and the discussion of the breakpoints left me hoping we would be told more later. This meant that I took a lot of time to get through this section.
For the residual covariances, I wonder if a couple of images displaying the matrix structures could help? I am not an expert here but, for the GLS approach, might this be a series of diagonal lines of reducing thickness the further away they are from the main diagonal(?). For the HAC approach, perhaps the contours would be deformed into a diagonally-oriented “Y” shape to indicate heteroskedasticity(?)
It is only in the Results section (Eq.19) where we find that the initial GLS is simply aiming to find a linear trend in the log of the EDA standard deviation, while fitting to a 3-harmonic annual cycle. A motivating statement in the Methods section might be something like “We use the generalised least squares approach to take account of seasonality while fitting a linear trend to the data, and we later partition this by the weather regimes”.
To make the Methods section accessible to more readers (including me), please discuss briefly how categorial covariates can be incorporated into GLS. The text (later, on lines 674-675) “Figure 7 shows the estimated marginal mean effect (cf. Eq. 13) of each of the WRs, from a model with the WR time series as the sole covariate” seems misleading, in the sense that I suspect the software package used will have replaced this “sole covariate” with dummy binary variables for all (but one) of these categories(?) The description of the marginal means calculation talks about “levels” which, in the readers’ mind, will have a very different connotation to (non-numerical) “regimes”.
2. In relation to Figure 1, the 95% confidence intervals about the change points seem very narrow. Do these confidence intervals make sense in the light on the constraints imposed on the length and number of segments (within the optimisation of the Bayesian Information Criterion, BIC)? Figure D1 suggests that five segments are also near-optimal based on BIC; how does the timing of the break points change with five segments?
In the first segment, the log variance decreases sharply; does this segment exist by virtue of consistent fitting, or is it due to the “minimum segment length” criterion? (The fit does not look brilliant in places).
Perhaps, in the light of the variance trends in each segment, the precise locations of the break points are not so important anyway, but rather it is good to account for some discontinuities in the record?
3. In relation to Figure 7, the authors highlight (and investigate) the negative correspondence between the weather regime anomalies and the spread residuals. Looking closely, the positive residuals actually appear to me to be shifted eastward somewhat relative to the troughs in the weather regimes. See, in particular: AT, ZO, and EuBL. From the literature, it is found that the largest growth in spread tends to be located along the leading edges of troughs (including warm conveyor belts), where moist processes are particularly active. Do the authors think that this might be a plausible explanation of the displayed results? (The negative spread residuals in the anticyclonic anomalies would also follow ipso facto).
4. In relation to section 5.3.1 and Figure 8. The authors suggest that poor correlations in the early record between variance-balanced residuals and moist variables [Lines 753-754] “contradicts the hypothesis of stochastic physics exerting a major influence”. They go on to state [Lines 757-759] “The positive trend with time … again, points to the assimilation uncertainty being dominated by varying observation uncertainties rather than stochastic physics”.
Have the authors considered the possibility that, in the early record, the variance-balanced residuals will be dominated by larger-scale state-dependent uncertainty growth associated with dry dynamics while, later on, better observations will help constrain these larger scales with the effect that the variance residuals will become relatively more sensitive to smaller-scale rapid uncertainty growth associated with moist physics? Such a possibility does not require stochastic physics to play a negligible role and does not require observation uncertainty (including the perturbation of moisture-sensitive observations) to play a dominant role.
This is not to say that the perturbation of observations is the optimal solution. While Isaken et al. (2010) demonstrate (and as alluded to in section 3.1.1), the estimation of background error covariances requires observation perturbations that correctly represent the true observation error covariances, I am convinced that this comes at the expense of increased spread. Machine-learning approaches (for example) to estimating and applying observation error covariances might provide a more optimal solution.
Minor general comments
1. [Line 289] “cf” is used widely to point to confirming information. Isn’t “cf” normally used to indicate opposing information?
2. [Figure 3] Is the unit in the centre panel really (log(σEDA))^2? Is there really a square here, and writing “EDA” when referring to the residual variance seems odd to me. This is actually a more general issue throughout the manuscript. Please could the authors justify or discuss in the text briefly?
3. [Line 617] “so interpretation has to be taken with a grain of salt”. There are quite a few colloquialisms used throughout the manuscript. Personally, I find these quaint and amusing, but maybe they will be confusing for the international target audience? For me, at [Line 559], I found it difficult to determine whether “On the other end of the spectrum” is a colloquialism or not.
Minor specific comments
1. [Lines 108-121] The statements in this paragraph see at odds with each other: environmental or model?
2. [Lines 147-149] “instead of seven continuous numerical time series, we simply use one categorical time series of eight categories”. This was a source of great confusion for me. I interpreted this text as defining the covariate x^WR. After some time, I now suspect that x^WR is deconstructed back to 7 binary variables within the regression(?)
3. [Line 180] “with α increasing by wavenumber.” Is this separate from the spectral smoothing of smaller scales prior to generation of B_EDA (Bonavita et al, 2012), or should it be “decreasing by wavenumber”?
4. [Line 205, Eq.3] Please give a reference to the second form for the analysis covariance.
5. [Lines 218-219] “The contribution of … in a similar way”. This sentence is unclear.
6. [Lines 223-225] I suggest changing the sentence “As uncertain observations … (Isaksen et al., 2010)” to something like: “As Isaken et al. (2010) demonstrate, the estimation of background error covariances requires observation perturbations that correctly represent the true observation error covariances”.
7. [Line 299] The word “slightly” does not covey the idea that differences would ultimately saturate at the full amplitude of atmospheric activity, at or before the intrinsic limit of predictability.
8. [Section 4.1] Some context to what the responses and covariates might be would be very helpful.
9. [Line 400] Change “working” to “parametrized”?
10. [Line 406] Change “levels” to “EMMs”?
11. [Line 412] Change “on the far side of zero” to “outside this interval”?
12. [Section 5.1] A key question: What happens at the break-points? The β change and/or the parameters defining the residual covariance matrix change? The reader would value this kind of insight.
13. [Line 472] What is the rationale for applying a minimum segment length?
14. [Line 474-475] “As mentioned in Sect. 4.3, the confidence intervals for the breakpoints are estimated with a HAC estimator, …” I do not readily see this mentioned in Section 4.3.
15. [Line 495] “Notably, from 1972 to 1978 a new Bcli was also used” … but this does not lead to a diagnosed break-point(?)
16. [Line 512] What is meant by “The pronounced seasonality in the time series becomes visibly clearer with increasing time”? Increasing time backwards or forwards? I see a strong seasonal cycle in the first segment, then no clear trend(?)
17. [Line 518] “seasonality is a property of the system itself”. What is meant here by “the system”? The data assimilation system, the earth system?
18. [Line 528] “without assuming the seasonal cycle as a covariate”. On reading for the first time, it was still unclear to me what the remaining covariates were in this case.
19. [Lines 539-541] “Crucially, these varying amplitudes are tied to the heteroskedastic nature of the time series; it would therefore be inconsistent to hold the seasonal coefficients constant across breakpoints while allowing residual variance and all other covariates to change”. I don’t quite follow here; surely the heteroskedastic nature is accounted for in the specification of residual variances?
20. [Figure 2] The lack of any appreciable k=1 annual cycle in segment 2 is curious. Do the authors have any explanations?
21. [Section 5.2] From the beginning, it would be good to know if this section combines the whole period, or represents a single segment, or differences between the beginning and end (as in Fig D2)?
22. [Line 553] Change “is of due” to “is due”(?)
23. [Figure D2] Why do the authors not plot log(σ1940/ σ2024), which would give the order of magnitude of the difference?
24. [Line 555] Change “Along our” to “Along with our”?
25. [Lines 562-563] Change “Finally, the scale of the spatial heterogeneity is considerably smaller than the level of heterogeneity in the scale of the total response-level variation” to “Finally, the scale of the spatial heterogeneity OF RESIDUALS is considerably smaller than THAT of the total response-level variation”?
26. [Figure 4 caption] Change “95% confidence intervals for the mean estimates” to “95% confidence intervals for the FITTED estimates”?
27. [Line 597] Change “1.5 orders of magnitude higher for the 1940s” to “1.5 orders of magnitude higher for the Gulf Stream location in the 1940s”?
28. [Line 602] “… data from the Gulf Stream location remains rather constant”. The curve for this region seems to decline too in Fig. 4. Possibly the authors are talking about data quantity(?)
29. [Lines 604-606] “Consistent with this, we hypothesize that the Arctic grid point is more affected by this (the onset of nearby radar rain-rate measurements around this time appears not to affect the response; cf. Fig. 2 in Soci et al. (2024))”. I am struggling to understand this sentence; maybe there is a double negative involved?
30. [Line 630] “As expected, the PSD peaks at the low-frequency end for both grid points”. Is this “expected” earlier in the manuscript text or appealing to previous studies?
31. [Lines 631-633] Change “The models also clearly erase local maxima at annual and semi-annual frequencies as well as power at a frequency of four months, though the original data features no prominent peak there” to “The models also clearly erase local maxima at annual and semi-annual frequencies; at four months the original data features no prominent peak there”?
32. [Line 642-643] “retrospectively justifying the implementation of three harmonics in the modelling of the seasonality”. Does the fact that the residuals at the annual and semi-annual periods drop quite a long way below their spectrally-adjacent values indicate that the annual cycle is still not well catered-for with the three harmonics?
33. [Lines 656-657] “Since we assume the segment-specific heteroskedasticity is an artifact of the observation and assimilation suite”. What is meant by “segment-specific”? The residuals will identify heteroskedasticity within and between segments(?)
34. [Line 660, Equation 20] There is a confusing double use of n here. Maybe use n-prime in the summation?
35. [Line 676, Equation 21] Maybe this short-hand equation requires a little more explanation, since x^WR is categorial? As mentioned before, everything pointed me in the wrong direction on first reading, because x^WR was originally defined a sole covariate with 8 “levels”. Please see my major comment 1.
36. [Lines 680-681] “(i.e. we added x^WR to the model in 19)”. Again, please see my major comment 1.
37. [Lines 683-684] “Importantly, all three representations yield very similar spatial patterns, …”. Differences between figures 7, D5 and D6 are almost imperceptible, and I would suggest removing two of these. This is one example of where the paper could be reduced in length.
38. [Line 701 and 704] Maybe replace “isohypses” with “contours”?
39. [Lines 703-713]. These two paragraphs were difficult for me to understand. One issue was the phrase [Line 709] “which variables attenuate the WR signal”; this meant little to me. Later, I came to understand (I think) that the attenuation discussed relates to the shift in variance from the weather regime covariate to the more physically-meaningful covariates, which thus help explain the WR partition.
40. [Lines 720-721] “Firstly, the observation uncertainty ingested in the perturbations varies not only on the long time scales isolated by means of our first-stage models”; as a reminder, maybe insert after this, “applied to monthly-mean fields”?
41. [Lines 722-723] “atmospheric moisture and the presence of hydrometeors in particular exert a state-dependent effect on the observation perturbations. Since zg500 is an integrative field and radiances and humidity measurements are important observation types”. Maybe add “humidity-sensitive” before “radiances”?
42. [Lines 724-725] “Perturbations for humidity observations are not applied to SYNOP and TEMP measurements, but are among the measurements with the longest history (Soci et al., 2024)”. Please check this. I think that SYNOP 2m relative humidities were perturbed(?)
43. [Line 732] “(across grid points)” … and days within the year? Maybe I am missing something here?
44. [Line 732] “implies” doesn’t seem like the right word here?
45. [Line 752] Change “weak constrain” to weak constraint”?
46. [Line 762] How is the layer-average relative humidity calculated?
47. [Lines 777-828] On first reading, this long discussion on the gradient and Laplacian of the geopotential seemed quite complicated and more of a distraction than informative. Maybe this could be summarized more?
48. [Line 830] “We also provide maps of correlation coefficients”. This is a strange way to start a subsection.
49. [Lines 834-836] “Why exactly a higher cloud cover in upper tropospheric layers is associated with lower assimilation uncertainty is outside the scope of this study – our focus is also more on the mid-latitudes than on the polar regions”. Possibly this polar cloud is shallower, longer-lived, and less associated with less-predictable latent heating?
50. [Line 872, Equation 24] Are the estimated marginal means (EMMs) internally calculated in variance space? (I think that would be better than in standard-deviation or log-standard-deviation space).
51. [Lines 882-886] “The covariates we add are local, have a finer structure and are more directly linked to physical processes, which is why they don’t “compete" with the WR time series for signal strength, but absorb it to the extent they share variance with it”. What is meant by “absorb”? I would have anticipated that the additional covariates would compete to some extent with the WR. After all, the WRs exist because they are associated with different dynamical/physical processes(?)
52. [Line 912] Change “deduct” to “deduce”?
53. [Line 929] What is the analogous equation to (25) for the “leave-one-out" strategy?
54. [Lines 932-933] “Results from “add-one" models mentioned above would look virtually identical to fig. 9”. Fig. 9 shows correlations, not ∆σ (?)
55. [Lines 937-938] “A large part of the variability in the uncertainty that is projected onto the WRs clustering is due to the uncertainties’ correlation with the geopotential gradient itself”. I think that the approach being used here aims to estimate the amount of the WR variance that can be accounted for by the various other covariates(?) Could some text like this be used to motivate the methodology?
56. [Lines 941-944] “The large and significant patch of negative values across northeastern North America and along its coast imply a similar conclusion. This is because the WR patterns exhibit minuscule variability in the zg500 gradient in that region whereas the region is generally quite variable in that respect”. I don’t fully understand the discussion around the negative attenuations. Is another approach to mask out regions where the WR patterns have small signal? This would certainly remove northeastern North America.
57. [Lines 951-952] “We consider it also a further argument against the Laplacian’s effect being mediated through its role in QG dynamics, since QG factors certainly have varying effect magnitudes across the domain”. To the extent that the Laplacian of the geopotential resembles minus the geopotential itself, would this not explain why the Laplacian is able to account for a lot of the WR variance?
58. [Lines 963-964] “… much of that uncertainty commonly attributed to model formulation rather than to the initial state …”. It is something inherent in the real world rather than just sensitivity to model formulation(?)
59. [Lines 978-969] “our target variable”. This is a bit vague; can you be more specific?
60. [Line 977, Equation 26]. I appreciate that inflow, ascent and outflow are discussed above, but it could be worth defining what is meant by superscripts “inf”, “asc”, and “out”.
61. [Lines 994-995] “Drawing attention to the differences in the range of values covered by the colour bar of fig. 12 compared to the colour bar in fig. 7”. Maybe some of the signals seen on the eastern flanks of WR troughs relate to the WCB signals seen here?
62. [Line 997] “lower frequencies and spatial locality”. Spatial locality: yes. I don’t follow the lower frequency argument though(?)
63. [Line 1005] Change “background forecast ensemble forecasts” to “background ensemble forecasts”?
64. [Figure 13] Please include the other (overlap?) colours in the key.
65. [Lines 1017-1018] “The procedure yields nine binary masks per WR”. Is this 3x3 = (inf, asc, out) x (0h, 12h, 24h) ?
66. [Line 1025] Change “source” to “inflow” for consistency?
67. [Lines 1031-1032] “sufficient for meaningful estimation of the WR specific influence of WCBs on assimilation uncertainty”. This last section, bringing WCBs and WRs together, adds further complexity and more uncertainty to the analysis. I wonder if it adds enough to the study to warrant its inclusion(?) By the way, Figure 14 does not seem to be mentioned in the text.
68. [Lines 1104-1106] “Rather, the presence of moisture exerts a state-dependent (positive) effect on the assimilation uncertainty through elevated observation uncertainties due to both increased observation perturbations and decreased observational coverage – in particular in regions with prevalent convective activity”. Please could the authors re-visit this conclusion in the light of their reply to my major general comment 4? Maybe this conclusion needs to be a little more nuanced?
69. [Lines 1112-1114] “On a more general level, the results presented herein point to short time scale assimilation uncertainties being driven by observation specific aspects more than physics of short-forecast error growth”. Again, this conclusion might need to be more nuanced following the authors’ reply to my major general comment 4? The authors do go on to state that “While regions and times associated with elevated uncertainties may relate to both – namely in convectively active regions – it has been shown, that this is not always the case” but this confuses me, since the convectively active regions are precisely what we are interested in(?) Could the authors state explicitly the regions and times where it is not the case?
70. [Lines 1115-1116] “We therefore argue considering flow-dependent assimilation uncertainty is vital, yet frequently overlooked in studies concerned with error growth or ensemble forecast uncertainty, especially from a WR perspective”. I totally agree and wonder if this is an avenue that the study could draw more conclusions from.
71. [Line 1152] Change “e_B = T” to “e_B = τ_B” (?)
72. [Line 1155-1156] “subject to the constraint that adjacent blocks in each permutation do not contain the same category”. I think I am not understanding something here; since these are permutations, how could adjacent blocks contain the same category?
73. [Line 1158] Change “A2” to “A1”?
74. [Figure D4] Are plots centred around the annual-mean values?