Beyond Linear Scaling: Dynamical Controls on European Precipitation Extremes
Abstract. This study investigates the scalability of extreme precipitation indices across the Mediterranean Basin and Europe using 0.11° EURO-CORDEX regional climate model projections. A pattern-scaling framework is applied to assess whether changes in extreme precipitation exhibit a robust relationship with increments in global mean temperature and to determine the time of emergence of scalable patterns during the twenty-first century. Scaling relationships are derived for five extreme precipitation indices using annual mean global temperature change from the driving global climate models.
A clear distinction emerges between ensemble-mean behavior and individual model responses. For most indices, the ensemble mean converges toward the late-century pattern relatively early in the century, suggesting that large-scale thermodynamic control linked to global warming governs the forced signal in precipitation extremes. Individual model realizations, however, display substantially greater spread and often lack a stable scaling relationship until much later in the century.
Projected changes reveal a pronounced north–south contrast across Europe, with increasing precipitation extremes in northern regions and progressively drier conditions across much of the Mediterranean Basin and southern Europe. However, the analysis shows that precipitation extremes do not conform uniformly to a linear scaling with global mean temperature. In contrast to temperature-based metrics, a robust linear relationship for precipitation emerges only gradually and remains weak or absent for the most extreme indices.
Crucially, the breakdown of linear scaling is not merely a limitation of the framework, but a physically informative signal indicating a transition in the controlling mechanisms of precipitation extremes. The forced thermodynamic response becomes clearly distinguishable from internal variability by around mid-century, but for the most extreme events, the deviation from scalable behavior reflects an increasing dominance of regional circulation changes and large-scale dynamical adjustments over thermodynamic constraints alone. In this regime, departures from scaling capture the growing influence of atmospheric dynamics rather than uncertainty.
By the late century, projected changes in several indices remain broadly consistent with large-scale patterns of change, yet the least scalable metrics highlight where dynamical processes override temperature-driven expectations. As a result, some precipitation metrics can be linked to warming more directly than others. This distinction has important implications for pattern-scaled projections, as the breakdown of scaling provides insight into when and where atmospheric dynamics become the primary control on European precipitation extremes, rather than representing noise around a simple thermodynamic response.
Review of Beyond Linear Scaling: Dynamical Controls on European Precipitation Extremes
By Ozturk et al.
This paper presents a very interesting study on how precipitation extremes are changing across Europe and Northern Africa, leveraging the EURO-CORDEX simulations. I think the work on pattern scaling is an especially interesting question. Generally, this is interesting research that I would like to eventually see published, however there do remain some issues to be resolved, which are described in my general and specific comments below.
-Travis Aerenson
General comments:
Statistical methods: Throughout the manuscript the authors seem to assess statistical significance via the signal to noise ratio, saying that a signal is robust when it exceeds its standard deviation. Unless I am missing something, this seems like a very low-bar statistically speaking. In a random sample with gaussian statistics ~32% of data points would exceed this threshold with no underlying trend. Hence, I do not agree with the author’s assessment of exceeding one standard deviation being a robust signal. Additionally, the authors should consider utilizing a field significance test. When applying statistical tests over many datasets independently (as they do in Figure 2 for example where each grid cell is assessed independently) there is a large probability that some of the data points are falsely detected signals. If for example, one applied a 95% confidence test to 1000 grid cells, if the data is gaussian with no underlying trend, one would expect to spuriously detect a signal in 50 grid cells. Statistical methods have been developed to account for this issue. Specifically, I refer the authors to Wilks (2016), which explains this issue and provides a solution that has been adopted in many studies since (including some looking at the same ETCCDI indices).
Additionally, it seems that a lot of the time of emergence analysis depends on the supplementary videos. I suggest that the authors add figures to the main text that show the time of emergence analysis and provide an explanation of the statistical scheme used for time of emergence. If it is simply when the trend exceeds one standard deviation, as stated above, I do not consider that sufficiently rigorous.
Wilks, D. S. (2016). “The stippling shows statistically significant grid points”: How research results are routinely overstated and overinterpreted, and what to do about it. In Bulletin of the American Meteorological Society (Vol. 97, Issue 12, pp. 2263–2273). American Meteorological Society. https://doi.org/10.1175/BAMS-D-15-00267.1
Here are a few papers of mine, the first two providing an example of the Wilks 2016 method and the last one an example of a statistically rigorous time of emergence analysis. I offer these as examples of the level of statistical rigor I consider appropriate and do not require the authors to cite my own papers in their revision.
Aerenson, T., Tebaldi, C., Sanderson, B., & Lamarque, J. F. (2018). Changes in a suite of indicators of extreme temperature and precipitation under 1.5 and 2 degrees warming. Environmental Research Letters. https://doi.org/10.1088/1748-9326/aaafd6
Aerenson, T., McCoy, D., Keys, P., & Caulton, D. (2026). Precipitation Over the Contiguous United States Is Coming From Farther Away Than in the Past. Geophysical Research Letters, 53(14), e2026GL122565. https://doi.org/10.1029/2026GL122565
Aerenson, T., Marchand, R., Chepfer, H., & Medeiros, B. (2022). When Will MISR Detect Rising High Clouds? Journal of Geophysical Research: Atmospheres, 127(2), e2021JD035865. https://doi.org/10.1029/2021JD035865
Clarity and writing style: At times in this manuscript, I found it difficult to discern exactly what the authors were showing. The authors should make sure that there is a clear explanation connecting the methods and results sections. Specifically, I think the scaling approach section needs to be expanded a bit, and it should be clarified specifically which figures are testing if pattern scaling is applicable to these indices. Additionally, the discussion of trend emergence is very reliant on supplemental videos. I believe that there should be sufficient results shown in the main manuscript to justify the discussion and conclusions, so I recommend the authors provide a figure summarizing the results of the videos.
Specific Comments:
L39: The acronym IPCC and AR6 should be defined.
L41-42: The statement that climate model projections indicate increasing precip at high altitudes and decreases across the subtropics should have a relevant citation.
L53-54: I find the wording of this sentence confusing, I recommend rephrasing with a more active voice.
L149: Although it assimilates many observations ERA5 is not itself an observation. I recommend the authors either integrate satellite observations into this analysis (such as IMERG satellite observations for example) or justify the validity of ERA5 ETCCDI indices either by comparing ERA5 extreme precipitation to surface observations or by citing relevant literature.
L157: This may be a little bit pedantic, but the ETCCDI indices have been adopted for longer than the last decade for example, Tebaldi et al (2006) was actually published two decades ago, and Sillmann and Roeckner (cited in the manuscript) was published 18 years ago.
Tebaldi, C., Hayhoe, K., Arblaster, J. M., & Meehl, G. A. (2006). Going to the extremes: An intercomparison of model-simulated historical and future changes in extreme events. Climatic Change, 79(3–4), 185–211. https://doi.org/10.1007/S10584-006-9051-4/METRICS
L180: I find it unclear what is meant by “across all model outputs”, does that mean that the mean is calculated combined from each model?
L211-212: I am having a difficult time connecting the “shifts” discussed in the text to where the shifts are shown in the figures. I think what I am missing is a map of climatological PRCPTOT and how it changes in each season so it is clear what the wet and dry regions are, and how they change during the simulations.
L214: I understand why the authors suggest discarding the precipitation changes over the Sahara, but I am wondering why even show it. I suggest the authors crop their figures to exclude the Sahara, as doing so will avoid confusion.
Figure 5: Is there assessment of the statistical significance of these changes/trends?
L258-259: The statement that no change in CDD with an increase in PRCPTOT indicates that the precipitation increase are not distributed across more days is not necessarily true. CDD is not the same as the number of dry days or number of wet days. There are many ways that the number of wet days could change while CDD stays the same, so this wording should be adjusted.
L260: I think it would be more accurate to replace the word anticipated with “found in model simulations” or something similar.
L299-300: The statement that regional climate model signals are GCM-driven needs to be quantified through statistical analysis. As it is written, it seems to simply be a visual test, which can be very misleading, especially with this style of figures where overlapping lines make it difficult to discern density.
L308: I believe the authors mean dipolar when they say bipolar.
L353: I think the manuscript would really benefit from some discussion on what timescales the linear approximation for precipitation extremes is valid (if any). I am not suggesting a lot of additional analysis here, merely a discussion of when/why this approximation works/does not work.
L354: Linear pattern scaling is by-definition a first-order approximation, the question really is on whether a first order approximation is appropriate for this data.
L363: I find it unclear what is meant by “late century” is that the late 20th or 21st century?
L368: The statement that ERA5 was selected for its high resolution and established suitability for climate analysis is out of place. It belongs with the introduction of ERA5 and should include citations.
L395-399: I find the comparison with the temperature indices very interesting; however this discussion would benefit from some more numbers to back it up. I think a summary table of the time of emergence of each precipitation index from this study and each temperature index from the previous paper would fit really well here.