Evaluation of global hydrological models for climate change impact assessment: How well can they translate interannual climate variability into streamflow variability?
Abstract. Global hydrological models (GHMs) are widely used in climate change impact assessments to estimate future changes in freshwater availability, droughts, and floods. Their suitability is often evaluated by comparing simulated daily or monthly streamflow with historical observations. However, since climate change impacts are usually expressed as changes relative to a reference period, biases in flow magnitude do not necessarily prevent GHMs from accurately representing hydrological change.
Here, we propose an alternative method for evaluating the suitability of GHMs for quantifying the impact of climate change on water resources. This method is based on the assumption that GHMs that are more effective in capturing the timing and magnitude of the annual streamflow anomaly during a historical period (which is mainly caused by climate variability) are likely to produce more plausible hydrological responses to climate change. Therefore, it compares simulated and observed absolute and relative streamflow anomalies instead of streamflow magnitudes themselves; it considers annually aggregated streamflow, which, compared with daily or monthly streamflow, is much less impacted by reservoir operations and human water use — two important drivers of streamflow that are difficult to model using GHMs.
In this study, we test this new evaluation method as an example for three GHMs (H08, MIROC-INTEG-LAND, and WaterGAP2.2e), forced by two climate datasets (20CRv3–ERA5 and 20CRv3–W5E5). We use streamflow observations from 589 gauging stations worldwide to evaluate the method. Magnitude-based evaluation shows large performance differences among the GHMs, with negative median Nash–Sutcliffe Efficiency (NSE) values for the two uncalibrated ones due to large biases. In contrast, all GHMs achieve positive NSE for annual streamflow anomalies. Regarding relative anomalies, the land surface model MIROC-INTEG-LAND performs worst with an NSE of about 0.3, whereas the two water resources models perform very similarly and achieve median NSE values above 0.5. Both models show very similar correlation, but WaterGAP simulates the standard deviation of the absolute and relative annual streamflow anomaly better than H08. The differences in performance between the two climate forcings are smaller than the differences among the models. In terms of the ability of GHMs to simulate the years in which extreme wet and dry anomalies occur, there is exact agreement between the observed and simulated extreme years at only 7–13 % of analysis locations. Meanwhile, the observed extreme year is identified among the five most extreme simulated years at 40–59 % of analysis locations. MIROC-INTEG-LAND also shows the lowest performance regarding extreme anomalies. GHM evaluation based on annual streamflow anomalies, particularly relative anomalies, is suitable for assessing how GHMs translate annual climate variability and, to a certain extent, climate change into hydrological changes.
However, the proposed GHM evaluation method does not consider the vegetation response to increased atmospheric CO2 concentrations and climatic changes that may strongly affect the hydrological response to climate change. Therefore, multi-model ensemble assessments of hydrological impacts of climate change should include even lower-performing GHMs, provided that these GHMs take the vegetation response into account.
Summary
This paper argues that, because simulations from Global Hydrologic Models tend to be biased and sub-yearly streamflow is hard to model accurately for a variety of reasons, assessing relative annual streamflow anomalies is a valuable addition to model evaluation practices. The paper then applies an example of this to simulations from 3 GHMs, using a set of some 500 gauges that have at least 30 years of data available. Practically, simulations are first de-biased (Q > Q*) and then rescaled by their own simulated mean (Q* > Q**). The authors then discuss the performance of the three GHMs on Q, Q*, and Q**.
Comments
I do not believe this paper achieves its stated goal. To the best of my understanding, the authors simply state that assessment of model performance on relative streamflow anomalies is helpful, then do such an assessment, and then conclude that this assessment of model performance is more robust than assessing model performance on annual streamflow directly. What's missing is the evidence that this is in fact more robust, or even helpful.
Using the single gauge in Figure 2 as an example, the simulations from WaterGAP are much closer to the observed annual values than those of the other two models. However, when relative anomalies Q** are used, WaterGAP emerges as the worst of the three, purely because the anomalies are calculated with respect to each model's own long-term mean flow. Because WaterGAP's mean flow is much closer to observations, its relative anomalies are larger than those of H08, despite being about an order of magnitude smaller in absolute terms. It's entirely unclear to me what I can learn from the Q** assessment that isn't misleading about the model's actual capabilities. If the authors wish to argue that the Q** analysis is more robust/meaningful than Q-based analysis, then they must show evidence that this is the case. Right now, I think Figure 2 mostly suggests that the Q** approach isn't helpful.
Stepping back a bit, there are parallels between the analysis done here and analyses done with global climate models. Because the climate models have biases, changes are often represented as percentage increase/decrease, rather than absolute values. The authors apply this here, but a fundamental difference between climate models and hydrology models is that the latter represents systems with considerable memory that is strongly controlled by the interaction between forcing and landscape. Evidence from the Australian Millennium Drought shows that under prolonged drying catchments can fundamentally change their runoff behaviour, and the idea of stationarity of hydrologic systems has been questioned for several decades. I therefore think that the assumption that relative changes predicted by GHMs are "safe" is at the very least unintuitive. Over the course of a future simulation model errors may become progressively larger as catchments move into consistently drier/wetter states. This stationarity of errors assumption needs to be at least explicitly stated, but evidence that this doesn't play a role (or one small enough to be negligible) would be good to add.
I also think the paper often conflates "improved model performance" and "higher metric values". The Q > Q** transformations may increase the scores found on NSE, r and gamma, but this does not mean model performance improves - it just means the scores are higher. The language used to describe the results needs to be tightened throughout the manuscript to be clear about what these increases in values actually represent, which is mainly that if you ignore specific errors you get higher scores on your performance metrics. This does not mean that the models themselves got any better.
Finally I'm uncertain about the approach of doing the analysis on the annual scale. The main argument seems to be that on annual timescales any interannual differences are mainly driven by differences in climate, and inaccuracies on the shorter timescales (due to incorrect/incomplete process representations) can be ignored. I have two comments here:
1. What I'm missing, particularly in the later analysis that focuses on the driest/wettest anomalies, is some indication of how much of that pattern is already visible in the forcing data. In other words, if the driest and wettest Q years already coincide with the driest and wettest years in the forcing data, what are we really measuring at this timescale? If the model doesn't get in the way too much?
2. To the best of my understanding, the models are still run at the daily timestep. Does aggregating daily simulations that allegedly have substantial errors to annual simulations really let us ignore these errors? The underlying assumption seems to be that the errors to some extent cancel out over a year, but I doubt this holds.
I have a number of further comments in the PDF that I won't repeat here. These mainly concern clarity, logic and some grammatical/editorial things.
Overall, I think this paper needs a substantial amount of work, mainly aimed at clarifying the assumptions that underpin the method, and the addition of evidence that shows that the method actually leads to more robust results. I have particularly strong concerns about the suggestion throughout the paper that ignoring magnitude errors is a way to perform more robust model assessment. Honest assessments of our models' capabilities and the sort of information they can and cannot provide is critical. If getting "good performance" requires the amount of aggregation, de-biasing and rescaling that is shown here, perhaps more effort needs to be spent on making the models themselves better instead.