the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
Using satellite observations to validate and improve reservoir storage simulations in global hydrological models
Abstract. Global hydrological models (GHMs) increasingly incorporate generic reservoir operation schemes (GROS) to simulate the regulation of rivers by dams. However, the reliability of GROS remains largely unvalidated on a global scale due to the historical scarcity of open in situ data. Here, we leverage the Global Reservoir Storage (GRS) satellite dataset to conduct the first comprehensive quantitative evaluation of reservoir storage simulations globally from five GHMs: H08, WaterGAP2-2e (WGP), MIROC-INTEG-LAND (MIL), CWatM (CWT) and LPJmL5-7-10-fire (LPJ). H08, WGP, MIL and LPJ adopted the process-based Hanasaki et al. (2006) reservoir operation scheme (H06), while CWT adopted the piecewise-function rule curve approach of Burek et al. (2013, 2020) (LIS). We address two primary questions: (1) how accurately do state-of-the-art GHMs reproduce global reservoir storage dynamics? and (2) are model deficiencies attributable to parametric rigidity (i.e., the adoption of globally uniform parameters) in GROS? We evaluated monthly reservoir storage series at 424 major dams (capacity ≥ 0.5 km³) over the historical period, 1999–2018. Performance was quantified using the Kling-Gupta Efficiency (KGE). Two post-hoc bias correction methods—linear scaling and variance-matching—were applied to the raw monthly storage simulations to evaluate whether simple, targeted statistical transformations could recover model skill. To comprehensively address parametric rigidity, we conducted a sensitivity analysis on H08 using its H06 scheme by varying two parameters: target storage level (TSL) and the degree of regulation threshold (DORT) and using LIS by varying the normal storage limit (LN). Our evaluation reveals that current GROS yield generally unsatisfactory performance, characterised by two distinct features. The first concerns seasonal amplitude in storage. MIL initially achieves the highest skill: 52.36 % of dams had a KGE > -0.41. However, KGE decomposition revealed this skill was largely due to dampened intra-annual variability rather than being driven by high correlation and/or low bias error. In contrast, the other GHMs often exhibit excessive seasonal drawdown, systematically overestimating storage amplitude. The second feature pertains to temporal dynamics in storage: within the group exhibiting exaggerated seasonal drawdown, H06-based models—H08, WGP and LPJ—significantly outperform the LIS-based CWT in temporal correlation. We demonstrate that when variance-matching bias correction is applied across all GHMs, two things happen: firstly, the performance of all GHMs becomes generally satisfactory (median KGE > -0.41), and secondly, the GHMs with exaggerated seasonal drawdown outperform MIL in terms of KGE, owing to their superior temporal correlation (H06-based GHMs) and mean bias estimation performance (except H08). By contrast, linear scaling yields only marginal improvements, indicating that correcting variability errors is substantially more effective than adjusting mean bias alone. Furthermore, sensitivity analyses confirm that exaggerated seasonal drawdown is primarily a result of parameter choices rather than inherent flaws in GROS. These findings highlight two critical insights: (1) one-size-fits-all parameters are a primary limitation in global reservoir modelling; and (2) satellite observations are a viable dataset for calibrating reservoir operation schemes in GHMs.
- Preprint
(7346 KB) - Metadata XML
-
Supplement
(10544 KB) - BibTeX
- EndNote
Status: final response (author comments only)
-
RC1: 'Comment on egusphere-2026-2043', Anonymous Referee #1, 21 Jun 2026
-
AC1: 'Reply on RC1', Okiria Emmanuel, 28 Jul 2026
Dear Referee,
We would like to thank you for taking the time to review our manuscript and for providing insightful feedback. We have carefully considered all of your comments.
Please find attached our detailed, point-by-point response as a PDF supplement. In this document, we address each of your concerns and outline the corresponding changes we will make to the revised manuscript.
Thank you again for your valuable contribution to our work.
Sincerely,
Okiria Emmanuel, on behalf of all
-
AC1: 'Reply on RC1', Okiria Emmanuel, 28 Jul 2026
-
RC2: 'Comment on egusphere-2026-2043', Anonymous Referee #2, 24 Jun 2026
This paper describes the evaluation of five different global hydrologic models using remotely sensed reservoir storage data. To evaluate this, the authors use monthly reservoir data for 424 major dams across the globe and calculate the KGE between monthly modelled storage and monthly remotely sensed storage. They also conduct a sensitivity analysis to the parameters within H08 to address rigidity of parameter values. Ultimately, the authors find that the generic reservoir operating schemes perform poorly when evaluating the KGEs. This is due in part to the dampened intra annual variability. This demonstrates that parameters are a key limiting factor in global reservoir modelling and that satellite observations can be used to calibrate and validate reservoir schemes.
Major Comments:
My first main concern after reading the abstract is that the study only evaluated 424 dams. As these are all global hydrologic models, I would assume there is sufficient overlap between the almost 7,000 large dams in GRanD and remotely sensed reservoir storage.Line 119-120: Averaging across the different forcings is totally valid, however, I would be interested as to which forcings the authors chose to average across and why these were chosen over running a new simulation with forcings that better represent the current climatic conditions (i..e W5E5).
Line 179: As Cooley et al 2025 already compared remotely sensed reservoir storage datasets against each other. I would be curious to know what the added benefit is of creating the GDS dataset. This could be better explained in this section by stating some of the limitations (such as temporal or spatial coverage) to the other datasets used.
Line 185 – 195: The selection of four dams in each continent in order to do a detailed timeseries analysis works, however, I would be more interested in a larger grouping of these dams. 4 dams out of the 7,000 to 22,000 that exist globally is quite small. I would suggest repeating this sensitivity with all the dams present in the study and then grouping based on other characteristics such as continent, main purpose, size, etc.
Line 285 – 290: Perhaps this is my reading, but I would assume that based on the equations in line 285 that dams would not be removed unless they did not have observed or simulated storage. I am still a bit unsure how the authors went from 7000 dams in each model to only a little over 400. It would be good to include in Table 1, how many dams each model has as maybe I am assuming there are more per model.
Line 346 to 355. I wonder if this section is a bit redundant as comparisons between remotely sensed storages have already been done. Additionally, the rational for including a new dataset (GDS) was not clearly explained so it can be hard for the reader to know why this addition is important.
Line 365: Now it is stated that the GRS dataset has the same coverage as GRanD. I would highly suggest the authors explain how they went from 7000 dams in the GRS dataset to only evaluating 400 globally and then only 4-7 for the sensitivity analysis. One suggestion could be to implement the same reservoir scheme for all the global and local dams in the selected models in order to have a larger subset to work with.
Line 605 – 609. I agree with the findings that the parameter rigidity could be one of the key limitations in a global reservoir scheme, however I would be interested to know what the authors suggest in order to remedy this issue as it seems highly unlikely that global schemes will end up calibrating their operations to each reservoir due to data limitations. In going off of this, I would suggest combining 4.2 and 4.3.
Minor comments:
Abstract line 36 – 37: A KGE of -0.43 is not necessarily satisfactory but instead should be described as no change as the model is not performing better or worse than before.
Line 118: Figure 1 is referenced a few pages before it shows up. I would suggest keeping figures close to where they show up in the manuscript as it’s easier for the reader to see.
Figure 3 and 4: I really like the colors and the graphics that are shown in Figures 2 , 3 and 4. That said, I would be interested to know if the dashed lines in Figure 4 mean the same thing as the dashed lines in the other figures. If they are also thresholds, it could be interesting to explain how the thresholds are calculated or chosen.
Line 410: This result is a key takeaway and I like the explanation here surrounding the methodology but also the explanation of what it means.
Figure 6,7,8: I think these figures also tell a nice story, but it could be good to try and combine them into one as then we can compare across dams. I also wonder if a heat map is a good way to look at the KGE scores per parameter value. Perhaps it could be nice to look at a line plot or a box plot in order to see the ranges?
Figures 9, 10, 11. I would also suggest combining these into one figure in order to allow for a better comparison. Ultimately, I would suggest combining all figures related to different analyses if the individual dams into one figure per analysis type. This way it is easier for the reader to compare across the different dams and evaluations.
Citation: https://doi.org/10.5194/egusphere-2026-2043-RC2 -
AC2: 'Reply on RC2', Okiria Emmanuel, 28 Jul 2026
Dear Referee,
We sincerely appreciate the time and care you devoted to reviewing our manuscript. We have carefully addressed each of your points. Please find attached our detailed, point by point response in PDF format.
Thank you once again for your generous and insightful contribution to improving our work.
Sincerely,
Okiria Emmanuel, on behalf of all co authors
-
AC2: 'Reply on RC2', Okiria Emmanuel, 28 Jul 2026
Viewed
| HTML | XML | Total | Supplement | BibTeX | EndNote | |
|---|---|---|---|---|---|---|
| 254 | 124 | 28 | 406 | 62 | 21 | 20 |
- HTML: 254
- PDF: 124
- XML: 28
- Total: 406
- Supplement: 62
- BibTeX: 21
- EndNote: 20
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
This manuscript examines two generic reservoir operation schemes implemented in five global hydrological models (GHMs). The authors first evaluate the storage simulation performance of five GHMs against satellite-derived reservoir storage products, using the Global Reservoir Storage (GRS) dataset as the primary benchmark across 424 dams globally. They then test two simple, post-hoc approaches: variance-matching and linear bias correction, to determine whether satellite-derived storage data can recover simulation skill, and conduct a parameter sensitivity analysis on a small number of dams. The study demonstrates that satellite-derived products can be used both for statistical correction and for parameter tuning of generic reservoir operation schemes, improving GHM reservoir simulation performance.
While the topic is relevant to this journal, I have serious concerns about the scientific contribution of the study in its current form. Specifically, three major issues critically affect the quality of the work:
1. Too much unrelated and unnecessary content.
The two central research questions of this study are: (1) How accurate are generic reservoir operation schemes in GHMs globally? and (2) Can reservoir simulation performance be improved by correcting against satellite products and by replacing globally uniform parameters? Substantial content does not serve these questions and could be removed, including the intercomparison of satellite-based storage products (Sect. 2.3.1 and 3.1) and the development of the new GDS dataset (Sect. 2.3.2). Previous research, such as Cooley et al. (2025), already performed a comprehensive global intercomparison of five satellite-derived reservoir storage products (GLWS, GRS, GloLakes, GRDL-Y and GRDL-L) against in situ observations. Moreover, the newly developed GDS dataset performs worse than the existing GRS product (Sect. 3.1), so introducing it here does not appear necessary.
2. The attempt to use satellite-derived products to improve GHM performance is too simple and lacks further exploration.
Testing simple correction methods with satellite data is a reasonable first step, but doing so introduces a satellite-data requirement for each reservoir, which undermines the central advantage of generic reservoir operation schemes. For the generic reservoir operation schemes, they can be applied in any reservoir, regardless of satellite data availability. So here, I would suggest the authors calibrate the key parameters against satellite products across a larger sample of reservoirs, and then examine whether the resulting optimal parameter values relate systematically to reservoir characteristics. Such relationships could then be transferred to reservoirs lacking satellite coverage, preserving the genericness of the scheme while still benefiting from satellite-based calibration.
3. Too few reservoir samples for the parameter tuning tests.
The detailed sensitivity analysis is conducted on only three dams, with four more dams added only as a supplementary check (Sect. S10). Even combined, seven dams out of 424 remain far too few to support general conclusions about suggested parameter ranges. It is easy to imagine that allowing adaptive, reservoir-specific parameters would improve simulation performance even without running these tests. So the tuning experiments on three dams mostly confirm an already-expected result.
These three issues fundamentally limit the scientific depth and impact of the work. Nonetheless, the authors’ efforts are appreciated, and I recommend rejecting this manuscript in its current form, while encouraging the authors to consider these points in future work to strengthen the study’s methodology and the structure of the paper.
Minor Comments
References
Cooley, S. W., Wang, J., Gao, H., Yao, F., Livneh, B., Li, Y., et al. (2025). Global intercomparison of satellite-derived variability in reservoir storage. Environmental Research Letters, 20(8), 084035. https://doi.org/10.1088/1748-9326/ade903
Knoben, W. J. M., Freer, J. E., & Woods, R. A. (2019). Technical note: Inherent benchmark or not? Comparing Nash–Sutcliffe and Kling–Gupta efficiency scores. Hydrology and Earth System Sciences, 23(10), 4323–4331. https://doi.org/10.5194/hess-23-4323-2019