the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
Are we approaching water budget closure at basin scale for the right reasons?
Abstract. The water cycle represents the continuous circulation of water through the atmosphere, the continents and the oceans. At the basin scale, this cycle can be decomposed into incoming and outgoing fluxes as well as changes of water stored within the basin. The relationship between these fluxes and changes in storage is called the water budget. Closing the water‑budget at basin scale, using observations, model outputs or a combination of both, provides a deeper understanding of water fluxes, storage changes as well as the underlying hydrological processes.
Satellite measurements, especially those from the Gravity Recovery And Climate Experiment (GRACE, 2002–2017) and GRACE Follow-On (GRACE-FO, 2018–present) missions, now allows to estimate the water budget closure level for large basins worldwide, including those poorly or not instrumented at all. One of the outgoing fluxes for exorheic basins, river discharge, is not available for many basins, particularly during the twenty‑first century due to the spatial and temporal heterogeneity of openly accessible in‑situ river discharge. To address this issue, several studies have attempted to close the basin‑scale water‑budget by estimating river discharge as the simple sum of runoff calculated from Land Surface Models. Despite well-documented significant errors in runoff estimation, uncertainties associated with these approaches are often neglected as river discharge is, in most cases, the flux with the smallest amplitude in the water budget. In this study, using an ensemble approach to assess the state‑of‑the‑art estimates of water cycle fluxes and stocks, we precisely evaluate the impact of river discharge errors on the conclusions drawn from the water budget analysis of 82 exorheic basins worldwide. When ensembles of products are used to estimate combination of precipitation, evapotranspiration and discharge that are the closest to the GRACE storage change variations, we demonstrate compensations between the components of the water budget can occur. Although river discharge is the flux with the smallest amplitude, its dispersion (36 %), when derived from aggregated runoff, is comparable to that of precipitation (44 %) and evapotranspiration (34 %) and a seemingly high level of water budget closure can be achieved with, for some basin, errors of more than 40 % in the annual volume assessment of precipitation and evaporation. Errors caused by these compensations are larger for continental, sub‑arctic and polar basins (around 15 %) than for tropical, arid and temperate basins (around 6 %). Moreover, applying a simple routing scheme to the runoff does not significantly reduce these errors. In this study, we mitigate and assess these compensations by using a combination of in situ and satellite-based discharge estimates. Future studies should explore methods that further limit the compensation effect when attempting to approach a water budget closure at basin scale.
- Preprint
(27928 KB) - Metadata XML
-
Supplement
(17164 KB) - BibTeX
- EndNote
Status: open (until 20 Oct 2026)
- RC1: 'Comment on egusphere-2026-4569', Anonymous Referee #1, 15 Sep 2026 reply
-
RC2: 'Comment on egusphere-2026-4569', Anonymous Referee #2, 18 Sep 2026
reply
This study presents an analysis of the water budget closure in 82 river basins worldwide, with a special focus on river discharge. The authors analyse a large number of precipitations, evapotranspiration, terrestrial water storage datasets, in conjunction with several models providing river discharge and runoff. The study confirms some known results, such as the possibility that the water balance is almost closed because of variable compensation and not accuracy of the individual components. The main novelty is to assess the computation of water discharge by aggregating gridded runoff datasets produced by land surface models, in comparison with discharge observations from gauges and altimetry.
In its current state, the paper is long and difficult to read. It contains too much details, some repetitions, while lacking a deep analysis and discussion of the results. This dilutes the scientific message of the paper and makes it difficult to extract the main conclusions that should be used by the community to improve water balance closure. The reviewer recommends a major revision to considerably improve the form of the paper while addressing the scientific questions below. The main questions concern the newly introduced metric (Coefficient of Relative Variation) that is not clearly defined and a lack of understanding on the datasets selection.
Form and organisation
1. The paper is very long and some parts contain too much details. In particular, the results section only start on page 21. Consider sharpening the claims.
2. Introductions to sections 3 and 4 are lengthy while not providing much information to the reader.
3. Interpretation of NSE (i.e., negative RMSE is worse than average) is repeated several times. State only once in the methods section.
4. Methods used to compute altimetry-based discharge (Section 3.2.1) are not explained in enough details to allow the reader to reproduce the method. Give the implementation details in the appendix.
5. Use quantifiers (highly reliable, significatively more reliable, much better, quite well) sparingly only when uncertainties can be properly quantified
6. Section 3.3 and its subsections: reliable can have several meanings. Consider writing “is in good agreement with”
7. Section 3.3.3 is unclear and maybe not necessary. “More reliable” than what? “Lower performance” than what? “Thus, OBD can be considered a reference discharge and is far more reliable than MBD.” Is not a logical conclusion of the above and if OBD is the reference dataset, OBD cannot be more reliable than MBD in reproducing OBD.
8. Section 3.3.4 is probably not necessary either
9. Section 3.4 and l.514-523: Filtering details can be moved to the appendix
10. CVR in Section 3.5 is a repetition of earlier sections
11. Summarise figures in a way that provides more information in less space. For instance, Figure 4 reports NSE and KGE but they yield the same interpretation.
12. The eight figures in the results section are violin plots. Consider providing more diverse representations
13. On figures, clip axes to meaningful values. It is not necessary to show or discuss any NSE < -1. Figure 5 should be clipped
14. The results section include several comparisons with previous works ; this belongs to the discussion. The results section should focus on an analysis of the figures, not merely a statement of what is shown in the figures.
Scientific questions
1. There are often several gauges for a given river. What is the sensitivity of the results with the choice of a different gauge? Related to this question, when selecting a large basin where the water balance closure is satisfying with a combination of datasets, how do the same datasets perform on sub-basins?2. The study focuses on years 2003 to 2019. For datasets selected as the best ones in 2003-2019, do they remain satisfying after 2020.
3. Equations (6, 7, 8) are not clearly defined: what are indices m, k, what is N_k? In the definition of the CVR, how do the authors choose the dataset that serves as a reference “the mean amplitude of one of the n time‑dependent vectors”? The reviewer is confused when averages/ standard deviations are computed over time or over several datasets.
4. L 448-449: How are the thresholds “low”, “average”, “high” discharge defined? If the reviewer understands correctly, the authors add random noise to the observed monthly discharge and compute the NSE between the noisy and the clean discharge. Unless in basins with low seasonal amplitude, NSE will always be very high. The reviewer doubts this is a robust quantification of measurement uncertainty.
5. The water budget closure yields “2880 NSE values for each discharge” model and each basin, correct? (l.511). Does it mean that the winning P and ET datasets can allow a good water budget closure in only one basin and be poor in all the others? Wouldn’t it be possible to include the annual volumes in the selection so that the datasets provide more than budget closure?
6. Related to the earlier question on CVR, its benefits are unclear in Section 3.5. Reporting the standard deviation of each variable and the mean of a reference variable conveys the same information as CVR while avoiding introducing a new metric. l. 671-674: what is the reference A in the CVR values reported?
7. l. 607: “OBD discharge can be considered a much better estimate of real discharge than MBD”. Is this really the conclusion of this analysis or is it more “OBD allows a better closure of the water balance than MBD”?
8. l. 667-668: "The dispersion of each component of the water budget equation reveals whether this component plays an important role in the overall budget.” The reviewer does not agree with this statement. The contribution of each component depends first on the mean to capture the correct average proportion of the water budget. Then, the dispersion influences how well closure can be achieved in different months. Can you clarify what are the novel findings about dispersion that go beyond the existing literature?
9. Section 4.3.4: Does the reviewer understand correctly that you re-evaluate annual volumes from the P and ET datasets that were selected as best closing the water budget for a chosen discharge calculation method? Do these results depend on a good selection of the P and ET datasets or do all datasets provide similar volumes once aggregated to the annual scale?
10. What are your overall recommendations for selecting a combination of datasets that satisfy the water budget closure while not relying only on compensation?
Minor remarks:
1. Check the spelling of all river names, especially in the figures and tables (e.g., Nil, Brahmapoutre have French spelling)
2. l.169: less thab
3. l.208: The sentence “Only one realization in geocenter (TN13) and one C20 (TN14) are selected in this study” is not clear to the reviewer
4. l.257: “GLEAM 4.2a and 4.2a”
5. L.491: “therefore indicates thus”
Citation: https://doi.org/10.5194/egusphere-2026-4569-RC2 -
RC3: 'Comment on egusphere-2026-4569', Bramha Dutt Vishwakarma, 19 Sep 2026
reply
Dear Authors,
Manuscript sets an expectation that improved water budget closure would be achieved in the light of discussion on the impact of compensation errors in various fluxes, however, the conclusions are nearly what the study started with as the major problem in water budget evaluation. The improvement in water budget closure is minimal and the misclosure is attributed to cancelation of errors.
Positives: the study presents a detailed assessment of discharge products and establishes that model-based discharge is still much behind the observed discharge. Machine learning based discharge appears to be doing great but a detailed discussion on them is missing. Proving that three-point filtering is indeed needed for WBE.
Major concerns: Compensation of errors has been highlighted, I wish the discussion would have been a bit more physical, for example, anthropogenic or climatological drivers for poor water budget closure over certain basins.
The major takeaway is already established: “A good water budget closure does not guarantee that all the fluxes are estimates accurately and it may occur due to cancellation of errors”. The use of CRV is an improvement but the discussion around various products and fluxes does not appear to establish something groundbreaking. The section is large and would benefit from reducing text and restructuring to highlight major takeaways.
Some product choices must be explained better. For example: why not take COST-G product available at GravIs (see ref below) along with SAGSA? It appears from Fig 6 that GSFC is doing better than SAGSA at global mean level. Similary GRUN is doing better in simulating discharge, hence, why not use GRUN and discard other discharges for WBE.
The text needs major improvement. The paper should be concise. At present it is too lengthy, which also makes hard for the readers to stay with the main story-line/narrative. Grammatical improvements along with better sentence formulation could help. I have provided such correction for two pages and recommend authors to take this as an example and work on the rest of the manuscript.
L18: water budget closure level --> water budget closure
L19: delete “including poorly and not instrumented at all”
L22: “simple sum of runoff from LSMs” appears to imply that estimate of runoff from several LSMs are added together. Please rephrase. I believe what authors meant was basin average of gridded simulation from one LSM.
L25: rephrase: “In this study, using an ensemble approach to assess the state-of-the-art estimates of water cycle fluxes and stocks, we precisely evaluate” to “In this study, we precisely evaluate”
L29-34: “Although ... evaporation”: please break it into two sentences and rephrase for clarity.
L40: “movement of water within the Earth system” might be better.
L42: At basin scale, the water cycle be indeed decomposed --> At the basin scale, the water cycle can be decomposed
L43: as well as change of water stored in the basin as well as change in water stored (in the beginning of the sentence, it is already said that we are talking about the basin scale)
L43: precipitations --> precipitation
L46: launches --> launch
L47: GRACE FO --> GRACE-FO
L47: observations of monthly variations in TWS
L49-50: detailed understanding of which aspect? Please be a bit more clear.
L52: the paragraph starts with “For this purpose”, but that purpose must be defined prior to this line. This sentence itself becomes too convoluted with texts such as “Evaluation as the evaluation of”.. and so on. Please simplify.
L55: “gravimetry-based” means satellite gravimetry or in-situ gravimeters based?
L79: “such as dams”
L83: secondly --> second,
L101: on the humid basins than on the arid basins
Some minor technical comments:
L23-24: please cite literature showing poor estimates of runoff from models.
L50: human impact is a real challenge but an opportunity as well. Goswami et al., 2024 showed that water budget was used to better estimate ET over highly irrigated Ganges.
L63: Lorenz et al. 2014 is another such studies that one could cite.
L210: the spherical harmonic products require filtering and a subsequent leakage correction. The filtering has been mentioned but no mention of the leakage correction. Kindly add that detail. See Klees et al., 2007, Vishwakarma et al., 2017.
L260: recently, water budget based estimates are gaining popularity. This has been exciting because the ET product from water budget are able to capture human signal better. Any irrigation activity will take water from storage and put it on the surface, contributing to ET and water budget based estimates are able to capture that (Goswami et al 2024)
Maybe it is worth mentioning it here, while maintaining that the products used here are only those obtained from empirical relations and modelling. This also begs interpretation around the Indus ET underestimation explained around line 750.
L520: it is appreciated that authors have shown the value in filtering the observations.
L590: routing has a small impact on improvement in modelled runoff performance. Can this be used to conclude that complex runoff/discharge modelling is probably irrelevant for catchment scale water budget evaluation?
L758: MBD has been changed to MDB and it remains so after this for a few pages. Please be consistent.
Ref:
Goswami, S., Rajendra Ternikar, C., Kandala, R., Pillai, N. S., Kumar Yadav, V., Abhishek, ... & Dutt Vishwakarma, B. (2024). Water budget-based evapotranspiration product captures natural and human-caused variability. Environmental Research Letters, 19(9), 094034.
Xiong, J., Abhishek, Xu, L., Chandanpurkar, H. A., Famiglietti, J. S., Zhang, C., ... & Vishwakarma, B. D. (2023). ET-WB: water balance-based estimations of terrestrial evaporation over global land and major global basins. Earth System Science Data Discussions, 2023, 1-47.
Lorenz, C., Kunstmann, H., Devaraju, B., Tourian, M. J., Sneeuw, N., & Riegger, J. (2014). Large-scale runoff from landmasses: a global assessment of the closure of the hydrological and atmospheric water balances. Journal of Hydrometeorology, 15(6), 2111-2139.
Klees, R., Zapreeva, E. A., Winsemius, H. C., & Savenije, H. H. G. (2007). The bias in GRACE estimates of continental water storage variations. Hydrology and Earth System Sciences, 11(4), 1227-1241.
Vishwakarma, B. D., Horwath, M., Devaraju, B., Groh, A., & Sneeuw, N. (2017). A data‐driven approach for repairing the hydrological catchment signal damage due to filtering of GRACE products. Water Resources Research, 53(11), 9824-9844.
GravIs portal for GRACE data: https://gravis.gfz.de/tws
Citation: https://doi.org/10.5194/egusphere-2026-4569-RC3 -
RC4: 'Comment on egusphere-2026-4569', Anonymous Referee #4, 20 Sep 2026
reply
The paper shows that model-based runoff can have a spread as large as that of precipitation and/or evapotranspiration, and that errors in modelled runoff estimates are often cancelled by errors in other water balance datasets when selecting a combination of products that better correspond to the GRACE datasets. This conclusion is reached via a comprehensive assessment that relies on diverse datasets for each variable: 12 evapotranspiration, 12 precipitation, and 20 total water storage timeseries. What distinguishes the current work from previous studies is that it uses a larger ensemble of precipitation, evapotranspiration, and runoff, while also adding altimetry-based river discharge; however, the conclusions reached remain largely overlapping. Accordingly, it is important to highlight what is truly new here, in terms of methods or new insights gained. Please find some questions and comments below.
Questions: Section 3: What is the selected Aj here, and on what basis is it selected? Ak is computed from the max and min in the record considered; how sensitive is this to, for example, the length of that record and whether there is trend/shifts in the record? Consequently, how sensitive are the conclusions to the metric used? And I was wondering if the authors could have just used the spread across the products for comparisons here.
Comments: It may be useful to shorten the paper to improve its readability. Here, I provide some suggestions by taking the abstract as an example. The abstract contains too many details. Providing less context makes it easier for the reader to find out the methods and contribution of the study. The order of sentences is important: line 34 should come before the results, as it is part of the methods. On the choice of words: Line 29: replace “smallest amplitude” with “smallest water balance term”, and instead of “dispersion,” is there a possibility to use “spread across the products”? On clarity: At what timescale is this error metric computed? Line 20: it may be useful to differentiate “measured river discharge” from the “modelled-derived river discharge,” as the term “river discharge” by itself can mean both things. Line 35 is confusing: is it the selection of the combination of the products that better correspond to GRACE datasets that needs further methods to limit the error compensation effects, or is it about the constrained water balance closure methods?
The text needs English revision; here I give a few examples from the introduction. Line 55: replace “variation” with “change”; Line 109: replace “satellite-based discharges” with “satellite -based discharge estimates”; Line 115: replace “is” with “are”.
Minor comments: Section 2: line 210: replace “impacts of a single basin” with “impacts on a single basin”; line 206: replace “Global Isostatic Adjustment (GIA) models” with “Glacial Isostatic Adjustment (GIA) models”
Citation: https://doi.org/10.5194/egusphere-2026-4569-RC4
Viewed
| HTML | XML | Total | Supplement | BibTeX | EndNote | |
|---|---|---|---|---|---|---|
| 172 | 93 | 15 | 280 | 23 | 12 | 16 |
- HTML: 172
- PDF: 93
- XML: 15
- Total: 280
- Supplement: 23
- BibTeX: 12
- EndNote: 16
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
This manuscript studies how the choice of river discharge data affects water-budget closure. The authors analyse 82 basins using different precipitation, evapotranspiration, runoff, and GRACE water-storage datasets. The study addresses an important question. The main result—that a good water-budget closure does not always mean that all water-budget components are correct—is useful. However, I have two main concerns about the NSE equation and the interpretation of (ΔP) and (ΔET). The manuscript also needs careful English editing.
1. Line 305- line 330: There may be a typing error in Eq. (1). In the standard NSE equation, the denominator is calculated using the observed values and the mean observed value. However, Eq. (1) appears to use the modelled values instead of the observed values. This also seems inconsistent with the explanation that NSE is zero when the mean observed value is used as the prediction. Please clarify whether this is only a typing error. If the standard equation was used in the calculations, please correct Eq. (1). Otherwise, the NSE results should be recalculated.
2. Line 510-535 and line 775- 820:
For each discharge dataset, the authors test 2,880 combinations of DTWS, P, and ET datasets and select the combination with the highest NSE. Therefore, POBD and PRRD may come from different precipitation products. (PPRD-POBD)/POBD shows the difference between two selected precipitation estimates. But the POBD is not true value. So the ΔP and ΔET can not define as estimation error. If the authors want to call them errors, they should provide independent evidence that POBD and ETOBD can be used as reference values.
3. Language error: I will take some examples, please carefully check the English throughout the manuscript.
Lines 29–32: Change “for some basin” to “for some basins.”
Lines 168–171: Change “less thab 40%” to “less than 40%.”
Lines 224–258: Change “Japonese” to “Japanese,” “Operationnal” to “Operational,”. “GLEAM 4.2a and 4.2a” should probably be “GLEAM 4.2a and 4.2b.”
Lines 460–478: Change “significatively” to “significantly.”