the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
Earth observation constrained calibration improves soil moisture drought representation: a multi-model analysis in the Rhine River basin
Abstract. Accurate characterisation of soil moisture drought is essential for operational water management and early warning systems. Yet, hydrological model simulations of drought often diverge substantially, even when forced with identical meteorological inputs. This study assesses the extent to which Earth Observation (EO) data can constrain model calibration and influence multi-model drought representation in the Rhine basin. Four hydrological and land-surface models (CLM, JULES, mHM, PCR-GLOBWB), simulated at ~1 km resolution, were calibrated using three strategies: (1 – baseline) discharge-only, (2 – EO-only) using satellite soil moisture (SM), evapotranspiration (ET), or both, and (3 – hybrid) calibration integrating discharge and EO constraints. Simulations were evaluated against the ESA CCI Soil Moisture satellite product (v9.1 COMBINED) as a quasi-independent large-scale benchmark and against in-situ observations from the International Soil Moisture Network (ISMN) as a site-scale temporal reference. To ensure comparability, all datasets were transformed into quantile-based Soil Moisture Index (SMI), and model outputs were aggregated to the 0.25° ESA CCI grid for spatial comparison. Across three major drought events (2015, 2018, 2019), EO-only calibration increased inter-model spatial agreement, quantified using the Inter-Model Agreement Index (IMAI), from 0.648±0.087 (baseline) to 0.663±0.039, while reducing event-to-event variability in agreement. Hybrid calibration showed lower ensemble coherence (IMAI = 0.605±0.137), reflecting competing spatial EO and discharge constraints together with reduced ensemble availability for this configuration. At the same time, EO constraints exposed greater spatial heterogeneity among individual model responses. Model behaviour differed structurally: CLM simulated more extensive severe-drought areas, mHM redistributed drought patterns into more localised clusters, JULES showed comparatively limited sensitivity to EO constraints, and PCR-GLOBWB simulated weaker drought intensities under EO calibration. Improvements observed for individual events were not uniform across all events, indicating event-dependent calibration responses. Calibration using combined SM+ET constraints produced intermediate ensemble agreement (0.647±0.034), whereas ET-only calibration yielded smaller and less consistent changes in spatial metrics. Site-scale evaluation at the Niederwerth station provided supporting evidence of temporal performance differences among calibration strategies, with CLM showing reduced RMSE (0.189 to 0.173) and increased correlation (0.70 to 0.76) under hybrid calibration; however, representativeness is limited to this location. Overall, EO-constrained calibration reduces ensemble spread in selected spatial diagnostics while simultaneously exposing structural differences among models. These findings indicate that EO data provide valuable spatial constraints on hydrological model behaviour and highlight trade-offs between improving agreement with observational benchmarks and maintaining inter-model coherence in drought representation.
Competing interests: At least one of the (co-)authors is a member of the editorial board of Hydrology and Earth System Sciences.
Publisher's note: Copernicus Publications remains neutral with regard to jurisdictional claims made in the text, published maps, institutional affiliations, or any other geographical representation in this paper. While Copernicus Publications makes every effort to include appropriate place names, the final responsibility lies with the authors. Views expressed in the text are those of the authors and do not necessarily reflect the views of the publisher.- Preprint
(9555 KB) - Metadata XML
- BibTeX
- EndNote
Status: final response (author comments only)
- RC1: 'Comment on egusphere-2026-1012', Anonymous Referee #1, 17 Jun 2026
-
RC2: 'Comment on egusphere-2026-1012', Anonymous Referee #2, 15 Aug 2026
In this paper, the authors examine to which Earth Observation (EO) data can constrain hydrological model calibration and influence the representation of soil moisture drought across different model structures. Four hydrological and land-surface models (CLM, JULES, mHM, PCR-GLOBWB) are forced with harmonised meteorological inputs over the Rhine River basin and calibrated using three strategies: discharge-only (baseline), EO-only (satellite-derived soil moisture and/or evapotranspiration), and hybrid (EO combined with discharge). Simulations are evaluated against the ESA CCI Soil Moisture product and against in-situ observations from a single ISMN station (Niederwerth), across three major drought events (2015, 2018, 2019), using an Inter-Model Agreement Index (IMAI) together with several temporal and spatial skill metrics. The authors conclude that EO-constrained calibration improves several dimensions of drought representation, exposing structural differences among models.
Overall, I found this study timely and relevant, as it addresses whether EO data add value to multi-model calibration for drought monitoring purposes, and how this affects inter-model consistency. The experimental design is ambitious, combining four models, three calibration strategies and nine experimental configurations, multiple evaluation metrics, and three drought events. Nevertheless, I have substantial concerns that need to be addressed before this manuscript can be considered for publication. In particular: (1) the manuscript never establishes why inter-model coherence should be expected or is a desirable property to begin with, given that the four models tested differ fundamentally in structure and process representation; (2) differences attributed to “model structure” are not disentangled from parameter equifinality, since – from my understanding of the manuscript – a single calibrated parameter set per model and objective function is used; (3) several conclusions drawn from the point-scale evaluation – limited to a single station (Niederwerth) – are generalised well beyond what the data supports; and (4) many spatial-pattern comparisons rely on qualitative visual inspection of maps (with claimed differences that are, in several instances, difficult to actually see) rather than on quantitative metrics. For these reasons, my overall recommendation is major revisions before this manuscript can be considered for publication in HESS.
Major comments
- The manuscript repeatedly treats inter-model coherence (or “agreement”) as a desirable outcome without ever justifying why this should be expected or sought after, given that CLM, JULES, mHM and PCR-GLOBWB differ fundamentally in structure, process representation and numerical implementation (L64–65, L261, L483). Unless a modular modelling system is used (e.g., Noah-MP, SUMMA), interpreting and achieving inter-model consistency seems an extremely difficult task. This point resurfaces in the Conclusions without a mechanistic explanation of why the model structures actually differ (L490: how different are the ET parameterisations across models? How is water stored and released from the soil column in each model?). Relatedly, the terminology used for this concept is inconsistent throughout the manuscript – e.g., “(IMAI)” (L13), “ensemble agreement” (L20), “ensemble coherence” (L14). I recommend that the authors explicitly motivate, early in the Introduction, why multi-model coherence (as opposed to individual-model realism) is a meaningful target for this analysis, and that they adopt one consistent terminology throughout the manuscript.
- A central issue is that differences in model responses to EO calibration are consistently attributed to structural differences, yet a single meta-objective function is used to calibrate each model, without any exploration of parameter equifinality. This means the analysis cannot actually disentangle whether the reported differences arise from genuine structural (i.e., process representation) differences, from the specific parameter set selected during calibration, or from a combination of both. This concern directly undermines several strong statements in the manuscript – e.g., L443 (“the progression Baseline → EO → EO+Q is consistent across all drought events, demonstrating that EO calibration exposes pre-existing structural uncertainty rather than creating new uncertainty”) and L492 (“rather than calibration artefacts”) – which are not supportable without an equifinality analysis (e.g., using multiple behavioural parameter sets per model). I also disagree with the framing in L489, which equates high inter-model agreement with forecast quality/sharpness; ensemble spread (statistical consistency) is a distinct and equally important attribute that is not addressed here. I recommend that the authors either conduct a limited equifinality analysis (even for a subset of models and events) or substantially soften/remove the structural-uncertainty claims made throughout the manuscript.
- Several statements in the Results generalise findings from the point-scale evaluation well beyond what is supported by the data, since this evaluation is limited to a single ISMN station (Niederwerth) within the Rhine basin (L391, L394). For instance, L391 states that results are “consistent with [the fact] that streamflow integration effectively improves point-scale dynamics”, when this was only demonstrated for soil moisture at Niederwerth. I suggest rewording these statements to explicitly reflect the limited spatial representativeness of the point-scale analysis (see my proposed rewording for L391 under “Suggested edits”).
- The description of the experimental and evaluation framework (Section 3) omits several details needed to assess and reproduce the results. Specifically:
- What optimisation algorithm was used, and what is the maximum number of iterations allowed per calibration experiment? (L191)
- What time step is associated with the model evaluations? (L195)
- Which KGE formulation (e.g., Gupta et al., 2009; Kling et al., 2012; Pool et al., 2018; Tang et al., 2021) was used during calibration? (L198)
- At what time scale (I assume monthly, given the use of SMI) are the Taylor-diagram metrics in Figure 6 computed? This description currently appears in the Results and should be moved to the Methods section (L379), together with the material presented around L400.
- I recommend including a table with the mathematical definitions of all evaluation metrics used (PCC, KGE, SPAEF, SSIM, SACC, sCoV, IMAI) and their associated time scales, placed immediately after Equation 1 (L213, L221).
- How are the nine experimental configurations aggregated into the three calibration categories analysed in the Results? Showing only the ensemble mean of each category risks losing important information (L147; L284, L362). Please clarify this procedure and consider showing the individual-experiment results behind each category-level summary, at least as supplementary material.
- Throughout the Results, differences among models and calibration strategies in the spatial maps (Figures 2–5) are described qualitatively, and in several instances the described differences are difficult, or impossible, to actually see in the figures (L321-322, L341, L344). I recommend complementing these qualitative descriptions with quantitative summaries. For instance, you could report empirical probability distributions of the metric of interest computed across all grid cells, for each combination of model structure and calibration strategy, in place of (or in addition to) the current maps (see my comment on Figure 4). Similarly, statements linking model disagreement to topography (e.g., Alpine/glacierised terrain) would be considerably strengthened by explicitly correlating a model-disagreement index with topographic descriptors (elevation, slope) and glacier cover (L300), rather than relying on visual inspection of the maps.
- The manuscript would benefit from citing and building on previous studies that have systematically compared hydrological/land-surface model structures and their process representations when discussing inter-model consistency (Vano et al., 2012; Fenicia et al., 2014; Mendoza et al., 2015; Saavedra et al., 2022; Spieler & Schütze, 2024 and many others). I also strongly recommend including a summary table describing the main process representations (e.g., runoff generation, snow/glacier dynamics, evapotranspiration formulation, soil-moisture storage) of CLM, JULES, mHM and PCR-GLOBWB. This would substantially help readers interpret the structural differences reported later in the manuscript.
Minor comments
- L46: Why “in contrast”? Do the authors consider the uncertainties in precipitation datasets to be negligible? Please clarify or rephrase.
- L74: Please clarify what “standard meteorological forcing” means – are the forcings identical across all model structures?
- L91: Please specify which model structure(s) are being referred to here.
- L97: Please clarify whether three precipitation forcing datasets were used, given that two baseline configurations rely on EMO-1 while the main experiments use HYPER-P.
- L105: Please state explicitly, at this point in the text, that the satellite benchmark is the ESA CCI Soil Moisture product (this is only specified a few sentences later).
- Figure 1: Please add elevation to the study-area map, to help readers identify areas of complex topography.
- L132: Please clarify what “conventional forcing” refers to.
- L137: Do the authors mean “precipitation dataset choice” with “forcing choice” here? Please clarify.
- L163: How is generalisability across model conceptualisations and parameterisation philosophies actually assessed or verified in this study?
- Table 1: Please clarify why JULES and PCR-GLOBWB are unavailable for certain experimental configurations.
- L216: “Best lag” between which variables? Please clarify.
- L232 (SSIM equation): This metric appears to depend on time; if so, please add “t” to the notation, for consistency with the other equations.
- L264: Does the sCoV metric also depend on time? Please clarify, as per the previous comment.
- L272: Is discharge (Q) also considered an “EO variable”? Please clarify the terminology.
- L297: “Which differences” – please specify which model-structure differences are being referred to.
- L299: Please indicate in Figure 1 the location of the Alpine and sub-Alpine regions mentioned in the text.
- L311 and L313: This description is hard to follow for readers not familiar with European geography; please clarify which regions/countries are meant.
- L324-325: Please clarify whether one month, or multiple months, per drought event were considered in this multi-event analysis.
- L340-341: This sentence appears incomplete – please review.
- L348: How do the authors know that the described spatial pattern is actually “coherent”? Please support this with a quantitative measure.
- L365: This statement does not appear to hold for the Baseline experiments, where PCR-GLOBWB provides better results – please check.
- Figure 7 (L407): The description of the quadrants in the ΔKGE–ΔSSIM plane should be moved to the figure caption rather than the main text.
- L467: Please clarify what is meant by “structural transparency.”
Suggested edits
- Abstract (L6): “∼1,km resolution” → “∼1 km resolution” (remove comma, add space).
- L279: “improvement” → “improvement/worsening”, since both temporal and spatial skill may be degraded.
- L336: Please confirm whether this should read “July/2019”.
- L349: Consider rewording to “EO-only calibration yields modest…” for clarity.
- L391: Reword to reflect the actual scope of the point-scale analysis, e.g.: “…consistent with [the fact] that streamflow integration effectively improves soil moisture simulations at the Niederwerth site” (rather than the current general statement about “point-scale dynamics”).
- L415: Delete one of the two consecutive periods (“..”).
References
Fenicia, F., Kavetski, D., Savenije, H. H. G., Clark, M. P., Schoups, G., Pfister, L., & Freer, J. (2014). Catchment properties, function, and conceptual model representation: Is there a correspondence? Hydrological Processes, 28(4), 2451–2467. https://doi.org/10.1002/hyp.9726
Gupta, H. V., Kling, H., Yilmaz, K. K., & Martinez, G. F. (2009). Decomposition of the mean squared error and NSE performance criteria: Implications for improving hydrological modelling. Journal of Hydrology, 377(1–2), 80–91. https://doi.org/10.1016/j.jhydrol.2009.08.003
Kling, H., Fuchs, M., & Paulin, M. (2012). Runoff conditions in the upper Danube basin under an ensemble of climate change scenarios. Journal of Hydrology, 424–425, 264–277. https://doi.org/https://doi.org/10.1016/j.jhydrol.2012.01.011
Mendoza, P. A., Clark, M. P., Mizukami, N., Newman, A., Barlage, M., Gutmann, E., et al. (2015). Effects of hydrologic model choice and calibration on the portrayal of climate change impacts. Journal of Hydrometeorology, 16(2), 762–780. https://doi.org/10.1175/JHM-D-14-0104.1
Pool, S., Vis, M., & Seibert, J. (2018). Evaluating model performance: towards a non-parametric variant of the Kling-Gupta efficiency. Hydrological Sciences Journal, 63(13–14), 1941–1953. https://doi.org/10.1080/02626667.2018.1552002
Saavedra, D., Mendoza, P. A., Addor, N., Llauca, H., & Vargas, X. (2022). A multi‐objective approach to select hydrological models and constrain structural uncertainties for climate impact assessments. Hydrological Processes, 36(1). https://doi.org/10.1002/hyp.14446
Spieler, D., & Schütze, N. (2024). Investigating the Model Hypothesis Space: Benchmarking Automatic Model Structure Identification With a Large Model Ensemble. Water Resources Research, 60(7). https://doi.org/10.1029/2023WR036199
Tang, G., Clark, M. P., & Papalexiou, S. M. (2021). SC-earth: A station-based serially complete earth dataset from 1950 to 2019. Journal of Climate, 34(16), 6493–6511. https://doi.org/10.1175/JCLI-D-21-0067.1
Vano, J. A., Das, T., & Lettenmaier, D. P. (2012). Hydrologic Sensitivities of Colorado River Runoff to Changes in Precipitation and Temperature. Journal of Hydrometeorology, 13(3), 932–949. https://doi.org/10.1175/JHM-D-11-069.1
Citation: https://doi.org/10.5194/egusphere-2026-1012-RC2
Data sets
4dHydro's Open Science Catalog 4dHydro Consortium https://4dhydro.eu/catalog/
Model code and software
CLM K. Oleson et al. https://github.com/HPSCTerrSys/CLM3.5
mHM L. Samaniego et al. https://github.com/mhm-ufz/mHM
PCR-GLOBWB E. H. Sutanudjaja et al. https://github.com/UU-Hydro/PCR-GLOBWB_model
Interactive computing environment
4DHydro WP6/SC2 Python E. Modiri and O. Rakovec https://codebase.helmholtz.cloud/4dhydro/wp6/sc2
Viewed
| HTML | XML | Total | BibTeX | EndNote | |
|---|---|---|---|---|---|
| 394 | 151 | 33 | 578 | 32 | 38 |
- HTML: 394
- PDF: 151
- XML: 33
- Total: 578
- BibTeX: 32
- EndNote: 38
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
The authors implemented Earth Observation data to constrain a set of different hydrological models and used three different strategies to evaluate the improvements in drought representation at both spatial and temporal scales. The topic is interesting and addresses current important gaps and research questions, which fit for the journal and are relevant for the community.
The paper is well structured and clearly written. I have one major comment in its current format and recommend some minor revisions to improve clarity and readability.
I reviewed the paper, and I would highlight the following concerns:
Major comment:
Minor comments:
I hope that my comments are helpful in improving the overall quality of the work for both the authors and the journal.