the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
Multiple-objective calibration of a conceptual hydrological model using satellite data of snow cover, soil moisture and limited streamflow observations
Abstract. One way to improve hydrological predictions in data-sparse regions is to assimilate satellite data of water cycle components into the calibration of hydrological models. This study evaluates the value of combining satellite snow cover (MODIS) and soil moisture (ASCAT) data with limited streamflow observations to improve hydrological model calibration at the regional scale. The study compares model performance of eleven calibration variants that differ in (a) whether they use satellite data only, (b) the type and temporal distribution of streamflow observations used, and (c) whether satellite and streamflow data are combined. The streamflow sampling strategies cover two scenarios: regularly spaced observations distributed over multiple years (one value per season or per month), and event-based strategies that mimic a single short-term gauging campaign in the wettest, driest, or average year (a peak-flow event with recession plus six bimonthly background samples). The analysis is performed for 213 catchments in Austria, grouped into 119 alpine and 94 lowland catchments. The results show that calibration to satellite data only provides reliable runoff simulations primarily in lower-elevation, drier, and more agricultural lowland catchments. In alpine catchments, adding any limited streamflow data substantially improves model efficiency. The combination of monthly streamflow observations and satellite data (Vmo+sat) results in the best overall runoff performance in lowland catchments, with a median validation runoff efficiency of 0.67. In alpine catchments, event-based streamflow-only strategies achieve median validation runoff efficiencies of 0.69–0.71, close to the regular monthly variant (Vmo, median 0.75) and substantially better than satellite-only calibration (Vsat, 0.26). In lowland catchments, Vmo+sat (0.67) outperforms all event-based variants. Adding satellite data to any streamflow-based variant reduces median snow-cover errors in alpine catchments by approximately a factor of five and consistently improves the simulated soil moisture, although it can reduce runoff efficiency in alpine catchments compared to streamflow-only calibration. These results support the practical value of short, targeted gauging campaigns combined with satellite remote sensing for hydrological modeling in data-sparse regions.
- Preprint
(1855 KB) - Metadata XML
- BibTeX
- EndNote
Status: open (until 29 Aug 2026)
-
RC1: 'Comment on egusphere-2026-2880', Anonymous Referee #1, 17 Aug 2026
reply
-
AC1: 'Reply on RC1', Asma Khalil, 24 Aug 2026
reply
Dear Editor and Reviewers,
We sincerely thank the Editor and the Reviewers for their careful evaluation of our manuscript and for their constructive and valuable comments. We have carefully considered all the comments and have revised the manuscript accordingly. We believe that these revisions have substantially improved the quality and clarity of the manuscript.
In the attached file, we provide a point-by-point response to the reviewer’s comments. The reviewer’s comments are reproduced in Italics, followed by our responses. All changes made in the revised manuscript are in Bold for ease of reference.
We hope that the revised manuscript and our responses adequately address the concerns raised by the reviewers, and we appreciate the opportunity to improve our work.
Sincerely,
Asma Khalil and all the Co-authors
-
AC1: 'Reply on RC1', Asma Khalil, 24 Aug 2026
reply
-
RC2: 'Comment on egusphere-2026-2880', Valentina Premier, 18 Aug 2026
reply
This work presents different calibration strategies for a conceptual hydrological model in the context of data scarcity across catchments with different hydroclimatic characteristics, including both Alpine and lowland catchments in Austria. The results suggest that incorporating remotely sensed information on snow cover and soil moisture into the calibration leads to only limited improvements in model performance, particularly for the lowland catchments.
The paper is well written, and the different sections are clearly structured and presented. The topic is relevant, and I find the proposed approach particularly interesting as a potential strategy for improving hydrological modelling in poorly monitored catchments.
These findings are broadly consistent with observations from our recent study (https://hess.copernicus.org/articles/30/1189/2026/), where we found that constraining the snow module with remote sensing data—through post-calibration replacement of the snowmelt coefficient—resulted in slightly improved or largely unchanged streamflow performance in the target catchments. Although streamflow performance did not substantially improve, we argued that the additional information could nevertheless improve the internal consistency of the model by providing a more realistic representation of processes and states such as snow cover and soil moisture.
An interesting difference, however, is that in our study the snowmelt parameter was replaced after calibration, while the remaining model parameters were left unchanged. We therefore hypothesized that a complete recalibration, allowing the other parameters to adjust to the additional constraint on the snow module, could potentially lead to improved streamflow performance. In the present study, however, the results appear to show that the inclusion of additional remotely sensed information can actually lead to a deterioration in streamflow performance, particularly for the Alpine catchments, where one might expect snow processes and therefore snow-cover information to play a comparatively important role.
I think this result deserves further discussion. In particular, the authors could better explain why constraining the model with snow-cover information does not translate into improved streamflow simulations in the more snow-influenced catchments, and why it may even lead to poorer performance. Could this arise from the multi-objective calibration strategy itself (used methodology, objective function definition, parameter range selection)? It would also be useful to investigate whether the response is related to the actual degree of snow influence of each catchment (e.g., fraction of precipitation falling as snow, snow-cover duration, or relative importance of snowmelt to runoff), rather than simply the Alpine/lowland classification. In fact, I would expect an improvement especially in Alpine catchments where snow is more dominant and in-situ measurements are sparser and hence, the use of satellite data should be an added value.
I am also wondering about the statement that the internal consistency is improved, which should be justified by the results shown in Figs. 5 and 6. I agree with the statement, but the results look a bit obvious. Of course, if the calibration function directly considers those metrics, adding the satellite observations and therefore those additional constraints will improve Osm and Osc. I have the impression that the robustness is evaluated using the same metrics that are included in the calibration, without showing a clear improvement in terms of streamflow prediction.
Specific comments:
L61/L83 The focus on high-flow events is understandable, but is it also justified in a context of drought conditions? In this context, I wonder how well low flows are represented by the model and by the different calibration strategies.
Table 1: Mean annual precipitation includes both liquid and solid precipitation. Given the importance of snow processes for some of the investigated catchments, it would be informative to provide an indication of the partitioning between rainfall and snowfall (e.g., fraction of annual precipitation falling as snow). This would also help the reader assess the degree to which each catchment can be considered snow influenced and would support the later interpretation of the results.
L115 For the interpretation of the results, and particularly in relation to the statement in Sect. 5.1 that poorer model performance may be associated with the sparse network of in situ stations used to interpolate the meteorological forcing, I suggest providing information on the number and/or density of available stations for each catchment. This could be included in an Appendix. More importantly, I suggest investigating this hypothesis more explicitly: is there evidence that catchments with poorer model performance are indeed characterized by a sparser station network? Such an analysis would provide stronger support for the interpretation proposed in Sect. 5.1.
L121 Why are you not using newer version 6.1?
L125 I am not an expert of this product, but so far as I know soil moisture products can be quite uncertain. Please provide the uncertainty of this product and discuss if this can also influence results.
Sec. 3.1 Is clustering just a trick to aggregate and visualize the results?
Table 2 How are the ranges of the parameters chosen? For DDF, for example, we allowed for a larger range. Do you think that constraining the range can also lead to a non-optimal calibration? Are the parameters calibrated for each variant falling within the ranges or at the bounds? In my opinion, it would also be interesting to provide information on how much these parameters vary from one variant to another and which of them are more sensitive to the different variants.
L167 How is a pixel of snow identified? MOD10A1 is a fractional snow product. Do you set a hard threshold and convert it to binary snow/snow free information?
L171 Having as calibration objective function something that counts for the Pearson correlation coefficient, could lead to results that are highly biased in terms of Soil Moisture Content (but still correlated)?
Table 3: It is not clear whether the satellite products are also considered as one observation per month/year or whether you consider all available satellite observations. It would also be meaningful to provide information on how many satellite observations are actually used for both snow and soil moisture, considering limitations such as cloud obstruction and the number of available acquisitions.
Fig 2 Could the bad performance of cluster 4 also be due to the fact that a small number of catchments are within that cluster?
Figure 3 Don’t you consider forest coverage as an indicator? And the soil type/effects of groundwater recharge? Is the hydrological regime also influencing the results? Catchments with higher/lower Q perform differently? Snow cover duration?
Fig 10 There is a strong underestimation for clusters 3 and 4 during the months when the snowmelt peak occurs (right?). Are these catchments characterized by a strong snow influence? Is snowmelt not fully explaining the observed discharge during this period? Are there also strong storm/rainfall events occurring at the same time? How does the hydrograph look during this period?
Sec 5.4 needs to be expanded and better discussed.
Citation: https://doi.org/10.5194/egusphere-2026-2880-RC2 -
AC2: 'Reply on RC2', Asma Khalil, 24 Aug 2026
reply
Dear Editor and Reviewers,
We sincerely thank the Editor and the Reviewers for their careful evaluation of our manuscript and for their constructive and valuable comments. We have carefully considered all the comments and have revised the manuscript accordingly. We believe that these revisions have substantially improved the quality and clarity of the manuscript.
In the attached file, we provide a point-by-point response to the reviewer’s comments. The reviewer’s comments are reproduced in Italics, followed by our responses. All changes made in the revised manuscript are indicated in Bold for ease of reference.
We hope that the revised manuscript and our responses adequately address the concerns raised by the reviewers, and we appreciate the opportunity to improve our work.
Sincerely,
Asma Khalil and all the Co-authors
-
AC2: 'Reply on RC2', Asma Khalil, 24 Aug 2026
reply
Viewed
| HTML | XML | Total | BibTeX | EndNote | |
|---|---|---|---|---|---|
| 46 | 25 | 7 | 78 | 5 | 3 |
- HTML: 46
- PDF: 25
- XML: 7
- Total: 78
- BibTeX: 5
- EndNote: 3
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
General Comments:
In this preprint, authors aim to address relevant and distinct research objectives with clear added value on top of their methodological predecessor [Tong et al. (2021)], specially by introducing event-based calibration experiments & limited flow observations combining with satellite snow-cover & soil moisture data. Comparison of results between alpine and lowland/hilly catchments from simulation of large sample data of Austria is indeed important strength of this study.
However, some of the main conclusions are difficult to interpret casually as experimental design does not isolate the effect of individual factors. Particularly, effect of adding satellite data cannot be separated from the effect of simultaneous reduction of weights assigned to the streamflow objective. Observation numbers and timing vary concurrently across sampling strategies; event campaigns are defined using retrospective information and uncertainty assessment is missing for supporting several small performance difference interpretations. Notably, most interesting regional patterns revealed remain inadequately explained with hydrological process-based understandings.
This manuscript has therefore high potential but needs careful major revision to enrich interpretation before consideration for publication. Below are some specific comments that Authors should consider to address and improve it for publication:
Specific Comments:
Introduction
Study Area & Data
Methods
Clustering
Calibration and Optimization
Results
Discussions
Limitations
Conclusions
Minor/Editorial Comments: