the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
Streamflow and satellite evapotranspiration are asymmetric calibration targets in water-limited catchments
Abstract. Reliable hydrological modelling in data-scarce regions is constrained by the scarcity of streamflow observations, which has driven increasing adoption of satellite-derived actual evapotranspiration (AET) as a supplementary calibration target. Whether AET provides parameter-constraining information complementary to streamflow, or merely trades discharge skill for evapotranspiration reproduction, remains contested. This study compared three calibration schemes: streamflow-only (Q-only), AET-only, and joint streamflow-AET calibration (Q+AET), applied to three conceptual rainfall-runoff models (GR6J, HBV, IHACRES) across 17 water-limited watersheds spanning the hydroclimatic gradient of the West African Sahel during 1991–2020. Parameter identifiability was assessed via variance reduction and distributional shift from prior to behavioural samples. Long-term water-balance estimates were evaluated against the Budyko-Fu framework, with the reference evaporative index anchored both on GLEAM AET and independently on water-balance closure. Results show that streamflow provided substantially more parameter-constraining information than AET. AET-only calibration reproduced evapotranspiration dynamics but left the production, soil-moisture and loss-module parameters governing runoff generation unconstrained. Q-only calibration constrained water partitioning sufficiently to recover AET dynamics in HBV and IHACRES, while AET in GR6J remained poorly constrained under streamflow-only conditioning. Joint calibration preserved streamflow skill of Q-only while matching or exceeding AET-only on AET simulation. It emerged as the only scheme to yield a simultaneously coherent long-term runoff ratio and evaporative index across all three model structures, which held under both Budyko reference anchors. The cost of multivariable calibration was governed by model structure: the direction of parameter displacement imposed by each variable (opposing in GR6J, weakly coupled in HBV and coincident in IHACRES) predicted whether joint calibration was structurally costly or essentially free. These results suggest that streamflow should remain the primary calibration target wherever observations are available, however, satellite-derived AET estimates remain valuable as an independent check on long-term water-balance plausibility rather than as a substitute calibration signal.
- Preprint
(4229 KB) - Metadata XML
-
Supplement
(4657 KB) - BibTeX
- EndNote
Status: final response (author comments only)
-
RC1: 'Comment on egusphere-2026-3981', Manfred Fink, 02 Sep 2026
-
AC1: 'Reply on RC1', Roland Yonaba, 17 Sep 2026
Technical corrections and clarifications
- I think table 3 is not necessary in manuscript because these Performance metrics are standard and the newest one (KGE) is already addressed in detail in the paper.
We agree. KGE is already defined in Eq. (1) and R², RMSE and PBIAS are standard. In the revised manuscript we will remove Table 3 from the main text and move it to the Supplement (as Table S2) for completeness of the formulations, retaining in Sect. 2.5.2 only a short sentence naming the four metrics, the variables and the periods over which they are computed.
- The lines in fig 8 should differ for different models (dashed, dotted ore something accordinly).
Agreed, and we thank the referee for catching this. Figure 8 currently maps calibration scheme to colour and model to point shape, but the nine LOESS curves are all drawn solid, so the three curves sharing a colour cannot be told apart. We will map model to line type (GR6J solid, HBV dashed, IHACRES dotted) while keeping colour for scheme. We also note that the current caption states that the curves pool the three models, which does not match the figure; the caption will be corrected accordingly.
- Please control the resolution of the graphics sometimes text is almost not readable fig 14 and especially S13 - S15.
We apologise for this and will correct it. All figures will be regenerated at 500 dpi (previously 350 dpi), with an increased base font size. For the 17-panel monthly time-series figures (S12-S14), the inset performance text is the main casualty of the panel density; we will enlarge it. Figure 14 will additionally be simplified following a suggestion from Referee #2 (colour for calibration scheme, symbol shape for model), which removes a large part of the clutter.
Questions
- The AET product from GLEAM takes differences from land cover into account, whereas PET is just interpolated grass reference PET from stations. Has this the mismatch consequences which should be mentioned?
This is a fair point and the mismatch is real. It is, however, a deliberate consequence of the design: as set out in lines 155-162, GLEAM produces potential and actual evaporation through a single evaporative-stress formulation, so taking PET from GLEAM as well would have made the AET/PET ratio an internal property of the product and the AET-based calibration schemes partly self-referential. Using station-based FAO-56 PET breaks that dependency at the cost of the surface mismatch the referee identifies. We will make the consequences explicit in the revised text, along the following lines:
- PET here is an index of atmospheric evaporative demand over a reference grass surface, not a land-cover-specific demand. The land-cover-dependent scaling is therefore absent from the forcing and must instead be absorbed by the calibrated parameters. This is explicit in IHACRES, where the efficiency factor e (range [0.01, 1.50]) rescales PET directly, and implicit in HBV (through the LP/FC threshold) and GR6J (through the filling function of the production store). It is a plausible part of why e is displaced toward the upper part of its range under both fluxes (Fig. 7).
- The mismatch affects the absolute value of the aridity index φ = PET/P and hence the position of each watershed along the Budyko x-axis. It cannot, however, generate the between-scheme contrasts that are the subject of the paper: φ is fixed by the forcing and is therefore identical for observed and simulated points and identical across the three calibration schemes.
- The mismatch is also the reason the AET ≤ PET consistency check was necessary (l. 163–165); exceedances were confined to about 2 % of the record. Evaporation efficiency AET/PET remains between 0.21 and 0.52 across the 17 watersheds (Table S1), so the reference-surface PET still bounds AET meaningfully over the whole gradient.
We will add two to three sentences to Sect. 2.2.2 and a corresponding item to the limitations paragraph in Sect. 4.
- The analyses of the north–south gradient is maybe also influenced by a systematically different catchment size in the different latitudes. Additionally in the Niger part you will find a lot of mixed signals because the large catchment include also others in a nested manner. Could these overshadow the analyses by latitudes only? Especially in the middle part the catchments tends to be larger then in the north or in the south.
We accept this criticism. Catchment area in our sample spans 1,113 to 75,907 km² and is not distributed evenly along the gradient, and several of our stations are indeed nested rather than independent. Referee #2 actually raises a closely related concern about record completeness co-varying with latitude, and we intend to treat the two together in the revision.
We propose the following, none of which requires re-running the models:
- Report the Spearman rank correlation between watershed area and latitude, and between area and the watershed-mean VR and KS, across the 17 watersheds. These are computed from quantities already stored and will be given in the text with a scatter panel added to Fig.
- Identify the nested stations explicitly in Table 1 and in the caption of Fig. 1, so that readers can see which points are not hydrologically independent.
- Rewrite Sect. 3.1.3 and the corresponding discussion to state plainly that latitude, aridity, catchment area and record completeness co-vary in this 17-watershed sample, and that the present design cannot separate their contributions. We will describe the pattern as a co-varying north–south gradient rather than attribute it to climate alone, and add it to the limitations.
We would stress that this concerns a secondary result. The central findings of the paper (which is the asymmetry of information content between the two fluxes and the physical coherence of the joint scheme) rest on the parameter-space diagnostics (Sect. 3.1.1, 3.1.2), the performance comparison (Sect. 3.2) and the Budyko analysis (Sect. 3.3), none of which depends on the latitudinal ordering.
We thank the referee again for these comments, which we believe will make the manuscript both clearer and more honest about what it can and cannot resolve.
Citation: https://doi.org/10.5194/egusphere-2026-3981-AC1
-
AC1: 'Reply on RC1', Roland Yonaba, 17 Sep 2026
-
RC2: 'Comment on egusphere-2026-3981', Anonymous Referee #2, 04 Sep 2026
General Comments (Overall quality):
This work investigated hydrological model performance across 17 watersheds in Burkina Faso. Three different hydrological models were calibrated against discharge (Q) only, remotely sensed actual evapotranspiration (AET) only, and these two metrics together. Results show the water balance is captured the best when both Q and AET are used jointly, but streamflow alone can carry enough information to yield acceptable AET simulations. Results show the utility of remotely sensed AET data to improve model simulations. However, stream discharge measurements maintained paramount importance if the models were to perform well across the water balance. The authors recommend that AET is valuable for validating the water balance of hydrological models, but that discharge should be the primary calibration target for hydrological models.
I enjoyed reading this manuscript and it stimulated a lot of thinking, which is always a rewarding sign of reading a nice study. The manuscript is overall well written, and I think could offer an interesting contribution to the literature. However, there is an objective function I have some questions about, and the discussion section needs substantial revision before this manuscript is ready for publication. Several results-section statements step into interpretation of results and should be moved to the discussion. As it stands the discussion would be stronger with sections that address main arguments of the manuscript rather than one lumped section.
I recommend major revisions, focused on the points below, plus specific minor and technical comments:
- It should be clearer why AET-only calibration overestimates Q and underestimates AET. Lines 595-602 hint at an answer, that matching AET is achieved at the cost of routing excess water to streamflow for IHACRES and HBV, but this should be stated in the discussion and connect clearly to the paper’s main claims.
- Related to point 1, I am struggling with the objective function for AET. I can understand why the KGE for the monthly values was included, however, I’m not convinced that the KGEs should be equally weighted here, especially when the values are square root transformed. The monthly KGE would more often outperform the daily KGE, exactly due to this reduction of noise and variation in the data. Additionally, AET is larger in the summer months and thus produces smaller errors in the square root transformed space and allows the model to under-predict summer AET without significantly impacting the objective function. I think these may be driving some of the overestimation of Q when using the AET only model and should be addressed.
- I’m not quite convinced that the latitudinal trend is NOT due to data-completeness. And discussion around this feels inconsistent at points in the manuscript. The authors note that the stronger streamflow constraint toward the semi-arid north (Lines 494-495) is not attributable to data completeness. However, the evidence cited in Lines 492-493 say records are more complete in the same northern watersheds where constraint is stronger. Additionally, Figure S3 shows data completeness follows the same north-south gradient and, to me, supports data completeness contributing to the trend.
Specific Comments (Individual scientific questions):
A threshold of 50% data gap seems rather high; I appreciate that there were checks on the distribution of the data. However, was the distribution of the gaps in the data also considered temporally such as high or low flow times or seasonally? I think this is what is being discussed in terms of the Sahelian hydrometrics, but this was not clear to me.
Do the number of parameters matter in terms of assessing if extra data adds constraining information?
It would be helpful to include figures showing the comparison of the model structures. This would help the reader significantly when interpreting why models may perform differently based on structure.
Lines 252 – 253: Why were the calibration and validation periods split sequentially rather than using k-fold cross-validation or a differential split-sample approach? Was this to keep the starting conditions consistent?
Additionally, given that climate and land-use may have changed over the study period, could the sequential split make it harder to discern if poorer model fits on validation data stems from the model itself or simply because of the different conditions during each time period?
Lines 446 – 448: The posterior distribution of LP under Q-only and Q+AEP look almost exactly the same to me? I’m not sure I’d compare it to the same behaviour as X1 in the GR6J model.
For the parameters that show little change from their prior distribution, could this be influenced by the performance threshold chosen to define “acceptable” parameter sets? If a different cutoff was used would these results hold?
Lines 477-478; 490-495: Points related to data completeness influencing the stronger parameter constraints along a latitudinal gradient would benefit from some rewriting for greater clarity. Attributing the trend to data gaps seems to be contradicted in lines 492-493 and Lines 494 – 495. At this point it seems the data-completeness argument is most logical with Figure S3 showing that data-completeness and constraint strength follow the same north-south gradient.
Figure 10 – 12: I think it would be really nice to see precipitation for each of these panels to see all the components of the water balance.
Technical Corrections (typing errors, etc.):
Graphical abstract: I think the graphical abstract covers the scope of the work well. However, do you actually consider the AET performance poor for all the models? The abstract almost says the opposite for HBV and IHACRES in Lines 25-26. A minor comment is the “over AET” is also coloured brown, which at first glance, made it appear to relate to the AET-only points.
Line 23: Should GLEAM be spelled out before it is used, or rather would it be more informative for a ready if it was mentioned as satellite-derived AET at this point.
Figure 1: I find the orange watershed names difficult to read on this map.
Table 1: The “n” in Elongation Ratio is missing
Lines 168 – 170: Here you say that stations were retained if their overall data-gap rate did not exceed 50%. Therefore, I’d suggest rephrasing the sentence in Lines 123 – 124 because this gives the impression there are no gaps in the data for the selected stations.
Line 126: comma after respectively
Lines 206-210: I recommend rearranging these sentences since the definition of the CMD acronym is after the first time it is used. Also, I don’t find the IHACRES acronym defined in the manuscript.
Line 272: “justifying” instead of “that justify”.
Figures 2 & 3: It may be worth checking that boxes without outlines are not surrounded by black outlined boxes. Based on line thicknesses it seems like there may be some boxes that don’t have an outline but look like they do because their neighbours are outlined.
I find that when many sentences begin with “Figure X shows….” Information is often repeated and less clear. I much prefer writing, and find it easier to follow, when sentences lead with the finding and cite figures parenthetically. For example, the two sentences in Lines 583-585 are mostly saying information that is repeated over the following two paragraphs.
Figure 14: Would it be more consistent with other figures if the Budyko plots showed the calibration as different colours and the models as different shapes?
Figure 15: The colours for the scenarios should stay consistent (Q-only: blue, AET-only : orange, and Q+AET : green). Additionally, the panel colours are the same as the river basin colours, which may cause confusion on first glance.
Figure S3: The river basin colours are not the same as in the main manuscript.
Citation: https://doi.org/10.5194/egusphere-2026-3981-RC2 -
AC2: 'Reply on RC2', Roland Yonaba, 18 Sep 2026
We are grateful to Referee #2 for the thorough review. The three main points are well taken. Each comment is reproduced below and followed by our response.
Response to the Discussion restructuring
We accept that a single undifferentiated Discussion is the weakest part of the manuscript. Section 4 will be reorganised into subsections:
- 1 Why AET-only calibration mis-partitions the water balance
- 2 Model structure and the coupling node between evaporative and runoff-generating pathways
- 3 Why our results differ from studies in which AET-only calibration succeeds
- 4 Implications for calibration practice in water-limited catchments
- 5 Limitations and future work
Interpretative statements currently in the Results, particularly those at lines 595-602, 636-648 and 663-671 will be moved into these subsections, to leave the Results section fully descriptive.
Response to the main points
- It should be clearer why AET-only calibration overestimates Q and underestimates AET. Lines 595-602 hint at an answer … but this should be stated in the discussion and connect clearly to the paper’s main claims.
Agreed. We will state the mechanism explicitly, in four steps:
- Under AET-only, the objective conditions only the production, soil-moisture and loss-module parameters; the routing and recession parameters remain at their prior distributions in all three models (Figs. 2b, 3b, 4). Nothing in the objective therefore controls either the volume or the timing of the water that is not evaporated.
- Because P and PET are fixed by the forcing and long-term storage change is small (|dS| ≤ 8.9 % of P, Table S1), closure requires that any deficit in simulated AET appear as an excess in simulated Q. The compensation is not a coincidence but an arithmetic consequence; and it is exactly what the mirror-image lower-right cluster in Fig. 15b, c shows.
- AET is systematically underestimated rather than matched because a modest negative bias is cheap under our AET objective function. This is a point that Referee #2 develops in comment 2 and which we address there.
- GR6J behaves differently because its production store mediates both fluxes through a single capacity parameter, AET-only pushes X1 upward while leaving the runoff-generating path entirely unconstrained, producing a degenerate, non-seasonal hydrograph (Fig. 10) rather than a compensating excess. This is the same structural conflict that makes Q-only leave AET unskilled in GR6J and it is the link to our central claim about coupling.
- I am struggling with the objective function for AET … I’m not convinced that the KGEs should be equally weighted here, especially when the values are square root transformed. The monthly KGE would more often outperform the daily KGE … Additionally, AET is larger in the summer months and thus produces smaller errors in the square root transformed space and allows the model to under-predict summer AET without significantly impacting the objective function.
We think the referee’s reading is correct, and we are defending our design as optimal. Two clarifications can be proposed:
Our rationale is that the monthly term was included because GLEAM daily AET is itself partly a modelled quantity and carries day-to-day retrieval noise that we did not want to dominate the fit. The square-root rather than log transformation was chosen because AET spans no orders of magnitude and contains no near-zero values. The equal weighting was chosen for structural symmetry with OF(Q) in Eq. (2), which likewise averages two KGE terms, so that OF(Q+AET) in Eq. (4) would not be implicitly weighted toward one flux. We accept that symmetry of form is not the same as symmetry of information content, and that the monthly term is the easier of the two to satisfy.
We do not think this drives the asymmetry: OF(AET) enters OF(Q+AET) unchanged. If the formulation of OF(AET) were sufficient to produce the runoff overestimation, the joint scheme would inherit it. It does not: Q+AET retains essentially all of the Q-only discharge skill (Fig. 9a, b) and its optima cluster at the origin of the Δ(Q/P)-Δ(AET/P) plane in all three models (Fig. 15). The objective-function design may well set the magnitude of the AET-only bias; the absence of any streamflow constraint is what sets its existence and direction. We will make this argument explicitly in Sect. 4.1.
- I’m not quite convinced that the latitudinal trend is NOT due to data-completeness … the evidence cited in Lines 492-493 say records are more complete in the same northern watersheds where constraint is stronger. Additionally, Figure S3 shows data completeness follows the same north-south gradient and, to me, supports data completeness contributing to the trend.
The referee is right and we withdraw the claim. The sentence at lines 492–495 does not follow: we intended to say that the northern stations are not the degraded ones, but as the referee observes, completeness increasing northward alongside constraint increasing northward is precisely the pattern a data-completeness explanation predicts. Figure S3 therefore does not exclude that explanation, and presenting it as if it did was an error. In the revision we will:
- delete the assertion that the trend is not attributable to record completeness;
- rewrite lines 477-478 and lines 490-500 to state that constraint strength, record completeness, aridity and catchment area all co-vary along the north-south axis in this sample, and that the present design cannot separate them (this also answers RC1’s comment on catchment size, and we will treat the two together);
- report the Spearman rank correlation between watershed-mean VR (and KS) and station gap rate, and the partial correlation of constraint with latitude controlling for gap rate, across the 17 watersheds. These use quantities already computed and will be added to the text with supporting panels in Fig. S3. Based on the outcome, we will temper or drop the latitudinal interpretation.
Two points survive this correction and we will make them carefully. First, the contrast between schemes is not affected: AET-only uses a gap-free target at all 17 watersheds and still shows no coherent latitudinal organisation, so the spatial structure we describe is specific to the streamflow signal whatever its cause. Second, Sect. 3.1.3 is a secondary result; our paper’s conclusions rest on Sects. 3.1.1, 3.1.2, 3.2 and 3.3, none of which uses latitude.
Response to specific comments
A threshold of 50 % data gap seems rather high … was the distribution of the gaps in the data also considered temporally such as high or low flow times or seasonally?
This is a reasonable concern and we admit that our text was not clear. Gaps in the network are dominated by whole-year and multi-year station outages associated with the documented decline of the Sahelian hydrometric network (Belemtougri et al., 2021), rather than by random dropouts within a monitored year; this is visible in the year-by-year structure of Fig. S2, where missing data tend to occupy entire years. We will state this explicitly, report the actual distribution of gap rates (median and range) rather than only the 50 % ceiling.
Do the number of parameters matter in terms of assessing if extra data adds constraining information?
A good question, which we will address in Sect. 4.2. Three elements: (i) VR and KS are computed per parameter and are scale- and dimension-independent, so a model with more parameters is not mechanically penalised by the metrics themselves; (ii) dimension does affect sampling: a 250,000-member LHS covers the 9-dimensional HBV prior less densely than the 6-dimensional GR6J and IHACRES priors, which bears on behavioral retention (Fig. S5) and on posterior precision, though not on the interpretation of the metrics; (iii) what matters for our argument is not the total parameter count but the partition into AET-exclusive, Q-exclusive and shared parameters, which is the coupling-node argument of objective (ii). We will make that distinction explicit in the revision.
It would be helpful to include figures showing the comparison of the model structures.
We agree and thank the referee for the suggestion. We will add a schematic figure placing the three structures side by side, with the shared soil-moisture or production state, the AET pathway and the runoff-generating pathway highlighted in each.
Lines 252-253: Why were the calibration and validation periods split sequentially rather than using k-fold cross-validation or a differential split-sample approach? Was this to keep the starting conditions consistent?
Partly, yes. A single continuous record with one 1989-1990 warm-up avoids repeated re-initialisation of storage states and this matters in a regime where soil storage is strongly depleted each dry season.
We can also mention two further reasons: first, the streamflow gaps are unevenly distributed in time, so k-folds would contain very unequal numbers of valid evaluation days and the folds would not be comparable across watersheds; second, a single fixed temporal framework was necessary to keep the design tractable across 3 models × 3 schemes × 17 watersheds × 250,000 parameter sets. We will state these reasons in the revision.
Additionally, given that climate and land-use may have changed over the study period, could the sequential split make it harder to discern if poorer model fits on validation data stems from the model itself or simply because of the different conditions during each time period?
Yes, and we accept this as a limitation. Land use and rainfall-runoff relations in this region have changed materially over 1991-2020 (the Sahelian paradox; Yonaba et al., 2021a, 2025b), so validation-period degradation confounds model deficiency with changed conditions. We will add this explicitly to Sect. 4.5 and note that a differential split-sample test is the natural follow-up. We would add that this affects all three schemes identically at a given watershed, so it does not affect the between-scheme comparison that is our subject.
Lines 446-448: The posterior distribution of LP under Q-only and Q+AET look almost exactly the same to me? I’m not sure I’d compare it to the same behaviour as X1 in the GR6J model.
The referee’s reading of Fig. 6 is correct and ours was overstated. For HBV-LP, the joint posterior essentially reproduces the Q-only posterior instead of lying between the two single-flux positions and the analogy with GR6J-X1 does not hold. We will rewrite the passage to say that LP is displaced in opposing directions by the two single-flux schemes but that the joint posterior follows the streamflow position, which is in fact more consistent with our central claim that streamflow dominates when both fluxes are used together.
For the parameters that show little change from their prior distribution, could this be influenced by the performance threshold chosen to define “acceptable” parameter sets? If a different cutoff was used would these results hold?
This could be tested: the GLUE ensemble is already run and all objective values are stored, so recomputing VR and KS at a different threshold is a re-filtering of existing output. We propose to recompute the aggregate constraint metrics (the Fig. 4 equivalent) at OF > 0.2 and OF > 0.5 and to add the result as a supplementary figure, stating whether the ranking of parameters and of schemes is preserved. We will also make clear that a stricter threshold reduces posterior sample size and lowers VR for every parameter, so the diagnostically meaningful quantity is a relative ordering, not the absolute level.
Lines 477-478; 490-495: Points related to data completeness … would benefit from some rewriting for greater clarity.
Agreed; this is covered by our response to main point 3 above.
Figure 10–12: I think it would be really nice to see precipitation for each of these panels to see all the components of the water balance.
We agree in principle. The difficulty is that these figures are already dense. Referee #1 has flagged legibility across the figure set, and adding a third series to 34 sub-panels risks making them worse. Our proposal is to add the interannual-mean daily P as an inverted bar series on a secondary axis in the upper (Q) sub-panel only, and to keep it if legibility holds at the increased resolution. We note that long-term P is already reported per watershed in Fig. S6a and Table S1.
Responses to suggested technical corrections
We accept all of these and are grateful for their precision.
- Graphical abstract. Panel A will be corrected to state that Q-only AET performance is poor for GR6J and acceptable for HBV and IHACRES, consistent with lines 25-26. The “OVER AET” and “UNDER AET” quadrant labels will be recoloured neutral so they no longer echo the AET-only point colour.
- Line 23. GLEAM will be introduced as “satellite-derived AET (GLEAM)” at first use in the abstract.
- Figure 1. The orange watershed labels will be replaced with a higher-contrast colour with a light halo.
- Table 1. “Elongatio Ratio” will be replaced with “Elongation Ratio”.
- Lines 123–124. “continuous daily streamflow records” will be rephrased to be consistent with the ≤ 50 % gap criterion stated at lines 168–170.
- Line 126. Comma added after “respectively”.
- Lines 206–210. The paragraph will be reordered so that CMD is defined at first use and IHACRES will be expanded at first use (i.e., Identification of unit Hydrographs And Component flows from Rainfall, Evaporation and Streamflow).
- Line 272. “that justify” will be replaced with “justifying”.
- Figures 2 & 3. We will check the outline weights so that cells without an outline are not read as outlined and add the outline convention to the legend.
- Prose style. We accept this and will rewrite sentences opening with “Figure X shows…” to lead with the finding and cite the figure parenthetically, beginning with lines 583–585, which duplicates the two paragraphs that follow.
- Figure 14 will be recoded so that colour denotes calibration scheme and shape denotes model, consistent with the other figures.
- Figure 15. The scheme colours will be harmonised to Q-only blue, AET-only orange, Q+AET green, and the panel/basin palette changed to avoid the clash.
- Figure S3. Basin colours will be harmonised with the main manuscript.
Citation: https://doi.org/10.5194/egusphere-2026-3981-AC2
Viewed
| HTML | XML | Total | Supplement | BibTeX | EndNote | |
|---|---|---|---|---|---|---|
| 204 | 79 | 28 | 311 | 60 | 22 | 17 |
- HTML: 204
- PDF: 79
- XML: 28
- Total: 311
- Supplement: 60
- BibTeX: 22
- EndNote: 17
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
In general, I like topic and I think that it is work out in comprehensive manner. The topic is not entirely new but the paper adds interesting aspects to multiresponse calibration/valiation topic for parsimonious hydrological models. Nevertheless I see potential to improve the manuscript with minor (basically technical) corrections/clarifications:
Furthermore, I have two questions whose answers do not necessarily need to be included in the manuscript, depending on what the answers are: