the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
Do reservoir-influenced gauges need explicit consideration in machine learning models? A case study with Hydra-LSTM
Abstract. Reservoirs fundamentally alter downstream river flow regimes, decoupling discharge from natural meteorological forcing and challenging standard hydrological prediction. While data-driven models, such as Long Short-Term Memory (LSTM) networks, show promise in regulated catchments, it remains unclear how training data composition across natural and regulated rivers influences model generalisability and behaviour. In this study, we investigate how the presence or absence of reservoir-influenced catchments in training data impacts model performance across different flow regimes and alters the physical drivers the models learn to rely on. Using carefully matched subsets of the CAMELS-GB dataset, we trained separate specialist LSTMs (reservoir and non-reservoir), a pooled Full LSTM, and a multi-headed Hydra-LSTM to investigate whether explicit architectural specialisation offers any advantage over pooled training alone. Models were evaluated on held-out test gauges using standard performance metrics and gradient importance analysis to interpret feature reliance. Our results demonstrate that exposure to reservoir-influenced catchments during training is essential. Models trained exclusively on natural catchments consistently overestimate the mean and variance of regulated flows. Conversely, training exclusively on reservoir-influenced data degrades performance on non reservoir-influenced rivers (KGE reduction of ≥ 0.1) giving importance primarily to anthropogenic static features, such as abstraction rates, at the expense of precipitation drivers. A single Full LSTM trained on combined data matched the performance of both specialist models in their respective domains, implicitly switching its feature reliance between regimes. The Hydra-LSTM performed comparably to the Full LSTM throughout, indicating that the shared body may act as a regulariser limiting over-specialisation, but that explicit architectural specialisation provides no further benefit under these conditions. We conclude that pooling training data across regimes is a highly effective strategy for general-purpose modelling. However, case studies highlight a fundamental limitation: purely meteorological inputs remain insufficient for predicting flows in heavily managed single-purpose reservoirs, where unobserved human operational decisions dominate the hydrograph.
- Preprint
(4953 KB) - Metadata XML
- BibTeX
- EndNote
Status: final response (author comments only)
-
RC1: 'Comment on egusphere-2026-2909', Saskia Salwey, 21 Jul 2026
-
AC1: 'Reply on RC1', Karan Ruparell, 18 Sep 2026
Thank you for your detailed comments and advice, they were all very helpful and I think will really improve the paper. I've tried to group and address all the points below:
RC: The novelty and research questions should be better situated in the existing literature. Why does this need to be tested?We have revised the Introduction (for example paragraph 3) to clarify the motivation and novelty of the study. Rather than asking only whether LSTMs can represent reservoir-influenced flow, we explicitly distinguish two questions:
- Exposure to reservoir-influenced training data: Do models need to be trained on reservoir-influenced gauges to generalise well to unseen regulated catchments?
- Explicit model specialisation: Given such training data, is a reservoir-specific model or architecture needed, or can a single model trained on pooled reservoir and non-reservoir gauges perform equally well?
Our results separate these two effects. Exposure to reservoir-influenced gauges is important, whereas explicit specialisation provides little additional benefit: models trained without reservoir-influenced gauges perform substantially worse on regulated catchments, while a single pooled model performs comparably to the corresponding specialist models. We now highlight this distinction more explicitly in the Introduction.
Prior work demonstrating that LSTMs can learn regulated flow behaviour has not controlled for training data composition, making it impossible to attribute performance differences to exposure rather than other dataset properties. We now state this directly.
RC: A third case study would be useful — one where the reservoir model clearly outperforms the non-reservoir model.
Agreed, and we will Spey at Boat Garten, which we have found the be a good representation of the average case. Here, the Reservoir LSTM achieves an NSE of 0.77, compared to -1.22 for the Non-Reservoir LSTM, with all other models ranging from 0.15 to 0.45. This gauge provides a clear illustration of what reservoir training adds in an intermediate case and we have added it as a third case study in the revised manuscript.
RC: Some discussion of LSTM performance relative to alternative hydrological modelling approaches would be useful. Salwey et al. (2024) improved NSE at Vyrnwy from -1.23 to 0.28 using simple operating rules.
This is an excellent point, and although running alternative hydrological models ourselves is outside the scope of this paper, we will add some discussion on \citet{salwey2024developing} and suggest further comparison for future work. The key comparison would be with the nationally consistent calibration, as the focus of this paper is on ungauged basins. In both settings, the models cannot fully reproduce the sharp releases in heavily managed reservoirs, or the lack of response after a rainfall event caused by the water being immediately stored. For a quantitative comparison on Vyrnwy, we will see if we can calculate the NSe on the same period. By making sure we are computing the NSE over the exact same dates, we can offer a fair comparison.
RC: More discussion on societal relevance and implications for modelling practice.
We have added a paragraph to the conclusion making the practical implications explicit: a single pooled model matches specialist performance in both regimes, which simplifies operational deployment for practitioners building national flood forecasting or water resource management systems.
RC: Figure 1 — which catchment boundaries are shown, and why can't elevation be read from the map?
The catchment boundaries shown are the overlap of upstream catchments in CamelsGB, so they represent the distinct areas covered by the study. We have clarified this in the text, and will see if we can add elevation to the plot as stated.
RC: None of the models reproduce the plateau releases at Vyrnwy — why?
We have added an explanation to the case study discussion. We believe that this is related to the fact that the models receive only meteorological inputs varying in time, and since this is an ungauged forecasting problem, it has trouble figuring out the exact release criteria. What it can do, however, is learn a better approximation of how much the variance of flow is controlled, depending on the different reservoir attributes provided.
NRC: Line 311 typo
Citation: https://doi.org/10.5194/egusphere-2026-2909-AC1
-
AC1: 'Reply on RC1', Karan Ruparell, 18 Sep 2026
-
CC1: 'Comment on egusphere-2026-2909', Jesús Casado Rodríguez, 24 Jul 2026
General Evaluation
The manuscript trains different deep learning models on the dataset CAMELS-GB searching for a procedure that is able to reproduce both natural and regulated catchments. It creates a paired dataset of natural and regulated catchments and trains two different architectures (LSTM and Hydra-LSTM) to three different data samples: all catchments, only regulated or only natural catchments. The results indicate that special architecture for reservoirs are not necessary, but the model must be exposed to regulated catchments.
The paper is interesting as it is an open question how to model regulated catchments in the typical lumped structure of the LSTM models used in hydrology. The comparison of models using a paired dataset is very interesting for this benchmarking. However, the methods and data used in the paper are not well explained. The two model architectures should be explicitly defined, and their differences should be made clear. There is no reference to what dynamic inputs are used in the models. The results lack an analysis of model performance against degree of regulation, which would help assessing whether this specific results in Great Britain could be extrapolated ot different hydrological regimes.
I recommend minor revision before publication.
Major comments
Lines 89-93. I understand the logic behind the pairing of the reservoir-influenced and natural catchments; the idea is to provide model trainings with the same amount of data. However, this procedure is limiting the amount of data the model is trained on. My understanding is that the capacity of deep learning models relies on the amount of data, compute and size of the model. Isn't this reduction of data availability limiting the capacity of the models? I understand that this is done here to make a fair comparison among models, but it could probably be mentioned. Even further, the Full LSTM or Hydra-LSTM: Main could be trained on the whole dataset.
Section 2.1 Data Sources and Preprocessing. There is no mention to the dynamic inputs in the dataset. There could be an Appendix B defining them. Are they just meteorological forcings from ERA5? Are there any temporal encoders used to help reproducing reservoir operations? Depending on the reservoir use, the operations are driven by temporal features. For instance, hydropower generation will be correlated with working days, but irrigation or water supply reservoirs have clear seasonal patterns. How are these models supposed to learn these patterns? Could these temporal encoders be fed directly to the reservoir head?
In connection to the following comment about the Methods, it would be interesting to explain where in the model are the dynamic and static inputs fed. From the original reference of Hydra-LSTM (Ruparell et al., 2025), this architecture can receive inputs directly on the head. Is that the case on the Hydra-LSTM:Reservoir head?Section 3.1 Models Trained. This section is a bit convoluted. The hypothesis behind the Hydra-LSTM models are introduced before the actual models. Even if there are citations to the specific models, I would introduce briefly the LSTM and Hydra-LSTM architectures and the particularities of the Hydra-LSTM model, so unfamiliar readers can follow. I would list or use bullet points to clarify the hypothesis. A figure summarizing the structural differences among models would be helpful.
Section 5 Discussion. As I will detail in the minor comments, the paper lacks an analysis of how heavily regulated are the catchments in the dataset, and how the different models perform in the most regulated samples. Modelling low regulated catchments, even in the presence of reservoirs, it is relatively easy for a general (Full) model, as the natural streamflow is not dramatically affected, so the performance metrics may be relatively good. The challenge is modelling heavily regulated catchments. Given the climate and orography of Great Britain, this may not be the best study case for heavy regulation, which does not invalidate the results, but it should be specified that extrapolation to other climates and regimes is not certain.
Minor comments
Line 28. Incomplete citations. The date is missing in "UK Centre for Ecology & Hydrology", and the cite "of Civil Engineers, 2015" is incomplete.
Line 35. Incomplete citation in "Yoshimi et al."; the date is missing.
Line 38. Reservoir operations may not be readily available in Great Britain, but national datasets exist in other countries like ResOpsUS (https://zenodo.org/records/6612040) or ResOpsBR+CARS (https://zenodo.org/records/16096623).
Line 54. I reckon there are more relevant references to process-based reservoir schemes:
Hanasaki, N., Kanae, S., & Oki, T. (2006). A reservoir operation scheme for global river routing models. Journal of Hydrology, 327(1–2), 22–41. https://doi.org/10.1016/j.jhydrol.2005.11.011<br>
Haddeland, I., Skaugen, T., & Lettenmaier, D. P. (2006). Anthropogenic impacts on continental surface water fluxes. Geophysical Research Letters, 33(8). https://doi.org/10.1029/2006GL026047<br>
Zajac, Z., Revilla-Romero, B., Salamon, P., Burek, P., Hirpa, F., & Beck, H. (2017). The impact of lake and reservoir parameterization on global streamflow simulation. Journal of Hydrology, 548, 552–568. https://doi.org/10.1016/j.jhydrol.2017.03.022<br>
Turner, S. W. D., Steyaert, J. C., Condon, L., & Voisin, N. (2021). Water storage and release policies for all large reservoirs of conterminous United States. Journal of Hydrology, 603. https://doi.org/10.1016/j.jhydrol.2021.126843<br>
Hanazaki, R., Yamazaki, D., & Yoshimura, K. (2022). Development of a Reservoir Flood Control Scheme for Global Flood Models. Journal of Advances in Modeling Earth Systems, 14(3). https://doi.org/10.1029/2021MS002944<br>
Salwey, S., Coxon, G., Pianosi, F., Lane, R., Hutton, C., Bliss Singer, M., McMillan, H., & Freer, J. (2024). Developing water supply reservoir operating rules for large-scale hydrological modelling. Hydrology and Earth System Sciences, 28(17), 4203–4218. https://doi.org/10.5194/hess-28-4203-2024<br>
Shrestha, P. K., Samaniego, L., Rakovec, O., Kumar, R., Mi, C., Rinke, K., & Thober, S. (2024). Toward Improved Simulations of Disruptive Reservoirs in Global Hydrological Modeling. Water Resources Research, 60(4). https://doi.org/10.1029/2023WR035433<br>
Line 81. Shouldn't the Scottish Environmental Protection Agency publication be cited?Figure 1. It could be interesting to show reservoir size and degree of regulation in this figure. It would help to understand the level of regulation of British rivers for comparison against other datasets, or to correlate degree of regulation with model performance---probably the higher the regulation the harder to model. If that's the case, the results in GB may not be applicable to heavily regulated rivers such as those in (semi)arid climates.
Line 147. To be more consistent with the LSTM models, I would name this Hydra-LSTM: Full. That would clarify the difference between Hydra-LSTM: Full and the specialists Hydra-LSTM: Res-Head and Hydra-LSTM: NonRes-Head. I understand that both models have the same structure, first Full is trained on the whole dataset. The body is frozen and only the head is trained independently on the two datasets.
Line 156. Improve readability as it is a bit repetitive.
Lines 163-167. This is a very interesting setup. I think it would be interesting to further explain the benefits of training sequence-to-sequence. Why it the window size 90? If feels like a short window only looking at 3 months of data. Think about reservoirs whose degree of regulation is of one or a few years.
Line 227. It would ease reading if the notation was consistent. The "Hydra reservoir head" was previously called "Hydra-LSTM:Res-Head".
Line 247. How is the head in Hydra-LSTM Reservoir Head trained? Is the body freezed and the head trained from scratch? Or is the whole model fine tuned? It is not specified.
The fact that Hydra-LSTM Res-head can't really reproduce better reservoir behaviour may imply that it didn't learn during this finetuning/retraining process. This seems to be a problem in both specialists' heads.Table 3. It is not clear to me what it means that "metrics were computed over a 90-day rolling window". Since this is a hindcast model, shouldn't it produce the complete time series for the test period and compare it against observations?
Figure 2. The labels in the legend are not consistent with the naming in the paper. It's understandable to which model it refers, but it hinders readability. It's not clear what it means "each point in the curve representing the average KGE in a different gauge"; there aren't points in the figure; what's the "average KGE" of a gauge?. There is a typo in "-a minimum KGE", remove the hyphen.
Lines 256 and 265. It would be interesting to show the degree of regulation or the
degree of disruptivity of these two reservoirs to get a clearer picture of how much the regulate streamflow. Definitions of these two attributes can be found in Shrestha et al. (2024), indicated above. In their work, they come up with threshols of these two values indicating reservoirs whose disruptivity is that low that they don't affect river discharge downstream. It could be the case that Gwyfrai is a non-disruptive reservoir, reason why all models perform similarly well.Line 278. The sentence "percentage of the upstream catchment area that is inland water" is not clear. Does it mean the percentage of the catchment area that is regulated by reservoirs?
Lines 283-285. The fact that the relevant static features are almost identical is expected as the models share most of the parameters, i.e., the body. It could also be an indicator that the training of the heads didn't work.
Line 286. There is an issue in this reference. Is it a Figure or a Table? It was first referenced as Table 5, but here it's referenced as a Figure, and the actual tables are named as a Figure.
Lines 294-295. This statement needs an analysis of the model performance versus degree of regulation (or another metric of human intervention). It could be that the median/average performance is good enough because the majority of reservoir-influenced gauges are not heavily regulated, as it is the case of Gwyfrai.
As a humid country, my guess is that the reservoir regulation in GB is not very strong. That doesn't mean that these results can be extrapolated to arid or semi-arid climates, where rivers are strongly regulated.Figure 5. I think that the reorganizing the table columns by type of model (LSTM vs Hydra) would help visualize that the feature importance is similar among each type of model, particularly for the Hydra models.
Line 305. Could you extract these worse-performing cases and analyse the degree of regulation? Are there good-performing gauges with a high degree of regulation?
Line 311. Typo. Replace "gayges" by "gauges".
Lines 343-345. To understand this, the Hydra architecture should've been mentioned in the Methods. The Hydra-LSTM has not been introduced, particularly the changes compared to the LSTM.
Line 350. Typo. "Howeverm".
Lines 363-365. The covariate for this pattern analysis is not geographic location as an indicator of climate or geology, but also the level of human regulation. As mentioned before, the paper lacks an analysis of the model performance related to the level of regulation.
Line 365. Typo. "indivdual".
Lines 377-378. Could the reservoir operations be learned from data, instead of being dynamic inputs? That data is available in some countries (ResOpsUS, ResOpsBR+CARS), or satellite estimates could be used.
Lines 387-388. I would restrict this statement to (semi)humid climates like Great Britain. Extrapolation to (semi)arid climates, where rivers are heavily regulated, is to be tested, as mentioned already in Line 396.
Line 399. I'm not sure that the cite (Mason and Dance, 2026) is relevant in the field of remote sensing for reservoir modelling. There is plenty of recent literature in that field:
Schwatke, C., Dettmering, D., Bosch, W., & Seitz, F. (2015). DAHITI - An innovative approach for estimating water level time series over inland waters using multi-mission satellite altimetry. Hydrology and Earth System Sciences, 19(10), 4345–4364. https://doi.org/10.5194/hess-19-4345-2015
Pekel, J. F., Cottam, A., Gorelick, N., & Belward, A. S. (2016). High-resolution mapping of global surface water and its long-term changes. Nature, 540(7633), 418–422. https://doi.org/10.1038/nature20584
Schwatke, C., Dettmering, D., & Seitz, F. (2020). Volume variations of small inland water bodies from a combination of satellite altimetry and optical imagery. Remote Sensing, 12(10). https://doi.org/10.3390/rs12101606
Donchyts, G., Winsemius, H., Baart, F., Dahm, R., Schellekens, J., Gorelick, N., Iceland, C., & Schmeier, S. (2022). High-resolution surface water dynamics in Earth’s small and medium-sized reservoirs. Scientific Reports, 12(1). https://doi.org/10.1038/s41598-022-17074-6
Khandelwal, A., Karpatne, A., Ravirathinam, P., Ghosh, R., Wei, Z., Dugan, H. A., Hanson, P. C., & Kumar, V. (2022). ReaLSAT, a global dataset of reservoir and lake surface area variations. Scientific Data, 9(1). https://doi.org/10.1038/s41597-022-01449-5
Hao, Z., Chen, F., Jia, X., Cai, X., Yang, C., Du, Y., & Ling, F. (2024). GRDL: A new global reservoir area-storage-depth data set derived through deep learning-based bathymetry reconstruction. Water Resources Research, (60). https://doi.org/10.1029/2023WR035781
Hou, J., van Dijk, A. I. J. M., Renzullo, L. J., & Larraondo, P. R. (2024). GloLakes: Water storage dynamics for 27000 lakes globally from 1984 to present derived from satellite altimetry and optical imaging. Earth System Science Data, 16(1), 201–218. https://doi.org/10.5194/essd-16-201-2024Citation: https://doi.org/10.5194/egusphere-2026-2909-CC1 -
AC3: 'Reply on CC1', Karan Ruparell, 18 Sep 2026
Thank you for your detailed feedback, it is very helpful and we are grateful for the time. I am also glad that you found the paired catchment approach interesting and useful. I've tried to group your comments and
RC: The matching procedure limits model capacity and should be mentioned. The Full LSTM could be trained on the whole dataset.Agreed on the acknowledgement, we have added this explicitly to the data section. However, we believe that limiting the Full LSTM to the catchments used by the two Specialist LSTMs is important. Holding the total catchment set constant across all configurations means the only variable between the Reservoir LSTM, Non-Reservoir LSTM, and Full LSTM is how data is organised during training, not how much is available. If we think of comparing one operational approach as ‘using all available data in one model’ and another as ‘using each basin in one of the specialist models’, training a Full LSTM on the complete unmatched dataset would not help us answer this. This is admittedly hidden in the current manuscript, and so we have changed it to explicitly give this reasoning.
RC: No mention of dynamic inputs. Are there temporal encoders for reservoir operations?
Thank you for pointing this out, we have added Appendix B listing all dynamic inputs with units and sources. All inputs are meteorological forcings from CAMELS-GB; no temporal encoders are used. We acknowledge this limits the model's ability to learn operation-driven patterns tied to working days or seasonal demand cycles, and note that routing temporal encoders directly into the specialist heads is an interesting future direction.
RC: Where are dynamic and static inputs fed in the model? Can inputs be fed directly to the Hydra head?
We have added a clarifying sentence to Section 3.1: dynamic inputs enter only through the shared Body; the heads receive only the Body's encoded output. Routing dynamic inputs directly into the specialist heads is flagged as future work.
RC: Section 3.1 is convoluted. Architectures should be introduced before hypotheses, with a figure.
Restructured. Both architectures are now described before the hypotheses, with the existing architecture figure referenced at the point of introduction.
RC: The paper lacks an analysis of performance versus degree of regulation.
We investigated available proxies in CAMELS-GB and found that reservoir contributing area (%) is the most informative metric, varying from near zero to 100% across the dataset. We have added an analysis of model performance against this variable and included it in the Discussion. Looking outside CAMELS-GB for additional regulation metrics is not feasible within this study, and we note that even within CAMELS-GB there are limitations: the dataset does not record when reservoirs ceased operation, meaning some gauges may include periods of reduced regulation that we cannot account for. We have added a caveat on extrapolation to heavily regulated arid and semi-arid environments.
RC: Line 28, 35 — incomplete citations.
The UKCEH citation has been completed. The Yoshimi et al. reference is a poster at the NeurIPS 2022 Workshop on Tackling Climate Change with Machine Learning and has been updated accordingly.
RC: Line 38 — national operational datasets exist in other countries.
Added: ResOpsUS \citep{steyaert2021resopsus} and ResOpsBR+CARS \citep{casado2025resopsbr} are cited as examples.
RC: Line 54 — additional process-based reservoir references.
We have added the suggested references where relevant to the discussion of process-based approaches.
RC: Line 81 — SEPA should be cited.
Added \citep{sepa_reservoir_register}.
RC: Line 147 — rename to Hydra-LSTM: Full for consistency.
Agreed and done throughout the manuscript.
RC: Line 156 — repetitive.
Condensed into a single paragraph merging justification and result for each variable.
RC: Lines 163-167 — explain seq2seq benefits and justify 90-day window.
Added. The seq2seq training allows stronger updates to the model per training batch, and leads to smoother training. We do not use a longer sequence so we can have more distinct batches to train the model and minimise overfitting. We initially chose a 90-day sequence to be able to look at high and low flow seasons distinctly, but also see the value in comparing with a 365 day window. We are currently rerunning our methods using a 365 day window, and would be happy to reframe our analysis using a 365-day window.
RC: Line 227 — notation inconsistency.
Fixed in the full notation pass.
RC: Line 247 — how is the Res-Head trained?
The Body is frozen and the head is initialised from scratch. Now stated explicitly. On the reviewer's suggestion that the head may not have learned: the Res-Head does outperform the stand-alone Reservoir LSTM on non-reservoir gauges (KGE 0.77 vs 0.63), which we attribute to the regularising effect of the shared Body, indicating the head training was effective.
RC: Table 3 — rolling window metrics are unclear.
Added a clarifying sentence: each evaluation day uses a fresh 90-day lookback for consistent state initialisation, avoiding bias toward later days in the record.
RC: Figure 2 — legend inconsistency, "each point" wording, hyphen typo.
Legend updated to match paper notation, caption reworded, hyphen removed.
RC: Lines 256, 265 — show degree of regulation for case study gauges.
We have added quantitative characterisation using CAMELS-GB attributes: reservoir storage capacity relative to mean annual precipitation and reservoir contributing area. These are described at the point of case study selection.
RC: Line 278 — "inland water" is unclear.
Reworded to clarify this refers to the proportion of upstream area classified as standing inland water bodies in the land cover data.
RC: Lines 283-285 — similar Hydra feature importance may reflect shared body, not learning.
This is an understandable uncertainty, and we have made changes to the explanation it in the paper. When looking at the behaviour in particular gauges, like Vynrwy, we see that the heads can behave quite differently, which reassures us that each of the heads are learning to behave differently.
RC: Line 286 — Figure or Table reference inconsistency.
Corrected.
RC: Lines 294-295, 305, 363-365 — performance versus degree of regulation analysis is needed.
Addressed via the reservoir contributing area analysis described above. We note the CAMELS-GB limitation on operational history and restrict our extrapolation claims to humid climates.
RC: Lines 377-378 — could reservoir operations be learned from data?
It seems not with meteorological temporal data, but datasets such as ResOpsUS and ResOpsBR+CARS may make this feasible in some regions. We have added this as a future work direction.
RC: Lines 387-388 — restrict extrapolation claim to humid climates.
Done.
RC: Line 399 — Mason and Dance 2026 is not the most relevant earth observation reference.
Replaced with more directly relevant remote sensing references from the suggested list, covering satellite altimetry and surface water extent datasets.
RC: Typos lines 311, 350, 365.
All corrected.
Citation: https://doi.org/10.5194/egusphere-2026-2909-AC3
-
AC3: 'Reply on CC1', Karan Ruparell, 18 Sep 2026
-
RC2: 'Comment on egusphere-2026-2909', Claudia Bertini, 12 Aug 2026
The manuscript “Do reservoir-influenced gauges need explicit consideration in machine learning models? A case study with Hydra-LSTM” investigates whether specialised specialized LSTM models perform better than pooled LSTMs in reproducing fully regulated and completely natural flows. The results show that pooled models are the better option to reproduce both the flow regimes tested compared to ad-hoc LSTMs.
I found the paper very interesting and of relevance to the scientific community. I believe it fits the aim and scope of Hydrology and Earth Systems Sciences (HESS) journal. While overall the manuscript reads well, I found some parts to be challenging to follow, and I suggest some modifications in the presentation. Here are my comments:
- Section 3.1 introduces first the hypotheses on all the models used, including the Hydra-LSTM, and then describes the configuration of the Hydra-LSTMs (from line 153 onwards). I know that reference to the paper describing the general Hydra-LSTM concept is provided beforehand, but since the configurations used here are slightly different from the original one, the description of the HydraLSTMs should come before the hypotheses description
- Section 3.1 (lines 153-161) provides the description of the HydraLSTMs, but it is not clear on which dataset is trained the Hydra Body (only static descriptors?)? and what is the difference of the training sets between Hydra Main and Hydra Body?
- Following the previous comment, in line 159 it is mentioned that the weights of the Body are frozen while training the Res Head, but what about the other heads? I suppose it is the same, but I think it should be mentioned explicitly.
- Table 1 should probably include the info about the Hydra Body and Hydra Non-Res as well, to be complete
- Section 3.2 on the Training Methods introduces the Seq2Seq training (line 163), but is not very clear how the switch happens from 90 days flow simulation to one day only. Also, line 256 adds up on the confusion because the NSE is computed on sequence of 90 days, while I had understood that the target of the LSTM models was one day only. I probably misunderstood, but I think this part should be explained a bit better, including stating clearly the target of the models.
- Section 4 presents the results of all models. This is the section I found most challenging to follow because of the confusion between Hydra Main and Body. For instance, Line 230 and line 259 mention the HydraLSTM but it is not clear which one, if the specialised ones, the Main? Lines 239-241: what about the performance of Hydra Res Head?
- Line 276 and 278: I think the reference to Table 5a and 5b should be changed into Figure 5a and 5b.
- Table 5 presents the performance of HydraLSTM specialists together. Do they have exactly the same scores?
- Figure 1 introduces the paired gauges employed and it seems that in some cases one catchment of the non-reservoir cluster is matched with two gauges of the reservoir cluster. Is this not creating an imbalance of samples? How many are the final samples in both specialised models? This is also helpful to know later on if the results of the Full LSTM are this good because there is (almost) a 1:1 proportion between the two regimes.
Citation: https://doi.org/10.5194/egusphere-2026-2909-RC2 -
AC2: 'Reply on RC2', Karan Ruparell, 18 Sep 2026
Thank you so much for your feedback and for your interest in the paper, I'm very happy to see that you found it interesting for the community. I also really appreciated your feedback, and in particular have made changes to the description of the Hydra to make it clearer without requiring a reader to go to the reference paper (and especially highlighting the differences from the reference paper as you pointed out). Here I have grouped and tried to address all the comments, I hope you find them satisfactory!
RC: The Hydra-LSTM description should come before the hypotheses.
Agreed and done. Section 3.1 now describes both architectures before the hypotheses are stated.
RC: It is unclear what the Hydra Body is trained on, and what distinguishes Hydra Main from Hydra Body.
We have added explicit signposting. The Body is trained jointly with the Main Head *Which we now call the Full Head) on the full combined dataset. The Full Head is the prediction component of that joint training, producing the Hydra-LSTM: Full configuration. The specialist heads are then trained with the Body frozen. The inputs are passed through the body, and then the outputs of the body are fed to the head to produce effect. The training loss is then compute and the parameters of the head are optimised to minimise the loss, while the parameters of the body are not optimised. Table 1 has been updated to make this explicit for all components.
RC: The Body is frozen during Res-Head training — what about the other heads?
The Body is frozen during training of both specialist heads. Both are initialised from scratch rather than fine-tuned from the Full Head. This is now stated explicitly.
RC: Table 1 should include Hydra Body and Hydra NonRes-Head.
Corrected. All components are now listed with an updated caption explaining the training sequence.
RC: The Seq2Seq section is unclear — how does the switch to one-day prediction happen, and why is NSE computed over 90 days?
We have rewritten the training methods paragraph to better explain this. During training, loss is computed across all 90 timesteps, which was a compromise between data stochasticity (requiring a longer days of data would have reduced the amount of data we could use during training, as gauges had intermittent streamflow data availability and dealing with missing data was outside the scope of this paper) and allowing the model to learn storage dynamics. We are currently testing using a 365-day NSE to better compare with literature. This is now stated clearly, including the distinction between the training and evaluation modes.
RC: Results section is hard to follow — Hydra Main vs Body vs Res-Head is ambiguous throughout.
We have worked to make this clearer in the text
RC: Lines 276 and 278 should reference Figure 5, not Table 5.
Corrected.
RC: Does Table 5 show models with identical scores?
No — the Hydra-LSTM: Specialists row represents a routing strategy: the Res-Head evaluated on reservoir gauges and the NonRes-Head on non-reservoir gauges, aggregated. We have clarified this in the caption and surrounding text.
RC: Figure 1 appears to show one-to-many pairings.
Every non-reservoir gauge is matched with exactly one reservoir gauge. The apparent one-to-many matches reflect geographic proximity between gauges. We have added an explicit statement to the caption confirming the strict 1:1 pairing.
Citation: https://doi.org/10.5194/egusphere-2026-2909-AC2
Viewed
| HTML | XML | Total | BibTeX | EndNote | |
|---|---|---|---|---|---|
| 142 | 60 | 23 | 225 | 15 | 12 |
- HTML: 142
- PDF: 60
- XML: 23
- Total: 225
- BibTeX: 15
- EndNote: 12
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
This paper investigates the importance of the type of training sets used to simulate flow in reservoir-influenced catchments using LSTM’s. The findings suggest that exposure to reservoir-influenced gauge records is essential for simulating flow in impacted catchments and that models trained only in natural catchments cannot capture the dynamics of a regulated flow regime. Interestingly, the results show that a full LSTM trained on a combined dataset could match the performance of both specialist models. The paper reads well and presents important and interesting findings but would benefit from a clearer explanation of the wider impact of the results as well as a more detailed explanation of where these finding sit in the broader literature. Below are some suggestions for how the paper might be improved.
Major comments:
The novelty of this study and its research questions should be better situated in the context of the existing literature. At the moment it is hard to understand where the authors research questions have come from and why these are questions that need to be answered. I would assume that a model trained on reservoir data only will not perform well in natural catchments (and visa versa), so could you better explain why this needs to be tested? Perhaps previous work has suggested that this is not the case? Or perhaps the novelty lies instead in understanding whether pooling reservoir and natural gauge data into a training set could be a more efficient approach? Paragraph 3 in the introduction could be a good place to discuss this in more detail.
I really like how the authors have chosen two case studies to focus on, this works very well. However, it might be useful to add a third case study to this section as a middle ground between the two the authors have currently presented. Currently it seems that at Vyrnwy neither the reservoir nor non-reservoir models are very good (I can see that the reservoir model has made improvements but the scores are still relatively low), and at Gwyfrai both models perform reasonably well. Could the authors perhaps pick a third case study where the reservoir model clearly makes large improvements on the non-reservoir model by recreating some of these very ‘reservoir’ behaviors that non-reservoir models simply can’t recreate?
I realize this is not the focus of the manuscript but I think this analysis would benefit from some discussion about the performance of the authors various LSTM models in comparison to alternative hydrological modelling approaches. I admit that I am biased because I spent quite a lot of time trying to improve the simulation of reservoir-impacted catchments in a semi-distributed hydrological model but I think that a comparison to the wider literature could be useful nonetheless. As an example, in Salwey et al. (2024) we simulate flow at Vyrnwy and improve the NSE from -1.23 to 0.28 by including simple reservoir operating rules, but similarly to the results in Figure 3 we can’t recreate the nuances of these specific rules enough to improve the simulations further. I’d be interested in some commentary on the pros and cons of the approach adopted in this paper in comparison to others in the literature. I think this would help users in the tricky process of selecting which model to use where!
It would be nice to include more discussion on the relevance of this papers findings for society, and for modelling practices more generally. At the moment the technical results are clear but the manuscript would benefit from better explaining the implications of its results.
Minor comments:
In general I really liked how the authors visualized the paired catchments in Figure 1, but I was slightly confused about which catchment boundaries have been marked on the map? How have these been chosen? Also, on L108 the authors mention that Figure 1 shows the catchment elevation but it is not clear to me how this can be read off the map.
I think it’s interesting how none of the models can recreate the constant flow releases present in the Vyrnwy timeseries, could the authors comment on why they think the model has not been able to learn this behavior? In my experience these plateaus on the hydrograph (which are also reflected in the flow duration curves) are very iconic of reservoir-impacted catchments!
L311 there is a typo here.