the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
Generalization of Deep Learning Models to Ungauged Glacierized Basins: Evidence from Alpine, Patagonian, and North American Catchments
Abstract. Glacierized high-mountain basins supply water to approximately two billion people yet remain among the most data-scarce hydrologic regions globally, making truly ungauged streamflow prediction a critical challenge. Deep learning (DL) offers a promising alternative to traditional regionalization, but fundamental questions remain about when and why DL models generalize to a target domain that is not merely ungauged but hydro-climatically distinct from the training data. We address two questions: under what training data do DL models generalize reliably to completely ungauged glacierized basins? And how does model architecture, including physics-informed DL, modulate sensitivity to these conditions? We systematically evaluate three architectures — Long Short-Term Memory networks (LSTM), Graph Neural Networks (GNN), and differentiable HBV (δHBV) — across four experiments that control for training dataset size, hydroclimatic representativeness, and inclusion of basins with glaciers, using 2,845 basins from the Caravan global dataset with 283 target glacierized basins. We perform 100-trial repeated K-fold cross-validation by holding out glacierized basins as test basins strictly in space and time. Hydroclimatic representativeness of training data- the degree to which training basins cover the target glacierized regime consistently dominates both training data size and architecture choice as the primary determinant of generalization skill. Including glacierized catchments in training provides the strongest representativeness signal, with all three architectures achieving median NSE between 0.66 and 0.71. When glacierized catchments are excluded, LSTM median NSE falls to −1.43 in the most dissimilar partition; larger dataset size only partially improves skill (median NSE −0.96), confirming that dataset size cannot substitute for representativeness. Non-glacierized mountain catchments partially improve skill, demonstrating that partial hydroclimatic representativeness – through inclusion of non-glacierized mountain basins – contributes to model performance in glacierized basins. Architecture differences are secondary: δHBV and GNN show greater resilience under data scarcity due to structural constraints, but no architecture compensates for lack of hydroclimatic representativeness in training data. These findings reframe model selection for ungauged glacierized basins, highlighting the importance of representative training data and the potential limits of “out of sample in landscape” performance of DL models, specifically for DL deployment in climate impact assessments of high-mountain water towers.
- Preprint
(2635 KB) - Metadata XML
-
Supplement
(2721 KB) - BibTeX
- EndNote
Status: final response (author comments only)
- RC1: 'Comment on egusphere-2026-3798', Iiro Seppä, 17 Jul 2026
-
RC2: 'Comment on egusphere-2026-3798', Anonymous Referee #2, 19 Aug 2026
Maharjan et al. investigated how training data, strategy, and the choice of deep learning model impact the efficiency of streamflow simulation in glacierized catchments for ungauged purposes. Their main findings are as follows: 1) Without glacierized catchments in the training set, streamflow efficiency is low; 2) different deep learning (DL) architectures handle small training sets and generalization differently; and 3) including glaciers does not improve streamflow simulations. This study is interesting, and the experiments are well designed. I only have minor comments.
1) I think the paper would benefit from being shortened. Some paragraphs are too long and could be shortened to strengthen the main message. The figures are often overly complex, likely because the authors wanted to be comprehensive. A few suggestions:
- Remove Experiment 4. I struggle to see how this experiment adds value compared to Experiment 3.
- Shorten the abstract to highlight the main hypotheses and results (e.g., articulate around the three aspects I mentioned above).
- Add two hydrographs to the main manuscript (not all subpanels), and remove some of the tables, as well as maybe figure 7 (e.g., move it to the supplementary information). In my opinion, hydrographs are helpful for readers to understand the main differences between model setups.
- I don't think figures 3 and 4 are both needed in the main manuscript.
- I would generally remove tables, especially when some of the figures already show some of the underlying numbers.
- The discussion is too long. For example, the discussion from lines 693 to 727 could be drastically shortened. For example, you could mention only two aspects: the added value of comparing your results to a model with an explicit glacier module, and how the performance reported in this study aligns with existing studies.
- I think one aspect is missing in terms of data selection and the underlying reasoning: some catchments are likely greatly affected by reservoirs. Could seasonal and annual dam storage and release have a similar signal to glacier runoff? If the authors do not want to invest time into dividing the catchment sets into several other groups, I think it would be nice to at least discuss that point.
- I missed a discussion on the time evolution of glacier coverage. Could this impact the results? I did not see this discussed anywhere.
- In the introduction, many references are grouped at the end of long sentences. While this is commonly done, it is better to use citations that are in the correct place and related to the topic of the corresponding sentence. For example, the reference "Thébault et al., (2026)" (line 160 and other instances) is not suited to the statement it refers to. Their study has nothing to do with "projections under future glacier change scenarios." I could not check all the references, but I suggest the authors verify that their references align with their placement.
- The distinction brought by Experiment 5 needs to be made clearer much earlier. I struggled to understand how the glacier data were included in the study before the explanations of the experiments.
- Surface net solar radiation is used for only two of the three DL architectures (Table S2a). An explanation of why this is the case is needed.
- Model efficiency is assessed for summer months in Section 4.2. However, I could not find any reference to this in Section 3.6. The reasoning behind this assessment should be made explicit.
- I think this study would have benefited from more "signature-based" metrics in glacierized catchments. In such catchments, I would expect good KGE values to be easily reached because streamflow is highly autocorrelated. However, I would expect low NSE values because it is difficult to beat the benchmark; the mean is already a good predictor of streamflow. It would be interesting to see some discussion on this topic and how it could influence the results.
Citation: https://doi.org/10.5194/egusphere-2026-3798-RC2 -
RC3: 'Comment on egusphere-2026-3798', Anonymous Referee #3, 21 Aug 2026
The authors investigate how different DL-based hydrological models generalize to ungauged glacierized catchments, and how sensitive this generalization is to the size and hydroclimatic representativeness of the training data. They conclude that including glacierized catchments in the training data is the most important control for effective generalization. The size of the training dataset was found to be of lesser importance. Among the three hydrological models (LSTM, δHBV and GNN), LSTM performed best under larger training dataset size, while δHBV showed higher robustness under lower dataset size. The scientific questions addressed are relevant and the manuscript succeeds in answering them comprehensively. However, the manuscript is very long, and several aspects of the methodology require either justification or alteration.
Below are two major comments and several minor comments that I believe would benefit the manuscript before publication:
Major comments
- The manuscript is currently too long for the message it aims to convey. Below are several suggestions to shorten it:
- Omit experiment 3: with experiment 1 as the baseline, experiment 2 and 4 deviate from it by respectively reducing the training dataset size and by excluding glacierized basins. I do not get the impression that experiment 3 adds information beyond 2 and 4. Referee #2 recommends omitting experiment 4, which I believe would also work. The fact that experiment 4 is not referenced in the discussion speaks for this option.
- Omit the annual NSE and keep only the summer NSE: since the aim is to improve glacier-influenced streamflow, evaluating only the summer months should be sufficient. I do not get the impression that the results are very different among the two metrics.
- Figure 4 could be reduced to just the G training dataset
- The High and Very High glacierized basin categories could be merged, considering that they contain only 4 and 2 basins respectively
- The introduction and discussion chapters are well-structured and comprehensive, but I believe each of the paragraphs can be condensed. I would also omit the part on the computation time in the discussion.
- Figure 2 is informative, but could be reduced in size
- Tables 2 and 3 can be moved to the appendix
- The k-fold sampling currently considers only basin splits, while the temporal training/testing split is kept constant across all experiments with 1988-2005 as the training period and 2008-2014 as the test period. I would recommend reducing the number of basin splits (currently 100) and adding several temporal splits where training, validation and testing periods are shuffled. Considering climate change and the occurrence of decadal droughts in two of the study regions (Garreaud et al., 2025; Hogan and Lundquist, 2024), there is a considerable risk that the assumption of stationarity between training and testing periods is violated. Testing for different combinations of training and testing periods would make the results more robust in this sense. As a side note, it is unclear from the manuscript why the years between 2005 and 2008 are not used, or why 1988 and 2014 define the lower and upper boundaries.
Minor comments:
- I would recommend using a minimum glacier cover before considering a basin glacierized. It is unclear from the manuscript what the exact distribution of the glacier percentages among the 283 glacierized basins is, but I suspect that a number of them have such low percentages that the glacier runoff ratio is negligible, and should therefore not be classified as glacierized basins.
- The authors might consider relaxing the requirement of full streamflow data availability over the study period. At least for the LSTM, Gauch et al. (2025) have shown different effective strategies of dealing with missing data. This could increase the number of training basins, especially in the currently underrepresented Chilean Andes. However, this might not be so straightforward for δHBV and GNN.
- The disaggregation of glacier runoff from monthly to daily timescale (L226-249) is done in a very similar way to the work of Wiersma et al. (2022). However, they use a Tbase of -5 °C and use a calibrated scaling factor to favor melt on warm days and decrease melt on colder days. While 0 °C is a physically sensible threshold, ERA5-Land forcing is known to be biased in mountain areas and therefore possibly calls for a different threshold. Please justify your choices and mention any validation against observed glacier runoff data.
- A definition of the exact summer months used for the NSE calculation is missing
- Why is the HydroAtlas glacier fraction not included as one of the Caravan static attributes (gla_pc_sse)?
- Why is ERA5-Land SWE included as a model input? While it might be the best-performing global SWE product (Mudryk et al., 2024) it can still show strong biases in mountain areas (Lundquist et al., 2026). I suspect that the data-driven models are capable of accounting for the seasonality of snow and glacier processes from the meteorological inputs without relying on the modeled ERA5-Land SWE output. Its use as an input should be justified in the manuscript if it is kept.
- Fig. 1 misses Danish catchments. Either add the catchments to the figure or leave CAMELS-DK out of the study.
- L198: define the minimum mountain cover for a catchment to be considered a mixed high mountain catchment.
- Fig. 6: mention in the caption what the colors represent
- Figs 3-5: either add the asterisk to the figure or write the final caption sentence in default formatting.
- L402: refer to the supplement for the KGE and RMSE results. I would argue against showing RMSE results, as NSE and RMSE are both squared residual metrics.
- The NSE is never formally defined. Please justify the choice of this metric as Schaefli2007 have shown that NSE values can be inflated in highly seasonal regimes such as glacial and nival regimes.
- L9: when and how?
- L17-18: unclear sentence
- L100: citation is missing, please check other citations as well
- Fig 1 caption: North American glacierized catchments
- L274-294: explain more intuitively for readers unfamiliar with GNNs
- 3.2: Consistently mention which example catchments are shown in which supplementary figure
- L786: future work
- L271: check sentence
- L623: missing point
Garreaud, R., Boisier, J. P., Alvarez-Garreton, C., Christie, D. A., Carrasco-Escaff, T., Vergara, I., Chávez, R. O., Aldunce, P., Camus, P., Suazo-Álvarez, M., Masiokas, M., Castro, G., Muñoz, A., Zambrano-Bigiarini, M., Fuster, R., and Godoy, L.: Hyperdroughts in central Chile: drivers, impacts, and projections, Hydrol. Earth Syst. Sci., 29, 5347–5369, https://doi.org/10.5194/hess-29-5347-2025, 2025.
Gauch, M., Kratzert, F., Klotz, D., Nearing, G., Cohen, D., and Gilon, O.: How to deal w___ missing input data, Hydrol. Earth Syst. Sci., 29, 6221–6235, https://doi.org/10.5194/hess-29-6221-2025, 2025.
Hogan, D. and Lundquist, J. D.: Recent Upper Colorado River Streamflow Declines Driven by Loss of Spring Precipitation, Geophys. Res. Lett., 51, https://doi.org/10.1029/2024gl109826, 2024.
Lundquist, J. D., Abel, M., Currier, W. R., and Jackson, D. L.: Snow what? Comparing snow products to reveal what works, where, and why, https://doi.org/10.22541/essoar.177100538.80471917/v1, 2026.
Mudryk, L., Mortimer, C., Derksen, C., Chereque, A. E., and Kushner, P.: Benchmarking of snow water equivalent (SWE) products based on outcomes of the SnowPEx+ Intercomparison Project, Cryosphere, 19, 201–218, https://doi.org/10.5194/tc-19-201-2025, 2024.
Wiersma, P., Aerts, J., Zekollari, H., Hrachowitz, M., Drost, N., Huss, M., Sutanudjaja, E. H., and Hut, R.: Coupling a global glacier model to a global hydrological model prevents underestimation of glacier runoff, Hydrol Earth Syst Sc, 26, 5971–5986, https://doi.org/10.5194/hess-26-5971-2022, 2022.
Citation: https://doi.org/10.5194/egusphere-2026-3798-RC3 - The manuscript is currently too long for the message it aims to convey. Below are several suggestions to shorten it:
Viewed
| HTML | XML | Total | Supplement | BibTeX | EndNote | |
|---|---|---|---|---|---|---|
| 122 | 55 | 20 | 197 | 39 | 10 | 13 |
- HTML: 122
- PDF: 55
- XML: 20
- Total: 197
- Supplement: 39
- BibTeX: 10
- EndNote: 13
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
Dear authors,
I think the study provides an interesting comparison of generalizability capability of three different machine learning architectures when making predictions in a climate that the models have either not seen during training, or have only received limited exposure. The result that machine learning model performance worsens if you remove catchments that represent the target climate is quite obvious. However, I’m strongly of the opinion that even obvious things ’that everybody knows’ should be tested rigorously and published, to create a sturdier foundation for further science. Additionally, the different generalization capability of the models under extreme target climate data scarcity is non-trivial knowledge. The findings could be especially relevant, if climate change creates climatic conditions that have never been seen anywhere during the historic record.
However, I think that the manuscript itself felt more like a draft than a finished article, and it still needs a lot polishing. I also found some of the methodological choices problematic for the robustness of the results. I hope the following comments will help improve the manuscript, because the core results of the article are interesting and worth publishing.
Best regards,
Iiro Seppä
General comments
The main issue I have is that the three models are forced with different meteorological inputs, which invalidates any direct comparison between δHBV and the two other models. Since δHBV is the most limited in inputs it can receive, the DL models should be forced with those inputs, and nothing else. Here are two suggestions on how to possibly fix the issue:
The best option is to rerun the DL experiments with same forcings. This is of course computationally expensive.
Second option is to only compare the models with same inputs. This would be technically easy, but limits the usefulness of results and requires very careful and clear writing and plot organization, and lots of warnings that the models cannot be compared. I would only choose this if rerunning the models is unattainable due to computational limitations
A worrying amount of citations where I wanted to check the original source did not have a corresponding reference. Please carefully check all the citations and ensure that they have a reference.
Please ensure that the sections don’t ’spill over’ into other sections. For example, there were part that belonged to methods in the results, new results in the discussion and introduction in methodology. End of introduction is a bit freer, but other than that the structure should be the following: 1. introducing the important literature and explaining what is the research gap. 2. Data and methods (methods should only contain things you did, often with brief explanation why with a citation, but no other literature review). 3. results (don’t include data description here, it goes to section 2.) 4, discussion. Don’t introduce new results here (computation time)
This study talks a lot about regional generalization. However, it only studied climatic generalization, so any regional claims should be limited to situations where you are citing literature that actually studied regional generalization.
The methods and results should be thoroughly revised to present the fifth experiment along with the rest. The fifth experiment seems currently completely separated from the rest without any good reason. In my opinion it has also been given too much space since its result could be summarized in a few sentences.
There are a lot of sections where the text feels like the article starts looking at something that is a diversion from the main focus. I think the most interesting part of the research can be summarized by figure 3 and table 5 (experiment 5 should be included in those), and I think that everything in the article should focus strongly to support that. Everything in the results or methods that is not relevant for the point of figure 3 should be moved to appendix or supplement in my opinion (for example performance considerations).
Specific comments and technical corrections are highlighted and commented to the attached pdf. If specific comments are clearly related to a general comment, I’m fine with one response collecting these comments together.
No need to respond to technical corrections if you agree with them.