Generalization of Deep Learning Models to Ungauged Glacierized Basins: Evidence from Alpine, Patagonian, and North American Catchments
Abstract. Glacierized high-mountain basins supply water to approximately two billion people yet remain among the most data-scarce hydrologic regions globally, making truly ungauged streamflow prediction a critical challenge. Deep learning (DL) offers a promising alternative to traditional regionalization, but fundamental questions remain about when and why DL models generalize to a target domain that is not merely ungauged but hydro-climatically distinct from the training data. We address two questions: under what training data do DL models generalize reliably to completely ungauged glacierized basins? And how does model architecture, including physics-informed DL, modulate sensitivity to these conditions? We systematically evaluate three architectures — Long Short-Term Memory networks (LSTM), Graph Neural Networks (GNN), and differentiable HBV (δHBV) — across four experiments that control for training dataset size, hydroclimatic representativeness, and inclusion of basins with glaciers, using 2,845 basins from the Caravan global dataset with 283 target glacierized basins. We perform 100-trial repeated K-fold cross-validation by holding out glacierized basins as test basins strictly in space and time. Hydroclimatic representativeness of training data- the degree to which training basins cover the target glacierized regime consistently dominates both training data size and architecture choice as the primary determinant of generalization skill. Including glacierized catchments in training provides the strongest representativeness signal, with all three architectures achieving median NSE between 0.66 and 0.71. When glacierized catchments are excluded, LSTM median NSE falls to −1.43 in the most dissimilar partition; larger dataset size only partially improves skill (median NSE −0.96), confirming that dataset size cannot substitute for representativeness. Non-glacierized mountain catchments partially improve skill, demonstrating that partial hydroclimatic representativeness – through inclusion of non-glacierized mountain basins – contributes to model performance in glacierized basins. Architecture differences are secondary: δHBV and GNN show greater resilience under data scarcity due to structural constraints, but no architecture compensates for lack of hydroclimatic representativeness in training data. These findings reframe model selection for ungauged glacierized basins, highlighting the importance of representative training data and the potential limits of “out of sample in landscape” performance of DL models, specifically for DL deployment in climate impact assessments of high-mountain water towers.
Dear authors,
I think the study provides an interesting comparison of generalizability capability of three different machine learning architectures when making predictions in a climate that the models have either not seen during training, or have only received limited exposure. The result that machine learning model performance worsens if you remove catchments that represent the target climate is quite obvious. However, I’m strongly of the opinion that even obvious things ’that everybody knows’ should be tested rigorously and published, to create a sturdier foundation for further science. Additionally, the different generalization capability of the models under extreme target climate data scarcity is non-trivial knowledge. The findings could be especially relevant, if climate change creates climatic conditions that have never been seen anywhere during the historic record.
However, I think that the manuscript itself felt more like a draft than a finished article, and it still needs a lot polishing. I also found some of the methodological choices problematic for the robustness of the results. I hope the following comments will help improve the manuscript, because the core results of the article are interesting and worth publishing.
Best regards,
Iiro Seppä
General comments
The main issue I have is that the three models are forced with different meteorological inputs, which invalidates any direct comparison between δHBV and the two other models. Since δHBV is the most limited in inputs it can receive, the DL models should be forced with those inputs, and nothing else. Here are two suggestions on how to possibly fix the issue:
The best option is to rerun the DL experiments with same forcings. This is of course computationally expensive.
Second option is to only compare the models with same inputs. This would be technically easy, but limits the usefulness of results and requires very careful and clear writing and plot organization, and lots of warnings that the models cannot be compared. I would only choose this if rerunning the models is unattainable due to computational limitations
A worrying amount of citations where I wanted to check the original source did not have a corresponding reference. Please carefully check all the citations and ensure that they have a reference.
Please ensure that the sections don’t ’spill over’ into other sections. For example, there were part that belonged to methods in the results, new results in the discussion and introduction in methodology. End of introduction is a bit freer, but other than that the structure should be the following: 1. introducing the important literature and explaining what is the research gap. 2. Data and methods (methods should only contain things you did, often with brief explanation why with a citation, but no other literature review). 3. results (don’t include data description here, it goes to section 2.) 4, discussion. Don’t introduce new results here (computation time)
This study talks a lot about regional generalization. However, it only studied climatic generalization, so any regional claims should be limited to situations where you are citing literature that actually studied regional generalization.
The methods and results should be thoroughly revised to present the fifth experiment along with the rest. The fifth experiment seems currently completely separated from the rest without any good reason. In my opinion it has also been given too much space since its result could be summarized in a few sentences.
There are a lot of sections where the text feels like the article starts looking at something that is a diversion from the main focus. I think the most interesting part of the research can be summarized by figure 3 and table 5 (experiment 5 should be included in those), and I think that everything in the article should focus strongly to support that. Everything in the results or methods that is not relevant for the point of figure 3 should be moved to appendix or supplement in my opinion (for example performance considerations).
Specific comments and technical corrections are highlighted and commented to the attached pdf. If specific comments are clearly related to a general comment, I’m fine with one response collecting these comments together.
No need to respond to technical corrections if you agree with them.