Technical note: Regional fine-tuning of LSTMs for improved streamflow predictions in ungauged catchments
Abstract. Predicting streamflow in ungauged basins (PUB) remains a central challenge in hydrology. Long short-term memory (LSTM) networks trained on large samples of catchments ("global" LSTMs) have emerged as a state-of-the-art approach for PUB, outperforming conceptual rainfall–runoff models with traditional regionalisation approaches. However, global LSTMs are spatially agnostic, relying solely on static catchment attributes to differentiate regional hydrological behaviour. This study introduces Regionalised Fine-Tuning (ReFT), a strategy that adapts a pretrained global LSTM to the region surrounding each ungauged target catchment by fine-tuning on a spatially weighted set of donor catchments using an inverse-distance weighting scheme. ReFT is evaluated on 218 catchments from the CAMELS-AUS dataset under a spatial out-of-sample cross-validation framework, comparing two fine-tuning configurations: updating all model parameters versus updating only the prediction head while keeping the recurrent backbone frozen. ReFT improves Nash–Sutcliffe Efficiency relative to the base global LSTM in more than 66 % of catchments, with the largest gains occurring for catchments of moderate baseline performance. The ReFT framework combines the broad process generalisation of large-sample deep learning with the local specificity of regional adaptation, providing an efficient route to improved streamflow predictions in data-sparse regions.
Technical note: Regional fine-tuning of LSTMs for improved streamflow predictions in ungauged catchments
Summary
This technical note introduces Regionalised Fine-Tuning (ReFT), a method for prediction in ungauged basins (PUB) that fine-tunes a pretrained continental-scale LSTM separately for each ungauged catchment using data from nearby gauged catchments, weighted inversely by distance. Evaluated on 218 CAMELS-AUS catchments under spatial out-of-sample cross-validation, ReFT is compared in two configurations (full-parameter and head-only fine-tuning) against the base LSTM, a regionalised GR4J, and the AWRA-L model. ReFT improves NSE over the base LSTM in more than 66% of catchments, with the largest gains where the base model already performed moderately well, and head-only fine-tuning outperforms full-parameter fine-tuning.
The idea of ReFT is simple and of interest to the researchers who work on the problem of PUB. However, the manuscript is largely lacking in methodological detail and references to relevant literature. There are very few citations to support the literature in the paper; the concepts and technical details of the method are not introduced; some key methodological choices are justified only by results not shown in the paper; and there is only a single evaluation metric. I therefore have some major and specific comments on the manuscript.
Major comments
Insufficient citation of claims throughout the manuscript. While reading, I noticed quite a few statements that seemed like they should have a reference attached but didn't. A few examples:
I'd suggest going through the manuscript claim by claim and checking whether each one is either supported by a reference or clearly framed as the authors' own interpretation.
Technical concepts are used without introduction or definition. Several methods central to the paper are invoked as if the reader already knows them:
Evaluation relies on a single metric (NSE). All results, figures, and conclusions rest exclusively on NSE, which is well known to emphasise high flows and to be sensitive to flow variability (e.g., Gupta et al., 2009; Knoben et al., 2019; Clark et al.). Whether ReFT's improvements persist under KGE or signature-based measures (low-flow and high-flow biases) is an open question with direct practical relevance.
The magnitude and significance of the headline result are not quantified. "Improves NSE in more than 66% of catchments" says nothing about how much. Please report the distribution of ΔNSE and other metrics which you will report.
Specific comments
References:
Fowler, K. J. A., Zhang, Z., and Hou, X.: CAMELS-AUS v2: updated hydrometeorological time series and landscape attributes for an enlarged set of catchments in Australia, Earth System Science Data, 17, 4079–4095, https://doi.org/10.5194/essd-17-4079-2025, 2025.
Knoben, W. J. M., Freer, J. E., and Woods, R. A.: Technical note: Inherent benchmark or not? Comparing Nash–Sutcliffe and Kling–Gupta efficiency scores, Hydrology and Earth System Sciences, 23, 4323–4331, https://doi.org/10.5194/hess-23-4323-2019, 2019.
Kratzert, F., Klotz, D., Brenner, C., Schulz, K., and Herrnegger, M.: Rainfall–runoff modelling using Long Short-Term Memory (LSTM) networks, Hydrology and Earth System Sciences, 22, 6005–6022, https://doi.org/10.5194/hess-22-6005-2018, 2018.
Gupta, H. V., Kling, H., Yilmaz, K. K. and Martinez, G. F.: Decomposition of the mean squared error and NSE performance criteria: Implications for improving hydrological modelling, J. Hydrol., 377(1-2), 80–91, doi:10.1016/j.jhydrol.2009.08.003, 2009.