the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
Technical note: Regional fine-tuning of LSTMs for improved streamflow predictions in ungauged catchments
Abstract. Predicting streamflow in ungauged basins (PUB) remains a central challenge in hydrology. Long short-term memory (LSTM) networks trained on large samples of catchments ("global" LSTMs) have emerged as a state-of-the-art approach for PUB, outperforming conceptual rainfall–runoff models with traditional regionalisation approaches. However, global LSTMs are spatially agnostic, relying solely on static catchment attributes to differentiate regional hydrological behaviour. This study introduces Regionalised Fine-Tuning (ReFT), a strategy that adapts a pretrained global LSTM to the region surrounding each ungauged target catchment by fine-tuning on a spatially weighted set of donor catchments using an inverse-distance weighting scheme. ReFT is evaluated on 218 catchments from the CAMELS-AUS dataset under a spatial out-of-sample cross-validation framework, comparing two fine-tuning configurations: updating all model parameters versus updating only the prediction head while keeping the recurrent backbone frozen. ReFT improves Nash–Sutcliffe Efficiency relative to the base global LSTM in more than 66 % of catchments, with the largest gains occurring for catchments of moderate baseline performance. The ReFT framework combines the broad process generalisation of large-sample deep learning with the local specificity of regional adaptation, providing an efficient route to improved streamflow predictions in data-sparse regions.
- Preprint
(966 KB) - Metadata XML
- BibTeX
- EndNote
Status: final response (author comments only)
-
RC1: 'Comment on egusphere-2026-2950', Anonymous Referee #1, 16 Jul 2026
-
AC1: 'Reply on RC1', Ashkan Shokri, 01 Oct 2026
We thank the reviewer for the careful and constructive review. We have addressed each comment below and describe the corresponding revisions to the manuscript.
Summary
Comment:
This technical note introduces Regionalised Fine-Tuning (ReFT), a method for prediction in ungauged basins (PUB) that fine-tunes a pretrained continental-scale LSTM separately for each ungauged catchment using data from nearby gauged catchments, weighted inversely by distance. Evaluated on 218 CAMELS-AUS catchments under spatial out-of-sample cross-validation, ReFT is compared in two configurations (full-parameter and head-only fine-tuning) against the base LSTM, a regionalised GR4J, and the AWRA-L model. ReFT improves NSE over the base LSTM in more than 66% of catchments, with the largest gains where the base model already performed moderately well, and head-only fine-tuning outperforms full-parameter fine-tuning.
The idea of ReFT is simple and of interest to the researchers who work on the problem of PUB. However, the manuscript is largely lacking in methodological detail and references to relevant literature. There are very few citations to support the literature in the paper; the concepts and technical details of the method are not introduced; some key methodological choices are justified only by results not shown in the paper; and there is only a single evaluation metric. I therefore have some major and specific comments on the manuscript.
Response:
We thank the referee for the careful and constructive review. We agree that the manuscript will benefit from additional methodological detail, stronger citation of claims, and a broader evaluation. We address each major and specific comment below and will revise the manuscript accordingly.
Major comments
RC1-M1
Comment:
Insufficient citation of claims throughout the manuscript. While reading, I noticed quite a few statements that seemed like they should have a reference attached but didn't. A few examples:
- In L38-40, the manuscript says global LSTMs treat the training set as a "homogeneous pool" and that regional differentiation "relies exclusively" on static attributes. This reads like an interpretation, so I think it would help to either cite something that supports this view or soften the wording a bit.
- In the discussion (L195-196), the authors state that spatial distance captures "large-scale gradients in climate and landscape properties," and later (L198-199) that spatial proximity captures unobserved geological or drainage controls. But I couldn't find any reference to the regionalization literature.
- L214-216 states that the structure of conceptual models has a "stabilising effect" in basins with poor predictability. This is an interesting hypothesis, but it is not supported by citation or by evidence shown in the paper itself.
I'd suggest going through the manuscript claim by claim and checking whether each one is either supported by a reference or clearly framed as the authors' own interpretation.
Response:
We thank the referee for highlighting this important issue and fully agree with the suggestion. We will carefully review the manuscript claim by claim and revise as follows:
- L38-40: We will soften the wording to: “Despite their success, global LSTMs are typically trained as a single model on a pooled multi-catchment dataset with shared parameters. In this setting, differentiation among catchments is primarily provided by static catchment attributes supplied as inputs (Kratzert et al., 2019), rather than by an explicit spatial regionalisation mechanism.”
- L195-196 / L198-199: We will revise this discussion and cite regionalisation studies that support proximity-based donor transfer, including inverse-distance weighting (Qi et al., 2022; Tong et al., 2022; Pool et al., 2021). For the point that proximity can add information beyond the static attributes already supplied to the LSTM, we will cite Merz and Blöschl (2004) and Merz et al. (2020), and frame this as a surrogate for unmeasured or incompletely represented controls rather than as evidence for specific omitted variables.
- L214-216: We agree that the “stabilising effect” is an interpretation rather than a result demonstrated in the paper. The observation we report is that the regionalised GR4J model outperforms the LSTM-based approaches in the lower tail of the performance distribution. We will revise the text to state this observation clearly and frame the explanation in terms of structural constraints as a possible interpretation of that pattern, not as an established general result. We will revise manuscript as follows: “In these difficult basins the regionalised GR4J model occasionally outperforms the LSTM-based approaches, including the ReFT-LSTMs. One possible interpretation is that, where hydrological signals are weak or strongly non-linear, the structural constraints of a conceptual model limit unphysical extrapolation relative to a more flexible LSTM.”
RC1-M2
Comment:
Technical concepts are used without introduction or definition. Several methods central to the paper are invoked as if the reader already knows them:
- Inverse distance weighting (Sect. 2.3) is presented without any background, motivation from prior literature, or citation.
- The smooth-joint NSE loss (L79) is named but not defined; a one-line equation or at least a verbal description with the citation is needed.
- The global LSTM training protocol is similarly underspecified: training/validation/test time periods, input normalisation, number of epochs, learning-rate schedule, and how the 10 seeds enter the fine-tuning stage (is ReFT applied to each seed independently?) are all missing.
Response:
We thank the referee for identifying the gaps. We will expand the Methods section as follows:
- Inverse distance weighting: We will briefly introduce IDW as a standard proximity-based regionalisation scheme, cite relevant literature (e.g. Shepard, 1968; and hydrological regionalisation studies including Shokri et al., 2026), and state that we use spatial distance to emphasise local hydrological information during fine-tuning.
- Smooth-joint NSE loss: We will add the basin-averaged NSE* equation of Kratzert et al. (2019), including the variance-normalisation rationale.
- Global LSTM training protocol: We will add a dedicated subsection in Methods specifying the training/evaluation periods, input normalisation, number of epochs, learning-rate schedule, and how the 10 seeds enter fine-tuning (ReFT applied independently to each seed; metrics reported as the median across seeds).
RC1-M3
Comment:
Evaluation relies on a single metric (NSE). All results, figures, and conclusions rest exclusively on NSE, which is well known to emphasise high flows and to be sensitive to flow variability (e.g., Gupta et al., 2009; Knoben et al., 2019; Clark et al.). Whether ReFT's improvements persist under KGE or signature-based measures (low-flow and high-flow biases) is an open question with direct practical relevance.
Response:
We thank the reviewer for this valuable point. In the revised manuscript we will report KGE for the comparisons (base LSTM vs ReFT, and head-only vs full-parameter) in the appendix and briefly note in the main text that the NSE-based conclusions hold under KGE.
RC1-M4
Comment:
The magnitude and significance of the headline result are not quantified. "Improves NSE in more than 66% of catchments" says nothing about how much. Please report the distribution of ΔNSE and other metrics which you will report.
Response:
In the revised manuscript we will include exceedance-probability plots of ΔNSE and ΔKGE (ReFT minus the base global LSTM) so that the distribution of improvements and degradations is shown explicitly across catchments.
Specific comments RC1-S1
Comment: Title / Abstract / Conclusion: Both "regionalisation" and "regionalization" also "generalization"/"generalisation" appear; please settle on one spelling convention consistently.
Response: We thank the referee for noting this inconsistency. We will adopt American English consistently throughout the manuscript.
RC1-S2
Comment: L66: state which 4 catchments were removed and the record-length threshold used, for reproducibility.
Response: We will add this information to the revised manuscript. The four catchments excluded were: A5040517, 803003, 219001 and 235205. Two were excluded because of incomplete streamflow coverage over the 1975–2014 study period (A5040517, availability 70%; 803003, availability 59%). The other two catchments were excluded because their streamflow regimes are substantially altered by water-resource infrastructure: 219001 (Rutherford Creek at Brown Mountain; regulated flows) and 235205 (Arkins Creek West Branch at Wyelangta; municipal water-supply diversion weirs).
RC1-S3
Comment:
Table 1: Clarify why this particular attribute subset was chosen relative to the full CAMELS-AUS attribute set (Fowler et al., 2021), and whether attributes were standardised.
Response: We thank the referee for this helpful suggestion. The attributes in Table 1 were selected so that they can be readily calculated for ungauged basins. In particular, we excluded all attributes that depend on streamflow observations (or other quantities that would not be available at an ungauged site). From the remaining candidates, a subset of climatic and geomorphological descriptors was then chosen by trial and error to keep the number of attributes minimal while not affecting LSTM performance. The same attribute set is used for the continental LSTM that serves as the ReFT base model (Shokri et al., 2026), ensuring consistency between the two studies.
Regarding standardisation, the static (and quasi-static) attributes in Table 1 were not standardised. These explanations will be added to the revised manuscript.
RC1-S4
Comment:
L75-76: hidden size 256 "follows Kratzert et al. (2018)"; They explored several configurations; state explicitly which configuration and report the dropout and other hyperparameter settings.
Response:
We thank the referee for this comment. We will add the architecture and training hyperparameters to the revised manuscript.
RC1-S5
Comment:
L129-130: clarify whether the 10 seeds refer to the global pretraining only or whether the fine-tuning is also repeated per seed, and whether the median is taken per catchment.
Response:
We thank the referee for pointing out this ambiguity. For each of the 10 random seeds, we first train the global LSTM and then apply ReFT fine-tuning to that pretrained model. This global-training-plus-fine-tuning procedure is repeated independently for all 10 seeds. For each catchment, performance metrics are then summarised as the median across the 10 seeds. We will clarify this procedure explicitly in Sect. 2.5 of the revised manuscript.
RC1-S6
Comment: L219-221: the computational-cost discussion would benefit from a concrete number (e.g., minutes per target catchment on given hardware).
Response: All runs were performed on an NVIDIA H100 GPU. Training each global LSTM took less than one day. ReFT uses only the nearest 5% of donors (about 9 catchments under SooS), so fine-tuning each target catchment is much faster, on the order of a few minutes. We will report accurate wall-clock times in the revised manuscript.
RC1-S7
Comment:
Appendix: References cited in Appendix A are missing from the reference list. "Frost and Shokri (2021)" and "Frost et al. (2021)" are cited (L128, L302–303) but do not appear in the References. Please add them and check the whole list for completeness.
Response:
We thank the referee for catching this omission. We will add the missing references and thoroughly audit all in-text citations against the reference list.
References:
Kratzert, F., Klotz, D., Shalev, G., Klambauer, G., Hochreiter, S., and Nearing, G.: Towards learning universal, regional, and local hydrological behaviors via machine learning applied to large-sample datasets, Hydrol. Earth Syst. Sci., 23, 5089–5110, https://doi.org/10.5194/hess-23-5089-2019, 2019.
Merz, R., and Blöschl, G.: Regionalisation of catchment model parameters, J. Hydrol., 287, 95–123, https://doi.org/10.1016/j.jhydrol.2003.09.028, 2004.
Merz, R., Tarasova, L., and Basso, S.: Parameter’s controls of distributed catchment models—How much information is in conventional catchment descriptors?, Water Resour. Res., 56, e2019WR026008, https://doi.org/10.1029/2019WR026008, 2020.
Pool, S., Vis, M., and Seibert, J.: Regionalization for ungauged catchments — Lessons learned from a comparative large-sample study, Water Resour. Res., 57, e2021WR030437, https://doi.org/10.1029/2021WR030437, 2021.
Qi, W., Chen, J., Li, L., Xu, C.-Y., Xiang, Y., Zhang, S., and Wang, H.: Impact of the number of donor catchments and the efficiency threshold on regionalization performance of hydrological models, J. Hydrol., 601, 126680, https://doi.org/10.1016/j.jhydrol.2021.126680, 2022.
Shepard, D.: A two-dimensional interpolation function for irregularly-spaced data, Proceedings of the 1968 23rd ACM National Conference, 517–524, https://doi.org/10.1145/800186.810616, 1968.
Shokri, A., Bennett, J. C., Robertson, D. E., Perraud, J.-M., Frost, A. J., and Lehmann, E. A.: Better continental-scale streamflow predictions for Australia: LSTM as a land surface model post-processor and standalone hydrological model, Hydrol. Earth Syst. Sci., 30, 757–777, https://doi.org/10.5194/hess-30-757-2026, 2026.
Tong, R., Parajka, J., Széles, B., Greimeister-Pfeil, I., Vreugdenhil, M., Komma, J., Valent, P., and Blöschl, G.: The value of satellite soil moisture and snow cover data for the transfer of hydrological model parameters to ungauged sites, Hydrol. Earth Syst. Sci., 26, 1779–1801, https://doi.org/10.5194/hess-26-1779-2022, 2022.
Citation: https://doi.org/10.5194/egusphere-2026-2950-AC1
-
AC1: 'Reply on RC1', Ashkan Shokri, 01 Oct 2026
-
RC2: 'Comment on egusphere-2026-2950', John Quilty, 19 Aug 2026
Please see attached for my comments.
Best regards,
John Quilty
-
AC2: 'Reply on RC2', Ashkan Shokri, 01 Oct 2026
We thank the reviewer for the careful and constructive review. We have addressed each comment below and describe the corresponding revisions to the manuscript.
Assessment
Comment:
This paper proposes Regionalized Fine-Tuning (ReFT) as a method for fine-tuning a global LSTM to a target catchment, using donor catchments selected by inverse-distance weighting. ReFT uses a weighted loss function based on this inverse-distance weighting, selecting a fraction (5%) of the “closest” catchments as donors. The authors test their approach using 218 catchments from CAMELS-AUS against a globally trained LSTM and two other conceptual benchmarks. The authors consider two fine-tuning scenarios: all parameters and only those from the output head.
The paper would benefit from, among other things: coverage of recent literature on the same topic; additional elaboration to improve reproducibility; justification of fine-tuning strategies (all parameters vs. output head only). Further, some results require additional explanation (e.g., polarizing results of full parameter fine-tuning).
The authors discuss the spatial concentration of catchments in CAMELS-AUS and how the results may not show the generality of the proposed method to other scenarios where catchments are more spatially distributed (e.g., CAMELS-USA). This calls into question whether the adopted dataset is the right one for evaluating the proposed method. Given this is a technical note, if the editor agrees, this may not be seen as a major hindrance to the paper’s suitability for this journal.
Overall, I think the paper has merit as a technical note and l look forward to seeing a revised version of the paper.
Response:
We thank the reviewer for the thorough and constructive review. We address each major and minor item below and will revise the manuscript accordingly.
Major items RC2-M1
Comment:
The authors should ensure they include recent papers in hydrology on fine-tuning DL models, contrasting earlier methods with their approach. Why is the presented approach necessary, what problems with existing approaches does it overcome?
Response:
In the revision we will add a brief discussion of recent fine-tuning and local adaptation of large-sample hydrological models. We will position ReFT relative to global LSTMs, classical donor-based regionalisation of conceptual models, and gauged-basin fine-tuning, and clarify the gap it addresses; in short “adapting a pretrained global LSTM using neighbouring gauged donors via a spatially weighted loss, without observations at the target”.
RC2-M2
Comment:
LSTM architecture: is this the right architecture for your dataset? I understand the adopted architecture is from earlier work and was selected by the original authors after much experimentation on CAMELS USA (531 catchments). However, that does not mean the architecture is necessarily the best suited for CAMELS-AUS (218 catchments). What experimentation was done to confirm the suitability of this architecture for CAMELS-AUS compared to potentially “lighter” architectures (fewer hidden states) or shorter sequences (e.g., 30, 180)?
Response:
We thank the referee for this point. We agree that an architecture selected for CAMELS-USA is not automatically optimal for CAMELS-AUS. For this study we used the LSTM configuration from Shokri et al. (2026), which was applied and tested on CAMELS-AUS rather than transferred unchanged from the CAMELS-USA literature. That work includes sensitivity analyses for hidden-state size and sequence length (e.g. Figure 8 for sequence length), which supported the settings used here (hidden size 256; sequence length 365 days). We will briefly note this in the revised manuscript. A broader architecture search is outside the scope of this technical note, which focuses on ReFT.
RC2-M3
Comment: What optimizer was used? How many epochs was the model trained for? More details on model development (global LSTM and fine-tuned models) would improve reproducibility.
Response: We agree that additional training details will improve reproducibility. Both the global LSTM and ReFT fine-tuning used the Adam optimiser. The global model was trained for 15 epochs and ReFT for 5 epochs. We will expand the Methods to report these settings together with the learning-rate schedule, batch size, and related hyperparameters.
RC2-M4
Comment:
Why use spatial separation alone instead of the “closeness” (e.g., cosine similarity or another appropriate distance metric) to a group of static features (which could include lat-lon-elevation)? For example, you could use the cosine similarity based on the global LSTM’s embedding vector (hidden state) to estimate closeness with a target catchment and its donors (e.g., Jahangir et al., 2026; https://doi.org/10.1016/j.envsoft.2026.106978). This could be interesting to check as it may validate your distance-based approach.
Response:
We thank the reviewer for this insightful suggestion and for drawing our attention to Jahangir et al. (2026). Attribute- or embedding-based measures of catchment similarity are a valuable and interesting extension of the present work.
For this technical note, however, we deliberately restricted donor similarity to spatial proximity. First, geographic distance is straightforward to compute and requires no additional choices about attribute selection, scaling, or embedding construction. Second, in a hydrological setting, spatial proximity is expected to capture much of, if not all, the information contained in climatic and landscape attributes because these controls typically exhibit spatial structure: catchment properties and hydroclimatic regimes are spatially correlated and often vary as regional gradients rather than as spatially independent fields. We will strengthen this rationale in the Discussion, cite Jahangir et al. (2026), and acknowledge that more complex similarity measures (e.g. attribute- or embedding-based donor weighting) can be investigated as future work.
RC2-M5
Comment:
How many epochs (or iterations) were used to train catchment-specific models using the ReFT loss? How did you check/evaluate overfitting?
Response:
We thank the reviewer for raising this point. ReFT was trained for 5 epochs. Key training hyperparameters for the global LSTM and ReFT are summarised below (Table r1); these will also be included in the revised Methods.
Table r1. Training hyperparameters for the global LSTM and ReFT.
Parameter
Global LSTM
ReFT
Optimizer
Adam
Adam
Learning rate
7×10⁻⁵
7×10⁻⁵
Epochs
15
5
Batch size
512
512
Dropout
0
0
Weight decay
1×10⁻⁴
1×10⁻⁴
Gradient clipping
0.5
0.5
IDW α / donor pool
—
α = 2; nearest 5% (~9 donors)
Fine-tuning modes
—
Full-parameter; head-only
Regarding overfitting, we used the following measures:
- Weight decay was applied during training. This penalises large parameter magnitudes and thereby helps limit overfitting, although it does not eliminate the risk on its own.
- In the head-only configuration, the recurrent backbone was frozen so that only the prediction head was updated, limiting the number of free parameters.
- More importantly, generalisation was assessed under spatial out-of-sample (SooS) cross-validation: ReFT never sees streamflow at the target catchment, so performance relative to the global LSTM on held-out targets is our primary check against overfitting to the donor pool.
RC2-M6
Comment:
Justify the fine-tuning configurations. Why these instead of other possibilities? Are there precedents that informed this selection? If so, they should be cited.
Response:
We will add additional justification of the fine-tuning configurations to the manuscript, as follows. When adapting a pretrained LSTM, fine-tuning strategies can be ordered by how much of the network is freezed. At one end of this spectrum the pretrained network is left entirely unchanged (no fine-tuning); at the other end every parameter is updated (full-parameter fine-tuning). Between these extremes lie partial-update strategies, of which prediction-head-only fine-tuning is among the most constrained: the recurrent backbone remains frozen and only the final mapping from latent states to streamflow is adapted.
We therefore compared:
- No fine-tuning, corresponding to the pretrained global LSTM applied unchanged to ungauged catchments (our base LSTM benchmark).
- Full-parameter fine-tuning (fully adaptive end), which updates all weights and allows both the temporal dynamics and the output mapping to change. This follows the catchment-level fine-tuning approach used for continental LSTMs (e.g. Shokri et al., 2026), applied here with a regional donor objective.
- Prediction-head-only fine-tuning (a highly constrained partial update), which freezes the recurrent backbone and updates only the final dense layer. The rationale is that the global LSTM has already learned useful rainfall–runoff temporal features, while regional differences may be expressed largely in the mapping from those features to streamflow. Restricting updates to the head also limits the number of free parameters when fine-tuning on a small donor set.
Comparing head-only and full-parameter ReFT tests whether regional adaptation for PUB requires updating the full network, or whether adapting only the streamflow head is sufficient. Other intermediate schemes were outside the scope of this technical note.
RC2-M7
Comment:
L115-116: Please describe the notion of nested catchments in more detail along with the issue of data leakage.
Response:
We thank the reviewer for giving us the opportunity to expand on this point. Nested catchments are gauges whose drainage areas are nested within the same river system (for example, an upstream headwater gauge and a downstream gauge that includes that headwater area). Streamflow at these sites is hydrologically dependent rather than independent: information observed at one gauge is partly shared with the other through the common upstream contributing area.
If nested gauges were split across training and validation folds under SooS cross-validation, the model could learn from a hydrologically coupled neighbour while being evaluated on the held-out site. That would inflate apparent PUB skill and constitute spatial data leakage. To avoid this, nested catchments were identified and blocked so that all members of a nested group remain in the same fold: either entirely in training or entirely in validation.
This explanation will be added to the revised manuscript (Sect. 2.5).
RC2-M8
Comment:
Please provide more details on the polarizing results when using ReFT to fine-tune all LSTM parameters.
Response:
Full-parameter ReFT produces both more large improvements and more large degradations relative to the base LSTM (ΔNSE beyond ±0.1) than head-only fine-tuning, with the largest disagreements between the two strategies concentrated in lower-performing catchments (Figs. 2–4).
In the revised manuscript we will state these quantities explicitly in the text and briefly restate the interpretation that full-parameter updates on a small donor pool allow stronger regional adaptation but also greater performance variability across catchments. We speculate the greater range of results when fine-tuning all LSTM parameters is a matter of greater degrees of freedom in the model leading to over-fitting, cf head-only fine-tuning.
RC2-M9
Comment: As you discuss, the spatial concentration of catchments in CAMELS-AUS may limit the generality of the findings. Why not apply the method to CAMELS-USA instead?
Response: We will clarify why CAMELS-AUS was used (consistency with our related continental LSTM / regionalisation work and benchmarks; Australia as a demanding and practically relevant PUB setting) and state more explicitly that the relative benefit of ReFT may differ in densely gauged regions such as CAMELS-USA. Multi-continent testing will be noted as future work. Given the technical-note format, we hope the editor agrees that CAMELS-AUS alone is acceptable provided these limitations are clearly acknowledged.
Minor items RC2-m1
Comment:
Check verb tense (e.g., are vs. were)
Response:
We thank the referee for noting this. We will carefully proofread the manuscript for consistent verb tense.
RC2-m2
Comment:
L147: check period placement.
Response:
We thank the referee for catching this. The sentence appears to have been corrupted by merging figure caption text into the body; we will correct the wording and punctuation at L147.
Citation: https://doi.org/10.5194/egusphere-2026-2950-AC2
-
AC2: 'Reply on RC2', Ashkan Shokri, 01 Oct 2026
Viewed
| HTML | XML | Total | BibTeX | EndNote | |
|---|---|---|---|---|---|
| 228 | 142 | 21 | 391 | 11 | 12 |
- HTML: 228
- PDF: 142
- XML: 21
- Total: 391
- BibTeX: 11
- EndNote: 12
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
Technical note: Regional fine-tuning of LSTMs for improved streamflow predictions in ungauged catchments
Summary
This technical note introduces Regionalised Fine-Tuning (ReFT), a method for prediction in ungauged basins (PUB) that fine-tunes a pretrained continental-scale LSTM separately for each ungauged catchment using data from nearby gauged catchments, weighted inversely by distance. Evaluated on 218 CAMELS-AUS catchments under spatial out-of-sample cross-validation, ReFT is compared in two configurations (full-parameter and head-only fine-tuning) against the base LSTM, a regionalised GR4J, and the AWRA-L model. ReFT improves NSE over the base LSTM in more than 66% of catchments, with the largest gains where the base model already performed moderately well, and head-only fine-tuning outperforms full-parameter fine-tuning.
The idea of ReFT is simple and of interest to the researchers who work on the problem of PUB. However, the manuscript is largely lacking in methodological detail and references to relevant literature. There are very few citations to support the literature in the paper; the concepts and technical details of the method are not introduced; some key methodological choices are justified only by results not shown in the paper; and there is only a single evaluation metric. I therefore have some major and specific comments on the manuscript.
Major comments
Insufficient citation of claims throughout the manuscript. While reading, I noticed quite a few statements that seemed like they should have a reference attached but didn't. A few examples:
I'd suggest going through the manuscript claim by claim and checking whether each one is either supported by a reference or clearly framed as the authors' own interpretation.
Technical concepts are used without introduction or definition. Several methods central to the paper are invoked as if the reader already knows them:
Evaluation relies on a single metric (NSE). All results, figures, and conclusions rest exclusively on NSE, which is well known to emphasise high flows and to be sensitive to flow variability (e.g., Gupta et al., 2009; Knoben et al., 2019; Clark et al.). Whether ReFT's improvements persist under KGE or signature-based measures (low-flow and high-flow biases) is an open question with direct practical relevance.
The magnitude and significance of the headline result are not quantified. "Improves NSE in more than 66% of catchments" says nothing about how much. Please report the distribution of ΔNSE and other metrics which you will report.
Specific comments
References:
Fowler, K. J. A., Zhang, Z., and Hou, X.: CAMELS-AUS v2: updated hydrometeorological time series and landscape attributes for an enlarged set of catchments in Australia, Earth System Science Data, 17, 4079–4095, https://doi.org/10.5194/essd-17-4079-2025, 2025.
Knoben, W. J. M., Freer, J. E., and Woods, R. A.: Technical note: Inherent benchmark or not? Comparing Nash–Sutcliffe and Kling–Gupta efficiency scores, Hydrology and Earth System Sciences, 23, 4323–4331, https://doi.org/10.5194/hess-23-4323-2019, 2019.
Kratzert, F., Klotz, D., Brenner, C., Schulz, K., and Herrnegger, M.: Rainfall–runoff modelling using Long Short-Term Memory (LSTM) networks, Hydrology and Earth System Sciences, 22, 6005–6022, https://doi.org/10.5194/hess-22-6005-2018, 2018.
Gupta, H. V., Kling, H., Yilmaz, K. K. and Martinez, G. F.: Decomposition of the mean squared error and NSE performance criteria: Implications for improving hydrological modelling, J. Hydrol., 377(1-2), 80–91, doi:10.1016/j.jhydrol.2009.08.003, 2009.