the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
A climate similarity-based transfer learning framework using global Caravan dataset for enhancing streamflow prediction in Ouémé River Basin (Benin, West Africa)
Abstract. Reliable streamflow forecasting is a fundamental component of flood risk management. However, the accuracy of such forecasts in data-scarce river basins remains one of the pressing challenges in most African catchments. While Long Short-Term Memory (LSTM) networks have demonstrated their effectiveness in rainfall-runoff modelling, their data-intensive requirements limit their direct applicability in poorly gauged catchments across sub-Saharan Africa. To address this issue, we propose a climate similarity-based transfer learning framework using the global Caravan hydrological dataset to improve local streamflow forecasting in Ouémé River Basin (ORB). Specifically, tropical-climate catchments are first identified from Caravan global dataset using the Köppen-Geiger climate classification. A composite climate similarity index (CI) is constructed from two normalized climate indices: annual mean precipitation and seasonality index. 200 basins were selected and ranked by CI. From the 200 ranked basins, a fixed subset of 50 basins was selected using stratified sampling across CI quartiles, including 30 basins for validation and 20 basins for independent testing. The remaining 150 basins were then used to define three cumulative pre-training subsets of increasing size (50, 100, and 150 basins), ranked by CI. Different LSTM models are pre-trained on each subset and fine-tuned on the ORB streamflow. The developed transfer learning models are evaluated against the LSTM model trained exclusively on local data as a baseline model. Pre-training improves prediction at all five stations, raising the median Kling–Gupta Efficiency from 0.65 for the local model to 0.75 for the best transfer configuration. The smallest, most climatically similar donor subset gives the best fine-tuning performance; adding less similar catchments does not help and slightly degrades it. The benefit is largest at stations with short or event-poor records and marginal at the well-gauged basin outlet, while peak flows remain underestimated across all configurations. These results show that, for data-scarce tropical basins, the climatic similarity of the transfer basins matters more than their number, and they point to similarity-guided donor selection combined with peak-weighted or hybrid formulations as a practical route to operational flood forecasting.
- Preprint
(1733 KB) - Metadata XML
- BibTeX
- EndNote
Status: final response (author comments only)
-
RC1: 'Comment on egusphere-2026-4749', Kai Li, 11 Sep 2026
-
AC1: 'Reply on RC1', Jerome Ahouandjinou, 08 Oct 2026
Dear Dr Li,
Thank you for your careful review and constructive suggestions. Please find attached our point-by-point responses, outlining the corrections and additional analyses we propose for the revised manuscript. These revisions have not yet been completed.
Kind regards,
Jérôme Enagnon Ahouandjinou, on behalf of the co-authors
-
AC1: 'Reply on RC1', Jerome Ahouandjinou, 08 Oct 2026
-
RC2: 'Comment on egusphere-2026-4749', Anonymous Referee #2, 23 Sep 2026
Review of: A climate similarity-based transfer learning framework using global Caravan dataset for enhancing streamflow prediction in Ouémé River Basin (Benin, West Africa)General Assessment
The manuscript addresses an important and timely question: whether transfer learning from large-sample hydrological datasets can improve streamflow prediction in data-limited regions. This is, in my view, one of the most valuable current directions in large-sample hydrology. Global datasets such as Caravan provide an unprecedented opportunity to leverage information from well-monitored regions to improve predictions in poorly gauged basins, and sub-Saharan Africa is precisely where such approaches are most needed and potentially most impactful.
The manuscript is also built around a potentially interesting contribution, namely the use of a climate-based similarity index for donor catchment selection prior to transfer learning. However, in its current form, the study design does not convincingly support the conclusions drawn. My concerns are methodological rather than editorial. They relate to (i) the composition and processing of the data, (ii) the selection and characterization of donor basins, (iii) the formulation and reproducibility of the similarity index, (iv) the specification and evaluation of the modeling framework, and (v) the adequacy of the validation strategy for supporting claims regarding flood prediction and transferability.
For these reasons, I recommend rejection in its current form. Nevertheless, I believe the central idea is valuable and that the dataset assembly effort provides a strong foundation for a future submission if the methodology is substantially revised and better justified.
Major Comments3.1 Data(a) Line 144 states that precipitation, air temperature, evapotranspiration, and streamflow discharge were used as hydroclimate forcing data. It is unclear whether observed discharge is included among the model inputs. If antecedent streamflow observations are used as predictors, then the study evaluates a form of autoregressive streamflow prediction rather than prediction based solely on meteorological forcing and catchment attributes. The manuscript should explicitly state the role of discharge in the input data and clarify how it enters the model.(b) Line 146 refers to "runoff conditions" among the selected static attributes, but I could not identify such variables in Table A1. Likewise, Line 147 refers to "groundwater and soil properties." Groundwater appears to be represented (e.g., gwt_cm_sav), but I could not identify any soil-related variables in Table A1. The description should therefore be corrected.
(c) A more substantive concern is the omission of soil and land-cover attributes from the similarity framework and model inputs. Catchment properties play a central role in controlling how meteorological forcing is transformed into runoff. The current attribute set contains primarily climatic and topographic descriptors, while largely excluding information on soils and land cover.This omission is particularly notable because HydroATLAS provides extensive descriptors within the Soils & Geology and Land Cover categories. Such information could be highly relevant for identifying hydrologically similar donor catchments and for improving transfer-learning performance. The authors should justify this choice or include those to their modeling framework which I believe that the results from larger catchment sample would be closer or better than smaller sample.
(d) Equation (1) may contain an error. The authors should verify the equation and clarify whether it was used during data preparation. If so, the implications of any correction for the reported analyses should be discussed.(e) A comprehensive summary table describing the five ORB stations is needed. At minimum, I would expect drainage area, record length, percentage of missing data within each split, mean annual precipitation, runoff ratio, mean streamflow, dominant soil class, and dominant land-cover class. Such information is necessary for understanding the hydrologic diversity of the target basins.(f) The manuscript states that the ORB attributes were derived by the authors using Google Earth Engine, while the donor attributes originate from Caravan's precomputed datasets. Differences in HydroATLAS versions, polygon processing, or aggregation methods may introduce inconsistencies between donor and target catchments. I recommend performing a PCA (or similar dimensionality-reduction analysis) using the selected attributes for the 200 donor catchments and projecting the five ORB basins onto the same feature space. This would allow readers to assess whether the target catchments are actually represented within the donor data manifold and whether transfer learning is being applied within a meaningful similarity domain.
3.2 Donor Pool and Composite Similarity Index(a) The selection of the 200 donor catchments requires substantially more explanation. Additionally, it is unclear whether these 200 catchments represent the complete tropical subset of Caravan or a subset selected according to additional criteria. If a subset was used, the selection procedure should be fully documented.
(b) The composite similarity index (CI) is potentially an interesting contribution of the manuscript. However, it is currently insufficiently justified.The manuscript assumes that donor selection should be based primarily on climate similarity, but the rationale for this choice is not adequately developed. In contrast, I think it shouldn't be on climatic similarities, as the whole 200 catchments from tropical condition.
Furthermore, annual precipitation and seasonality are used as the principal descriptors, while aridity index, despite being available in the study, is omitted from the similarity metric. Given its relevance to hydrologic behavior, this omission requires justification.
More broadly, the manuscript does not explain why similarity should be defined using climate descriptors alone. Potential alternatives include:
- Soil similarity
- Land-cover similarity
- Catchment area similarity
- Topographic similarity
- Flow-regime similarity
- Multi-attribute similarity metrics
Because donor selection constitutes the central methodological contribution of the manuscript, alternative versions of the similarity index should be evaluated. This would help determine whether performance gains arise specifically from the proposed index or simply from donor selection in general.
(c) I found Equations (2)-(4) difficult to follow and could not reproduce the construction of the composite index from the information provided.The manuscript should provide a complete, reproducible description of the index, including all variables and weighting steps. At present, an independent reader would have difficulty implementing the method.
3.3 Modeling Framework(a) The description of the LSTM models lacks several details that are necessary for reproducibility.While the manuscript reports parameters such as hidden size and dropout rate, it does not clearly specify:
- Sequence length (lookback window)
- Number of LSTM layers
- Number of fine-tuning epochs
- Treatment of missing target observations
The manuscript should also explicitly confirm whether LSTM1-LSTM4 share an identical architecture and differ only in training strategy and transferred information.
(b) Please provide the complete hyperparameter search space used during tuning, along with the optimization procedure employed.(c) Deep-learning results can be sensitive to random initialization. The manuscript reports only a single realization of each model configuration.I recommend evaluating multiple random seeds and reporting performance variability. This would demonstrate whether the observed differences exceed stochastic training variability and would strengthen confidence in the reported conclusions.
(d) The fine-tuning protocol requires additional justification and detail.The manuscript should explain:
- Whether any layers were frozen
- Which layers were updated during fine-tuning
- How learning rates were selected
- Why the chosen strategy was considered appropriate for the size of the target dataset
Without these details, the transfer-learning workflow is difficult to evaluate and reproduce.
(e) A conceptual hydrologic benchmark should be included.Because the manuscript proposes a transfer-learning framework for streamflow simulation, comparison against a conceptual rainfall-runoff model would provide important context for assessing practical value. At present, the evaluation is restricted to machine-learning variants.
3.4 Training and Evaluation PeriodsThe study uses approximately sixteen years of training data (1996-2012), which is not an exceptionally short hydrologic record. Consequently, the manuscript should be careful when framing the problem as one of severe data scarcity.
My larger concern is the testing period. The test period spans only four to five years (2016-2020), which appears limited for supporting conclusions regarding flood prediction and extreme-event behavior because only a small number of major events may be represented.
If the manuscript intends to make claims about flood-relevant performance and transferability under extremes, a longer evaluation period would provide a more robust assessment, I would suggest at least 8 years.
Additional Major Comments4.1 Forecasting versus SimulationThe manuscript repeatedly discusses forecasting and flood early-warning applications. However, the presented experiments appear to be retrospective streamflow simulations driven by observed meteorological inputs. Forecasting and simulation involve fundamentally different information availability and operational constraints. The manuscript should clearly distinguish between these concepts and avoid implying forecasting capability unless actual forecasting experiments are performed.
4.2 Evaluation MetricsThe evaluation focuses primarily on aggregate performance statistics and does not adequately assess high-flow behavior. Given the motivation of flood prediction, I encourage the authors to report metrics specifically targeting high flows and extreme events, such as:
- Peak-flow bias
- High-flow volume bias (FHV)
- Q95 flow errors
- Peak timing errors
- Event-based flood metrics
Such analyses would provide stronger evidence regarding the practical usefulness of the proposed framework for flood-related applications.
Citation: https://doi.org/10.5194/egusphere-2026-4749-RC2 -
AC2: 'Reply on RC2', Jerome Ahouandjinou, 08 Oct 2026
Dear Reviewer,
Thank you for your detailed assessment and methodological recommendations. Please find attached our point-by-point responses, explaining how we propose to address your concerns in the revised manuscript. The proposed corrections and additional experiments have not yet been completed.
Kind regards,
Jérôme Enagnon Ahouandjinouon behalf of the co-authors
Viewed
| HTML | XML | Total | BibTeX | EndNote | |
|---|---|---|---|---|---|
| 339 | 125 | 54 | 518 | 44 | 50 |
- HTML: 339
- PDF: 125
- XML: 54
- Total: 518
- BibTeX: 44
- EndNote: 50
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
This manuscript presents a climate-similarity-based transfer learning framework using the Caravan dataset to improve LSTM-based streamflow prediction in the Ouémé River Basin. The study addresses a relevant problem and has potential practical value for data-scarce regions. However, the current experimental design does not adequately disentangle the effects of donor-set size and climatic similarity, leaving the central conclusion insufficiently supported. The construction and calibration of CI, the fine-tuning implementation, and several equations and references also require clarification or correction. I therefore recommend major revision, with particular attention to methodological reproducibility and the evidence supporting the claimed value of similarity-based donor selection. My specific comments are as follows.
1. The current experimental design cannot disentangle the effects of donor-set size and climatic similarity
The study constructs cumulative donor sets of 50, 100, and 150 catchments based on their CI rankings. As the number of donor catchments increases, the climatic similarity and other characteristics of the donor-set composition also change. Therefore, the current experiments can compare the performance of three specific donor sets but cannot independently attribute performance differences to donor-set size or climatic similarity. The conclusion that “similarity outweighs volume” is thus not sufficiently supported.
I recommend adding, at a minimum, repeated random-selection experiments with a fixed donor-set size. For example, 50 donor catchments could be randomly selected multiple times from the candidate pool after excluding the fixed validation and test catchments, and compared with the 50 catchments selected using CI. The target-basin data, training procedures, and evaluation conditions should remain consistent across these experiments. This comparison is essential for establishing the added value of CI-based selection. Uncertainty arising from donor sampling and random network initialization should be assessed separately, with station-level paired performance differences and performance distributions across repeated experiments reported.
As an additional comparison, the authors could consider donor selection based on predictive performance, as measured by KGE. However, they should clarify whether the selection metric is calculated for the donor catchments themselves or for the target basin, and ensure that the final test data are not used for donor selection.
The authors already acknowledge that the optimal donor-set size may be smaller than 50, and the abstract and conclusions should be consistent with this limitation. At present, the 50-catchment set should be described only as “the best overall configuration among those tested under the current training setup,” rather than as an established optimum or a universal threshold. Sensitivity experiments involving a few smaller donor sets could strengthen the analysis, but they cannot replace comparisons of selection strategies at a fixed donor-set size.
2. Key equations contain errors or inconsistencies in units
The streamflow conversion in Equation (1) is inconsistent with the stated units of catchment area. If SS is expressed in square metres, the conversion factor of 1000 should appear in the numerator when converting volumetric discharge to millimetres per day. If SS is expressed in square kilometres, the existing equation is correct, but the stated area unit must be revised. Please specify the area units used in the actual implementation and confirm that the target-basin and Caravan streamflow data are on a consistent scale.
The Spearman correlation formula in Equation (2) is missing the squared rank-difference term. The standard formula for observations without tied ranks uses d2, not d.
Please clarify whether these issues are limited to the written equations or also affect the actual calculations. If they affect the implementation, the relevant data, CI weights, and donor rankings should be recalculated, and the affected experiments should be repeated.
3. The output-layer fine-tuning strategy and implementation of the local baseline require further explanation
The manuscript states that the LSTM body is frozen during fine-tuning and that only the 256 weights and one bias of the output regression layer are updated. Thus, the approach involves output-layer adaptation rather than retraining the entire LSTM network. Please explain the rationale for this strategy and ensure that the discussion of local adaptation is consistent with the actual scope of parameter updates, distinguishing between “recombining pre-trained features” and “relearning the LSTM’s internal temporal representations.”
To ensure reproducibility and a clear baseline comparison, please specify in the main text or supplementary material whether the five target stations are fine-tuned separately or jointly; whether the local baseline uses the same multi-station training arrangement; the input sequence length; how static attributes are incorporated; how missing data are handled; how inputs and outputs are standardized and how this standardization is handled across training stages; and the rules used to select model checkpoints. Dynamic inputs should be clearly distinguished from prediction targets, particularly because Section 2.2.1 currently also lists “streamflow discharge” as forcing data.
Figure 7 already compares direct transfer, fine-tuned models, and the locally trained baseline. The discussion should distinguish between two types of improvement: gains relative to direct transfer reflect the benefits of local adaptation, whereas gains relative to training from scratch on local data more directly indicate the added value of pre-training.
If the authors intend to generalize the donor-ranking conclusions to fine-tuning strategies more broadly, I recommend a targeted comparison involving different extents of layer freezing. Otherwise, the conclusions should be explicitly limited to the current output-layer-only fine-tuning setup.
4. Several key in-text citations do not match the corresponding reference-list entries and require systematic checking
Based on the references currently listed in the manuscript, at least the following clear mismatches are present.
Song et al. (2025) is cited to support the limited performance of deep-learning hydrological models in reproducing extreme events, but the corresponding reference is a medical study on the diagnosis of cerebral edema.
Yao et al. (2022) is cited to support the reduced effectiveness of transfer learning in headwater catchments with strong groundwater contributions, but the corresponding reference is a review of transfer learning for machinery diagnostics and prognostics.
Yang et al. (2023) is described as a study of soil moisture estimation on the Qinghai–Tibet Plateau, whereas the corresponding reference concerns runoff prediction using a dynamic spatiotemporal graph neural network.
These issues concern the accuracy of the supporting evidence, not merely reference formatting. Please systematically check that the cited studies support the associated claims and revise both the main text and the reference list accordingly. The same paper by Frame et al. is also listed twice as 2022a and 2022b; these duplicate entries should be merged.
5. Remove the duplicated sentence
On page 8, line 210, “as the validation and test set throughout all experiments (Figure 3)” repeats the ending of the preceding sentence and should be deleted.
6. Shorten the general background in the Methods section
The opening of Section 2.4 mainly describes the general advantages and architecture of LSTMs, with limited relevance to the specific model setup used in this study. Please condense material that overlaps with the Introduction and use the space to clarify the model implementation, input organization, and training procedures.
7. Correct the figure references
In Section 3.1, the references to “Figure 5a” and “Figure 5b” in the descriptions of the climate-deviation distributions and weight calibration should be changed to “Figure 4a” and “Figure 4b,” respectively.
8. Use consistent model-stage terminology and clarify the samples represented in Figure 6
Section 2.4 defines LSTM1–3 as the three fine-tuned models, whereas Figure 6 uses the same labels for the pre-trained configurations. Using the same labels across stages is acceptable, but the pre-trained and fine-tuned model states should be clearly distinguished. Figure 6 should also specify that 50, 100, and 150 refer to the numbers of training donor catchments, whereas validation and testing use fixed sets of 30 and 20 catchments, respectively. This clarification is needed to avoid ambiguity about the samples represented by the different boxplots.