the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
Extracting Interpretable Representation of Catchment Hydrological Processes with Deep Learning
Abstract. Traditional approaches in hydrological modeling typically rely on parameter calibration within prior empirical functional forms and an integrated subsurface bucket. While effective, these fixed mathematical representations can sometimes limit flexibility in capturing complex and localized behaviors. To complement these approaches, we proposed a temporal difference loss regularization (TDLR) framework. Using meteorological and streamflow observations, this framework enables a Long Short-Term Memory (LSTM) network to extract interpretable functional representations of distinct hydrological processes in parallel within a mass-conserving, conceptual bucket architecture. Applications of TDLR across three hillslope catchments (Shihmen, Deji, and Jiji) yielded stable re-extracted function representations under the given structural constraints, offering strong internal consistency between extracted representations of processes. For the integrated subsurface bucket, the LSTM extracted patterns consistent with the transition from saturation-excess to infiltration-excess runoff. Furthermore, results from the Shihmen catchment captured localized variations that plausibly suggest preferential flow dynamics during heavy rainfall – patterns that are often challenging to represent using fixed functional forms. The framework also presents daily variations consistent with transitions from water- to energy-limited regimes in evapotranspiration. Canopy interception exhibited the expected asymptotic plateau, showing high sensitivity to precipitation and its covariance with wind speed. Additionally, the approach captured dynamic routing relationships that align with local morphological and meteorological characteristics. In our study area, the reconstruction accuracy of the extracted functional representations compares favorably with or exceeds both baseline hybrid models and pure black-box LSTMs. By demonstrating that catchment-scale behaviors can be effectively represented as emergent combinations of stable, data-driven relationships, this study provides a complementary perspective for exploring functional representations to advance hydrological modeling.
- Preprint
(3043 KB) - Metadata XML
- BibTeX
- EndNote
Status: final response (author comments only)
- RC1: 'Comment on egusphere-2026-2483', Anonymous Referee #1, 06 Jul 2026
-
RC2: 'Comment on egusphere-2026-2483', Anonymous Referee #2, 13 Jul 2026
This manuscript presents a temporal difference loss regularization framework, TDLR, that trains an LSTM within a conceptual bucket structure to infer catchment-scale hydrological response functions. The topic is interesting and relevant. The manuscript has some strengths: the framework is well motivated and the repeated-training experiment is a useful check on numerical stability. However, I think the current version sometimes interprets model-derived internal variables more strongly than the evidence supports. The results show plausible and stable latent process representations under the chosen model structure, but they do not yet fully establish that these are identifiable or physically unique hydrological process functions.
I therefore recommend major revision.
Major comments
- My main concern is that the identifiability of the inferred process functions is not yet well established. Many internal quantities, including canopy storage change, subsurface storage change, evapotranspiration partitioning, runoff partitioning, and routing parameters, are inferred mainly from streamflow error and the imposed bucket constraints. It is therefore possible that different combinations of internal process responses could lead to similar hydrograph performance. The repeated training with 30 random seeds in Sect. 4.5 is helpful, but I would interpret it mainly as evidence of optimization stability under the same data, architecture, loss function, and model structure. It is not yet strong evidence of physical identifiability. The authors also note in Sect. 4.1 that accuracy alone cannot fully verify identifiability. The paper would be stronger if the authors softened wording around “identified process functions” and added some targeted checks, such as synthetic experiments with known process functions, validation against independent ET/soil moisture/groundwater/baseflow data, or sensitivity tests with alternative bucket structures.
- The residual term L_t needs a more cautious treatment. In Sect. 3.1 and Sect. 4.3, L_t is defined as a mixture of interbasin groundwater flow, measurement error, and structural mismatch. The manuscript also acknowledges that IGF cannot be definitively separated from these other components. Later, however, the residual patterns are interpreted quite strongly as evidence for IGF or preferential subsurface pathways. This interpretation may be plausible, but it is not unique. A systematic residual could also reflect precipitation bias, discharge uncertainty, PET/AET uncertainty, missing processes, or model structural error. I would suggest referring to L_t more neutrally as a water-balance closure residual unless independent groundwater evidence is provided. Even a simple uncertainty analysis for precipitation, discharge, and ET would make this section more convincing.
- Sect. 4.2 and Fig. 5 are central to the paper, but the interpretation needs some tightening. The LSTM uses a 64-day sequence of multiple meteorological variables, while many process interpretations are based on one-dimensional projections against current precipitation. These projected curves are useful summaries, but they are not the full learned functions. This matters because the projected relationships are shaped by the observed covariation among precipitation, wind speed, radiation, humidity, temperature, seasonality, and antecedent wetness. For example, the canopy response in Shihmen is discussed in terms of wind-rainfall interaction, but other correlated factors may also contribute. The authors do not necessarily need to resolve all of this in one paper, but the limitation should be made clearer. Conditional partial dependence, controlled perturbation tests, or two-variable response surfaces would help. If these analyses are not added, the physical interpretation of the projected curves should be phrased more cautiously.
- The evaluation setup is not yet transparent enough. Sect. 3.2 indicates that observed previous-day streamflow is used during training, while simulated previous-day streamflow can be used during reconstruction. Sect. 4.1 then presents reconstruction results using an initial streamflow value. This leaves some uncertainty about whether the reported results are teacher-forced one-step predictions, fully recursive continuous simulations, or something in between. Please provide exact training, validation, and testing periods, clarify warm-up and initialization, and report whether normalization and hyperparameter tuning are based only on training data. It would also be useful to report fully recursive simulation performance separately from any teacher-forced setting. The benchmark comparison in Sect. 4.4 also needs more detail. The performance gap between TDLR and dPL-HBV is large, especially for Jiji, so readers need enough information to judge whether the comparison is fair. The claim would be stronger if the authors described the dPL-HBV setup, parameter ranges, forcing inputs, spin-up, optimization budget, and whether all models use the same data split and evaluation protocol.
- The manuscript sometimes states that TDLR learns process functions “without relying on prior functional forms” and presents the method as a new paradigm for solving the closure problem. I agree that the method avoids some fixed empirical process equations, but it still imposes substantial prior structure: the bucket layout, process ordering, Muskingum routing, parameter bounds, non-negativity constraints, 64-day input window, and time-invariance assumption. A more precise framing would be that TDLR learns flexible process-closure relationships within a prescribed mass-balanced conceptual structure. This is still a meaningful contribution, but it would avoid overstating the method. Similarly, conclusions about the general non-transferability of fixed functional forms should be limited to the three tested catchments unless broader tests are added.
Minor comments
- Please define and use “functional representation,” “projected relationship,” “response function,” and “constitutive relationship” more consistently.
- NSE alone gives a limited view of model behavior. KGE, bias, RMSE, high-flow error, low-flow error, peak timing error, and water-balance error would provide a fuller evaluation.
- Fig. 5 would benefit from uncertainty bands and sample-density information, especially in the high-precipitation tails where the fitted curves may be less well constrained.
- Physical interpretations such as preferential fracture flow, backwater effects, and canopy shaking should be framed as hypotheses unless supported by independent observations.
- The manuscript would benefit from language editing and some tightening, especially where model output, fitted projection, physical interpretation, and independent evidence are discussed together.
Citation: https://doi.org/10.5194/egusphere-2026-2483-RC2
Viewed
| HTML | XML | Total | BibTeX | EndNote | |
|---|---|---|---|---|---|
| 216 | 92 | 19 | 327 | 13 | 13 |
- HTML: 216
- PDF: 92
- XML: 19
- Total: 327
- BibTeX: 13
- EndNote: 13
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
Liu and colleagues provide an interesting piece of work, where they explore the possibility of extracting interpretable representation of catchment hydrological processes using deep learning methods. The topic has been highly explored recently and it is worthy the study, but I have some major issues that need to be addressed before the paper might be considered for publication at HESS.
Major comments:
Kratzert, F., Gauch, M., Klotz, D., and Nearing, G.: HESS Opinions: Never train a Long Short-Term Memory (LSTM) network on a single basin, Hydrol. Earth Syst. Sci., 28, 4187–4201, https://doi.org/10.5194/hess-28-4187-2024, 2024.
Minor comments:
L58-60: Did these authors say that?
L75: perhaps the authors meant: “there is currently” instead of “there was”?
Figure 4: Why use Qtrue and in the legend inside the figure use Qobs? Please stick to one.
L341: Competitive to what?A benchmark? Some NSE literature threshold?
There are many places where some claims, which are backed by literature or by the results themselves, are made, but such references are never exposed. Please make sure that you do not have any hanging claim. reference the paper that backed you up, or the figure. Some examples:
L361: Word competitive used again.
Figure 5: Please be verbose about what the y-axis variable means. This can help the readers to map directly to the interpretations without the need to come back to where they really mean.
L409: “(ii) smooth…” and ”“(iii) increasing first, then…” I have difficulties seeing it in the figure. Can you confirm that this is what you meant? Also reference which catchment each of these patterns refer to exactly.
L509: which plateau? I have difficulties to see a plateau from the figure.
L540-542: But TDLR also uses Qt-1, doesn't it?
Table 2: Is this the reference for NSE?
Table 2: what does “static/dynamic” mean here?
Figure 7: why are the lines in Figure 7 not the same as the ones in Figure 5, or are the differences just a scaling effect?