the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
Sensitivity-Aware Gradient Estimation (SAGE) for Rapid Continental-Scale Training of Hydrologic Models
Abstract. We introduce SAGE (Sensitivity-Aware Gradient Estimation), a new framework for scalable and physics-consistent training of hydrologic models that leverages analytic forward sensitivities to enable exact and efficient gradient-based learning of model parameters from catchment attributes. Unlike existing approaches that rely on finite-difference approximations, automatic differentiation, or surrogate emulators, SAGE propagates exact derivatives through physically based dynamical systems using analytically derived sensitivity equations. This eliminates the need for repeated model evaluations, substantially reduces computational cost, and preserves the interpretability and structural integrity of process-based hydrologic models. We demonstrate SAGE in a large-sample hydrology experiment using the CAMELS data set, comprising 531 hydrologically valid catchments across the contiguous United States. A feedforward neural network maps static catchment attributes to the parameter space of a conceptual rainfall-runoff model, while exact gradients of the loss function with respect to network weights are computed through analytic sensitivity propagation of the governing ordinary differential equations. Compared to conventional training strategies based on numerical differentiation or automatic differentiation, SAGE achieves machine-precision agreement with reference gradients while reducing computational cost by several orders of magnitude. To assess cross-basin model performance, we further introduce a new integrated distributional skill score based on the empirical cumulative distribution function of Nash-Sutcliffe efficiency (NSE) values across basins. Rather than summarizing performance using a single quantile such as the median NSE, the proposed score quantifies the distance between the observed basin-wise NSE distribution and the ideal degenerate distribution at NSE = 1. This distributional skill score provides a more robust and informative measure of large-sample model skill and enables objective comparison of learning strategies at continental scale. Together, SAGE and the proposed Vrugt-Frame loss score form a unified framework for both training and evaluating physics-based hydrologic models in large-sample settings and offer a new pathway toward continental-scale, attribute-conditioned calibration that is both computationally tractable and physically interpretable.
- Preprint
(2602 KB) - Metadata XML
- BibTeX
- EndNote
Status: final response (author comments only)
-
RC1: 'Comment on egusphere-2026-693', Ngo Nghi Truyen Huynh & Pierre-André Garambois (co-review team), 18 Jun 2026
- AC1: 'Reply on RC1', Jasper Vrugt, 18 Aug 2026
-
RC2: 'Comment on egusphere-2026-693', Anonymous Referee #2, 20 Jul 2026
This manuscript presents SAGE, a framework that combines neural-network-based parameter regionalization with analytically derived model sensitivities for large-sample hydrologic model training. The method is demonstrated using six conceptual rainfall–runoff models across 531 CAMELS basins.
Overall, I believe the method could make a useful contribution. I appreciate the practical idea of retaining an established C++ hydrologic model while still enabling end-to-end gradient-based training. This provides a pathway for future development of mature, non-Python hydrologic models, including modifying or replacing selected process components for hypothesis testing and scientific discovery. My main concern is that some of the broader claims regarding efficiency, predictive accuracy, and interpretability are not yet fully demonstrated. I have the following comments.
Major comments
1. The efficiency claim needs to be supported by a matched benchmark. The manuscript discusses the runtime and memory limitations of reverse-mode automatic differentiation (AD). However, I did not see a matched comparison with an equivalent PyTorch or JAX implementation. Some of the efficiency discussion relies on general statements such as training taking “days,” but no direct empirical evidence is provided under the same experimental setup. Ideally, the two approaches would use the same model equations, ANN architecture, basin set, training period, loss function, hardware, and convergence criterion. A comparison using one representative hydrologic model would already be informative. I suggest reporting wall-clock runtime, peak memory use, convergence behavior, and final model performance for both approaches. This would provide a clearer basis for evaluating the actual efficiency gain offered by SAGE.
2. The practical scope of SAGE should be clarified. One potential advantage of SAGE is that it can work with hydrologic models implemented in compiled languages without requiring the full model to be rewritten in PyTorch or JAX. However, this advantage comes with the need to derive and maintain the analytical sensitivities. It would therefore be helpful for the authors to explain this trade-off more clearly and when they consider SAGE the preferable choice to guide users. For more complex or spatially distributed models, is the sensitivity derivation still manageable?
The manuscript also mentions future hybrid physics-aware formulations (line 1044), which is a direction I find appealing for scientific discovery. Could the authors clarify how a neural network component would be incorporated into the SAGE framework? Does it require additional manual derivations? A clearer discussion of these practical limits would help readers understand the types of models and applications for which SAGE is most suitable.
3. The physical model benchmarking performance needs more clarification. The individually calibrated model results in Part A are obtained using L-BFGS-B. Can the SAGE analytical gradients also be used for individual-basin calibration? If so, it would be useful to include a Part A benchmark using SAGE, or at least clarify whether the results in Figure 5 were obtained with SAGE-derived gradients. This would help demonstrate that SAGE itself can reproduce strong physical-model benchmarks across CONUS. Otherwise, applying SAGE only in the regional static-parameter setting makes it difficult to tell whether SAGE's underperformance relative to Part A comes from the regional/static parameterization or from a difference in optimization procedure.
In addition, I do not think the current benchmarking accuracy supports the statement that the SAGE-trained conceptual models are comparable to the data-driven LSTM benchmark, given that the reported best median performance is approximately 0.66 compared with 0.73–0.75 for the LSTM (Lines 780-785). While the authors later acknowledge that “the SAGE-trained conceptual models remain, on average, below the best machine-learning benchmarks reported for rainfall-runoff simulation.”, I suggest softening the earlier statement and describing the performance difference more directly.
On a related note to both accuracy and point 2: if static attributes alone are insufficient to reach high accuracy, does SAGE support a more expressive parameterization network? For example, an RNN taking both dynamic forcing and static attributes, the way dPL uses an LSTM ahead of the physical model skeleton? Is SAGE able to accommodate more advanced architectures and additional inputs for parameter estimation, or is it inherently limited to static-attribute mappings?
4. The interpretability enabled by the explicit Jacobians should be demonstrated more directly. The manuscript relates the interpretability of SAGE to the availability of explicit Jacobians. However, the current analysis mainly presents spatial maps of normalized parameter values (Figure 8), which can also be produced by other parameter-learning approaches independent of SAGE’s sensitivity computation. I suggest including an example that uses the explicit Jacobians for model diagnosis. This would provide a clearer demonstration of the practical interpretability gained from the SAGE formulation.
Minor comments
- Lines 367-369: The calibration and validation periods appear to overlap in water year 1999 (training on WY 1999-2008 while validation over WY 1989-1999). Please clarify the exact years used for spin-up, training, and validation. Do the authors mean WY 2000-2008 for training, and WY 1990-1999 for validation?
- The statement that "calibrated parameters are often effective quantities" appears twice (first near line 912, again near line 1030). I'd suggest giving the full explanation only where it first appears and removing the repeated version.
- Line 1040: "sandwich-adjusted uncertainty estimates" would benefit from brief context, or a plain-language note on why this approach delivers "reliable parameter and predictive uncertainties within a single Markov chain Monte Carlo run."
Citation: https://doi.org/10.5194/egusphere-2026-693-RC2 - AC2: 'Reply on RC2', Jasper Vrugt, 18 Aug 2026
Model code and software
SAGEhydrology Jasper A. Vrugt https://doi.org/10.5281/zenodo.18488836
Viewed
| HTML | XML | Total | BibTeX | EndNote | |
|---|---|---|---|---|---|
| 1,041 | 994 | 76 | 2,111 | 64 | 83 |
- HTML: 1,041
- PDF: 994
- XML: 76
- Total: 2,111
- BibTeX: 64
- EndNote: 83
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
Please find our review in the attached file.
Best regards,
Pierre-André Garambois and Ngo Nghi Truyen Huynh