Preprints
https://doi.org/10.5194/egusphere-2026-4258
https://doi.org/10.5194/egusphere-2026-4258
11 Oct 2026
 | 11 Oct 2026
Status: this preprint is open for discussion and under review for Geoscientific Model Development (GMD).

Reliability diagnostics for learned input-to-parameter relationships in differentiable hybrid models: RELIP v1.0 with a large-sample hydrologic evaluation

Xin Jing, Na Wei, Xue Yang, and Jungang Luo

Abstract. Predictive performance remains the usual basis for assessing geoscientific models, but neural-parameterized mechanistic models expose another object that may be interpreted: the learned mapping from inputs to parameters. RELIP v1.0 evaluates whether that mapping is reproducible. It proceeds through four linked levels: predictive adequacy, stability across random seeds and loss functions, consistency of dominant controls across parameter-learning formulations, and compactness of the full relationship matrix. We apply RELIP to differentiable HBV parameter learning for 531 CAMELS-US catchments, comparing deterministic, Monte Carlo dropout, and distributional formulations under three losses and five seeds. The three formulations produced closely comparable streamflow skill, so predictive metrics alone offered little basis for choosing among them. RELIP distinguished them at the level of relationship structure. The distributional formulation had lower cross-seed variability in Spearman correlations, more reproducible dominant controls, and the most compact relationship matrix. Seven of fourteen mechanistic parameters formed a shared dominant-control core. Parameter spread also contained structured information, although several gradients were coupled to parameter means or boundary proximity. These results identify learned parameterization structure as a separate target of model assessment and provide a practical diagnostic for formulations that appear similar in prediction.

Publisher's note: Copernicus Publications remains neutral with regard to jurisdictional claims made in the text, published maps, institutional affiliations, or any other geographical representation in this paper. While Copernicus Publications makes every effort to include appropriate place names, the final responsibility lies with the authors. Views expressed in the text are those of the authors and do not necessarily reflect the views of the publisher.
Share
Xin Jing, Na Wei, Xue Yang, and Jungang Luo

Status: open (until 06 Dec 2026)

Comment types: AC – author | RC – referee | CC – community | EC – editor | CEC – chief editor | : Report abuse
Xin Jing, Na Wei, Xue Yang, and Jungang Luo
Xin Jing, Na Wei, Xue Yang, and Jungang Luo
Metrics will be available soon.
Latest update: 11 Oct 2026
Download
Short summary
Computer models can predict river flow well while learning different links between landscape features and model settings. We developed a four-stage reliability test and applied it to 45 runs across 531 United States river basins. Prediction scores were similar, but one approach produced more repeatable relationships. The results show that accurate prediction alone does not justify interpretation. Learned relationships should be tested for stability before they support scientific conclusions.
Share