Reliability diagnostics for learned input-to-parameter relationships in differentiable hybrid models: RELIP v1.0 with a large-sample hydrologic evaluation
Abstract. Predictive performance remains the usual basis for assessing geoscientific models, but neural-parameterized mechanistic models expose another object that may be interpreted: the learned mapping from inputs to parameters. RELIP v1.0 evaluates whether that mapping is reproducible. It proceeds through four linked levels: predictive adequacy, stability across random seeds and loss functions, consistency of dominant controls across parameter-learning formulations, and compactness of the full relationship matrix. We apply RELIP to differentiable HBV parameter learning for 531 CAMELS-US catchments, comparing deterministic, Monte Carlo dropout, and distributional formulations under three losses and five seeds. The three formulations produced closely comparable streamflow skill, so predictive metrics alone offered little basis for choosing among them. RELIP distinguished them at the level of relationship structure. The distributional formulation had lower cross-seed variability in Spearman correlations, more reproducible dominant controls, and the most compact relationship matrix. Seven of fourteen mechanistic parameters formed a shared dominant-control core. Parameter spread also contained structured information, although several gradients were coupled to parameter means or boundary proximity. These results identify learned parameterization structure as a separate target of model assessment and provide a practical diagnostic for formulations that appear similar in prediction.