the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
Technical note: Machine learning metamodelling for global sensitivity analysis
Abstract. Global sensitivity analysis (GSA) plays a central role in hydrologic modelling by supporting model understanding, diagnosis, and decision-making through the identification of influential and non-influential parameters and their interactions. Variance-based methods provide a rigorous framework for GSA but are often computationally expensive, as their estimation requires a large number of model evaluations. Metamodelling has therefore been widely adopted as a strategy to alleviate this issue, with recent advances in machine learning (ML) offering new opportunities to construct accurate and flexible surrogates for complex models. This technical note examines the practical relationship between Sobol’ total-effect indices (Ti) and feature importance measures derived from ML metamodels within a hydrologic modelling context. Building on theoretical results that link Ti to permutation variable importance (PVIi) under independence assumptions, we provide systematic numerical evidence using three conceptual hydrologic models of varying complexity (HBV, HyMod, and VIC) applied to three headwater catchments in northern Germany, together with three ML metamodels: a random forest (RF), a neural network (NN), and a linear model (LM). The three metamodels were trained on Monte Carlo samples and used to estimate sensitivities through PVIi and SHapley Additive exPlanations (SHAPi). The results demonstrate that RF and NN metamodels reliably reproduce both the ranking and relative magnitude of Ti using PVIi across all hydrologic models, providing clear empirical support for the theoretical connection between the two measures. In contrast, the performance of LM-based estimates depends strongly on the degree of linearity in the underlying model response. Mean absolute SHAPi values exhibit a consistent monotonic relationship with Ti and preserve parameter rankings, while sample-specific SHAPi values enable a distributed evaluation of sensitivities across both the parameter space and the target variable space. Overall, this study highlights ML metamodelling as a computationally efficient and conceptually sound framework for GSA in hydrologic modelling and beyond.
- Preprint
(2425 KB) - Metadata XML
-
Supplement
(3711 KB) - BibTeX
- EndNote
Status: final response (author comments only)
-
RC1: 'Comment on egusphere-2026-1787', Anonymous Referee #1, 30 May 2026
- AC1: 'Reply on RC1', Patricio Yeste, 22 Jul 2026
-
RC2: 'Comment on egusphere-2026-1787', Anonymous Referee #2, 24 Jun 2026
The present contribution elaborates on the numerical validation of the theoretical relationship between variance-based Sobol’ total index and Permutation Variable Importance (PVI), considering diverse hydrological models and three (humid and groundwater driven) catchments and Random Forest, Neural Network and Linear Model as surrogates. Additionally, the Authors explore empirically the connection between Sobol’ total index and Shapley Additive exPlanations (SHAP) index.
The work is sound and convincing, and proves how surrogate modelling can contribute to sensitivity analysis for in the hydrological context. I only have few very minor comments.
Comment
The math behind SHAPi should be better defined in the work (maybe in an appendix).
Comment
In figure 6 the label in panel (a) should be S_m instead of S.
Comment
Why Section 3.4 Sensitivity measures is not included in Section 2? Where there are all the others definitions of sensitivity?
Comment
Non-aggregated SA measure has also been proposed by Dell’Oca et al., 2020, i.e., sensitivity can be assessed across the output range of variability.
Aronne Dell'Oca, Alberto Guadagnini, Monica Riva, 2020. Copula density-driven metrics for sensitivity analysis: Theory and application to flow and transport in porous media. Advances in Water Resources, 145, https://doi.org/10.1016/j.advwatres.2020.103714.Citation: https://doi.org/10.5194/egusphere-2026-1787-RC2 - AC2: 'Reply on RC2', Patricio Yeste, 22 Jul 2026
-
RC3: 'Comment on egusphere-2026-1787', Anonymous Referee #3, 09 Aug 2026
This technical note explores the interest of using machine learning metamodels (surrogates) to evaluate global sensitivity indices. It compares several methods of ML (RF, NN) with linear regression and Sobol SHAPi indices. The paper is clear, includes a lot of work, and is worth published as a technical note, although too long see in HESS Manuscript types description: Manuscripts of this type should be short (a few pages only).
My major comments are thus on the length of the paper, on the hypothesis of this work, and on the chosen methods (evaluation only on the learning sampling, independance of parameters, choice of SHAPi to get Sobol indices)..
In particular, two things are said many times that are debatable and are linked together:
- This is said most often that the aim is not to extrapolate the surrogate outside of the learning sample, and thus not to validate it on a test (independent) sampling. This choice is much debatable, because this is sure that you will get very high R2 on the learning sample, at least with many points, and especially if you evaluate on the same points that were used to learn (whatever the number of points in the sample). We lose the interest of a surrogate that could be used on another sampling (even in same conditions). This is said and honestly assumed by the authors, but I disagree with this point.
- The interest of surrogates runs being cheaper is highlighted but, to build the surrogate, the authors use a sampling (matrix A) that was used to get the Sobol indices as well, and run it on it only (for the reason given above). So, the interest of reducing the cost is not exploited and not convincing here, since you don’t use the surrogate much more to get any other indices or any other simulations, explore uncertainties, calibration or else. One convincing interest could have been to learn on one catchment, validate on another one (or an independent sampling) and/or use this surrogate on the other catchments. But if I understood well, you learnt the ML surrogates separately on each catchment and you run them on these same catchments, just to evaluate PVI.
However , I think that the aim to evaluate the validity of equation 8 in this conceptual hydrological context is much interesting and has to be done as you did. This may be a more important point of the paper than claiming cheaper cost thanks to ML wich is not convincing the way you use them. Plus, to discuss about the cost, that would be very interesting to give the numerical cost of all methods, runs, for the 3 models and 3 catchments (at least to justify the use of 3 catchments that we don’t see in the results which are given only on one).
I have some more detailed comments and questions below.
------------------
Section 1 Introduction
------------------
L 35-39 : true, but then why choosing only one criteria (KGE). Plus, this is the case for any model that is a bit complex, not only in hydrology.
L42 : are these 3 models really expensive and high-dimensional ? We would need some values of numerical cost. Conceptual hydrological models are usually much less expensive than physically-based hydrological models.
L45 – building a high-quality surrogate is also expensive, we need a large sampling. If the model runs are not too expensive, this must be taken into consideration in the decision-making process.
------------------
Section 2
------------------
The paper could be reduced on sections describing the methodology 2,1, 2,2, 2,3, keeping a way to understand the link between the Sobol indices and the PVI from ML. I appreciate the way indices are described in section 2, to be intuitive along with the mathematical equations.
L91 : Why calling total Sobol index Ti and not Sti ?
L95-100 : Basing the GSA on independant parameters is a huge asssumption but often made. However, for a more hydrological complex model with, for ex. hydraulic caracteristics curves, it is not correct to assume independance. GSA with dependent parameters is explored in papers that may be cited, e.g. Kucherenko et al., 2012, Mara et al., 2015, Il Idrissi et al. 2021, or during the sampling process like in Lambert et al., 2025, e.g.).
L105 This is a bit surprising to refer to Puy at al. (2022) which is a R package paper (that you don’t even use in this study) more than a seminal article in GSA methods (such as Saltelli et al., Da Veiga et al, 2021, etc.).
L 140 – 144 : This relation was given, to my opinion, by Gregorutti et al. (2017). Please check or precise the role of each of these two references.
------------------
Section 3. Numerical Exp.
------------------
I think objective 1 is particularly new and interesting in this study, the obj. 2 should be more convincing, by adding numerical cost comparisons and using the surrogate in other contexts : as said before, I don’t feel like the authors gained any cost here, since they use the same sampling for SHAP indices and to build the metamodels.
L 192 : the choice of hydrological models and of the daily time step which smooths many processes may be important if you find contradictory results and rules comparing to Rouzies 2023, where the model is more complex / semi-conceptual? The models you chose are conceptual
L200 : with p the number of model parameters
210+Table 4 : is it useful since in the end you only show one catchment in the rest of the paper ?
L224 : why looking at the sensitivity of KGE if you don’t plann to do calibration after ? Why not studying a model output directly ?
L226 : uniform distributions is not so obvious for hydrological parameters. For example Ksat is usually a lognormal distribution. At least, it must be said that this is a high hypothesis and not the most common choice in hydrology.
L244 “Accordingly, model configurations that would typically be regarded as overfitting in a predictive setting are acceptable here, provided that they do not distort the resulting sensitivity measures.” => You check this because you have the SHAPi values, but the aim is to show that ML methods can replace the classical Sobol indices calculations (obj. 2 of the paper), and in the case you d’ont have the “real” Sobol indices, you are not able to assess that.
L251 : “cheap framework” : no since you have run the sampling from the matrix A? to me, unless I miss something, this is the same cost.
L256-270 : The properties of SCHAP that you cite are indeed interesting for the KGE criteria (so, in a context of calibration, that may be said explicitely). However, the choice of SHAP confused me a bit. To me they are specifically adapted in GSA to deal with dependent inputs (Owen and Prieur, 2019) and you choose these indices, assuming voluntarily independent inputs here. Can you add a sentence on this ?
------------------
Section 4. Results and discussion
------------------
I would suggest to put Figure 1 BEFORE figure 2 : we first validate the surrogates before using them for SA.
L323 again, this is logival to have very good performance since the ML is evaluated on the points the surrogate was built. This has to be explicited when discussing Figure 2.
On Fig 1 I am surprised that the ratio is not larger due to equation 8 for any model ? what does it mean in terms of variance of Y?
In Fig 2 caption, please put the configurations of NN and RF you plot. OR, it would be even better to put here all the configurations to see the interest of the one you choose and the sensitivity of the ML to the configuration. This can be done by boostrap easily. Again, I am surprised that the ratio is so small between them all.
Rouzies et al (2023) compared Sobol (from Polynomial Chaos Expansion) vs RF (and HSIC) measures on a hydrological model as well. They look at the ratio from equation 8 too and find much less evidence of it. This is a more complex hydrological model and it may allow to enrich your discussion, since you compare several complexities of hydrological models.
Fig 4: uncertainties on the estimated indices would help to see the confidence we have by choosing one method again another. It may allow to prefer ML methods of the uncertainty of indices is much smaller, for example..
Last part on SCHAPi with Fig 6 may be removed. It is very interesting and shows the interest of using SCHAPi against other (aggregated) indices, but the paper is too long, and this is independent from the rest. This gives a zoom on SCHAPI specifities although the rest of the paper focuses on the interest of ML to evaluate sensitivity. Plus, text and caption repeat the information.
------------------
References of this review report
------------------
M. Il Idrissi, V. Chabridon and B. Iooss, Developments and applications of Shapley effects to reliability-oriented sensitivity analysis with correlated inputs, Environmental Modelling & Software, 143, 105115, 2021
Da Veiga, S.; Gamboa, F.; Iooss, B. & Prieur, C. Basics and Trends in Sensitivity Analysis Society for Industrial and Applied Mathematics, 2021, doi:10.1137/1.9781611976694
Kucherenko, S., A. Tarantola, and P. Annoni (2012). Estimation of global sensitivity indices for models with dependent variables. Comput. Phys. Comm. 183, 937–946. doi:10.1016/j.cpc.2011.12.020
Lambert, G.; Helbert, C. & Lauvernet, C. Quantization-Based Latin Hypercube Sampling for Dependent Inputs With an Application to Sensitivity Analysis of Environmental Models Applied Stochastic Models in Business and Industry, 2025, 41, e2899 , doi:10.1002/asmb.2899
Mara, T. A. & Tarantola, S. Variance-based sensitivity indices for models with dependent inputs Reliability Engineering & System Safety, 2012, 107, 115-121, doi: 10.1016/j.ress.2011.08.008
Owen, A.B., Prieur, C., 2017. On Shapley value for measuring importance of dependent inputs. SIAM/ASA J. Uncertain. Quantification 5, 986–1002. doi:10.1137/16M1097717.
Rouzies, E., Lauvernet, C., Sudret, B., and Vidard, A.: How is a global sensitivity analysis of a catchment-scale, distributed pesticide transfer model performed? Application to the PESHMELBA model, Geosci. Model Dev., 16, 3137–3163, doi:10.5194/gmd-16-3137-2023, 2023.
Citation: https://doi.org/10.5194/egusphere-2026-1787-RC3
Viewed
| HTML | XML | Total | Supplement | BibTeX | EndNote | |
|---|---|---|---|---|---|---|
| 395 | 151 | 25 | 571 | 67 | 30 | 28 |
- HTML: 395
- PDF: 151
- XML: 25
- Total: 571
- Supplement: 67
- BibTeX: 30
- EndNote: 28
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
The paper presents use hydrological models and generate simulations, then train ML surrogate models and perform sensitivity & explainability analysis of parameters. Yes, these parameters are often already calibrated or studied in hydrological modelling literature, but sensitivity analysis is still meaningful because the parameters are treated as uncertain within feasible ranges. This approach is actually a good framework instead of using only ML models such as RF/ANN/LSTM/Transformer and explained those ML models. Hydrological simulations can be computationally expensive if we use sensitivity analysis (SA) which are heavily dependent on the sample size and number of parameters. For hydrological models like HBV/VIC, the traditional SA becomes extremely expensive. Mainly due to the high number of interactions, and uncertain input factors causing the curse of dimensionality problem. On the other side, run hydrological model a limited number of times, train ML surrogate, use ML surrogate for sensitivity analysis, because ML inference is very fast. This is standard surrogate modeling logic. From my understanding, the authors attempt to encourage hydrological modelers to consider surrogate modelling approaches over advanced standalone ML frameworks, mainly due to the computational efficiency offered by surrogate models based on the current experiment.
#Q1--Authors should provide the explanation for choosing these two models (RF & ANN), why not “XgBoost instead of RF” and “LSTM instead of ANN?
#Q2-- If these parameter ranges are already well established, what new hydrological insight is obtained from surrogate explainability? Or just comparing SHAP/PVI with traditional Sobol SA.
#Q3-- It would be valuable to discuss whether the reported agreement between SA and XAI SHAP remains consistent under substantially different parameter ranges. Based on the SA, RF and ANN provide the same feature rankings based on this experiment. When this setup changes, why can RF and ANN still produce similar sensitivity rankings? Even though RF uses tree splits, recursive partitioning, & ensemble averaging whereas ANN uses weighted neurons, nonlinear activations and gradient optimization. RF and ANN will behave similarly and may converge toward similar functional approximations but does not necessarily imply methodological equivalence or robustness. Results may vary across random seeds, training samples, architectures, hyperparameters and basin characteristics. I understand this is a big ask when the paper is fully concentrated on SA application (PVI & Sobol) and its compare with XAI. But one paragraph should be included to make the reader understand “Do these tools truly provide valuable insights into where ML models “align with” or “diverge from” theoretical expectations or the hydrological system understanding as per these papers- https://doi.org/10.1029/2024WR037398, https://doi.org/10.1016/j.envsoft.2026.107007
#Q4--For a moment, I considered the idea of ML as a surrogate, but do that ML models have that real practical solution where I don’t know which surrogate best reproduces hydrological model behavior? Can I trust the ML model? Since the sensitivity structure may depend strongly on the prescribed parameter bounds, it is unclear whether the reported consistency between RF/NN-based importance measures and XAI would remain stable under broader or alternative parameter ranges.
#Q5—In figure-1, although sanity check has been performed, aims to demonstrate robustness across configurations, the near-complete overlap of the importance curves makes the interpretation difficult. Quantitative stability measures would strengthen the conclusions drawn from this figure. Additional robustness analyses using different random seeds and sampling strategies would strengthen the experiment. For robustness checking, Kendall τ / Spearman ρ, top-k Jaccard, rank variance / entropy, and bootstrap confidence interval metrics could be considered. Even inclusion a few of these metrics may be sufficient to better justify the robustness claims.
#Q6: From lines 385-395 authors mentioned “The ability of SHAPi values to characterize how parameter influence varies across both the parameter space and the target….. such as DELSA, while naturally integrating with ML metamodelling frameworks.” The discussion comparing SHAPi and distributed sensitivity approaches such as DELSA (Rakovec et al. (2014) and Razavi & Gupta (2015, 2016)) may require additional nuance.
--Here, the technical note appears to implicitly position SHAP-based XAI and GSA (Sobol & PFI) as equivalent frameworks. However, while SHAP provides valuable sample-specific feature attributions, it fundamentally operates as a local prediction-explanation method rather than a global sensitivity propagation framework. maybe we can say as complementary in nature rather than same approach.
--Although aggregation in GSA (DELSA &VARS as cited by authors) may obscure localized behaviour, SHAP-based global importance derived from mean absolute SHAP values should not be interpreted as mathematically equivalent to classical GSA. In simpler term SHAP and GSA are different but may provide similar feature ranking and complementary to each other.
#Q7: Positioning SHAPi as a computationally efficient analogue to distributed sensitivity approaches may overstate the equivalence between these methodologies. In particular, SHAP computational efficiency is highly dependent on the underlying ML architecture (e.g., TreeSHAP versus ANN-based explainers) and may scale substantially with increasing feature dimensionality and sample size. I would refer this paper-https://doi.org/10.1016/j.envsoft.2026.107007
#Q8: Authors should clarify the SHAP implementation details used for different surrogate architectures (e.g., TreeExplainer for RF; linear explainer for LM model; Kernel/Deep Explainer for ANN or kernel explainer for all ML models). Please specify the SHAP explainer used for each ML model, as this information is important for interpreting both the computational cost and the resulting feature attributions.
#Q9: Since the computational efficiency and resulting feature attributions may depend strongly on the chosen explainer methodology. PVI provides a computationally efficient sensitivity-analysis framework; however, SHAP-based methodologies are generally more computationally demanding than PVI, except in the case of optimized implementations such as TreeSHAP. Both KernelSHAP and DeepSHAP are computationally more expensive compared to TreeSHAP. Authors should explain these in detail to strengthen the conclusion.
#Q10: If it is Kernel explainer for ANN, then authors should clearly mention that as KernelSHAP is highly sensitive to choices such as background dataset size, number of explained examples, and nsamples. Therefore, the study must clearly provide information such as how the background set was selected, n_samples and other KernelExplainer arguments, the number of explained points.
Q11: It is unclear why SHAP-derived global importance rankings were not included in Figure 4 alongside the PVI-based comparisons with Sobol Ti. Including normalized mean absolute SHAP rankings could help readers better assess the extent to which SHAP-based interpretations align with classical GSA results.
#Q12: Why are Sections 2.1 & 2.2 presented separately when GSA & Sobol appear to describe the same methodology. Similarly different section numbering for PVI & estimation of PVI?