Beyond the simple mean: an alternative way to improve multi-model bottom-up wetland CH4 estimates
Abstract. Wetlands are the largest natural source of atmospheric methane, yet bottom-up estimates of their emissions remain highly uncertain due to structural differences, parametric uncertainties among process-based models and strong environmental heterogeneity. In addition, most wetland emission models are not systematically calibrated against methane flux measurements, limiting their ability to capture realistic spatial and temporal dynamics. The increasing availability of eddy-covariance towers measuring CH4 fluxes across diverse wetland types now offers new opportunities to evaluate and constrain model ensembles using observational data. Multi-model ensembles are commonly used to quantify uncertainty, but the widespread use of simple model averaging implicitly assumes that all models contribute equally and optimally across sites, an assumption that is rarely justified. Here, we present a data-driven framework that moves beyond the simple mean to derive site-adaptive ensemble estimates and to characterize spatial patterns of model disagreement, and how they are reshaped by ensemble weighting, in wetland CH4 emissions.
Using flux observations from a global network of 44 wetland sites, ranging from boreal arctic to tropical wetland ecosystems, and simulations from sixteen global wetland biogeochemistry models, we estimate site-specific optimal ensemble weights via a Bayesian model averaging framework fitted by Expectation–Maximization (EM-BMA). To improve robustness, weights are stabilized through resampling, and sites are clustered based on their stabilized weight signatures in compositional space, yielding groups of locations with similar model skill structures. Within each cluster, we perform cross-validated predictions and compare EM-BMA against simple model averaging (SMA) using standard performance metrics (R2, normalized RMSE, and mean bias).
To interpret these clusters and to delineate the conditions under which different model combinations are best fitting the site measurements, we relate cluster membership to environmental predictors. We characterize cluster-specific validity domains in predictor space using low-dimensional projections and geometric and probabilistic envelopes, and we identify the most influential predictors using a machine-learning classifier with back-projected feature importance.
We show that optimal model combinations vary across sites and that Bayesian model averaging outperforms simple model averaging in cross-validation. When propagated beyond the measurement sites using environmental predictors and applied to model diagnostic outputs, the resulting global wetland CH4 emission estimates differ only slightly from those obtained with simple averaging (less than 5 %). However, substantial differences emerge at local and regional scales, highlighting the importance of accounting for spatial heterogeneity in model skill. This framework therefore provides a transparent and reproducible alternative to equal-weight ensemble means for improving bottom-up wetland CH4 estimates across heterogeneous environments.
This manuscript presents a site-adaptive ensemble framework for estimating wetland methane emissions from 16 process-based models and observations at 44 wetland sites. The study addresses an important problem. Equal weighting remains common in multi-model wetland CH₄ assessments despite substantial spatial variation in model performance. The attempt to combine site observations, multi-model simulations, compositional clustering, environmental classification, and spatial extrapolation is potentially valuable. The manuscript is generally well organized, the workflow is transparent, and the recognition that model skill may vary among environmental regimes is scientifically reasonable.
However, several methodological issues currently limit confidence in the reported predictive improvement and especially in the global extrapolation. The most important concerns are the potential dependence between cluster construction and cross-validation, the very limited observational basis for global classification, the treatment of zero and non-emitting observations, the mismatch between site measurements and coarse-grid model outputs, and insufficient quantification of uncertainty in spatial cluster assignment. These issues require substantial additional analysis and clarification. I therefore recommend major revision.
Major comments
Line 99: Add a comma after “Table 1”: “These models, presented in Table 1, differ…”
Lines 151–153: GPP does not itself provide “proxies of methane emissions.” Suggested wording: “GPP was included as a proxy for vegetation productivity and potential carbon-substrate supply…”
ine 153: “vegetation activities” should be “vegetation activity.”
Line 230: Replace “As evoked in Sec. 2.3” with “As described in Sect. 2.3.”
Line 296: “This confirmes” should be “This confirms.”
Lines 318–320: Replace “This confirms” with “This suggests.”
Lines 455–457: “allow the propagation” could be revised to “enable the propagation.”
Lines 457–458: “Bayesian model averaged emissions” should be hyphenated as “Bayesian-model-averaged emissions,” or simply “BMA emissions.”