the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
On combining climate models into weighted ensembles
Abstract. Several methods have been proposed and used to refine estimates of future climate change based on combined output from comprehensive climate models. While previously the so-called model democracy approach was used to combine model predictions, where every model is given equal weight, it is now widely accepted that using model weights that account for model performance and model independence is necessary to obtain more reliable results. However, most existing approaches rely, implicitly or explicitly, on a similar statistical basis, while describing things in different ways. Here we distinguish between approaches that are based on the performance of individual models (individual performance weighting) and approaches that are based on the performance of the weighted ensemble as a whole (ensemble performance weighting). At the same time, we formulate both in probabilistic Bayesian terms to make their application and comparison straightforward. Using simple constructed examples, we demonstrate that the ensemble performance weighting approach implicitly accounts for co-dependencies among models, which arguably makes the computation of independence weights for the purpose of model weighting obsolete. We also show that a set of weighted models within the ensemble weighting approach will naturally tend to artificially reduce uncertainty and that this is strongly influenced by the choice of the prior distribution over weight vectors. The distinction between individual and ensemble performance weighting is both methodological and conceptual. Formulating both approaches in general probabilistic Bayesian terms as done here, can serve as a common basis for future developments with regard to ensemble model weighting in Earth system science.
- Preprint
(2323 KB) - Metadata XML
- BibTeX
- EndNote
Status: open (until 29 Aug 2026)
-
RC1: 'Comment on egusphere-2026-2320', Anonymous Referee #1, 30 Jun 2026
reply
-
AC1: 'Reply on RC1', Britta Grusdt, 24 Jul 2026
reply
Thank you very much for your review. We address each of your points below.
In general, the manuscript is well written and comprehensible. The purpose of the study is met. Nevertheless, it would be beneficial for modelers and other people dealing with climate model data but not proficient in Bayesian statistics, to give a brief introduction regarding the most important Bayesian concepts used inside the manuscript (prior, posterior, MCMC, Dirichlet distributions, ...).
Thank you. We will clarify key concepts in the Introduction and otherwise add a brief section in the appendix to give an overview of the Bayesian concepts we use in the paper.
I recommend to publish the article on condition of adding the aforementioned explications as well as implementation of the following minor revisions:Line 67, 88, Box 1, 95, 100, 101, 110, 117, 184, ... [general]: I would recommend to refrain from the term "prediction" when talking about climate models
Yes, we will update this and use ‘model output’ or ‘simulated data’ instead of ‘model predictions’.Â
Box 1: it would be good, to include something like i=[1,N] or i=1,..,N in Box 1; in Equation 4 N is used but not defined, the first explanation of N is in line 136 (although it is quite intuitive what N should mean)
We agree that, even though it is quite intuitive, it’s better to explicitly define it in Box 1, which we will do in the revised manuscript.
equation 5: superscript v missing for y in P(y|...)
Right, we will add it.
Lines 114–120: please add reference(s) for this paragraph
We will add a reference to a classical textbook on Bayesian data analysis here.L140: please add more information and reference regarding Dirichlet distribution (if not done in the general explanations regarding Bayesian statistics)
As suggested, we will add a new section in the appendix with a short introduction to the statistical concepts used in the paper. In this section we will also add information about the Dirichlet distribution and a reference to the classical textbook on Bayesian Data analysis from Andrew Gelman et al.
Figure 1: a) although clear from the text, please add information that model 1 and 2 are hidden below model 1 in the two lower rows
We agree and will add this information to the caption.
Line 194: please inform the reader that i.i.d. = independent and identically distributed
We will do so.
Lines 217/218: please include reference for "common pri distribution ... Invserse-gamma"
We will add a reference to the Gelman textbook.
Line 266: Reference for ERA5 missing (question mark in brackets)
Thank you, this will be fixed.
Line 278: This (always same weight for M1 irrespective of number of M2 copies) only holds true for ensemble performance weighting (not for individual performance weighting) [as also can be seen in Figure 6]
That’s right, we will rewrite this sentence to make that clear.
Figure 7: is each point in Fig. 7a a different location and Fig. 7b the spatial variance? From equation 8 and the discussion I deduce that this is not the case. So I don't understand the single points. In Figure 6 there is just one weight for each model and now we have plenty of them? Or is this the probability distribution of the weights again?
That’s right. Here we show the sampled individual weights, i.e. the entries of the sampled weight vectors, while in Fig.6 we only show their mean values. We will rewrite the caption of Fig. 7 to make this clearer.
Line 332: please write out in full SSP5-8.5
We will do so.
Line 333: please use (en) dash instead of hyphen for time span
We will do so.
Figure 8: Please add unit to x-axis. Is this K per W/m² or K per doubling of CO2?
We will do so.
Figure 9: "Projected" instead of "Predicted" please, please write out in full SSP5-8.5
We will do so.
Line 351: I cannot find any "thin dashed green lines in Fig. 9"
We will fix this and will refer to the shaded regions, which we actually meant. We had previously used thin dashed green lines but overlooked this in the text.
Line 368: "ECS as performance metric" this is no metric but a variable considered for weighting (which is then applied to temperature data) -> a basis for assessing model performance
Right, we will rephrase this formulation.Lines 393–396: This is a good point. The other way round, also holds true: models with very similar model structures can lead to different projections, e.g. if they are based on the same ocean model but the atmosphere is different, …
That’s right. Two models with similar structure that yield different outputs (as in your example), are independent enough to generate such different outputs and are thus upweighted. However, two models that are structurally more different, but generate more similar outputs, are downweighted even though this should give us more confidence in the output. The worst case, however, in terms of the resulting uncertainties, is to not downweight models with similar outputs and similar structures as this would make the uncertainty range overly narrow.ÂPlease check: sometimes, blank spaces between commas ("e.g.,see") or dots ("vs.0.5820") and the following words/numbers are missing
We will do so.Citation: https://doi.org/10.5194/egusphere-2026-2320-AC1
-
AC1: 'Reply on RC1', Britta Grusdt, 24 Jul 2026
reply
-
RC2: 'Comment on egusphere-2026-2320', Anonymous Referee #2, 29 Jul 2026
reply
I consider this paper a successful attempt to put clarity in the filed of ensemble weighting.
What is probably reducing the value of this manuscript is the fact that what starts as a nice and well-presented and exhaustive theoretical framework, appears tailor made for the mere application to climate modelling.
I my view that remains a mere application, an important one but yet an application, whereas all the treatment presented applies to any kind of ensemble treatment for any kind of data.
In this respect I would change the title to reflet this aspect thus distinguishing between the general character of the topic and the application to climate models.
In ACP, other examples exist of publications that intended to put clarity on the topic of ensemble modelling regardless of the fields of application:
ACP - Est modus in rebus: analytical properties of multi-model ensembles (https://acp.copernicus.org/articles/9/9471/2009/)
ACP - Pauci ex tanto numero: reduce redundancy in multi-model ensembles (https://acp.copernicus.org/articles/13/8315/2013/)
ACP - Seeking for the rational basis of the Median Model: the optimal combination of multi-model ensemble results (https://acp.copernicus.org/articles/7/6085/2007/)
Many aspects of this paper are common with these publications although their fields of applications are different.
Distinguishing the theoretical treatment from the application case would increase the value and breath of your research. Making the case that your approach as a vaster field of applications will increase that value and breath even further.
More technical aspects were exhaustively covered by the fellow reviewer, and I have nothing more to add as the paper is very well written. Attending to his comments will further improve the manuscript, its readability and accessibility by non-experts in Bayesian statistics.
Citation: https://doi.org/10.5194/egusphere-2026-2320-RC2
Viewed
| HTML | XML | Total | BibTeX | EndNote | |
|---|---|---|---|---|---|
| 59 | 15 | 7 | 81 | 3 | 4 |
- HTML: 59
- PDF: 15
- XML: 7
- Total: 81
- BibTeX: 3
- EndNote: 4
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
The manuscript "On combining climate models into weighted ensembles" addresses model ensembles with specific weights per model based on two approaches: individual performance weighting and ensemble performance weighting. It is argued that, as regarding ensemble performance weghting weights of individual models depend on the tuning of the whole ensemble, co-dependencies of the models are implicitely accounted for. The intention of the manuscript is to delineate the two concepts based on Bayesian probabilites particularly regarding considering inter-model dependencies. Also different methodologies inside the two approaches are analysed.
In general, the manuscript is well written and comprehensible. The purpose of the study is met. Nevertheless, it would be beneficial for modelers and other people dealing with climate model data but not proficient in Bayesian statistics, to give a brief introduction regarding the most important Bayesian concepts used inside the manuscript (prior, posterior, MCMC, Dirichlet distributions, ...).
I recommend to publish the article on condition of adding the aforementioned explications as well as implementation of the following minor revisions:
Line 67, 88, Box 1, 95, 100, 101, 110, 117, 184, ... [general]: I would recommend to refrain from the term "prediction" when talking about climate models
Box 1: it would be good, to include something like i=[1,N] or i=1,..,N in Box 1; in Equation 4 N is used but not defined, the first explanation of N is in line 136 (although it is quite intuitive what N should mean)
equation 5: superscript v missing for y in P(y|...)
Lines 114–120: please add reference(s) for this paragraph
L140: please add more information and reference reagrding Dirichlet distribution (if not done in the general explanations regarding Bayesian statistics)
Figure 1: a) although clear from the text, please add information that model 1 and 2 are hidden below model 1 in the two lower rows
Line 194: please inform the reader that i.i.d. = independent and identically distributed
Lines 217/218: please include reference for "common pri distribution ... Invserse-gamma"
Line 266: Reference for ERA5 missing (question mark in brackets)
Line 278: This (always same weight for M1 irrspective of number of M2 copies) only holds true for ensemble performance weighting (not for individual performance weighting) [as also can be seen in Figure 6]
Figure 7: is each point in Fig. 7a a different location and Fig. 7b the spatial variance? From equation 8 and the discussion I deduce that this is not the case. So I don't understand the single points. In Figure 6 there is just one weight for each model and now we have plenty of them? Or is this the probability distribution of the weights again?
Line 332: please write out in full SSP5-8.5
Line 333: please use (en) dash instead of hyphen for time span
Figure 8: Please add unit to x-axis. Is this K per W/m² or K per doubling of CO2?
Figure 9: "Projected" instead of "Predicted" please, pleasse write out in full SSP5-8.5
Line 351: I cannot find any "thin dashed green lines in Fig. 9"
Line 368: "ECS as performance metric" this is no metric but a variable considered for weighting (which is then applied to temperature data) -> a basis for assessing model performance
Lines 393–396: This is a good point. The other way round, also holds true: models with very similar model structures can lead to different projections, e.g. if they are based on the same ocean model but the atmosphere is different, ...
Please check: sometimes, blank spaces between commas ("e.g.,see") or dots ("vs.0.5820") and the following words/numbers are missing