the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
Sea Level Emulator Intercomparison Project (SLEIP) Phase 1: Assessing emulated multi-century global mean sea level projections
Abstract. Simplified sea level modelling approaches calibrated against reference data, also referred to as sea level emulators, can explore future sea level rise and associated uncertainties with low computational cost. Here, we introduce the Sea Level Emulator Intercomparison Project (SLEIP) to assess available sea level emulators and identify future research needs. The SLEIP protocol has a multi-century focus with 2300 as the end year to explore more of the impact-relevant time horizon and also investigate sea level responses under temperature overshoot. SLEIP covers a total of 13 sea level datasets from the participating emulators BRICK, FACTS (7 individual emulator workflows), FRISIA, MAGICC, MP25, ProFSea and SURFER. The sea level components of participating emulators range from statistical models fitted to projections or observational data to physical parametrizations, to coarse spatio-temporal physical models. They also differ in whether and, if yes, how they account for low-confidence ice-sheet processes. To remove the influence of different climate forcings on the sea level responses, we run the participating emulators with both their native climate forcings and a common forcing from the MAGICC climate model. We provide an analysis of the emulator output for the 2300 projection horizon as well as the historical period. SLEIP fills a gap in the IPCC AR6 sea level assessment by providing 2300 projections for three additional policy-relevant pathways. In the MAGICC-forced configuration, the 2300 probability boxes (p-boxes: mean of medians, lowest 17th to highest 83rd percentile) for total global mean sea level rise relative to 1995–2014 are 0.93 m (0.30–3.15 m) under the 1.5 °C-consistent SSP1-1.9, 1.36 m (0.56–10.85 m) under the strong overshoot scenario SSP5-3.4-OS, and 2.08 m (0.97–11.00 m) under SSP2-4.5, which has been used as a proxy for current climate policies. Native-forced projections generally agree closely with the MAGICC-forced results, indicating that structural differences between sea level models dominate over climate forcing, with the Antarctic ice sheet as the largest source of uncertainty. The relative consistency of thermal expansion and glacier contributions largely reflects shared parametrizations and common calibration datasets rather than independent agreement. Under SSP5-3.4-OS, an overshoot sea level rise penalty of roughly 0.1 to 0.3 m persists by 2300 compared to SSP1-2.6, even after global mean surface air temperature has returned to the SSP1-2.6 level by 2150. This first phase of SLEIP provides a lay of the land of current sea level emulation. Future phases will build on this with ScenarioMIP-CMIP7 scenarios, targeted sensitivity experiments, and a potential extension toward regionally resolved projections.
- Preprint
(6254 KB) - Metadata XML
-
Supplement
(6792 KB) - BibTeX
- EndNote
Status: open (until 28 Oct 2026)
-
RC1: 'Comment on egusphere-2026-3874', Marcus Sarofim, 10 Sep 2026
reply
-
RC2: 'Reply on RC1', Marcus Sarofim, 12 Sep 2026
reply
An addendum to my review regarding the available glacier SLE. I realized that the glacier inventories may be referenced to different baseline periods: I presume that the 0.32 m estimate from Farinotti et al. (2019) is based on present day (maybe 1998-2016?), whereas Wigley-Raper and SURFER may be baselined to earlier dates. W-R (https://agupubs.onlinelibrary.wiley.com/doi/10.1029/2004GL021238) states that V0 is the "initially available ice volume", referenced to a GSIC steady-state point which is before the "late 19th century". In addition, W-R note that "the 2-sigma uncertainty range for V0 is about 40 ± 10 cm" (I did not find the 41 cm from line 480 of the manuscript referenced in the paper, so that is also worth checking). So it might be worth going into the models that use W-R to determine their exact assumptions about glacier ice availability, and then either rebaselining everything to the same period in the 2000s, or reporting the original baseline date with each number. Another element of caution is that not all models may include the Greenland and Antarctic peripheral glaciers (RGI regions 05 and 19) - in fact, W-R explicitly exclude them - so that might be a useful consistency check. (And this is more evidence for why model intercomparison projects are both very difficult and very worthwhile)
Citation: https://doi.org/10.5194/egusphere-2026-3874-RC2
-
RC2: 'Reply on RC1', Marcus Sarofim, 12 Sep 2026
reply
Viewed
| HTML | XML | Total | Supplement | BibTeX | EndNote | |
|---|---|---|---|---|---|---|
| 238 | 100 | 24 | 362 | 39 | 27 | 31 |
- HTML: 238
- PDF: 100
- XML: 24
- Total: 362
- Supplement: 39
- BibTeX: 27
- EndNote: 31
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
Review of Nauels et al. 2026, Sea Level Emulator Intercomparison Project (SLEIP) Phase 1: Assessing emulated multi-century global mean sea level projections
This paper documents the first phase of an intercomparison project comparing the behavior and structure of 7 different sea level rise emulators (one of which has 7 workflows, making for a total of 13 datasets). This is an immensely valuable project which was also clearly the product of a substantial amount of work from a diverse group of collaborators.
Reduced complexity models are widely used in impact and risk assessments, so the first systematic intercomparison of sea level emulators will provide valuable clarity. Two high-value applications are the social cost of GHGs (see p. 2, line 60-62) and the probability of sea level exceedances. Being able to see the details of all the models in a single place will help the community choose the appropriate model for their purposes, and understand how much their results might vary with other model choices (e.g., the median model projection across emulators for SSP5-8.5 in 2300 can range from 2 m to 10 m, p. 15 line 435).
I will note that I would not have been able to identify many of the issues I raise in my review without the documentation and open-source code that the authors provide, so I wish to commend them all on that.
The paper is definitely worthy of publication.
The primary concern I have (though even this concern is fairly minor) is whether enough work was done to distinguish results that are specific to the climate driver (MAGICC) rather than the sea level rise modules – I will document a couple examples below where I have identified some specific issues, but if the authors determine they have merit, this may require some additional review for more such examples.
I also wonder if the authors could do more to apply their own judgments rather than primarily describing the model results. This kind of insight might be challenging to include given that it could be seen as criticizing individual models, and may not be consistent with the philosophy behind intercomparison papers, but I still think it would be of great value to the community if certain component modules could be identified as being clearly superior or inferior to others. I list a couple of examples below where such an inferior/superior judgment call could potentially be made (though I do not do it here), and the authors may be able to identify others.
I will conclude with minor line-by-line comments.
Climate Driver:
The paper abstract states that “[u]nder SSP5-3.4-OS, an overshoot sea level rise penalty of roughly 0.1 to 0.3 m persists by 2300 compared to SSP1-2.6, even after global mean surface air temperature has returned to the SSP1-2.6 level by 2150.” I focus in on this sentence because:
Therefore, I think section 4.4 should talk about whether different climate-drivers might produce different overshoot results. One approach to addressing this question would be to make a table listing the peak overshoot temperature by native climate model and then the resulting % difference in overshoot response relative to the MAGICC-driven version, in 2150 and 2300 (i.e., relative to the results from Figure 8h). The data from all these runs should already exist, so hopefully this is not too large an ask.
If it turns out that using native models yields results outside the 0.1 to 0.3 range, then the abstract would either need to be clear that the 0.1 to 0.3 m is a MAGICC specific result, or the abstract would need to expand the total range of penalties and make it clear that this is a result across both climate drivers and sea level rise models.
Then on page 28, line 696, it is stated that “Native-climate-forcing p-box projections are generally in close agreement with the MAGICC-forced results across all scenarios and time horizons, suggesting that sea level process uncertainty rather than climate modelling uncertainty mainly spans the overall projection envelope”, but if my instinct is correct about this particular 0.1 to 0.3 range, that might not be true for overshoot even if it holds for all the monotonic scenarios.
Component module comparisons:
GSIC and Wigley-Raper: The authors identify that the Wigley-Raper glacier model used by BRICK and FRISIA has a total volume constraint of 0.41 m, larger than the estimated 0.32 m from Farinotti et al. 2019 (Page 18, line 477-480). Do the authors consider this a reasonable disagreement between models, or is Farinotti right and the Wigley-Raper volume wrong? Also note that the BRICK volume constraint is an uncertain parameter (page 8 line 202, and see the 0.45 m of melt at the 95th percentile in Fig 4 SSP5-8.5) (given uncertainty in measurements, and ability to fit to historical trends, adjustable parameters are good practice, but be careful to document what is a constant and what isn’t).
In addition, by construction, the Wigley-Raper model can never stabilize when temperatures are higher than its equilibrium (which is 0.15 degrees below pre-industrial, as noted on p. 8, line 200). This means that even in overshoot scenarios, the Wigley-Raper GSIC component will continue melting at a non-zero rate as long as ice remains to be melted. This kind of behavior could be judged as being unphysical. Page 18 line 474-480 also misses the opportunity to talk about sea level rise in 2300 from cooler scenarios: based on this difference in structure, I would guess that the W-R models would yield more SLR at low temperatures and long timescales (and a smaller difference between the coolest and warmest scenarios in 2300, which might be an interesting metric for a lot of the SLR components).
Related, in the SURFER section, p. 14 line 409, it states that “Mountain glacier melt relaxes toward an equilibrium as a function of surface temperature anomaly, with a sea level potential of 0.5 m”: that should be added as another exception on the page 18 line 477 sentence in terms of total meltable ice available (and 0.5 m is even further from 0.32).
Thermal Expansion Approaches: p. 11, line 325-328: I would be interested in whether the authors could or would quantify the change resulting from moving from a linear relationship with OHC to a more complex (and possibly superior?) layer-specific thermal expansion coefficient approach (as in MAGICC or SURFER). I’m guessing that the change will be small, so this might not qualify for a clear “is one approach better than another” finding… but I want to ask the question. Page 19, line 508-510 does highlight the interesting result that using a multi-layer approach in overshoot scenarios enables thermosteric contraction, but again, doesn’t really answer the question “is it a meaningful enough improvement that it would be a clear advance over other models” (or are both approaches plausible so it is worth exploring the model structure space?).
The other thermal expansion approach that deserves more attention is the MP25 model that drives thermal expansion directly from GSAT without an OHC intermediary. P. 13, line 370 says, “Thermal expansion is parameterised as a linear function of global mean temperature.” However, page 21 line 547-548 says, “MP25 is an exception, parameterising thermal expansion as directly proportional to the cumulative GSAT integral and thus the only emulator with a linear relationship by construction.” Table 2 (page 6, line 155) says “MP25 uses GSAT rather than OHC as direct input for thermal expansion.”
First: be clear whether TE is a function of GSAT or a function of cumulative GSAT. The three sentences should be consistent. Second: Figure 8d shows MP25 off by itself with a much later and much less overshoot recovery. Given that MP25 is only calibrated to 2100, this might be a reason to express caution about the use of the MP25 TE approach in post-2100 exercises (though its slow initial response to warming is also worth noting).
I also note that ProFSea dips below the zero axis in Figure 8d. Given that ProFSea stands out here, it would be worth a sentence explaining this behavior – particularly since FRISIA has a very similar approach to TE and doesn’t yield the same dip.
AIS:
In contrast to the above component deep dives which identified some areas in which some models might be robustly superior or inferior, while there are substantial differences between Antarctic ice sheet models, it is much less clear to me whether any are “better” or “worse”. But again, if there are some aspects of some of the AIS components that are known to be unphysical, that would be useful to highlight.
One More Judgment Call
Page 30 line 724: “SURFER is an exception: its total sea level trajectory runs outside the observational range.” If a model does poorly with observations, should it be included in future projection p-boxes? This is potentially also true for BRICK and historical Antarctic melt. Of course, FACTS and ProFSea don’t have historical simulations, so it might be unfair to create a test for some models that other models can’t even take. Moreover, apparently while SURFER’s thermosteric component overestimates at smaller time scales it agrees well with more complex models on millennial timescales (p. 14, line 408), so this may be a case of “different horses for different courses” where each model has things it is good at and things that it is worse at, and poor observational fits do not necessarily imply worse projections.
But on the other hand, since the Table p-boxes are based on the min and max p17 and p83 values, that means that one outlier model can define the range, so some quality control could be useful if these tables are meant to inform plausible values rather than just describing the model range (see again MP25’s outlier nature for thermal expansion).
Line by Line comments:
p. 2, line 65-68: 1) I would move the “however” to be the first word of the sentence for clarity. 2) I think there is a missing word or phrase after “beyond”.
p. 5, line 122: Each emulator receives the 600 member MAGICC ensemble, but I think it should be specified either here or within the model sections exactly how that ensemble is combined with each emulator’s set of uncertain parameters (e.g., on page 10, line 282, it states that there are 2,000 samples per workflow: are those 2,000 samples matched total randomly with the 600 MAGICC members, or are the MAGICC members drawn without replacement until all 600 are gone and then reset, or are each of the 2,000 samples run with all 600 MAGICC members?).
p. 5, line 129-133: I think I understand what you are getting at, but I think you are using the wrong example. I would want something like, “(i.e., a set that was consistent with high CS in their native model being paired with a low CS set of MAGICC parameters)”.
p. 6, line 155: Table 2 lists BRICK’s land water storage component as a dash. But section 3.1, p. 8, states that “The land water storage component of BRICK is based on mass balance trends between 2003-2013 from Dieng et al. (2015). This component assumes that the mean and variance of the distribution of land water storage annual sea level trends remains stationary throughout the span of the sea-level projections.” Would this be sufficient to qualify for a “Type 1” Land water storage component?
p. 15, line 440: “are often closely together”: I think this is not worded correctly.
p. 20, line 530: I find these plots interesting! It seems like having high temperatures early leads to more influence on SLR than lower temperatures over a longer time period. I would hazard a guess that thermosteric sea level rise actually shows the opposite behavior – longer time with low warming gives more time for OHC to rise – but that the higher temperatures drive a lot more melt from land ice. Now, having put that marker down, I’ll look at Figure 7 (btw, these figures would not work in print, the ability to zoom a pdf is the only thing that makes them useable). Having done that, it looks like I’m wrong, if I’m reading across graphs correctly (one more column in Figure 7 that shows a graph for each component of median SLR per scenario against GSAT would be a nice complement to Figure 6).
p. 24, line 610: “(B) Median total sea level response for all MAGICC-forced emulators 610 under both scenarios; the overshoot penalty is the vertical gap (dashed to full line) between scenario pairs at any given time.” “(H) Overshoot penalty per emulator at 2150 and 2300, defined as the median total SLR difference between SSP5-3.4-OS and SSP1-2.6. Colours identify individual emulators as in the legend.” Technically these are two different definitions of overshoot penalty: (B) is the difference between medians, and (H) is the median of differences. They will produce different numbers, particularly if the underlying distributions are heavily skewed as would be true for any of the workflows with MICI.
p. 30, line 740: Forster et al. is listed as having a 2006-2025 GMSL uncertainty of plus/minus 2.47… but when I go to Table 11 of Forster et al, it lists the 2006-2025 GMSL as 3.66 [3.42 to 3.92]. Is this uncertainty in table 6 an order of magnitude error?
COI declaration: I have been communicating with various authors on research projects, including about SLR models, but I believe that this review reflects an unbiased assessment of the submitted article.