the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
Parametric Uncertainty in Prediction of AMOC weakening by an Earth System Model
Abstract. The Atlantic Meridional Overturning Circulation (AMOC) is a critical tipping element in the Earth system. While parametric uncertainty in climate models represents a dominant source of uncertainty in future projections, particularly regarding non-linear tipping phenomena, it has been largely ignored or overlooked in previous studies based on the CMIP6 model ensemble which have predominantly focused on structural uncertainties. To address this fundamental gap, this study investigates AMOC predictability by performing an uncertainty quantification (UQ) that explicitly accounts for parametric uncertainty. Crucially, this work represents the first application of an Observing System Simulation Experiment (OSSE) framework to an Earth System Model (ESM)-based AMOC predictability study under parametric uncertainty. Utilizing the Earth system model of intermediate complexity, LOVECLIM, we conduct an OSSE by varying 25 key uncertain parameters under a freshwater hosing scenario. To mitigate the prohibitive computational cost of coupled model simulations, we construct a high-fidelity surrogate model combining Principal Component Analysis (PCA) and Gaussian Process Regression (GPR) and realize efficient Surrogate-model Uncertainty Quantification, which effectively utilizes Perturbed Parameter Ensembles (PPEs). Our pseudo-observation experiments demonstrate that assimilating contemporary Sea Surface Salinity (SSS) data effectively constrains critical parameters governing freshwater transport, showing the best efficacy in reducing future AMOC projection uncertainty. Even though Sea Surface Temperature (SST) or other atmospheric observations reduce parametric and simulation accuracy in the observation period, the current AMOC and future AMOC prediction uncertainty is not reduced very well.
- Preprint
(6698 KB) - Metadata XML
-
Supplement
(6551 KB) - BibTeX
- EndNote
Status: final response (author comments only)
- RC1: 'Comment on egusphere-2026-3915', Anonymous Referee #1, 09 Sep 2026
-
RC2: 'Comment on egusphere-2026-3915', Anonymous Referee #2, 22 Sep 2026
General comments
In this study, a methodology to estimate the parametric uncertainty of AMOC’s strength projections is introduced. The methodology hinges on Observing System Simulation Experiments (OSSE), which entail producing synthetic realizations of climate states with an intermediate-complexity model. In synthesis, these realizations are employed to estimate the model’s posterior parameter distribution under different assimilation scenarios. Posterior projections are compared to the model true state to estimate the extent to which assimilating different observations reduces uncertainty. Key findings include: (i) Sea Surface Salinity (SSS) constraints model parameters more tightly than Sea Surface Temperature (SST), and (ii) large uncertainties in AMOC strength remain even under assimilation constraints.
This work focuses on an important regulator of global climate, the AMOC, develops a novel methodology to accomplish its objectives, and the conclusions claimed by the authors may have implications for the interpretation of experiments run with higher-complexity models. For these reasons, I think this study fits well within ESD’s aims and scope.
However, I believe this work suffers from two significant drawbacks. The first issue concerns form: the manuscript contains numerous grammar and syntax errors, to the extent that reading requires considerable effort. Some examples are listed at the end of the Technical Corrections. Other aspects related to form include the accuracy of the notation used in formulas and the visualisation of the results. The second issue is that, in my opinion, one of the authors’ main conclusions (SSS constrains projections better than SST) depends heavily on a specific methodological choice which is not well motivated (Specific Comment 6b). Other choices (Specific Comments 2b and 5) may also influence the results and would warrant further explanation.
My overall assessment is that this manuscript addresses important scientific questions relevant to ESD, but would require careful revision before consideration for publication.
Specific comments
- Sections 2-3: I found Sections 2 and 3 challenging due to the abundance of technical content combined with the lack of a preliminary illustration of the methods. In other words, I think it would be easier for readers to digest the details were they provided with a broad picture of the whole methodological framework beforehand. The Abstract partially addresses this (L9-17), but I would suggest that a dedicated subsection be added at the beginning of Section 2. Including a description of the OSSE framework (similarly to Section 2.1 in Kubo and Sawada, 2025) and a brief mention of the model scenario would, in my opinion, make it easier to link with the building blocks which make up subsequent sections.
- L137-155: The rationale for applying Principal Component Analysis (PCA) is clear and well explained. Nevertheless, I think the discussion may benefit from an improved explanation of the following points:
- Dimensionality: in my understanding, C is a symmetric, real-valued, (NxN) matrix. Thus, its eigenvector decomposition is such that the matrix V is square (NxN), and not (NxM) as stated in the manuscript. Similarly, there are N eigenvalues ω, and not M as stated. If this comment is accepted, please make sure dimensions are consistent throughout, including whether the “x”s are row or column vectors.
- Standardisation: the rationale for standardising the data upon division by 𝜎, as defined by the authors, is not explained in the manuscript. To my eye, the standardisation proposed is not equivalent to converting the covariance matrix C to a correlation matrix R, because in that case each variable x_n (with n=1, …, N) is divided by the square root of its own variance, and not by that of a globally averaged variance. The point of performing PCA based on the correlation matrix is to weigh variables endowed with different variances equally (Wilks, Section 11.1.2), and I am not sure this is what the authors’ transformation achieves. Here, this choice may be important as x contains information about multiple variables (L143), each potentially endowed with a different scale of variability. Finally, considering again that x contains information about multiple variables, I wonder if the formula for 𝜎 is dimensionally correct.
- It is not clear to me why 𝜎 is described as the “global mean of amplitude of annual fluctuations” (L147), considering that the authors stated earlier that x (and hence X) has been averaged over a period of 50 years (L131-132). The reference to annual mean variability seems to allude to the amplitude of random fluctuations imposed on nature run observations (L271): please clarify.
- The paragraph would benefit from a reference if available.
- L156-164: Please add a reference for Gaussian Process Regression, and explain why this specific model was chosen.
- Section 2.2.3: This is perhaps the most challenging part of the manuscript: the main thrust of the argument comes across, but several aspects of the illustration could, in my opinion, be improved:
- References would be most welcome for the Mahalanobis distance (L168; check spelling) and the Cholesky decomposition (L187). Also, the claim behind Equation 3 (“the following equation is proved by Bayes theory”) could be backed up by a reference.
- Since Equation 2 stems from Equation 3, I would consider presenting them in the reverse order. Furthermore, it is unclear why Equation 2 is introduced by “the likelihood function p(uobs|x) was modelled as the following equation” since it is in fact derived from Equation 3: my understanding is that the authors’ modelling choice consists in the representation of the observation error, and that it is from its Normality that Equations 3 and 2 follow (please clarify if this is the case). Finally, I would move the mention to the Mahalanobis distance (currently L168) to where the likelihood is actually discussed.
- For ease of reading, I would consider converting the in-line equation at L176 into a numbered equation, clearly identifying prior, likelihood, and posterior.
- Please provide more details on the notation used for the right hand side of Equation 1 and the dimensionality of the objects appearing therein.
- Notation: the truncated matrix of eigenvectors is named Vd in Section 2.2.3, but is named V here. It is unclear why “x” appears at L177 and in Equation 4. Please check the dimensionality of matrices is consistent with Section 2.2.1.
- L265-270: How many synthetic observations from the nature run have been used to estimate the likelihood and derive the posterior parameter distribution? My reading from L265-266 and L269-270 is that 51 annual-mean observations were used, one per year for model years 1950–2000. Please clarify if this is indeed the case. If it is, u_obs represents (the latent-space representation of) an annual-mean variable in Equation 2, whereas m represents an estimate of the mean of the distribution of a 50 years-mean variable (L131-132). I wonder if the authors could comment on the pros and cons of combining variables derived with different choices of time aggregation.
- L269-272: The authors opt to represent the amplitude of observation errors as a fraction of the “mean annual variability of all the grids”.
- The wording is unclear: is this the standard deviation of the 51 annual-mean samples along the time dimension, spatially averaged over all grid points?
- The same observation error is applied to all grid points, implying that errors wield a larger impact at grid points where the variability is smaller. Figure S13 shows that the spatial patterns of signal-to-noise ratio are different for SSS and SST, especially at high latitudes: the authors argue this explains the superior performance of SSS compared to SST (L417). In my view, this challenges the remarks made in the conclusions (L502-503) and Short Summary (L28-29): the fact that SSS observations better constrain AMOC predictions is a direct consequence of the authors’ modelling choices, and does not straightforwardly apply to real-world data. Have the authors considered generalising their definition to accommodate a spatially varying error matrix?
Technical corrections
- L36 “European and global”
- L46-50: This paragraph may give the impression that CESM data will be used for subsequent analysis, which is not the case. Please clarify.
- L51-60: I suggest adding a 1-line explanation of how Emergent Constraint works at the beginning of the paragraph.
- L95 Keetz et al., 2025 is not in the References.
- L104-106: Please clarify which opinion Carslaw et al., 2026 expresses.
- L118-120: Please clarify how the active core components considered in this study were selected, if relevant in relation to Kim et al. 2022. I would also consider adding a brief explanation of why the terrestrial vegetation component is necessary.
- Table 1: it is mentioned at L238 that “All the parameters are standardised into [0, 1] in our experiments”, but that is not the case in Table 1.
- L302-304: Please avoid conflicting notation with previous equations
- L325: “coefficients of determination” instead of “determinant coefficients”
- L325: “5-fold cross-validation” instead of “5-folded Cross validation”
- Figure 3: This plot is very hard to read because the details are small and get blurred if one zooms over the figure (or prints it). I recommend that the authors find an alternative visualisation strategy. Since the two-dimensional marginal distributions are not explicitly referenced in the text, perhaps the authors may consider showing the one-dimensional distributions only.
- Figure 4: This figure contains elements to which I struggle to attribute an obvious meaning, namely “(w/ IQR)” and “(All)” in the titles and “Combo” in the legend. Also, I wonder if the title is necessary at all given it is the same for all panels, and the information it delivers could be moved to the caption avoiding repetition. This last point applies to other figures, e.g. Figure 6.
- Figure 4 (and anywhere an envelope is shown): Please explain what the envelope is.
- Figure 4: it is not explained why bias decreases with noise for SSS. Also, panel 4c is not discussed.
- L586: I wonder if the comment about figures having been updated is necessary here.
- Grammar: L42-43; Syntax: L12, L42-43, L43-44, L48, L72, L112, L399-400, L454-455; Wording: L93 (“a lot”), L159 (“kernel trick”), L471 (“not so much”)
References
Kubo, A. and Sawada, Y., 2025. Predictability of climate tipping focusing on internal variability in the Earth System, Geophysical Research Letters, 52, e2024GL113146.
Wilks, D., 2006. Statistical methods in the atmospheric sciences, second ed. Elsevier.
Citation: https://doi.org/10.5194/egusphere-2026-3915-RC2
Viewed
| HTML | XML | Total | Supplement | BibTeX | EndNote | |
|---|---|---|---|---|---|---|
| 172 | 82 | 55 | 309 | 36 | 41 | 40 |
- HTML: 172
- PDF: 82
- XML: 55
- Total: 309
- Supplement: 36
- BibTeX: 41
- EndNote: 40
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
General evaluation:
This is an interesting paper, both methodologically, in studying parametric uncertainty
of an Earth system Model of Intermediate Complexity (EMIC), and in terms of its
application to AMOC weakening. The authors use LOVECLIM and construct a
perturbed-parameter ensemble in which 25 model parameters are varied. To make
Bayesian parameter estimation computationally feasible, model output is reduced
using PCA, after which GPR is used to emulate the dependence of the retained
components on the model parameters. Posterior parameter distributions are
subsequently sampled using REMC.
An OSSE setup is used in which synthetic observations of SST, SSS, and two
atmospheric variables are used separately to constrain the uncertain model
parameters. The posterior parameter samples are then propagated forward to evaluate
their effects on present-day climatology and long-term AMOC projections under a
freshwater-hosing experiment. The main result is that SSS observations provide
substantially stronger constraints (compared with SST or the atmospheric variables
considered here) on several parameters controlling freshwater transport into the
North Atlantic. Consequently, assimilation of SSS leads to a stronger reduction of
uncertainty in future AMOC behaviour.
The topic is timely and relevant for Earth System Dynamics.
However, the paper reads like a first draft and contains so many errors in language
and formulation (see a limited list below) that it is difficult to follow precisely
what is done. I would encourage the authors to subject future submissions to more
careful internal and language review before submission. I also have several major
concerns about the methodology, some of which may affect the validity of the main
result and require substantial additional work. A major revision of the paper will
therefore be necessary before I can recommend publication in Earth System
Dynamics.
Major Comments:
1. lines 175–180 and lines 550–559.
Equation (2) introduces a covariance inflation coefficient alpha, which is set to 100. This is a
quite a consequential modeling choice: multiplying the covariance by 100 substantially
broadens the likelihood and therefore directly determines how strongly the
observations can constrain the parameter posterior. Later, the authors themselves
identify covariance inflation as one of the main reasons for the remaining large
posterior uncertainty.
Why is alpha = 100 chosen? The authors should (i) provide a clear statistical
motivation for this inflation, (ii) show the results for a reasonable range of
alpha, and (iii) demonstrate whether the main conclusion—in particular the much
stronger constraint obtained from SSS compared with SST—is robust to this choice.
Without such a sensitivity experiment, it is difficult to interpret the reported
posterior widths as meaningful estimates of parametric uncertainty.
2. lines 157–164, lines 324–335 and lines 550–559
The manuscript
evaluates the GPR mainly using R-squared values of the PCA components and argues that the
first few components are reproduced sufficiently well. This demonstrates that the
emulator captures a substantial fraction of the large-scale variance, but this is not
necessarily sufficient for posterior inference. Posterior estimation depends on
accurately resolving potentially narrow regions of parameter space, rather than only
reproducing the dominant global variance of the training ensemble. Indeed, the
authors later acknowledge that 1000 training simulations may be insufficient to
resolve narrow, highly correlated posterior structures in the 25-dimensional
parameter space.
A more direct validation of the emulator should be provided to show that it is
fit for purpose. For example, one could compare emulator predictions with withheld
LOVECLIM simulations specifically in high-likelihood/posterior regions, and evaluate
whether the predictive uncertainty of the GPR is calibrated. Ideally, posterior
inference based on the emulator should also be compared for a subset of cases against
direct LOVECLIM evaluations. At minimum, the manuscript should demonstrate that the
reported differences between SSS and SST are larger than uncertainties introduced by
the emulator itself.
3. lines 137–155, in particular lines 150–155.
The
manuscript states that retaining the dominant principal components 'effectively
filters out internal variability that is independent of the parameters. However, PCA
ranks spatial patterns according to
variance; it does not distinguish parameter-induced variance from internal
variability. A mode dominated by internal variability can in principle explain
substantial variance, while a parameter-sensitive mode relevant for AMOC
predictability could explain relatively little total variance.
The retained PCA modes determine the likelihood and hence which parameter
combinations can be constrained. It should therefore be quantified how the number of
retained PCs affects the posterior and the SSS–SST comparison. Repeated simulations
with identical parameters but different initial conditions would provide a much more
direct way of separating internal variability from parameter-induced variability. If
such simulations are unavailable, the limitations of the current PCA approach should
at least be discussed explicitly.
4. lines 267–274 and lines 485–503
The manuscript's explanation for why SSS outperforms SST relies partly on differences in
signal-to-noise ratio across regions. This makes the assumed noise model directly
relevant to the central result. If the noise normalization is changed, the relative
information content of SST and SSS may also change.
The authors should either perform additional experiments with more realistic,
spatially varying errors or substantially qualify the interpretation. At minimum,
they should show that the ranking of SSS and SST is robust to alternative reasonable
definitions of observational error and to regional rather than global assimilation.
5. lines 467–472
The tipping time is defined as the
first time at which AMOC strength falls below 2.2 Sv. However, it is not clear why
2.2 Sv represents a dynamical tipping point rather than an operational threshold for
a strongly weakened AMOC. Under a prescribed, continuously increasing freshwater
forcing, crossing a low AMOC threshold does not by itself demonstrate the existence
of a bifurcation or an irreversible transition. This distinction is important because
the manuscript repeatedly uses the terms 'tipping', 'collapse' and 'tipping
predictability.
6. lines 282–294, especially lines 288–294
For computational
reasons, posterior LOVECLIM simulations are initialized from the spun-up state of the
nearest prior parameter sample rather than being spun up independently. The authors
assume that this state is sufficiently close to the equilibrium state associated with
the posterior parameter vector. Given the long adjustment time scales of the ocean
and the importance of small differences in AMOC state for the subsequent
threshold-crossing time, this assumption could influence the results.
Please quantify the parameter-space distance between posterior samples and their
selected initial-condition samples, and validate this procedure by fully spinning up
a representative subset of posterior simulations. The resulting differences in AMOC
strength and 'tipping time' should be compared with the posterior uncertainty
discussed in the manuscript.
7. lines 10–20, lines 97–107, and lines 505–515
The finding that
SSS is more informative than SST means that SSS is more informative within the
assumed LOVECLIM parameter space and under the particular forcing and
observation-error assumptions used here. It does not directly demonstrate that
present-day SSS observations will constrain the real-world AMOC, or a CMIP-class
model, to the same degree.
I suggest that the authors formulate the central conclusion explicitly in terms of
potential information content in a perfect-model experiment and discuss how
structural model error could alter the ranking of observational variables. This
distinction is especially important because the manuscript contrasts its estimates
with CMIP6 inter-model uncertainty and emergent constraints. Those uncertainty
sources are fundamentally different and cannot simply be added or compared as if
they represented equivalent probability distributions.
8. lines 5–20 and Section 5.1, lines 505–519, especially lines 510–515
It is argued
that the large parametric uncertainty found here implies that existing CMIP6-based
estimates may underestimate the 'whole uncertainty' in AMOC projections. I do not
think this conclusion follows directly from the presented results. The LOVECLIM
prior parameter ranges are prescribed by the authors and generate a particular amount
of spread. A large spread in this PPE does not demonstrate that CMIP6 models contain
an additional uncertainty component of the same magnitude.
Moreover, structural and parametric uncertainties are not necessarily independent:
part of what appears as structural uncertainty across CMIP models may itself arise
from different parameter choices or parameterizations. The comparison with Bonan et
al. (2025) is interesting, but should therefore be presented as qualitative context
rather than evidence that CMIP6 uncertainty has been underestimated by a particular
amount. The relevant statements in Section 5.1 and the Abstract should be revised to
make this distinction clear.
Minor comments:
The manuscript would benefit substantially from careful English-language editing; I
will not list any typos below.
Throughout: please use consistent capitalization for "Earth system model",
"surrogate model", "Gaussian", "North Atlantic", and names of observational
variables.
l. 1: Title suggestion: "Parametric uncertainty in predictions of AMOC weakening
with an Earth system model of intermediate complexity."
Introduction: literature on parameter-sensitivity studies using EMICs/ESMs is
missing here; for example, Boot and Dijkstra, "The effect of freshwater biases
on AMOC stability across the model complexity spectrum", Clim. Dyn. 64, 409 (2026).
l. 19: "other" -> "several".
l. 20: delete "current AMOC and future".
l. 32: "deep ocean" -> "ocean current" (this is not the deep current).
l. 37: consider using the Lenton et al. (2025), Global Tipping Points Report 2025, reference; delete "current".
l. 56: "relationship" -> "constraint".
l. 62: after "projections", references are needed.
l. 63: "climate" -> "AMOC".
l. 94: "AMOC climate" -> "AMOC".
l. 99 ff.: LOVECLIM is more precisely described as an Earth system model of
intermediate complexity (EMIC); please use this terminology consistently rather
than implying equivalence with a comprehensive CMIP-class ESM.
l. 168: "maharanobis" -> "Mahalanobis".
l. 189: Please check Eq. (4). If Sigma_total = LL^T, the usual quadratic form involves
solving Lz = (u_obs - m) and evaluating z^T z, rather than multiplying the residual by L.
The notation as currently written appears inconsistent with the Cholesky factorization.
l. 192–215: Consider moving the detailed REMC description to an appendix, as it is quite
technical and not needed to follow the main methodology. In particular, the 'temperature'
terminology may be confusing in the context of this paper, so it should be explained more clearly.
l. 222: "freshwater is poured" -> "a freshwater flux is applied". Is any compensation applied
to conserve global mean salinity?
l. 231–233: This is unclear; specify the CO2 ranges and/or radiative forcing.
Table 1: "Adjestment" -> "Adjustment" (for both P16 and P17).
l. 317: "k" -> "N"? Please check the notation.
Figure 3 is very difficult to interpret; please clarify substantially or omit it.
l. 381: The statement concerning P19 needs references.
l. 425–427: Unclear statement: please define what is meant by "thermohaline balance" here.
l. 662–668: duplicated reference; please check the reference list.