the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
Aggregating signals of Earth system dynamics across space, time, models, and variables
Abstract. Model Intercomparison Projects (MIPs) provide standardised computer simulations of the Earth system, offering unique opportunities to systematically detect and assess features and dynamics, such as abrupt shifts, across diverse models and variables. Recent advances combine time-series analysis with spatiotemporal clustering to identify dynamically connected regions within individual datasets. Yet, extending this notion of connectivity across the model and variable dimensions of MIP output remains an open challenge. Here, we present a conceptual workflow that addresses this by introducing two aggregation strategies for "detect-then-cluster" pipelines: "Detect-Cluster-Aggregate-Cluster" (DCAC) and "Detect-Aggregate-Cluster" (DAC), enabling systematic synthesis of spatiotemporal signals across multiple datasets. These aggregation algorithms are evaluated and tuned using a customisable Analytic Hierarchy Process (AHP) framework, which allows users to encode prior knowledge about dataset reliability. In anticipation of output from the Tipping Points Modelling Intercomparison Project (TIPMIP) and other MIPs within the Coupled Model Intercomparison Project (CMIP), we implement the proposed aggregation methods using the "Tipping and Other Abrupt Events Detector" (TOAD) package. To demonstrate feasibility, we apply the methods to CMIP6 simulations of Amazon rainforest dynamics, detecting and clustering abrupt vegetation shifts first across multiple variables, where a shared signal indicates a coherent ecosystem response, and then across multiple models, where a shared signal reflects model alignment. Our case study reveals that this aggregation helps distinguish such shared behaviour from dynamics that are specific to individual variables or models, patterns that typically remain obscured when datasets are analysed in isolation. These results illustrate that conclusions about abrupt dynamics depend critically on how information is synthesised across time, space, models, and variables. While showcased here in the context of tipping points, the proposed aggregation framework provides a structured and transferable foundation for multimodel and multivariate risk assessments of diverse Earth-system processes within MIPs.
- Preprint
(5075 KB) - Metadata XML
- BibTeX
- EndNote
Status: final response (author comments only)
- RC1: 'Comment on egusphere-2026-3389', Anonymous Referee #1, 30 Aug 2026
-
RC2: 'Comment on egusphere-2026-3389', Anonymous Referee #2, 30 Sep 2026
The manuscript presents potentially useful methods for aggregating signals across models and variables. However, the central methodological and scientific contribution is often difficult to identify, and the manuscript would benefit from substantial streamlining. I also encourage the authors to clarify more explicitly the relevance of the work to Nonlinear Processes in Geophysics, as it currently reads primarily as a general MIP data-analysis framework.
Major Comments
1. Streamline the manuscript and clarify the high-level contribution- Introduction: The manuscript is unnecessarily long, which makes the main contribution difficult to identify. The Introduction should state the key methodological problem and the main contribution more directly; the detailed bullet points at the end could be replaced by a short, high-level summary.
- Sections 2 & 3: Both sections should also be substantially shortened. In particular, the distinction between DAC and DCAC can be explained more concisely using Figure 1, while technical details of AHP could be moved to an appendix.
2. Quantitative validation using synthetic data
The current synthetic experiment mainly illustrates how the methods work, rather than validating their performance. The authors should test more challenging cases, such as signals present in only a subset of datasets, different noise levels, and spatial or temporal misalignments. Quantitative evaluation against known ground truth (e.g., Precision, Recall, or IoU) would help demonstrate when DAC, DCAC, and the different aggregation functions perform reliably.3. Distinguish "aggregation" from "consensus"
The manuscript should more clearly distinguish aggregation from cross-model consensus. For example, Cluster 2 in Figure 6 is mainly driven by an abrupt shift in GFDL-ESM4, while the other models show different responses. Since MAM can preserve strong signals from only a subset of models, an aggregated signal does not necessarily indicate multi-model agreement. The terminology and interpretation should be revised accordingly.4. Clarify the choice of aggregation method
The manuscript should provide clearer guidance on when to use DAC or DCAC and how to choose among Mean, Median, and MAM. Since different methods are emphasized in the multivariate and multimodel examples, it is currently unclear how these choices should be made for a new dataset. The authors should explain how the appropriate method depends on the research question or characteristics of the data.5. Justify the AHP framework and assess its sensitivity
- Sensitivity: The variable and metric weights in Tables D1 and D2 involve subjective choices. The authors should examine whether the selected cluster maps are robust to alternative weights, including equal weighting.
- Role of AHP: Since the pairwise comparisons are constructed directly from score ratios (\(S_{ij}=s_i/s_j\)), it is unclear what AHP adds beyond normalization and weighted scoring. This should be clarified.
- Terminology: The term “optimal” should be used cautiously, as the selected maps are simply the highest-ranked candidates under the chosen metrics and weights.
6. Clarify key methodological choices
- DCAC frequency: Eq. (4) combines variations across datasets and DBSCAN parameter settings into a single consensus measure, although these represent different sources of variability. The rationale and normalization of this quantity, including the “cluster frequency” in Figure 3, should be clarified.
- GWL transformation: Remapping model time to smoothed GWL may affect gradient-based abrupt-shift detection and may obscure differences in ecological response timescales across models. The potential influence of this transformation should be discussed.
Minor Comments
- Eq. (B5): Verify the denominator \(n(n-1)\) if the summation is over unique unordered pairs.
- Distance metric: Please justify the use of Ward’s hierarchical clustering with \(d = 1-r^2\) in Eq. (B6).
Citation: https://doi.org/10.5194/egusphere-2026-3389-RC2
Viewed
| HTML | XML | Total | BibTeX | EndNote | |
|---|---|---|---|---|---|
| 182 | 65 | 29 | 276 | 22 | 18 |
- HTML: 182
- PDF: 65
- XML: 29
- Total: 276
- BibTeX: 22
- EndNote: 18
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
The paper by De Maeyer et al “Aggregating signals of Earth system dynamics across space, time, models, and variables” proposes several methods of combining complex climate data for studying tipping points, with an example of Amazon rainforest.
I find the paper structure strange and unsuitable for this journal: the metrics and relevant plots and tables are placed in the appendix; this is typical for Nature submissions, and similar journals, but in fact not easy to follow, and better to change into regular structure: Intro – Method – Data – Results.
I find the title of the paper misleading: the paper is not only about aggregation, and space and time are included in the variables that are mentioned along those. Furthermore, aggregation may be not the best term here. I would call the approach assemblage: aggregation usually leads to a single statistics, speaking mathematically - which is what is shown in section 2.1 but is more narrow than what the title states. In Fig.1, aggregation is mentioned as an alternative of clustering, and this contradicts the paper title.
The meaning of “aggregation” (assemblage) should be explained much earlier than it currently is in page 4. Similarly, clusters are mentioned early but explained much later, and this complicates reading of the paper. When clustering is mentioned, it is necessary to explain whether it is spatial or temporal; later there are “spatiotemporal” clusters – this requires clarification. Clustering is discussed in general terms for many pages, and only in page 14 the DBSCAN technique is mentioned.
As I understand it, by “detection time series” (time series analysis of dynamics of interest) the authors mean, in particular, early warning signal indicators – why is it not mentioned explicitly?
In line 200, relative importance vector is introduced but a reference is not provided.
In section 4.2, the processed data represents 150 annual mean datapoints. Considering the context of “abrupt shifts”, annual resolution may be too crude. Also, it would be useful to include an actual vegetation map of the current state of the Amazon rainforest. This would be interesting to compare with the three clusters in Fig.5.
FURTHER COMMENTS
> why sorting is mentioned in Eq.1? Median is a standard metric
> acronyms DAC and DCAC are defined multiple times
> the caption of Fig.2 should mention what data are analysed
> in line 135, what values are meant as “high counts”?
> Fig.3, panel (c) – are frequencies normalised? From the x-range between 0 and 100, this is unlikely. Then why y-axis has label “cluster frequency”? The caption mentions “counts” but the values are below 1 – what processing was applied to integer counts?
> In Eqs.7,8, only three datasets are used – isn’t it a formula for an arbitrary number of datasets and should be expressed by a general sum?
> Tables 1 and D1 can be merged, and the order of their lines can be rearranged according to the decreasing weights
MINOR COMMENTS
> In the 4th affiliation, System should start with a capital letter. The same in line 21 of the main text.
> in line 184, word “priority” can be omitted (a priory is sufficient)
> in Table 1, 4th line from bottom, CO2 should be with 2 as index
> line 381, clarify the meaning of “a clear, and consistently abrupt” – if this was meant as an Oxford comma, that is incorrect.
> in the caption of Fig.6, panel (c), explain that the colours of the curves correspond to the colours of the clusters.
> line 493, 564, 568 start with a dot