the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
An Earth system deep learning classifier for tipping point detection
Abstract. Tipping points are thresholds at which a system, often abruptly and irreversibly, transitions from a stable state to a contrasting one. Crossing such critical boundaries poses a risk to Earth system stability and may have catastrophic consequences. This is especially relevant, as current climate change is destabilizing Earth subsystems, potentially bringing them closer to tipping points. Thus, it is important to be able to detect approaching tipping points in the Earth’s system, which can be achieved through calibration on palaeo-records. Recently, new deep learning (DL) methods have been established that are able to confidently and quantitatively identify different types of critical transitions characterised by their abruptness and (ir)reversibility. Based on this, we develop a new (simplified) DL classifier focusing on the quantitative detection of catastrophic tipping points (fold bifurcations) in the Earth system. Our approach reduces computational demand and improves performance, especially for short timeseries. We first test the new classifier's performance on synthetic data and subsequently on different existing Cenozoic proxy records. Our DL results are compared to the results from previous studies applying generic early warning signals (EWS), which can detect approaching transitions qualitatively but cannot distinguish bifurcation types (abruptness and (ir)reversibility of the transition). Our DL classifier enables us to identify how abrupt and (ir)reversible an approaching transition is, which is important for tipping point risk assessment and mitigation. Results are generally consistent between generic EWS from previous studies and our DL approach and fit with what is known from the geological context. We note that some results are dependent on the length of the classifier used and the time interval investigated before the bifurcation. We implement an out-of-distribution (OOD) detection method to reduce the misclassification of non-catastrophic bifurcations as catastrophic tipping points. Combined with the binary DL classifier, this approach enables reliable, quantitative detection of catastrophic tipping points in Earth system records.
- Preprint
(1208 KB) - Metadata XML
- BibTeX
- EndNote
Status: final response (author comments only)
- RC1: 'Comment on egusphere-2026-2223', Chris Boulton, 28 May 2026
-
RC2: 'Comment on egusphere-2026-2223', Laiming Zhang, 24 Jul 2026
Overall, this manuscript builds on the work of Bury et al. (2021) by developing a binary CNN–LSTM classifier specifically designed to identify fold-bifurcation dynamics. The classifier is combined with an energy-based out-of-distribution detection method in an attempt to move beyond general early-warning signals towards the identification of bifurcation type. The research question is important, the methodological approach is promising, and the applications to palaeoclimate and palaeoceanographic records are potentially valuable.
However, the manuscript currently needs to define more clearly what can and cannot be inferred from the model output. The robustness of the empirical results with respect to classifier length, analysis-window selection, and the choice of event intervals also requires further evaluation. In addition, the construction of the synthetic training data and the extent to which patterns learned from low-dimensional synthetic systems can be transferred to complex Earth-system records should be explained in greater detail. My main comments are outlined below.
Major comments
1. The relationship between fold bifurcations and irreversible Earth-system tipping points needs to be explained more carefully
The manuscript repeatedly associates fold bifurcations with abrupt and irreversible Earth-system tipping points. However, the mathematical meaning of a local fold bifurcation is that a stable and an unstable equilibrium collide and disappear. This may produce an abrupt transition and, under an appropriate bistable and global branch structure, may also lead to hysteresis. A local fold bifurcation alone, however, does not directly demonstrate strict irreversibility in an empirical system.
This distinction is particularly important here because the empirical analyses mainly use time series preceding the prescribed transition point. They do not examine the post-transition state, recovery trajectory, or system response when the external forcing weakens or reverses. The conclusions supported by the present analysis are therefore more accurately framed as follows: the pre-transition time series contain dynamical patterns consistent with those generated before synthetic fold bifurcations. This is not equivalent to demonstrating that the corresponding Earth-system transition was irreversible.
I recommend that the authors distinguish more clearly among fold bifurcation, hysteresis, difficult-to-reverse behaviour, and strict irreversibility. They should explain to what extent potential hysteresis or limited reversibility can be inferred from pre-transition time series alone, and discuss the limitations of that inference. Statements such as “detect irreversible tipping points” should be moderated, for example to “identify time-series patterns consistent with an approaching fold bifurcation”. Similarly, wording such as “how abrupt and irreversible” should be avoided because the model provides a class or class probability rather than a quantitative measure of the degree of abruptness or irreversibility.
2. The results are sensitive to classifier length and analysis-window selection, but no systematic sensitivity analysis is presented
One of the most important results in the manuscript is that different input lengths can lead to opposite classifications for the same event. This is not a minor shift in probability, but a change in the inferred bifurcation class. The PETM and the Oligocene–Miocene Transition provide clear examples. The discussion offers several possible explanations, but these are currently largely post hoc and are not tested systematically.
I suggest that the authors use the existing classifiers to analyse a set of nested pre-transition windows for each sufficiently long record. Fold probability could then be shown as a continuous function of both the effective number of observations and the physical duration represented by the window. This would reveal whether the classification is stable over a range of window lengths or changes abruptly as earlier parts of the record are included.
The current comparison between length-500 and length-1500 changes both the empirical window and the classifier itself. It therefore remains unclear whether the differing results arise from the amount of empirical data included or from differences between models trained at different sequence lengths. These two effects should be separated as far as possible.
In addition, 500 data points represent very different physical durations in the PETM, CENOGRID, and sapropel records. The phrase “500 points” therefore combines sampling resolution, physical duration, and the intrinsic timescale of the system. Window length should consequently be reported in both number of observations and actual time span.
The manuscript also needs a reproducible criterion for selecting classifier length. When the length-500 and length-1500 classifiers disagree, it is not clear how a future user should decide which result is more reliable or how the two outputs should be combined. If a single selection rule cannot be provided, this length dependence should at least be treated explicitly as an important source of uncertainty rather than resolved through event-specific mechanistic interpretations alone.
3. The empirical validation lacks non-event control intervals
The empirical applications focus on predefined major transitions and analyse the intervals immediately preceding them. This is useful for comparison with earlier EWS studies, but it does not provide a realistic estimate of the background false-positive rate in geological records.
Even if high fold probabilities occur before known events, it remains necessary to determine whether similar signals also occur in ordinary background intervals, in non-stationary intervals that are not associated with a recognised tipping event, or in intervals affected mainly by proxy noise or preprocessing choices.
I recommend that the authors include several non-event control windows for each type of empirical record. These should be matched as closely as possible to the event windows in terms of duration, sampling resolution, and preprocessing. Suitable controls could include adjacent background intervals, intervals without recognised major transitions, or a small number of randomly selected ordinary intervals. The full workflow should then be applied in the same way, preferably without using event labels during the analysis.
Such controls would allow the authors to estimate the background false-positive rate and determine whether high fold probabilities are specifically concentrated before known transitions. This is especially important for the sapropel application, where S3–S9 are all classified as folds and share similar proxy characteristics, depositional settings, and processing procedures. If adjacent non-sapropel intervals also produce high fold probabilities, the event specificity of the model output would need to be reconsidered.
4. The origin, construction, and representativeness of the synthetic training data require fuller explanation
The description of the training data relies heavily on Bury et al. (2021). Referring readers to the earlier paper is reasonable, but the present manuscript should still be sufficiently self-contained for readers to understand what data were used, what the classifier was trained to recognise, and how the synthetic systems differ from the empirical records.
At minimum, the manuscript or Supplement should provide the general form of the randomly generated two-dimensional dynamical systems, explain how system parameters were sampled, define the fold and null classes explicitly, identify which state variable was retained, and describe how the control parameter approached the bifurcation. The distributions or ranges of the parameter-drift rate and stochastic-noise intensity should also be reported, because both can affect variance, autocorrelation, the duration of critical slowing down, and classification performance.
The hierarchy of the synthetic dataset also needs clarification. Did each underlying dynamical system generate multiple stochastic realisations? Which elements varied among those realisations? Was the train/validation/test split performed at the level of individual time series, or were all realisations from the same underlying system kept within a single subset?
If different realisations of the same underlying dynamical system occur in both the training and test sets, the reported performance may overestimate generalisation to genuinely unseen systems. Where the dataset structure allows, a group-wise evaluation in which underlying dynamical systems are held out from training would provide a stronger test. If such an analysis is not feasible, the manuscript should at least describe the current splitting procedure clearly and discuss its implications for interpreting model performance.
Finally, the manuscript should discuss more fully the transferability of patterns learned from random, low-dimensional synthetic systems to complex Earth-system records. Palaeoclimate time series are influenced by multivariate coupling, non-stationary external forcing, age-model uncertainty, irregular sampling, proxy-specific processes, and sedimentary or diagenetic effects. The authors should explain which pre-bifurcation features are expected to be sufficiently universal across systems, which empirical characteristics may fall outside the training distribution, and how these differences limit the interpretation of the PETM, CENOGRID, and sapropel results.
Specific comments Abstract and Introduction
Lines 23–24:
The classifier appears to assess whether an input series is more consistent with fold or null dynamics. It does not quantify how abrupt or how irreversible a transition is. Please revise the wording accordingly.Line 36:
“The Earth system” would be more conventional than “System Earth”.Lines 61–66:
Generic EWS are based on quantitative measures such as autocorrelation and variance. Their limitation is not that they are qualitative, but that they do not generally identify the type of bifurcation. Please revise this comparison and state the added value of the DL approach more precisely.Around Line 76:
The statement that transcritical bifurcations are relatively rare in nature should be qualified. Please clarify whether this refers to mathematical non-genericity or to their demonstrated frequency in empirical natural systems, which is difficult to establish directly from observations.Lines 82–87:
The phrase “two bifurcation types of interest” is inaccurate because the null class does not represent a bifurcation. “Two classes of interest” would be more appropriate.Methods
Lines 99–109:
Please add a concise but self-contained description of the synthetic training data, including the general model form, parameter sampling, class definitions, control-parameter drift, stochastic forcing, and dataset hierarchy.Line 114:
Please explain the rationale for the unusual 0.95/0.04/0.01 train/validation/test split. It would also be helpful to report the variability in model performance across different random initialisations or data splits.Lines 143–145:
Please explain why nearest-neighbour interpolation was selected and discuss its potential effects on temporal structure. This method can introduce repeated values and step-like patterns, potentially altering autocorrelation and the form of the model input. A sensitivity comparison with at least one alternative interpolation method would strengthen the analysis.Line 213:
Using only 100 transcritical time series to estimate the softmax misclassification rate seems limited given the larger synthetic test dataset. Please explain how these series were sampled, report an uncertainty interval, and, if feasible, repeat the estimate using a larger sample.Figures and Results
Figure 5:
The figure combines the synthetic OOD distributions with empirical results that are not discussed until later in the manuscript. This interrupts the sequence of the results. Please consider separating the synthetic validation and empirical applications into different figures, or revising the placement and order of the figure references.Figure 5 caption:
The caption should briefly explain how the yellow threshold was determined, or direct the reader to the relevant part of the Methods.Table 5:
The main text refers to median noise levels, whereas the table heading reports means. Please verify the statistic used and make the terminology consistent.Section numbering:
Section 3.3.3 is followed directly by Section 3.6. Please check whether Sections 3.4 and 3.5 are missing or whether the numbering should be corrected.Interpretation of the empirical results:
In several places, agreement between the model output and an existing geological interpretation is described as “corroboration”, for example in relation to ETM3 being the weakest event. Please provide appropriate references and distinguish between consistency with an established interpretation and independent validation of that interpretation.Line 475:
“Other additional mechanisms” is redundant. Please use either “additional mechanisms” or “other mechanisms”.Citation: https://doi.org/10.5194/egusphere-2026-2223-RC2
Viewed
| HTML | XML | Total | BibTeX | EndNote | |
|---|---|---|---|---|---|
| 306 | 146 | 21 | 473 | 33 | 32 |
- HTML: 306
- PDF: 146
- XML: 21
- Total: 473
- BibTeX: 33
- EndNote: 32
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
I found the manuscript to be an interesting read and a useful follow-on from the Bury et al., 2021 paper, with a focus that switches to important considerations for the Earth system, such as are the abrupt shifts the DL model might predict the approach to irreversible or not. I think the work presents a great step forward in the literature, but I have some concerns about the robustness of the results and I think it would benefit from stepping away from the Bury et al., paper by explaining certain aspects in more detail rather than leading the reader to read the other paper. I have detailed my comments below.
One of my main comments is that I do not think the training set is explained well enough. I think the time series essentially are the ones from the ‘fold’ and ‘null’ categories in the other paper, but perhaps an equation explaining the derivation would help the reader understand what is being used in this training specifically. I get the impression that the underlaying equation can be used to simulate a number of different bifurcations but this information would help the reader understand properly.
To a certain degree I can understand why the transcritical time series were used as an OOD test but currently it is difficult to look past what feels like a strong focus on fold and transcritical bifurcations. An instant question would be why not include transcritical time series in the training set as null cases and I assume this is because it would be difficult to determine where to draw the line on what to include. I would be interested to know how well a DL model trained on a binary output of those two differs from the current model.
Regarding the OOD tests, I would be interested in seeing how other time series are classified e.g. different types of bifurcations (non-catastrophic Hopf), white to red noise processes where the memory in the system is increased, or stable time series where the noise level is increased. Furthermore, including catastrophic Hopf bifurcations in an OOD test could tell you how reliant the model is on only fold bifurcations being considered catastrophic shifts.
I have a few minor comments regarding readability and presentation which I have detailed below.
Line 24 – The classifier only identifies if a transition is abrupt and irreversible, not ‘how’ abrupt or irreversible it is.
Line 36 – Consider ‘The Earth system’ rather than ‘System Earth’.
Line 65 – I’d argue that generic EWS do not really have a qualitative nature, they use basic statistical measure and trends can be quantified. Perhaps rephrase to properly describe the added benefit of DL.
Line 85 – ‘Two bifurcation types of interest’ suggests that the ‘null’ category is a bifurcation in this instance.
Line 213 – 100 time series seems quite low, why was this?
Mentions of Fig. 5 – I would consider how Figure 5 is referenced as it is used a lot throughout the manuscript. Maybe do not reference it until the results section so it comes after Fig. 4. It is also confusing as it has results on it regarding the palaeo data that are not discussed in the text until much later.
Figure 5 caption – Should briefly explain how the threshold (yellow line) was calculated or point the reader to the main text to find out.
Table 5 – To me (both in the table and in the main text) it reads that time series with higher noise are likely to be above the energy threshold regardless of if they are ID or OOD, meaning that it is classifying noisier time series as fold bifurcations regardless. Higher energy scores according to Fig. 5 should be associated with fold bifurcations yet the text on line 294 says they are associated with greater confidence in identification. Line 514 says something similar.
Section 3 – I think some specific referencing would be beneficial throughout the explaining the context of the results. Line 314 for example saying that your results corroborate ETM3 being the weakest event needs a reference.
Line 475 – ‘other’ should come after ‘additional mechanisms’ not before.
I would suggest on the heatmaps (Fig. 6-8) on having one colour bar that goes from blue to red with grey in the middle. I am struggling with ~0.6 looking quite similar to me on both bars. Also be careful around these values on using the black fold text (Figure 7 for example), although solving the colour bar problem might fix this.