the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
On using neural networks to predict mean Age-of-Air from long-lived tracers
Abstract. Climate models predict changes in the Brewer-Dobson circulation under a changing climate, which could have profound effects on tracer distributions and the radiative budget. Age-of-Air is an important concept for describing transport in the stratosphere and to understand and quantify global atmospheric circulation patterns such as the Brewer-Dobson circulation. Being an unobservable quantity, it must be inferred from other, directly observable quantities such as long-lived trace gases. It is therefore essential to have accurate ways of determining Age-of-Air through observations. These observations are subject to measurement noise, which is a long-known source of uncertainty when deriving Age-of-Air, as such uncertainties can affect the derived Age-of-Air significantly. We present a novel approach of using neural networks to derive Age-of-Air from long-lived trace gases. Multi-layer perceptrons can be used to predict model Age-of-Air with accuracy as little as a single month. The networks can be optimally trained according to expected measurement uncertainties. An unsupervised autoencoder is presented which is capable of achieving similar predictability almost without relying on model Age-of-Air inputs. This study presents an overview of these new approaches and discusses their capabilities, their accuracies and precisions mostly from a technical perspective regarding input parameters, regularization and predictions outside the training domain. Our approach allows us to derive Age-of-Air with accuracy of up to a single month under considerable measurement noise and over a wide altitude range. This accuracy is even retained when predicting values from a completely different period.
- Preprint
(7949 KB) - Metadata XML
- BibTeX
- EndNote
Status: final response (author comments only)
- RC1: 'Comment on egusphere-2026-4183', Anonymous Referee #1, 29 Aug 2026
-
RC2: 'Comment on egusphere-2026-4183', Anonymous Referee #2, 27 Sep 2026
I commend the authors for taking a first deep dive into investigating the potential for using neural networks to derive age-of-air from long-lived trace gases. On novelty alone the manuscript receives high marks for its innovation and potential for serving as a important reference moving forward. There are several technical points and issues with the manuscript presentation, however, that require attention. While I consider these major and critical for improving the manuscript, I think they are doable.
Major Comment 1:
I appreciate that the study is ambitiously establishing an approach for constraining age-of-air (AoA) using neural networks that cannot easily rely on methodologies established in past references. Thus it is inevitable that some pieces may be missing and/or questionable. While, overall, I am impressed with the amount of attention paid to methodological sensitivities, I think there are still some concerns that I would like addressed:
a) First, a major concern is that normalizing each dataset independently may artificially improve temporal generalization (Section 2.2, lines 119-180 and Section 3.3, lines 458-468). In particular, is it possible that the network’s ability to predict AoA in 2021, despite being trained only on 2011 data, partly reflects the fact that you have normalized tracer concentrations using that year’s minimum and maximum values? It seems like this pre-processing step is building in information about the underlying tracer distribution, no? Could the authors please address this possibility and, if so, state this caveat clearly in the text? It could be worth doing an experiment where you train the network using normalization parameters derived exclusively from the 2011 training dataset and then apply those values, unchanged, to 2012, 2016 and 2021. Do you get similar results?
b) Certain sensitivities in both the supervised and unsupervised networks are not explored/addressed. For example, it’s not clear to me why the supervised networks use three hidden layers with 300 neurons each. Would it be possible to do a sensitivity experiment comparing one, two and three hidden layers and/or 30, 100 and 300 neurons per olayer? In other words, some discussion of sensitivity to selected architecture is missing but seems important to include.
Major Comment 2:
I’m not sure I agree with the asseriton that the autoencoder approach is entirely unsupervised. If I understand things correctly, the AE itself is unsupervised as it yields the (1-D) latent representation of the LLTs, but then the ensuing calibration steps are critical in mapping the latent representation to AoA (e.g., Figure 13d). Perhaps more care can be taken throughout to make sure that the AE approach is not misrepresented as divorced from AoA information.
Minor Comments:
Is there a way to methodologically infer the information content of the individual LLTs? 6 are used but it is not clear which better constrain the AoA versus others. Having access to this information would be important in prioritizing measurements, etc. Perhaps a statement can be made about this in the conclusions.
Figures 3 and 10: I am not sure that having both figures is necessary as they both illustrate the AE architecture, no? It seems like Figure 10 is just a split up version of Figure 3. Please either better justify the need for both figures or remove Figure 10.
Technical Errors:
Line ~18 Error: Missing space before citation: “transport(Brewer…”
Correction: “transport (Brewer…”
Figure 1 caption Error: “Adapated” Correction: “Adapted”
Figure 1 caption Error: “where it may reside for around one year Ray et al. (2017)” is missing citation punctuation.
Correction: “…around one year (Ray et al., 2017).”
Lines ~35–36 Error: “high and mid latitude subsidence”
Correction: “high- and mid-latitude subsidence.”
Lines ~58–59 Error: “chemical species which have lifetimes…”
Correction: Grammatically preferable: “chemical species that have lifetimes…”
Lines ~65–66 Error: “Other LLTs are also subject to other depletion mechanisms”
Correction: Redundant “other”; delete one.
Lines ~97–99 Error: “an overview is given over the simulations”
Correction: “an overview is given of the simulations.”
Line ~101 Error: CLaMS is expanded as “Chemical Lagrangian model of the Stratosphere,” whereas elsewhere it is called “Chemical Lagrangian Model of the Atmosphere.”
Correction: Make terminology consistent. The standard name is “Chemical Lagrangian Model of the Stratosphere.”
Lines ~104–105 Error: “physics-informed” appears broken in the PDF extraction.
Correction: Check typesetting and ensure “physics-informed” is properly hyphenated.
Lines ~105–106 Error: Citation construction: “provided by … (ECMWF) Hersbach et al. (2020)
Correction: “…ECMWF (Hersbach et al., 2020).”
Lines ~107–108 Error: “coordinate Konopka et al. (2019)”
Correction: “…coordinate (Konopka et al., 2019).”
Line ~111 Error: “initiated at 1st January 1979”
Correction: “initiated on 1 January 1979.”
Line ~115 Error: “data is interpolated”
Correction: “data are interpolated.”
Lines ~124–125 Error: “normalized … to range (0, 1)”
Correction: If extrema map to 0 and 1, this should be “normalized … to the range [0,1].” Parentheses mathematically imply that 0 and 1 are excluded.
Line ~126 Error: “The values in each dataset therefore are uniform, random samples across different altitudes, latitudes, and longitudes.”
Technical issue: Random sampling of CLaMS parcels does not imply a uniform distribution in altitude, latitude, and longitude.
Correction: State that samples are drawn randomly without deliberate spatial weighting, if that is what is intended.
Lines ~144–145 Error: “the … technique AirCores which are used”
Correction: For example, “the AirCore measurement technique, which is used…”
Line ~145 Error: “2%, 3% and 5%, which is the order of magnitude…”
Correction: “which are of the order of magnitude…”
Lines ~166 Error: “There vectors”
Correction: “These vectors.”
Lines ~166–167 Error: “fed-forward” used as a verb.
Correction: “fed forward.”
Lines ~193–194 Error: “batchsize”
Correction: “batch size.”
Lines ~198–200 Error: “An overview into the advantages…”
Correction: “An overview of the advantages…”
Line ~200 Error: “Tensorflow-Probability”
Correction: Official name/capitalization is “TensorFlow Probability.”
Lines ~211–213 Error: “It consists of two types of neural networks itself”
Correction: “It itself consists of two neural networks: an encoder and a decoder.”
Lines ~216–217 Error: “The decoder is the pseudo-inverse of the encoder.”
Technical issue: This is too strong/incorrect. The decoder is trained to approximate an inverse mapping over the learned
representation; it is not generally the mathematical pseudoinverse of the encoder.
Lines ~218–220 Error: “The constraint of an AE is that only such encodings can be determined which are invertible…”
Technical issue: An undercomplete encoder mapping N→n with n
Line ~220 Error: “mirror-symmetrical”
Correction: More standard: “mirror-symmetric.”
Lines ~221–224 Error: Comma splice in the description of the encoder and decoder; also “right hand side.”
Correction: Separate into two sentences or use a semicolon; use “right-hand side.”
Figure 3 caption Error: “Its layers mirror that of the encoder.”
Correction: “Its layers mirror those of the encoder.”
Lines ~225–226 Error: “the probabilistic AE finds an embedding of
distributions”
Technical issue: More precisely, the encoder represents each input by a parameterized latent distribution.
Line ~245 Error: Noisy model inputs are said to “correspond directly to noisy measurements.”
Technical issue: This is too literal unless the Gaussian perturbation reproduces the observational error model.
Correction: Say they “represent” or “simulate” noisy measurements.
Lines ~263–267 Issue: Definitions of “accuracy” and “precision” are nonstandard.
Correction: Make explicit that these are paper-specific definitions. In particular, “precision” here means sensitivity to perturbed inputs rather than the usual statistical meaning.
Line ~267 Error: “An over-fitted model is highly accurate, but tends to have poor precision.”
Technical issue: Overfitting can produce low training error, but does not imply high accuracy on unseen data.
Correction: Refer to high training performance/low training error rather than high accuracy.
Line ~267 Error: “…since it does not reflect any X outside of its training domain.”
Correction: Meaning is unclear. Likely: “…does not generalize to X outside its training domain.”
Figure 5 caption Error: “RMSE between between…”
Correction: Delete duplicated “between.”
Line ~344 Issue: “The precision of predicting unmodified inputs …
decreased…”
Technical issue: Under the paper’s stated definition, for an unperturbed input Y_delta = Y_p, precision should be zero. This may actually refer to accuracy/performance. Check terminology against Fig. 6.
Lines ~492–493 / corresponding discussion
Error: “the model’s epistemic uncertainty reflects the aleatoric uncertainty”
Technical issue: The terminology is confused. Measurement noise is generally aleatoric/data uncertainty; model/parameter uncertainty is epistemic.
Lines ~553–554 Error: “The trained models inherent the bias…”
Correction: “The trained models inherit the bias…”
Lines ~564–565 Issue: “A cyclic time coordinate … implies similarities between seasons that are not present.”
Technical issue: Cyclic encoding imposes continuity between the end and beginning of the cycle (e.g.,
December/January), not generally “similarities between seasons.”
Figure 8 caption Error: “design as showed in Fig. 6”
Correction: “design as shown in Fig. 6”
Figure 8 caption Error: “levels.Panel a)”
Correction: “levels. Panel a)”
Line ~582 Error: “expectedly leads…”
Correction: “As expected, including more training data leads…”
Lines ~603–604 Issue: If months with poor performance are removed, the resulting statistic is called “average annual performance.”
Technical issue: It is no longer an annual average if selected months are excluded.
Correction: Call it an average over the retained months.
Line ~470 Error: “certain months may have less performance, or more performance”
Correction: “poorer performance or better performance.”
Line ~482 Error: Double parentheses: “(e.g. (Volk et al…”
Correction: “(e.g., Volk et al., 1997; …)”
Line ~485 Error: “The general concept is described in Sect. 3.4.”
Correction: Incorrect cross-reference. The AE concept/architecture is described in Sect. 2.3.2.
Lines ~485–486 Error: “lower dimensional”
Correction:“lower-dimensional.”
Line ~486 Error: “without relying on any critical assumption on that input data”
Correction: “without relying on any critical assumptions about the input data.”
Lines ~491–492 Error: The manuscript says the supervised MLPs demonstrate that a 1-D encoding between the 6-D LLTs and 1-D AoA most
likely exists.
Technical issue: Successful regression from six variables to one target does not demonstrate the existence of an invertible 1-D latent representation sufficient to reconstruct the six tracers. Is that not correct?
Lines ~496–498 Issue: AE architectures are said to be shown in Figure 10.
Correction: Check the figure reference. Figure 3 already explicitly presents the AE architectures, so this may be an incorrect figure
number. See minor comment above.
Lines ~605–606 Error: Sentence ends “…CLaMS model AoA” without
punctuation.
Correction: Add period.
Line ~617 Error: “real life measurements”
Correction: “real-life measurements.”
Lines ~618–620 Error: The probabilistic MLP is described as providing “well-calibrated estimates of the epistemic uncertainty.”
Technical issue: Predictive variance from a Gaussian output does not by itself
establish calibrated epistemic uncertainty.
Lines ~637–641 Error: Inconsistent forms: “AoA-values,” “in-situ AoA,”
“model-AoA.”
Correction: Standardize, preferably “AoA values,” “in situ AoA,” and “model AoA.”
Line ~645 Error: “the here presented methodology”
Correction: “the methodology presented here.”
Line ~647 Error: “deployed in small scale”
Correction: “deployed on a small scale.”
Line ~650 Error: “Python packages Tensorflow”
Correction: “the publicly available Python package TensorFlow” (and TensorFlow Probability, if applicable).
Line ~665 Error: “National Oceanic and Atmospheric Administration (NOOA)”
Correction: “National Oceanic and Atmospheric Administration (NOAA).”
Lines ~671 Error: “Finally” is repeated in close succession.
Correction: Delete or change one occurrence.
Lines ~671–672 Error: “Tensorflow community”
Correction: “TensorFlow community.”
Lines ~671–672 Error: “library, which were essential”
Correction: “library, which was essential.”
Citation: https://doi.org/10.5194/egusphere-2026-4183-RC2
Data sets
Supplement for "On using neural networks to predict mean Age-of-Air from long-lived tracers" Jan Kaumanns, Florian Voet, Felix Plöger, Peter Preuße, Jörn Ungermann, and Michaela Imelda Hegglin https://doi.org/10.5281/zenodo.21333750
Viewed
| HTML | XML | Total | BibTeX | EndNote | |
|---|---|---|---|---|---|
| 190 | 105 | 39 | 334 | 33 | 29 |
- HTML: 190
- PDF: 105
- XML: 39
- Total: 334
- BibTeX: 33
- EndNote: 29
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
This paper discusses the use of neural networks to produce stratospheric mean age of air from a set of long-lived trace gases. I know very little about neural networks so I can’t evaluate if the methods used are the best suited for this task. But the basic techniques are well described and at a level that a non-expert can at least gain an understanding of how they are used. Neural networks are certainly a powerful data analysis tool and it’s great to see how they can be implemented in the context of AoA, at least for model output.
As the authors mention, the estimation of AoA from trace gas measurements is not fully constrained and involves certain assumptions and corrections to the data. While measurement uncertainty, or noise in the context of the neural network, is emphasized in this study, I think this is only one of a number of important aspects of how well we can calculate AoA. Ideally, the use of a neural network can provide an increased robustness to the conversion of multiple trace gas distributions to AoA distributions. I’m not necessarily convinced by this paper that the neural network technique can improve on existing AoA estimates from trace gases but it’s important to explore this possibility and hope to improve the calculations in the future.
I realize this paper is an initial discussion of the use of neural networks to calculate AoA from model output but I was hoping for some exploration of the possibilities and potential issues with the use of a more limited observational dataset. Sparse sampling, missing data or fewer different trace gases measured are examples of the complications involved with a measurement dataset. A new satellite data set is mentioned and a brief description of how to use in situ measurements near the tropopause to help calibrate the neural network, although that section wasn’t entirely clear to me.
In general, I support the publication of this paper with consideration of the minor comments listed below. The methods are described well and the results are clearly explained. Some of the figures could use better labels to help the reader as mentioned below.
Specific comments:
Line 387: change ‘inherent’ to ‘inherit’
Line 390: Do you mean ‘zonal’ here or ‘meridional’?
Fig. 8: The y-scale is too large to see differences between these plots. They also should be labeled differently so it’s easy to see what they represent rather than ‘Performances 2011’ on each one. For instance, 8a could be labeled ‘Trained with time dependence’, 8b ‘Trained with latitude as a feature’, etc.
Lines 436-7: Some missing parentheses and the e.g. statement is awkward.
Lines 438-9: I assume this is referring to the chemistry climate model predictions of AoA and the trends. I’m not sure about such a broad statement and what ‘relying on new observations’ means in this context. Mostly free running CCMs with greenhouse gas emission trends predict a decreasing trend of AoA. The emission trend could be called new observations but I don’t think that’s what’s referred to here.
Lines 440-50: I would expect some of this variability in performance to be due to the phase of the QBO relative to the seasonal cycle. The QBO has a significant impact on AoA and each cycle can be phased somewhat differently at different levels with the seasonal cycle. This could be checked fairly easily.
Line 567: ‘monotonic’ instead of ‘monotonous’
Line 589: ‘dependent’ instead of ‘depending’
Fig. 15: Again, some headers on the plots, in this case the year would be appropriate, would help the reader.
Line 612: ‘observationally-derived AoA’