the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
Implementation and evaluation of the lognormal prior probability distribution in a variational atmospheric inversion framework
Abstract. In this study, we investigate the use of a lognormal prior probability distribution in atmospheric inverse modelling. We present the formal implementation in a variational inversion framework and analyze how the choice of statistical optimization parameter (mean, median, or mode) affects the inversion outcome. Using a case study of inverse modelling of sulfur hexafluoride (SF6) in Europe, we evaluate the performance of the lognormal implementation through both synthetic and real data experiments, and compare the results to inversions using a normal prior probability distribution. We estimate the posterior uncertainties using a Monte Carlo approach and examine their distribution.
We find that optimizing for the mean or the mode can produce improved emission estimates under the condition of a strong observational constraint, however, this can lead to unstable and strongly biased inversion results under a weak constraint. In contrast, optimizing for the median consistently improves emission estimates and leads to physically plausible results across all tested cases, providing the most reliable option.
We show that inversions using a lognormal prior distribution produce a similar posterior emission pattern as when using a normal prior distribution, however, avoid non-physical negative emission values and occasionally allow for stronger positive emission adjustments. Posterior uncertainties can be estimated using interpercentile ranges from an ensemble of inversions with prior emission errors following a lognormal distribution. Due to the strong asymmetry of posterior distributions with respect to the sign of the inversion increments, error reduction is better assessed in log space, where it provides a clearer measure of the constraints imposed by the observations.
Competing interests: At least one of the (co-)authors is a member of the editorial board of Geoscientific Model Development.
Publisher's note: Copernicus Publications remains neutral with regard to jurisdictional claims made in the text, published maps, institutional affiliations, or any other geographical representation in this paper. While Copernicus Publications makes every effort to include appropriate place names, the final responsibility lies with the authors. Views expressed in the text are those of the authors and do not necessarily reflect the views of the publisher.- Preprint
(21626 KB) - Metadata XML
- BibTeX
- EndNote
Status: final response (author comments only)
-
RC1: 'Comment on egusphere-2026-2125', Anonymous Referee #1, 14 Jun 2026
-
AC1: 'Reply on RC1', Martin Vojta, 28 Aug 2026
We thank Reviewer #1 for the assessment and comments. In the following we reply in italic fonts while new or changed text is in bold fonts.
The reviewer's main concern is the novelty of the manuscript, given that the concept of the lognormal distribution is, of course, not new, and that it has already been used in previous inversion studies. We agree that the properties of a lognormal distribution are already well known and that they can be found in textbook material from a statistics course, as Reviewer #1 rightly points out. However, we argue that the novelty of this work does not lie in the mathematical properties of the distribution itself, but in systematically investigating the implications of adopting a lognormal prior in atmospheric inverse modeling, by presenting the formal implementation and analysing the consequences, practical challenges, and potential pitfalls that arise when the distribution is used in a real inversion framework. While several of the inversion results can indeed be understood from the mathematical properties of the lognormal distribution, we believe it is important to demonstrate how these properties manifest themselves in a realistic atmospheric inversion framework, rather than interpreting this consistency as evidence of a lack of novelty. We therefore do not believe that the fact that some findings are consistent with theoretical expectations diminishes their scientific value nor should this discourage us from investigating practical implications.
In the following, we further elaborate on these points and explain why we believe this work provides interesting insights for the inversion community. However, we agree that some of the bullet points may not have sufficiently conveyed our intended message and we therefore revise the conclusion section of the manuscript to meet the comments of the reviewer.
- Formal implementation (1):
We provide the reader with all the key expressions required to implement a lognormal prior probability distribution in a variational inversion framework for the different statistical optimization parameters (mean, median, and mode). While (as stated) the underlying cost functions were originally derived by Fletcher and Zupanski (2006) in the context of data assimilation, our manuscript provides a complete set of equations needed for their direct implementation in variational atmospheric inversion frameworks. We believe this lowers the barrier for other groups wishing to adopt and evaluate lognormal priors in atmospheric inverse modeling. We also describe how prior uncertainties can be specified in log space.
- Investigating different statistical optimization parameters (2):
As stated in the manuscript, previous atmospheric inversion studies have already employed lognormal prior distributions. However, to the best of our knowledge, all of these studies effectively minimized the cost function corresponding to the median (as also noted by the reviewer). In the majority of cases, the optimized statistic is not even explicitly acknowledged, and, to our knowledge, none of these studies provides a justification for why the median should be the preferred optimization parameter or even acknowledges that alternative choices exist (please note that Fletcher and Zupanski originally identified the mode as the preferred choice). To our knowledge, this is the first inversion study that systematically investigates optimization with respect to the mean, median, and mode and compares the resulting inversion behaviour. This comparison demonstrates why optimization for the median provides the safest choice for inverse modelling, while we tried to show in detail that optimizing for the mean or mode can lead to problems and spurious inversion results under certain conditions. We believe this is an important methodological contribution, as it provides practical guidance for future studies adopting lognormal priors and can help other researchers to avoid potential pitfalls.
- Comparison between normal and lognormal prior distribution (3)
While both normal and lognormal prior distributions have been employed in previous inversion studies, to the best of our knowledge, this is the first study to systematically compare the inversion results obtained with the two different prior formulations. We believe this comparison is valuable because it provides a quantitative assessment of how the choice of prior distribution influences inversion results. We agree with the reviewer, that the non-negativity of the control variables follows directly from the mathematical definition of the distribution, however, we believe it is nevertheless worthwhile to demonstrate this advantage and more importantly, to assess how the different theoretical properties of the distributions influence inversion results in practice.
- Investigation of the posterior distribution and error reduction (4):
To the best of our knowledge, this is the first atmospheric inversion study to systematically investigate the shape of the posterior distribution and the associated error reduction when employing a lognormal prior distribution. We believe that neither the form of the posterior distribution nor the behaviour of the uncertainty reduction follows trivially from the well-known properties of the lognormal distribution, and therefore merits investigation. Our posterior distribution results from the combination of a Gaussian likelihood and a lognormal prior, giving rise to a considerably more complex behaviour. For example, one might expect the posterior distribution to become approximately Gaussian as the observational constraint becomes stronger due to the Gaussian likelihood. However, the lognormal prior simultaneously acts to retain a certain degree of skewness, the magnitude of which is hard to predict and depends on the relative influence of the observations and the prior. As demonstrated for our inversion application, posterior distributions become generally closer to Gaussian however can retain a pronounced skewness, such that the commonly used error reduction evaluated in linear space can no longer be interpreted as a meaningful measure of information content. Had the posterior been close to Gaussian throughout the inversion, this issue would largely disappear. We further consider the closer examination of error reduction in log space to provide useful insights.
Conclusion/Comments:
Emissions being lognormally distributed has been known for decades. This is unsurprising. The assumption that the associated uncertainties (or errors) are therefore lognormally distributed is often assumed. This has been done by lots of previous work. I do not find anything compelling one way or the other in Votja et al to change what people are already doing regarding the assumption of lognormality of errors.
We agree that the generally lognormal shape of emission distributions is already known (as also stated in the introduction), and that the illustrated shape of SF6 emission distributions is primarily an illustrative example rather than a novel finding (although, to our knowledge, the distribution of gridded European SF6 emissions has also not previously been explicitly shown for both EDGARv8 and GAINS inventories). Nevertheless, we believe that explicitly illustrating the strongly skewed distribution of gridded prior emissions remains useful. In inversion studies, the distribution of the prior values is almost never shown, and the substantial departure from a Gaussian distribution may therefore not be immediately apparent, particularly to a broader readership. We have therefore revised the text to make this clearer:
Gridded emission inventories, which are often used as prior information in inversions, can exhibit strongly skewed, approximately lognormal distributions. We use European SF6 emissions as an illustrative example to demonstrate the pronounced skewness that can arise in gridded emission inventories. For both the EDGARv8 and GAINS inventories, the emission distributions contain a large number of small values and very few large values, resulting in a substantial departure from a Gaussian distribution. In such cases, it is likely that the associated uncertainties are also better represented by a lognormal rather than a normal distribution.
The fact that the mean, median, and mode are distinct for a lognormal distribution is textbook material that is covered in basic statistics.
This is, of course, true, and we agree that this point does not warrant a separate bullet point. We initially included it to highlight an aspect that, although textbook material, is rarely explicitly acknowledged in previous inversion studies employing lognormal priors. As it is not a specific outcome of this work, we have removed it as a standalone point and incorporated it as the opening sentence of the next point.
For a lognormal distribution mean, median, and mode do not fall together making it necessary to choose explicitly which of these parameters to optimize for in the inversion. Building on the cost functions derived by Fletcher and Zupanski (2006) for these different statistics, we provide the complete set of equations required for their direct implementation in atmospheric inversion frameworks.
This material is covered in textbooks on inverse modeling. The median of the posterior distribution is typically used when lognormal distributions are used for either the P(x) or P(y|x).
Please see general comment (2)! We have consolidated the bullet points on optimizing the mode, median, and mean into a single bullet point.
Optimizing for the mode and the mean can yield improved emission estimates when the observational constraints are strong. However, under weak observational constraints, these choices can lead to unstable and strongly biased inversion results caused by a decrease in the cost in the state space due to the extra logarithm term in the cost functions. In contrast, optimizing for the median consistently improved prior emissions across all tested cases, and thus provides the most robust estimator to optimize for.
This seems entirely unsurprising. A lognormal is by definition going to yield positive constraints and it has a long tail. Again, not a new finding
Please see general comment (3)
This seems unsurprising. Doing perturbations in log-space is appropriate if one has log-normal distributions
Please see general comment (4): We revised the bullet point accordingly!
Posterior uncertainties can be estimated using a Monte Carlo ensemble of inversions, each with prior emissions perturbed by lognormally distributed errors. Although the posterior distribution generally becomes more Gaussian than the prior, pronounced skewness is retained in many grid cells, with its extent strongly dependent on the sign of the inversion increment. Negative increments typically produce narrow, nearly Gaussian posteriors, whereas positive increments tend to yield broader and more strongly skewed, lognormal-like distributions. To better account for this retaining asymmetry, we recommend characterizing posterior uncertainty using interpercentile ranges derived from the ensemble, rather than the standard deviation.
We also revised the last sentence of the abstract to:
Although the posterior distributions generally become more Gaussian than the priors, substantial skewness can remain, with their asymmetry depending on the sign of the inversion increment, making error reduction in log space a more meaningful measure of observational constraints than in linear space.
Fairly textbook. The information content is going to be skewed if you are working in log-space. The information content the authors are considering is formulated from Gaussians, so things are skewed in log-space.
Please see general comment (4). We revised the bullet point to:
The error reduction in linear space is highly asymmetric with respect to the sign of the inversion increment, limiting its interpretation as a meaningful measure of information content. We therefore recommend assessing error reduction in log space, where it is primarily governed by emissions sensitivity and exhibits substantially weaker asymmetries, making it better suited to reflect the strength of the observational constraints.
Citation: https://doi.org/10.5194/egusphere-2026-2125-AC1
-
AC1: 'Reply on RC1', Martin Vojta, 28 Aug 2026
-
RC2: 'Comment on egusphere-2026-2125', Anonymous Referee #2, 01 Jul 2026
Review of egusphere-2026-2125:
Implementation and evaluation of the lognormal prior probability distribution in a variational atmospheric inversion framework
General comments
This paper presents a study on the use of lognormal probability distributions as part of a variational inversion system. The derivations for cost functions using lognormal prior PDFs and the mean, median or mode of the posterior solution are clearly presented and the impact of this choice on posterior emissions is tested using synthetic data tests with different prior and observational uncertainties. The impact of using lognormal and normal prior PDFs on posterior emissions and posterior emission uncertainties is tested using one year of SF6 observations to estimate SF6 emissions across Europe. An ensemble of inversions is used with Monte Carlo analysis to determine the posterior emission uncertainties.
As noted in the paper, the use of lognormal PDFs in atmospheric inversions is not a novel concept, but this investigation of the impact of a lognormal distribution on the inversion results could act as a useful reference for future inversion studies, and the method and experiments are explained clearly and in enough detail to enable other researchers to reproduce the method.
The impact of each posterior option (mean, mode, median) on the synthetic data inversions are presented clearly in combined figures, which allows for easy interpretation of the results.
Overall, despite the limited new science presented here, I think this manuscript meets the scope of GMD by presenting a summary of the technical aspects involved with choosing prior PDFs for inversions, which could act as a useful reference for future inversion studies. Therefore, I think this manuscript could be accepted for publication, after the minor comments below have been addressed.
Specific comments
Posterior emissions are analysed across the whole of Europe, despite the authors noting that emission sensitivity from the observations is poor for most of eastern and southern Europe. To address this, a more detailed comment should be included about the uncertainty of results beyond the region of good sensitivity. Or, results should only be analysed and presented over a smaller region that does have good sensitivity to emissions. Showing spatial emission results over a smaller region would also make all of the spatial emission figures easier to read.
Some topics that would be unfamiliar to non-experts (such as using atmospheric models to link surface emissions to atmospheric concentrations) are treated as assumed knowledge, which is a reasonable choice considering the scope and target audience of this paper. If the authors want to make this paper relevant to a wider audience, they would need to introduce some terms (such as atmospheric transport) in more detail, but I will leave this to the authors discretion.
Technical corrections and questions
Line 13: missing word after “however”? Should this read “…however, they avoid…”
Line 19: this line may be clearer if “based on its measurement in the atmosphere” is replaced with “based on measurements of its concentration in the atmosphere”.
Line 23: The repeat of “(or mole fractions)” could be removed, as this is already included on line 22.
Line 24: replace “determining” with “determine”.
Line 27: what does “variation, especially due to climate” mean here? Spatial or temporal variation or trends in emissions?
Line 37: “and prior estimate of x” should be added after “observations, y”.
Line 49: is there a reference for the line “The prior emission estimates for many atmospheric species have a lognormal distribution, strongly suggesting that the errors follow a lognormal distribution as well”?
Line 51: “cost function” has not been introduced yet. A line could be added earlier in this section, describing how a cost function is used to find a most likely estimate of x.
Line 79: add a line here, explaining why the PDF needs to be expressed in terms of the error, otherwise the reader has to wait for the presentation of the cost function for this explanation.
Line 81: add brackets to these equations (e.g. ln(ε) = ln(x) – ln(xb)) to maintain consistency with where this equation is used again later (e.g. line 101).
Line 108: I think the footnote here could be moved to after equation 13, as this is useful information and it would not break the flow of the text here. However, this is just a personal preference, so I’ll leave this to the authors’ discretion.
Line 115: some discussion of the choice of mean/median/mode in relation to atmospheric inversions could be included here. How does this choice differ compare to the data assimilation context in the quoted references?
Line 125: ‘transport function H(x)” is referenced here, when H has only been referred to as an “operator” before this point. To resolve this, more detail on H could be added at line 90, explaining how this could contain output from a Lagrangian or Eulerian transport model.
Line 127: reference or derivation for this equation? Other equations have been derived in detail, (e.g. with clear references to what substitutions and transformations are being used) but there is less detail on the derivation of Equation 15 from Equation 14.
Line 132: why does using this substitution improve the convergence rate?
Line 158: some more explanation of how this equation was derived would be helpful.
Line 161: you could add a real-world example of what this correlation means, presumably spatial correlation between emissions from different grid cells?
Line 165: add a reference for the statement “ land sinks are negligibly small”.
Lines 166 and 167: add the full names of the EDGAR and GAINS databases.
Line 169: for reproducibility, it might be useful to explain how you produced the normal and lognormal fit lines, discussed at this point in the text and shown in Figure 1.
Line 183: the footnote here could be included in the caption for Figure 2 instead.
Line 185: when reading this section, I wondered what observational uncertainty you were using, but this is covered later in sections 3.3 and 3.4. So it might be helpful to add a line here stating where observational error is defined.
Line 193: add a reference for the statement “SF6 is almost inert up to the middle stratosphere”.
Line 233: did you use any methods to test for convergence, in either the synthetic or real data tests?
Line 243: you could add some information on whether these prior and observation uncertainties are realistic, compared to the uncertainties ‘traditionally’ placed on these terms in inversions. Where does model error (e.g. uncertainties from the transport model) fit into these definitions of uncertainty?
Line 253: what inversion period are you using? Just one inversion over the whole of 2020, or multiple shorter inversion periods across 2020?
Line 254: EDGAR’s distribution seems to be closer to a lognormal, so why was GAINS chosen as the prior for these tests, instead of EDGAR?
Line 257: why was this observation error of 0.09 ppt chosen? And is it a reasonable assumption to apply this error value to all observations from all sites? I know this choice is currently under discussion in some inverse modelling groups, so I don’t know the ‘correct’ answer for this myself, but it might be useful to comment on why these observational uncertainties were chosen for this work in particular.
Line 273: “Synthetic data experiment setup” is confusing as a subheading under the “Results” heading, should “setup” be removed?
Line 313: missing space in “Fig.5j”.
Line 320: you note that the results are influenced by the specific random realization. Did you try rerunning these tests with different random perturbations, to see how the results varied?
Line 360: Does this claim that negative emission increments favours error reduction apply in all areas or just these 4 cases and in the two representative grid cells discussed in the next paragraph? Did you investigate this for all grid cells?
Figures 9 and 10: these figures may be easier to interpret if fewer histogram bins are used.
Line 476: should “hx” be “H(x)” here?
Citation: https://doi.org/10.5194/egusphere-2026-2125-RC2 -
AC2: 'Reply on RC2', Martin Vojta, 28 Aug 2026
We thank Reviewer #2 for the very constructive and detailed feedback. We included almost all the suggestions which substantially improved our manuscript! In the following we reply in italic fonts while new or changed text is in bold fonts. We also append a pdf version for better readabilty.
Specific comments
Posterior emissions are analysed across the whole of Europe, despite the authors noting that emission sensitivity from the observations is poor for most of eastern and southern Europe. To address this, a more detailed comment should be included about the uncertainty of results beyond the region of good sensitivity. Or, results should only be analysed and presented over a smaller region that does have good sensitivity to emissions. Showing spatial emission results over a smaller region would also make all of the spatial emission figures easier to read.
Yes, we agree! We initially chose the full European domain because it includes both regions that are well constrained by the observations and regions with poor observational coverage. In eastern and southern Europe, where the emission sensitivity is low, the improvements through the inversion is limited and the posterior emission are more uncertain. This is, for instance, clearly to see in our synthetic experiments, where perturbations in the prior can almost not be corrected in these regions. For the purpose of this study, we wanted to investigate the behaviour of the algorithm under both strong and weak observational constraint as limited observational coverage is a common challenge in atmospheric inversions. In particular, we wanted to study the transition from the lognormal prior towards a more normally distributed posterior, and its dependence on emission sensitivity. As illustrated in Fig. 8, this transition occurs predominantly in well-constrained regions, which is consistent with our expectation that normally distributed observational uncertainties exert a relatively stronger influence on the shape of the posterior distribution where the observational constraint is larger.
To make this clear we added:
Emission sensitivities are substantially lower in eastern and southern Europe, and thus, the improvements achieved by the inversion in these regions are limited and posterior emission estimates are more uncertain.
In contrast, in regions with weaker observational constraints, the transition is weak or absent, so that the posterior largely retains the lognormal shape of the prior.
Some topics that would be unfamiliar to non-experts (such as using atmospheric models to link surface emissions to atmospheric concentrations) are treated as assumed knowledge, which is a reasonable choice considering the scope and target audience of this paper. If the authors want to make this paper relevant to a wider audience, they would need to introduce some terms (such as atmospheric transport) in more detail, but I will leave this to the authors discretion.
Thank you for this comment. We agree that some topics are treated as assumed knowledge. As the topic is rather specific, the primary target audience of this study is the atmospheric inversion community, particularly researchers interested in implementing and using lognormal priors in atmospheric inversions. We therefore feel that a more detailed introduction to these established concepts would go beyond the scope of the paper, and we would prefer to retain the current level of detail. We nevertheless appreciate the suggestion and agree that a more extensive introduction could be useful for a broader audience.
Technical corrections and questions
Line 13: missing word after “however”? Should this read “…however, they avoid…”
Yes, thank you - >done
Line 19: this line may be clearer if “based on its measurement in the atmosphere” is replaced with “based on measurements of its concentration in the atmosphere”.
Yes, agreed - >done
Line 23: The repeat of “(or mole fractions)” could be removed, as this is already included on line 22.
Yes, agreed, thank you - >done
Line 24: replace “determining” with “determine”.
Yes - >done
Line 27: what does “variation, especially due to climate” mean here? Spatial or temporal variation or trends in emissions?
We included “spatiotemporal“
Line 37: “and prior estimate of x” should be added after “observations, y”.
done
Line 49: is there a reference for the line “The prior emission estimates for many atmospheric species have a lognormal distribution, strongly suggesting that the errors follow a lognormal distribution as well”?
Yes, we cited Super et al. 2020 which seems appropriate, who studied uncertainties of emission inventories, suggesting lognormal distributed errors.
Line 51: “cost function” has not been introduced yet. A line could be added earlier in this section, describing how a cost function is used to find a most likely estimate of x.
Yes, that is a great idea: We added:
“The maximum of the posterior distribution P (x | y), corresponding to the most probable solution, can be obtained by minimizing a cost function J given by the negative logarithm of the posterior probability”
after introducing Bayes’ Theorem.
Line 79: add a line here, explaining why the PDF needs to be expressed in terms of the error, otherwise the reader has to wait for the presentation of the cost function for this explanation.
Yes, that is a good idea, thank you. We added:
“The probability distribution is expressed in terms of the errors ε, as the uncertainty is associated with deviations from the prior estimate. For the lognormal distribution these errors are geometric, meaning……”
Line 81: add brackets to these equations (e.g. ln(ε) = ln(x) – ln(xb)) to maintain consistency with where this equation is used again later (e.g. line 101).
Yes, done!
Line 108: I think the footnote here could be moved to after equation 13, as this is useful information and it would not break the flow of the text here. However, this is just a personal preference, so I’ll leave this to the authors’ discretion.
Thank you, we agree and changed the text accordingly.
Line 115: some discussion of the choice of mean/median/mode in relation to atmospheric inversions could be included here. How does this choice differ compare to the data assimilation context in the quoted references?
Yes, that’s an excellent idea! We added:
While these results were derived in the context of data assimilation, the choice of estimator is also relevant to atmospheric inversions, where xb represents a prior emission estimate rather than a model forecast and the observational constraint can vary substantially across the domain. The performance of the mean, median, and mode may therefore depend on both the accuracy of the prior emission estimate and the strength of the observational constraint.
which also bridges to our synthetic tests, using different uncertainties!
Line 125: ‘transport function H(x)” is referenced here, when H has only been referred to as an “operator” before this point. To resolve this, more detail on H could be added at line 90, explaining how this could contain output from a Lagrangian or Eulerian transport model.
Agreed, we added/changed:
….. where H is an operator that maps the state to the observation space. For atmospheric inversions, H represents the forward model that relates the emissions to the observed concentrations, using an atmospheric transport model, such as a Lagrangian particle dispersion model (LPDM) or a Eulerian transport model.
Line 127: reference or derivation for this equation? Other equations have been derived in detail, (e.g. with clear references to what substitutions and transformations are being used) but there is less detail on the derivation of Equation 15 from Equation 14.
Yes, we add more details to the derivation (see pdf)
Line 132: why does using this substitution improve the convergence rate?
The substitution acts as a preconditioning of the optimization problem by transforming the prior covariance into the identity matrix. This makes the quadratic part of the prior curvature equal in all transformed directions (I think of it as geometrically turning an elongated prior ellipse into a circle), thereby improving the conditioning of the cost function, generally accelerating the convergence.
Line 158: some more explanation of how this equation was derived would be helpful. &
Line 161: you could add a real-world example of what this correlation means, presumably spatial correlation between emissions from different grid cells?
Yes, thank you, we agree with both points and revised the section to be more detailed. (see pdf)
Line 165: add a reference for the statement “ land sinks are negligibly small”.
Yes, we cited Harnisch and Eisenhauer, 1998!
Lines 166 and 167: add the full names of the EDGAR and GAINS databases.
Done ->
Figure 1 represents the a priori SF6 flux distribution in Europe provided by the well-established bottom-up inventories from the Emissions Database for Global Atmospheric Research (EDGAR version 8; EDGAR, 2023; Crippa et al., 2023) and the Greenhouse Gas and Air Pollution Interactions and Synergies (GAINS) model (Purohit and Höglund-Isaksson, 2017), regridded to a spatial resolution of 0.25°.
Line 169: for reproducibility, it might be useful to explain how you produced the normal and lognormal fit lines, discussed at this point in the text and shown in Figure 1.
Yes, thank you , we add:
The orange dashed lines show the maximum-likelihood fits of the respective distributions, with the resulting probability density functions scaled to match the histogram counts.
to the capture of Figure 1.
Line 183: the footnote here could be included in the caption for Figure 2 instead.
Agreed, done!
Line 185: when reading this section, I wondered what observational uncertainty you were using, but this is covered later in sections 3.3 and 3.4. So it might be helpful to add a line here stating where observational error is defined.
Alright, we added:
The observational uncertainties are defined in Sects. 3.3 and 3.4.
Line 193: add a reference for the statement “SF6 is almost inert up to the middle stratosphere”.
Done -> we cited Kovács et al., 2017
Line 233: did you use any methods to test for convergence, in either the synthetic or real data tests?
We assessed convergence by monitoring the evolution of the cost function. We initially set the maximum number of iterations to 20, which is often sufficient, but observed considerable changes in the cost function and therefore increased the limit to 40 iterations. At 40 iterations, the cost function changed only marginally. For the real-data experiments, we nevertheless increased the maximum number of iterations to 100 as a precaution, to ensure that the differences investigated in our analysis were not caused by insufficient convergence.
Line 243: you could add some information on whether these prior and observation uncertainties are realistic, compared to the uncertainties ‘traditionally’ placed on these terms in inversions. Where does model error (e.g. uncertainties from the transport model) fit into these definitions of uncertainty?
&
Line 257: why was this observation error of 0.09 ppt chosen? And is it a reasonable assumption to apply this error value to all observations from all sites? I know this choice is currently under discussion in some inverse modelling groups, so I don’t know the ‘correct’ answer for this myself, but it might be useful to comment on why these observational uncertainties were chosen for this work in particular.
Yes, we agree. The definition of uncertainties (and particularly the role of model error) is indeed an active area of research, and different choices are still being discussed within the inversion community. We are happy to discuss this, as it is also an aspect that we would like to investigate further in future work.
In general, allowing the observational error to vary in space and time based on model information or observational variability certainly has the potential to improve inversion results, as demonstrated by recent studies such as Steiner et al. (2024). Different inversion groups have proposed different approaches for estimating such errors. However, given the complexity of the problem, a simple estimate may in some cases be preferable, particularly when the aim is to isolate and investigate the general behaviour of an inversion system.
In our case, the main purpose of considering a broad range of uncertainties was to test the algorithm under different levels of observational constraint. For this purpose, we considered a single uncertainty value to be the most appropriate. The range of uncertainties tested in this study is largely within the range of values traditionally used or explored in previous inversion studies. To clearly demonstrate the artefacts that can arise under weak observational constraints, however, some combinations of relatively small prior uncertainties and large observational uncertainties result in particularly weak observational constraints and thus rather represent the lower end of the range considered. Similarly, somewhat larger prior uncertainties could have been explored (looking at the Gain metric), but an important objective was to show the behaviour under weak observational constraints. For the real-data inversion, the selected observational uncertainty was primarily guided by tests in previous studies (e.g., Vojta et al., 2025). We, however, also tested other uncertainty values, although these tests were limited to 40 iterations, and found no substantial differences in the general results.
We added/changed:
To evaluate the performance of the implementation of the lognormal distribution, we perform multiple inversions testing various prior uncertainties of the prior state variables (0.18, 0,34, 0.47, 0.59, 0.69) and observation uncertainties (0.02, 0.04, 0.06, 0.08, 0.1 ppt). These values largely fall within the range explored in previous inversion studies (e.g. Vojta et al., 2025) and allow us to investigate the inversion behaviour under different levels of observational constraint.
Line 253: what inversion period are you using? Just one inversion over the whole of 2020, or multiple shorter inversion periods across 2020?
We performed, one inversion over the whole of 2020 - we included:
For all performed inversions, the temporal interval was set to one year.
in Section 3.2.5
Line 254: EDGAR’s distribution seems to be closer to a lognormal, so why was GAINS chosen as the prior for these tests, instead of EDGAR?
Indeed, looking at the lognormal fit, EDGARv8 appears somewhat closer to the underlying distribution. On the other hand, GAINS provides a broad range of emission values with fewer extreme high-emission values, which motivated our initial choice of GAINS for our experiments. We do not, however, claim that there is a clear criterion for preferring one inventory over the other.
Line 273: “Synthetic data experiment setup” is confusing as a subheading under the “Results” heading, should “setup” be removed?
Yes, indeed, thank you !
Line 313: missing space in “Fig.5j”.
Yes, thank you – done!
Line 320: you note that the results are influenced by the specific random realization. Did you try rerunning these tests with different random perturbations, to see how the results varied?
Yes, we also tested a few additional random realizations (although not in the same detail) to ensure that the observed behaviour was not specific to a single realization. While the specific values of the Gain metrics were sensitive to the realization, mainly depending on whether larger perturbations occurred in regions that were better or less well constrained by the observations, the overall behaviour and the artefacts observed under weak observational constraints remained consistent. There were also realizations in which these artefacts were less apparent. however, the main purpose of these tests was to raise awareness that such artefacts can arise from the additional terms in the cost function associated with the mean and mode.
Line 360: Does this claim that negative emission increments favours error reduction apply in all areas or just these 4 cases and in the two representative grid cells discussed in the next paragraph? Did you investigate this for all grid cells?
Yes, yes, we also investigated other grid cells and found that this behaviour generally applies across the domain, although not always in the same magnitude. I think this can be best seen when comparing Figure 11a to 11c, where the pattern of small or even negative error reduction is largely consistent with regions of positive increments, and vice versa.
Figures 9 and 10: these figures may be easier to interpret if fewer histogram bins are used.
Thank you for this suggestion. We tested different numbers of histogram bins, but unfortunately reducing the number of bins did not substantially improve the readability of the figures. We therefore retained the original binning
Line 476: should “hx” be “H(x)” here?
No, in this case we intentionally used hx to denote a linear transport function in one dimension. To make this clear, we added:
To be consistent with the approximation around xMAP let’s consider now the cost function for the mode (see Eq. 14) in one dimension, assuming a linear transport function h
References:
Harnisch, J. and Eisenhauer, A.: Natural CF4 and SF6 on Earth, 25, 2401–2404, https://doi.org/10.1029/98GL01779, 1998
Kovács, T., Feng, W., Totterdill, A., Plane, J. M. C., Dhomse, S., Gómez-Martín, J. C., Stiller, G. P., Haenel, F. J., Smith, C., Forster, P. M., García, R. R., Marsh, D. R., and Chipperfield, M. P.: Determination of the atmospheric lifetime and global warming potential of sulfur hexafluoride using a three-dimensional model, Atmospheric Chemistry and Physics, 17, 883–898, https://doi.org/10.5194/acp-17-883-2017, 2017.
Super, I., Dellaert, S. N. C., Visschedijk, A. J. H., and Denier Van Der Gon, H. A. C.: Uncertainty Analysis of a European High-Resolution Emission Inventory of CO2 and CO to Support Inverse Modelling and Network Design, 20, 1795–1816, https://doi.org/10.5194/acp-20-1795-2020, 2020.
Steiner, M., Cantarello, L., Henne, S., and Brunner, D.: Flow-dependent observation errors for greenhouse gas inversions in an ensemble Kalman smoother, Atmos. Chem. Phys., 24, 12447–12463, https://doi.org/10.5194/acp-24-12447-2024, 2024.
Vojta, M., Plach, A., Thompson, R. L., Purohit, P., Stanley, K., O’Doherty, S., Young, D., Pitt, J., Arduini, J., Lan, X., et al.: Quantifying European SF6 emissions from 2005 to 2021 using a large inversion ensemble, Atmospheric Chemistry and Physics, 25, 15 197–15 243, https://doi.org/10.5194/acp-25-15197-2025, 2025
-
AC2: 'Reply on RC2', Martin Vojta, 28 Aug 2026
Viewed
| HTML | XML | Total | BibTeX | EndNote | |
|---|---|---|---|---|---|
| 80 | 49 | 16 | 145 | 16 | 15 |
- HTML: 80
- PDF: 49
- XML: 16
- Total: 145
- BibTeX: 16
- EndNote: 15
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
Attached.