the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
GenGHG v1.0: A generative machine learning emulator of greenhouse-gas atmospheric transport around global emission hotspots
Abstract. Dense greenhouse-gas (GHG) observations, particularly from satellites, require fast atmospheric transport models that link surface emissions to these observations. We introduce GenGHG, a generative machine learning (ML) model that emulates footprints from the Stochastic Time-Inverted Lagrangian Transport (STILT) model. These footprints estimate how a unit of emissions would alter a downwind atmospheric measurement. Unlike existing deterministic ML emulators, which give a single prediction, GenGHG predicts an ensemble of plausible footprints under given meteorological forcing conditions to represent stochastic atmospheric transport. GenGHG retains ML-level computational efficiency, with each member generated in less than 2 s on a single GPU, and we evaluate its generative advantage over deterministic ML across a benchmark of 60 urban areas worldwide. Results show that GenGHG preserves footprint total mass substantially better than deterministic ML, with a mean relative bias of +2.67 % compared with −22.38 %, and also better preserves the footprint-value distribution. Grid-level accuracy also shows that GenGHG better captures spatial footprint patterns, with larger gains in more dispersive transport regimes. We further test GenGHG’s advantage in a synthetic inverse modeling experiment, where transport from GenGHG yields methane emission estimates that closely follow the STILT transport reference and outperform deterministic ML. Finally, we highlight GenGHG’s compatibility with any global meteorology product at a 0.25° resolution. This flexibility, combined with its low computational cost, makes it straightforward to run footprints using multiple meteorology products and subsequently evaluate the possible effects of meteorological uncertainties.
- Preprint
(11501 KB) - Metadata XML
- BibTeX
- EndNote
Status: final response (author comments only)
- RC1: 'Comment on egusphere-2026-4304', Anonymous Referee #1, 16 Sep 2026
-
RC2: 'Comment on egusphere-2026-4304', Anonymous Referee #2, 03 Oct 2026
Wang et al. present a generative ML framework to emulate the STILT footprints. The framework generates an ensemble of footprints from different stochastic ML realizations, applies mass conservation screening to the STILT footprints, and performs an OSSE to evaluate if GenGHG could recover prescribed methane fluxes. It is interesting to see the model evaluated against DerGHG across different dispersive and terrain regimes, under cases with/without STILT mass violations, and within an OSSE for flux inversion. GenGHG generally shows improved performance relative to DeterGHG across these cases, demonstrating the potential value of the generative framework.
My major comment lies in the lack of evaluation of temporally resolved footprints. The current evaluation appears to focus primarily on time-integrated footprints. Have the authors considered evaluating the ability of GenGHG to reproduce the temporal evolution of footprints, for example at hourly resolution? Temporally resolved footprints can be particularly important for applications involving rapidly varying or bidirectional surface fluxes, such as biogenic CO₂ exchange, where daytime uptake and nighttime respiration have opposite signs. Evaluating whether GenGHG preserves this temporal structure would help establish its applicability to a broader class of inverse problems.
A more general suggestion is to provide a clearer discussion of the use cases or ranges of applicability for GenGHG, which would be useful for users. In particular, it would be helpful to discuss the spatial and temporal scales over which GenGHG is expected to remain reliable. The current applications appear relatively near-field and are especially relevant to urban or source-region studies. Within what transport distances, footprint ages, or domain sizes would the authors recommend using GenGHG, perhaps based on the behavior shown in Figure 3de (potentially degraded far-field performance?)? Would the current framework also be suitable for regional or global inversions, or would additional training data and/or model development be required? If aiming for regional and global scales, it would be useful to discuss how the framework would handle background concentrations and trajectory endpoints. Defining the intended range of applicability would help readers assess whether GenGHG is appropriate for their applications.
Minor comments
L115-L125: Some of the terminology used in the flow-matching description may be difficult to interpret for atmospheric science readers. In particular, terms such as “target velocity” and “prior samples” could be defined more explicitly. For example, does “velocity” around Lines 116–117 refer to the velocity in the generative/data space, rather than atmospheric wind velocity? A short intuitive explanation would help avoid confusion.
L244-L247: I did not quite follow these two explanations for the GenGHG vs. DerGHG differences. What do the authors mean by “biased structural collapse after transformation…”.
L255-257: I may miss the details about the mass violation tests - my understanding is that GenGHG was trained using STILT footprints that already pass the dmas screening (or not?). If so, it seems a bit weird and unfair to compare the ML results with it would be helpful to add the screened STILT footprint line in Figure 3c next to the dashed lines for the badly-fail ones. Moreover, the L255-256 sort of suggests that GenGHG was able to learn this mass violation from the unfiltered total footprints. I would suggest clarifying the training data and avoid over-stating the performance of their model if I interpreted correctly.
L263-265: The degraded performance for highly dispersive footprints is interesting. Could this partly reflect limitations of the assumed Gaussian prior distribution used to initialize the generative process, particularly when the target footprint is very broad or multimodal? Some discussion of whether the prior distribution contributes to the difficulty in these cases would be useful.
Sect. 3.3: I am curious about the effects of dmas screening under different meteorological inputs. Does the dmas screening result in more or fewer usable footprints relative to the GFS?
The manuscript seems to miss several citation/references, e.g., none for FLUXPART and NAME whenever mentioned.
Citation: https://doi.org/10.5194/egusphere-2026-4304-RC2
Viewed
| HTML | XML | Total | BibTeX | EndNote | |
|---|---|---|---|---|---|
| 190 | 168 | 23 | 381 | 17 | 16 |
- HTML: 190
- PDF: 168
- XML: 23
- Total: 381
- BibTeX: 17
- EndNote: 16
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
The article describes an innovative method of emulating LPDM footprints, leveraging flow-matching based generative models that have been shown successful in meteorological and other applications. I believe this paper is of excellent quality, and almost ready for publication.
The central idea is well motivated, proposing a generative model that builds on documented limitations and strategies of existing deterministic emulators. Accelerating LPDM emulation is a timely task to leverage the large amounts of satellite measurements. The results are evaluated through a nice range of metrics, and compared against the outputs of a deterministic version of the model as a benchmark. The dataset is also a contribution in its own right, given the large amount of locations and dates included. The synthetic out-of-sample inversion experiments illustrate well the downstream application, and the cross-product comparison is a useful demonstration of the potential of a fast emulator where the full physics runs are prohibitive.
There are three main areas that I believe need a bit more development:
Other questions the authors might consider addressing:
Typos and smaller comments