Accelerating greenhouse gas retrievals with neural network-based forward models
Abstract. Greenhouse gas (GHG) retrievals rely on repeated evaluations of computationally expensive physics-based forward models, which limit the feasibility of near-real-time retrievals and timely detection of emission hotspots. A promising alternative is to replace these forward models with fast machine learning emulators trained to approximate their input-output mapping, while retaining the overall retrieval algorithm.
Here, we assess the feasibility of this approach in the context of the Sentinel-5 mission, systematically comparing two different emulation strategies: an end-to-end approach, which directly approximates the full forward model with neural networks, and a hybrid approach, which combines fast non-scattering simulations with a neural network-based correction for atmospheric scattering effects. We comprehensively validate each emulator in the full retrieval chain, evaluating their impact on the accuracy of retrieved XCO2 and XCH4.
Our results show that a hybrid approach is needed to meet the stringent accuracy requirements on XCH4 and XCO2. While the end-to-end emulator achieves large speed-ups exceeding a factor of 300, it introduces considerable errors of 7.22 ppb for XCH4 and 4.25 ppm for XCO2 compared to full-physics retrievals, and fails to generalize to high-emission scenarios beyond the training range. In contrast, the hybrid approach can effectively leverage the information provided by the non-scattering approximation, reducing emulator-induced retrieval errors to less than 1.5 ppb for XCH4 and 0.5 ppm for XCO2, while still being an order of magnitude faster than full-physics retrievals and maintaining robust performance for high-emission scenarios.
Together, these results pave the way for operational deployment of neural network-based forward models in GHG retrievals from Sentinel-5, and more broadly demonstrate the potential of hybrid machine learning emulators to facilitate timely and accurate processing of the rapidly growing data volumes from modern satellite missions.
Competing interests: At least one of the (co-)authors is a member of the editorial board of Atmospheric Measurement Techniques.
Publisher's note: Copernicus Publications remains neutral with regard to jurisdictional claims made in the text, published maps, institutional affiliations, or any other geographical representation in this paper. While Copernicus Publications makes every effort to include appropriate place names, the final responsibility lies with the authors. Views expressed in the text are those of the authors and do not necessarily reflect the views of the publisher.
Review of “Accelerating greenhouse gas retrievals with neural network-based forward models” by Fiona Lippert et al.
General comments
This new study investigates the use of neural network-based forward model emulators to accelerate greenhouse gas retrievals from Sentinel-5 measurements. The authors compare a fully data-driven end-to-end emulator with two hybrid approaches that combine a fast non-scattering radiative transfer model with a neural network correction for scattering effects. The different approaches are evaluated comprehensively in terms of spectral accuracy, Jacobians, retrieval performance, computational efficiency, sensitivity to training data volume, and generalization to methane concentrations outside the training range. The results clearly demonstrate the advantages of the hybrid approach. It achieves retrieval accuracies close to the full-physics reference, substantially improves robustness and data efficiency compared with the purely data-driven emulator, and still provides significant computational speed-ups.
Overall, I find this to be an excellent and very convincing study. The methodology is sound, the experimental design is comprehensive, and the results are robust and remarkably clear. In particular, I appreciate that the authors evaluate the proposed emulators not only in terms of spectral accuracy, but throughout the complete retrieval chain, including Jacobians, retrieval accuracy, computational performance, sensitivity to training data volume, and extrapolation beyond the training range. The manuscript is also very well written and provides sufficient detail to understand and reproduce the approach. I therefore recommend acceptance subject to minor revisions, as outlined below.
Specific comments
Line 10: I suggest expressing the emulator-induced XCH4 and XCO2 errors also, or perhaps preferably, as relative errors in percent, as these may be easier for readers to assess than the absolute values in ppb and ppm. In addition, reporting values such as 7.22 ppb and 4.25 ppm with three significant digits may suggest a level of precision that is not warranted for these summary error metrics; fewer significant digits might be more appropriate.
Lines 155–160: The atmosphere is represented by only 20 homogeneous layers that are equidistant in pressure. This may well be sufficient for the present retrieval setup, but it would be useful to briefly justify this choice, for example by referring to previous sensitivity studies or established RemoTeC/LINTRAN configurations showing that this vertical resolution is adequate for XCH4 and XCO2 retrievals.
Lines 165–170: The aerosol representation involves several simplifying assumptions, including spherical particles, fixed refractive indices, and a prescribed vertical distribution. These assumptions may be well established for the RemoTeC retrieval framework, but it would be helpful to briefly motivate them here or refer more explicitly to previous studies demonstrating that this simplified aerosol representation is adequate for the present XCH4/XCO2 retrieval application.
Lines 207–219: The logarithmic transformation of the radiances and the PCA-based dimensionality reduction appear to be important and well-motivated design choices for reducing the complexity of the learning problem. Could the authors briefly comment on how important these steps are in practice? For example, did tests without the logarithmic transformation or PCA result in noticeably less stable training or reduced emulator accuracy? A short qualitative statement based on their model-development experience would already be helpful.
Lines 305–310: Reducing the 20-layer temperature profiles to only two principal components is an efficient way to limit the dimensionality of the neural network input. At the same time, this may remove finer-scale vertical temperature structure. Could the authors briefly comment on whether such information is expected to be relevant for the Sentinel-5 retrieval geometry considered here, and whether two principal components were found to be sufficient in sensitivity tests or previous retrieval studies?
Lines 430–435: Since the Jacobian evaluation is an important part of demonstrating that the emulators can be used reliably within the iterative retrieval, I suggest moving one representative Jacobian figure from Appendix A into the main text. In addition, Figures A1–A3 show noticeably larger deviations for the aerosol central height (z_aer). The manuscript notes that this parameter also has the lowest instrument sensitivity; could the authors briefly discuss whether the reduced Jacobian agreement is mainly a consequence of this weak sensitivity and whether it has any noticeable impact on the retrieval performance?
Line 470: The runtime comparison is based on measurements on a single CPU core. It may be useful to briefly note that the reported relative speed-ups are implementation- and hardware-dependent, in particular since neural network inference may benefit considerably from GPU acceleration or batched processing.
Lines 538–540: The suggestion that emulator-based retrievals could ultimately exceed the accuracy of operational full-physics retrievals is interesting, but appears somewhat speculative based on the present results. The authors should qualify the statement as a possible future perspective.
Lines 555–560: Given that the present evaluation is based entirely on synthetic measurements, I suggest slightly qualifying the statement that the hybrid emulators are candidates for implementation in operational Sentinel-5 processing. The authors appropriately discuss the need for evaluation with real measurements later in this section, so a small wording adjustment here would make the distinction between demonstrated feasibility and operational readiness clearer. Also, “this emulators” should read “these emulators”.
Technical corrections
Line 160: “reach the required accuracy required”: please remove the duplicated “required”.
Line 525: A full stop is missing after “training distribution”.
Figure B1 caption: “aerosol layer hight” should read “aerosol layer height”.