the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
Physically Structured Target Design Improves the Robustness and Transferability of Neural Network Methane Retrievals
Abstract. Artificial neural networks (ANNs) offer a computationally efficient alternative to conventional physics-based satellite retrievals of atmospheric methane (CH₄). However, the impact of target design on retrieval generalization and robustness remains unclear. Here, we develop two ANN-based column-averaged dry-air mole fraction of methane (XCH₄) retrieval frameworks from GOSAT-2 observations. The models are trained with full-physics (ANN-FP) or proxy (ANN-Proxy) retrieval products and evaluated against 2021–2022 observations and TCCON measurements. While both ANNs reproduce their training targets during training, ANN-Proxy is more stable out of sample, whereas ANN-FP shows increasing bias and drift. Mechanistic analyses using feature attribution, local gradient geometry, and environmentally conditioned perturbation modes indicate that ANN-Proxy concentrates over 97 % of its attribution within CH₄ and carbon dioxide (CO₂) absorption windows and maintains a coherent sensitivity field (mean gradient-direction similarity 0.96 versus 0.72 for ANN-FP). Furthermore, ANN-Proxy sensitivity directions are less aligned with aerosol- and albedo-induced perturbation modes, indicating reduced environmental coupling. The proxy-oriented target shapes the learned mapping to enhance robustness and transferability. These results suggest that physically structured proxy targets provide a practical strategy for rapid, stable, and operationally deployable neural-network methane retrievals in future carbon-monitoring satellite missions.
- Preprint
(1853 KB) - Metadata XML
- BibTeX
- EndNote
Status: open (until 23 Oct 2026)
- RC1: 'Comment on egusphere-2026-3549', Anonymous Referee #1, 15 Sep 2026 reply
Viewed
| HTML | XML | Total | BibTeX | EndNote | |
|---|---|---|---|---|---|
| 169 | 66 | 29 | 264 | 39 | 28 |
- HTML: 169
- PDF: 66
- XML: 29
- Total: 264
- BibTeX: 39
- EndNote: 28
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
Review of “Physically Structured Target Design Improves the Robustness and Transferability of Neural Network Methane Retrievals” submitted for publication to AMT by Rongjin Son et al.
This manuscript investigate the impact of the training target on the robustness of neural-network XCH₄ retrievals from GOSAT-2 observations. The authors have compared an ANN trained to reproduce the full-physics (FP) retrieval with a proxy-oriented framework based on the CH₄/CO₂ ratio. They show substantially better stability for the proxy-oriented approach, and complement the retrieval evaluation with TCCON comparisons and several analyses of the learned models, including SHAP attribution, gradient geometry, and empirical aerosol- and albedo-related perturbation modes.
The manuscript is very clearly written, and I have not identify any flaw that would invalidate the empirical results. The degradation of ANN-FP relative to ANN-Proxy when moving from the 2020 training period to 2021–2022 appears quite clear. However, I have some reservations regarding the interpretation of the comparisons. I do not think that the experimental design fully supports the central conclusion that the improved robustness results from the choice of the training target itself.
My main concern is that the two experiments differ in several respects simultaneously. ANN-FP and ANN-Proxy do not simply use different supervised targets: they also have different architectures, different information pathways, and different sets/organization of input variables. In addition, ANN-Proxy predicts the CH₄/CO₂ ratio, from which XCH₄ is subsequently reconstructed using a model-based XCO₂ field. The comparison therefore contrasts two retrieval frameworks rather than two otherwise identical neural networks differing only in their target. The authors themselves acknowledge this issue in Section 4.2, noting that the frameworks differ “not only in their supervised targets but also in the information pathways emphasized during learning,” and later state that the observed behavior may arise from “the proxy-oriented retrieval framework as a whole rather than from the retrieval target alone.”
This distinction is important because the title, abstract, and conclusions place considerable emphasis on target design as the mechanism responsible for the improved robustness. A more convincing demonstration would require an ablation experiment in which the FP and proxy targets are learned using the same architecture and the same input information. Ideally, target choice and architecture/input design would be varied independently. Without such an experiment, it seems difficult to determine how much of the observed improvement originates from the proxy target itself and how much results from the architecture and information content provided to the two networks.
This issue also affects the interpretation of the mechanistic analyses. For example, the fact that ANN-Proxy concentrates more than 97% of its SHAP attribution in the CH₄ and CO₂ absorption bands is presented as evidence that the proxy target leads to a more physically constrained representation. However, ANN-Proxy is explicitly constructed around separate CH₄ and CO₂ spectral branches and has a different set of auxiliary inputs from ANN-FP. Therefore, at least part of the difference in feature attribution may be imposed by the model design rather than emerging as a consequence of the target. The same caution applies, I think, to the interpretation of the gradient-space analyses.
More generally, I am not entirely convinced that the principal empirical result is unexpected. The motivation for proxy retrievals is precisely that the CH₄/CO₂ ratio reduces sensitivity to common light-path perturbations. The manuscript itself explains that this cancellation is one of the main advantages of the proxy approach. It is therefore not surprising that a neural network designed and trained to emulate such a retrieval exhibits less sensitivity to aerosol and surface-reflectance variability than a network emulating a full-physics product. I think the manuscript would benefit from a clearer discussion of what is genuinely learned from the present experiments beyond this expected property of proxy retrievals.
Finally, the claims regarding operational or onboard applicability appear somewhat stronger than what is demonstrated here. ANN-Proxy is indeed smaller and faster than ANN-FP, but the quoted inference time concerns the neural network itself. The final XCH₄ retrieval additionally requires a model-based XCO₂ field. This external information requirement should be considered when discussing onboard implementation, particularly since the manuscript itself notes that proxy retrievals depend on the accuracy of the CO₂ reference field. Also, there is no indication that onboard processing power is really the main issue in the context of methane observation.
Overall, I find the empirical comparison potentially useful, but I think the conclusions need to be more carefully aligned with what the experimental design actually demonstrates. At present, the results provide convincing evidence that the proxy-oriented retrieval framework as implemented here is more stable under temporal extrapolation than the implemented ANN-FP framework. They do not, in my view, yet demonstrate that target design itself is responsible for this improvement. An appropriate set of ablation experiments, or alternatively a substantial reframing of the conclusions around the comparison of the two complete retrieval frameworks, would correct the main weakness of the manuscript.
Technical comment
Line 34 “particularly in FP retrieval”. Is there any indication that there are less systematic errors for other types of retrievals ?
Line 54 : I do not think that this reference is realy appropriate for the statement
Figure 1 : Color scale should be white for zero and colorized for 1 or larger. Do not have white for ≈50 samples.
Also, it would make sense to have different ranges for the two color scales to judge where the spatial distributions are exactly the same, or not. If the distributions are exactly the same (which is expected for random sampling), a single figure is sufficient
Figure 6 : Please provide more explanation on the figure. What are the vertical bars ? How are the days selected ?