the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
Performance validation of the GHGSat Methane Constellation via controlled releases
Abstract. Satellite remote sensing has become an important tool for detecting, quantifying, and attributing methane emissions, yet a rigorous, condition-dependent characterization of point-source imager performance has not previously been established, despite being essential to support its use in regulatory and voluntary reporting frameworks. We present a comprehensive performance assessment of the GHGSat constellation of high-resolution methane-imaging satellites based on a multi-campaign controlled release dataset spanning 2021 to 2026 that combines self-organized campaigns with independent third-party experiments. We develop a probabilistic, environmental conditions-dependent detection model in which the probability of detection depends on the plume signal to noise ratio (SNR), which is in turn a function of emission rate, wind speed, retrieval noise, and spatial resolution. We find that an SNR of 1.72 ± 0.39 is required to achieve a 50 % probability of detection, which translates to a detection limit Q50 = 99.0 ± 5.3 kg h⁻¹ in median environmental conditions encountered in controlled releases (3 m s⁻¹ wind speed, 7 mmol m⁻² column density noise, 27 m resolution). The estimated Q50 is shown to converge stably as data accumulate and to remain consistent when blind validation samples are added to self-organized releases. Quantification accuracy is evaluated through parity analysis of estimated versus metered emission rates, yielding an ordinary least squares slope of 0.93 ± 0.03 and R² of 0.92 using reported model winds, improving to 0.96 ± 0.02 and R² = 0.95 with a co-located anemometer. A comparison of emission rate estimates based on ERA5, IFS, and HRRR winds shows that quantification accuracy is largely insensitive to the choice of operationally available wind product, with a residual underestimation at low emission rates that correlates with wind model spatial resolution. We further demonstrate a bias correction combining a local concentration background correction with an empirically recalibrated effective wind speed, which recovers the low-rate sources that were previously underestimated and removes most of the residual bias without degrading accuracy at higher emission rates. These results establish a transparent, statistically grounded baseline for the GHGSat constellation's detection and quantification performance and provide a methodological framework that can be extended as additional controlled-release data become available.
Status: final response (author comments only)
- RC1: 'Comment on egusphere-2026-3662', Anonymous Referee #1, 07 Aug 2026
-
RC2: 'Comment on egusphere-2026-3662', Anonymous Referee #2, 25 Aug 2026
This is a very interesting and important analysis about the best developed satellite methane system, and I enjoyed reading it. It is a comprehensive analysis that brings together numerous studies. The manuscript has a solid basis in that it extends from previously published studies, but I feel that a number of communication-related issues would make it clearer and stronger:
Introduction: I suggest removing the final 6.5 lines of the Introduction because this information is repeated later in the Results and Discussion:
“Comparing multiple variants of this model based on the Akaike and Bayesian information criteria, we find the PoD to be well described by a log-normal function of the ratio of source rate (q) to retrieval noise (n), spatial resolution, and wind speed (u), with a linear offset parameter (u₀). The best-fit model obtained through maximum likelihood implies a 50% probability of detecting an emission rate of 99.0 ± 5.3 kg h⁻¹ at a wind speed of 3 m s⁻¹ and a median CH₄ column error of 7.0 mmol m⁻². We demonstrate that this estimate converges stably over time and remains highly consistent as independent blind validation is added to the self-organized releases. We also evaluate quantification accuracy through a parity analysis comparing GHGSat-estimated emission rates across all campaigns.”
Table 1: The reference “Tarek 2026” is not included in the reference list.
Anon. reference: This protocol appears to be available online through the Carbon Management Canada website. If this is the same version, it could be properly referenced using the following web address: https://cmcghg.com/wp-content/uploads/2025/09/MDAQ-Test-Protocol-Canada.pdf
Section 3.1 is technically appropriate but very jargon-heavy for the broad community that may be interested in GHGSat outcomes. It is difficult to read without frequently stopping to think because the language is overly technical and some sentences are particularly challenging. For example: “The spatial resolution metric 𝑎 is the 1-sigma width of the optical point spread function (PSF) in ground coordinates, scaled to account for pixel stretch at off-nadir viewing angles,” and “We note that the standard error based on the fit residuals provides a more accurate and conservative estimate of the local column precision than its ‘model noise’ counterpart which simply propagates sensor shot noise and readout noise into retrieval parameter space and therefore yields only a lower bound that can be exceeded due to correlated noise and poor model fit.” While I appreciate the economy of highly technical writing, some plain-language explanation would help here. These are not overly difficult concepts, and the authors could communicate them more clearly.
Z is derived from a heuristic, but plume detection uses IME. Could the authors explain more clearly how a peak-based heuristic can be used to estimate PoD for an IME-based detection pipeline?
Why is the false-positive side of detection performance omitted? As presented, the model characterizes only sensitivity rather than providing a complete performance specification.
Section 3.2: Does the bootstrap method resample individual events, or is resampling blocked by campaign or site? Because the data are grouped by site and campaign, releases from the same site may share correlated noise and other characteristics.
Figure 4 and surrounding text: Since resolution (a) does not improve the model on its own, please state explicitly that a in the equation is assumed rather than fitted.
The differences among the predictor/link combinations are small, but the linear-offset model is adopted as the primary model. Perhaps a likelihood-ratio test against the Jacod heuristic model could be used to establish formal significance. How many parameters separate the two models?
The text at the bottom of page 16, lines 364–373, could also be less jargon-heavy. The differences among the detection models were small and were smaller than the uncertainty caused by the limited set of experiments. Therefore, if the study were repeated using a somewhat different collection of methane releases, the model scores could appear in a different order.
Line 445: The range of 3 to 13 pixels is quite wide. Could the authors provide a sensitivity check?
The manuscript frequently refers to a local anemometer. At the CMC site (Pseudo), the local onsite anemometer was 2.25 m above the ground, whereas wind measured at 10 m would be expected to be approximately 25% faster. Since Figure 6 presents results from Pseudo, Blind, and other campaigns, the authors must be referring to more than one local anemometer, presumably one for each study. At what heights were these other anemometers deployed? Were they all 2.25 m above the ground, or were some closer to the 10 m reference height of the global wind product? These are integrated into the modeling, so deployment characteristics and differences should be clear.
Ueff is presented in Figure 7. I am always interested by this adjustment factor. Based on the wind power law, a Ueff of 0.4 using an input wind of 10 m would correspond to a wind location well within the roughness layer. Realistically, many plumes are probably affected by terrain, vegetation, and other obstacles. However, one might expect Ueff to be closer to 1 in these relatively ideal controlled-release studies, after accounting for release height. In the CMC work, the vent stack heights were 3.4 and 4.8 m. Not very high and near the roughness layer. In the Stanford studies, I believe the release height was closer to 7 m. Do release height and proximity to the roughness layer affect Ueff and your model?
The study develops a sensitivity outcome for these experiments, but the authors also develop a correction factor that extends these optimal results to less optimal real world conditions. The sensitivity is expected to be 1.6x worse in the real world. I think the implications in the real world should be clearer, not only in the abstract but in the conclusions as well. In the abstract, only the optimal value is listed. Many people read only the abstract, and may therefore get the wrong impression.
Citation: https://doi.org/10.5194/egusphere-2026-3662-RC2
Viewed
Since the preprint corresponding to this journal article was posted outside of Copernicus Publications, the preprint-related metrics are limited to HTML views.
| HTML | XML | Total | BibTeX | EndNote | |
|---|---|---|---|---|---|
| 91 | 0 | 2 | 93 | 0 | 0 |
- HTML: 91
- PDF: 0
- XML: 2
- Total: 93
- BibTeX: 0
- EndNote: 0
Viewed (geographical distribution)
Since the preprint corresponding to this journal article was posted outside of Copernicus Publications, the preprint-related metrics are limited to HTML views.
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
Ramier et al. present a detailed accounting of GHGSat’s performance validation work during 2021-2026, specifically with respect to detection sensitivity and quantification accuracy. The validation assessment is based in single- and pseudo-blind controlled releases by the authors/GHGSat and third parties. The work largely leverages published methods but also presents unique insights into the relevant parameters affecting sensitivity in addition to a novel bias correction method for emission rate quantification.
I have some minor concerns about the underlying methodology of the paper. Specifically, the use of an idealized predictor; the limited assessment of alternative model formulations; and the extensibility of the results to other regions of observation (i.e., not the controlled release locations).
The paper is generally well-written, describes novel methods and results, and will be a relevant contribution to the readership of AMT. I recommend that it be accepted for publication after the comments below are addressed.
Main comments:
Abstract
Methods
Detection performance
Quantification accuracy
Minor comments:
Page 2, line 28 – Ayasse et al. (2024; doi: 10.1021/acs.est.4c06702) evaluated the detection performance of EMIT.
Page 3, line 52 – provide citation for the atmospheric lifetime of methane and the GWP.
Page 3, line 64 – can the authors comment on the steady-state count of GHGSat-C satellites? This might be interesting and relevant to the reader.
Page 6, figure 1 – I suggest including this figure in an SI. Are there similar figures of the other release sites?
Page 7, eq. 1 – should this be a summation over pixels in the plume “mask” rather than an integral?
Page 8, line 177 – there are many examples of measurement-based and -informed inventories in the literature that consider detection sensitivity. Consider citing some additional sources.
Page 11, line 269 – I’m surprised that the Rice distribution did not work as well given its physical basis. I was hopeful; kudos for trying. Could the AICc of the various fits be provided in an SI (or added to Table 4) so that the reader could get a sense of how differently the various models performed?
Page 12, line 282 – cite Conrad et al. (10.1016/j.rse.2023.113499).
Page 15, line 325 – Q90 seems to vary much more than Q50 suggesting that the tails of the various POD fits can be quite different.
Page 16, line 358 – I’m very fond of this incremental assessment of each variable’s predictive power and the insights. Consider adding it to the abstract if space allows.
Page 17, figure 5 – there is an interesting spike in Q90 early in 2025. Can the authors comment on whether this is a new satellite performing very differently vs. additional controlled release data (I think the latter)?
Page 18, line 402 – why is “true” in quotes? Considered replacing with metered as in the following paragraph.
Page 18, line 417 – clarify which specific wind data are used here and clarify in the caption of Figure 6 that the top and bottom rows are modelled and local wind speeds, respectively.
Page 21, line 457 – can “markedly” be quantified? E.g., the scatter at rates less than some kg/h reduce by some amount.
Page 23, line 511 – this is very interesting. Is this related to the skewness of the wind speed distribution?