the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
Facility-scale quantification and monitoring of ammonia (NH3) emissions using ASTER multispectral thermal infrared observations
Abstract. Ammonia (NH3) is an important atmospheric pollutant affecting air quality, ecosystems, and climate, but current satellite observations remain limited in their ability to resolve individual emission sources. Hyperspectral thermal infrared sounders such as the Infrared Atmospheric Sounding Interferometer (IASI) and the Cross-track Infrared Sounder (CrIS) provide broad spatial coverage and high spectral sensitivity, but their kilometer‑scale footprints limit direct facility‑scale source attribution. Here, we investigate whether high‑spatial‑resolution multispectral thermal infrared imaging can detect NH3 plumes at facility scale.
We develop a physically based retrieval framework for the Advanced Spaceborne Thermal Emission and Reflection Radiometer (ASTER), combining radiative transfer calculations with lookup table inversion. The method exploits the differential sensitivity of ASTER bands 13 and 14 to NH3 absorption in the ν2 band near 930–970 cm⁻¹ and retrieves NH3 column enhancements at 90 m spatial resolution. Surface emissivity is taken from a long‑term ASTER emissivity climatology, while scene‑level emissivity products are used diagnostically to identify plume‑related band behavior. Sensitivity tests show that NH3 absorption remains measurable after convolution with the ASTER spectral response functions, but retrieval performance depends strongly on thermal contrast between the surface and the NH3‑bearing layer.
The retrieval is applied to ASTER observations over three industrial NH3 point sources: Khor Al Zubair, Tolyatti, and Piesteritz. Khor Al Zubair provides the clearest demonstration, with repeated source‑connected plume structures under favorable arid conditions. Tolyatti and Piesteritz show that detection is also possible in more heterogeneous environments, although only under suitable thermal contrast and surface conditions. Source‑rate estimates derived with the Integrated Mass Enhancement method are interpreted as instantaneous effective estimates for successful plume scenes, not as annual mean emissions or continuous facility‑average emissions. Where independent constraints are available, ASTER‑derived source‑rate statistics are consistent in magnitude with published satellite and airborne estimates.
These results demonstrate that multispectral thermal infrared imagers can provide high‑resolution information on NH3 plume structure and source location, complementing coarse‑resolution hyperspectral satellite observations. The approach is best suited for large, persistent sources and episodic plume mapping rather than routine monitoring, because ASTER sampling is limited by revisit frequency, cloud cover, and thermal contrast. The framework supports retrospective analysis of archival ASTER scenes and informs future high‑resolution thermal infrared imaging concepts for NH3 point‑source detection.
- Preprint
(30429 KB) - Metadata XML
- BibTeX
- EndNote
Status: final response (author comments only)
-
RC1: 'Comment on egusphere-2026-2893', Anonymous Referee #1, 25 Jun 2026
-
AC1: 'Reply on RC1', Lidewij Tijhuis, 19 Aug 2026
Response to Reviewer 1
We thank the reviewer for the careful reading of the manuscript and for the constructive comments. We particularly appreciate the suggestions regarding repetition and manuscript structure, which have led to substantial revisions.
Reviewer comments are reproduced in bold. Line numbers refer to the original submission.
Comment on repetitiveness and length
The paper is repetitive and unnecessarily long, with five itemized examples: the effective-emissions caveat, thermal contrast, ERA5 winds, comparison versus validation, and the band 13/14 setup. Sections 3.1.1 and 3.1.2 removal.
We agree with the reviewer's assessment and have substantially restructured the manuscript rather than simply shortening it. The five examples cited reflect a broader issue of repetition throughout the text. We therefore addressed the issue systematically and report the outcome for each item below.
Effective-emissions qualification: reduced from ten occurrences to three (Abstract, IME definition, and Conclusions). The Methods section now carries the qualification explicitly: "All source rates reported in this paper, including the detected-scene medians and means, are statistics over such snapshots and are not annual mean or continuous facility-average emissions; this qualification is not repeated at each occurrence below."
Thermal contrast: reduced from six occurrences to one discussion in Sect. 3.1, with the detailed quantitative treatment retained in Appendix A.
Wind source and its limitations: reduced from five occurrences, including a figure caption, to a single discussion in Sect. 2.3.4, which is cross-referenced thereafter.
Comparison with published estimates is not validation: reduced from five occurrences to a single statement immediately before Table 3, where it is most relevant.
Band-13/Band-14 rationale: reduced from seven occurrences to four (Abstract, Introduction, Sect. 2.1, and Sect. 2.3). We believe these are the locations where the information is needed for context and methodology.
Three structural changes account for most of the reduction in length. First, the opening portion of the Discussion has been removed because it largely repeated material already presented in the Results section. The Discussion now begins directly with the implications and limitations of the work. Second, Sections 3.1.1 and 3.1.2 have been condensed into a single paragraph in Sect. 3.1, and the associated figures have been moved to Appendix A as suggested. Third, the Conclusions have been rewritten to avoid repeating the Discussion and instead focus on the broader implications for instrument design and future observations
Comment, technical corrections
Line 9, missing space. Line 33, inventories "exhibit" temporal variability. Figure 1, font size. Figure 2, "Specie" and the band legend. Line 101, retrieved emissivity. Line 104, remove the last sentence.
We thank the reviewer for these corrections and have adopted all of them. The missing space at line 9 has been corrected. Line 33 has been revised to state that inventories generally fail to capture strong temporal variability. Figure 1 has been regenerated with larger annotation text. Figure 2 now reads "Species"; the thermal infrared bands have been removed from the legend because they are already labelled on the axes, and the figure has been reformatted accordingly. Line 101 now refers to band-correlated variability in the radiance fields and retrieved emissivity. The final sentence at line 104 has been removed.
Line 149: perhaps mention the lapse rate that was used
No fixed lapse rate is used in the LUT construction. We have therefore clarified the temperature-profile construction procedure in Sect. 2.3.1 rather than quoting a single representative lapse rate.
A reference LOTOS-EUROS profile provides the vertical structure, with the upper part of the profile retained unchanged. Below a prescribed anchor level, the profile is replaced by a smooth quadratic fit and rescaled between two constraints: the upper anchor temperature and the prescribed air temperature at the lowest model level (approximately 60 m). The temperature gradient within the NH3 layer is therefore determined by the profile rescaling and varies throughout the LUT. Because the scaling depends on the prescribed air temperature, warmer states generally produce steeper gradients. Consequently, no single lapse rate applies across the lookup table, and this is now stated explicitly.
Line 169/327: No mention is made in the paper of the water vapor continuum, which is apart from clouds and emissivity the next largest contributor to changes in the baseline slope. Did the authors run simulations at different levels of humidity?
Regarding the water-vapor continuum, TFIT currently does not include an explicit continuum treatment such as MT_CKD. We have therefore not quantified the influence of continuum absorption in the present study. Because continuum absorption varies smoothly across the spectral window sampled by ASTER bands 13 and 14, we expect part of its contribution to cancel in the split-window ratio. However, this expectation has not yet been tested explicitly and is now identified as a limitation of the current retrieval framework.
To address the reviewer's question regarding humidity sensitivity, we performed an additional sensitivity analysis. Using a dedicated sensitivity LUT, the water-vapor profile was scaled by factors of 0.7, 1.0, and 1.3 while holding all other state variables fixed. The perturbed ratio was then inverted using the unperturbed LUT so that the resulting effect could be expressed directly as an error in the retrieved NH3 column.
A 30 % error in water vapor produces a retrieval error of approximately 4-7 % at a thermal contrast of 10 K, decreasing to below 4 % at thermal contrasts of 40 K or greater. Across all sampled states, the median absolute error is 3.8 % and the maximum error is 7.4 %.
The influence on the plume-free baseline is smaller still. At the lowest NH3 node, a 30 % humidity perturbation changes the split-window ratio by approximately one quarter of the per-pixel ratio noise at the median and 0.39 times the noise at worst. Thus, even a substantial humidity error shifts the background by less than the instrumental noise and by roughly one eighth of the 2σ plume-detection threshold.
Interestingly, the sign of the effect is opposite to what might be expected intuitively: increased water vapor is interpreted as a reduction in NH3, because band 14 is more sensitive to water-vapor absorption than band 13.
It should be noted that this analysis quantifies sensitivity around one reference state rather than the operational retrieval error. The operational LUT does not include humidity as a retrieval dimension, and no humidity field is ingested. A single water-vapor profile is therefore applied to all scenes and sites. Given the climatic differences between our arid, temperate, and continental sites, the true mismatch may be larger than 30 %, potentially producing errors of order 10-15 %. Since humidity information is already available from the same meteorological archive used for temperature, extending the LUT to include humidity is a logical next step. Water-vapor sensitivity is now included in the uncertainty budget in Appendix B.
Line 266: "This is a general limitation of ASTER TIR...". This is not specific to ASTER, the retrieval sensitivity to TC is general to all IR sounders
We agree that the original wording was too specific. The revised manuscript now states that the thermal-contrast dependence is a general characteristic of thermal infrared sounding of near-surface gases rather than a limitation unique to ASTER. The suggested reference has been added.
Figure 6: are the units correct? As a ratio, is should be dimensionless. I would also add a contour at sqrt(2)sigma/B14, with sigma the radiance noise and B14 calculated at some temperature. This would delimit the detectable TC/NH3 region nicely.
The reviewer is correct that the quantity plotted is dimensionless, and the caption now states this explicitly.
We have also added detectability contours corresponding to one and three times the propagated ratio noise. These contours define the marginal and robust detectability regimes and therefore provide a clearer indication of the useful retrieval domain than a single threshold curve.
Line 410: IASI box-model estimate used a lifetime of 12 hours, but this can easily be converted to the more realistic lifetime of 2-3 hours, giving 36 kt, well-aligned with the ASTER value. An advantage of ASTER and other high spatial resolution sounders is that they do not require an estimate of the chemical lifetime (I do not think this is mentioned in the manuscript).
We now discuss this point explicitly in the Tolyatti section.
Because the box-model source rate scales inversely with the assumed lifetime, the published estimate of 7.4 kt yr-1 for a 12 h lifetime corresponds to 35.5 kt yr-1 when scaled to the 2.5 h lifetime used by Dammers et al. This places it just above our detected-scene median of 27.1 kt yr-1, rather than a factor of four below it. Expressed on a common lifetime basis, the two IASI approaches bracket the ASTER estimate within roughly a factor of three.
We also now discuss an advantage of high-resolution plume imaging. Because the IME method integrates mass directly within a resolved plume, no explicit chemical lifetime enters the ASTER calculation. The independence is not complete, however. For the Khor Al Zubair example shown in Fig. 8, the plume mask spans approximately 5 km. At the observed wind speed of 2.2 m s-1, this corresponds to roughly 0.6 h of transport. Consequently, an NH3 lifetime of a few hours would still imply some loss before the plume leaves the integration domain. In that respect, the reported source rates remain a lower bound.
General: Are all three point sources associated with the production of fertilizers? Did you have a look for NH3 emissions from livestock housings/feedlots? It would be good to address both questions in the revised manuscript.
All three sites are fertilizer-production complexes associated with ammonia and urea manufacturing, and this is now stated explicitly in Sect. 2.4. Because the industrial process is common across all three facilities, differences in retrieval performance are more plausibly attributed to observing conditions, such as thermal contrast, cloud cover, and surface heterogeneity, than to differences in source type.
We did not investigate livestock emissions in the present study. The objective was to develop and evaluate the retrieval framework using large, well-established industrial NH3 point sources, and we now state this explicitly in the manuscript.
The primary limitation is expected to be source strength. Dairy housing emits approximately 13.8 kg NH3 per per-animal-place yr-1. Even adopting a generous ASTER detection threshold of 1.5 kt yr-1, slightly below the smallest source rate detected in this study, a facility would require on the order of 105 animal places to reach that emission level. Such source sizes are uncommon and largely restricted to the very largest feedlot operations. Consequently, industrial facilities represent the most suitable initial target class for demonstrating the methodology.
Citation: https://doi.org/10.5194/egusphere-2026-2893-AC1 -
AC2: 'Reply on RC1', Lidewij Tijhuis, 19 Aug 2026
Response to Reviewer 1
We thank the reviewer for the careful reading of the manuscript and for the constructive comments. We particularly appreciate the suggestions regarding repetition and manuscript structure, which have led to substantial revisions.
Reviewer comments are reproduced in bold. Line numbers refer to the original submission.
Comment on repetitiveness and length
The paper is repetitive and unnecessarily long, with five itemized examples: the effective-emissions caveat, thermal contrast, ERA5 winds, comparison versus validation, and the band 13/14 setup. Sections 3.1.1 and 3.1.2 removal.
We agree with the reviewer's assessment and have substantially restructured the manuscript rather than simply shortening it. The five examples cited reflect a broader issue of repetition throughout the text. We therefore addressed the issue systematically and report the outcome for each item below.
Effective-emissions qualification: reduced from ten occurrences to three (Abstract, IME definition, and Conclusions). The Methods section now carries the qualification explicitly: "All source rates reported in this paper, including the detected-scene medians and means, are statistics over such snapshots and are not annual mean or continuous facility-average emissions; this qualification is not repeated at each occurrence below."
Thermal contrast: reduced from six occurrences to one discussion in Sect. 3.1, with the detailed quantitative treatment retained in Appendix A.
Wind source and its limitations: reduced from five occurrences, including a figure caption, to a single discussion in Sect. 2.3.4, which is cross-referenced thereafter.
Comparison with published estimates is not validation: reduced from five occurrences to a single statement immediately before Table 3, where it is most relevant.
Band-13/Band-14 rationale: reduced from seven occurrences to four (Abstract, Introduction, Sect. 2.1, and Sect. 2.3). We believe these are the locations where the information is needed for context and methodology.
Three structural changes account for most of the reduction in length. First, the opening portion of the Discussion has been removed because it largely repeated material already presented in the Results section. The Discussion now begins directly with the implications and limitations of the work. Second, Sections 3.1.1 and 3.1.2 have been condensed into a single paragraph in Sect. 3.1, and the associated figures have been moved to Appendix A as suggested. Third, the Conclusions have been rewritten to avoid repeating the Discussion and instead focus on the broader implications for instrument design and future observations
Comment, technical corrections
Line 9, missing space. Line 33, inventories "exhibit" temporal variability. Figure 1, font size. Figure 2, "Specie" and the band legend. Line 101, retrieved emissivity. Line 104, remove the last sentence.
We thank the reviewer for these corrections and have adopted all of them. The missing space at line 9 has been corrected. Line 33 has been revised to state that inventories generally fail to capture strong temporal variability. Figure 1 has been regenerated with larger annotation text. Figure 2 now reads "Species"; the thermal infrared bands have been removed from the legend because they are already labelled on the axes, and the figure has been reformatted accordingly. Line 101 now refers to band-correlated variability in the radiance fields and retrieved emissivity. The final sentence at line 104 has been removed.
Line 149: perhaps mention the lapse rate that was used
No fixed lapse rate is used in the LUT construction. We have therefore clarified the temperature-profile construction procedure in Sect. 2.3.1 rather than quoting a single representative lapse rate.
A reference LOTOS-EUROS profile provides the vertical structure, with the upper part of the profile retained unchanged. Below a prescribed anchor level, the profile is replaced by a smooth quadratic fit and rescaled between two constraints: the upper anchor temperature and the prescribed air temperature at the lowest model level (approximately 60 m). The temperature gradient within the NH3 layer is therefore determined by the profile rescaling and varies throughout the LUT. Because the scaling depends on the prescribed air temperature, warmer states generally produce steeper gradients. Consequently, no single lapse rate applies across the lookup table, and this is now stated explicitly.
Line 169/327: No mention is made in the paper of the water vapor continuum, which is apart from clouds and emissivity the next largest contributor to changes in the baseline slope. Did the authors run simulations at different levels of humidity?
Regarding the water-vapor continuum, TFIT currently does not include an explicit continuum treatment such as MT_CKD. We have therefore not quantified the influence of continuum absorption in the present study. Because continuum absorption varies smoothly across the spectral window sampled by ASTER bands 13 and 14, we expect part of its contribution to cancel in the split-window ratio. However, this expectation has not yet been tested explicitly and is now identified as a limitation of the current retrieval framework.
To address the reviewer's question regarding humidity sensitivity, we performed an additional sensitivity analysis. Using a dedicated sensitivity LUT, the water-vapor profile was scaled by factors of 0.7, 1.0, and 1.3 while holding all other state variables fixed. The perturbed ratio was then inverted using the unperturbed LUT so that the resulting effect could be expressed directly as an error in the retrieved NH3 column.
A 30 % error in water vapor produces a retrieval error of approximately 4-7 % at a thermal contrast of 10 K, decreasing to below 4 % at thermal contrasts of 40 K or greater. Across all sampled states, the median absolute error is 3.8 % and the maximum error is 7.4 %.
The influence on the plume-free baseline is smaller still. At the lowest NH3 node, a 30 % humidity perturbation changes the split-window ratio by approximately one quarter of the per-pixel ratio noise at the median and 0.39 times the noise at worst. Thus, even a substantial humidity error shifts the background by less than the instrumental noise and by roughly one eighth of the 2σ plume-detection threshold.
Interestingly, the sign of the effect is opposite to what might be expected intuitively: increased water vapor is interpreted as a reduction in NH3, because band 14 is more sensitive to water-vapor absorption than band 13.
It should be noted that this analysis quantifies sensitivity around one reference state rather than the operational retrieval error. The operational LUT does not include humidity as a retrieval dimension, and no humidity field is ingested. A single water-vapor profile is therefore applied to all scenes and sites. Given the climatic differences between our arid, temperate, and continental sites, the true mismatch may be larger than 30 %, potentially producing errors of order 10-15 %. Since humidity information is already available from the same meteorological archive used for temperature, extending the LUT to include humidity is a logical next step. Water-vapor sensitivity is now included in the uncertainty budget in Appendix B.
Line 266: "This is a general limitation of ASTER TIR...". This is not specific to ASTER, the retrieval sensitivity to TC is general to all IR sounders
We agree that the original wording was too specific. The revised manuscript now states that the thermal-contrast dependence is a general characteristic of thermal infrared sounding of near-surface gases rather than a limitation unique to ASTER. The suggested reference has been added.
Figure 6: are the units correct? As a ratio, is should be dimensionless. I would also add a contour at sqrt(2)sigma/B14, with sigma the radiance noise and B14 calculated at some temperature. This would delimit the detectable TC/NH3 region nicely.
The reviewer is correct that the quantity plotted is dimensionless, and the caption now states this explicitly.
We have also added detectability contours corresponding to one and three times the propagated ratio noise. These contours define the marginal and robust detectability regimes and therefore provide a clearer indication of the useful retrieval domain than a single threshold curve.
Line 410: IASI box-model estimate used a lifetime of 12 hours, but this can easily be converted to the more realistic lifetime of 2-3 hours, giving 36 kt, well-aligned with the ASTER value. An advantage of ASTER and other high spatial resolution sounders is that they do not require an estimate of the chemical lifetime (I do not think this is mentioned in the manuscript).
We now discuss this point explicitly in the Tolyatti section.
Because the box-model source rate scales inversely with the assumed lifetime, the published estimate of 7.4 kt yr-1 for a 12 h lifetime corresponds to 35.5 kt yr-1 when scaled to the 2.5 h lifetime used by Dammers et al. This places it just above our detected-scene median of 27.1 kt yr-1, rather than a factor of four below it. Expressed on a common lifetime basis, the two IASI approaches bracket the ASTER estimate within roughly a factor of three.
We also now discuss an advantage of high-resolution plume imaging. Because the IME method integrates mass directly within a resolved plume, no explicit chemical lifetime enters the ASTER calculation. The independence is not complete, however. For the Khor Al Zubair example shown in Fig. 8, the plume mask spans approximately 5 km. At the observed wind speed of 2.2 m s-1, this corresponds to roughly 0.6 h of transport. Consequently, an NH3 lifetime of a few hours would still imply some loss before the plume leaves the integration domain. In that respect, the reported source rates remain a lower bound.
General: Are all three point sources associated with the production of fertilizers? Did you have a look for NH3 emissions from livestock housings/feedlots? It would be good to address both questions in the revised manuscript.
All three sites are fertilizer-production complexes associated with ammonia and urea manufacturing, and this is now stated explicitly in Sect. 2.4. Because the industrial process is common across all three facilities, differences in retrieval performance are more plausibly attributed to observing conditions, such as thermal contrast, cloud cover, and surface heterogeneity, than to differences in source type.
We did not investigate livestock emissions in the present study. The objective was to develop and evaluate the retrieval framework using large, well-established industrial NH3 point sources, and we now state this explicitly in the manuscript.
The primary limitation is expected to be source strength. Dairy housing emits approximately 13.8 kg NH3 per per-animal-place yr-1. Even adopting a generous ASTER detection threshold of 1.5 kt yr-1, slightly below the smallest source rate detected in this study, a facility would require on the order of 105 animal places to reach that emission level. Such source sizes are uncommon and largely restricted to the very largest feedlot operations. Consequently, industrial facilities represent the most suitable initial target class for demonstrating the methodology.
Citation: https://doi.org/10.5194/egusphere-2026-2893-AC2
-
AC1: 'Reply on RC1', Lidewij Tijhuis, 19 Aug 2026
-
RC2: 'Comment on egusphere-2026-2893', Anonymous Referee #2, 06 Jul 2026
This study by Tijhuis et al. demonstrates the use of a multispectral thermal infrared imaging instrument, ASTER, to map facility-scale ammonia emissions. Using the high-resolution radiative transfer model TFIT and the ASTER spectral response functions, the authors construct a lookup table (LUT) of ASTER-simulated radiances, varying the ammonia column, surface emissivity, air temperature, and surface temperature. Given a multi-year mean ASTER-derived emissivity product, a reanalysis air temperature, and the ASTER-retrieved surface temperature, the ammonia column is taken from the LUT entry that produces a band ratio (L13-L14/L14) closest to that of the observations (which are corrected with an RMA regression using the LUT radiances on a scene-by-scene basis). These retrievals are carried out over three industrial sites, followed by emission retrieval via the IME method. The authors are cautious with their emission estimates, for example because they are dependent on an assumed ammonia vertical distribution. This is a nice study, and, to my knowledge, the first demonstration of ammonia retrievals for this class of instrument, which serves as a nice contribution to the field. Though I strongly encourage the authors to edit the paper for conciseness and clarity throughout, my primary comments as follows are minor and for the authors to consider at their discretion.
General comment:
- From the perspective of emissions monitoring… how well known is the thermal contrast considered to be? This would be of relevance for interpreting the scenes where ammonia is not seen (e.g., in the “no Q” scenes of Figure 9). For these, is there a class of scenes where the thermal contrast and thus sensitivity is believed to be high such that the scene can be confidently labeled a null detect? Would the studied sources be expected to be intermittent?
Specific comments:
- In the scene-based radiance correction described in section 2.3.2, which LUT-simulated radiances are selected to match with the observed radiances?
- Equation 5: where does knowledge of the near-surface air temperature come from? Is the ammonia column the only free parameter?
- Line 214: what do you use for L in the IME equation?
- Line 225: are there specific potential co-emitted species to be concerned about with the sources studied here?
- Line 229: if it is known, consider saying what types of industrial processes are occurring at each of these sources
- Table 2: why are the total number of scenes less than would be expected for a 25-year data record with the stated 16-day revisit time (line 76)?
- Figure 6: can you add an explanation of the contour behavior on the left side of the plot (at less than 1e17 molecules/cm2 of NH3)? Also, is this not unitless?
- Line 435: why use only the Spring estimates from Noppen et al. (2023)?
Technical corrections:
- Line 9: missing space before “Surface”
- Line 47: please check that Varon et al. (2019) and Cusworth et al. (2022) are appropriate references here for the use of Sentinel-2 data
- Line 57: please check the wavenumber to wavelength conversion for 970 cm-1
- Figure 1: subplots are inconsistent with source figures in the manuscript (e.g., inset emission rate in the middle subplot; mean emission rate in the right subplot)
- Figure 2: “Specie” to “Species”
- Figure 3: please check that the magnitude of radiance values is correct given the stated units
- Table 1: consider removing TC as it is not independently varied
- Line 200: I believe this is an unintended sentence break
Citation: https://doi.org/10.5194/egusphere-2026-2893-RC2 -
AC3: 'Reply on RC2', Lidewij Tijhuis, 19 Aug 2026
Response to Reviewer 2
We thank the reviewer for the generous assessment and for a set of comments that were unusually precise about where the manuscript was underspecified. Several of them exposed steps that we had performed but not described, and the revision is clearer for it.
Reviewer comments are reproduced in bold. Line numbers refer to the original submission.
General comment on conciseness
I strongly encourage the authors to edit the paper for conciseness and clarity throughout.
We agree and have substantially streamlined the manuscript. The opening page of the Discussion has been removed, Sects. 3.1.1 and 3.1.2 have been condensed into a single paragraph with their figures moved to the Appendix, and several recurring statements now appear only once.
General comment on thermal contrast, null detections and intermittency
From the perspective of emissions monitoring, how well known is the thermal contrast considered to be? This would be of relevance for interpreting the scenes where ammonia is not seen. For these, is there a class of scenes where the thermal contrast and thus sensitivity is believed to be high such that the scene can be confidently labeled a null detect? Would the studied sources be expected to be intermittent?
This comment highlights an important distinction that was only implicit in the original manuscript. We have now made it explicit by classifying every cloud-free scene into one of five categories: detected, null detection, insensitive, unassigned enhancement, and unusable.
Thermal contrast is not measured directly. In our retrieval it is defined as the difference between the ASTER surface temperature and the prescribed near-surface air temperature. Consequently, its uncertainty is inherited from those inputs, particularly the air temperature. Rather than reporting an uncertainty on thermal contrast itself, we quantify its impact on the retrieval. An error of 1 K in air temperature changes the retrieved NH3 column by approximately 4-10 %, with the strongest sensitivity occurring at low thermal contrast. This dependence is now discussed explicitly and included in the uncertainty budget.
We agree that some non-detections can be interpreted more confidently than others. We therefore distinguish between null detections and insensitive scenes. A null detection occurs when thermal contrast and retrieval sensitivity are sufficient for a plume of the magnitude typically observed at that site to have been detected, yet no source-connected enhancement exceeds the detection threshold. These scenes can be interpreted as genuine non-detections and provide an upper limit on the source rate. By contrast, insensitive scenes occur when thermal contrast is too weak for a plume to have been detectable regardless of emission strength and therefore provide no information on source activity.
This distinction is particularly important at Piesteritz, where thermal contrast is frequently weak. Of the 51 cloud-free scenes, only three qualify as null detections, with minimum detectable source rates of 2.1, 5.5, and 21.4 kt yr-1, respectively. The majority of cloud-free scenes are instead classified as insensitive.
With respect to intermittency, we do not possess independent operational data for the facilities that would allow us to determine whether emission variability arose from process changes, maintenance activities, or normal production fluctuations. As a result, we cannot distinguish an intermittent source from a source observed under insensitive conditions. This limitation is one of the motivations for the classification scheme above. Only null detections provide evidence that emissions were below a quantifiable upper limit at the time of the overpass; insensitive scenes cannot be interpreted as evidence for either low emissions or source shutdown. We have clarified this point in the revised manuscript.
Specific comment 1, the scene-based radiance correction
In the scene-based radiance correction described in section 2.3.2, which LUT-simulated radiances are selected to match with the observed radiances?
We agree that this step was ambiguous in the submitted manuscript and have clarified it in Sect. 2.3.2. The radiance correction uses the LUT radiance corresponding to the lowest NH3 state (effectively an ammonia-free atmosphere) as the reference. For each pixel, this reference radiance is obtained by interpolating the LUT to that pixel's prescribed emissivity, near-surface air temperature, and ASTER surface temperature. The regression is therefore performed between the observed radiances and a pixel-specific background simulation that already reflects the local surface and atmospheric state. It is not performed against a single scene-average spectrum or a fixed reference atmosphere.
Because the reviewer's question naturally raises the issue of what impact the correction may have on the retrieved NH3 signal, we have also expanded the discussion of this step in the revised manuscript. The correction is applied independently to each band and therefore does not strictly preserve the split-window ratio. To assess the magnitude of this effect, we compared the corrected and uncorrected radiances and quantified the resulting perturbation to the band ratio. The induced ratio change is typically comparable to the instrumental noise and substantially smaller than the plume signals associated with successful detections. We now discuss this limitation explicitly and identify the scene-wide correction as one of the weaker elements of the current retrieval framework.
Specific comments 2 and 3, the retrieval and IME inputs
Equation 5: where does knowledge of the near-surface air temperature come from? Is the ammonia column the only free parameter? Line 214: what do you use for L in the IME equation?
The near-surface air temperature is taken from the ECMWF operational IFS analysis at the acquisition time and location, interpolated to the scene.
The ammonia column is indeed the only free parameter. Emissivity, air temperature and surface temperature are all prescribed, and the inversion is a one-dimensional search along the ammonia axis of the table for the entry whose band ratio lies closest to the observed one. We have added the consequence, which the submitted text left unsaid: any error in a prescribed quantity aliases directly into the retrieved column, since the retrieval has no other degree of freedom in which to absorb it. This is the reason the uncertainty budget in Appendix B lists air temperature and emissivity as bias terms rather than random ones.
For the length scale in the integrated mass enhancement we use the square root of the segmented plume area, so that the length adapts to the plume actually detected rather than to a fixed distance from the stack. This is now stated explicitly at the equation.
Specific comments 4 and 5, the sources themselves
Line 225: are there specific potential co-emitted species to be concerned about with the sources studied here? Line 229: if it is known, consider saying what types of industrial processes are occurring at each of these sources.
All three sites are fertilizer production complexes associated with ammonia and urea manufacturing. We now state this explicitly in Sect. 2.4. Because the sites represent broadly similar industrial processes, differences in retrieval performance between them are more likely attributable to observing conditions, such as thermal contrast, cloud cover, and surface characteristics, than to differences in source type.
Regarding potential co-emitted species, the main species considered was ethylene, whose strongest thermal infrared absorption band lies close to the NH3 feature used in our retrieval. We do not expect substantial ethylene emissions from ammonia or urea production processes, although this assessment is based on the expected composition of emissions rather than on a dedicated analysis. We have therefore revised the manuscript to clarify that possible interference from other species associated with the Haber-Bosch production chain has not been comprehensively evaluated.
In addition, sensitivity tests performed during retrieval development indicated that even relatively large ethylene columns produce only a small effect on the ASTER band-integrated radiances after convolution with the instrument spectral response functions. This suggests that ethylene is unlikely to generate an NH3 like signal of comparable magnitude in the split-window observable used here, although we have not performed a full multi-species interference assessment.
We also note that, at Khor Al Zubair, nearby oil and gas flaring plumes are clearly visible in the VNIR imagery and in the thermal infrared radiances, yet they are not retrieved as NH3 enhancements (Fig. 8). This provides observational evidence that emissions from these neighboring sources do not generate a false NH3 signal at a detectable level in our retrieval.
Specific comment 6, the number of scenes in Table 2
Table 2: why are the total number of scenes less than would be expected for a 25-year data record with the stated 16-day revisit time?
The quoted 16-day revisit time refers to the ASTER orbital repeat cycle rather than to a guaranteed acquisition frequency. ASTER is a tasking-based instrument and does not acquire imagery continuously. Observations are collected according to acquisition priorities and instrument duty-cycle constraints, so many potential overpasses do not result in archived scenes. The counts reported in Table 2 therefore represent the total number of scenes available in the ASTER archive for each location before any screening by us.
We have revised the manuscript to clarify this distinction and now separately report the number of archived scenes and the number of cloud-free scenes retained for retrieval processing. For Khor Al Zubair, Tolyatti, and Piesteritz, 75, 67, and 136 scenes were available in the archive, respectively, of which 56, 25, and 51 remained after cloud screening
Specific comment 7, Figure 6
Figure 6: can you add an explanation of the contour behavior on the left side of the plot (at less than 1e17 molecules/cm2 of NH3)? Also, is this not unitless?
The reviewer is correct that the quantity plotted is dimensionless. The units given in the original caption were incorrect and have been removed.
We have also added text explaining the behavior of the contours at low NH3 columns. Below approximately 1017 molecules cm−2, the magnitude of the split-window ratio falls below the detection threshold for all thermal contrasts shown. In this region, the contour is controlled by the noise floor rather than by NH3 sensitivity, which causes it to turn over instead of continuing to follow an NH3 dependent trend. The contour should therefore be interpreted as marking the boundary of the detectable domain rather than as an NH3 isopleth.
Following a related suggestion from Reviewer 1, we have also added detectability contours corresponding to one and three times the propagated ratio noise to make the detectable region more explicit.
Specific comment 8, the Noppen et al. comparison
Line 435: why use only the Spring estimates from Noppen et al. (2023)?
We use the spring results because they are more tightly constrained. The autumn campaign was conducted under weaker thermal-contrast conditions (0-4 K), whereas the spring campaign experienced stronger thermal contrast (3-12 K). Noppen et al. show that the uncertainty associated with plume-rise assumptions is consequently much larger in autumn than in spring. For example, their native-resolution CSF estimate ranges from 2.2-12 kt yr−1 in autumn, compared with 2.2-5.4 kt yr−1 in spring.
The two flights differ in thermal contrast. The autumn flight was flown under near-neutral stability, with a thermal contrast of only 0 to 4 K between the surface and the boundary layer, whereas the spring flight was flown under slightly unstable conditions with a thermal contrast of 3 to 12 K. This matters because Noppen et al. show that it directly affects how well constrained their flux estimate is: the spread between their non-rising and buoyant plume scenarios is a factor of about 5 for the autumn flight but only a factor of about 2 for spring, which they attribute explicitly to the weaker autumn thermal contrast. Correspondingly, their native-resolution CSF estimate for autumn ranges from 2.2 to 12 kt per year, against 2.2 to 5.4 kt per year in spring.
We have revised the manuscript to state explicitly that the spring results are used because they provide the better-constrained comparison.
Technical corrections
Line 9, missing space before Surface. Line 47, please check that Varon et al. (2019) and Cusworth et al. (2022) are appropriate references for the use of Sentinel-2 data. Line 57, please check the wavenumber to wavelength conversion for 970 cm-1. Figure 1, subplots are inconsistent with source figures in the manuscript. Figure 2, Specie to Species. Figure 3, please check that the magnitude of radiance values is correct given the stated units. Table 1, consider removing TC as it is not independently varied. Line 200, I believe this is an unintended sentence break.
We are grateful for these and have adopted all. The missing space at line 9 has been corrected. The references at line 47 have been checked and corrected. Figure 1 has been regenerated so that the subplots agree with the source figures, including the inset emission rate in the middle panel and the mean emission rate in the right panel. Figure 2 now reads Species. The radiance magnitudes in Figure 3 have been checked against the stated units. Thermal contrast has been removed from Table 1, since the reviewer is right that listing a derived quantity alongside the independently varied ones invites the reader to assume it was gridded. The sentence break at line 200 was indeed unintended and has been repaired. An updated the 10.338 micrometers to 10.31 micrometers that it should have been.
Citation: https://doi.org/10.5194/egusphere-2026-2893-AC3
-
RC3: 'Comment on egusphere-2026-2893', Anonymous Referee #3, 21 Jul 2026
The authors of the manuscript "Facility-scale quantification and monitoring of ammonia (NH3) emissions using ASTER multispectral thermal infrared observations" present an proof of concept that uses ASTER multiband thermal data to detect ammonia point sources. The work is interesting and novel, complementing NH3 observations by spectrometers like CrIS and IASI. The authors use a split window method for NH3 estimation followed by IME to measure total source flux. The images are quite convincing that the detection works, and the method appears fairly conventional. However, I have some specific questions about the quantification aspect of the workflow, and the uncertainties that it includes or does not;
1) How fragile is the assumption that the plume is contained within the bottom 500 m of atmosphere, especially as NH3 is buoyant and plumes can extend for tens of kilometers? Is the result insensitive to errors in plume height?2) Which meteorological database does the acquisition air temperature estimate come from, and what is the uncertainty? Of course the air temperature is critical because, as indicated, the absorption depth is proportional to the temperature contrast so air temperature error will alias directly into IME. A buoyant plume at elevation could also be in air that is a different temperature than the at-surface estimate would suggest. How serious an issue is this?
3) What do the black bars signify in Figures 9 & 11 - are they 95% confidence intervals, one-sigma error bars, or something else? Also, how exactly is the flux retrieval uncertainty calculated? Does it account for the uncertainties listed by the authors: "radiometric noise in the retrieved column, background subtraction, plume masking, wind speed, wind direction, wind-field representativeness, the Ueff parameterization, surface emissivity assumptions, and the vertical distribution of NH3?" It seems like windspeed alone could lever the solution around by a factor.
4) The most confusing aspect of the workflow to me is initial radiometric adjustment. Here I admit some naiveté about ASTER; maybe such a scene-wise adjustment is always needed. But it seems like this sort of manipulation could impact the accuracy of the flux estimate. Note that, because the correction is applied independently on all bands, the absorption depth for the split window approach will not, in general, be preserved. The authors state that the correction accounts for "residual calibration differences, imperfect representation of surface emissivity, uncertainties in the atmospheric temperature profile, and simplifications in the forward model." But modifying the measurement data to fit the model seems like a suspect way to deal with this parameter uncertainty, since our measurement is the best connection we have to the real world. A more conventional approach would be to fit the model parameters to fit the data, weighted by uncertainty in calibration and prior uncertainty in the model parameters. This would preserve the physical interpretability and allow the residual uncertainty in model parameters to be forward-propagated to the solution. As it stands, the uncertainty induced by the radiometric update seems very difficult to interpret or quantify.
5) A related issue - it seems unlikely that these errors are all uniform across the scene, as the radiometric update assumes. Fitting the update on disjoint portions of the scene independently might be a way to confirm that the update is, in fact, scene-wise constant, or to understand the errors induced by such an assumption.Citation: https://doi.org/10.5194/egusphere-2026-2893-RC3 -
AC4: 'Reply on RC3', Lidewij Tijhuis, 19 Aug 2026
Response to Reviewer 3
We thank the reviewer for engaging closely with the quantification chain. The questions on the radiometric correction in particular have led to the most substantial changes in the revision, and we have taken the opportunity to state several limitations that the submitted manuscript left implicit.
Reviewer comments are reproduced in bold. Line numbers refer to the original submission.
Question 1, the assumption that the plume occupies the lowest 500 m
How fragile is the assumption that the plume is contained within the bottom 500 m of atmosphere, especially as NH3 is buoyant and plumes can extend for tens of kilometers? Is the result insensitive to errors in plume height?
We agree that the retrieval is sensitive to the assumed vertical distribution of NH3, and we have now quantified this sensitivity. The retrieval assumes that NH3 is uniformly distributed between the surface and 500 m altitude. Redistributing the same total column over alternative vertical profiles changes the retrieved NH3 column by between -4% and +63%, and this sensitivity is now included explicitly in the uncertainty budget presented in Appendix B.
The resulting uncertainty is notably asymmetric. Raising the plume above the assumed layer produces a substantially larger positive bias than the negative bias produced by lowering it. This asymmetry arises because the radiative signal depends on the thermal contrast between the emitting surface and the air containing the NH3.
This term is not propagated into the reported error bars. We now state this explicitly in the manuscript. Unlike the propagated uncertainty terms, the plume-height effect represents a scene-dependent systematic bias whose sign and magnitude are not known a priori. Combining it in quadrature with random errors would therefore be misleading. We regard plume-height uncertainty as the largest unquantified contribution to the source-rate estimates and discuss it explicitly in the revised Discussion section.
The influence of plume extent on the IME calculation is more limited because the integration is performed over the segmented near-source plume rather than over the full downwind plume extent. For the Khor Al Zubair example shown in Fig. 8, the plume mask spans approximately 5 km. At the observed wind speed, this corresponds to roughly 0.6 h of transport time. Whether the plume persists for tens of kilometers beyond the integration domain therefore does not directly affect the IME calculation. However, any chemical loss occurring within the integrated plume remains unaccounted for, meaning that the reported source rate should be interpreted as a lower bound in that respect.
Question 2, the air temperature and its uncertainty
Which meteorological database does the acquisition air temperature estimate come from, and what is the uncertainty? Of course the air temperature is critical because the absorption depth is proportional to the temperature contrast so air temperature error will alias directly into IME. A buoyant plume at elevation could also be in air that is a different temperature than the at-surface estimate would suggest. How serious an issue is this?
The air temperature used in the retrieval is taken from the ECMWF operational IFS analysis at the time and location of the ASTER acquisition. The submitted manuscript incorrectly referred to ERA5 in several locations, and this has now been corrected throughout, including in the data-availability statement. This correction is not merely editorial: the operational analysis underwent changes in horizontal resolution over the ASTER record, meaning that the spatial representativeness of the temperature field is not constant throughout the time series. We now note this explicitly in the methodology section.
We agree that errors in air temperature propagate directly into the retrieved NH3 column. We therefore quantified this sensitivity explicitly. A 1 K error in the prescribed air temperature changes the retrieved NH3 column by approximately 4-10%, with the largest sensitivity occurring under weak thermal-contrast conditions where the split-window ratio is least well constrained. Because NH3 column amount is the only free parameter in the inversion, there is no additional retrieval parameter that can absorb this discrepancy. This contribution now appears as a dedicated entry in the uncertainty budget.
We also agree that a buoyant plume may reside in air whose temperature differs from the near-surface value prescribed in the retrieval. The air temperature is imposed at the lowest model level, approximately 60 m above the surface. A plume transported above this level will generally reside in cooler air than assumed by the inversion. This effect biases the retrieved NH3 column high and therefore acts in the same direction as the plume-height sensitivity discussed in Question 1. We therefore do not treat the two effects as independent uncertainty terms.
At present, however, we cannot robustly quantify the magnitude of this particular bias. The temperature profiles used in the LUT are not constructed from a fixed lapse rate. Instead, a reference profile defines the vertical structure, which is subsequently rescaled between an upper-level anchor and the prescribed near-surface temperature. Consequently, the temperature gradient through the NH3-containing layer varies across the LUT and generally becomes steeper at warmer temperatures. We therefore do not report a representative magnitude for this bias, as such a value would not be directly supported by the retrieval framework. The sign of the effect is understood and is now discussed explicitly in the manuscript, but its magnitude remains an unquantified systematic uncertainty.
Question 3, the error bars and the flux uncertainty
What do the black bars signify in Figures 9 and 11, are they 95% confidence intervals, one-sigma error bars, or something else? Also, how exactly is the flux retrieval uncertainty calculated? Does it account for the uncertainties listed by the authors? It seems like windspeed alone could lever the solution around by a factor.
The black bars represent one-standard-deviation uncertainties. This was not stated clearly in the submitted manuscript and is now specified explicitly in the figure captions.
The uncertainty calculation is based on a quadrature sum of relative uncertainty terms. Importantly, it does not include all uncertainty sources listed in the original manuscript. Three terms are propagated: radiometric noise in the retrieved NH3 column, which averages down over the plume mask; uncertainty in the background subtraction; and uncertainty in the effective wind speed. Other uncertainty sources, including wind direction, the effective-wind-speed parameterization, surface emissivity, and the vertical distribution of NH3, are not propagated. The submitted manuscript listed these sources together in a way that could be interpreted as implying that all were included in the reported error bars, which was not the case. We have revised the text accordingly and now distinguish explicitly between propagated and non-propagated uncertainties. A complete inventory is provided in Appendix B.
We agree that wind speed is the dominant contributor to the reported uncertainty. The relative uncertainty assigned to the effective wind speed is 50% for all scenes. This corresponds to the upper end of the 15-50% range reported by Varon et al. (2018) for IME-based flux estimates when meteorological database winds are used in place of source-scale wind measurements, with the largest uncertainties occurring under low-wind conditions similar to those encountered in our dataset.
Because source rate scales linearly with effective wind speed and because the wind term dominates the quadrature sum, the uncertainty bars in Figs. 9 and 11 are typically close to ±50% of the retrieved source rate. We now state this explicitly in the methodology section. These uncertainty bars should therefore be interpreted primarily as an estimate of wind-related uncertainty rather than as a complete error budget for the emission estimate.
Questions 4 and 5, the scene-wise radiometric correction
The most confusing aspect of the workflow to me is initial radiometric adjustment. It seems like this sort of manipulation could impact the accuracy of the flux estimate. Note that, because the correction is applied independently on all bands, the absorption depth for the split window approach will not, in general, be preserved. Modifying the measurement data to fit the model seems like a suspect way to deal with this parameter uncertainty, since our measurement is the best connection we have to the real world. A more conventional approach would be to fit the model parameters to fit the data, weighted by uncertainty in calibration and prior uncertainty in the model parameters. As it stands, the uncertainty induced by the radiometric update seems very difficult to interpret or quantify. A related issue, it seems unlikely that these errors are all uniform across the scene, as the radiometric update assumes. Fitting the update on disjoint portions of the scene independently might be a way to confirm that the update is, in fact, scene-wise constant.
We address these comments together because they concern the same methodological issue. We agree with the reviewer's underlying concern that a retrieval framework based on fitting model parameters under appropriate prior constraints would be more rigorous than the empirical radiometric correction adopted here. The revision does not claim otherwise: instead of asserting that the correction is harmless, it now quantifies the effect and carries it explicitly in the uncertainty budget.
The reviewer correctly notes that independently correcting each band does not, in general, preserve the split-window absorption depth. We therefore quantified the extent to which the correction alters the retrieval observable, and this analysis is now reported in Sect. 2.3.2. Because the correction is applied per band, the residual difference between the bands enters the split-window ratio directly. Expressed relative to the propagated noise-equivalent column, its median magnitude is 1.05 at Khor Al Zubair, 1.61 at Piesteritz and 2.72 at Tolyatti. The perturbation introduced by the correction is therefore comparable to the radiometric noise at the most favorable site and reaches a few times the noise level at the least favorable one.
This effect is not negligible and we do not treat it as such. Its magnitude nevertheless remains far smaller than a factor-level distortion of the retrieval, and it is insufficient to explain the detected plume signals on its own.
We also report two additional diagnostics. First, the residual is systematic rather than random in sign, being negative in 107 of 120 scenes, so it biases against detection rather than manufacturing plumes. This behavior is consistent with a systematic radiometric or forward-model offset rather than random measurement noise. Second, the magnitude of the residual shows no relationship with the retrieved source rate (Spearman ρ = −0.07, p = 0.59), so the correction is not preferentially enhancing the largest retrievals that underpin the study conclusions.
The reviewer also suggests testing the assumption of a scene-wide correction by fitting separate corrections to disjoint portions of the scene. We have done this. Refitting on a 3 × 3 grid of disjoint tiles gives a tile-to-tile spread of 1.10 times the noise-equivalent column, comparable to the scene-wide residual itself. The assumption of a spatially constant correction is therefore an approximation, and the spatial residual is now carried as a systematic term of similar size in the uncertainty budget in Appendix B.
The scene-wide correction is an approximation, but a quantified one: it leaves a residual of order the radiometric noise, signed against detection, and uncorrelated with the retrieved source rates, which remain dominated by the assumed effective wind speed. Improved characterization of the correction is listed among the priorities for future work.
Citation: https://doi.org/10.5194/egusphere-2026-2893-AC4
-
AC4: 'Reply on RC3', Lidewij Tijhuis, 19 Aug 2026
Viewed
| HTML | XML | Total | BibTeX | EndNote | |
|---|---|---|---|---|---|
| 147 | 74 | 18 | 239 | 13 | 13 |
- HTML: 147
- PDF: 74
- XML: 18
- Total: 239
- BibTeX: 13
- EndNote: 13
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
This paper reports the first NH3 satellite measurements made from two broadband infrared channels. The authors show that, even though the detection threshold of ASTER is high, under favourable conditions NH3 plumes from industrial facilities can be measured and quantified. The spatial resolution of ASTER is a key element here, as it can probe within-plume airmasses undiluted (unlike CrIS or IASI), pushing the ammonia concentrations above the detection threshold. The results of this paper are very important for the design of future NH3 satellite instruments and the monitoring of mega-emitters.
The science in the paper is of high quality, robust and sound, and I have mainly minor comments and suggestions (see below). My main comment however, and the reason why I recommended major revision, is one of style. Overall, the paper is very repetitive and therefore unnecessarily long. For a reader it is difficult to maintain focus when the same points are repeated over and over; this ultimately dilutes the message rather than reinforcing it. As an example, the first page of the Discussion contains no new information, and is just repeating what was said before, and it is then repeated again in the conclusion. Discussion + conclusion can probably be covered in ~1 page. Some examples of near-identical repetitions are given below. I estimate the paper could be cut by roughly a third without any loss of content, with a gain in clarity and impact.
1. The sentence on effective emissions (x10, even without counting the figure captions)
Lines 16-18: "source-rate estimates derived with the Integrated Mass Enhancement method are interpreted as instantaneous effective estimates for successful plume scenes, not as annual mean emissions or continuous facility-average emissions"
Lines 204-205: "the derived source rates should be interpreted as effective emissions over the observed plume extent rather than as total emissions under all transport conditions"
Lines 344–345: "In contrast, the ASTER values reported here are successful-scene instantaneous IME estimates and should not be interpreted as annual mean emissions"
Lines 374–375: "The source-rate statistics reported here describe the distribution of instantaneous source-rate estimates under favorable observing conditions and should not be interpreted as annual mean emissions."
Lines 377–379: "The successful-scene mean and interquartile range should therefore be interpreted as descriptive statistics of ASTER plume snapshots rather than as a long-term facility-average emission estimate."
Lines 392–393: "The resulting values should therefore be interpreted as effective source rates for the observed plume extent, not as continuous annual emissions."
Lines 407–408: "The resulting source-rate statistics should be interpreted as successful-scene instantaneous estimates, not as annual mean emissions."
Lines 427–428: "This value should be interpreted as an instantaneous successful-scene statistic rather than an annual mean emission."
Lines 490–491: "The source rates reported here should therefore be interpreted as instantaneous effective source-rate estimates for the observed plume extent, not as annual mean emissions or continuous facility-average emissions."
Lines 539–540: "The successful-scene means reported here should therefore not be interpreted as annual mean emissions."
2. On the TC
Lines 153–154, 275, 283-284, 303-304, 454-455, 528
3. ERA5 winds
Lines 218–222, 357–360, 483–486 + the caption of figure 8
4. Comparison with published estimates is not validation
Lines 338-339, 385-386, 443-445, 466 + the caption of table 3
5. The band-13/band-14 setup is re-explained six–seven times.
Lines 56-57, 85-88, 92-93, 255-256, 263-265, 448-450, 524-526
etc..
Minor comments:
Line 9: space missing after "resolution."
Line 33: "Emission inventories .. exhibit strong temporal" I think you mean the inventories fail to capture the strong temporal variability, rather than that they exhibit it.
Figure 1: please increase the font size in these figures (especially the numbers)
Figure 2: Specie should be species (always plural). I would not put the TIR bands in the legend, as they are indicated on the figure. This would allow making the figure a bit wider and easier to read.
Line 101: "This can introduce band-correlated variability in emissivity and radiance fields." I would add retrieved emissivity (since emissivity is a property of a surface not from the measurement)
Line 104: Remove the last sentence (implied above)
Line 149: perhaps mention the lapse rate that was used
Line 169/327: No mention is made in the paper of the water vapour continuum, which is apart from clouds and emissivity the next largest contributor to changes in the baseline slope. Did the authors run simulations at different levels of humidity?
Section 3.1.1 and 3.1.2: These sections do not bring a lot of new information. Figure 6 and Figure 7 would probably be enough for the discussion on detectability/TC
Line 266: "This is a general limitation of ASTER TIR...". This is not specific to ASTER, the retrieval sensitivity to TC is general to all IR sounders (e.g. https://doi.org/10.34133/remotesensing.0142 )
Figure 6: are the units correct? As a ratio, is should be dimensionless. I would also add a contour at sqrt(2)sigma/B14, with sigma the radiance noise and B14 calculated at some temperature. This would delimit the detectable TC/NH3 region nicely.
Line 410: IASI box-model estimate used a lifetime of 12 hours, but this can easily be converted to the more realistic lifetime of 2-3 hours, giving 36 kt, well-aligned with the ASTER value. An advantage of ASTER and other high spatial resolution sounders is that they do not require an estimate of the chemical lifetime (I do not think this is mentioned in the manuscript).
General: Are all three point sources associated with the production of fertilizers? Did you have a look for NH3 emissions from livestock housings/feedlots? It would be good to address both questions in the revised manuscript.