the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
Urban atmospheric CO2 plumes from space – Part 2: Representation of fine-scale structures
Abstract. Quantifying urban CO2 emissions using spaceborne total column observations requires atmospheric transport models capable of resolving fine-scale spatial and temporal variability across metropolitan areas. Using the Grand Paris region as a testbed, we evaluated the sensitivity of simulated CO2 fields to spatial resolution and model physics within the Weather Research and Forecasting model coupled to a high-resolution fossil fuel emissions inventory. Simulations were conducted at mesoscale (900 m) and Large-Eddy Simulation (LES) resolutions (300 m and 100 m), and evaluated against dense urban CO2 observations from the Paris Picarro and rooftop mid-cost sensors network. Horizontal resolution strongly affects plume morphology: high-resolution simulations produce sharper gradients, higher peak mixing ratios and more realistic temporal variability. LES simulations capture localized enhancements and intra-urban variability more accurately than the 900 m mesoscale run, which systematically underestimates spatial contrasts by a factor of 2–3. Aggregation experiments indicate that averaging LES simulations to 900 m significantly modifies plume magnitudes and spatial gradients. Domain-mean relative errors for near-surface CO2 reach 29–34 % in winter and 59–61 % in summer, while column-integrated XCO2 shows 21–22 % errors in winter and 58–59 % in summer, with maxima exceeding 70 %. While LES improves the representation of urban CO2 variability, it increases sensitivity to local wind and inventory uncertainties, meaning misplaced point sources can cause large local biases. These findings emphasize the need for inversion frameworks able to incorporate scale-dependent transport and inventory uncertainties to ensure unbiased top-down emission estimates and accurate spatial attribution within city domains.
- Preprint
(17817 KB) - Metadata XML
- BibTeX
- EndNote
Status: open (until 29 Sep 2026)
- RC1: 'Review of egusphere-2026-2581', Anonymous Referee #2, 17 Sep 2026 reply
Viewed
| HTML | XML | Total | BibTeX | EndNote | |
|---|---|---|---|---|---|
| 136 | 85 | 61 | 282 | 67 | 69 |
- HTML: 136
- PDF: 85
- XML: 61
- Total: 282
- BibTeX: 67
- EndNote: 69
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
General
This study investigates the performance of weather and CO2 simulations based on the WRF model with a large dataset of meteorological and CO2 observations. The key focus is on the influence of horizontal model resolution on the performance.
The investment in evaluating transport model performance is commendable and the general modeling approach is sound. The large and diverse sensor network analyzed here is a great basis for simulation evaluation. However, in my opinion, both the scientific analysis of the model results and the presentation have severe shortcomings.
The manuscript is much too long and for many sections it's unclear what the key messages are. These shortcomings also make the manuscript very hard to follow. Many sections contain repetitive statements and concluding statements that are not always central to the topic of the manuscript. Mostly, the analysis of CO2 results (Sections 3.4-3.9, 4, 5) suffers from these problems, less so the sections about meteorology (Sections 3.1-3.3).
In some sections, especially towards the end of the manuscript, I found myself making a large number of specific comments because I disagree with interpretations of the figures or tables, find the language too vague, and towards the end simply didn't understand what some statements mean. I wonder whether I simply did not understand so many statements. In that case, the authors should be able to clarify what is meant with revisions that address each point. However, the way I see it, the combination of issues - large number of specific comments, a lack of clarity about messages, possible misinterpretations and repetitious statements - could make addressing all specific comments counterproductive because it risks making the manuscript even longer than it already is. Instead, it should be significantly shorter. I guess that the points could be made in half the length. Therefore, in my opinion, the authors should consider a fresh start of the analysis and the text, especially of the CO2 results, discussion and conclusions. Here is how I would go about that:
- Identify the performance metrics they want to investigate (e.g.: metrics that are relevant for inverse modeling - mean CO2 bias was omitted, absolute differences instead of relative differences were emphasized)
- Evaluate meteorology and CO2 consistently across cases, such that CO2 performance can be linked to meteorology performance (e.g.: different resolutions / summer vs. winter / day vs. night - currently, the winds cannot be linked to the CO2 day/night differences)
- Collect the key messages and present only the results needed to demonstrate them. Do not include every possible way of looking at the data if it all yields the same conclusions (which is often the case now - I think most of the "qualitative sections" 3.4-3.7 can be removed)
- Link CO2 performance to meteorology performance. Be thorough in the attribution (currently often PBLH is cited for CO2 performance, wind is neglected, as are the accuracy of CO2 fluxes and boundary conditions)
- Always keep the goal in mind (e.g. characterize transport model errors relevant for / find a good setup for inverse modeling). Only include analyses/statements that contribute to this goal.
Despite all the shortcomings, the overall modeling and analysis methods are sound. Therefore, I think that this manuscript has the potential to be revised such that it makes a valuable contribution towards characterizing and improving CO2 transport modeling for the Paris region. Therefore, I recommend considering the manuscript for publication after major revisions of the analyses and their presentation. I would consider reviewing a revised version of the manuscript if it's significantly shorter than the current version.
Modeling approach
A key difference between the simulations is that the mesoscale (900m) domain was nudged towards observations using the fdda technique, while the LES domains (300m, 100m) were not. I suspect that this "free-running WRF" configuration of the LES domains could be the reason for higher wind speed and direction errors in the LES simulations compared to the 900m simulation, and thus be very important also for the CO2 performance. The authors do not acknowledge or discuss this difference in the model setup in the analysis of the results, recommending in the conclusions to abandon high-resolution wind simulations, which I do not agree with based on the presented results. Furthermore, the authors attribute performance differences to other differences among the simulations (PBLH), which in my opinion risks misinterpretation of the results. These shortcomings should be addressed by eliminating nudging as a difference among the simulations. I'm aware that the authors refrained from nudging in the LES domains deliberately, as to not distort the turbulence simulation inside the boundary layer. I don't know whether this is a valid concern because observation nudging settings allow setting it to act smoothly across larger distances, so perhaps the influence is acceptable. But if it's unacceptable, the disturbance of the boundary layer dynamics could be minimized by employing nudging to the driving meteorology instead of observations because it can be restricted to model layers above the PBL in WRF. The alternative to new simulations is to incorporate the nudging differences into the interpretation of meteorology and CO2 results throughout the manuscript, where it is currently neglected.
About the 300m simulation: Doubrawa and Munoz-Esparza 2020 (doi:10.3390/atmos11040345) found that 333m is too coarse for an LES simulation and should rather employ the Shin-Hong PBL scheme. Here, the 300m simulation performed worst w.r.t. observations - suggesting that the findings of Doubrawa and Munoz-Esparza may be applicable to the case studied here, and that the 300m simulation could perform better when not run in LES mode, but with the Shin-Hong PBL scheme.
Apparently, the anthropogenic CO2 emissions that were used as input were constant in time. In reality, anthropogenic emissions are variable in time, which could contribute to biases in the simulations compared to the observations. Temporal profiles (diurnal, weekly, seasonal) that approximate the temporal variability of CO2 emissions are available in the literature and should be taken into account.
It's unclear how model XCO2 at night were obtained. The text and the interpretation mentions TCCON averaging kernels, but those are not defined for night-time observations. If anything, it would make sense to apply averaging kernels of active sensors like MERLIN to the night-time results.
Analysis of CO2 results
The topic of the study is the impact of model resolution on model performance. But in the analysis of CO2 and XCO2 results, the texts often compare summer against winter, and several sections conclude with the statement that concentrations are higher in winter than in summer because of the lower boundary layer height. This observation does not tell the reader much about which simulation performs best. It should suffice to state this result once.
A major shortcoming of the discussions of CO2 results is that it often fails to take into account wind speed differences among the resolutions, seasons and daytime, as well as flux differences and boundary conditions in time (at least the biogenic fluxes must have had a temporal resolution). Instead, in many cases only the influence of the PBL height is mentioned. Thus, some results may be misinterpreted.
The authors repeatedly state that shallow boundary layers lead to higher discrepancies because it leads to higher signals. In my opinion, relative error is a much more important measure because it relates more directly to flux estimation errors than concentrations themselves. So, what I would find more informative - and would stress more - is a relative metric like E_sym, or mean bias in percent. There is some analysis of E_sym here, but only to compare model resolutions, not models and data.
The accuracy of flux estimation is the key motivation for this study. But in the analysis of CO2 and XCO2, mean biases are not discussed, although they have a significant impact on the mean bias of flux estimation. Several figures indicate that there are mean biases both between models and data (Fig. 9) and among the different model resolutions (day vs. night, Figures 12 and 14). In my opinion, mean biases should be analyzed to enhance the value of the study.
The authors miss the opportunity to link the performance of the modeled meteorology to the performance of the modeled CO2 observations. Suggestions: One could investigate the relationship between CO2 probability density function and wind and PBLH performance. The peak of the CO2 pdf is too low compared to observations, a bit less so for the 900m simulation in winter (Fig. 10a-d). Winds are always overestimated except for the 900m simulation in winter (Table 2) - perhaps that overestimation could contribute to the underestimation of CO2 enhancements in winter? However, this argument doesn't really translate to summer, where differences among simulated CO2 pdfs are small despite similar wind speed bias differences. Perhaps because the PBLH biases are different in summer than in winter? At the one station shown, the PBLH of the 900m simulation is always higher than that of the 300m and 100m simulations (Fig. 5). That would result in underestimated enhancements by the 900m simulation, contrary to the results in Fig. 10. Perhaps the broadening due to the lower resolution counteracts this effect by increasing the probability to intersect point source plumes.
Sections 3.4-3.7 contain qualitative descriptions of case studies but lack quantitative analyses. Especially 3.4 and 3.7 are often vague and I don't agree with all statements. They are not always about simulation performance, but often about differences among the cases (summer/winter), which does not contribute much to the topic of the manuscript. My specific comments often ask for numerical quantities, but I suggest to not expand them overall, but remove the sections that only contain repeated conclusions or significantly shorten them while adjusting the focus towards quantitative analyses.
For Sections 3.8 (statistical analyses) and 3.9 (aggregation), while the figures are informative, the text is lengthy, sometimes vague, and in my opinion does not always accurately describe what I see in the figures. I make concrete suggestions below, but overall, the text could and should be much shorter.
In Sect. 3.9, the authors interpret the "impact of aggregation" of results. It's unclear to me what is meant by "impact of aggregation", because the section compares the simulations of different native resolutions, with all their performance differences, and it seems that the higher resolutions are aggregated simply to enable the comparison. To me, studying the "impact of aggregation" would entail a comparison of the performance e.g. w.r.t. data of native vs. aggregated simulation results. Though I'm not sure that such an analysis would be very valuable - I believe that the differences among the simulations should be very similar regardless of aggregation. See also my first comment on Sect. 3.9.
Specific comments
Lines 47f: "We further compare native kilometer-scale simulations with spatially aggregated LES outputs to quantify potential biases introduced by spatial averaging." What is the motivation for the aggregation analysis? The sentence sounds as if the authors expected a degradation of performance due to averaging. Then why would one try spatially averaging a higher resolution simulation, when a lower resolution simulation is cheaper? Please clarify.
- Figure 1b: Stations close to d05 boundary. Probably the effective resolution is lower than the 100m because turbulence needs to build up from the inflow edge (for wind-driven turbulence) - see the works abound the cell perturbation method for WRF-LES (e.g. doi:10.1175/MWR-D-18-0077.1). Consider including this in the discussion.
Line 101: "The outer domain (8.1 km) lies within the gray zone of convection" - So do d02 and perhaps d03. Cite a reference for which resolutions the gray zone spans.
Lines 105f: "No FDDA is applied in the LES domains to prevent numerical instabilities associated with explicitly resolved turbulence". Observation nudging can be set to act smoothly over wider areas, so I suspect that it may be possible to employ it in LES mode as well. But I don't know the literature about this topic. Please provide a reference that justifies avoiding nudging here? As stated in the general comments, I think this difference - and neglecting it in the interpretation of results - is a major shortcoming of the study.
Lines 125-129: These statements - as well as Sections 3.4-3.8 - neglect another consequence of ignoring the vertical distribution of plumes: plume displacement (e.g. Brunner et al. 2019, www.atmos-chem-phys.net/19/4541/2019/).
Lines 140ff: In Sect. 2.3, please state the number of observation locations not only for the weather stations, but also the other datasets.
Lines 140ff: State the data selection especially for the CO2 observations. I guess data of all 24 hours of the day were used. Any selection according to quality criteria?
Lines 184f: "Discrepancies between modeled and observed PBLH should not be interpreted at face value, given the fundamental methodological differences underlying each estimate". Perhaps instead of "at face value", a more specific formulation along the lines of "Discrepancies ... do not necessarily indicate model bias, given ...". But do both methods not attempt to quantify the same thing? For a well developed boundary layer, the results should be close, no?
Lines 279-281: "Although LES explicitly resolves turbulent structures, their intermittent spatial and temporal characteristics may not coincide with point measurements, thereby increasing error metrics even when variability is physically realistic". Yes, but the PSD - a measure of the variability - doesn't match (in summer). Also, please explain why the PSDs in summer are overestimated by the higher resolution simulations. I'm surprised by the result. If anything, I would have expected an underestimation by the coarser resolution, which is the case in winter. The order of model PSD at high frequency is expected (PSD of coarse model < PSD of fine model), so it seems like there could be an underlying bias common to all three simulations in summer.
Lines 282ff: The discussion in Sect. 3.2 completely omits a point that could, in my opinion, be key for explaining the results: Rather than the horizontal resolution, I think the reason why the 900m (d03) winds might perform better than the higher resolution (d04, d05) winds is that d03 was nudged towards observations, whereas d04 and d05 were free-running WRF. Previous studies showed the importance of tying WRF meteorology to meteorological observations. For example, Ho et al. 2024 (doi:https://doi.org/10.5194/gmd-17-7401-2024) found that WRF nudged towards ERA5 outperformed free-running WRF (their domain size and resolution and nudging technique were different, but I believe the result could be similar here).
Lines 282ff: Only winter was analyzed in this section, what about wind profiles in summer?
Lines 316ff: Mostly the PBLH of one station was analyzed, though the start of the section claims it's four. What about the other three stations inside d03? In Fig. 5d, it looks like "MEUD" could show a different behavior than the station shown here. Average diurnal cycles for the other stations may provide insight about the different CO2 pdfs in Section 3.8.1.
Lines 316ff: The section misses a concluding statement about the performance differences of the simulations. The summer in particular seems hard to interpret, though. It seems that no simulation consistently performs best w.r.t. PBLH in summer.
Lines 342f: "The 900 m simulation systematically overestimates PBLH along most of the transect, consistent with its positive daytime bias in winter". To me, it looks more like that overestimation is limited to three out of the seven stations, whereas for the other four (i.e. a small majority), the difference among (all) simulations and observations are much smaller.
Lines 352-392: Please comment on a striking difference among the resolutions: in the 900m run, the plumes are much shallower than in the 300m and 100m runs. It looks like there is a big difference in vertical mixing between the simulations.
Line 355: I think that the term "aggregation error" does not apply here, because it refers to coarse resolution state vectors in greenhouse gas flux estimation. The description given here fits "representation error", so I would stick to that. For the terminology, see also Ciais et al., 2010 (doi:10.1007/s10584-010-9909-3)
Lines 361-370: I suggest to strongly shorten this paragraph. The only relevant information in here is that the transect was chosen so that it intersects strong point sources and their plumes.
Lines 377-382: Here follows the interpretation of the cross sections (Fig. 6c and d), but I don't find it very useful. The paragraph mostly compares summer and winter, rather than the different model resolutions. Looking at the figures, I don't agree with all of the statements. I would prefer switching to a numerical analysis of the case study. No significant conclusion is drawn about the different model resolutions. Let's go through it statement by statement. These comments _could_ be commented on point by point by the authors, but it would make the text even longer and they may also choose their own approach rather than follow all of these suggestions point by point.
- "Under winter conditions, characterized by a relatively shallow boundary layer (PBLH reaching approximately 500–600 m), CO2 accumulates within the confined lower layer, leading to enhanced concentrations throughout the boundary layer depth." - Ok.
- "The limited vertical mixing strengthens horizontal gradients ..." - The statement ("strengthens") seems to be about the difference between summer and winter. Stronger gradients are indeed expected with lower PBLH (winter), but the difference in dynamic range between winter and summer is hard to judge based on Fig. 6c and d because of the different color scales. I suggest to use the same dynamic range in both color scales, but it would be even better to quantify these horizontal gradients to corroborate the statements (e.g. provide histograms of all pixels within the pbl, compare IQR or SD or min-max and how heavy the tail at the high end is)
- "... and accentuates differences between model resolutions." - Vague: I see the differences in model resolution as clearly in the summer case as in the winter case. See the suggestion for a quantitative analysis above.
- "In contrast, during the summer case, the deeper and more convective boundary layer promotes stronger vertical mixing, ... " - The plumes seem indeed to reach higher in the summer case. But again, it's hard to say - different dynamic range, different background mole fractions within the PBL. This statement should be corroborated by quantification. For example, compare the height a.g.l. below which, say, 90% of the plume mass is contained.
- "... resulting in a more vertically homogeneous distribution of CO2 within the PBL." - I don't follow. If anything, I see more structure downwind of the plumes throughout the PBL in the summer cases.
- "Consequently, elevated concentrations remain primarily confined near the emission source, while background concentrations appear more uniform away from the plume core " - Vague again. What is meant by "confined" - plume mass is always transported with the wind. The statement seems to be an interpretation of visual contrast differences between the two figures, which don't share the same dynamic range. So I think the statement could be removed. But if the author consider this feature to be important enough to warrant additional work and space, it could be visualized by plotting the difference between cross-sections (900m-300m, 900m-100m, 300m-100m) to see how large the spatial extent and magnitude of differences is, and to see differences between summer and winter. Also, the horizontal distribution is of course strongly affected by wind speed, but wind speed isn't mentioned here.
Lines 406f: "Because the transect does not necessarily pass exactly through the grid cell of maximum concentration at higher resolution ..." - Or if the grid cell with the source "starts" later along the transect (i.e., if the source is not in the "first" third of the coarser grid cell)
Lines 413-415: "However, near the source region, discrepancies between modeled and observed values can be large when a station lies within a neighboring grid cell rather than the one intersected by the transect." - Add which stations. From Figures 6 and 7, I guess HBO and TF1. Also EIF (far from any source, but in the plume).
Lines 393-441: Consider shortening Sect. 3.5. In my opinion, the section is much too long. The transect coordinates should be moved to the previous section, where the transects are first shown. Apart from that and the station location description, I think all relevant points would be covered by the following paragraph:
Coarser model resolutions distribute the tracer over more space (Fig. 6 a and b). For observations close to sources and inside a plume, these differences can be large, as shown by Fig. 7: The transects in Fig. 7 demonstrate that higher-resolution simulations produce sharper and higher peaks. Peaks in the coarser resolution simulation can be broadened also towards the upwind direction of the source due to numerical diffusion. Simulated observations show the largest differences when the station is close to a source and/or inside a simulated plume (Fig. 7; HBO, TF1 and EIF in winter and NEY in summer). Since these model differences are related to geometry and numerical diffusion, they would also occur independently of differences in transport model physics and dynamics.
Lines 451-458: As the paragraph states, the features described here were shown before ("consistent with the plume and transect-based CO2 gradient analysis"), so it seems that the analysis does not add much value.
Lines 463-465: Add that wind direction performance is worse in the 300m and 100m domains than in the 900m domain.
Lines 471f: "under certain conditions": Be specific. E.g., "depending on plume displacement". Also, it's possible that the emission inventory erroneously introduces peaks. I wonder whether the cases where modeled peaks are not observed at all are sometimes cases where a nearby power plant (e.g.) was offline, e.g. for maintenance.
Lines 473-498: Similarly to Sect. 3.4-3.6, Sect. 3.7 contains little quantitative information. As in Sect. 3.4, the qualitative descriptions of figures offered here are often vague and in some cases I don't fully agree with them. It describes differences between simulations and observations but makes no statement about which simulation performs better. The conclusions (lines 491-498) could stand alone without the remainder of the section before it. I suggest to significantly shorten the section by focusing on the differences between the simulations and the quantitative performance difference to the observations. As in Sect. 3.4, I add individual comments on the section below. However, addressing them could lengthen the section further. Therefore, the authors may choose to not address all the individual comments but write a new shortened version.
Line 482: I only see a larger variability of the 100m simulation in the 16 January case.
Line 483: "On 11 January, the simulations alternate between over- and underestimation relative to the TCCON observations". Not really, unless the statement refers only to the fluctuations in the morning (much smaller than later variability) - unclear whether this is meant. I would rather say that the morning observations are reproduced well, then they mistime the XCO2 rise around 12:00 by about half an hour.
Lines 486f: "On 16 January, all model configurations systematically underestimate XCO2 compared with the observations". Except at 11:00 and from 14:30 to 15:00, where the match seems to be pretty good. Why not say "the simulations fail to reproduce the observed peak between 11:00 and 14:30".
Lines 486-488: "This behaviour is consistent with the underestimation of near-surface CO2 previously identified at JUS in Section 3.6, suggesting that differences in near-surface concentrations may also influence the integrated column". Add a cross-reference to the figure panel (Fig. 8a). Also, I'm not sure surface CO2 and XCO2 show the same signal: it seems to me that the ground CO2 and XCO2 observations have a different timing: XCO2 starts at 11:00 rises until 12:00 (Fig. 9a). However, ground CO2 drops in the same period. Assuming the x axis is correct in both plots, I think the observed peaks in CO2 and XCO2 - both not (entirely) reproduced by the observations - are not the same signal. Could the later XCO2 peak stem from higher in the atmosphere than the earlier CO2 peak? Then the offsets in the simulations would have different origins instead of indicating the consistency claimed in the statement.
Lines 489f: "During the summer cases (07 and 09 June), the differences between model resolutions are relatively small. All simulations slightly underestimate XCO2 compared with TCCON, although the magnitude of the bias remains limited". On the contrary, they overestimate the observations. So, is there a simulation that performs better? The model differences are "relatively small" compared to what? The bias "remains limited" by what criterion? In some cases, the bias between observations and models is much larger than the simulation differences. On 7 June, it's close to 1 ppm sometimes (I guess? hard to see without a difference plot), which is not negligible e.g. for regional inversions and larger than the CO2M accuracy requirement for XCO2. The two peaks observed on 9 June are completely missed by the simulations, which is not mentioned in the text. This comment is an example for why I think the qualitative result sections (3.4-3.7) are not very informative. Rather than lengthening them by answering the questions I pose here, focusing on the statistical analyses conveys more information with fewer words.
Line 509: State the concentration of the peak in the pdfs of the observations and the 900m simulation, so it can be compared to the 430 ppm of the 300m and 100m simulation.
Line 512: Also state the concentration of the peak of the pdf of the observations.
Lines 514f: "finer resolution suppresses the high-mixing ratio episodes that broaden the observed distribution". This statement needs to be refined. It contradicts the paragraph about the logarithmic plots, which correctly states that the pdf of the 100m simulation reproduces the highest mixing ratios best.
Lines 516-519: The text does not convey that the simulated pdfs are skewed towards lower values, while the minima of the distributions quite accurately coincide with that of the observations. This means that there is a mean bias in the enhancements. Could it be related to mean bias in PBLH? Wind? Emission inventory?
Line 526: To put the 5*10^-4 into perspective, state somewhere in the methods the total number of observations.
Lines 526-528: "They likely arise from the accumulation of local emissions combined with insufficient mixing in stable boundary layers, a well known challenge in winter, where fine-resolution simulations sometimes overestimate episodic peaks and may generate artificial rare event variability". I agree that accumulation in shallow boundary layers is more of a challenge in winter, but the observed overestimation at the high end by the 100m simulation seems to occur much more often in summer than in winter (Fig. 10f and 10h, not 10e and 10g). So the interpretation of the figure doesn't seem to fit.
Line 532: "These distributional analyses are consistent with the plume, transect, and time series results". This statement shows that not all of these analyses have to be included in the paper.
Lines 533-534: "narrower central distributions, reflecting high resolution grids’ ability to preserve localized emission maxima" - I don't understand the connection of the first part to the second part of the statement. Why do "narrower central distributions" reflect the preservation of "localized emission maxima"? It rather seems to me that the narrower distributions (Fig. 10a-d) indicate an average low bias of the enhancements, regardless of the improved ability to model the highest emission peaks (Fig. 10e-h).
Line 535: "indicating that the central tendency is captured regardless of grid spacing" - Vague and too optimistic. What is meant by "central tendency"? I would summarize the summer results as "In summer, all simulations result in pdf maxima that are too low and too narrow, with little difference among simulations except that the finer resolutions capture the frequency of rare high-concentration events better".
Lines 535-537: "Across both seasons, ... the 100m simulation occasionally introduces minor spurious peaks at very low probabilities". According to Fig. 10, it does that only during summer. In winter, the maximum of the distribution is captured quite well (Fig. 10e,g). I guess the statement takes into account Fig. 8 - that the matching distribution doesn't mean individual peaks are always matched. Please refine the statement to be in line with the figures.
Lines 504-539: Sect. 3.8.1. - the analysis of Fig. 10 - fails to recognize that all simulations systematically underestimate the observations, as can be seen in Fig. 10a-d. It would be good to include a table with overall metrics for CO2, similar to Table 2. The 900m simulation apparently has a slightly smaller underestimation in summer. Why is that? Can this CO2 performance difference be linked to the meteorological performance?
Line 552: "summer, higher-resolution simulations improve correlations and reduce RMSD compared with the 900 m configuration" - Please state by how much. The improvement seems small in Fig. 11 - not sure it's statistically significant.
Lines 552-555: "For urban stations, as illustrated by the time series analysis in section 3.6, higher-resolution simulations improve the representation of peak events and short-term variability, but may also introduce occasional overestimations, which can in turn influence the statistical metrics represented in the Taylor diagrams". Please be more specific: How do the occasional overestimations influence the statistical evaluation? E.g. "influence" -> "reduce"
Line 556: "For the mid-cost network, a similar pattern is observed". The section goes on describing only features that are different from those exhibited at the Picarro stations. State first what exactly the similarity is.
Lines 556-557: "since most stations are located within Paris, the variability exhibits a broader spread rather than forming distinct groups". I don't follow. Above, the authors state that for Picarro stations within Paris the standard deviation is clustered around the true value, forming a tight group in the Taylor diagram. So why would the reason for the broader spread at the mid-cost sensor stations be that they are also all within Paris. Should they not cluster more tightly around the true value as well than the Picarros outside of the center as well?
Line 563: "particularly in summer". No, only in summer.
Lines 564f: "In winter, correlations and RMSD remain relatively similar across resolutions, and higher-resolution simulations can occasionally exhibit slightly larger RMSD" - RMSD is a summary metric - what ist meant by "occasionally" slightly larger RMSD? Perhaps for some stations? Not sure I would still call that "slightly" as there is quite a big spread (Fig. 11c). Also, judging from Fig. 11, at the midcost network the small degradation in winter is similar to the small improvement in summer, whereas the text sounds as if the summer improvement is larger. Please refine the sentence. The paragraph would benefit from adding the corresponding numbers to the text.
Lines 540-577: Overall, the section improves upon the previous sections by providing a statistical analysis of the data, although it would be better to include numbers in the text because they are not always clear to read off the figures. However, most conclusions drawn in this section are the same as already drawn in previous sections. This illustrates what I mentioned above - the whole results section could (and should) be much more concise. I recommend that the authors significantly shorten them.
Lines 578ff: The title of Sect. 3.9 states that the section is about the effect of aggregation. However, only the performance differences among simulations after aggregation to the same resolution are compared. Different metrics are used than in the analysis at the native resolutions (Sect. 3.8). Therefore, the analyses shown here present the combined impact of differences among the native simulations and aggregation. The isolated effect of aggregation remains unclear - in contrast to what the title says. Besides, I wonder what exactly the error of "aggregation" is supposed to be - it's pretty clear to me that the differences arise from the native simulations (winds, PBLH, ...), and I view the aggregation as merely a postprocessing step that makes the results comparable. In fact, the analysis sometimes refers to differences in simulated PBL height, indicating that the authors are aware that they are not studying the impact of aggregation here. Nonetheless, in many places, the interpretation is centered around the "effects of aggregation". Perhaps there is a misunderstanding on my part about what is meant by "effects of aggregation". Please clarify.
Line 583: What is "near-surface numerical noise"? I'm not aware of such a thing.
Lines 597-604: Here the authors claim that Fig. 12a (simulated winter daytime CO2 at 35m) is consistent with "a well-developed and efficiently mixed planetary boundary layer (PBL), as indicated by comparable daytime PBL heights across resolutions, with the 900 m configuration showing a slight overestimation". The only way this statement makes sense to me is in the comparison to night-time, because it doesn't make sense for the differences among resolutions. The higher PBLH of the 900m simulation should result in lower mixing ratios than in the 300m/100m simulations because the same sources are diluted over a larger volume - as recognized in lines 606-608 about the night-time results. However, Fig. 12a shows that the opposite is the case in the daytime. Why?
Lines 609-613: "The native 900 m simulation, by contrast, produces a smoother and weaker enhancement due to stronger effective mixing and spatial averaging. These pronounced differences arise because, while vertical resolution is consistent across simulations, the effective transport differs with horizontal resolution. Higher-resolution simulations better resolve local flow structures and turbulence intermittency, which limits vertical mixing in shallow nocturnal boundary layers and promotes CO2 accumulation near emission sources". These statements first mostly repeat the observation from the previous paragraph (the 300m/100m simulations have higher CO2 mixing ratios than the 900m simulation) and then gives a different explanation than before that is, in my opinion, inaccurate. Specifically why do "better resolve[d] local flow structures and turbulence intermittency" "limit vertical mixing"? I think the whole statement could be removed.
Lines 613-615: "It is worth noting that... " What does this statement have to do with the simulation differences? Perhaps that the CO2 is trapped below the 35m level shown here? The authors could easily confirm that instead of speculating.
Lines 615-617: "Additionally, the 35 m observation height corresponds to the second model level, making it sensitive to reduced mixing in the lowermost layer, whereas the coarser 900 m grid effectively
reduces concentration gradients due to horizontal averaging across the larger grid cells". I don't see what the first part of the sentence has to do with the section or the second part ("whereas ...") of the sentence, and the second part of the sentence has been stated numerous times already. I would cut this out.
Line 625f: "Differences are modest during well mixed daytime conditions, but become pronounced at night" - How is "modest" defined? I think it's important to also state that, while concentration differences might be higher during the night, relative differences (E_sym) have little difference between day and night. Because it implies that flux estimation error could actually be similar.
Table 3: Please include mean bias.
Line 626: The terms "plume confinement and "plume amplification" are confusing. "Confinement" is used throughout the manuscript sometimes to describe a small horizontal extent and sometimes a low PBLH. Please replace "confinement" with more precise terms throughout the manuscript (or define it properly). Similarly, "amplification" implies a comparison ("amplified compared to ..."). Here, it probably means "higher concentrations due to shallower PBLH", which would be much easier to understand.
Lines 627, 638, 644: "Fig. 13" instead of "Fig. 15", I guess?
Lines 640-643: "This behavior is consistent with the shallower nocturnal boundary layer discussed previously: enhanced confinement increases overall concentration levels in the aggregated high resolution runs, producing broader domain wide differences relative to the native 900 m simulation rather than sharply localized discrepancies". I can't follow this sentence. Previously, shallower PBL heights were correctly identified as a reason for higher concentrations. That argument was also used to explain that concentrations have larger differences at night (Lines 605ff). Here now we see that the night-time E_sym is smaller than the daytime E_sym, and this is supposed to be explained by the shallow nocturnal boundary layer as well? I guess the statement is rather about the difference _structures_, not their absolute magnitudes. Even then, I fail to see why a shallower boundary layer should result in "broader domain wide differences". Wind speed and direction differences should play a significant role in this comparison, but isn't mentioned in the section so far.
Line 644: "amplified" again - just say "larger" instead of "seasonally amplified".
Line 645: "up to 70%". I'm not sure anymore that the authors are still describing near-surface CO2, which is the topic of the section. First, they consistently refer to Fig. 15 (XCO2) instead of Fig. 13 (surface CO2, nowhere referenced in the text). At first I thought it's just a mix-up of the cross-reference. But here, they say the surface CO2 E_sym (summer, daytime) is "up to 70%". According to Table 3, the maximum is 88.4%, while that of XCO2 is 73.18% - much closer to the claimed value here.
Line 646: What's "background mixing". Perhaps it should say "vertical mixing"?
Lines 648-650: "This suggests that under certain stable conditions, aggregation can partially smooth extreme local maxima in a way that reduces proportional differences relative to the native coarse resolution simulation". Once again, the results don't distinguish between the effects of aggregation and the effects of the different simulation dynamics. Therefore, the authors can't claim that the result is due to aggregation. The result here could also be due to wind direction differences not being the same during day/night/winter/summer. Plus, I'm not even sure what is meant by "extreme local maxima" (of E_sym? of the concentration?) or what it could have to do with "certain stable conditions". Therefore, I don't see the what value this statement adds to the manuscript.
Line 661: What's "coarser mixing" and why was it not given as a reason for the simulation differences in line 373
Lines 669ff: How were XCO2 values computed, especially for the night? For the comparison with TCCON earlier, averaging kernels from TCCON were used. But they are not defined for the night (in the absence of sunlight). It should also be stated somewhere that night-time XCO2 could be relevant for active sensors (e.g. Merlin, DQ-1), because it's confusing to look at night-time XCO2 when the only reference to column measurements so far was passive sensors (TCCON), which require sunlight.
Line 672: "Correlation coefficients indicate good overall agreement" - I'm less enthusiastic about correlations around 0.8. Relative differences (E_sym) are 20-60%, which is a bit less than for surface CO2, but could still have a large impact on flux estimation.
Lines 673f: "reflecting a stronger sensitivity to boundary-layer stability" - Once again: or it reflects other differences, especially winds.
Lines 680f: "XCO2 retrievals are less sensitive to near-surface layers due to the shape of the averaging kernel" - For TCCON, that statement depends on solar zenith angle (see Laughner et al., 2024). But once again, TCCON kernels aren't applicable to night-time simulations. What is the shape of the averaging kernels used here?
Lines 680ff: The paragraph does, once again, not consider wind speed as a factor that may explain concentration differences. On average, wind speed should be lower at night than during the day, thus not explaining the difference (it would have the opposite effect - it could rather contribute to explaining the surface CO2 differences between days and nights). Makes it even more curious what the shape of the averaging kernel is. I'll stop here pointing out the remaining instances where the omission of wind speed is critical.
Line 720-723: "These larger Esym values are primarily a consequence of the much weaker XCO2 signal in summer ...". I don't see why Esym is necessarily larger with weaker XCO2. Should smaller signals not also result in smaller differences, if all else were the same? Could the higher Esym in summer not be explained by transport differences - specifically, could transport differences among the simulations be larger in summer than in winter? Was this investigated? I also wondered if background variability plays a role, but apparently the background was subtracted for the analysis (Fig. 12).
Lines 741f: I don't agree that this double-penalty effect is specific to the gray zone. It should become more important for even higher resolutions. The authors seem to claim here that fully resolving Eddies would improve comparisons to in situ data, but due to the random nature of turbulence, this is not the case.
Lines 751-756: These statements create the impression that model differences are larger in winter than in summer. While this is true for concentrations, it's the opposite for relative concentration differences (E_sym), which are larger in summer than in winter and more important for flux estimation.
Lines 756-757: "This seasonal dependence is consistent with observational evidence reported in Mitchell et al. (2018) ..." The fact that shallow boundary layers enhance concentrations of surface-released pollutants has been well known long before 2018.
Lines 758-762: I think that the different PBLH methodologies only influence the comparison to PBLH observations, but these lines sound as if it could also influence other results (winds, CO2?). Please clarify. As far as I know, PBLH in WRF is a diagnostic variable and doesn't influence other simulation results. Also, these lines break the train of thought (the lines before and after belong together).
Lines 756-764: "As a result, differences between model resolutions become less pronounced. ... These results reinforce the conclusion that the benefits of high-resolution modeling are strongest under stable conditions when transport is less diffusive and plume structures are more distinct". I fail to see how the results demonstrated that. "Benefits" should mean "more accurately reproducing observations". The statement can be about results in 3.8.1, 3.8.2 or 3.9. Sect. 3.8.1 showed that in winter (more stable conditions), large parts of the CO2 pdf are reproduced better by the coarser simulation than the finer ones (except for rare high peaks, both in summer and in winter). Sect. 3.8.2 says that RMSD in winter is similar across all simulations, whereas it gets better for the high resolutions in summer. Variability is better in the higher resolutions both in summer and in winter (to a similar extent? Looks like it in the Taylor diagrams). Thus, the lines here contradict both Sect. 3.8.1 and 3.8.2. As for Sect. 3.9, it compares simulation results to each other but not to data, so it allows no conclusion about model performance.
Line 772: "may partly reflect representativeness effects rather than purely model deficiencies" - What is a "representativeness effect" rather than a model deficiency? I would remove the sentence - the key statement about high resolution has been made in the sentence before.
Lines 773-778: "Comparisons with the ground-based observation networks reveal a trade-off between simulated variability and conventional error metrics ... " - I find this statement hard to understand. Also, it seems to be mostly yet another repetition of statements already made elsewhere in the manuscript. How about replacing these lines with something like this: "Comparisons with the ground-based observation networks reveal that better matching simulated variability does not necessarily translate to better average matches of the data (e.g. RMSD)".
Lines 785f: "resolution effects propagate vertically through the atmospheric column". More like "resolution differences also affect column-averaged CO2" (the study doesn't demonstrate a "propagation" of effects from the surface to higher altitudes)
Line 795: This is not what "aggregation error" means in the community. See my comment on line 355.
Lines 797-799: "The persistence of these differences indicates that sub-grid plume variability contributes significantly to the simulated concentration field and cannot be fully represented by simple spatial averaging" - I don't follow. I would say the differences are due to the different winds, pbl heights, numerical diffusion, ... of the different resolutions. I don't see what "sub-grid plume variability" has to do with it.
Lines 799-801: "The larger Esym values observed in summer primarily reflect the weaker absolute signal associated with enhanced convective mixing. When plume amplitudes are small, even modest absolute differences introduced by aggregation translate into larger Esym values.". I disagree. All else being the same, I would expect that smaller plume enhancements also translate into smaller plume enhancement differences. I suspect that dynamic differences among the simulations (winds, pblh) are larger in summer than in winter. However, meteorology differences among the simulations were only analyzed with respect to data, not against each other (similar to the CO2 analysis in Sect. 3.9), so that remains speculative.
Lines 801-802: Another repetition of statements that were made before.
Lines 805f: "If the model fields are aggregated to match these footprints without accounting for sub-grid plume variability, significant representation errors may be introduced". This is not what "representation error" means. And the authors did not demonstrate whether the fine resolution simulations perform better with respect to data with or without aggregation, because that's not how Sect. 3.9 is set up. But as stated before, I'm not so sure the isolated effect of aggregation is relevant, as the differences shown here are probably mostly related to dynamic differences among the simulations. I also don't see what is meant by "sub-grid plume variability" here.
Lines 826f: "careful consideration must be given to spatial resolution, observational representativeness, and scale mismatch". And what exactly are these "careful considerations"?
Line 827: "the interaction between model resolution and observational footprint remains a key challenge" I don't understand what "interaction" means here.
Line 834: "may amplify biases related to PBL dynamics" I don't understand what is meant here. I guess it has to do with the higher concentrations in shallower compared to deeper boundary layers again. See my comments above on why I think that relative metrics are more important.
Lines 837-839: "Importantly, our results indicate that the differences observed between LES and mesoscale simulations are not primarily due to model physics, but rather to aggregation errors associated with coarse spatial averaging of heterogeneous emission fields". This statement is simply not supported by the analyses. It remains unclear to me what the "aggregation error" is supposed to be in this context. See my comments on Sect. 3.9.
Lines 840-842: "Thus, the discrepancy between resolutions reflects the inability of coarse grids to aggregate high-resolution emission variability, rather than a deficiency of the transport model itself". I don't understand what this sentence means.
Line 846: "... excessive surface CO2 accumulation" - Makes sense, but I don't remember this was shown anywhere in the manuscript. Figure 10 rather seems to indicate an overall underestimation of surface CO2, though it doesn't distinguish between day and night.
Line 848: "Our results highlight the importance of improved turbulence schemes and assimilation of vertical profile data". Which results specifically demonstrate that turbulence schemes need improvement? What's "vertical profile data"? Does this mean wind lidar data and that it was used for observation nudging in the 900m domain? Then this statement would finally be an acknowledgment that nudging could contribute to the differences between the 900m and the 300m/100m simulations.
Line 848: "Vertical and horizontal resolution effects are tightly coupled" The study did not study varying vertical resolution.
Line 849: "errors in near-surface transport propagate to the total column XCO2" The study did not demonstrate that XCO2 discrepancies are due to surface transport, nor any "propagation" from the surface upwards.
Lines 851f: "Model–observation agreement depends on how sensor locations align with grid cells". The study did not investigate potential effects of sub-gridscale position of observations on model-data-mismatch.
Lines 858-862, 865f: I disagree with the conclusion that a hybrid approach of coarse meteorology and fine-scale CO2 transport is the best way forward. It suggests that 100m and 300m simulations are in general incapable of simulating meteorology as good as a 900m simulation, whereas I believe that the better wind performance of the 900m simulation could be because it employed nudging (fdda) while the finer domains did not.
Technical corrections / minor comments
Line 10: "... modifies plume magnitudes and spatial gradients [compared to the native 900m run]"
Lines 11-12: Specify what the ranges mean. Often a range indicates min-max, but this is not the case here.
Line 21: "([e.g., ] Ye et al., 2020)"
Line 24: " ([e.g., ] Che et al., 2024)"
Line 32ff: Using the singular for "large Eddy simulation" feels off. I think the plural should be used for "LES". Or "LES models".
Line 37f and 76f: On citing Lauvaux et al. (2012): The WRF-GHG module, which is often used for tracer dispersion simulation, is "Beck et al., The WRF greenhouse gas model (WRF-GHG). techreport 25, Max Planck Institute for Biogeochemistry, 2011". How does the Lauvaux et al. "modification" of WRF differ from WRF-Chem / WRF-GHG? The reference doesn't give any details.
Line 63: Please clarify what is meant by "intensive" periods
Line 69: Please clarify what is meant by "complex land-atmosphere interactions"
Line 70: "representative" for what?
Line 72: I think instead of "seasonal variability", "temporal" or "spatio-temporal" variability would fit the statement better.
Line 115: Fix the reference "(Agency, 2024)"
Line 134: Fix the reference "(Eric, 2021)" -> "(Vermote, 2021)"
Line 168: Rather than "the vertical distribution of CO2", "the vertical average of CO2"
Line 210: "stratified" - rather "divided"?
Line 247: "Previous study" -> "A previous study"
Line 262: Just "metrics" instead of "verification metrics"?
Line 265: Please clarify what is meant by "second level timescales"
Fig. 5d: In the title of the panel, add that it shows winter results.
Line 355: Please clarify what is meant by "plume integrity"
Line 361: "The transect was defined ..." -> "The transects were defined ..."
Line 448: "Figures 8c and 8c" -> "Figures 8c and 8d"
Line 458: "." missing
Line 509: "distributions" -> "distribution"
Line 520: "the differences between model resolutions [at high mixing ratios] become more pronounced" (they are less pronounced in the center)