Status: this preprint is open for discussion and under review for Atmospheric Measurement Techniques (AMT).
An aircraft-based Bayesian inversion of Berlin CO2 emissions with explicit background treatment
Lukas Pilz,Christopher Lüken-Winkels,Alina Fiehn,Anke Roiger,and Sanam N. Vardag
Abstract. Quantifying greenhouse gas (GHG) emissions from urban areas is critical for assessing progress toward climate goals. Aircraft-based measurements provide valuable information on urban emission fluxes; however, their interpretation is often limited by uncertainties in atmospheric background concentrations. Here, we present an emission estimate of Berlin city emissions that explicitly includes background concentrations from a Eulerian background run as state vector elements in a Bayesian inversion framework. Using data from the 20th of July flight of the 2018 3DO aircraft campaign of the [UC]2 (“Urban Climates Under Change”) project over Berlin and high-resolution GHG and meteorological simulations using Weather Research and Forecasting (WRF) model, we derive a CO2 emission estimate for the city of Berlin. We find an optimized Berlin emission rate of 13.4 ± 4.6 MtCO2a−1. The emissions rate is about 48 % higher than the prior total annual emissions of the TNO anthropogenic bottom-up inventory (time-scaled emissions, 10.8 MtCO2a−1) and VPRM biogenic (time-varying, -1.7 MtCO2a−1) and about 30 % of the estimate from the traditional mass balance approach (44 ± 24 MtCO2a−1), while reducing the uncertainty by over a factor of five. We verify our posterior background concentration with independent concentration measurements above the boundary layer. This study provides an improved independent estimate of urban emissions and demonstrates that explicit treatment of background concentrations in inversions improves aircraft-based urban flux estimates.
Received: 30 Jun 2026 – Discussion started: 12 Aug 2026
Publisher's note: Copernicus Publications remains neutral with regard to jurisdictional claims made in the text, published maps, institutional affiliations, or any other geographical representation in this paper. While Copernicus Publications makes every effort to include appropriate place names, the final responsibility lies with the authors. Views expressed in the text are those of the authors and do not necessarily reflect the views of the publisher.
In the present study, the authors describe their Bayesian approach to constrain urban scale CO2 emissions in the city of Berlin using 2018 aircraft campaign observations. They explicitly solve one of the major inverse problems – background concentration estimation. The results suggest that the proposed method performs well for the studied campaign. It gives a flux estimate that agrees better with emission inventory data and has lower uncertainty in comparison to a standard but more simplistic mass balance approach. The optimized background concentrations are verified using independent concentration measurements above the boundary layer.
The manuscript is well structured and written in a clear, plain language. Overall, the work is in the journal scope and provides a relevant contribution to urban scale flux estimation methods. The major concern is a relatively small amount of data used for inverse experiments and to draw the conclusions about method performance. The authors should discuss this limitation and provide additional justification for why their conclusions would hold when more data is used. The study is worth publishing after minor revisions and clarifications about the described method.
Specific comments
Line 8: ‘The emissions rate is about 48% higher than the prior total annual emissions …’
I would suggest explicitly stating the emission rate over the given day, since CO2 fluxes have significant interannual variations.
Line 15: ‘Urban areas account for the majority of anthropogenic CO2 emissions …’
Please add relevant references.
Lines 38-41: ‘However, the spatial representativeness of the networks is constrained by the limited number of stations and their proximity to heterogeneous near surface sources within the urban canopy, which can strongly bias inferred city-wide fluxes and if not carefully treated can pose issues, as they are not able to directly and homogeneously sample the urban plume downwind of the city (Li et al., 2026).’
This sentence is quite long and contains several statements. I would suggest splitting it into several sentences. I agree with the statement that airborne observations can provide better spatial coverage in comparison to ground-based in-situ networks. However, the argument about proximity to heterogeneous near-surface sources biasing inferred fluxes is not clearly described. It would only hold if we are optimizing total city emissions and do not allow the inversion to correct intra-urban distribution. Otherwise, the stronger CO2 signal in in-situ observations should help us to better constrain those sources and their spatial gradients. I think the problem is rather related to atmospheric transport errors inside the canopy layer and the inability of the models to capture the air flow (plume direction, vertical mixing, wind speed etc.). However, this does not diminish the value of in-situ measurements. Rather, it incentivizes researchers to improve atmospheric models.
Lines 57-62: This is an important paragraph and could be expanded. I would suggest the authors mention that there is a difference in background definition in Eulerian and Lagrangian models. Eulerian transport models produce background concentration directly, since it can be transported explicitly as a separate tracer forced only by initial and boundary conditions. In contrast, Lagrangian particle dispersion models output only a sensitivity (footprint) field relating flux to concentration enhancement. In the latter case, the background concentration is not a native model output and must instead be assigned externally, e.g. by sampling a separate concentration field at the trajectories’ point of origin. Since this study uses a Lagrangian model (FLEXPART-WRF) for footprint calculation, it is important to make this distinction explicit, and to discuss the uncertainty arising from background definition in Lagrangian models and how it compares to other sources of errors as well as the emission signal. I would also suggest citing other studies and providing their estimates, e.g. Karion et al. (2021) find summer monthly-mean scatter of 1 to 3 ppm between different background definition methods, with hourly RMSE reaching up to 8 ppm. These errors, at least partially, represent an additional source of background uncertainty on top of errors coming from boundary condition products and model errors.
Lines 85 – 87: ‘Flights on the other days are not considered, as the urban plume was either not captured or the boundary layer was not clearly separated (18th, 23rd, 24th, 25th July).’
Most of the days were discarded, do the authors expect the limited number of usable flight days to be a common limitation for similar campaigns? Could the authors discuss how flight planning could be optimized to maximize the number of observations suitable for inversion, and for the future work, how could one estimate the number days required to constrain fluxes monthly/seasonally?
Line 96: Why were the MYJ and the YSU boundary-layer schemes selected?
Lines 99-100: The VPRM setup needs a clarification. Could you include the following information: Which VPRM code version was used? What input products for satellite indices (MODIS or Sentinel-2), temperature, radiation, land cover data did you use? Also, what is the spatio-temporal resolution of the resulting fluxes?
Line 103: Why did you decide to use that specific CAMS experiment? Have you tried using different CAMS products?
Lines 112-115: ‘Because we also perform a background estimation in parallel to the emissions estimation, we need prior background concentration information for all measurements. These prior background concentrations are provided by MACRO-2018, which includes a first estimate thereof using the previously mentioned background concentration field from outside the Europe domain.’
Please provide a description of how the background concentration is sampled for each observation here in the methods section or in the appendix.
Line 125: Did you consider re-scaling annual totals to correspond to the official self-reported Berlin emissions?
Line 137: How was the uncertainty estimate of 6.8 MtCO2a-1 derived? Is it just 75% of the prior that you describe in the Section 2.5.1?
Line 157: Why did you decide to also include separate regions 3 and 4 for Germany upwind and downwind? Does it have an advantage over having only three regions 0 – 2, i.e. by increasing the size of Berlin upwind and downwind to cover 3 and 4 regions, correspondingly?
Line 170: Does ‘CO2_BCKfield’ include the forcing from Berlin emissions?
Lines 187-194: Please provide the formula you used for computing the gain and average kernel matrices, and comment on how it follows from this formula that values outside [-1, 1] are considered non-physical.
Lines 236-237: ‘The model reproduces the general structure well up to about 2km but underestimates and smooths the mixing-layer height, likely due to limited vertical resolution (cf. Table 1).’
The wind performance seems to be the same between both altitude ranges, I suggest mentioning this in the text. Also, based on the wind profiles (Figure 2), it seems like the vertical structure is quite different between the schemes. Could you please provide possible reasons for this difference.
Lines 272-273: ‘This is achieved without increasing its correlation to the observations (prior and posterior correlation both 0.46).’
You mention this statement several times in the text, could you please provide a few comments in the text on why it is important to keep the correlation between measurements and background unchanged. The background concentration can also include plumes of elevated concentration in it because CO2 is a near-conservative tracer on synoptic scales, so the enhancements can be advected from outside of the model domain as well. How do you make sure those data are not used for the inversions?
Lines 292-293: Why are the gain matrix values for cell 1 significantly smaller than for cell 0? Whereas for the diagonal elements of the average kernel matrix they are much closer to each other. Also, what is the meaning of the sign of the gain matrix values? Why does the sign alternate throughout the day?
Figure 4. It would be good to include a similar figure, but that also includes posterior estimates when the background was not optimized, to directly show what the impacts are on the concentration timeseries.
Lines 323-324: Please add comments on how you think the inversion is able to effectively separate the background and emission signals from observations. Is it because some parts of the observation were unaffected by local sources and were correlating only with background adjustments?
Lines 339-341 and 362: ‘This substantial uncertainty reduction highlights the ability of the Bayesian Inversion framework with explicit background estimation to more efficiently extract the available information from the aircraft observations.’
‘We reduce the flux uncertainty by around a factor of five compared to the one in Klausner et al.
(2020)…’
How realistic do you think the posterior uncertainty estimate is? Can it be decreased by the inversion too much? Are there any other sources of uncertainty that could not be included in your study? Please add some comments to the Discussions section.
Technical corrections
Table 1. Please provide variable name descriptions in the caption.
Figure 2. The difference in wind speed profiles between the two model schemes is quite significant. Could you please change zorder on the profile plots and bring the smooth model lines to the front, otherwise it is difficult to see them behind fluctuating observations.
Figure 3 e). What do the 'Posterior' units represent? Please add an explanation in the caption.
Please make sure that the posterior uncertainty is consistent throughout the manuscript. In conclusions (line 360), it is specified as ± 4.3, while in the abstract (line 8) you write ± 4.6.
References
Li, J., Li, P., Han, P., Cheng, Z., Li, J., Zhang, T., Chen, D., Zheng, Y., Zeng, N., and Zhang, G. (2026). Advances in the design of urban CO₂ emission monitoring networks: a review. Carbon Research, 5, 3. https://doi.org/10.1007/s44246-025-00239-z
Karion, A., Lopez-Coto, I., Gourdji, S. M., Mueller, K., Ghosh, S., Callahan, W., Stock, M., DiGangi, E., Prinzivalli, S., and Whetstone, J. (2021). Background conditions for an urban greenhouse gas network in the Washington, DC, and Baltimore metropolitan region. Atmospheric Chemistry and Physics, 21, 6257–6273. https://doi.org/10.5194/acp-21-6257-2021
Klausner, T., Mertens, M., Huntrieser, H., Galkowski, M., Kuhlmann, G., Baumann, R., Fiehn, A., Jöckel, P., Pühl, M., and Roiger, A. (2020). Urban greenhouse gas emissions from the Berlin area: A case study using airborne CO₂ and CH₄ in situ observations in summer 2018. Elementa: Science of the Anthropocene, 8, 15. https://doi.org/10.1525/elementa.411
Lukas Pilz,Christopher Lüken-Winkels,Alina Fiehn,Anke Roiger,and Sanam N. Vardag
Data sets
Companion data to "An aircraft-based Bayesian inversion of Berlin CO2 emissions with explicit background treatment"Lukas Pilz, Christopher Lüken-Winkels, Sanam N. Vardag https://doi.org/10.5281/zenodo.20964395
Model code and software
Software for "An aircraft-based Bayesian inversion of Berlin CO2 emissions with explicit background treatment"Lukas Pilz, Christopher Lüken-Winkels, Sanam N. Vardag https://doi.org/10.5281/zenodo.20964324
Lukas Pilz,Christopher Lüken-Winkels,Alina Fiehn,Anke Roiger,and Sanam N. Vardag
Viewed
Total article views: 34 (including HTML, PDF, and XML)
HTML
PDF
XML
Total
BibTeX
EndNote
28
5
1
34
4
3
HTML: 28
PDF: 5
XML: 1
Total: 34
BibTeX: 4
EndNote: 3
Views and downloads (calculated since 12 Aug 2026)
Cumulative views and downloads
(calculated since 12 Aug 2026)
Viewed (geographical distribution)
Total article views: 18 (including HTML, PDF, and XML)
Thereof 18 with geography defined
and 0 with unknown origin.
Cities need reliable estimates of their carbon dioxide emissions to track climate progress. We combined aircraft measurements with meteorological computer simulations, improving the accuracy of the results. For Berlin, we found emissions were higher than a widely used inventory but much lower, and far more precise, than a traditional calculation. This approach can provide more trustworthy estimates to support climate action and policy.
Cities need reliable estimates of their carbon dioxide emissions to track climate progress. We...
Review comments of Pilz et al. (2026)
In the present study, the authors describe their Bayesian approach to constrain urban scale CO2 emissions in the city of Berlin using 2018 aircraft campaign observations. They explicitly solve one of the major inverse problems – background concentration estimation. The results suggest that the proposed method performs well for the studied campaign. It gives a flux estimate that agrees better with emission inventory data and has lower uncertainty in comparison to a standard but more simplistic mass balance approach. The optimized background concentrations are verified using independent concentration measurements above the boundary layer.
The manuscript is well structured and written in a clear, plain language. Overall, the work is in the journal scope and provides a relevant contribution to urban scale flux estimation methods. The major concern is a relatively small amount of data used for inverse experiments and to draw the conclusions about method performance. The authors should discuss this limitation and provide additional justification for why their conclusions would hold when more data is used. The study is worth publishing after minor revisions and clarifications about the described method.
Specific comments
Line 8: ‘The emissions rate is about 48% higher than the prior total annual emissions …’
I would suggest explicitly stating the emission rate over the given day, since CO2 fluxes have significant interannual variations.
Line 15: ‘Urban areas account for the majority of anthropogenic CO2 emissions …’
Please add relevant references.
Lines 38-41: ‘However, the spatial representativeness of the networks is constrained by the limited number of stations and their proximity to heterogeneous near surface sources within the urban canopy, which can strongly bias inferred city-wide fluxes and if not carefully treated can pose issues, as they are not able to directly and homogeneously sample the urban plume downwind of the city (Li et al., 2026).’
This sentence is quite long and contains several statements. I would suggest splitting it into several sentences. I agree with the statement that airborne observations can provide better spatial coverage in comparison to ground-based in-situ networks. However, the argument about proximity to heterogeneous near-surface sources biasing inferred fluxes is not clearly described. It would only hold if we are optimizing total city emissions and do not allow the inversion to correct intra-urban distribution. Otherwise, the stronger CO2 signal in in-situ observations should help us to better constrain those sources and their spatial gradients. I think the problem is rather related to atmospheric transport errors inside the canopy layer and the inability of the models to capture the air flow (plume direction, vertical mixing, wind speed etc.). However, this does not diminish the value of in-situ measurements. Rather, it incentivizes researchers to improve atmospheric models.
Lines 57-62: This is an important paragraph and could be expanded. I would suggest the authors mention that there is a difference in background definition in Eulerian and Lagrangian models. Eulerian transport models produce background concentration directly, since it can be transported explicitly as a separate tracer forced only by initial and boundary conditions. In contrast, Lagrangian particle dispersion models output only a sensitivity (footprint) field relating flux to concentration enhancement. In the latter case, the background concentration is not a native model output and must instead be assigned externally, e.g. by sampling a separate concentration field at the trajectories’ point of origin. Since this study uses a Lagrangian model (FLEXPART-WRF) for footprint calculation, it is important to make this distinction explicit, and to discuss the uncertainty arising from background definition in Lagrangian models and how it compares to other sources of errors as well as the emission signal. I would also suggest citing other studies and providing their estimates, e.g. Karion et al. (2021) find summer monthly-mean scatter of 1 to 3 ppm between different background definition methods, with hourly RMSE reaching up to 8 ppm. These errors, at least partially, represent an additional source of background uncertainty on top of errors coming from boundary condition products and model errors.
Lines 85 – 87: ‘Flights on the other days are not considered, as the urban plume was either not captured or the boundary layer was not clearly separated (18th, 23rd, 24th, 25th July).’
Most of the days were discarded, do the authors expect the limited number of usable flight days to be a common limitation for similar campaigns? Could the authors discuss how flight planning could be optimized to maximize the number of observations suitable for inversion, and for the future work, how could one estimate the number days required to constrain fluxes monthly/seasonally?
Line 96: Why were the MYJ and the YSU boundary-layer schemes selected?
Lines 99-100: The VPRM setup needs a clarification. Could you include the following information: Which VPRM code version was used? What input products for satellite indices (MODIS or Sentinel-2), temperature, radiation, land cover data did you use? Also, what is the spatio-temporal resolution of the resulting fluxes?
Line 103: Why did you decide to use that specific CAMS experiment? Have you tried using different CAMS products?
Lines 112-115: ‘Because we also perform a background estimation in parallel to the emissions estimation, we need prior background concentration information for all measurements. These prior background concentrations are provided by MACRO-2018, which includes a first estimate thereof using the previously mentioned background concentration field from outside the Europe domain.’
Please provide a description of how the background concentration is sampled for each observation here in the methods section or in the appendix.
Line 125: Did you consider re-scaling annual totals to correspond to the official self-reported Berlin emissions?
Line 137: How was the uncertainty estimate of 6.8 MtCO2a-1 derived? Is it just 75% of the prior that you describe in the Section 2.5.1?
Line 157: Why did you decide to also include separate regions 3 and 4 for Germany upwind and downwind? Does it have an advantage over having only three regions 0 – 2, i.e. by increasing the size of Berlin upwind and downwind to cover 3 and 4 regions, correspondingly?
Line 170: Does ‘CO2_BCKfield’ include the forcing from Berlin emissions?
Lines 187-194: Please provide the formula you used for computing the gain and average kernel matrices, and comment on how it follows from this formula that values outside [-1, 1] are considered non-physical.
Lines 236-237: ‘The model reproduces the general structure well up to about 2km but underestimates and smooths the mixing-layer height, likely due to limited vertical resolution (cf. Table 1).’
The wind performance seems to be the same between both altitude ranges, I suggest mentioning this in the text. Also, based on the wind profiles (Figure 2), it seems like the vertical structure is quite different between the schemes. Could you please provide possible reasons for this difference.
Lines 272-273: ‘This is achieved without increasing its correlation to the observations (prior and posterior correlation both 0.46).’
You mention this statement several times in the text, could you please provide a few comments in the text on why it is important to keep the correlation between measurements and background unchanged. The background concentration can also include plumes of elevated concentration in it because CO2 is a near-conservative tracer on synoptic scales, so the enhancements can be advected from outside of the model domain as well. How do you make sure those data are not used for the inversions?
Lines 292-293: Why are the gain matrix values for cell 1 significantly smaller than for cell 0? Whereas for the diagonal elements of the average kernel matrix they are much closer to each other. Also, what is the meaning of the sign of the gain matrix values? Why does the sign alternate throughout the day?
Figure 4. It would be good to include a similar figure, but that also includes posterior estimates when the background was not optimized, to directly show what the impacts are on the concentration timeseries.
Lines 323-324: Please add comments on how you think the inversion is able to effectively separate the background and emission signals from observations. Is it because some parts of the observation were unaffected by local sources and were correlating only with background adjustments?
Lines 339-341 and 362: ‘This substantial uncertainty reduction highlights the ability of the Bayesian Inversion framework with explicit background estimation to more efficiently extract the available information from the aircraft observations.’
‘We reduce the flux uncertainty by around a factor of five compared to the one in Klausner et al.
(2020)…’
How realistic do you think the posterior uncertainty estimate is? Can it be decreased by the inversion too much? Are there any other sources of uncertainty that could not be included in your study? Please add some comments to the Discussions section.
Technical corrections
Table 1. Please provide variable name descriptions in the caption.
Figure 2. The difference in wind speed profiles between the two model schemes is quite significant. Could you please change zorder on the profile plots and bring the smooth model lines to the front, otherwise it is difficult to see them behind fluctuating observations.
Figure 3 e). What do the 'Posterior' units represent? Please add an explanation in the caption.
Please make sure that the posterior uncertainty is consistent throughout the manuscript. In conclusions (line 360), it is specified as ± 4.3, while in the abstract (line 8) you write ± 4.6.
References
Li, J., Li, P., Han, P., Cheng, Z., Li, J., Zhang, T., Chen, D., Zheng, Y., Zeng, N., and Zhang, G. (2026). Advances in the design of urban CO₂ emission monitoring networks: a review. Carbon Research, 5, 3. https://doi.org/10.1007/s44246-025-00239-z
Karion, A., Lopez-Coto, I., Gourdji, S. M., Mueller, K., Ghosh, S., Callahan, W., Stock, M., DiGangi, E., Prinzivalli, S., and Whetstone, J. (2021). Background conditions for an urban greenhouse gas network in the Washington, DC, and Baltimore metropolitan region. Atmospheric Chemistry and Physics, 21, 6257–6273. https://doi.org/10.5194/acp-21-6257-2021
Klausner, T., Mertens, M., Huntrieser, H., Galkowski, M., Kuhlmann, G., Baumann, R., Fiehn, A., Jöckel, P., Pühl, M., and Roiger, A. (2020). Urban greenhouse gas emissions from the Berlin area: A case study using airborne CO₂ and CH₄ in situ observations in summer 2018. Elementa: Science of the Anthropocene, 8, 15. https://doi.org/10.1525/elementa.411