the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
Spatiotemporal dynamics of atmospheric CO2 across China revealed by long-term, high-resolution satellite-derived data
Abstract. Understanding the spatiotemporal dynamics of atmospheric carbon dioxide (CO2) is fundamental for advancing climate change research and designing effective mitigation strategies. Yet current analyses are constrained by two key limitations: sparse observations that hinder intra-urban assessment and relatively short monitoring periods that limit long-term consistency. To overcome these challenges, we developed a long-term atmospheric CO2 hindcast modeling framework that generates daily 1-km column-averaged dry-air mole fraction of CO2 (XCO2) across China for 2000–2020. The framework adapts the proven PM2.5 hindcast approach to CO2 estimation by training an Extremely Randomized Trees model on the residuals between OCO-2 observations and CarbonTracker simulations. The model integrates a comprehensive set of physically interpretable predictors—including MAIAC aerosol optical depth, NO2, peroxyacetyl nitrate, meteorological variables, and land-use indicators—linking CO2 variability to co-emitted tracers and boundary-layer processes. Rigorous evaluation demonstrated high reliability (cross-validation R2 = 0.94–0.97, RMSE = 0.82–1.29 ppm; independent validation R2 = 0.82–0.97). The resulting long-term, high-resolution dataset reveals distinct carbon hotspots and their evolution: the North China Plain remained persistently elevated with rapid increases during 2000–2010, while southern China exhibited accelerated growth after 2010. Enhancement analyses identified consistent intra-regional hotspots in southeastern Beijing-Tianjin-Hebei and northern Zhejiang, with emissions declining after 2012 and rebounding after 2018. During the Wuhan COVID-19 lockdown, urban cores showed sharper reductions than suburban areas. The proposed XCO2 hindcast modeling framework and the resulting dataset provide a valuable foundation for advancing carbon-neutrality assessments and guiding climate policy across multiple spatial scales.
- Preprint
(1855 KB) - Metadata XML
-
Supplement
(1771 KB) - BibTeX
- EndNote
Status: final response (author comments only)
-
RC1: 'Comment on egusphere-2025-5647', Anonymous Referee #2, 11 Mar 2026
-
AC1: 'Reply on RC1', Qingqing He, 11 Aug 2026
We thank the reviewer for the detailed comments and for recognizing that the topic is timely and that the dataset may be valuable. We have carefully considered all points raised and revised the manuscript thoroughly based on the comments/suggestions. Below we provide point-by-point responses.
Response to Referee #1
General comment
This manuscript presents a long-term, high-resolution satellite-derived XCO2 dataset over China using a machine learning approach. While the topic is timely and the dataset potentially valuable, the manuscript in its current form suffers from several fundamental methodological concerns that undermine confidence in the validity and interpretability of the results. Key issues include: (1) an inadequately described and physically questionable methodology (2) an overstated and poorly validated claim that the model improves upon existing products; (3) a misuse of the XCO2 enhancement framework, which cannot be straightforwardly interpreted as a proxy for CO2 emissions without accounting for wind speed and other confounding factors; and (4) several factual errors, inconsistent terminology, and unclear or internally inconsistent figure descriptions. Taken together, these issues represent substantial weaknesses that require major revision before the manuscript can be considered suitable for publication. Given the number and severity of the concerns raised below, the manuscript is not recommended for publication in its current form.
Response: We thank the reviewer for this thorough and demanding review. We have revised the manuscript substantially in response. In brief: (1) the methodological description has been expanded, including the physical rationale for the predictor selection and a clarification of what our cross-validation does and does not demonstrate; (2) claims of improvement over existing products have been removed and replaced with a controlled evaluation of all datasets against the same independent observations under identical criteria; (3) the enhancement framework has been reformulated as an XCO2 anomaly, with emission-attributing language removed and its interpretive bounds stated explicitly; and (4) terminology, tables, and figure descriptions have been made consistent throughout. Each specific comment is addressed individually below.
Specific Comments
(1) Lines 37–40:
The argument that atmospheric CO2 data can be used to assess mitigation effectiveness and inform sustainable development is stated too broadly and lacks specificity. Connecting column-averaged XCO2 retrievals to surface emissions is scientifically challenging due to CO2's long atmospheric lifetime, its associated large and variable background signal, a low signal-to-noise ratio relative to emission-driven enhancements, and the difficulty of separating anthropogenic signals from biospheric fluxes. The authors should provide a more rigorous and nuanced discussion of how their dataset could be used for these purposes—acknowledging these limitations.
Response: Thank you for this important comment. Our intention was not to suggest that column-averaged XCO2 observations alone can directly quantify surface emissions or mitigation effectiveness. We have therefore revised the relevant sentence to avoid implying a straightforward relationship between XCO2 and near-surface carbon emissions. It now reads like:
“Developing an XCO2 dataset that combines fine spatial detail, long temporal span, and daily temporal resolution is therefore critical for characterizing atmospheric carbon hotspots and anomalies and for tracking their long-term dynamics, thereby providing an observational foundation for climate research and for studies seeking to link atmospheric CO2 patterns to underlying emission and uptake processes.”In addition, following the reviewer’s suggestion, we have added a discussion of these limitations in the final paragraph of Section 4. The revised text now explicitly clarifies both the potential applications of our estimated XCO2dataset and the constraints that should be considered when interpreting it. In particular, we emphasize that such applications are inherently limited by the nature of XCO2 itself, which is a column-averaged dry-air mole fraction.
(2) Line 50:
The manuscript states that OCO-2 has a "daily revisit capability." This is incorrect. According to the OCO-2 mission documentation, the satellite operates on a 16-day repeat cycle. This should be corrected, and the relevant citation should be verified accordingly.
Response: We thank the reviewer for this comment. We have revised the sentence to read:
“Among the available satellite sensors, OCO-2 retrievals are most frequently employed because of their relatively fine footprint (~1–2 km) and daily data acquisition, with a 16-day orbital repeat cycle.”
(3) Lines 99–101:
The citation of Crisp et al. (2017) as a reference for the CarbonTracker (CT) XCO2 product is inappropriate. Crisp et al. (2017) describes the algorithm theoretical basis for OCO-2 Level 2 retrievals, not the CarbonTracker data assimilation system. The authors should replace this with an appropriate reference for CarbonTracker CT2022 (e.g., Peters et al. or the relevant NOAA/GML documentation).
Response: Thank you for your comment. We have replaced this inappropriate citation with Peters et al. (2007; https://doi.org/10.1073/pnas.070898610).
(4) Lines 101–104:
The description of how the coarse-resolution CarbonTracker data (3° × 2°) are resampled to a 0.01° grid is insufficient. Also, the authors should clarify how the final XCO2 estimate is reconstructed—specifically, whether the predicted residual is added back to the CT XCO2 value.
Response: We appreciate the reviewer’s valuable comment on our methodological processing workflow. We adopted the bilinear resampling method to regrid CarbonTracker (CT) XCO2 to our target standard grid. Our machine learning model was trained to predict the residual difference between OCO‑2 XCO2 retrievals and CT XCO2outputs. The final high-resolution XCO2 estimates were reconstructed by superimposing the predicted residual onto the baseline CT XCO2 values. We have incorporated the above methodological details into Sections 2.1.1 and 2.2.2 in the revised manuscript, respectively.
(5) Lines 115–117:
The use of daily-mean meteorological and air pollution data to predict XCO2 at the satellite overpass time (approximately 13:30 local time) requires physical justification. Atmospheric properties such as planetary boundary layer height, temperature, humidity, and trace gas concentrations can vary substantially throughout the day. Using daily-mean values rather than values contemporaneous with the satellite overpass may introduce systematic biases. The authors should either provide a physical justification for this approach or conduct a sensitivity analysis to evaluate how the temporal sampling of input predictors affects model performance for the target overpass time.
Response:
We appreciate this valuable comment. The application of daily-mean meteorological and air pollution data is supported by two complementary perspectives: the physical character of XCO2 as a column-integrated quantity, and the structure of our residual-learning framework. Please see the details below:
(1) Physical rationale: XCO2 near local noon is close to its daily-mean value
Unlike near-surface CO2, which responds directly to the diurnal cycle of boundary-layer mixing and local fluxes, XCO2 is the dry-air mole fraction integrated over the full atmospheric column. Vertical redistribution of CO2 within the boundary layer alters near-surface concentrations substantially but leaves the column amount largely unchanged. XCO2 therefore exhibits far weaker intraday variability than surface CO2, and the OCO-2 overpass is deliberately positioned at a time of day when this variability is minimal. This is supported empirically:
- NASA/JPL states this explicitly in describing the mission’s measurement approach: "since XCO2 measurements tend to be near their daily average value at this time of day, the Observatory data will be highly representative of the region where they were acquired" (https://ocov2.jpl.nasa.gov/science/measurement-approach/).
- In addition, He et al. (2022b, https://doi.org/10.1029/2022GL098435) show that hourly XCO2 at 13:00 is well correlated with the daily-mean XCO2 concluding that "the mid-day XCO2 could accurately reflect the general patterns of XCO2 (their Figure S1).
The key implication is not simply that midday XCO2 resembles the daily mean, but that the target quantity itself effectively varies on daily rather than sub-daily timescales. Predictors describing the daily-scale environmental state are therefore temporally well matched to the quantity being modeled. An instantaneous meteorological value at 13:30 characterizes a single moment, whereas a column-integrated residual reflects transport and surface-flux influences accumulated over the preceding hours to days; daily aggregates arguably represent these integrated conditions more faithfully than a snapshot.
(2) Modeling rationale: predictors describe daily environmental context, not the instantaneous atmospheric state
Our estimation framework is not using daily meteorology and air pollution predictors to directly reconstruct 13:00 OCO-2 XCO2. Instead, it is a residual-learning approach in which the response variable is the difference between OCO-2 XCO2 and the co-located, overpass-matched CarbonTracker XCO2. CarbonTracker thus supplies the first-order atmospheric XCO2 field, including its sub-daily transport structure, and the machine-learning model is tasked only with explaining the systematic, fine-scale departures of OCO-2 from that field. Within this framework, the predictors act as descriptors of each day’s environmental regime rather than as instantaneous physical forcing, and they are not required to share the temporal resolution of the response.
In the revised manuscript, we have also revised Section 2.1.2 to clarify the role of these predictors and the rationale for their temporal aggregation.
(6) Lines 163–167:
The leave-one-year-out cross-validation strategy described here appears to withhold one year from within the 2015–2020 period for evaluation. However, the authors claim this approach allows assessment of model performance for the pre-2015 period. This logic is not convincing: withholding a year from the middle of the training period does not simulate the extrapolation challenge posed by hindcasting to years before the OCO-2 era, which may involve structural differences in the predictor-XCO2 relationships. The authors should clarify the validation approach and, if the goal is to assess pre-2015 performance, adopt a more appropriate out-of-sample evaluation strategy (e.g., training exclusively on post-2015 data and evaluating against ground-based observations in pre-2015 years).
Response: We thank the reviewer for this careful and constructive comment. We agree that leave-one-year-out (LOYO) cross-validation cannot by itself establish model performance for the pre-2015 period. We would respectfully note that the strategy the reviewer proposes (i.e., training exclusively on post-2015 data and evaluating against ground-based observations in pre-2015 years) corresponds to the design already adopted in our work. Because OCO-2 observations begin in September 2014, the model can only ever be trained on 2015–2020 data; no pre-2015 satellite record exists that could be included in training or used for direct hindcast validation. We have explained these points in Sections 2.2.2 and 2.3.
In response to the reviewer’s comment, we have now separately evaluated the pre-2015 hindcast period (2000–2014), rather than reporting only the comparison for the entire 2000–2020 period, with the validation statistics summarized in the table below:
- Please note that the ground-based observations are fully independent of the model: no surface or column measurements from any site are used as predictors or as training targets at any stage.
- The estimated XCO2 closely tracked the long-term trend and interannual variations observed at WLG, with an R2 of 0.97 for 2000–2020 (Fig. 1g) and a similarly high 0.94 for the pre-2015 period (Table S7), indicating that the model captures the temporal dynamics of XCO2 consistently both within and outside the OCO-2 era. We acknowledge that WLG is an imperfect reference, as it is a surface in-situ measurement at a high-altitude baseline station rather than a column observation, and we have stated this limitation explicitly in Section 2.3.2.
- Accordingly, we interpret this comparison as a test of long-term trend and interannual consistency rather than of absolute column accuracy. Its principal value here is that it is the only continuous, independent, well-calibrated record spanning the full 2000–2020 study period.
- Within the OCO-2 data constraint, we believe that LOYO combined with independent long-term ground-based evaluation (below) therefore represents the strongest assessment available.
Table S7. Validation statistics of the reconstructed XCO2 against WLG in-situ measurements for the pre-OCO-2 hindcast period (2000–2014) and the full period (2000–2020).
Time Period
R2
RMSE (ppm)
MAE (ppm)
No. of matchups
2000-2014
0.94
2.77
2.35
4284
2000-2020
0.97
2.90
2.48
6356
We would, however, respectfully argue that LOYO provides meaningful, if indirect, evidence bearing on hindcast reliability:
- We would note the fundamental constraint in this study: OCO-2 observations begin in September 2014, so no satellite column measurements exist before 2015 against which hindcasts could be directly validated.
- Hindcasting to the pre-2015 period requires that the predictor–residual relationship learned from the OCO-2 era transfer to years whose meteorological conditions, flux anomalies, and interannual variability were not represented in training. LOYO tests precisely this property: by withholding an entire year, the model is evaluated on a full annual cycle of conditions it has never encountered, rather than on randomly held-out samples interpolated among their temporal neighbors. Averaged across all withheld years, it therefore quantifies the stability of the learned relationship under year-to-year variation, which is a necessary condition for hindcast validity and is informative about it.
- We have accordingly revised Sections 2.3 and 3.1.2 so that LOYO is described as evidence for the temporal transferability of the learned relationship rather than as a direct assessment of pre-2015 accuracy, and we now present it as one of several complementary evaluations rather than as the primary basis for our hindcast claims.
(7) Lines 183–185:
The use of a 10:30–16:30 time window to average ground-based observations for comparison with OCO-2 retrievals (overpass time ~13:30) is overly broad and requires justification. Atmospheric CO2 concentrations at surface sites exhibit pronounced diurnal variability driven by boundary layer dynamics and biospheric fluxes, meaning that measurements taken in the early morning or late afternoon may differ substantially from those at solar noon. Averaging over a six-hour window centered loosely on the overpass time may introduce significant biases in the validation. The authors should either narrow the averaging window (e.g., ±1–2 hours around the overpass time) or provide a sensitivity analysis demonstrating that this choice does not materially affect the validation statistics.
Response: We thank the reviewer for this comment. We would first clarify a point of scope: the 10:30–16:30 averaging window applies only to the TCCON comparisons, which are column-averaged dry-air mole fractions (XCO2) obtained by direct solar absorption spectrometry, rather than near-surface in-situ concentrations. Because such redistribution largely conserves the total column amount, XCO2 exhibits substantially weaker diurnal variation than near-surface CO2.
Following the reviewer’s suggestion, we evaluated the sensitivity of the validation statistics to the averaging window. We recomputed all comparison statistics using ±1 h, ±1.5 h, and ±3 h windows centered on the overpass time, in addition to the full daily average. As shown in table below, differences among the window configurations are negligible. Please note that the Mt. Waliguan record is provided as daily values, so no sub-daily averaging window is applied there and this comment does not bear on that comparison. We have added Table S4 and revised the description in Section 2.3.2 in the revised manuscript correspondingly.
Table S4. Sensitivity analysis of validation performance using different time averaging windows centered on OCO-2 overpass time (~13:30) at HF and XH sites within a 1 × 1 spatial matching window.
Station
Time windows
R2
RMSE (ppm)
MAE (ppm)
No. of matchups
HF
±1 h
0.926
1.22
0.97
214
±1.5 h
0.926
1.23
0.97
243
±3 h
0.923
1.22
0.97
294
daily mean
0.920
1.22
0.96
341
XH
±1 h
0.821
1.69
1.29
498
±1.5 h
0.817
1.72
1.29
520
±3 h
0.817
1.69
1.29
578
daily mean
0.816
1.69
1.30
602
(8) Lines 196–199:
The XCO2 enhancement method, while widely used in exploratory analysis, cannot be directly interpreted as an indicator of surface CO2 emissions without accounting for atmospheric transport, particularly wind speed and direction. An XCO2 enhancement above a background value reflects a combination of upstream emissions, atmospheric dilution (governed by wind speed), boundary layer height, and biospheric signals. High XCO2 enhancements under calm wind conditions may not correspond to higher emissions than lower enhancements under strong winds. The authors should either (1) incorporate a wind-speed correction or apply a more physically rigorous emission estimation framework, or (2) explicitly characterize the temporal variability of wind conditions over the study regions and discuss the extent to which this limits the interpretation of enhancements as emission proxies.
Response: Thank you for your thoughtful comment. We agree that the XCO2 enhancement cannot be interpreted as a quantitative indicator of surface CO2 emissions. To remove the implied link to emissions at its source, we now refer to the quantity throughout the manuscript as the XCO2 anomaly rather than the XCO2 enhancement. This follows the terminology of Hakkarainen et al. (2016), whose method we adopt, and describes the quantity accurately in the revised Section 2.4.2: a departure from a same-day regional reference value. The resulting anomaly therefore represents the spatial departure of each grid cell from its same-day regional context, rather than an absolute XCO2 level or a temporal enhancement. This construction is what makes the quantity informative about persistent spatial structure while remaining agnostic about its physical attribution on any individual day.
In addition, we have characterized the temporal variability of wind conditions over the study regions. Figure S9 shows the seasonal and interannual distribution of 10-m wind speed, east component, and north component over the three urban agglomerations (BTH, YRD, PRD), based on ERA5 reanalysis. Mean wind speeds over these regions are 2.02–2.25 m s⁻¹, with moderate seasonal variation and no significant long-term trend over 2000–2020 (-0.013 to -0.053 m s⁻¹ decade⁻¹, all p > 0.05). Because wind conditions show no systematic long-term trend over the study period, the multi-decadal changes in anomaly patterns that we report are unlikely to be artefacts of changing transport, although we note that this argument concerns aggregate behavior and does not license emission attribution for individual days or grid cells. We have added this analysis and discussion to Section 4.
(9) Figure 1g: A visible data gap appears in the estimated XCO2 time series for approximately 2002–2003. Is it attributable to missing input data (e.g., predictor variables), or a deliberate data exclusion decision? This should be addressed explicitly in the text.
Response: Thank you for pointing this out. We have clarified the cause of this apparent gap in Section 3.1.2 in the revised manuscript. The visible gap in Fig. 1g during 2002–2003 is not due to data exclusion or missing predictor variables, but rather reflects a period when no valid in-situ XCO2 records were available at the WLG station during that time.
Section 3.1.3:
This section discusses feature importance but focuses almost exclusively on MAIAC AOD, neglecting several other highly ranked predictors. Notably, total column water vapour emerges as the most important predictor in both the YRD and PRD regions, yet this is not discussed. The authors should provide a physical explanation for this result—is this relationship physically meaningful or potentially spurious? Similarly, Day of Year appears to be among the most important features across regions, raising the question of whether the model's representation of XCO2 seasonality is driven primarily by this temporal index rather than by physically meaningful predictors. The implications of this for spatial generalization and for hindcast periods should be discussed.
Response: We thank the reviewer for this constructive comment, and we have substantially expanded the results in Section 3.1.3 and the discussion in the third paragraph of Section 4 to cover the other highly ranked predictors.
First, we consider the high importance of total column water vapor to be physically meaningful. It ranked first in both YRD and PRD, where local SHAP values indicate a predominantly negative association: humid conditions correspond to OCO-2 XCO2 below the CarbonTracker field. This is consistent with the monsoon climate of these two coastal regions, in which high column water vapor can indicate maritime air masses that are relatively CO2-depleted compared with continental air over eastern China, and with the known sensitivity of the 2.06 µm CO2 retrieval band to humidity along the optical path.
For day of year, it ranked second in YRD and PRD and third in BTH; however, this importance does not indicate that XCO2 seasonality is reproduced by a temporal index.
- Because our response variable is the OCO-2–CarbonTracker residual, in which the seasonal cycle of XCO2 is already represented by CarbonTracker. It therefore reflects a seasonally recurring component of the discrepancy between the retrievals and the modelled field, consistent with the seasonality of biospheric flux errors and of retrieval conditions such as solar zenith angle and aerosol loading.
- Note that this predictor is spatially invariant on any given day and therefore cannot itself generate spatial structure; the spatial gradients in our estimates derive from spatially varying predictors such as MAIAC AOD, the meteorological fields, and the spatial covariates.
- It is also a bounded, cyclic variable whose full range is sampled within every training year, so hindcast prediction involves no extrapolation in this dimension, in contrast to an absolute time index (e.g., year), which would require the model to project beyond its training range.
Day of year and total column water vapor thus carry complementary information, the former representing the mean seasonal cycle and the latter the spatially resolved, day-specific expression of monsoon influence.
(10) Lines 270–272:
There is a typographical error in the figure caption: "Figure 2. of each …" should be corrected.
Response: Thank you for pointing out this error. We have corrected the mistake in the Figure 2 caption accordingly.
(11) Lines 276–280:
The comparison of R² and RMSE values across different models to argue that the present model outperforms previous studies is methodologically inappropriate. Model skill metrics are highly sensitive to the spatial domain, temporal coverage, and resolution of the evaluation dataset, as well as the choice of validation sites and periods. Without a controlled, identical evaluation framework applied to all compared models, such inter-model comparisons cannot support strong claims of superiority. The authors should reframe this comparison more cautiously, noting that direct performance comparisons are not possible across studies with differing spatiotemporal coverage and evaluation protocols.
Response: We thank the reviewer for this comment. We have added a comparison under controlled, identical evaluation framework in the revise manuscript (Please see the table below). That is, we evaluated the original CarbonTracker and CAMS products against the same independent ground-based observations, over the same period, at the same sites, using an identical ±3 h temporal matching window and identical evaluation metrics. Critically, we also evaluated our own estimates at each of the comparison products' native spatial resolutions, by aggregating our 0.01° field to the 3° × 2° CarbonTracker grid and to the 0.75° × 0.75° CAMS grid. Under this framework, our estimates agree more closely with the independent ground-based observations than either reanalysis product at every matched spatial scale and at all three stations:
HF
R² 0.93, RMSE 1.13
R² 0.93, RMSE 1.22
R² 0.92, RMSE 1.24
R² 0.89, RMSE 1.48
R² 0.92, RMSE 1.22
294
XH
R² 0.81, RMSE 1.73
R² 0.74, RMSE 2.18
R² 0.82, RMSE 1.68
R² 0.58, RMSE 2.61
R² 0.82, RMSE 1.69
578
WLG
R² 0.97, RMSE 2.91
R² 0.87, RMSE 5.55
R² 0.97, RMSE 2.92
R² 0.95, RMSE 3.08
R² 0.97, RMSE 2.90
6356*
* CAMS N = 5667 at WLG, as the CAMS record begins in 2003 while the CarbonTracker and ground-based records span the full study period.
We have revised Section 3.1.1 based on these comparison results presented above and added the above table in the revised manuscript (Table 2). In addition, we have removed all language claiming superiority to publications using different spatial domains, time periods, validation sites, and matchup protocols. These values are now presented in Table S8 solely to indicate the range of performance reported in the literature, accompanied by an explicit statement that they are not directly comparable to our results and that no conclusions about relative model skill are drawn from them. We have additionally renamed Section 3.2.1 from "Improved estimation accuracy" to "Estimation accuracy comparison" so that the heading no longer asserts the conclusion.
(12) Lines 294–296:
The claim that previous machine-learning studies producing daily 1-km XCO2 estimates have been limited to post-2015 data is made without citation. Please add appropriate references to support this statement.
Response: We fully agree with this comment and have supplemented supporting references while refining the original statement to avoid overgeneralization. To achieve a rigorous and non-overgeneralized description, we have revised the original text to replace the fixed “1-km” wording with the general term “high-resolution”. We have cited representative machine-learning XCO2 mapping works (Li et al., 2023, https://doi.org/10.1016/j.scitotenv.2023.164921; Wang et al., 2023; https://doi.org/10.5194/essd-15-3597-2023; Wu et al., 2024; https://doi.org/10.1016/j.scitotenv.2024.176171) to back up the observation that most existing high-resolution daily XCO2 datasets are limited to post-2015 coverage.
(13) Lines 303–305 and 309–310:
The claim that the high-resolution product captures intra-urban XCO2 variations "with greater accuracy" than CT is not substantiated. To support this claim, the authors must provide a direct comparison of CT XCO2 and the model-estimated XCO2 against independent validation data (i.e., data not used in training), with evaluation metrics reported for both. Without this, the claim that the machine learning model improves upon CT at fine spatial scales remains unverified.
Response: We thank the reviewer for this comment. In the revised manuscript, we have both carried out the requested comparison and revised the claim to match what the evidence supports:
- As described in Comment 11, we have evaluated CarbonTracker XCO2, CAMS XCO2, and our estimated XCO2 against the same independent ground-based observations, with metrics reported for each. We emphasize that these observations are fully independent of the model: no surface or column measurements from any site are used as predictors or as training targets at any stage. To ensure that no comparison is confounded by scale, our estimates were aggregated to each comparison product’s native grid, and each product was evaluated on identical matchup days.
- At every matched spatial scale, our estimates agree more closely with the ground-based observations than the corresponding reanalysis product — most substantially at WLG against CarbonTracker (RMSE 2.91 vs 5.55 ppm) and at XH against CAMS (RMSE 1.68 vs 2.61 ppm).
- This addresses the reviewer’s requirement that any claim of improvement upon CarbonTracker be supported by a direct, like-for-like comparison against independent data.
We agree that the specific assertion that our product captures intra-urban variation "with greater accuracy" was not supported. The revised passage no longer claims greater intra-urban accuracy, and instead states that our product resolves spatial variation that the coarse-resolution products cannot represent, while being more accurate than those products at every scale at which direct comparison is possible.
- We note that the comparison is not, in fact, one of relative accuracy at intra-urban scales: a single CarbonTracker grid cell spans approximately 300 × 220 km and a single CAMS cell approximately 80 km, so an entire urban agglomeration falls within one or a few cells and intra-urban variation cannot be represented by these products at all. The distinction is therefore one of representational capability rather than of accuracy.
- We have also verified that generating this fine-scale detail does not degrade accuracy (Fig. 1e-f vs Table 2). We further note, in the revised text, that no dense within-city observational network exists in China against which intra-urban gradients could be validated directly, and that the site-level comparisons in Table 2 therefore represent the strongest available evidence on this point rather than a direct validation of intra-urban structure.
(14) Figure 3a:
A sharp spatial gradient in XCO2 is visible around the Taklamakan Desert region, which appears physically implausible for a remote arid region with minimal anthropogenic activity. The authors should investigate whether this feature is present in the OCO-2 retrieval data or the CarbonTracker product, and if not, identify which input variable(s) are driving this artifact.
Response: We appreciate this observation. A gradient in a remote arid region is not necessarily implausible, because regional XCO2 patterns can be influenced by large-scale transport, topography, surface pressure, and land-atmosphere exchange, not only by local anthropogenic emissions. We would note that the gradient in our product is not unique and is consistent with these independent studies that used different machine learning algorithms and data sources (Please see the figure attached):
- We observed similar features that have been reported in previous XCO2 mapping studies over China. Zhang and Liu (2023; https://doi.org/10.1016/j.scitotenv.2022.159588) developed a contiguous XCO2 dataset over China using a CNN-based machine learning approach with multi-satellite observations (SCIAMACHY, GOSAT, OCO-2) and also found elevated XCO2 values in the Taklamakan Desert region (their Figure 9a). They noted that this feature may be associated with the lack of vegetation and topographic factors, while also acknowledging some overestimation (approximately 1 ppm) compared to the CT model (their Figure 7d-e).
- He et al. (2022; https://doi.org/10.1029/2022GL098435) similarly reported that Northwest China exhibits internal spatial gradients in XCO2, with the Tarim Basin showing relatively higher values than surrounding areas despite being the lowest overall region.
- In addition to the machine-learning estimates, the CT and CAMS data for this region also show relatively higher levels than other western regions.
Nevertheless, we agree that this feature deserves clarification to avoid misinterpretation. In the revised manuscript, we have added a brief explanation noting that this pattern may be influenced by topographic effects (low elevation of the Tarim Basin vs. high elevation of the Tibetan Plateau), limited vegetation uptake, and large-scale transport, while also acknowledging the uncertainty as discussed in Zhang and Liu (2023).
(15) Lines 351–352:
The terms "CO2," "XCO2," and "mixing ratio/concentration/level" are used inconsistently throughout the manuscript. XCO2 is a column-averaged dry-air mole fraction expressed in parts per million (ppm), not a concentration in the physical chemistry sense (e.g., mol/m³). The authors should adopt consistent and scientifically precise terminology throughout the manuscript and avoid referring to ppm values as "concentrations."
Response: We fully agree with this suggestion. In the revised manuscript, we have used “XCO2” or “column-averaged dry-air mole fraction of CO2” consistently when referring to the satellite-derived/modelled variable. We have avoided referring to XCO2 in ppm as “concentration” where that could be confused with a physical volumetric concentration such as mol/m3. For general references to atmospheric carbon dioxide, we explicitly distinguish bulk atmospheric CO2 from column-averaged XCO2 to eliminate inconsistent wording.
(16) Figures 5c and 5d:
The interpretation of XCO2 enhancements as indicators of urban CO2 emissions in these figures is problematic. If a single background XCO2 value is used for each region, the spatial patterns shown primarily reflect the climatological XCO2 gradient across the region, which integrates regional wind transport, biospheric fluxes, and the regional XCO2 gradient—not local emission differences between cities. For example, the larger enhancements observed in the southern portion of the BTH region do not necessarily imply greater emissions than those in Beijing; they may simply reflect more favorable transport or boundary layer conditions. The authors should clarify the background correction methodology and substantially revise the interpretation of these figures.
Response: We thank the reviewer for raising this important point. We first clarify the background methodology: the background is not a single fixed value per region, but is recomputed for every day as the regional median (Section 2.4.2). This removes the secular increase in XCO2, the regional seasonal cycle, and synoptic-scale variability affecting the region as a whole, so that the anomaly represents each grid cell’s departure from its same-day regional context rather than a climatological XCO2 level.
We accept, however, the substance of the comment. Because the background is a single value per region on each day, systematic gradients in horizontal transport within a region are not removed, and a multiyear mean of daily anomalies therefore describes where column CO2 persistently accumulates rather than the magnitude of emissions at individual locations. The reviewer’s example is well taken: the higher anomalies in southeastern BTH reflect both the regional distribution of emissions and the proximity of those areas to the intensively industrialized parts of Shandong and Henan, from which column CO2 is advected.
We have revised the manuscript accordingly. Throughout the Results (Sections 3.3.2 and 3.3.3) the term “enhancement” is replaced by “anomaly”, the qualifier “anthropogenic” is removed where anomalies are described, and the results are now reported descriptively without emission attribution. In the Discussion we have moderated the causal language used in interpreting the hotspot patterns, and have added a dedicated passage stating the limitations governing their interpretation — that the anomaly integrates transport, terrain, and biospheric influences in addition to emissions, and that quantitative emission attribution would require inversion or co-emitted tracer approaches. We have additionally characterized wind conditions over the study regions (Fig. S9), as described in our reply to Comment 8.
(17) Lines 383–385:
The description of Figure 5d is inconsistent with its caption. The text refers to "a striking difference in the spatial patterns of XCO2 enhancements between the two years" and reports specific percentage reductions during the COVID-19 lockdown in Wuhan, but the figure caption states it shows the "20-year mean XCO2 enhancements for the YRD region." It is unclear which figure the text is actually referring to. The authors should ensure that all in-text figure references are accurate and that figure captions correctly describe the displayed content.
Response: Thank you for identifying this inconsistency. The discussion of spatial distribution of percentage changes in XCO2 enhancement between the two years in the Wuhan refers to Figure 5a, not Figure 5d. We have corrected the figure reference and ensure that the caption and in-text description are consistent.
-
AC1: 'Reply on RC1', Qingqing He, 11 Aug 2026
-
RC2: 'Comment on egusphere-2025-5647', Anonymous Referee #3, 01 Jul 2026
This manuscript develops a long-term (2000–2020), daily 1-km XCO2 dataset over China using a machine-learning framework trained with OCO-2 observations and CarbonTracker data. The study addresses an important topic because long-term, high-resolution atmospheric CO2 products remain limited, particularly prior to the launch of OCO-2. The modeling framework is well designed, and the validation results demonstrate good predictive performance using both cross-validation and independent observations. The enhancement analysis also produces interesting spatial and temporal patterns over major urban agglomerations.
Overall, I think this is a solid contribution that is suitable for ACP after moderate revision. My main concern is the interpretation of the enhancement results, while several other issues related to clarity and presentation should also be addressed.
Major comment
My primary concern is the interpretation of XCO2 enhancement.
Throughout the manuscript, the enhancement maps are frequently interpreted as "carbon emissions" or "changes in carbon emissions." Although the enhancement patterns clearly reveal anthropogenic influences, the enhancement itself remains an atmospheric concentration signal that is influenced not only by emissions but also by atmospheric transport, boundary-layer dynamics, biospheric exchange, and the definition of the background. Therefore, I suggest that the authors carefully review the manuscript and moderate statements that directly equate enhancement with emissions unless additional atmospheric transport or inverse modeling is incorporated. This issue affects not only the terminology but also some of the causal and policy interpretations in the Abstract, Sections 2.4.2 and 3.3, the Discussion, and the Conclusions.
In addition, the distinction between atmospheric XCO2, XCO2 enhancement, and carbon emissions should be maintained consistently throughout the manuscript. In several places these concepts appear interchangeable, which may confuse readers.
Minor comments
1. The caption of Figure 1 is unusually long. Some details, particularly the precise positions of the color bars, are not necessary for interpreting the figure. The caption could be shortened while retaining the definitions of panels (a)–(h), the validation datasets, and the principal information represented by each panel.
2. Caption of Fig. 2 caption beginning around line 270 is grammatically incomplete: “Figure 2. of each predictor on XCO2 levels quantified using the SHAP method…”
Please also insert a space in “Fig.2” in Section 3.1.3.
3. The caption of Fig. 3 lists panels (a), (b), and (c), but the panel identifier “(d)” is missing before the GOSAT interpolation. The sentence structure is also difficult to follow because the daily, monthly, and annual columns are described in reverse order. I suggest revising it to clearly identify all four rows or products and the three temporal aggregations.
4. The title of Section 3.3.2 contains duplicated wording: carbon emission.
In addition, the paragraph following Fig. 5 states, “As shown in Fig. 5d,” when discussing spatial changes during the Wuhan lockdown. According to the Figure 5 caption, panel (d) presents the multiyear mean enhancement in the YRD, whereas the Wuhan percentage changes are presented in panel (a). This reference should therefore be checked and likely changed to Figure 5a.
5. Specific grammar and wording corrections
The manuscript would benefit from careful language editing. Examples include:
1) Lines 119–122: The sentence beginning “We obtained annual land cover classification data…” joins several independent clauses with commas. It should be divided into separate sentences or restructured with parallel verbs.
2) Line 223 should be revised to “with more than 65% of grid cells containing only one observation.”
3) Lines 309–310: “These of accuracy and spatiotemporal patterns underscore…” is incomplete. It may have been intended to read: “These comparisons of accuracy and spatiotemporal patterns underscore…”
4) Line 466: Insert a space before the citation in “previous studies(Lu et al., 2025).”
5) Lines 503–505: The sentence here is missing a noun after “Wuhan-specific.” It should be revised to “consistent with previous Wuhan-specific studies.” There is also an extra closing parenthesis in “40–60%)).”
6) Line 526 should be revised to “applying it to hindcast XCO2 for earlier years”.
6. Undefined dataset abbreviations
Most regional and methodological abbreviations are appropriately defined. However, several dataset or reanalysis abbreviations appear without their full names, particularly MERRA2-GMI and EAC4 in Section 2.1.2. ERA and ERA5 are also introduced only as product names, and CAMS is used later without first providing its complete name. Please define these terms at their first occurrence and then use the abbreviations consistently.
Citation: https://doi.org/10.5194/egusphere-2025-5647-RC2 -
AC2: 'Reply on RC2', Qingqing He, 11 Aug 2026
Response to Referee #2
Major comment
My primary concern is the interpretation of XCO2 enhancement.
Throughout the manuscript, the enhancement maps are frequently interpreted as "carbon emissions" or "changes in carbon emissions." Although the enhancement patterns clearly reveal anthropogenic influences, the enhancement itself remains an atmospheric concentration signal that is influenced not only by emissions but also by atmospheric transport, boundary-layer dynamics, biospheric exchange, and the definition of the background. Therefore, I suggest that the authors carefully review the manuscript and moderate statements that directly equate enhancement with emissions unless additional atmospheric transport or inverse modeling is incorporated. This issue affects not only the terminology but also some of the causal and policy interpretations in the Abstract, Sections 2.4.2 and 3.3, the Discussion, and the Conclusions.
In addition, the distinction between atmospheric XCO2, XCO2 enhancement, and carbon emissions should be maintained consistently throughout the manuscript. In several places these concepts appear interchangeable, which may confuse readers.
Response: We agree that the XCO2 enhancement cannot be interpreted as a quantitative indicator of surface CO2 emissions. To remove the implied link to emissions at its source, we now refer to the quantity throughout the manuscript as the XCO2 anomaly rather than the XCO2 enhancement. This follows the terminology of Hakkarainen et al. (2016), whose method we adopt, and describes the quantity accurately in the revised Section 2.4.2: a departure from a same-day regional reference value. The resulting anomaly therefore represents the spatial departure of each grid cell from its same-day regional context, rather than an absolute XCO2 level or a temporal enhancement. This construction is what makes the quantity informative about persistent spatial structure while remaining agnostic about its physical attribution on any individual day.
In the revised manuscript, we have implemented three major revisions:
- Uniformly distinguish three independent concepts (atmospheric XCO2, relative XCO2 enhancement, surface carbon emissions) in the Abstract, Section 2.4.2, Section 3.3, Discussion and Conclusion, and avoid interchangeable use of these terms;
- Remove all direct expressions that treat enhancement as equivalent to carbon emissions or emission variations. We replace phrases such as “carbon emission hotspots” with neutral descriptions including “persistent positive XCO2 anomaly hotspots” and “emission-related XCO2 spatial anomalies”;
- Add explanatory notes in Section 2.4.2 and Section 4 to explicitly clarify the interpretive limitations of XCO2 anomaly: the signal only reflects integrated atmospheric perturbations driven by both anthropogenic and natural processes, rather than direct quantitative proxies of surface CO2
Minor comments
- The caption of Figure 1 is unusually long. Some details, particularly the precise positions of the color bars, are not necessary for interpreting the figure. The caption could be shortened while retaining the definitions of panels (a)–(h), the validation datasets, and the principal information represented by each panel.
Response: We appreciate this valuable suggestion. In the revised manuscript, we have shortened the caption of Figure 1 by removing details that are not essential for interpreting the figure, such as the precise positions of the color bars. At the same time, we have retained the definitions of panels (a)–(h), the validation datasets, and the key information conveyed by each panel. Because Figure 1 is a multi-panel figure that integrates the main validation results of our study, its caption remains relatively detailed to ensure that the figure can be understood clearly and independently.
- Caption of Fig. 2 caption beginning around line 270 is grammatically incomplete: “Figure 2. of each predictor on XCO2 levels quantified using the SHAP method…” Please also insert a space in “Fig.2” in Section 3.1.3.
Response: Thank you for pointing out this grammatical error and formatting issue. In the revised manuscript, we have corrected the incomplete sentence structure of Figure 2’s caption to form a complete and logical statement. Meanwhile, all instances of “Fig.2” throughout Section 3.1.3 will be revised to “Fig. 2” with a standard space added between the abbreviation and number.
- The caption of Fig. 3 lists panels (a), (b), and (c), but the panel identifier “(d)” is missing before the GOSAT interpolation. The sentence structure is also difficult to follow because the daily, monthly, and annual columns are described in reverse order. I suggest revising it to clearly identify all four rows or products and the three temporal aggregations.
Response: We are grateful for this careful reminder. We have supplemented the missing panel label (d) corresponding to GOSAT interpolation in the caption of Figure 3. In the revised caption, we have reordered the temporal descriptions to follow the left-to-right layout of the figure.
- The title of Section 3.3.2 contains duplicated wording: carbon emission.
In addition, the paragraph following Fig. 5 states, “As shown in Fig. 5d,” when discussing spatial changes during the Wuhan lockdown. According to the Figure 5 caption, panel (d) presents the multiyear mean enhancement in the YRD, whereas the Wuhan percentage changes are presented in panel (a). This reference should therefore be checked and likely changed to Figure 5a.
Response: Thank you for your comment. First, we have deleted the repeated phrase “carbon emission” in the title of Section 3.3.2. Second, we corrected the misreference “Fig. 5d” to “Fig. 5a” in the paragraph discussing Wuhan lockdown spatial variations, consistent with the actual content of each subgraph defined in Figure 5’s caption.
- Specific grammar and wording corrections
The manuscript would benefit from careful language editing. Examples include:
1) Lines 119–122: The sentence beginning “We obtained annual land cover classification data…” joins several independent clauses with commas. It should be divided into separate sentences or restructured with parallel verbs.
2) Line 223 should be revised to “with more than 65% of grid cells containing only one observation.”
3) Lines 309–310: “These of accuracy and spatiotemporal patterns underscore…” is incomplete. It may have been intended to read: “These comparisons of accuracy and spatiotemporal patterns underscore…”
4) Line 466: Insert a space before the citation in “previous studies(Lu et al., 2025).”
5) Lines 503–505: The sentence here is missing a noun after “Wuhan-specific.” It should be revised to “consistent with previous Wuhan-specific studies.” There is also an extra closing parenthesis in “40–60%)).”
6) Line 526 should be revised to “applying it to hindcast XCO2 for earlier years”.
Response: We sincerely thank you for listing these detailed linguistic defects. We have conducted a full round of language editing across the whole manuscript, and all six grammatical and wording issues you mentioned have been fully revised one by one as recommended.
- Undefined dataset abbreviations
Most regional and methodological abbreviations are appropriately defined. However, several dataset or reanalysis abbreviations appear without their full names, particularly MERRA2-GMI and EAC4 in Section 2.1.2. ERA and ERA5 are also introduced only as product names, and CAMS is used later without first providing its complete name. Please define these terms at their first occurrence and then use the abbreviations consistently.
Response: We appreciate your careful check on abbreviations. We have supplemented the full official names for all undefined dataset abbreviations at their first appearance, including MERRA2-GMI, EAC4, ERA, ERA5, and CAMS. All subsequent appearances of these terms then uniformly adopt the corresponding abbreviations to maintain consistency throughout the manuscript.
Finally, we would like to thank the reviewer once again for their time, effort, and constructive comments, which have greatly helped us improve the manuscript.
Citation: https://doi.org/10.5194/egusphere-2025-5647-AC2
-
AC2: 'Reply on RC2', Qingqing He, 11 Aug 2026
Viewed
| HTML | XML | Total | Supplement | BibTeX | EndNote | |
|---|---|---|---|---|---|---|
| 807 | 564 | 72 | 1,443 | 211 | 65 | 93 |
- HTML: 807
- PDF: 564
- XML: 72
- Total: 1,443
- Supplement: 211
- BibTeX: 65
- EndNote: 93
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
"Spatiotemporal Dynamics of Atmospheric CO2 across China Revealed by Long-Term, High-Resolution Satellite Data"
General Comments
This manuscript presents a long-term, high-resolution satellite-derived XCO2 dataset over China using a machine learning approach. While the topic is timely and the dataset potentially valuable, the manuscript in its current form suffers from several fundamental methodological concerns that undermine confidence in the validity and interpretability of the results. Key issues include: (1) an inadequately described and physically questionable methodology (2) an overstated and poorly validated claim that the model improves upon existing products; (3) a misuse of the XCO2 enhancement framework, which cannot be straightforwardly interpreted as a proxy for CO2 emissions without accounting for wind speed and other confounding factors; and (4) several factual errors, inconsistent terminology, and unclear or internally inconsistent figure descriptions. Taken together, these issues represent substantial weaknesses that require major revision before the manuscript can be considered suitable for publication. Given the number and severity of the concerns raised below, the manuscript is not recommended for publication in its current form.
Specific Comments
Lines 37–40:
The argument that atmospheric CO2 data can be used to assess mitigation effectiveness and inform sustainable development is stated too broadly and lacks specificity. Connecting column-averaged XCO2 retrievals to surface emissions is scientifically challenging due to CO2's long atmospheric lifetime, its associated large and variable background signal, a low signal-to-noise ratio relative to emission-driven enhancements, and the difficulty of separating anthropogenic signals from biospheric fluxes. The authors should provide a more rigorous and nuanced discussion of how their dataset could be used for these purposes—acknowledging these limitations.
Line 50:
The manuscript states that OCO-2 has a "daily revisit capability." This is incorrect. According to the OCO-2 mission documentation, the satellite operates on a 16-day repeat cycle. This should be corrected, and the relevant citation should be verified accordingly.
Lines 99–101:
The citation of Crisp et al. (2017) as a reference for the CarbonTracker (CT) XCO2 product is inappropriate. Crisp et al. (2017) describes the algorithm theoretical basis for OCO-2 Level 2 retrievals, not the CarbonTracker data assimilation system. The authors should replace this with an appropriate reference for CarbonTracker CT2022 (e.g., Peters et al. or the relevant NOAA/GML documentation).
Lines 101–104:
The description of how the coarse-resolution CarbonTracker data (3° × 2°) are resampled to a 0.01° grid is insufficient. Also, the authors should clarify how the final XCO2 estimate is reconstructed—specifically, whether the predicted residual is added back to the CT XCO2 value.
Lines 115–117:
The use of daily-mean meteorological and air pollution data to predict XCO2 at the satellite overpass time (approximately 13:30 local time) requires physical justification. Atmospheric properties such as planetary boundary layer height, temperature, humidity, and trace gas concentrations can vary substantially throughout the day. Using daily-mean values rather than values contemporaneous with the satellite overpass may introduce systematic biases. The authors should either provide a physical justification for this approach or conduct a sensitivity analysis to evaluate how the temporal sampling of input predictors affects model performance for the target overpass time.
Lines 163–167:
The leave-one-year-out cross-validation strategy described here appears to withhold one year from within the 2015–2020 period for evaluation. However, the authors claim this approach allows assessment of model performance for the pre-2015 period. This logic is not convincing: withholding a year from the middle of the training period does not simulate the extrapolation challenge posed by hindcasting to years before the OCO-2 era, which may involve structural differences in the predictor-XCO2 relationships. The authors should clarify the validation approach and, if the goal is to assess pre-2015 performance, adopt a more appropriate out-of-sample evaluation strategy (e.g., training exclusively on post-2015 data and evaluating against ground-based observations in pre-2015 years).
Lines 183–185:
The use of a 10:30–16:30 time window to average ground-based observations for comparison with OCO-2 retrievals (overpass time ~13:30) is overly broad and requires justification. Atmospheric CO2 concentrations at surface sites exhibit pronounced diurnal variability driven by boundary layer dynamics and biospheric fluxes, meaning that measurements taken in the early morning or late afternoon may differ substantially from those at solar noon. Averaging over a six-hour window centered loosely on the overpass time may introduce significant biases in the validation. The authors should either narrow the averaging window (e.g., ±1–2 hours around the overpass time) or provide a sensitivity analysis demonstrating that this choice does not materially affect the validation statistics.
Lines 196–199:
The XCO2 enhancement method, while widely used in exploratory analysis, cannot be directly interpreted as an indicator of surface CO2 emissions without accounting for atmospheric transport, particularly wind speed and direction. An XCO2 enhancement above a background value reflects a combination of upstream emissions, atmospheric dilution (governed by wind speed), boundary layer height, and biospheric signals. High XCO2 enhancements under calm wind conditions may not correspond to higher emissions than lower enhancements under strong winds. The authors should either (1) incorporate a wind-speed correction or apply a more physically rigorous emission estimation framework, or (2) explicitly characterize the temporal variability of wind conditions over the study regions and discuss the extent to which this limits the interpretation of enhancements as emission proxies.
Figure 1g:
A visible data gap appears in the estimated XCO2 time series for approximately 2002–2003. Is it attributable to missing input data (e.g., predictor variables), or a deliberate data exclusion decision? This should be addressed explicitly in the text.
Section 3.1.3:
This section discusses feature importance but focuses almost exclusively on MAIAC AOD, neglecting several other highly ranked predictors. Notably, total column water vapour emerges as the most important predictor in both the YRD and PRD regions, yet this is not discussed. The authors should provide a physical explanation for this result—is this relationship physically meaningful or potentially spurious? Similarly, Day of Year appears to be among the most important features across regions, raising the question of whether the model's representation of XCO2 seasonality is driven primarily by this temporal index rather than by physically meaningful predictors. The implications of this for spatial generalization and for hindcast periods should be discussed.
Lines 270–272:
There is a typographical error in the figure caption: "Figure 2. of each …" should be corrected.
Lines 276–280:
The comparison of R² and RMSE values across different models to argue that the present model outperforms previous studies is methodologically inappropriate. Model skill metrics are highly sensitive to the spatial domain, temporal coverage, and resolution of the evaluation dataset, as well as the choice of validation sites and periods. Without a controlled, identical evaluation framework applied to all compared models, such inter-model comparisons cannot support strong claims of superiority. The authors should reframe this comparison more cautiously, noting that direct performance comparisons are not possible across studies with differing spatiotemporal coverage and evaluation protocols.
Lines 294–296:
The claim that previous machine-learning studies producing daily 1-km XCO2 estimates have been limited to post-2015 data is made without citation. Please add appropriate references to support this statement.
Lines 303–305 and 309–310:
The claim that the high-resolution product captures intra-urban XCO2 variations "with greater accuracy" than CT is not substantiated. To support this claim, the authors must provide a direct comparison of CT XCO2 and the model-estimated XCO2 against independent validation data (i.e., data not used in training), with evaluation metrics reported for both. Without this, the claim that the machine learning model improves upon CT at fine spatial scales remains unverified.
Figure 3a:
A sharp spatial gradient in XCO2 is visible around the Taklamakan Desert region, which appears physically implausible for a remote arid region with minimal anthropogenic activity. The authors should investigate whether this feature is present in the OCO-2 retrieval data or the CarbonTracker product, and if not, identify which input variable(s) are driving this artifact.
Lines 351–352:
The terms "CO2," "XCO2," and "mixing ratio/concentration/level" are used inconsistently throughout the manuscript. XCO2 is a column-averaged dry-air mole fraction expressed in parts per million (ppm), not a concentration in the physical chemistry sense (e.g., mol/m³). The authors should adopt consistent and scientifically precise terminology throughout the manuscript and avoid referring to ppm values as "concentrations."
Figures 5c and 5d:
The interpretation of XCO2 enhancements as indicators of urban CO2 emissions in these figures is problematic. If a single background XCO2 value is used for each region, the spatial patterns shown primarily reflect the climatological XCO2 gradient across the region, which integrates regional wind transport, biospheric fluxes, and the regional XCO2 gradient—not local emission differences between cities. For example, the larger enhancements observed in the southern portion of the BTH region do not necessarily imply greater emissions than those in Beijing; they may simply reflect more favorable transport or boundary layer conditions. The authors should clarify the background correction methodology and substantially revise the interpretation of these figures.
Lines 383–385:
The description of Figure 5d is inconsistent with its caption. The text refers to "a striking difference in the spatial patterns of XCO2 enhancements between the two years" and reports specific percentage reductions during the COVID-19 lockdown in Wuhan, but the figure caption states it shows the "20-year mean XCO2 enhancements for the YRD region." It is unclear which figure the text is actually referring to. The authors should ensure that all in-text figure references are accurate and that figure captions correctly describe the displayed content.