the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
Generative reconstruction of high-resolution historical climate fields from station observations
Abstract. Reconstructing complete, high-resolution daily meteorological fields over multidecadal periods from station observations is severely underdetermined: stations are sparse, unevenly distributed, and temporally inconsistent, especially over complex terrain. Because this ill-posedness requires prior assumptions, the choice of prior determines how well extremes are preserved. Conventional products, whether interpolation, multi-source fusion, or reanalysis, rely on prescribed background-error covariances that attenuate localized extremes as coverage declines. We introduce Generative REconstruction (GRE), a sequential assimilation framework that replaces prescribed covariances with a generative prior learned from satellite-era gridded fields. The prior, a diffusion model trained jointly on precipitation, pressure, wind, and temperature, encodes multivariate spatial structure including heavy-tailed extremes. At each step, posterior sampling conditions the prior on available stations and a forecast background, producing an ensemble whose spread reflects observational coverage. Applied over Southwest China (1951–2024), GRE yields daily 0.1° ensemble reconstructions of all four variables. Conditioned on 80 % of stations, GRE achieves R = 0.93 and RMSE = 5.8 mm/day for daily precipitation versus 0.69 and 7.1 for the baseline product, recovering rather than smoothing localized peaks. The joint prior propagates observational information across variables, constraining fields not directly observed. Over three decades, reconstructed precipitation extremes remain consistent with the reference climatology while ensemble spread contracts as gauges densify, evidencing calibrated long-term uncertainty. The forecast background adds most value under sparse coverage and persistent synoptic conditions, maintaining bounded physically consistent uncertainty throughout the early record.
- Preprint
(9822 KB) - Metadata XML
-
Supplement
(7287 KB) - BibTeX
- EndNote
Status: open (until 23 Oct 2026)
-
CEC1: 'Comment on egusphere-2026-4341 - No compliance with the policy of the journal', Juan Antonio Añel, 09 Sep 2026
reply
-
AC1: 'Reply on CEC1', Jie Chao, 17 Sep 2026
reply
Dear Editor,
Many thanks for your careful check, and apologies for the initial non-compliance. We have addressed each point below, and would be very happy to adjust further if any detail still falls short of the policy.
(1) CMFD v2.0
The CMFD v2.0 used in our work is preserved at the National Tibetan Plateau / Third Pole Environment Data Center (TPDC), with the persistent DOI https://doi.org/10.11888/Atmos.tpdc.302088 and the CSTR identifier https://cstr.cn/18406.11.Atmos.tpdc.302088. We will cite these identifiers explicitly in the "Code and Data Availability" section of the revised manuscript.(2) CMA observations and reconstructed output data
Following your suggestion, we have archived the 1951 subset of daily station observations (precipitation, 10-m wind speed and 2-m temperature) and the corresponding reconstructed 0.1° daily climate fields over China, together with the model code, on Zenodo, under the title "Official PyTorch Code of 'Generative reconstruction of high-resolution historical climate fields from station observations'". The Zenodo concept DOI that resolves to all archived versions is https://doi.org/10.5281/zenodo.21525754. The "Code and Data Availability" section of the revised manuscript will cite this DOI accordingly.(3) Pysolar
We apologise for the missing version. The version used in this work is pysolar==0.13, retrieved from the official Python Package Index at https://pypi.org/project/pysolar/0.13/. A software citation entry for Pysolar will be added to the reference list of the revised manuscript.We hope this resolves the concerns you raised. Please do let us know if any further revision is needed, and we will be glad to make it at your convenience.
With our thanks again for your time,
Yours sincerely,
Jie Chao
on behalf of all co-authorsCitation: https://doi.org/10.5194/egusphere-2026-4341-AC1 -
CEC2: 'Reply on AC1', Juan Antonio Añel, 22 Sep 2026
reply
Dear authors,
Thanks for your reply. Unfortunately, for the case of the CMFD v2.0 data it does solve the pending issues. The sites that you cite to access them continue having the problems that I mentioned in my previous comment. They do not have a long-term retaining policy or a data management policy. Therefore, we can not accept them to host the data used in your manuscript, and you must store the CMFD v2.0 data in a repository that complies with the policy of the journal. Due to this, your manuscript continues to not comply with the policy of the journal, and can not be considered for publication as it is.
Juan A. Añel
Geosci. Model Dev. Executive Editor
Citation: https://doi.org/10.5194/egusphere-2026-4341-CEC2 -
AC3: 'Reply on CEC2', Jie Chao, 23 Sep 2026
reply
Dear Editor,
Thank you very much for your clarification, and please accept our apologies for not fully addressing this issue in our previous response.
We now understand the journal's requirement more clearly. To address this issue and comply with the data policy of Geoscientific Model Development, we have now deposited the exact subset of CMFD v2.0 used in our manuscript in Zenodo as a permanently archived dataset:
https://doi.org/10.5281/zenodo.21525754
We have revised the Code and Data Availability section of the manuscript accordingly and included both the Zenodo DOI for the archived subset and the original CMFD v2.0 DOI.
We sincerely appreciate your patience and your guidance on this matter. We are sorry that our previous response did not fully resolve the issue, and we are very grateful for the opportunity to correct it properly. Should there be any further concerns, we would be very grateful for your guidance and will continue to make any necessary improvements.
Best regards,
Jie Chao
on behalf of all authorsCitation: https://doi.org/10.5194/egusphere-2026-4341-AC3 -
CEC3: 'Reply on AC3', Juan Antonio Añel, 23 Sep 2026
reply
Dear authors,
Thanks for addressing the pending issues. I have checked the repositories and we can consider now the current version of your manuscript in compliance with the code policy of the journal.
Juan A. Añel
Geosci. Model Dev. Executive Editor
Citation: https://doi.org/10.5194/egusphere-2026-4341-CEC3
-
CEC3: 'Reply on AC3', Juan Antonio Añel, 23 Sep 2026
reply
-
AC3: 'Reply on CEC2', Jie Chao, 23 Sep 2026
reply
-
CEC2: 'Reply on AC1', Juan Antonio Añel, 22 Sep 2026
reply
-
AC1: 'Reply on CEC1', Jie Chao, 17 Sep 2026
reply
-
RC1: 'Comment on egusphere-2026-4341', Anonymous Referee #1, 17 Sep 2026
reply
The authors developed a sequential assimilation framework to reconstruct historical daily precipitation at 0.1-degree resolution over southeastern China, with auxiliary meteorological variables including pressure, surface temperature, and surface wind speed. The authors demonstrated the value of this framework to reconstruct extreme precipitation events throughout the early record with limited or no observation data. The topic is important and the method is sound. I only have minor comments below.
General comments
- How are external influences outside the training domain accounted for the reconstruction of precipitation?
- For extreme events reconstruction in historical period, how can the gridded extremes not observed be validated in any way? In other words, are those extreme events which were not observed due to limited ground station or satellite observations coverage realistic or supported by other data?
Citation: https://doi.org/10.5194/egusphere-2026-4341-RC1 -
AC2: 'Reply on RC1', Jie Chao, 22 Sep 2026
reply
We appreciate the reviewer’s comments regarding external influences on the reconstruction and the validation of historical extremes, and we address these points below with corresponding revisions to the manuscript.
1. We agree that the treatment of external influences out of the reconstruction domain is a limitation of the current framework.
GRE explicitly incorporates topographic and seasonal information through the DEM and solar angle, but potentially relevant large-scale climate signals, such as ENSO, are not currently included as conditioning information. This may limit the ability of the reconstruction to adapt to historical large-scale circulation states that differ from those represented in the training climatology. We therefore regard the inclusion of such large-scale climate indices as an important direction for future development.
Other contemporaneous atmospheric information surrounding the reconstruction domain is not directly available as independent historical observations for much of the reconstruction period. If suitable long-term boundary or surrounding-domain information becomes available, it would be valuable to investigate whether explicitly incorporating it can further constrain the reconstruction.
We discuss the incorporation of additional large-scale conditioning information as an extension of the current framework in Sect. 4.Add in Section 4 “Conclusions”:
*GRE explicitly incorporates topographic and seasonal forcing through the DEM and solar angle, while other historically available large-scale information, such as SST indices or large-scale circulation fields, is not currently used as an explicit conditioning variable. In addition, the exact historical atmospheric state surrounding the reconstruction domain is not completely observed and can only be approximated retrospectively. Consequently, the diffusion prior may not fully adapt to large-scale circulation anomalies that are poorly represented in the satellite-era training period.*
2. For historical extremes, we appreciate the reviewer’s concern and agree that this is an important point for interpreting the reconstruction. In particular, it is helpful to distinguish between assessing whether GRE can recover an extreme event when the corresponding observation is withheld, and independently verifying the exact historical value at a location where no direct observation exists. In the latter situation, the exact value is inherently difficult to establish without additional independent evidence.
Our current evaluation therefore uses controlled withholding experiments, in which observations are available as reference data but are intentionally excluded from GRE. In Sect. 3.2, a subset of stations is withheld from assimilation and used exclusively for evaluation. Sect. 3.3 provides a more targeted assessment of extremes by withholding precipitation observations for two observed events of 74 and 151 mm day$^{-1}$, and then further withholding all current-day observations. These experiments do not fully resolve the uncertainty associated with truly unobserved historical events, but they provide a practical test of whether GRE can retain information about an extreme when its direct observational constraint is removed.
For locations and periods with very limited observational coverage, we therefore believe it is more appropriate to interpret an individual reconstructed grid-cell extreme as a plausible posterior realization rather than as an independently verified historical value. Its interpretation should be considered together with the available observations, the learned multivariate prior, the temporal background, and the ensemble uncertainty. As shown in Sect. 3.5, the spread of precipitation extremes is substantially larger during the sparsely observed early period and decreases as the gauge network becomes denser. We will revise Sect. 3.5 to make this limitation and interpretation clearer, and to avoid overstating the degree to which individual historical extremes can be independently verified.Add in Section 3.5 “Calibrated Ensemble Uncertainty Through Network Evolution”:
*For locations and periods without independent observations, the exact magnitude and location of an individual grid-cell extreme cannot be independently verified. Such values are therefore interpreted probabilistically: the GRE ensemble represents plausible historical states consistent with the available observations, the learned prior, and the forecast background, rather than a uniquely determined historical field. The larger ensemble spread during the sparsely observed early period reflects this remaining uncertainty.*
We thank the reviewer for the valuable comments and suggestions, which have greatly improved our manuscript.Yours sincerely,
Jie Chao
On behalf of all authorsCitation: https://doi.org/10.5194/egusphere-2026-4341-AC2
Viewed
| HTML | XML | Total | Supplement | BibTeX | EndNote | |
|---|---|---|---|---|---|---|
| 172 | 85 | 32 | 289 | 38 | 23 | 23 |
- HTML: 172
- PDF: 85
- XML: 32
- Total: 289
- Supplement: 38
- BibTeX: 23
- EndNote: 23
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
Dear authors,
Unfortunately, after checking your manuscript, it has come to our attention that it does not comply with our "Code and Data Policy".
https://www.geoscientific-model-development.net/policies/code_and_data_policy.html
You have archived (or provide as links to access) the data used and produced in your work in several sites that do not fulfil GMD’s requirements for a persistent data archive. We refer here specifically to the sites where are hosted the CMFD)version 2.0 data, the observations provided by the China Meteorological Data Service Centre, and all the output data (which you have not shared). The mentioned sites do not comply with the policy of the journal because:
- They do not appear to have a published policy for data preservation over many years or decades (some flexibility exists over the precise length of preservation, but the policy must exist).
- They do not appear to have a published mechanism for preventing authors from unilaterally removing material. Archives must have a policy which makes removal of materials only possible in exceptional circumstances and subject to an independent curatorial decision,
If we have missed a published policy which does in fact address this matter satisfactorily, please post a response linking to it. If you have any questions about this issue, please post them in a reply.
Also, you have used the Pysolar library; however, you have not identified its version and you do not provide a valid repository for it.
The GMD review and publication process depends on reviewers and community commentators being able to access, during the discussion phase, the code and data on which a manuscript depends, and on ensuring the provenance of replicability of the published papers for years after their publication, and your manuscript should have not been accepted for Discussions due to the above mentioned issues. Please, therefore, publish your code and data in one of the appropriate repositories and reply to this comment with the relevant information (link and a permanent identifier for it (e.g. DOI)) as soon as possible. We cannot have manuscripts under discussion that do not comply with our policy.
Later, if the Topical Editor decides to continue with the review or publication process of your manuscript and you are requested to upload a new version of it, then The 'Code and Data Availability’ section of your manuscript must also be modified to cite the new repository locations, and corresponding references added to the bibliography.
I must note that if you do not fix this problem, we cannot continue with the peer-review process or accept your manuscript for publication in GMD.
Juan A. Añel
Geosci. Model Dev. Executive Editor