the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
Generative reconstruction of high-resolution historical climate fields from station observations
Abstract. Reconstructing complete, high-resolution daily meteorological fields over multidecadal periods from station observations is severely underdetermined: stations are sparse, unevenly distributed, and temporally inconsistent, especially over complex terrain. Because this ill-posedness requires prior assumptions, the choice of prior determines how well extremes are preserved. Conventional products, whether interpolation, multi-source fusion, or reanalysis, rely on prescribed background-error covariances that attenuate localized extremes as coverage declines. We introduce Generative REconstruction (GRE), a sequential assimilation framework that replaces prescribed covariances with a generative prior learned from satellite-era gridded fields. The prior, a diffusion model trained jointly on precipitation, pressure, wind, and temperature, encodes multivariate spatial structure including heavy-tailed extremes. At each step, posterior sampling conditions the prior on available stations and a forecast background, producing an ensemble whose spread reflects observational coverage. Applied over Southwest China (1951–2024), GRE yields daily 0.1° ensemble reconstructions of all four variables. Conditioned on 80 % of stations, GRE achieves R = 0.93 and RMSE = 5.8 mm/day for daily precipitation versus 0.69 and 7.1 for the baseline product, recovering rather than smoothing localized peaks. The joint prior propagates observational information across variables, constraining fields not directly observed. Over three decades, reconstructed precipitation extremes remain consistent with the reference climatology while ensemble spread contracts as gauges densify, evidencing calibrated long-term uncertainty. The forecast background adds most value under sparse coverage and persistent synoptic conditions, maintaining bounded physically consistent uncertainty throughout the early record.
- Preprint
(9822 KB) - Metadata XML
-
Supplement
(7287 KB) - BibTeX
- EndNote
Status: open (until 23 Oct 2026)
- CEC1: 'Comment on egusphere-2026-4341 - No compliance with the policy of the journal', Juan Antonio Añel, 09 Sep 2026 reply
Viewed
| HTML | XML | Total | Supplement | BibTeX | EndNote | |
|---|---|---|---|---|---|---|
| 115 | 57 | 21 | 193 | 26 | 19 | 19 |
- HTML: 115
- PDF: 57
- XML: 21
- Total: 193
- Supplement: 26
- BibTeX: 19
- EndNote: 19
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
Dear authors,
Unfortunately, after checking your manuscript, it has come to our attention that it does not comply with our "Code and Data Policy".
https://www.geoscientific-model-development.net/policies/code_and_data_policy.html
You have archived (or provide as links to access) the data used and produced in your work in several sites that do not fulfil GMD’s requirements for a persistent data archive. We refer here specifically to the sites where are hosted the CMFD)version 2.0 data, the observations provided by the China Meteorological Data Service Centre, and all the output data (which you have not shared). The mentioned sites do not comply with the policy of the journal because:
- They do not appear to have a published policy for data preservation over many years or decades (some flexibility exists over the precise length of preservation, but the policy must exist).
- They do not appear to have a published mechanism for preventing authors from unilaterally removing material. Archives must have a policy which makes removal of materials only possible in exceptional circumstances and subject to an independent curatorial decision,
If we have missed a published policy which does in fact address this matter satisfactorily, please post a response linking to it. If you have any questions about this issue, please post them in a reply.
Also, you have used the Pysolar library; however, you have not identified its version and you do not provide a valid repository for it.
The GMD review and publication process depends on reviewers and community commentators being able to access, during the discussion phase, the code and data on which a manuscript depends, and on ensuring the provenance of replicability of the published papers for years after their publication, and your manuscript should have not been accepted for Discussions due to the above mentioned issues. Please, therefore, publish your code and data in one of the appropriate repositories and reply to this comment with the relevant information (link and a permanent identifier for it (e.g. DOI)) as soon as possible. We cannot have manuscripts under discussion that do not comply with our policy.
Later, if the Topical Editor decides to continue with the review or publication process of your manuscript and you are requested to upload a new version of it, then The 'Code and Data Availability’ section of your manuscript must also be modified to cite the new repository locations, and corresponding references added to the bibliography.
I must note that if you do not fix this problem, we cannot continue with the peer-review process or accept your manuscript for publication in GMD.
Juan A. Añel
Geosci. Model Dev. Executive Editor