the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
A Terrain-Aware Residual U-Net Framework for Hourly Kilometre-Scale 2 m Temperature Downscaling over Austria
Abstract. High-resolution hourly 2 m air temperature fields are essential for weather, climate, and impact applications in complex terrain, but coarse atmospheric products cannot directly resolve the local thermal structure produced by Alpine topography. This study develops and evaluates a terrain-aware residual U-Net for kilometre-scale hourly 2 m temperature downscaling over Austria and the surrounding Alpine region. The model learns the correction between bilinearly interpolated ERA5 reanalysis and the Integrated Nowcasting through Comprehensive Analysis (INCA) high-resolution temperature analysis using dynamic atmospheric predictors, digital elevation model-derived terrain descriptors, and cyclical time features. Performance is evaluated over an independent 2023–2025 test period against INCA and compared with two reference approaches: 1) interpolated ERA5 and 2) a physically interpretable lapse rate baseline with monthly orographic and bias corrections. In addition to conventional deterministic skill metrics, the evaluation examines whether the downscaled fields reproduce physically meaningful Alpine temperature structure, including spatial error patterns, hourly lapse rate variability, elevation-dependent diurnal cycles, high-frequency spatial variability, ridge-valley temperature contrasts, and heat- and cold-event biases. The residual U-Net substantially improves predictive skill relative to both reference methods. It reduces root-mean-square error from 2.39 °C for interpolated ERA5 and 1.97 °C for the lapse rate baseline to 1.22 °C and reduces mean absolute error from 1.69 °C and 1.37 °C for ERA5 and the lapse rate baseline, respectively, to 0.89 °C. These improvements correspond to root-mean-square error reductions of approximately 49 % relative to ERA5 and 38 % relative to the baseline. The model also weakens terrain-locked error structures, produces smaller and less spatially coherent biases, and better reproduces observed lapse rate variability, high-elevation diurnal cycles, small-scale temperature variability, nocturnal ridge-valley contrasts, and event-scale warm and cold biases. A source-transfer experiment using the NASA MERRA-2 reanalysis predictors without retraining shows that the learned correction remains beneficial relative to raw MERRA-2, but with reduced skill compared to the ERA5-driven model, indicating sensitivity to the driving reanalysis distribution. Overall, the results demonstrate that residual deep-learning downscaling can generate physically realistic INCA-scale hourly temperature fields in Alpine terrain, while also highlighting the need to treat such products as high-resolution emulations of the reference analysis rather than direct observations.
- Preprint
(2203 KB) - Metadata XML
-
Supplement
(7202 KB) - BibTeX
- EndNote
Status: open (until 31 Oct 2026)
-
CEC1: 'Comment on egusphere-2026-4059 - No compliance with the policy of the journal', Juan Antonio Añel, 08 Aug 2026
reply
-
AC1: 'Reply on CEC1', Kelsey Ennis, 09 Aug 2026
reply
Below is the edited Code and Data Availability section with the requested information. If you require additional information and/or the revised manuscript with the added dataset citations, please let us know.
Code and data availability
The U-Net architecture and downscaling Python scripts used in this paper are archived under the MIT licence at https://doi.org/10.5281/zenodo.21864594 (Ennis, 2026). INCA data (1 km spatial resolution, hourly temporal resolution) are available from GeoSphere Austria (2021, https://doi.org/10.60669/6akt-5p05). ERA5 hourly data on pressure levels (Hersbach et al., 2023a, https://doi.org/10.24381/cds.bd0915c6) and ERA5 hourly data on single levels (Hersbach et al., 2023b, https://doi.org/10.24381/cds.adbb2d47) from 1940 to present are available from the ECMWF/Copernicus Climate Data Store. MERRA-2 tavg1_2d_lnd_Nx land surface diagnostics (GMAO, 2015a, https://doi.org/10.5067/RKPHT8KC1Y1T) and tavg1_2d_slv_Nx single-level diagnostics (GMAO, 2015b, https://doi.org/10.5067/VJAFPLI1CSIV), both V5.12.4, are available from the NASA Goddard Earth Sciences Data and Information Services Center (GES DISC).
Citation: https://doi.org/10.5194/egusphere-2026-4059-AC1 -
CEC2: 'Reply on AC1', Juan Antonio Añel, 10 Aug 2026
reply
Dear authors,
Unfortunately, your proposed solution does not solve the outstanding issues with your submission. The sites that you have linked (GeoSphere Austria, NASA servers, and the Copernicus Data Store) do not comply with the requirements expressed in the previous comment.
Therefore, the situation with your submission is not solved. Please, provide a new Code and Data Availability section that contains suitable repositories for the data used in your work.
Juan A. Añel
Geosci. Model Dev. Executive Editor
Citation: https://doi.org/10.5194/egusphere-2026-4059-CEC2 -
AC2: 'Reply on CEC2', Kelsey Ennis, 10 Aug 2026
reply
We have uploaded and published all data (INCA, ERA5, MERRA2) used in the paper to our Zenodo repository. Please see the accordingly updated Code and Data Availability statement below, and let us know if anything remains incorrect and/or if you need additional information. This is our first time submitting to this journal, and we are new to the code/data availability requirements. We appreciate your patience and more specific instructions if the current changes are insufficient.
Code and data availability
The U-Net architecture and downscaling Python code (released under the MIT licence), together with the Austria-domain subsets of INCA, ERA5, and MERRA-2 data (released under CC-BY 4.0) used in this paper, are archived at https://doi.org/10.5281/zenodo.21865319 (Ennis, 2026). ERA5 (2012–2025) served as the predictor dataset used to train and evaluate the U-Net, and INCA (2012–2025) provided the high-resolution 2 m temperature data used for model validation. MERRA-2 (2023–2025) was prepared as an independent predictor dataset for a source-sensitivity experiment. The original data source archives, listed here for completeness, are the GeoSphere Austria INCA v1-1h-1km dataset (GeoSphere Austria, 2021, https://doi.org/10.60669/6akt-5p05), the ECMWF/Copernicus Climate Data Store ERA5 hourly data on pressure levels (Hersbach et al., 2023a, https://doi.org/10.24381/cds.bd0915c6) and single levels (Hersbach et al., 2023b, https://doi.org/10.24381/cds.adbb2d47) from 1940 to present, and the NASA Goddard Earth Sciences Data and Information Services Center (GES DISC) MERRA-2 tavg1_2d_lnd_Nx land surface diagnostics (GMAO, 2015a, https://doi.org/10.5067/RKPHT8KC1Y1T) and tavg1_2d_slv_Nx single-level diagnostics (GMAO, 2015b, https://doi.org/10.5067/VJAFPLI1CSIV), both V5.12.4.
Citation: https://doi.org/10.5194/egusphere-2026-4059-AC2 -
CEC3: 'Reply on AC2', Juan Antonio Añel, 11 Aug 2026
reply
Dear authors,
Thanks for addressing this issue so quickly. I have checked the repositories and we can consider now the current version of your manuscript in compliance with the code policy of the journal.
Juan A. Añel
Geosci. Model Dev. Executive Editor
Citation: https://doi.org/10.5194/egusphere-2026-4059-CEC3
-
CEC3: 'Reply on AC2', Juan Antonio Añel, 11 Aug 2026
reply
-
AC2: 'Reply on CEC2', Kelsey Ennis, 10 Aug 2026
reply
-
CEC2: 'Reply on AC1', Juan Antonio Añel, 10 Aug 2026
reply
-
AC1: 'Reply on CEC1', Kelsey Ennis, 09 Aug 2026
reply
-
RC1: 'Comment on egusphere-2026-4059', Anonymous Referee #1, 19 Aug 2026
reply
General
The manuscript describes 2m temperature downscaling experiments in the Alpine region using ML methods, with the physics-based INCA high-resolution hourly analysis as a training and verification dataset. The material is well presented and the paper clearly written. It shows that many (but not all) of the features in INCA fields can be emulated with ML methodology.
Suggested additions
1) As INCA is used as the 'truth' here, it would be good to include a few more sentences on its characteristics (Haiden et al, 2011). For example, that it starts from NWP forecasts as a first guess and corrects this first guess using interpolated station observations, where the interpolation takes into account the stratification of the atmosphere by using potential temperature as a vertical coordinate. In essence, when there is stable stratification, the vertical scale of the corrections is reduced. Also, unlike most other data assimilation methods it exactly reproduces the observed values at the station locations (within the accuracy possible on a 1 km grid).
2) It is great that the evaluation part looks beyond mean scores and checks the high-frequency spatial variability etc. However, it may be instructive to show one challenging example (e.g. when a front passes through) of a sequence of hourly INCA and U-Net analyses (and their difference) side by side to illustrate their behaviour. This would help confirm that the U-Net produces physically plausible analyses even in cases where the diurnal temperature evolution deviates more strongly from the climatological mean.
Minor items
L186: How were the predictors selected? Was there a trial-and-error process? Please describe. Also, is there a way to say how much weight the U-Net gave to each of them overall? I am for example wondering about 300 hPa zonal wind, whether this parameter actually improves the analysis.
L417: How sensitive is the result to the exact choice of the coefficients? Did you make some experiments on this?
L691-693: It is expected that U-Net has larger errors when starting from a large-scale analysis like MERRA that has larger errors than ERA5. It may be worth pointing out that the relative reduction of RMSE and MAE provided by U-Net relative to the respective large-scale analysis is also slightly reduced for MERRA. You could express the error reduction in %. This gives the reader a quantitative idea of the transferability of the training.
L771: Fig. 12a,b,d instead of Fig. 10a,b,d?
Citation: https://doi.org/10.5194/egusphere-2026-4059-RC1
Viewed
| HTML | XML | Total | Supplement | BibTeX | EndNote | |
|---|---|---|---|---|---|---|
| 180 | 80 | 25 | 285 | 50 | 29 | 24 |
- HTML: 180
- PDF: 80
- XML: 25
- Total: 285
- Supplement: 50
- BibTeX: 29
- EndNote: 24
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
Dear authors,
Unfortunately, after checking your manuscript, it has come to our attention that it does not comply with our "Code and Data Policy".
https://www.geoscientific-model-development.net/policies/code_and_data_policy.html
In the Code and Data Availability section you do not provide suitable repositories for the INCA, ERA5 and MERRA2 data used in your work. The sites that you link:
- They do not appear to have a published policy for data preservation over many years or decades (some flexibility exists over the precise length of preservation, but the policy must exist).
- They do not appear to have a published mechanism for preventing authors from unilaterally removing material. Archives must have a policy which makes removal of materials only possible in exceptional circumstances and subject to an independent curatorial decision,
If we have missed a published policy which does in fact address this matter satisfactorily, please post a response linking to it. If you have any questions about this issue, please post them in a reply.
The GMD review and publication process depends on reviewers and community commentators being able to access, during the discussion phase, the code and data on which a manuscript depends, and on ensuring the provenance of replicability of the published papers for years after their publication. Please, therefore, publish your code and data in one of the appropriate repositories and reply to this comment with the relevant information (link and a permanent identifier for it (e.g. DOI)) as soon as possible. We cannot have manuscripts under discussion that do not comply with our policy.
Later, if the Topical Editor decides to continue with the review or publication process of your manuscript and you are requested to upload a new version of it, then The 'Code and Data Availability’ section of your manuscript must also be modified to cite the new repository locations, and corresponding references added to the bibliography.
I must note that if you do not fix this problem, we cannot continue with the peer-review process or accept your manuscript for publication in GMD.
Juan A. Añel
Geosci. Model Dev. Executive Editor