the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
NUKLEUS – A First Kilometre Scale Multi-model Climate Ensemble for Germany: Evaluation
Abstract. This study presents the evaluation of NUKLEUS, the first kilometre-scale, multi-model convection-permitting regional climate ensemble (CPM) for Germany. Three state-of-the-art regional climate models (ICON-CLM, COSMO-CLM, and REMO) were run at ~3 km horizontal resolution using a two-step downscaling chain driven by ERA5 reanalysis. The ensemble provides high-resolution climate information for Central Europe. In particular, we evaluate the CPM simulation results for Germany with a focus on six representative pilot regions selected within the German RegIKlim programme. Temperature, precipitation, global radiation, and near-surface wind, including their spatial patterns, annual and diurnal cycles, distributional characteristics, are compared to observational datasets. Moreover, selected climate indices relevant for heat and precipitation extremes are analysed. Overall, the ensemble demonstrates substantial added value compared to coarser-scale regional climate modelling, particularly in capturing regional climatic features and fine-scale variability. Temperature is reproduced with small biases (mostly within ± 0.5 K), with ICON-CLM performing best, while COSMO-CLM shows a weak cold bias and REMO a warm bias in parts of southern Germany. The CPM models realistically capture daily temperature distributions, though REMO underestimates minimum-temperature extremes. The annual cycle of precipitation is generally well represented, but all CPM models tend to overestimate totals in several regions, e.g., REMO exhibits a distinct spatial bias pattern with stronger deviations along topographic gradients. Extreme precipitation frequencies are generally overestimated, while regional contrasts such as stronger extremes in mountainous regions are preserved. Diurnal cycles show deficiencies specific to the models, including timing errors of afternoon precipitation peaks and misrepresentation of nocturnal precipitation revivals. For global radiation, ICON-CLM achieves the smallest biases, benefiting from its modern radiation scheme (ecRad), whereas COSMO-CLM and REMO show region specific over- and underestimations linked to cloud representation. 10 m wind speed diurnal cycles are best simulated by ICON-CLM, which captures the nocturnal wind minimum, while COSMO-CLM and REMO generally overestimate nighttime wind. Climate indices reveal underestimation of heat-related metrics (summer days and hot days) and systematic overestimation of heavy precipitation indices, although spatial patterns and regional differences are reproduced. In summary, we conclude that NUKLEUS provides valuable climate information for Germany, supporting climate-impact assessments and adaptation planning at municipal to regional scales.
- Preprint
(2669 KB) - Metadata XML
-
Supplement
(3325 KB) - BibTeX
- EndNote
Status: open (until 17 Aug 2026)
-
CEC1: 'Comment on egusphere-2026-1024 - No compliance with the policy of the journal', Juan Antonio Añel, 26 Jun 2026
reply
-
AC1: 'Reply on CEC1', Kevin Sieck, 09 Jul 2026
reply
Dear Juan Antonio Añel,
thank you for your comment and apologies for the late response. We are working on both issues which takes a bit more time than expected. We will post our solutions here as soon as possible.
Kind regards
Kevin Sieck
Citation: https://doi.org/10.5194/egusphere-2026-1024-AC1 -
CC1: 'Reply on CEC1', Florian Ehmele, 11 Aug 2026
reply
Dear Juan A. Añel,
on behalf of all co-authors, I would like to give an update on this. Analogous to the companion paper by Braun et al. (https://doi.org/10.5194/egusphere-2026-2517), we initiated the publication of the simulation data associated with the manuscript. Currently, the corresponding summary of the data is available via this link (https://www.wdc-climate.de/ui/entry?acronym=DKRZ_LTA_1203_ds00005), which will later be replaced by a permanent handle, when the archiving process is finished. We will provide this as soon as possible.
For the publication of the ICON model code, we also found a solution and will implement it and provide the information also as soon as possible.
Kind regards,
Florian EhmeleCitation: https://doi.org/10.5194/egusphere-2026-1024-CC1
-
AC1: 'Reply on CEC1', Kevin Sieck, 09 Jul 2026
reply
-
RC1: 'Referee comment to "NUKLEUS - A First Kilometre Scale Multi-model Climate Ensemble for Germany: Evaluation" by Sieck et al. (preprint for GMD, DOI: 10.5194/egusphere-2026-1024)', Anonymous Referee #1, 10 Aug 2026
reply
# SUMMARY
The manuscript by Sieck et al. presents the evaluation of a regional climate model (RCM) ensemble with the COSMO, ICON (in their regional modelling versions, COSMO-CLM and ICON-CLM), and REMO RCMs at convection-permitting resolution (about 3km grid spacing). Model runs are part of the NUKLEUS simulation experiment, a first multi-model km-scale climate change ensemble with a focus on Germany ("CEU-3" domain). The primary goal of NUKLEUS is to provide applicable and usable information for Vulnerability, Impacts, Adaptation, and Climate Services (VIACS) use cases and applications. The limited area modelling experiment design as presented here follows a one-way double nest dynamical downscaling approach for timeslices. In the initial step of the overall NUKLEUS experiment, the experiment design is presented; the model setups and configurations are defined; based on a dynamical downscaling of the ERA5 reanalysis and an extensive evaluation with observational reference data, the fidelity of the overall setup, and model configurations is assessed for a 10-year hindcast timespan from 2005 to 2014. As the evaluation paper, this is the inital paper, it constitutes the basis of the NUKLEUS simulations and is the start of a series of complementing NUKLEUS papers. Historical simulations and climate projections will be covered in seperate studies by Braun et al. (2026, "NUKLEUS – A first kilometer-scale convection-permitting multi-model climate ensemble for Germany: Characteristics of the historical simulations 1961–1990", https://doi.org/10.5194/egusphere-2026-2517), and Beier et al. (2026); the added-value of the 3km simulations vs the 12 km EUR-12 driving simulations will be a publication by Baumann and Paeth (2026).
# GMD REVIEW ITEMS TO CHECKIn the following, GMD review criteria are individually addressed by the reviewer, item by item (see https://www.geoscientific-model-development.net/peer_review/review_criteria.html); some are just comments to the editor, others are also meant as suggestions to the authors for improveents or clarifications:
"1. Does the paper address relevant scientific modelling questions within the scope of GMD? Does the paper present a model, advances in modelling science, or a modelling protocol that is suitable for addressing relevant scientific questions within the scope of EGU?"
The paper is suitable in form and content for a publication with GMD.
The authors present a new simulation experiment, a multi-model regional climate change ensemble with state-of-the-art RCMs at convection permitting resolution. Such simulations are still relatively rare, especially when done in a concerted manner.
There is much proven added value when going from convection-parametrized to convection-permitting resolutions at km-scale grid spacings for the reproduction of small scale heterogeneity, surface atmosphere interactions, convective processes, reproduction of dynamical features, etc.
The data produced with the experiment design or modelling protocol supports an advancement of regional climate modelling science as this is the first km-scale ensemble for the region of Germany. The data generated may be used for further processes and VIACS studies.
The experiment thereby produces evaluated (not optimized) model configuration and setup information, gives evidence of model fidelity, provides a simulation protocol synchronized with relevant stakeholders, and certainly also uncovered and helped resolve many HPC-related technical obstacles.
The experiment design as well as the analysis follows best practice approaches without being overly innovative (one-way double nest, reanalysis downscaling, NWP model configurations). The meticulously done evaluation, with suitable reference data, systematically pinpoints model deficits for ensuing model improvement or data preparation (bias adjustments).
As convection-permitting simulations seem to become a new standard, experiments such as NUKLEUS help to gain experience in all aspects of running and maintaining such simulations (runtimes, configurations, data volumes, reference data, model behaviour, deficits in inputs and parametrisations) and thereby promote such simulations.
The papaer does not make an attempt for an in-depth explanation of biases. This may be a weakpoint but can be considered valid, as the focus is strictly on evaluation. Likewise no computational aspects or "lessons learned" are addressed. This in essence manifests itself with simulation information (i.e., how were the simulaitons run, configurations, input data, see review item 6).
"2. Does the paper present novel concepts, ideas, tools, or data?"
The authors follow a well proven experiment design, starting off with evaluation run as a basis for a follow up climate change scenario simulation.
Model systems are state-of-the-art; many users will most likely be interested in the performance of the ICON-CLM, which can be considered the next generation model over COSMO and REMO.
The generated data is novel, see review item 1.
The specific focus on and evaluation for six pilot regions across the model domain, based on a stakeholder co-design approach, can be considered novel (here more information would be interesting on the process, L.100ff, but certainly beyond the scope of the paper). In the evaluation, the specific characteristics of the pilot regions are addressed in the interpretation of the biases.
Aside from only assessing thermal and hygric quantities, 10m wind speed and global radiation seem clearly tailored to renewable energy, using km-scale information for wind energy potential analyses.
It can be considered a novel concept in regional climate modelling to run coordinated km-scale climate change ensembles. NUKLEUS facilitates that such experiments become the new normal in regional climate modelling.
"3. Does the paper represent a sufficiently substantial advance in modelling science?"
Yes, given the combined configuration refinements or tests, newly produced datasets, and the thorough model evaluation and testing.
Modelling science could be more advanced if the authors provide more information on computational and big data aspects, as this is often a hindrance for such model runs, and leassons learned from the setup and configuration.
As mentioned above, the experiment is novel. Within the GCM realm, I am struggling whether this is to be considered a "model experiment description paper", as the information on the exact expeirment details may be a bit vague (see review item 6), also I would not consider this a generic benchmarking experiment; to me this appears more a "model evaluation paper" that provides "thorough performance and behavior evaluations of previously published models".
The discussion remains in places vague when it comes to addressing root causes of the biases and model behaviour and concludes that despite the substantial improvements we see, bias adjustment is still needed for many VIACS applications. This is OK, as the study focussses strictly on evaluation and a thourough assessment is beyond the scope, but perhaps make this very explicit to not raise false expectations with the readers.
"4. Are the methods and assumptions valid and clearly outlined?"
The overall experiment design is clearly explained, more details would help in reproducibility (see other review items). The dynamical downscaling, with a one-way double nest, is an established procedure. The evaluation follows in terms of variables considered and diagnostics and statistical measures common practice, incl. also the ETCCDI indices. I was surprised that you did not show any Taylor diagrams? The reference data are well chosen.
In addition to the evaluation for the six pilot regions, more information on all of Germany would have been interesting. The reviewer assumes that the pilot regions are considered represenative enough. Why do you show the EF and TSA regions in the main text? The STU domain would have had a similar number of meteorological stations as EF and be similar in size.
As the paper shall define a spefic simulation protocol, serve as a benchmarking, why was this specific timespan chosen?
The reader inevitably will ask why the 12km intermediate nest, for the well proven and widely used EURO-CORDEX domain, is not also covered to show the added value on-the-fly, see also a comment with the review item 5 on this. So perhaps mention the added value paper also earlier.
"5. Are the results sufficient to support the interpretations and conclusions?"
In principle yes. The rigorous evaluation uncovers systematic biases for some of the most relevant stakeholer variables. Overall there is a good agreement with observations. The NUKLEUS experiment certainly also provides "a substantial advancement in high-resolution climate information for Germany" (L.517f). The conclusions on model improvement mainly rely on a comparison with evaluation studies from literature.
Whether substantial added value is provided by the km-scale runs is to be expected, but remains from the analysis IMHO unclear, so "...offering improved representation of convective processes, regional climate features, and local extremes compared to coarser regional or global models." (L.518f) would need to be rephrased; the 12km intermediate nest vs 3km comparison is done in a seperate companion paper.
Although I think the statement as such is valid, that "Overall, the ensemble demonstrates strong potential as a foundation for actionable climate information, especially in the context of regional climate adaptation efforts within the RegIKlim program." (L.522ff), because the stakeholder requirements wrt quality criteria for their applications (a core motivation for the production of the ensemble) are not mentioned, there is no proof the data is actually fit for purpose, however, the authors cite Pinto et al. (2026, https://doi.org/10.1002/joc.70304) in the context of climate indices. Maybe rephrase this and/or further elaborate.
From the evaluation, it seems that ICON-CLM, not surprisingly perhaps, may be the best performing model, did this have implications for the further experiment design of NUKLEUS?
"6. Is the description sufficiently complete and precise to allow their reproduction by fellow scientists (traceability of results)?"
As this is a "model experiment description paper", the simulation experiments needs to be reproducible. At the time of writing this review, the discussion is open on how model source codes, input data, results datasets may be made publicly avaialable to ensure reproducibility (https://doi.org/10.5194/egusphere-2026-1024-CEC1 and https://doi.org/10.5194/egusphere-2026-1024-AC1). A similar discussion is ongoing with the companion paper on the historical simulations by Braun et al. (https://doi.org/10.5194/egusphere-2026-2517-CEC1). At this point the information provided in "Code and data availability" may be considered not sufficient to meet the requirements in as defined in https://www.geoscientific-model-development.net/policies/code_and_data_policy.html.
Aside from formal GMD requirements, a reproducible experiment is highly desirable to help the community to pick up this type of experiment and get started more quickly with their own km-scale workflows.
As the reviewer understands, there are several aspects and possible solution: (i) https://www-regiklim.dkrz.de will provide FAIR open access to the NUKLEUS data, right now data is free but access is restricted; in any case a Zenodo-based table (with DOI) may provide a listing of URIs to make the model results data files used in the study findable; (ii) the specific model versions may be made available as software packages (with DOI) via Zenodo; (iii) the input data to the models (static and external parameter fields, etc.) may also be shared through Zenodo with a DOI; (iv) the model configurations (i.e. parametrisation settings, etc.) as well as the scripts for pre- and postprocessing are packaged in case of ICON-CLM and COSMO-CLM via the SPICE and STARTER-Package integrated workflow engines, available through Zenodo; perhaps this has to be more clear in the main text as well as in the "Code and data availability" section; are the REMO configurations available somewhere public?; (v) evaluation analysis tools are at least to some degree part of FREVA package, FOSS available through https://zenodo.org/records/1325149 and https://github.com/freva-org/freva-legacy; (vi) most reference data (see also the discussion for Braun et al. companion paper, https://doi.org/10.5194/egusphere-2026-2517-CEC1) is available as FAIR, free, open acess reasearch data from their respective repositories.
How to deal with this situation of especially (i), the reviewer leaves open to the editor.
Maybe structure the information in the Code and Data Avilability section more clearly, so a reader not familiar with COSMO-CLM or ICON-CLM procedures around Starter Package and SPICE can easily understand.
Despite the principle availability of the configurations through the workflow engines for COSMO-CLM and ICON-CLM, perhaps add tabulated information to the supplement, which highlight the most relevant non-default configuration settings, also related to the km-scale resolution. Despite the description in Section 2.1 (e.g., L.150ff), the exact configurations are referred to being close to the NWP (L.150, L.190), but as a foundation paper for the follow-up companion papers, it might be very helpful to provide more information.
"7. Do the authors give proper credit to related work and clearly indicate their own new/original contribution?"
- In principle yes, especially from L.77ff onwards similar other experiments, like the CORDEX FPSCONV, are mentioned. Perhaps expand this overview by other similar modelling activities to NUKLEUS in Europe, like the simulation experiments with HCLM at 3km over Fenno-Scandinavia (e.g., Lind et al., 2022, https://doi.org/10.1007/s00382-022-06589-3) or the 2.2 km UKCP18-CPM UK ensemnle (Kendon et al., 2020, https://doi.org/10.1175/JCLI-D-20-0089.1). Also, previously km-scale time-sclice climate projections and evaluation runs for Germany or sub-regions exist, though no large ensembles, perhaps also mention those in this context, like Fosser et al. (2016, https://doi.org/10.1007/s00382-016-3186-4) or Knist et al. (2020, https://doi.org/10.1007/s00382-018-4147-x) and others.
"8. Does the title clearly reflect the contents of the paper? The model name and number should be included in papers that deal with only one model."
- Yes. Given there is a companion paper ("NUKLEUS – A first kilometer-scale convection-permitting multi-model climate ensemble for Germany: Characteristics of the historical simulations 1961–1990", see above), perhaps indicate this earlier on than just in the outlook. Although the ensemble is available for all of Germany, the title implies that the evaluation excercise covers all of Germany while in fact only the pilot regions are considered in detail.
"9. Does the abstract provide a concise and complete summary?"
- Yes.
"10. Is the overall presentation well structured and clear?"
- The manuscript follows a very clear, easy-to-follow structure. For example, the results sections 3.1 and 3.2 on temperature and precipitation are identically structured (spatial bias maps, annual cycles, bias plots, 2d bias plots (with MB, MAD), freq diststributions (with PSS + MAQD), diurnal cycles). Few things might be be improved as follows.
- Perhaps in L.86ff call this part simulation setting or framework, as this gives the overall experiment design.
- In L.130ff, there is a certain imbalance in model description sections between ICON-CLM, and COSMO-CLM and REMO, which are much shorter. This may be due to the fact they models are better known and ICON will be likely superseeding the models, but as all three RCMs are used in the NUKLEUS ensemble, perhaps the the model desciptions can be somewhat similar in structure and extend; see also review item 6.
"11. Is the language fluent and precise?"
- The manuscript reads well and fluent. Language is clear and precise. There's only a few specific remarks at the end of this review.
"12. Are mathematical formulae, symbols, abbreviations, and units correctly defined and used?"
- Yes. But perhaps provide more complete variable names, L.74ff, like total precipitation, 2m air temperature, global radiation, 10m wind speed.
"13. Should any parts of the paper (text, formulae, figures, tables) be clarified, reduced, combined, or eliminated?"
- Fig.1: Expand the figure to each side to show subset maps of the pilot regions incl., e.g, the DWD stations and land use and topographic information (isolines) perhaps; you argue that spatial heterogeneity and small scale features may be better resolved, this can be shown here. Also the Pilot regions are crucial for the evaluation, e.g., urbanisation may be important. Also add some information on the size, I think this is missing. Maybe this is described in another NUKLEUS paper, but as this paper deals with the "experiment desciption", this informaiton is valuable for the reader.
- Fig.2 and Fig.11: Overplot the six pilot region boundaries for orientation.
- As the paper follows a rigorous structure, see review item 10, this reaults in many figures. To me some figures might be combined as they are linked. E.g. consider combining temperature analysis Fig.3+4, Fig.5+6, Fig.7+8, Fig.9+10; precipitation analysis: Fig.12+13, Fig.14+15, Fig.16+17; rediation and widn speed perhaps also accordingly (Figs.18-21).
"14. Are the number and quality of references appropriate?"
- In principle yes, there are a number of statements in the text, which may be backed up by additional references, though, e.g. L.35f.
"15. Is the amount and quality of supplementary material appropriate? For model description papers, authors are strongly encouraged to submit supplementary material containing the model code and a user manual. For development, technical, and benchmarking papers, the submission of code to perform calculations described in the text is strongly encouraged."
- Supplementary informaition is highly important as the text contains only information on two pilot regions out of six. So the supplementary material is a logic extension / addons to the plots from the main text. Perhaps consider reordering the supplementary figures according to the sequence in the main text, move S4 to the end. See review item 6, perhaps it would be good to have tables with model parameterisation listings.
# SPECIFIC COMMENTSThe are on various parts of the manuscript, which do not fir anywhere else.
- L.25ff: Perhaps also mention the motivate that Europe experienceing accelerated strong climate change and also refer to governmental programmes in which NUKLEUS was run, in addition to the longer explanation in "Method and Data".
- L.29f: Urban planning should be in a separate sentence, the context, especially through "particularly", with the hydrometeorological extremes sounds misleading.
- L.49ff: In this context the computational cost vs domain size vs experiment simulation length may be addressed more in detail as this is an important aspect, cite, e.g., papers such as Schär et al. (2020, https://doi.org/10.1175/BAMS-D-18-0167.1).
- L.50ff: Perhapse reorder this part of the paragraph to distinguish the continental and global short runs vs the longer smaller domain experiments, still often for time slices only. There are all in their own right frontier simulations.
- L.52: Provide some examples.
- L.56ff: RCM emulators -> advote a seperate section, more lit, his is a bit too weak, this is an old text
- L.66: "pioneering" refers here to the regional model runs, maybe rephrase.
- L.71: In the context above, what is a "mini-ensemble"
- L.76f: Perhaps identify variables under investigation more clearly, like total prepcipitation, near surface air temperature, etc., although it certainly is intuitive.
- L.115: Give examples for the application of the models.
- L.254: Put station data in seperate paragraph to emphesize this data product.
# MINOR (TECHNICAL) CORRECTIONS- L.290, sentence: "...is slightly from..."
- L.293, typo: "gird"
- Tab.1, caption: TXin90, TNin90 not actually listed in table
- L.296, grammar: "present"
- L.297ff, remove sentence, superfluous: "The latter are..."
- Fig.6, caption, replace: "meaning of ..." -> "See Fig.4 for cell colour coding."
- L.352, references: refer to the Supplement Figs, check that all supplementary Figs are referenced in the main text
- L.483, typo: dates -> data
- L.516, sentence: "...will be the focus..."
Citation: https://doi.org/10.5194/egusphere-2026-1024-RC1
Viewed
| HTML | XML | Total | Supplement | BibTeX | EndNote | |
|---|---|---|---|---|---|---|
| 190 | 85 | 18 | 293 | 27 | 13 | 15 |
- HTML: 190
- PDF: 85
- XML: 18
- Total: 293
- Supplement: 27
- BibTeX: 13
- EndNote: 15
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
Dear authors,
Unfortunately, after checking your manuscript, it has come to our attention that it does not comply with our "Code and Data Policy".
https://www.geoscientific-model-development.net/policies/code_and_data_policy.html
Namely, there are two outstanding issues in your manuscript which need to be addressed. First, the exact version of the ICON model used for your work should be stored in a private repository and cited. Other versions of the model do not ensure the replicability of the results presented in your work. This can be for example a Zenodo private repository.
Also, you do not provide repositories for the input datasets used in your work, and you must do it.
Additionally, it is necessary that you provide better information and details about the data files produced in your work. Currently, the DKRZ hosts more than 614,000 files that respond to the NUKLEUS keyword and it is not possible to identify the ones that correspond to your work.
The GMD review and publication process depends on reviewers and community commentators being able to access, during the discussion phase, the code and data on which a manuscript depends, and on ensuring the provenance of replicability of the published papers for years after their publication. Please, therefore, publish your code and data in one of the appropriate repositories and reply to this comment with the relevant information (link and a permanent identifier for it (e.g. DOI)) as soon as possible. We cannot have manuscripts under discussion that do not comply with our policy.
Later, if the Topical Editor decides to continue with the review or publication process of your manuscript and you are requested to upload a new version of it, then The 'Code and Data Availability’ section of your manuscript must also be modified to cite the new repository locations, and corresponding references added to the bibliography.
I must note that if you do not fix this problem, we cannot continue with the peer-review process or accept your manuscript for publication in GMD.
Juan A. Añel
Geosci. Model Dev. Executive Editor