the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
A PMP-inspired Evaluation Framework for Assessing Deep-Learning Earth System Models
Abstract. In recent years, Deep-Learning Earth System Models (DL-ESMs) have emerged as promising, computationally efficient complements to traditional Earth system models. Here, we present an evaluation framework for testing DL-ESMs from an Earth system model-development perspective using standardized diagnostics from the PCMDI Metrics Package (PMP). This framework allows DL-ESMs, including Ai2’s ACE2 and Google’s NeuralGCM, to be assessed with metrics that quantify their ability to reproduce climatology, major modes of variability, monsoon behavior, and precipitation variability relative to observational reference datasets and CMIP-class benchmarks. By evaluating DL-ESMs with tools commonly used for traditional models, we extend their assessment beyond short-range forecast skill and toward longer Earth system–relevant applications. The results identify encouraging strengths in several large-scale fields and modes of variability while also highlighting persistent challenges in precipitation, tropical variability, and long-run stability for some model versions. This evaluation is a critical step toward building trust in DL-ESMs, guiding future model development, and clarifying their fitness for Earth system science applications.
Status: open (until 14 Aug 2026)
- RC1: 'Comment on egusphere-2026-3295', Anonymous Referee #1, 27 Jul 2026 reply
-
RC2: 'Comment on egusphere-2026-3295', Anonymous Referee #2, 30 Jul 2026
reply
Overview
This manuscript applies the PCMDI Metrics Package (PMP) to four deep-learning atmospheric emulator configurations, and benchmarks them against CMIP6 models and observational references. The premise is correct that forecast-oriented scores are not sufficient to establish whether a data-driven model is fit for Earth-system purposes, and the natural way to close that gap is the diagnostic battery routinely used for CMIP-class models. The breadth of diagnostics is a genuine strength, and several findings are worth publishing.
My concerns are as follows: Several methodological choices are not documented in enough detail to be reproduced or assessed, most importantly the CNN-based ad-hoc machine learning method, which underpins all of Section 4.2, and the stability screening. The comparison against CMIP6 is worth refinement. And no uncertainty is reported anywhere, although the ensembles range from ten members to one.
Note on referencing: the preprint carries no line numbers, so comments below are keyed to page, figure, or section number.
Major Comments
- Scope and novelty. The paper applies an already-published package (Lee et al., 2024, GMD) to four existing models, so the technical novelty needs to be stated and delivered. On my reading, the new elements are (i) the workflow that makes DL-ESM output ingestible by PMP and (ii) the CNN psl reconstruction; neither is described reproducibly.
- "PMP-inspired" understates what was done — PMP itself is used, so "PMP-based" is more accurate. Also, all four configurations are atmosphere-only emulators with prescribed SST and SIC, so "Earth System Models" without qualification overstates their scope. Please qualify at first use.
- Ad-hoc machine learning strategy for sea-level pressure (p. 6). A residual CNN predicts psl from several inputs, but no architecture, training period, loss function, or held-out validation is given. This matters because all four modes in Section 4.2 are psl-based, so that section evaluates the DL-ESM convolved with an undocumented emulator. Please provide the full specification and an independent validation, and address the circularity: if the network is trained toward ERA5 psl, scoring it against ERA5 in Fig. 2 partly measures the emulator. Applying the same network to a CMIP6 model's surface pressure and comparing against that model's native psl would bound the effect.
The claim that the psl results are "encouraging" is also contradicted by Fig. 2, where the NeuralGCM psl cell is saturated at the top of the color scale, and Fig. 3, where both DL-ESM psl markers lie far outside the CMIP6 violin. Please revise and discuss the consequences for Section 4.2.
- Stability screening and member selection (p. 5). "Stable" is never defined. Please state the criterion and give, per model and per diagnostic, the number of usable members and the exact start and end dates:
- Figure 1. Panels (c) and (d) are both labeled "GPCP 3.2" but report different stats. Please explain and harmonize. The analysis period is annotated on only two of the six panels.
- Figure 2. Normalizing by the median of the CMIP6 models together with the DL-ESMs means the benchmark moves as models are added. Please normalize by the CMIP6-only median, or report both, so that a DL-ESM score can be read as relative to the CMIP class.
- Figure 3. Several axes extend to negative values, but RMSE is non-negative; centering each axis on the multi-model median has pushed the limits into physically impossible territory. Please clip at zero or use a logarithmic scale.
- Figure 4. The caption states that the CMIP6 results are "historical or equivalent," whereas Figs. 7 and 9 label the experiment "amip." Comparing SST-forced DL-ESMs against freely coupled historical runs for modes of variability is not like-for-like. Please state per figure which experiment was used.
- Uncertainty quantification. The manuscript only includes some acknowledgments of sampling uncertainty, and it is qualitative. There are no confidence intervals or significance tests accompanying any metric. Statements such as "NeuralGCM shows slightly superior ability to reproduce NAO" may not survive. Please add ensemble-based or bootstrap intervals, and soften comparisons that do not.
- Discussion of model errors. Explanations are offered as alternatives rather than tested: NeuralGCM-evap's precipitation errors are attributed variously to the shorter stable analysis period, a shorter training period, extrapolation, and SST sensitivity, with no attempt to discriminate among them. I would encourage the authors to exploit this further. If they prefer to keep the scope purely diagnostic, which is defensible, please moderate the claims in Section 5 about guiding model development and delivering process-specific benchmarks.
Minor Comments
- Please make the abstract quantitative: state that the configurations are atmosphere-only, give the evaluation period, and give at least one number per finding.
- Table 1. The "Boundary Conditions" column contains only a citation and is identical in every row; please name the dataset, and consider adding the training period and the analysis period used here.
- Figure 8. The orange horizontal line is not defined in the caption.
- Figure 10. The caption describes blue stars, red dots, and green triangles, but the legend uses dots for all three categories.
- For example, "Eart System modeling" (p. 3) should read "Earth"; "among the the PMP's ETMoV metrics" (p. 11) has a duplicated word, etc. Please proofread the whole manuscript before resubmission.
- Cross-references. "Section 2" and "Sect. 3" are used alternatively.
- "DL-ESM," "AI emulator," and "AI-based model" are used for the same objects; please settle on one.
Citation: https://doi.org/10.5194/egusphere-2026-3295-RC2 -
RC3: 'Comment on egusphere-2026-3295', Anonymous Referee #3, 31 Jul 2026
reply
This manuscript presents a realistic and timely framework for comparing and validating the performance of different Deep Learning Earth System Models (DL-ESMs), which provides a very useful contribution to the Earth system modelling community. The manuscript is generally well written and clearly structured. However, there are a few key points and technical details that could be further clarified or expanded upon to strengthen the manuscript.
As the paper is submitted to GMD, it would be beneficial for the authors to further highlight the model development aspects of this work. Given that the core evaluation package (PMP) has been published previously, could the authors better articulate what new developments were implemented in PMP for this study, and what new developments were in the pre- and post-processing of DL-ESM outputs? At present, the manuscript reads more like an application of PMP rather than a development framework, and clarifying this distinction would help align it firmly with GMD's scope.
Additionally, a few general aspects regarding the core content could be improved:
- Explanations for some key metrics are missing. Furthermore, the visual presentation of some figures could be enhanced (e.g., increasing font sizes for readability and reducing large blank spaces without useful information).
- The rationale behind selecting specific realizations and study periods throughout the paper is not always clear. While the authors note "exposing model-specific limitations such as missing output variables, post-processing requirements, and stability constraints", it would be clearer to explicitly outline the criteria for choosing these different periods and realizations.
- While ta-200, ta-850, va-200, va-850, zg-200, and zg-850 are listed as key variables, there is limited discussion of them in the main text.
- The “stability analysis” seems to be missing.
- Adding a schematic diagram showing the complete workflow—including pre-processing, post-processing of DL-ESM outputs, and the integration with PMP—would greatly assist readers in visualizing the overall framework.
- Overall, deepening the discussion of the results in a few key sections would help maximize the paper's impact.
Major revisions
- Page 4, second paragraph. Regarding the statement, “These models differ both in architecture and in the way they represent atmospheric evolution,” while the structural differences between ACE2 and NeuralGCM are clear, could the authors briefly elaborate on the specific structural differences among the three NeuralGCM variants?
- Page 4, last paragraph. Could you please clarify whether the random perturbations are applied to all model parameters or only a subset? Additionally, specifying their statistical distributions (e.g., Gaussian vs. uniform) could be very useful.
- Page 5, first paragraph. Why the way of generating initial condition ensemble for ACE2 and NeuralGCM are different? The first one is perturbated by choosing between different days and the latter is by choosing from different months.
- Page 5, second paragraph. Could the authors please clarify how "long-run" stability is defined and what the "end-date threshold" refers to (e.g., is it a specific precipitation threshold)? Rewording this paragraph would clear up any potential confusion.
- Table 2. Sea level pressure is listed as one of the variables to be evaluated, while in Page 2 second paragraph, the “global mean log logarithm of surface pressure” is fixed. What’s the impact of this correction on the evaluation? And what’s the main purpose of this correction?
- Figure 1. It is not clear how the RMSE is calculated (e.g., across all grid cells over time versus time series of global averages). In Figure 1c, the longitudinal tick labels appear partially cut off, and the time period for the map seems to be missing.
- Page 8. Regarding the sentence “while the remaining RMSE (0.80 mm day−1) highlights persistent tropical amplitude errors”, could the authors clarify how the "persistent" is defined or demonstrated here?
- Page 10, line 8. The “lack of stability” needs to be explained.
- Page 10. Could the authors explain the “spatiotemporal RMSE”, and how it is different from other RMSE calculations.
- Figure 4. What criteria were used to select the specific ensemble members presented in this calculation?
- Figure 5. For SAM_MAM, NAO_DJF, and PNA_JJA, both two DL-ESMs show quite large biases compared to the medians, and the directions of the biases are consistent for these two DL-ESMs. Could the authors offer some insight or hypotheses as to what might be driving these consistent biases?
- Figures 5, 6 and 9. Do the r1i1p1f1, r53i1p1f1, and r100i1p1f1 have the same meaning as in CMIP data naming conventions? Could you explain why these specific realizations were selected for display, and why the f tag is omitted in Figure 9?
- Page 14, last paragraph. The differences in EWR among three DL-ESMs and obs are compared here. However, based on Figure 7, none of the three DL-ESMs reproduce the dominance of wavenumber 1 and 3 appeared in the observations. Rather, all of them show the dominance of wavenumber 2. It would be great if the discussion could address this wavenumber shift alongside the EWR values. Also, based on EWR, NeuralGCM-evap better replicates the obs. But between ACE2 and NeuralGCM-precip, which one is better? Why were different time periods selected for the DL-ESMs versus the obs in this analysis?
- Page 16, first paragraph. Regarding the sentence “These comparisons can be potentially affected by the ensemble-size differences, because the DL-ESM ensembles contain 10 members, whereas, several CMIP6 models contribute to the analysis with fewer members”, it would be helpful to discuss how sensitive the DL-ESMs' mean EWR is to the chosen ensemble size.
- Figure 8. In terms of May-October EWR, the three DL-ESMs appear to outperform some CMIP6 models. It would be helpful to have some figures of the wavenumber-frequency power spectra for CMIP6 models.
- Page17. In the sentence “They also produce more widespread false alarms across the tropical oceans”, could the authors provide quantitative counts for these false alarms to complement the qualitative description?
- Figures 9 and 10. Could you please clarify why different study periods were selected for Figure 9 versus Figure 10?
- Page 20. Regarding the sentence “However, the NeuralGCM-precip precipitation/convection formulation appears to improve behavior at synoptic timescales relative to longer-period variability”, could the authors present additional visual evidence (such as spectral power ratios across frequencies) to further back up this?
- Page 20. In sentence “an apparently reasonable global amplitude can mask poor geographic placement”, could you elaborate on what defines a "reasonable" amplitude here? Does ACE2's amplitude contribute significantly to its high spatial correlation, and does this suggest that visual inspection of spatial maps remains necessary alongside summary metrics?
- Page 21. In the sentence “NeuralGCM is comparable to the best-performing CMIP6 models, whereas ACE2 lies closer to the lower-skill edge of the CMIP6 distribution for this field”. Could you explain what is this lower-skill edge?
- Page 23. In the sentence “NeuralGCM-evap and NeuralGCM-precip show greater sensitivity to initial conditions, random seeds, and run length”, where is the discussion to support this?
- Page 23, third paragraph. Please ensure that appropriate literature references are added to complete this paragraph.
Minor revisions
- Page 2 line 28. It would be helpful to include a few concrete examples of standard weather-forecast scores, along with a brief note on their respective strengths and limitations.
- Page 3 line 24. Typo: please change to "Earth System Modeling".
- Page 9, line 11. A summary table for all the annual and seasonal climatology fields could be helpful. Additionally, increasing the font size in the Supplementary Figures would improve readability.
- Page 9, line 13. Could you provide an explicit formula or example showing how the global spatial RMSE is calculated?
- Figure 2. Please consider increasing the font size for clarity. Rotating the matrix so that models are on the y-axis might also make the figure easier to read.
- Page 10 line 10. In the sentence “We also notice the good performance of ACE2 across all the output variables”, the performance of ta-200 and ta-850 appears somewhat weaker; the authors may wish to nuance this sentence slightly.
- Page 10 line 11. As "Root Mean Square Error" has already been introduced, the acronym "RMSE" can be used directly here.
- Figure 6. Could you please clarify the precise definition of "NAO skill"? It would also be helpful to discuss why the values simulated by both DL-ESMs over the North Pacific are notably higher.
- Figure 11. Adjusting the y-axis limits to zoom in on the active data range (as there are currently no data above 1 or below 10-3) would enhance visual clarity.
- Figure 13. Consider using insets or re-scaling the plot area to reduce unused white space and highlight the data.
Citation: https://doi.org/10.5194/egusphere-2026-3295-RC3
Viewed
Since the preprint corresponding to this journal article was posted outside of Copernicus Publications, the preprint-related metrics are limited to HTML views.
| HTML | XML | Total | BibTeX | EndNote | |
|---|---|---|---|---|---|
| 42 | 0 | 1 | 43 | 0 | 0 |
- HTML: 42
- PDF: 0
- XML: 1
- Total: 43
- BibTeX: 0
- EndNote: 0
Viewed (geographical distribution)
Since the preprint corresponding to this journal article was posted outside of Copernicus Publications, the preprint-related metrics are limited to HTML views.
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
This study adopts the PCMDI Metrics Package to evaluate the forecasting performance of four deep learning-based prediction models for weather systems at various spatial scales, with intercomparisons against conventional CMIP models. The authors confirm that the AI-based models possess promising predictive skill yet exhibit distinct systematic error features, which they attribute to the limited time span of the training dataset. Although the findings can offer references for the evaluation research of artificial intelligence forecasting models, the analysis lacks sufficient depth. Most discussions merely describe intuitive features shown in the figures without linking them to the inherent disparities among different AI architectures. At minimum, the authors ought to discuss potential causes underlying performance gaps across multiple versions of the NeuralGCM model. Such insufficient and cursory analysis weakens the overall value of this work, and further in-depth elaboration on this part is strongly recommended.
Specific comments:
“These efforts offer important precedents, but they do not fully address the need for a comprehensive evaluation of DL-ESMs using the same diagnostics routinely applied to CMIP-class Earth System models.”
The authors are advised to clearly respond to their own research question: which extra evaluation metrics are supplemented to enable thorough evaluation of the DL-ESM model.
PCMDI: Do not assume that all readers are familiar with this abbreviation; its full name should be provided.
“We generated ensemble members by varying … For the stochastic NeuralGCM-evap and NeuralGCM-precip configurations, we additionally introduced perturbations through different random seeds.”
What are the underlying reasons for these different ensemble generation approaches? Please provide a detailed explanation of their respective purposes.
Table 2: What is the rationale for selecting the variables in Table 2? This selection serves as one of the key bases for justifying the claimed comprehensive evaluation.
“It also introduces a modest generalization test, because the boundary-condition source differs from the ERA5 training environment.”
While this generalization test is more suitable for physical models that strive for physical realism, AI models essentially deliver statistical predictions constrained by existing datasets. The bias correction schemes natively integrated inside the AI model may be incompatible with the boundary conditions set in this experiment. Hence, it remains debatable whether the evaluation metrics can truly characterize the real predictive and simulating performance of the AI model.
Figure 1: Both panels (c) and (d) are titled GPCP3.2. What are the specific differences between them? Do they correspond to different averaging periods?
Figure 4: Why are the results from the other two NeuralGCM model versions not presented?
Figure 6: A more thorough analysis is recommended to interpret the advantages of NeuralGCM, particularly whether the incorporated physical module governs the winter NAO pattern.