the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
Forecast biases of extratropical cyclones classified by their diabatic heating intensity in operational physics-based and machine learning weather prediction models
Abstract. Extratropical cyclones (ETCs) are strongly influenced by moist processes, rendering the impact of diabatic heating critical for forecast performance. We systematically evaluate short-range (12-hour) forecast biases of wintertime maritime ETCs over the North Atlantic, North Pacific, and Southern Ocean for the period 2023–2024. Employing a cyclone-centred composite framework, cyclones are categorised into strong and weak diabatic heating groups. We compare forecasts from the ECMWF high resolution operational 9-km Integrated Forecasting System (IFS) and the data-driven Artificial Intelligence Forecasting System (AIFS).
Both the higher resolution of IFS and the data-driven AIFS significantly reduce the ETC propagation biases previously identified in ERA5. However, both models exhibit more pronounced errors in cyclones with strong diabatic heating. In the physics-based IFS, while the previous severe dry and cold biases are largely improved, the model still underestimates cyclone intensity and warm sector wind speeds. Furthermore, IFS displays a distinct spiral-shaped positive bias in the 850–500 hPa temperature difference, suggesting a misrepresentation of the vertical distribution and depth of diabatic heating. AIFS generally yields similar but smaller biases in most fields compared to IFS. Despite these improvements, AIFS features a notable physical inconsistency as it demonstrates a weaker mean sea level pressure (MSLP) bias but a stronger 10 m wind bias compared to IFS. This is related to an underestimation of the near-surface ageostrophic wind speed.
- Preprint
(4185 KB) - Metadata XML
-
Supplement
(4827 KB) - BibTeX
- EndNote
Status: open (until 13 Aug 2026)
- RC1: 'Comment on egusphere-2026-3727', Anonymous Referee #1, 27 Jul 2026 reply
-
RC2: 'Comment on egusphere-2026-3727', Anonymous Referee #2, 11 Aug 2026
reply
This manuscript provides a concise overview of errors in the 12-hour forecast from the IFS and AIFS for a subset of cyclones in the extratropical storm tracks of the Northern and Southern hemisphere. They do a good job of identifying a series of variables linked to cyclone intensity and generally to diabatic heating. The results are concise and clear – though perhaps to the point of missing opportunities to expand their explanations or depth of work (which would broaden the impact of the paper). I’ve summarized the comments below as a combination of major and minor comments.
Major comments:
- Depth of referenced literature: In general, I felt like the authors narrowed their literature review/connection to prior work more than I would like. For example, there’s a large body of work that has explored the impacts of diabatic heating on forecast errors – I would suggest expanding the citations here beyond those of the author(s). In addition, I some of the issues in, for example, boundary layer representation between a physics model and AI model have been explored, even if with other models.
- Attribution to diabatic heating: Correct me if I’m wrong, but can’t you get the explicit diabatic heating terms from the IFS? I know this isn’t available for the AIFS, but some of your leaps (attributing variables to physical processes) could be shored up by showing some of these terms at least for the IFS.
- Extent of results: I understand (and appreciate) the narrow focus of this work (on 12-hour biases). However, the role of diabatic heating can be an integrated process (to an extent) on the atmosphere, and plays a distinct role in the temporal evolution of systems (eg. feedbacks on cyclone growth via baroclinic processes). It may be beyond what the authors can do, but given the brevity of the manuscript it feels like there’s an opportunity to look at what these errors look like either at longer forecast lead times, or at different steps in the cyclone evolution.
Minor comments:
- L20-21: Please add a ‘by’ between research and Yu.
- L23-25: This last sentence is worded a bit awkwardly (particularly at the end ‘…investigate the forecast performance for ETCs of the …’) – please consider reworking.
- L27: Please rework as ‘Joos and Forbes (2016) and Binder et al. (2016)’
- L41: Please remove the ‘Ben’ from your citation.
- L47-48: The sentence is unclear/has some grammar issues as written – please aim to rework.
- L49: ERA5?
- L75 and elsewhere: You often are missing using a definite article (eg. ‘the’) in front of IFS or AIFS (eg. it should the ‘the AIFS’ or ‘the IFS’).
- L82-83: You may also want to note that you appear to have employed a land mask within these domains (based of Fig. 1).
- Section 2.3: Given how new the Yu et al. 2026 methodology is, I would aim to add more detail on how the method works here. Though this may have been addressed in the other manuscript (if so, you may want to say so), it would be good to elaborate on why you focused on just the amount of diabatic heating compared to, say, anomalous diabatic heating relative to the background climatology (which would likely remove the latitudinal bifurcation of the majority of your cases). I know you get at this a bit with L92-95, but I think a bit more justification/reasoning could be applied here.
- L100-104: Understood on the biases being similar amongst the three basins. It may not be relevant, but are the impacts of these heating biases on the jet (and thus the baroclinic evolution of the system) the same? The North Atlantic jet tends to operate more as an eddy-driven jet, while the North Pacific and Southern Ocean jets tend to be more a merged eddy-driven and subtropical jet. Does this influence the interpretation of the results given the time evolution aspect of your study?
- L108-109: While I agree the warm sector is likely in this vicinity, many oceanic storms at peak intensity have also started to undergo occlusion or seclusion processes. Is it possible that this could also be a residual of the diabatic heating process from an earlier warm sector that has since been displaced? Or is this analysis based off your figures 5 and 6?
- In general, when referencing your quadrant diagrams – please keep in mind that a quadrant can only be ¼ of your domain. For example, in section 3.2 you discuss the ‘right quadrant’ – this should either be written as ‘right quadrants’ or ‘right hemisphere’.
- L130-131: Is the PGF really weakened? Yes there are pressure differences, but it’s unclear that the pressure gradient itself is reduced (for example, the blue lines are all biased ‘inside’ of the black lines, but the horizontal distance between your blue isobars and black isobars doesn’t appear to be different).
- L140-145: This is a nice summary of the issue at hand (r.e. boundary layer representation). Maybe it’s beyond the scope of the project, but can you elaborate more on the differences between the two models for the boundary layer? Are the other studies that have looked at boundary layer processes between physics models and AI models?
- L177: Have you also explored biases in vertical velocity for these event? While you are generally diagnosing diabatic heating, you need vertical motion (and condensation) to unlock the heating process. Your diagnosis here captures the moisture availability, but not the vertical velocity.
- Section 3.6: The looks more like a position/tilt bias rather than an intensity of trough bias, in particular for the IFS. Also, wouldn’t positive (negative) biases in diabatic heating result in positive (negative) biases in geopotential heights (through a variety of arguments – thickness, thermal wind, QG, PV, etc.) – maybe link into this here as well?
Citation: https://doi.org/10.5194/egusphere-2026-3727-RC2
Viewed
| HTML | XML | Total | Supplement | BibTeX | EndNote | |
|---|---|---|---|---|---|---|
| 58 | 33 | 4 | 95 | 17 | 6 | 7 |
- HTML: 58
- PDF: 33
- XML: 4
- Total: 95
- Supplement: 17
- BibTeX: 6
- EndNote: 7
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
Recommendation: Major Revisions
Yu et al. (2026) examines the short-term forecast errors for ETCs in the updated ECMWF IFS and AIFS for strong-diabatic and weak-diabatic events. Overall, the authors find smaller biases in the IFS compared to the previous version for a number of chosen metrics, and largely smaller biases for AIFS compared to IFS. Generally speaking, strong heating cases show larger biases compared to weak heating cases.
This study is well motivated by the impactful nature of ETCs and does well to build on previous work by the authors by expanding the study regions and using the most current forecast systems, including an ML model. Several aspects of the methods and results, however, need additional justification, clarification, and/or quantification before recommending publication in Weather and Climate Dynamics.
General Comments
Specific Comments
L8: Worth specifying here that “ERA5” in this context refers to the previous version of IFS used in ERA5 rather than the ERA5 reanalysis. This made more sense after reading the full paper but readers may initially be confused.
L19: Another reference to add here is Büeler and Pfhal (2017): https://doi.org/10.1175/JAS-D-17-0041.1
L24: Need to define AIFS here.
L25: Please include references here for IFS and AIFS.
L43: Can remove AIFS definition since acronym used already in L24.
L47–48: This sentence reads a bit strange. Perhaps “whether” should be “the” instead?
L49: Should “ERA” be “ERA5”?
L61–62: What is the temporal resolution for forecast output?
L65: Need to define DJF and JJA.
L67–68: Please provide additional context for choices of forecast times and variables. In particular, I’m curious why 300-hPa geopotential height is preferred over 500-hPa height and why jet-level winds aren’t considered.
L69: What other forecast verification or bias quantification metrics were considered and how do they compare to RMSE?
L80–81: If possible, please include a table of these tracking criteria in supplemental material.
L85–86: Please include the classification criteria/threshold for strong vs. weak heating here. Also, what are the sample sizes for each classification?
L87: I assume the diabatic heating is a direct output of IFS and AIFS? If so, are there any limitations to how it’s computed in the models?
L87: Is there any sensitivity to this radius selection?
L98–99: Can you clarify how large the propagation biases are in ERA5 vs. IFS and AIFS as compared to the reanalysis? Visually, the propagation biases in Fig. S1–S3 arguably aren’t drastically different across the three forecast datasets. The intensity biases, however, are substantially reduced in IFS and AIFS compared to ERA5.
L109–110: Can you quantify the magnitude of this positional shift? It’s very difficult to estimate from the figure.
L112: Should “propagation bias ERA5” be “propagation bias of ERA5”?
L113: Please quantify (for propagation speed) and clarify these improvements. For example, the average bias and RMSE for IFS in Fig. 2a is larger than for ERA5 in Fig. S1a (though the color shading indicates the intensity bias is noticeably weaker).
L115: I’m finding it difficult to directly compare Fig. S1–S3 to Fig. 2 given the difference in baseline analysis (i.e., ERA5 reanalysis vs. IFS analysis). How do the basin-specific forecast differences compare when using the IFS analysis as reference?
Fig. 2 (and others): What are the contour values for the blue contours?
Fig. 3: Can you increase the quiver density a bit and add quivers for the forecast composite?
L143–145/Fig. 4: Can you report the correlation coefficient and plot the linear regression lines (or some other relevant metric) to quantify (1) the match between forecast and analysis and (2) the difference in match strength between IFS and AIFS?
L152: Can you quantify? Narrower by how much? The difference in contour placement is difficult to estimate on the figure.
L153: TCW acronym already defined.
L177: Is this the primary explanation? How do the vertical differences in humidity compare to support this hypothesis? Could differences in trough depth also explain an upper-level cold bias?
L184–185: Could you show this case-to-case variability in the supplemental figures like MSLP in Fig. S7?
L223: Being in the gray zone, does the IFS use a convective parameterization scheme?