the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
Physics vs. AI inWeather Prediction: Evaluating GraphCast, AIFS, and FuXi against an Observation-Corrected WRF Model for Flash Floods
Abstract. Data-driven artificial intelligence (AI) weather models are increasingly positioned as alternatives or complements to physics-based numerical weather prediction (NWP) systems, yet systematic model evaluation studies that compare both paradigms under controlled experimental conditions remain limited. Here we present a structured model evaluation framework applying standardised verification metrics to assess a regional WRF configuration against three publicly available AI models – GraphCast, AIFS Single 1.0, and FuXi – over three high-impact flood events in Luxembourg and the Greater Region (2016, 2018, and 2021), comprising 90 simulation days and over 7300 matched forecast-observation pairs. The WRF model employs a three-dimensional variational (3D-VAR) data assimilation scheme ingesting GNSS Zenith Total Delay and conventional observations on a 6-hourly Rapid Update Cycle. All systems are initialised from ERA5 reanalysis to ensure consistent initial conditions. Categorical precipitation scores at the 1 mm threshold and continuous temperature metrics at surface stations serve as verification targets. Data assimilation measurably improves WRF's categorical precipitation skill (Critical Success Index: 0.306 to 0.341) while leaving near-surface temperature largely unchanged, demonstrating the value of the assimilation scheme. AIFS achieves the highest detection rate (POD 0.765) and net categorical skill (CSI 0.370), GraphCast and AIFS reduce temperature RMSE by ~10 % relative to assimilated WRF, and WRF alone reproduces the observed mesoscale precipitation structure of the catastrophic July 2021 flood. The per-event breakdown reveals no skill degradation for AI models on out-of-sample events, suggesting that meteorological regime rather than training-data overlap governs model performance. These results provide a replicable model evaluation methodology for benchmarking emerging data-driven NWP systems against observation-corrected regional models.
- Preprint
(6650 KB) - Metadata XML
- BibTeX
- EndNote
Status: open (until 14 Sep 2026)
-
RC1: 'Comment on egusphere-2026-3655', Víctor Galván Fraile, 31 Jul 2026
reply
-
AC1: 'Reply on RC1', Haseeb Ur Rehman, 07 Aug 2026
reply
Dear Editor and Referee,
Please find attached our PDF file, which includes responses to the comments raised.
-
AC1: 'Reply on RC1', Haseeb Ur Rehman, 07 Aug 2026
reply
-
RC2: 'Comment on egusphere-2026-3655', Anonymous Referee #2, 11 Aug 2026
reply
General comments
1 The comparison between WRF and AI models is interesting, but the experimental design may introduce an inherent imbalance in the initial conditions. The WRF experiments use ERA5 as boundary and initial conditions followed by a regional data assimilation system incorporating additional observations (SYNOP, TEMP, TAMDAR, GNSS ZTD), whereas the AI models are initialized directly from ERA5 analyses.
Although the authors discuss this issue in the conclusions, its impact deserves a more quantitative assessment. In particular, the comparison between observation-corrected WRF and AI models may partly reflect differences in the quality of the initial atmospheric state rather than differences between dynamical and AI forecasting approaches. I suggest including a more explicit comparison between WRF without DA and AI models, or quantifying the contribution of improved initial conditions relative to the forecasting model itself.
2 The AI models are re-initialised from ERA5 every 6 hours rather than being run in a fully autoregressive mode. This design avoids error accumulation but does not represent the operational use case of multi-step AI forecasting.
Please clarify why this evaluation strategy was selected and discuss how the results would change under free-running forecasts. Since one of the major advantages claimed for AI models is computational efficiency during long-range forecasting, it would be valuable to evaluate both single-step forecast skill and multi-step rollout degradation.
3 The WRF configuration uses a single 12 km domain. However, flash floods are often strongly influenced by convection-permitting processes and local topography. The manuscript should discuss whether the chosen WRF resolution is sufficient to represent convective rainfall mechanisms in this region. A comparison with higher-resolution convection-permitting WRF simulations would help determine whether the observed advantage of WRF for rainfall structure is related to physical modelling itself or simply to the spatial resolution difference.
4 The manuscript correctly mentions that AIFS benefits from ECMWF operational analyses. However, this issue deserves more emphasis because AIFS is not purely comparable with GraphCast and FuXi. AIFS uses fine-tuning with IFS analyses from 2016–2022, which overlaps with all three evaluation events. Therefore, the comparison does not only represent architecture differences but also differences in training data and operational assimilation systems.
5 Table 2 provides basic information about AI models, but important details for reproducibility are missing. Please consider providing
exact model checkpoints/releases used;
inference settings;
precipitation variable definition;
any preprocessing applied to ERA5 inputs;
interpolation method from model grid to station locations.
6 Figure 9 provides qualitative comparison of rainfall fields. However, visual comparison alone is subjective. Please consider adding spatial verification metrics, such as Fractions Skill Score (FSS), object-based precipitation verification, spatial correlation...This would strengthen the conclusion that WRF better represents rainfall structures.
Citation: https://doi.org/10.5194/egusphere-2026-3655-RC2 -
AC2: 'Reply on RC2', Haseeb Ur Rehman, 20 Aug 2026
reply
Dear Referee and Editor
please find attached our response (in pdf) to remarks from referee 2.
-
AC2: 'Reply on RC2', Haseeb Ur Rehman, 20 Aug 2026
reply
Data sets
Observational rainfall data of the 2021 mid-July flood event in Belgium– Part 2. Radar product RADFLOOD21 E. Goudenhoofdt et al. https://doi.org/10.5281/zenodo.7740059
Model code and software
Python scripts for WRF vs. AI weather model evaluation over Luxembourg flood events H. u. Rehman https://doi.org/10.5281/zenodo.20794937
Video supplement
Evaluating RADAR vs Physics based NWP vs AI Models H. u. Rehman https://doi.org/10.5446/73607
Viewed
| HTML | XML | Total | BibTeX | EndNote | |
|---|---|---|---|---|---|
| 182 | 55 | 13 | 250 | 11 | 5 |
- HTML: 182
- PDF: 55
- XML: 13
- Total: 250
- BibTeX: 11
- EndNote: 5
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
The manuscript presents an evaluation of a regional physics-based numerical weather prediction model (WRF) and three global data-driven weather models (AIFS, GraphCast and FuXi) over three high-impact flash-flood events in Luxembourg and the Greater Region. The study analyses the impact of data assimilation on WRF forecast skill, particularly for precipitation, and compares the different forecasting systems against surface observations using categorical verification scores for precipitation (based on a threshold of 1 mm per 6 hours) and continuous statistics for 2-metre temperature.
The overall objective of comparing the strengths and weaknesses of physics-based and AI-based weather forecasting approaches is timely and highly relevant for guiding future developments in numerical weather prediction. I consider the manuscript suitable for publication in Geoscientific Model Development after revisions addressing the comments below.
General comments:
Specific comments:
Line 32: I suggest reframing this paragraph to motivate your analysis. In the first part, it is described one study that finds improvements from AI vs dynamical models for short-term regional forecasts, while latter is mentioned other which finds that extreme events are not so well captured by AI models. I would suggest using these different results to motivate your study, to try to clarify the strengths and weaknesses of each modelling framework by comparing them in three extreme events that took place in your research region.
Line 44: Is the choice of the WRF configuration, described as being tuned for this region, based on a previous study? If so, please provide the corresponding reference. In addition, how sensitive are the results to the selected physical parametrisations (Table 1), particularly the cumulus parametrisation?
Fig. 1: I suggest adding levels to the topography to better illustrate the reader the influence that the terrain can play in the study area.
Line 113: Why the WRF simulations use a 30-day window and not a shorter one? If there is a reason to this election, please justify it in the text.
Fig. 6: Why are the verification scores for WRF (with DA) different from those shown in Fig. 4?
Fig. 9: Could the authors include the WRF simulation without data assimilation in this figure? This would help to assess the impact of data assimilation on the spatial distribution and intensity of the predicted precipitation for this event.
Line 212: The discussion of the computational efficiency is very valuable. It would be helpful to include a quantitative comparison of the computational requirements (e.g., execution time) for the WRF and AI forecasts.