Physics vs. AI inWeather Prediction: Evaluating GraphCast, AIFS, and FuXi against an Observation-Corrected WRF Model for Flash Floods
Abstract. Data-driven artificial intelligence (AI) weather models are increasingly positioned as alternatives or complements to physics-based numerical weather prediction (NWP) systems, yet systematic model evaluation studies that compare both paradigms under controlled experimental conditions remain limited. Here we present a structured model evaluation framework applying standardised verification metrics to assess a regional WRF configuration against three publicly available AI models – GraphCast, AIFS Single 1.0, and FuXi – over three high-impact flood events in Luxembourg and the Greater Region (2016, 2018, and 2021), comprising 90 simulation days and over 7300 matched forecast-observation pairs. The WRF model employs a three-dimensional variational (3D-VAR) data assimilation scheme ingesting GNSS Zenith Total Delay and conventional observations on a 6-hourly Rapid Update Cycle. All systems are initialised from ERA5 reanalysis to ensure consistent initial conditions. Categorical precipitation scores at the 1 mm threshold and continuous temperature metrics at surface stations serve as verification targets. Data assimilation measurably improves WRF's categorical precipitation skill (Critical Success Index: 0.306 to 0.341) while leaving near-surface temperature largely unchanged, demonstrating the value of the assimilation scheme. AIFS achieves the highest detection rate (POD 0.765) and net categorical skill (CSI 0.370), GraphCast and AIFS reduce temperature RMSE by ~10 % relative to assimilated WRF, and WRF alone reproduces the observed mesoscale precipitation structure of the catastrophic July 2021 flood. The per-event breakdown reveals no skill degradation for AI models on out-of-sample events, suggesting that meteorological regime rather than training-data overlap governs model performance. These results provide a replicable model evaluation methodology for benchmarking emerging data-driven NWP systems against observation-corrected regional models.
The manuscript presents an evaluation of a regional physics-based numerical weather prediction model (WRF) and three global data-driven weather models (AIFS, GraphCast and FuXi) over three high-impact flash-flood events in Luxembourg and the Greater Region. The study analyses the impact of data assimilation on WRF forecast skill, particularly for precipitation, and compares the different forecasting systems against surface observations using categorical verification scores for precipitation (based on a threshold of 1 mm per 6 hours) and continuous statistics for 2-metre temperature.
The overall objective of comparing the strengths and weaknesses of physics-based and AI-based weather forecasting approaches is timely and highly relevant for guiding future developments in numerical weather prediction. I consider the manuscript suitable for publication in Geoscientific Model Development after revisions addressing the comments below.
General comments:
Specific comments:
Line 32: I suggest reframing this paragraph to motivate your analysis. In the first part, it is described one study that finds improvements from AI vs dynamical models for short-term regional forecasts, while latter is mentioned other which finds that extreme events are not so well captured by AI models. I would suggest using these different results to motivate your study, to try to clarify the strengths and weaknesses of each modelling framework by comparing them in three extreme events that took place in your research region.
Line 44: Is the choice of the WRF configuration, described as being tuned for this region, based on a previous study? If so, please provide the corresponding reference. In addition, how sensitive are the results to the selected physical parametrisations (Table 1), particularly the cumulus parametrisation?
Fig. 1: I suggest adding levels to the topography to better illustrate the reader the influence that the terrain can play in the study area.
Line 113: Why the WRF simulations use a 30-day window and not a shorter one? If there is a reason to this election, please justify it in the text.
Fig. 6: Why are the verification scores for WRF (with DA) different from those shown in Fig. 4?
Fig. 9: Could the authors include the WRF simulation without data assimilation in this figure? This would help to assess the impact of data assimilation on the spatial distribution and intensity of the predicted precipitation for this event.
Line 212: The discussion of the computational efficiency is very valuable. It would be helpful to include a quantitative comparison of the computational requirements (e.g., execution time) for the WRF and AI forecasts.