A data assimilation scheme to improve groundwater state estimation in the Aqui-FR modelling platform
Abstract. Groundwater is a key resource for human activities, and anticipating its evolutions months in advance is a major challenge for stakeholders. Hydrological model for subsurface flows can be used for groundwater level forecasts. However, due to uncertainties in the model's forcings and parameters, forecast initial state estimation may be inaccurate. We propose the implementation of a sequential data assimilation (DA) scheme within the Aqui-FR modelling platform, aiming at improving groundwater state estimation over a regional scale for future seasonal forecasting system.
We assimilated in situ groundwater level observations into a regional hydrogeological model, using a Localized Ensemble Kalman Filter (LEnKF). Two localization methods are assessed to evaluate the best way to propagate data assimilation increment from observation sites into the model space. A distance based method is compared to a correlation method, based on a variogram analysis.
Both method show good performances to improve groundwater head simulations, with a root-mean-square error (RMSE) reduction of 90% compared to a reference simulation without DA [open loop (OL) run]. Experiments with validation observations sites show that the correlation method lead to a more robust DA analysis, with less degradation of the simulation compared to OL run and measurements.
Hindcast experiments using reanalysis of atmospheric forcing suggest that state assimilation, in context of inertial aquifers, can help improve forecast within a six-month range. The persistence of DA correction varies within the model space domain and may be due by an initial calibration that could be improved. After a three months lead time, 75% of assimilated observation sites still show an improvement of RMSE compared to OL.
This is generally a well-written and well-organized manuscript. It treats a worthwhile topic which remains of scientific interest – data assimilation of groundwater levels into distributed hydrological models, with a perspective on seasonal groundwater level forecasting. After reviewing the manuscript, I have a few major comments that I hope will improve the manuscript and its relevance to the potential audience.
Improve the readability of the introduction, especially wrt. separation of multivariate DA and DA in integrated modelsyou’re your discussion of different scales of previous studies around line 85 to 105 vs. the mention of the high computational burden in line 89 and 98. A better separation of the different aspects you want to point out in literature would be helpful. Also, it could lead more directly to your own study, which could mean to cut back on the discussion of multivariate DA. Moreover, I suggest to formulate your own research objectives (from line 101) more clearly. More specific comments below.
Performance improvements of DA: E.g. in Figure 3, Table 2, and the surrounding text: Performance improvement at assimilated sites is fine to show. However, quite expected and not that relevant in applications. I suggest to include performance at the validation wells (and adapt the text in section 3.1 accordingly). Similarly, in Figure 5 / section 3.2: For a sound evaluation of the DA improvements, I think that e.g. a cross validation through all sets of 20%/80% of observations used in assimilation/validation (or validation/assimilation) with their NIC score could be great. Results could e.g. be shown as a histogram of NIC values. In general, the text should focus more on the validation performance – this is the really exciting part of groundwater data assimilation in such a distributed model – can it improve it beyond the actual observation?
Section 4.1: As the test of the localization method is one of your main objectives: I feel it is worthwhile exploring different radii, i.e. vary the 12km. As you state yourself, this value is set based on a variographic analysis, but I asumme it is uncertain whether this is the optimal value. Similar for the threshold of 0.7 in the variogram method.
Section 3.2 / 4.3: One very relevant question in forecasting is whether the modelling system is able to beat the naïve forecasts, namely climatology and persistence. I.e. is the model (with or without DA) better at forecasting the different lead times than a simple climatology (of the observations) or assuming persistence (of the observations)?
There remain quite a few typos and grammatical errors in the manuscript. Below, I only state the first five I stumbled across; I advise a thorough proof-reading.
Besides this, a few more minor comments: