the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
A data assimilation scheme to improve groundwater state estimation in the Aqui-FR modelling platform
Abstract. Groundwater is a key resource for human activities, and anticipating its evolutions months in advance is a major challenge for stakeholders. Hydrological model for subsurface flows can be used for groundwater level forecasts. However, due to uncertainties in the model's forcings and parameters, forecast initial state estimation may be inaccurate. We propose the implementation of a sequential data assimilation (DA) scheme within the Aqui-FR modelling platform, aiming at improving groundwater state estimation over a regional scale for future seasonal forecasting system.
We assimilated in situ groundwater level observations into a regional hydrogeological model, using a Localized Ensemble Kalman Filter (LEnKF). Two localization methods are assessed to evaluate the best way to propagate data assimilation increment from observation sites into the model space. A distance based method is compared to a correlation method, based on a variogram analysis.
Both method show good performances to improve groundwater head simulations, with a root-mean-square error (RMSE) reduction of 90% compared to a reference simulation without DA [open loop (OL) run]. Experiments with validation observations sites show that the correlation method lead to a more robust DA analysis, with less degradation of the simulation compared to OL run and measurements.
Hindcast experiments using reanalysis of atmospheric forcing suggest that state assimilation, in context of inertial aquifers, can help improve forecast within a six-month range. The persistence of DA correction varies within the model space domain and may be due by an initial calibration that could be improved. After a three months lead time, 75% of assimilated observation sites still show an improvement of RMSE compared to OL.
- Preprint
(9790 KB) - Metadata XML
- BibTeX
- EndNote
Status: final response (author comments only)
- RC1: 'Comment on egusphere-2026-1504', Anonymous Referee #1, 22 Jul 2026
-
RC2: 'Comment on egusphere-2026-1504', Anonymous Referee #2, 26 Aug 2026
The manuscript titled “A data assimilation scheme to improve groundwater state estimation in the Aqui-FR modelling platform” presents a method for model prediction correction using techniques adopted from optimal control theory. Specifically, the work invokes a Kalman filtering approach to adjust model predictions based on available data (hence, producing a mean behavior in accordance with data considering expected system/measured variance). While it seems that their updated models outperform standard openloop models, I do find that certain elements of the manuscript are challenging to follow.
One point of contention is that it’s unclear exactly which type of hydrological model is being used in the study (i.e. can the reader assume unsaturated soil conditions or is that neglected due to scale?). Some specificity regarding the open loop model would help the reader to understand the parameter space used for the fitting as well. This leads into further details that could help the readers to understand the process. Ground water us specified as a state variable here in L199, but more examples would help for different parameters (what are observations used for forecasting etc,). There are some notations that come out of nowhere that are unclear for readers (e.g. what is the Marthe model in L 212?). Specificity would help the reader, so a subsection on the open loop model would be really good.
The introduction of the local weights is also difficult to follow. Do the weights in eq5 integrate to 1? There’s little to no explanation on eq 7. I understand that there are two localization methods being used, but their differences remain unclear to me. Explicitly showing some differences and the form of the weight matrices would be good here too. For example, what should the reader understand about L_xy if we don’t really know what L looks like?
Coming to the results, they mostly show improved performance in most regions. There is Figure 4 f that looks odd, but that just needs some elaborating. The organization of some of the results are hard to follow. Figure 5 needs sub labels (a-l) along with more legible axis labels. I don’t know what we’re supposed to take from this. Lastly, I don’t understand why increasing the lead time increases the error in Figure 6. I see that in figure 7 it reduces the median of the RMSE monthly, but is that even significant reduction between 3 and 6 months?
I guess my final concluding statement would be for the authors to spell out the novelty a bit more explicitly. I trust that its technical and state of the art, but isn’t this being used already? Really hammer in and make it clear what is different and what we should take home from this work. There are few references to particular results in the discussion (none that I can see actually). Please specify figures or tables that can help the reader to navigate the discussion in a manner that is supported by your findings. I would recommend the manuscript be revised in accordance with my comments.
I've left some specific comments in an annotated manuscript attached here. There's nothing in there that I haven't specified here, but the specific location might be helpful.
Data sets
A data assimilation scheme to improve groundwater state estimation in the Aqui-FR modelling platform – Data and code to reproduce figures Adrien Manlay https://zenodo.org/records/18851277
Viewed
| HTML | XML | Total | BibTeX | EndNote | |
|---|---|---|---|---|---|
| 194 | 110 | 21 | 325 | 14 | 14 |
- HTML: 194
- PDF: 110
- XML: 21
- Total: 325
- BibTeX: 14
- EndNote: 14
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
This is generally a well-written and well-organized manuscript. It treats a worthwhile topic which remains of scientific interest – data assimilation of groundwater levels into distributed hydrological models, with a perspective on seasonal groundwater level forecasting. After reviewing the manuscript, I have a few major comments that I hope will improve the manuscript and its relevance to the potential audience.
Improve the readability of the introduction, especially wrt. separation of multivariate DA and DA in integrated modelsyou’re your discussion of different scales of previous studies around line 85 to 105 vs. the mention of the high computational burden in line 89 and 98. A better separation of the different aspects you want to point out in literature would be helpful. Also, it could lead more directly to your own study, which could mean to cut back on the discussion of multivariate DA. Moreover, I suggest to formulate your own research objectives (from line 101) more clearly. More specific comments below.
Performance improvements of DA: E.g. in Figure 3, Table 2, and the surrounding text: Performance improvement at assimilated sites is fine to show. However, quite expected and not that relevant in applications. I suggest to include performance at the validation wells (and adapt the text in section 3.1 accordingly). Similarly, in Figure 5 / section 3.2: For a sound evaluation of the DA improvements, I think that e.g. a cross validation through all sets of 20%/80% of observations used in assimilation/validation (or validation/assimilation) with their NIC score could be great. Results could e.g. be shown as a histogram of NIC values. In general, the text should focus more on the validation performance – this is the really exciting part of groundwater data assimilation in such a distributed model – can it improve it beyond the actual observation?
Section 4.1: As the test of the localization method is one of your main objectives: I feel it is worthwhile exploring different radii, i.e. vary the 12km. As you state yourself, this value is set based on a variographic analysis, but I asumme it is uncertain whether this is the optimal value. Similar for the threshold of 0.7 in the variogram method.
Section 3.2 / 4.3: One very relevant question in forecasting is whether the modelling system is able to beat the naïve forecasts, namely climatology and persistence. I.e. is the model (with or without DA) better at forecasting the different lead times than a simple climatology (of the observations) or assuming persistence (of the observations)?
There remain quite a few typos and grammatical errors in the manuscript. Below, I only state the first five I stumbled across; I advise a thorough proof-reading.
Besides this, a few more minor comments: