the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
A hybrid physics–ML framework for integrating groundwater dynamics into land surface modeling
Abstract. Three-dimensional groundwater dynamics play a critical role in regulating land–atmosphere interactions, yet resolving three-dimensional subsurface flow processes at large scales remains computationally prohibitive. Here we present a hybrid coupling framework that enables the integration of three-dimensional groundwater processes into land surface modeling at substantially reduced computational cost. The framework replaces the physics-based groundwater solver with a deep learning surrogate while preserving the original coupling interface, providing a practical pathway for incorporating groundwater dynamics into Earth system simulations. A key feature of the framework is an error-control strategy based on a free-drainage lower bound, which approximates the treatment of subsurface processes in conventional land surface models where groundwater feedback is largely neglected. The hybrid solution is considered acceptable as long as its deviation remains within this free-drainage bound, with a user-defined threshold providing additional control over acceptable error levels, enabling flexible, application-dependent control of model fidelity. The framework is demonstrated in a ~34,000 km² watershed in the Pearl River Basin, China, achieving an approximately 20× speedup while maintaining strong agreement with the physics-based reference. Over a full-year hourly simulation, the median water table depth error is within 0.5 m and the domain-averaged latent heat flux reaches a Kling–Gupta efficiency of 0.965. This study demonstrates the feasibility of hybrid surrogate–physics coupling for representing groundwater processes and provides a flexible foundation for multi-timescale, on-demand simulations, with potential for extension to diverse hydroclimatic settings and integration into Earth system modeling frameworks.
Status: final response (author comments only)
-
RC1: 'Comment on egusphere-2026-2332', Anonymous Referee #1, 17 Jun 2026
-
AC4: 'Reply on RC1', Chen Yang, 15 Jul 2026
We thank the referee for the careful reading of our manuscript and for the constructive comments. We appreciate the positive evaluation and the valuable suggestions. We will carefully address these comments and revise the manuscript accordingly. Our brief responses are provided below.
C#1: We will clarify the applicability of the proposed framework and revise the title and abstract accordingly.
C#2: We will further clarify the experimental design and revise this part if necessary.
C#3: We will include quantitative computational time statistics to better demonstrate the computational efficiency of the proposed framework.
Citation: https://doi.org/10.5194/egusphere-2026-2332-AC4 - AC5: 'Reply on RC1', Chen Yang, 03 Sep 2026
-
AC4: 'Reply on RC1', Chen Yang, 15 Jul 2026
-
CEC1: 'Comment on egusphere-2026-2332 - No compliance with the policy of the journal', Juan Antonio Añel, 21 Jun 2026
Dear authors,
Unfortunately, after checking your manuscript, it has come to our attention that it does not comply with our "Code and Data Policy".
https://www.geoscientific-model-development.net/policies/code_and_data_policy.html
In the "Code and Data Availability" section of your manuscript you do not provide the data used to train your model. The GMD review and publication process depends on reviewers and community commentators being able to access, during the discussion phase, the code and data on which a manuscript depends, and on ensuring the provenance of replicability of the published papers for years after their publication. Please, therefore, publish the mentioned data in one of the appropriate repositories and reply to this comment with the relevant information (link and a permanent identifier for it (e.g. DOI)) as soon as possible. We cannot have manuscripts under discussion that do not comply with our policy.
Later, if the Topical Editor decides to continue with the review or publication process of your manuscript and you are requested to upload a new version of it, then The 'Code and Data Availability’ section of your manuscript must also be modified to cite the new repository locations, and corresponding references added to the bibliography.I must note that if you do not fix this problem, we cannot continue with the peer-review process or accept your manuscript for publication in GMD.Juan A. AñelGeosci. Model Dev. Executive EditorCitation: https://doi.org/10.5194/egusphere-2026-2332-CEC1 -
AC1: 'Reply on CEC1', Chen Yang, 24 Jun 2026
Dear Editor,
Thank you for your reminder.
The data have been deposited in a public repository (DOI: https://doi.org/10.11888/Terre.tpdc.303513) and will become publicly accessible within 24 hours. The archive includes the data used for model training, additional data for result analysis, model checkpoints and results, training log files, a template and instructions for running the model, and figure-generation scripts that run directly with the archived data.
The README also links to the updated code archive on Zenodo (DOI: https://doi.org/10.5281/zenodo.20791904), which contains the model code and more detailed documentation. The actively maintained GitHub repository is available at: https://github.com/aureliayang/ParFlow-nn.
Best regards,
ChenCitation: https://doi.org/10.5194/egusphere-2026-2332-AC1 -
CEC2: 'Reply on AC1', Juan Antonio Añel, 25 Jun 2026
Dear authors,
Unfortunately your proposed solution does not solve the outstanding issues. The Tibetan Plateau Data Center does not comply with the requirements of the journal to deposit data, namely:
- It does not appear to have a published policy for data preservation over many years or decades (some flexibility exists over the precise length of preservation, but the policy must exist).
- It does not appear to have a published mechanism for preventing authors from unilaterally removing material. Archives must have a policy which makes removal of materials only possible in exceptional circumstances and subject to an independent curatorial decision,
- It does not appear to issue a persistent identifier such as a DOI or Handle for each precise dataset.If we have missed a published policy which does in fact address this matter satisfactorily, please post a response linking to it. If you have any questions about this issue, please post them in a reply.
Therefore, please, deposit the data in a repository we can accept, and reply here with the corresponding information.
Juan A. Añel
Geosci. Model Dev. Executive EditorCitation: https://doi.org/10.5194/egusphere-2026-2332-CEC2 -
AC2: 'Reply on CEC2', Chen Yang, 25 Jun 2026
Dear Dr Añel,
Thank you for the clarification. We will deposit the data in a repository that meets the journal’s requirements as soon as possible and will provide the repository information here promptly.
Best regards,
The authorsCitation: https://doi.org/10.5194/egusphere-2026-2332-AC2 -
AC3: 'Reply on AC2', Chen Yang, 01 Jul 2026
Dear Editor,
Thank you for your patience.
We have now deposited the data in Science Data Bank and updated the DOI accordingly.
The data have been deposited in a public repository (DOI: https://doi.org/10.57760/sciencedb.41621). The archive includes the data used for model training, additional data for result analysis, model checkpoints and results, training log files, a template and instructions for running the model, and figure-generation scripts that run directly with the archived data.
The README also links to the updated code archive on Zenodo (DOI: https://doi.org/10.5281/zenodo.20791904), which contains the model code and more detailed documentation. The actively maintained GitHub repository is available at: https://github.com/aureliayang/ParFlow-nn.
Best regards,
Chen Yang
Citation: https://doi.org/10.5194/egusphere-2026-2332-AC3
-
AC3: 'Reply on AC2', Chen Yang, 01 Jul 2026
-
AC2: 'Reply on CEC2', Chen Yang, 25 Jun 2026
-
CEC2: 'Reply on AC1', Juan Antonio Añel, 25 Jun 2026
-
AC1: 'Reply on CEC1', Chen Yang, 24 Jun 2026
-
RC2: 'Comment on egusphere-2026-2332', Anonymous Referee #2, 23 Aug 2026
General comment:
This manuscript develops a groundwater surrogate that combines static spatial information with dynamic hydrological inputs and is designed to replace ParFlow within a coupled land-groundwater modeling framework. The approach is demonstrated over a ∼34,000 km² basin in the Pearl River Basin, China, and achieves an approximately 20-fold speedup while maintaining generally good agreement with the physics-based reference model. The proposed design preserves the coupling interface with the land-surface model and therefore has potential value for efficient hybrid land-groundwater simulation.
However, the Results and Discussion section should be substantially reorganized. The authors should structure this section into clearly defined thematic subsections and provide more detailed discussion of the unresolved issues revealed by the results, particularly regarding spatial errors, long-term stability, and the advantages and limitations of the Hybrid model relative to the physics-based PF-CoLM system.
Detailed comments:
2 Methodology
2.1 Physics-based CoLM/ParFlow model
Line 118-120: The rooting-zone depth (3.43 m) and aquifer thickness (100 m) differ from those commonly used in previous ParFlow applications. Could the authors clarify the rationale for this vertical discretization and whether it was mainly designed to match the CoLM configuration?
2.2 ParFlow deep-learning surrogate model
Lines 176-177 and 234-235: The manuscript emphasizes that, unlike the original FSTR using meteorological forcing directly, the modified surrogate uses CoLM-derived root-zone net fluxes as external inputs. However, these fluxes are themselves responses to meteorological forcing processed through land-surface processes. The authors should clarify the advantage of using root-zone fluxes instead of meteorological forcing directly and how this flux-conditioned design differs from a conventional forcing-driven surrogate model. Does this modification mainly enable coupling with CoLM by preserving the ParFlow-CoLM interface, or does it also improve surrogate accuracy, stability, or physical consistency?
Lines 250-252 and 255-256: The surrogate is trained only on WY2019–2020 using 120 h temporal windows. The authors should justify the choice of these two years and clarify whether the training dataset sufficiently represents wet, dry, and extreme hydrological conditions.
Lines 250-256: The surrogate is trained exclusively using meteorological forcing and ParFlow simulations from the Pearl River Basin. It is therefore unclear whether the resulting surrogate represents a general approximation of ParFlow dynamics or is specific to the hydrological and spatial characteristics of the training basin. The authors should clarify this point.
Lines 267-268: Pressure heads are standardized using their mean and standard deviation for each vertical layer. Please discuss whether this standardization affects the representation of vertical hydraulic gradients, particularly under very dry or saturated conditions.
2.3 Additional architectural experiments
Lines 277-278 and 302-303: The purpose and added value of the temporal self-attention module are not sufficiently clear. The authors should clarify whether it is primarily intended to improve long-term dependency learning, prediction accuracy, autoregressive stability, or computational efficiency. If the improvement is limited, this section should be streamlined and greater emphasis placed on the core surrogate and coupling framework.
2.4 ParFlow surrogate-CoLM physics hybrid coupling
Lines 342-343: Given that the hybrid model is evaluated only on WY2021, please discuss whether a single evaluation year is sufficient to demonstrate long-term robustness across different hydroclimatic conditions.
Lines 358-360: Does the reported speedup apply only to the ParFlow-CoLM coupled simulation, excluding the ParFlow spin-up? If a full ParFlow spin-up is still required for application to a new basin, the overall computational benefit of the surrogate may be considerably smaller. Clarify this limitation and discuss the end-to-end computational cost.
3 Results and discussion
I do not recommend presenting the Results and Discussion primarily in the order of individual figures. In its current form, this section reads largely as a figure-by-figure description, making it difficult to identify the main scientific findings and distinguish the presentation of results from their interpretation.
The authors should reorganize this section into clearly defined subsections, with each subsection addressing a specific aspect of the results. The authors may determine the specific subsection structure according to the content. For example, Lines 373-391 could be consolidated into a subsection focusing on differences in spatial patterns between the hybrid and physics-based simulations. Such a structure would substantially improve the readability and scientific focus of this section.
The manuscript evaluates WTD, soil moisture, and latent heat flux, but runoff is not discussed. Please clarify why runoff is excluded and include an evaluation of runoff performance, given its importance in assessing the coupled hydrological system.
Figure S2 and Lines 347-348: The hybrid model is trained five times using different random initializations. Clarify what the labels in Fig. S2a represent and what specifically differs among these five realizations.
Line 356: The checkpoint with the lowest validation RMSE was selected for subsequent evaluation. Identify the selected checkpoint and report the corresponding random initialization and training epoch/iteration. Please also clarify how sensitive the reported evaluation is to the checkpoint selection.
Lines 378-380: The authors attribute the finer-scale heterogeneity to soil texture and land-surface properties. However, if the hybrid model exhibits stronger or finer-scale spatial variability than the physics-based simulation, the authors should clarify why this occurs and whether these additional spatial features are physically meaningful or artifacts introduced by the surrogate.
Line 385: Large errors appear to be concentrated in the southeastern part of the domain. The authors should provide a more detailed discussion of the possible causes of this spatial error pattern.
The Results and Discussion should focus more clearly on the performance gain of the hybrid model relative to the original PF-CoLM system. The central contribution of this manuscript is the development of a surrogate model to replace or accelerate ParFlow within the coupled framework. Therefore, the key comparisons should be hybrid vs. PF-CoLM and hybrid vs. observations/reference datasets.
Lines 412-419 and 443-449: Considerable emphasis is currently placed on comparisons between PF-CoLM and the PFFD/no-lateral-flow configuration. While PFFD can serve as a useful secondary benchmark for illustrating the effects of simplified groundwater treatment, these comparisons do not directly demonstrate the added value of the hybrid model and should not dominate the main Results and Discussion.
I therefore suggest substantially reducing the discussion of PFFD vs. PF-CoLM and retaining PFFD only as a secondary reference case.
Figure S4d: I do not recommend including the Spearman-correlation comparison between the free-drainage and lateral-flow ParFlow configurations. Since both are physics-based ParFlow configurations representing different subsurface treatments, this comparison provides limited information for evaluating the hybrid model.
Lines 421-423 and 438-440, Figure 7: The authors should provide a quantitative discussion of the latent heat flux errors shown in Figs. 7b and 7c, particularly for the comparison between hybrid model and ERA5-Land. The key performance metrics and their values shown in Fig. 7 should also be reported explicitly in the text to allow readers to assess model performance without repeatedly referring to the figure.
Lines 450-452: Although the hybrid model provides substantial computational speedup, its improvement in simulation performance appears relatively modest. For example, the physics-based model already shows a high KGE against ERA5-Land for latent heat flux, with only a small improvement from hybrid. The authors should clearly demonstrate and discuss the specific advantages of hybrid over the physics-based model beyond computational efficiency.
Figure 9: The river-network-related spatial patterns remain clearly visible in the physics-based simulation but are much weaker in the hybrid results, including the bias and correlation maps. Could this indicate a deficiency of the hybrid model in reproducing the river-network features captured by the physics-based model? The authors should explicitly discuss this issue and explain its possible causes.
Lines 481-485 and Figure 8: Figure 8 appears to provide limited additional information, as the upward propagation of subsurface-state errors into surface fluxes has already been demonstrated and discussed earlier. The authors should either clarify the distinct contribution of this figure or consider removing/streamlining it to avoid repetition.
Lines 510-511: Large hybrid-Physics discrepancies in the southeastern region are mentioned repeatedly, but the manuscript does not provide a detailed explanation for this persistent spatial error pattern. The authors should discuss the possible causes of the larger errors in this region.
Line 525: The progressive deterioration is attributed to heavy precipitation, but the supporting evidence is unclear. Please specify whether this interpretation is based on Fig. 6 or Fig. 10 subplot a and cite the corresponding figure. A spatial precipitation map should also be provided to further examine whether regions of large hybrid errors coincide with areas of high precipitation.
Figures 5, 9, and 11: In the full-year simulation shown in Fig. 9, the hybrid model does not reproduce the river-network features in soil moisture as clearly as the physics-based model. The same behavior is also evident in Fig. 5, whereas Fig. 11 shows much clearer river-related WTD patterns during a selected short-term heavy-rainfall event. The authors should discuss the possible reasons for this difference between long-term and short-term predictions.
Figures 6 and 10: For the same period (hours 5761-6480), the median WTD bias is approximately 0.3-0.4 m in the continuous full-year simulation (Fig. 6) but remains below about 0.1 m in the segmented 720-h prediction (Fig. 10). Please explain why the same period produces such different error magnitudes under the two prediction methods.
If Fig. 9 represents a long-term simulation and Fig. 11 represents a short-term heavy-rainfall event, the results seem to suggest that the hybrid model loses clear river-network features during long-term simulation but can reproduce them during short-term high-rainfall conditions. The authors should discuss whether this indicates that the hybrid model is more suitable for short-term heavy-rainfall simulations than for long-term hydrological simulations.
Lines 573-574: I agree that the hybrid model appears promising for short-term flood simulations, as it provides substantial computational speedup while reproducing river-network features and overall model performance comparable to the physics-based simulation. However, the current results do not convincingly demonstrate its suitability for long-term or decadal simulations. Long-term application requires substantial training data, and the hybrid model appears to lose clear river-network features, whereas the physics-based model maintains good performance. Under these conditions, the main demonstrated advantage of the hybrid model appears to be computational efficiency. The authors should discuss this limitation more explicitly.
Conclusion
The hybrid model is developed and evaluated only for the Pearl River Basin. The authors should clarify whether the trained surrogate model can be directly applied to other basins, or whether new ParFlow simulations, training data, and retraining would be required for each new domain. The manuscript should also distinguish between the general applicability of the proposed framework and the transferability of the trained surrogate model.
Citation: https://doi.org/10.5194/egusphere-2026-2332-RC2 - AC6: 'Reply on RC2', Chen Yang, 03 Sep 2026
Viewed
Since the preprint corresponding to this journal article was posted outside of Copernicus Publications, the preprint-related metrics are limited to HTML views.
| HTML | XML | Total | BibTeX | EndNote | |
|---|---|---|---|---|---|
| 248 | 0 | 1 | 249 | 0 | 0 |
- HTML: 248
- PDF: 0
- XML: 1
- Total: 249
- BibTeX: 0
- EndNote: 0
Viewed (geographical distribution)
Since the preprint corresponding to this journal article was posted outside of Copernicus Publications, the preprint-related metrics are limited to HTML views.
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
In the paper entitled "A hybrid physics–ML framework for integrating groundwater dynamics into land surface modeling", the authors used a deep learning model to replace a physics-based solver in the ParFlow model. Overall this paper is well written. The structure is clear and description of the model is complete. However, there are still some issues to be clarified before the paper can be considered for a future publication in GMD.
1. In my understanding, because the training of the DL model is performed using the inputs and outputs of the simulations for the Pearl River Basin, China, the updated hybrid model is probably moslty suitable for this specific region. The authors may need to clarify it in the title and the abstract of the paper. Moreover, the name of the model ParFlow may also need to appear in the title and the abstract as well.
2. L251, in my mind, the rest of WY2021 should be used for model testing instead of the full year, because the beginning of the year has been used for validating the model.
3. L360-361, the authors should use some numbers such as the computing time for one-month simulation, to indicate the speedup influence of the replacement more clearly.