the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
Advancing identification of drivers of groundwater head change using commonly available observed hydroclimate data
Abstract. Groundwater depletion in arid and semi-arid regions is a pressing global challenge, driven by intensive extraction for irrigation and compounded by climate variability. However, distinguishing between the impacts of anthropogenic pumping and climate variability on groundwater dynamics remains difficult due to the lagged response of groundwater levels and the scarcity of long-term abstraction records. This uncertainty limits effective groundwater management. Previous studies classify climate-influenced sites using time-series models based primarily on calibration fit rather than predictive performance. Here, we advance groundwater driver attribution by explicitly predicting groundwater heads 2 to 8 years ahead in addition to calibration fit, using commonly available climate data and observed groundwater heads. We undertake this by assessing calibration and predictive performance across 92 wells in North Gujarat, western India, using the HydroSight time-series model. By integrating probabilistic forecasting metrics, particularly the Continuous Ranked Probability Score (CRPS), with traditional calibration measures – Coefficient of Efficiency (CoE), we identified climate-dominated wells with greater spatio-temporal consistency. CRPS effectively differentiated groundwater wells influenced by climate from those affected by pumping, revealing distinct regional patterns in groundwater dynamics. We identified 37–51 % of wells as climate-dominated across multiple prediction periods, primarily in eastern districts, validated by field observations and regional assessments. This study advances groundwater assessment methods by demonstrating limitations of conventional calibration-based approaches and advocating predictive skill evaluation. Our data-driven classification uniquely relies on observed groundwater head data to assess climate influence at finer resolution, offering insights at a scale not previously explored. These findings support improved groundwater management strategies and guide sustainable use policies in data-scarce regions.
- Preprint
(2802 KB) - Metadata XML
-
Supplement
(1146 KB) - BibTeX
- EndNote
Status: final response (author comments only)
- RC1: 'Comment on egusphere-2026-3346', Anonymous Referee #1, 30 Jul 2026
-
RC2: 'Comment on egusphere-2026-3346', Anonymous Referee #2, 07 Aug 2026
Summary
The authors present an analysis of regional well hydrographs that attempts to identify the degree of climate influence on groundwater heads over time. The technique employed involves the application of a time series model that incorporates a simplified soil and ET model that translates land surface conditions and meteorology into a lagged recharge. This time series model can be calibrated to observed well hydrographs. The quality of the calibration over varying portions of the observed time series is used as a proxy for climate influence under the assumption that the better the calibration, the more directly the climatic inputs are driving changes in groundwater levels.
General comments
The topic addressed in the manuscript is an important one - regions of groundwater decline unfortunately tend to suffer a lack of instrumentation and observational record that would be necessary to identify the role of groundwater extraction versus other hydroclimatic influence. The testing of different calibration and evaluation periods and their impact on the ability to discern climate influence is useful and especially relevant given this paucity of groundwater data. The figures presented and organization of the manuscript provide the reader with a sensible explanation of methods and results. I also appreciate the commitment to transparency and reproducibility with the provided data and code. However, while the structure and intent of the analysis is overall reasonable, I think some additional justification and explanation should be added to bolster the claims and interpretability of the results.
One key issue that I would like to see addressed is the core assumption that goodness-of-fit measures derived from fitting an incomplete model to well hydrographs provide a reliable indicator that something is missing from the model (and that that missing component can be specifically identified). The concept seems broadly plausible, but a demonstration of the idea in action, ideally in the study domain, would improve the confidence we can put in the results. This might include showing an example where a model with known groundwater pumping influence is fit with and without the pumping as part of the model and showing how the degree of mismatch/misfit in the "without pumping" version correctly identifies the missing component - something like Figure 1 from Fan et al 2023, cited in the manuscript.
Another general item that would improve the overall flow and interpretation of the work would be to include a more comprehensive conceptual model of the system at the outset of the manuscript (in the background and setting/methods description). As a reader, it is helpful to know what is known about key stresses in the study area, the accepted understanding of hydrogeology and groundwater use (are there particular aquifers that are used more than others, for example) and to get a sense for why the proposed method should be expected to be useful. The current manuscript includes some very helpful context in terms of which wells and areas are known to be influenced by groundwater pumping (starting at line 718). Similarly, information about land use change, source and prevalence of surface water irrigation in certain areas, and land cover change all seem like important parts of a conceptual model of the region that would be part of the application of the model, rather than post hoc comparisons. Following from the first general comment, this might include example well hydrographs that show what well that have typical or expected climate influence or mixed influences would look like.
Specific comments
Line 45: "..determine whether a decline is caused by overexploitation or climate conditions.." I would suggest framing this less as an absolute either/or distinction and more as a question of understanding the combination of factors. This could be tied into a conceptual model description as well.
Line 130: The geological description is good (and description of aquifers at line 140), but it would be helpful to connect this to the setting for groundwater use and observation. Are most wells in a particular unit (Quaternary/alluvial sediments? are some wells deeper in certain areas?). This would be a good place to extend the conceptual model description.
Lines 136-139: The statements here indicate that agriculture is dominant but that surface water supported irrigation is rare, making groundwater irrigation the key source of irrigation for the region. It is a bit difficult to reconcile this with the explanations provided in the discussion section stating that certain wells are known to exist near dams, canals, and areas of surface water irrigation. Land use change is also invoked as an explanation in the discussion section but the background info states that there has been little change (L137-138: "..land use and land cover (LULC) changes have been minimal.."). I'd suggest reconciling these statements, perhaps through a more complete conceptual model/description of the region. It is certainly reasonable to expect that certain areas may have characteristics that diverge from the broader region - it would be helpful for that to be explicitly stated if that is the case.
Lines 149-150: "..of which they estimate 95% of the groundwater is extracted." Does this mean that 95% of the water is extracted for domestic, industrial, and irrigation purposes? The wording makes this a bit unclear. Suggest revising for clarity.
Line 167: These additional methods for estimating groundwater pumping may not be directly applicable to local conditions, but do they provide any useful information that can bound or provide context for the analysis? Does this information suggest a priori that certain wells in certain regions should be expected to be more/less influenced by groundwater pumping? Some additional explanation as to why this information is not at all useful (if that's the case) would be helpful here.
Line 213: Do you expect the resolution of the precip and temperature data used here to affect the fit of the models and the interpretation of the results? In other words, is precipitation (especially events that would be expected to support recharge) spatially distributed in a way that would not be well represented in a 0.25 degree gridded product? A brief comment on this would be helpful.
Line 297: Explain more about the objective function used in the calibration versus the CoE measures used in the analysis. Does the measure driving the calibration differ from the measures used for interpretation? Are they the same?
Line 331: Some additional explanation for why a CoE of 0.6 is justified in capturing climate fluctuations is needed here, especially since Fan et al 2023 used 0.8. What is different in this study? What about the groundwater dynamics makes supports the claim that CoE=0.6 reflects climate is the dominant driver of variation?
Line 380: "drives" should be "drivers", perhaps?
Page 14: Description of CRPS - where does the simulated CDF come from? It is not entirely clear from what source the predicted distribution at each time step is derived. Similarly, what is the source of the error bands on the time series plots? Some additional explanation is needed.
Section 3.2 (page 17): The results suggest that the ET-constrained approach is more appropriate. The inclusion of the ET-unconstrained option does not add substantially to the main goals of the manuscript, so I suggest removing this component from the main document and moving it to the supplementary information - it is sufficient to state in the main document that this ET-unconstrained approach was tested and found not to be hydrologically plausible. This will streamline the manuscript and improve overall flow and effectiveness.
Line 512: "This highlights a key limitation of relying solely on calibration performance: a good model fit does not always mean the underlying processes are realistic." This statement casts a bit of doubt on the proposed method and analysis. What about the approach being described, beyond the ET constraints, ensures that the fitted models are reliably reflecting hydroclimatic factors? A little more discussion on this point, in light of the highlighted sentence, would be appreciated and would serve to support of the broader claims of the analysis.
Figure 6: Maybe it is just my color perception, but the color scheme for the dots on the maps are difficult to discern. I'd suggest exploring options for other color schemes.
Figure 7: This figure was a bit blurry/pixelated on my PDF, so detailed inspection was a bit of a challenge. I'd encourage a higher resolution or larger size on revision.
Figure 9: I'd suggest that the subplot axis labels are renamed something more immediately interpretable.
Results discussion - One bit of context that would help with interpretation is a statement describing the extent to which any climate/meteorological shifts occurred during the evaluation periods. Was hydrology/climate essentially stationary over the entire analysis period?
Discussion starting at Line 585 and Figure 9: Following from my previous comment, are there any hydroclimatic changes (drought, especially wet years) that should also be considered in the comparison of calibration and evaluation periods? If not, it would be helpful to state that. If there's a chance some more extreme wet/dry conditions could influence the evaluation it would be worth explaining.
Line 640: The discussion of a canal system in a part of the domain is key information that would be suitable as background (and a conceptual model) provided at the beginning of the paper.
Figure 10. There is lots of interesting information on these plots, but the very small symbol size makes it difficult to discern. I'd suggest making the symbols larger and perhaps using a more distinct marker shape for the p>0.05 and p<0.05.
Line 725: Statements here regarding surface water irrigation conflict with the background and motivation information provided at the beginning of the manuscript (i.e that groundwater is the dominant irrigation method, implying it is the key controlling factor of concern).
Citation: https://doi.org/10.5194/egusphere-2026-3346-RC2
Viewed
| HTML | XML | Total | Supplement | BibTeX | EndNote | |
|---|---|---|---|---|---|---|
| 46 | 15 | 8 | 69 | 9 | 4 | 6 |
- HTML: 46
- PDF: 15
- XML: 8
- Total: 69
- Supplement: 9
- BibTeX: 4
- EndNote: 6
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
This is a valuable and well-written manuscript addressing an important practical problem: distinguishing climate-driven groundwater-head variability from changes caused by pumping or other anthropogenic influences when abstraction or irrigation records are unavailable, which is often the case in many regions globally. The emphasis on independent predictive performance, rather than calibration fit alone, is meaningful in my opinion, although this is not new. The comparison of alternative CoE formulations with a normalized CRPS is also potentially useful, and the North Gujarat case study application is relevant for groundwater management in data-scarce regions.
Nevertheless, the central attribution claim is currently stronger than the analysis can fully support. The method demonstrates whether a climate-only model can reproduce groundwater heads during selected calibration and evaluation periods. This is not necessarily equivalent to demonstrating that a well is “climate dominated.” Good predictive performance may occur despite compensating model errors, correlated pumping and rainfall, stable pumping, or groundwater dynamics inherited from conditions before the evaluation period. Conversely, poor performance does not uniquely diagnose pumping or other artificial impacts: it may arise from meteorological-input error, observation error, model structural inadequacy, aquifer heterogeneity, land-use change, artificial recharge, or parameter uncertainty and so on. There are multiple reasons why a model is not able to reproduce the observations. The missing drivers are one (certainly important) piece but not the full story. The manuscript recognizes some of these issues, but they are not propagated systematically into the classification.
More detail comments
The paper currently appears to classify a well as climate-dominated as a model driven only by precipitation and PET performs adequately in calibration and prediction. This is an operational definition, but it is not yet stated precisely enough. Several different interpretations are possible. I) Climate explains most of the temporal variance in groundwater head. Ii) Climate is the largest individual driver iii) Anthropogenic effects are negligible iv) A climate-only model is predictively adequate. V) No detectable non-climatic driver emerges during evaluation.
These statements are not equivalent. For example, a well may be influenced by substantial but nearly constant pumping, while climate still controls most temporal fluctuations. Such a well could be predicted well by a climate model after its average pumping effect is absorbed into an intercept or slowly varying state. Similarly, pumping demand may be correlated with rainfall deficits, meaning that the climate model can indirectly reproduce part of a pumping signal. I believe using these more defined terms and clustering of the results may help to make the work more defensible than done in the current version.
This is the most important addition required in my opinion, and I am fully aware that this may create (significant) extra work, but I believe it is needed to show that the workflow really helps identify the different drivers and responses. The manuscript introduces or adapts several metrics and argues that CRPS is more reliable, but its ability to identify known drivers is not demonstrated under controlled conditions. Figure 2 illustrates bias and variance conceptually, but it does not test attribution when the true generating processes are known.
A synthetic experiment would allow the authors to quantify false-positive and false-negative classification rates. I recommend constructing groundwater-head series from a known climate-response model and then progressively introducing realistic complications.
Then add one source of uncertainty or non-climatic forcing at a time. For example, a scenario with perfect climate-only case. Then you could, for example, correct precipitation and PET, correct model structure, known parameters, no pumping, low observation error. This should establish the best attainable distributions of calibration CoE, prediction CoE variants and CRPS ratio. Afterward, scenario 2 comes into play. Here parameter estimation uncertainty can be incorporated, e.g., estimate parameters from finite calibration records, vary calibration length and observation frequency, repeat calibration from multiple initial populations or random seeds. This would determine whether wells can be misclassified merely because parameters are weakly identified. Subsequently, scenarios regarding meteorological-input uncertainty or model structural error, or pumping effects, land-use, and recharge change could be considered (this is just some ideas; likely this list can be extended).
I believe all of these scenarios are essential because the proposed method is intended to distinguish climate from pumping. Constant or rainfall-correlated pumping may be especially difficult to identify and should be acknowledged explicitly. This more systematic experiment would give us a transparent answer to the central question: under what conditions do CoE, modified CoE and CRPS correctly identify the known driver, and under what conditions do they fail?
The final classification uses calibration CoE ≥ 0.2 and CRPS ratio ≤ 0.8, whereas an earlier calibration-only assessment uses CoE ≥ 0.6. The manuscript explains the motivation qualitatively, but the threshold remains mostly unsupported. A relatively small change in either cutoff may materially alter the reported percentage of climate-dominated wells.
The paper presents CRPS as a central innovation, yet it is not sufficiently transparent how the predictive ensemble or cumulative distribution was generated. The authors should specify: I) whether uncertainty arises from parameter uncertainty, residual uncertainty, state uncertainty, forcing uncertainty, or a combination; ii) how many predictive samples were generated at each time step; iii) whether parameter posterior distributions were used or only an approximate error model around one optimum; iv) whether temporal correlation in forecast errors is preserved, etc.
CRPS can reward an overly broad predictive distribution because wider intervals may contain more observations, although excessive spread is penalized. Therefore, the manuscript should also assess forecast calibration and sharpness separately. Useful diagnostics include for example, empirical coverage of 50%, 80% and 95% prediction intervals; average interval width; comparison with a persistence or climatological forecast.
The analysis largely proceeds as if the selected model is correctly calibrated and its climatic inputs are known. Yet the proposed attribution depends on several interacting uncertainties. This should at least be discussed (less discussion is needed when the synthetic example with the different scenarios would certainly be used). Moreover, it is acknowledged that gridded meteorological uncertainty exists, but it is concluded that its influence is modest because affected wells would tend to perform poorly. That reasoning is too strong. Input and structural errors can be absorbed by calibrated parameters, particularly in these kinds of models. Consequently, both false negatives and false positives are possible.
Using the final 2, 4 and 8 years as evaluation periods is helpful, but these periods are nested and not independent. The 2-year evaluation is contained in the 4-year period, which is contained in the 8-year period. Therefore, agreement across horizons does not represent three independent validations.
The quadrant interpretation tends to associate poor evaluation performance with an emerging external driver, especially pumping. This is plausible but not diagnostic. Other explanations include changing recharge caused by land cover or irrigation return flow; changes in river or canal leakage; local abstraction at neighboring wells; errors in meteorological forcing; nonstationary aquifer response;
Rejecting ET-unconstrained models because of implausible recharge ratios is an important strength. However, plausibility is judged mainly using a recharge-to-precipitation threshold of about 20%. Recharge can be spatially and temporally variable, and local focused recharge, irrigation return flow or losing streams may exceed a diffuse-recharge benchmark. Perhaps some clarification would help here.
Four model structures are calibrated, and subsequent analyses focus on ET-constrained models. However, it is not fully clear how the final model for each is selected or whether the classifications from the one- and two-layer models agree. I believe that, in addition, a model ensemble would provide a more robust basis for classification.
The manuscript concludes that CoE-based prediction metrics can be misleading and that CRPS provides greater spatial and temporal consistency. Spatial consistency, however, is not necessarily evidence of correctness; spatially correlated forcing errors, geology or model bias can also produce coherent patterns. Thus, as proposed above already, metrics should be compared against known truth in the proposed synthetic experiment and against simple predictive benchmarks in the regional application. A climate-driven model should outperform these alternatives before climatic attribution is asserted. In particular, persistence may be difficult to beat in strongly autocorrelated groundwater series. Good absolute CRPS is less informative if a simple head-only model performs equally well.
Additional points
Selection among the 92 wells
Only 92 of 493 available wells were retained. The paper should examine whether the selected wells are spatially, hydrogeologically or operationally representative. Wells with longer and cleaner records may disproportionately occur in particular districts or aquifer settings. This could influence the apparent east–west pattern. Furthermore, the analysis assumes that observations represent comparable unconfined-aquifer responses. However, screen depth, well depth, changes in construction and hydraulic connection can strongly affect groundwater dynamics. If possible, you should provide available metadata and discuss possible mixing or at least discuss this as an additional uncertainty.
Serial correlation and effective sample size
CoE, CRPS averages and threshold comparisons are based on highly autocorrelated observations. The effective number of independent observations may be much smaller than the nominal number. Confidence intervals should account for temporal dependence, for example through block bootstrap procedures, or did I miss something here?
CRPS normalization
Normalizing mean CRPS by observed-head standard deviation permits cross-well comparison, but the denominator can be unstable for wells with low variability or short records right?
Minor comments
1.Use one spelling consistently: “modeling” or “modelling.”
2. The text sometimes uses “prediction,” “evaluation,” and “validation” interchangeably (at least in myopinion). Define each term and use it consistently.
3. Improve the legibility of maps and hydrographs. Well numbers and map symbols are difficult to read.
4. Include uncertainty or sample size in figure captions, not only threshold categories.
5. The conclusion should explicitly state that unsuccessful climate-only prediction cannot by itself identify the missing driver as pumping.