the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
Groundwater-Level Response of Annual CO2 Budgets from a Dutch Eddy Covariance Network: Comparison with European Temperate Peatlands across Land Use and Sites
Abstract. Peatlands in the Netherlands hold significant cultural, agricultural, and ecological value but also contribute substantially to national greenhouse gas (GHG) emissions as a result of drainage. This necessitates urgent mitigation, leading to the implementation of various land and water management strategies. As part of the Dutch national GHG research program (NOBV), eddy covariance (EC) measurements of carbon dioxide (CO2) fluxes were conducted at 20 sites, encompassing managed peat meadows, wet nature areas, and paludiculture systems over periods of 1 to 3 years. A novel mobile EC set-up was utilized at some locations to enhance spatial coverage.
Net Ecosystem Exchange (NEE) of CO2 demonstrated pronounced seasonal dynamics across different land-use types. Wet and highly productive systems exhibited both elevated uptake and emissions, leading to intervals of net CO2 uptake. Managed pastures were typically net emitters on a daily timescale, except during early summer when assimilation peaked. Although annual net ecosystem carbon balance (NECB) estimates were subject to uncertainty due to data gaps and adjustments for harvest, grazing, and manure application, systematic differences between land-use types remained evident. To assess the sensitivity of CO2 emissions to groundwater dynamics, ecological response functions (ERFs) were employed to relate NECB to groundwater depth, and these relationships were compared with findings from the literature. Across all sites, NECB exhibited only a weak association with mean annual groundwater level when referenced to the soil surface (ERF slope approximately 0.5 t CO2 ha-1 yr-1 cm-1, R2 = -0.08), and an even weaker association with mean summer groundwater level. When groundwater depth was referenced to the clay layer, the inferred ERF sensitivity increased, resulting in steeper slopes of about 0.8 t CO2 ha-1 yr-1 cm-1 (R2 = 0.26), although the overall explanatory power remained limited. These ERF-based sensitivities align with previously published ERFs, including studies that report non-linear groundwater–CO2 responses and threshold behavior.
Groundwater management interventions produced mixed effects on NECB. Of the pasture-oriented measures, only the Active Water Infiltration System (AWIS) resulted in a detectable shift in NECB beyond background interannual variability, while other interventions were not distinguishable from within-land-use variability. By contrast, clearer categorical patterns emerged across land-use and soil classes, and paludiculture systems showed lower net CO2 emissions than drained managed grasslands, and sites with a clay cap consistently emitted less CO2 than pure peat under comparable management. Overall, these results indicate that although mean groundwater metrics have limited predictive power at the site-year level, CO2 emissions at landscape scales are primarily governed by land use and the presence of a clay-layer, which defines the baseline against which water-management measures operate.
- Preprint
(7219 KB) - Metadata XML
- BibTeX
- EndNote
Status: final response (author comments only)
-
RC1: 'Comment on egusphere-2026-3151', Anonymous Referee #1, 19 Aug 2026
-
AC2: 'Reply on RC1', Laurent Bataille, 26 Sep 2026
We thank Referee #1 for a thorough and critical review. The comments have helped us state the paper’s central argument more precisely and to correct several genuine issues. Before providing point-by-point replies, we summarise the main changes, address the four cross-cutting concerns, and disclose one data correction.
Erratum (Ilperveld, ILP_PT): During the discussion phase, we identified that an incorrect input data file had been used for the Ilperveld site. We will correct the ILP_PT fluxes and all derived quantities (including its seasonal NEE in Fig. 6, its points in Figs 7 and 9, its annual budgets in Table C1, and the fitted slopes and R² values) in the revised manuscript. Ilperveld represents 2 of approximately 41 site-years. We anticipate that the ERF slopes, their confidence intervals, and all qualitative conclusions will remain essentially unchanged whether the site is corrected or excluded, and we will confirm this with a leave-one-site-out analysis in the revised manuscript.
Cross-cutting responses
(1) The central finding addresses predictive power rather than the importance of water level itself. Our previous wording conflated these points, and we will clarify the distinction throughout. Water level is a primary mechanistic control on peat oxidation. However, our data indicate that mean-annual groundwater level referenced to the soil surface is a weak statistical predictor of annual NECB across a heterogeneous set of managed sites. CO₂ release is determined by the amount of aerated peat organic matter, and surface-referenced water level is an incomplete proxy for this quantity. Evans et al. (2021) reported a strong relationship (R² = 0.90 for 16 UK sites, 0.68 across 65 pooled studies) using “effective water table depth,” defined as the depth of organic matter exposed to aerobic decomposition. Similarly, Aben et al. (2024), in the same Dutch network, found that annual exposed carbon explained more variance in NECB than water-table depth. Our clay-referenced depth is the Dutch stratigraphic analogue of Evans’ effective water-table depth, and referencing to it improves our relationship, as predicted by this framework. The R² values are out-of-sample (cross-validated). The surface-water-level value of approximately −0.08 is negative because that regression predicts annual NECB slightly worse out of sample than the sample mean; this reflects weak predictive power rather than a computational artefact. Clay-referencing increases the value to approximately 0.26. The within-site interannual value (R² ≈ 0.37, Section 3.6) is a fit statistic rather than a cross-validated one, and we will report all three metrics on a consistent basis to enable direct comparison. We will make this reconciliation explicit in the Abstract, Results, Discussion, and Conclusions.
(2) Data and methods availability. The unit of analysis is the annual carbon budget, and every annual budget underlying each figure is fully tabulated (Appendix C, Table C1). Table C1 does not yet include the annual-mean surface and clay-referenced water levels per site-year; we will add both columns to enable reproduction of the ERF fits directly from the paper. Eddy-covariance measurement and processing follow standard, cited procedures (Section 2.2.1), and the gap-filling approach, including its two-component uncertainty, is described in detail (Section 2.3, Fig. 5, Appendix B). Every step of the analysis is documented. The manuscript does not include the raw half-hourly flux dataset or the detailed benchmarking of the machine-learning gap-filler; these are addressed in two companion papers in preparation (a data paper for the half-hourly NOBV flux dataset and a methods paper for the gap-filler), and the half-hourly data will be released there. The two sentences referencing these “in preparation” works caused confusion. We will remove them from the Introduction, provide precise references in the Code and Data Availability section.
(3) Site comparability: lateral fluxes and grazing. Lateral (dissolved organic) carbon export is expected to constitute a minor fraction of NECB in drained systems, although it increases with drainage and disturbance (Evans et al., 2021; Rosset et al., 2022). It is significant primarily where the profile is saturated and water flows through the site, which is why we estimated it only at fully inundated sites. At the remaining sites, we cannot quantify it with the available data. The choice is therefore between setting it to zero, as we do, and introducing a poorly constrained term of ambiguous sign uniformly, which would add noise without reducing bias. We will state this rationale, flag the constructed wetlands (ONL, CAM) as higher-uncertainty, and address the term in the uncertainty discussion rather than as a bias of known sign. We will also report the NECB of the inundated sites with the lateral term set to zero, ensuring comparability across all sites. Grazing is treated as approximately carbon-neutral at the parcel scale over annual and multi-annual budgets. Carbon consumed by cattle within the footprint is either respired there, and thus already included in the measured NEE, or returned as excreta. Only the fraction exported in milk and meat leaves the parcel unmeasured. We will provide its order of magnitude and specify which sites were grazed in Appendix F. This approach follows the 2006 IPCC Guidelines, which assume zero annual net CO₂ from livestock because the carbon fixed by plants returns as respired CO₂ (IPCC, 2006, Vol. 4, Ch. 10, Sect. 10.1). We have added this citation to Appendix F and revised the main text for consistency: the Conclusions no longer list grazing pressure and stocking rates among the management pathways linking water-table regulation to NECB, and instead name only the terms that enter the budget (mowing frequency, fertilization intensity, and manure application).
(4) Site classification. Two labels mislead, and we will correct classification throughout (details under L280).
Specific comments
[L15] Summer groundwater level not shown. We examined the summer-mean relationship, and it is weaker than the annual-mean one: summer-mean surface water level is an even poorer predictor of annual NECB than the annual mean, which is itself weak. We will add it (slope, R²) as a supplementary panel referenced in Section 3.2. In summer both gross primary production and ecosystem respiration are large and largely covary, so they mask the water-level signal in the net balance more strongly than at the annual scale. The weaker summer relationship is consistent with our central point.
[L34] Restoration and policy context. We wish to clarify the scope of this study. Our research focuses on degraded, actively farmed peatlands, where the objective is to reduce emissions while maintaining agricultural production. Full restoration, defined as returning drained peat to a prior (semi-)natural bog or fen, removes land from agriculture and is essentially not implemented at scale in the Dutch agricultural peat landscape. Even the wettest systems in our dataset do not constitute restoration in this sense: paludiculture represents a newly established wet agrosystem (a productive crop grown under raised water levels), and the constructed wetlands (e.g., Onlanden) are newly created ecosystems. Neither represents a return to a pre-existing natural state. The land use, water management, and history of each site are described in Appendix A (Table A1); none of the sites was designed or is managed as an ecological restoration project. Restoration is therefore outside the land-use scope of this study. Dutch national strategy similarly focuses on raising water levels within continued agricultural use, paludiculture, and clay amendment, rather than on ecological restoration. We will state this scope explicitly to ensure that the focus on managed peatlands is understood as a deliberate boundary.
[L42] “Groundwater management plays a crucial role.” We will revise this to “is generally assumed to play a crucial role,” include an appropriate citation, and specify that our analysis evaluates the predictive strength of this assumption at the annual, site-year scale (see cross-cutting response 1).
[L50] Non-linearity sentence unclear. We will revise this passage. Non-linear (Gompertz) ERFs represent a saturating response: each additional centimetre of drawdown below the shallow zone results in a smaller incremental CO₂ release, as oxygen supply to deeper peat is limited by diffusion. We previously expressed this ambiguously and will clarify the direction explicitly.
[Fig. 1] “Effective linear slope.” This is the slope of a straight line fitted through the Gompertz-predicted NECB values across the stated groundwater interval (not an arithmetic mean of segment slopes). We agree the term is unclear and will replace it with “interval-mean slope”, define it in the caption, and state the interval bounds.
[Fig. 3] Fluxes and pools. We will simplify the diagram and its caption: correct the missing bracket, define all flux terms (including F_manure, F_harvest, F_grazing, F_lat), and clarify that NEE integrates GPP, autotrophic and heterotrophic respiration, and peat oxidation across the entire aerated profile, including the root zone. We will replace “deeper degraded pools” with “older, more decomposed peat” and acknowledge that the distinction between peat and root-litter is gradational.
[L90] Lateral fluxes and grazing. See cross-cutting response 3.
[L96] No-accumulation assumption. We agree and have added the citation. The Tier 1 approach of the 2006 IPCC Guidelines assumes no change in biomass or dead organic matter in grassland remaining grassland, with growth balanced by grazing and decomposition (IPCC, 2006, Vol. 4, Ch. 6, Sect. 6.2.1); the Wetlands Supplement (IPCC, 2014, Ch. 2) leaves this unchanged for drained organic soils, and the Dutch national inventory applies the same assumption (van Baren et al., 2025). After the sentence at L96 we now add: “This no-accumulation assumption is the Tier 1 steady-state convention for grassland biomass and litter (IPCC, 2006), also applied in the Dutch national inventory (van Baren et al., 2025).” The assumption stays restricted to pastures, since the paludiculture and wetland sites are transitions that do accumulate organic material (L105), and the following sentence on interannual litter variability (L98) remains as the residual uncertainty around it.
[L101] Manure and carbon. Correct. We will revise “might imply the addition of carbon” to “is an addition of carbon,” specifying that it is in a chemically distinct, more labile form than peat or humus, and retain the point regarding turnover time.
[L132] Underlying data not shown. See cross-cutting response 2. We tabulate the annual budgets in full (Table C1), describe the processing and gap-filling with propagated uncertainty, and we will further strengthen the Code and Data Availability section.
[L140 / L165 / L200] Mobile EC and ML gap-filling. The mobile setup alters the sampling design (one standard EC tower rotating between paired sites), but not the measurement itself, which remains standard EC. This design follows from a constraint of the NOBV programme: the sites are operational commercial farms, so we cannot impose a rigid experimental design for mowing, grazing, fertilisation, or water management. The network must accommodate farmers’ practices, which means covering more sites than there are EC systems and monitoring both members of a control–treatment pair under similar weather and management conditions. Rotating one tower between a pair sharing a continuous meteorological station was the practical solution. The resulting multi-week gaps are addressed by the gap-filling framework described here. Cross-validation is blocked in three-week groups matching the rotation interval (leave-one-group-out), so validation reflects the actual gap structure, and uncertainty is quantified and propagated to annual totals. We will make this explicit and consistently use the term “leave-one-group-out cross-validation (LOGO-CV)” (correcting instances currently reading “LOOCV/LOO”).
[L249] Dipwell truncation. We screened the raw half-hourly water-level record of all 41 site-years for a plateau at the sensor floor. This affected a single site, only during the exceptional 2022 drought: at the Aldeboarn reference field (ALB_RF), the water table reached the dipwell installation depth (about 1.2 m) on 1 September 2022, and the record was truncated until 7 September (7 days). That is 1.9% of that site-year’s water-level record and 0.05% of the network-wide record. The annual mean of ALB_RF 2022 shifts by 0.4 cm if the truncated days are assigned a level 20 cm below the floor, and by 0.8 cm at 40 cm below, so by less than 1 cm for any plausible value. Its influence on the annual means and on the ERF is therefore negligible, and we will state these numbers in the revised Methods. We did not fill the affected period because groundwater level below the installation depth is unobserved: unlike flux gaps, there is no target signal to train or validate against, so imputing it would introduce unquantifiable error. Truncation of this kind compresses the deep end of the water-level range and would weaken the fitted relationship, so the direction of the referee’s concern is right. At 7 days on one site-year, however, its size is far too small to matter.
[L280; Fig. 6 caption; Fig. 7] Classification. We agree. ILP_PT is in early-stage establishment, and its vegetation resembles extensive grassland rather than Sphagnum paludiculture. We will relabel it “semi-natural / extensive grassland (early-stage Sphagnum establishment)”, show it as its own category rather than under biomass-productive paludiculture, and correct the Fig. 6 caption sentence. The “wet pasture” at −50 cm is Demmerik, a semi-natural pasture with an uncontrolled water level; we will relabel it and clarify that “wet” refers to management and vegetation, not to a shallow annual mean water table. We will audit all site labels for consistency. The ILP_PT data correction noted above also bears on this comment.
[Fig. 6] Clay cover and reduced oxidation. We will test this interpretation with the harvest data: if pastures on thick clay show harvest and biomass comparable to those on thin clay at similar drainage, a reduced-productivity explanation is unlikely and the lower emissions point to reduced peat oxidation (Nijman et al., 2024). We will add this comparison and soften the causal wording either way.
[Fig. 7] Net uptake / near-neutral at deeply drained sites. There is no contradiction with process understanding once “deeply drained” is read as surface water level rather than exposed peat. Among the dairy pastures, the site-years with the deepest annual-mean water levels yet near-neutral or net-sink balances lie under a thick marine clay cap: BUW (65 cm clay) at −52 and −41 cm, BUO (45 cm clay) at −27 and −38 cm. Referenced to the base of the clay, these levels sit within a few centimeters of the peat (about −4 to −8 cm), so almost no peat organic matter is exposed to oxidation, and these sites are not deeply drained peat. The sink has two causes, and we will attribute it to both: low exposure of peat under the cap, and management. BUO and BUW are extensively managed pastures with no biomass export recorded in the field team’s management reports, on which the harvest and grazing terms of all sites are based, whereas HOH, with the same cap thickness and a similar clay-referenced level, is mown intensively (13 to 16 t CO₂ ha⁻¹ yr⁻¹ exported) and is a net source of 7 to 9 t CO₂ ha⁻¹ yr⁻¹. The contrast between these sites is the combined effect of clay cover and management intensity, which is the paper’s central finding. The cap alone does not explain it. This reading agrees with Paul et al. (2024), who found that a 40 cm mineral cover on drained agricultural peat in the Swiss Rhine Valley did not reduce soil carbon loss while the water table remained below the cover, and concluded that coverage mitigates CO₂ only in combination with a raised water table. At our thick-cap sites the water table sits at the base of the cap, which is the condition under which a cover matters and the quantity the clay-referenced depth measures. Pure-peat, deeply drained sites (the maize fields) remain large sources, as the literature predicts. To our knowledge, this behaviour does not appear in Evans et al. (2021) or Tiemeyer et al. (2020) because neither compilation separates peat pastures under a thick marine clay cap, a class that covers a substantial share of the Dutch coastal peat area (Aben et al., 2024). The difference is in site composition, not measurement or data quality. We will make this explicit, show the clay-capped pastures as a separate symbol class in Fig. 7, and add Paul et al. (2024) to the Discussion. We will also note that some apparent “net uptake” is at the NEE level (productive grassland taking up CO₂ that then leaves as harvest, accounted in NECB), that several are short records, and that the near-neutral points overlap zero within uncertainty. We do not claim that deeply drained peat is a stable sink.
[L294] Non-linear function tested? Yes. We fitted Gompertz and sigmoid forms to the dataset during the analysis, and neither improved on the linear fit, so we chose the linear form for parsimony and for comparability with the RMA-based studies in Table 2. We did not show these tests in the manuscript and will add them in the appendices. A refit on the current site-year table, before the Ilperveld correction, gives the same result: relative to the linear least-squares fit, the Gompertz form raises both the AIC and the BIC for the surface-referenced and the clay-referenced relationship. The weak predictive power therefore reflects the aggregation of sites with different water-level-to-exposure relationships, not the choice of functional form.
[L311] Section 3.4 too short. We will expand Section 3.4 into a proper results narrative on land-use and management effects, drawing in relevant material currently in the Discussion.
[L316] Section 3.5 (control–treatment). We agree the section is too short to be informative. We will not drop the material, because the paired contrasts explain part of the scatter in our dataset: much of the between-site variability in Fig. 7 arises from sites deliberately paired to differ in one management factor (water infiltration, fertilization, harvest regime). We will merge Section 3.5 into the expanded Section 3.4, where each pair is described with its treatment and where its main finding, that management interventions often shifted NECB via biomass export or organic-matter import rather than via NEE, has the necessary context. We will retain the underlying figure (relabelled to show NECB components by site) and add the missing treatment definitions (AWIS, PWIS, FI, HALN, BAU, BAS, and each paired contrast) to the Methods.
[L333] Saturation claims. We agree that the interannual analysis (Fig. 9) spans a narrower range than the cross-site analysis (Fig. 7). Fig. 7 is a spatial (cross-site) relationship and spans the full hydrological range (≈ 80 cm) because it compares sites with different baselines. Fig. 9 is a temporal (within-site, interannual) relationship and, by construction, spans only the year-to-year change at a given site (≈ 25 cm). We will label the two as spatial and temporal sensitivity, clarify that the temporal analysis describes the response to hydrological change at a site and not the shape of the response across the full range, and remove any saturation-range inference the interannual span cannot support. The classification relabelling above removes the apparent conflict between observed water levels and land-use classes.
[L351] Uncertainty in the slopes; are they distinguishable? We will report bootstrap confidence intervals on our pooled and per-land-use slopes and state, for each published slope in Table 2, whether it falls inside them. We will also state the fitting method (reduced major axis) explicitly in the Methods and report ordinary least-squares slopes alongside, since the two estimators differ when the correlation is weak, and the published slopes were obtained with different estimators. Where our interval excludes a published slope, we will say so and discuss the likely reason (estimator, groundwater range, land-use composition) rather than claim comparability.
[Table 2; Figs 10–12] Presented in the Discussion. We will move Table 2 and Figures 10–12 into a Results subsection (“Comparison with published ERFs”) with the necessary descriptive context, and keep the Discussion for interpretation.
[L388] Per-land-use slopes. We do compute per-land-use (and per-country) slopes; they are the basis for the statement that slopes are broadly comparable while intercepts differ, but they are not shown. We will add them, with confidence intervals, as a supplementary table and reference it wherever we make that claim.
[L398] Measurement method (EC vs chamber). We agree this cannot be demonstrated and did not intend to claim it. We will reword it as a caveat: because monitoring method co-varies with country, land-use composition and site selection (Fig. 12), a methodological contribution cannot be disentangled from these factors, and we flag it only as a possibility. We will also say where a difference could arise: the comparison groups differ in the use of manual chambers, whose annual budgets rest on infrequent campaigns interpolated with light- and temperature-response models at plot scale, versus continuous EC integrating over a field-scale footprint. Manual chamber budgets could sit higher than EC budgets from comparable sites for these reasons, but in the present compilation this cannot be separated from site selection, so we will present it as an open question and make it consistent with the statement at L405.
[L412 / L427] Biomass removal. We will separate two roles. As an explicit budget term, harvest export is quantified in Eq. 1 and shown in the retained components figure, so we do derive its direct contribution to NECB. As an indirect effect, repeated harvest or mowing alters productivity and subsequent respiration over time, which is not recoverable from a single annual budget; this is what we were loosely describing, and we will label it as a hypothesised mechanism. For paludiculture specifically, we will state that its higher field-scale NECB arises because biomass carbon leaves the field, and that whether this export is a net climate cost depends on the downstream fate of the harvested biomass (long-lived products or fossil substitution versus decomposition), which lies outside our field-scale system boundary. We will not conclude that paludiculture “emits more”.
[L426] Absolute water levels. We agree and will replace qualitative drainage terms (“deeply drained”, etc.) with absolute annual-mean water-level ranges (cm) in the cross-country comparisons.
[L445] Water level as a “weak predictor.” See cross-cutting response 1. We do not dismiss water level; we distinguish its role as a first-order mechanistic control from the modest predictive power of the mean-annual surface value for annual NECB across heterogeneous managed sites. We will acknowledge the literature identifying water level as the principal driver, and clarify that what improves prediction is information on exposed peat (the clay reference, stratigraphy, and land use) in addition to water level, not instead of it. The water-level relationship also re-emerges within sites: once the fixed between-site baseline (clay, stratigraphy, land use) is differenced out in the interannual analysis, year-to-year changes in water level track changes in NECB more coherently than the pooled cross-site fit suggests. The weak pooled relationship therefore reflects between-site confounding rather than a lack of a water-level effect. Evans et al. (2021) obtained their strong relationship using exposed-peat (effective water-table) depth across a wide bog-to-cropland gradient; on the narrower, managed-only range studied here, the predictive power of a pooled surface-water-level regression is lower, even along the same relationship. We will revise L445, L363, and the Conclusions accordingly.
Citation: https://doi.org/10.5194/egusphere-2026-3151-AC2
-
AC2: 'Reply on RC1', Laurent Bataille, 26 Sep 2026
-
AC1: 'Comment on egusphere-2026-3151', Laurent Bataille, 19 Aug 2026
Erratum for one site - After submission, we noted that an incorrect data file had been used for the Ilperveld site. The analyses and results reported for this site will therefore change, and corrected results will be provided in the final version of the manuscript. Because this site represents only a very small fraction of the dataset, the general results and conclusions are not significantly impacted.
Citation: https://doi.org/10.5194/egusphere-2026-3151-AC1 -
RC2: 'Comment on egusphere-2026-3151', Inge Wiekenkamp, 06 Sep 2026
I have read the manuscript entitled “Groundwater-Level Response of Annual CO₂ Budgets from a Dutch Eddy Covariance Network: Comparison with European Temperate Peatlands across Land Use and Sites,” written by Laurent Bataille and coauthors and considered for publication in Biogeosciences. I think the topic is timely and highly relevant, and the manuscript presents an impressive dataset together with a substantial number of analyses. However, I also think that several aspects of the manuscript could be improved to strengthen its clarity, focus, and overall presentation. I have provided general and detailed comments below, together with suggestions for improvement.
General Comments:
1) Title: I think the title, “Groundwater-Level Response of Annual CO₂ Budgets from a Dutch Eddy Covariance Network: Comparison with European Temperate Peatlands across Land Use and Sites,” is quite lengthy, and it is not immediately clear what the article is about at first glance. I suggest shortening the title and making it more concise and intuitive, while still capturing the main focus of the study.
At present, the title seems to combine two different focal points: (1) the relationship between groundwater levels and CO₂ fluxes, and (2) the comparison of the Dutch data with data from temperate peatlands across Europe. Given that the manuscript focuses substantially on the relationship between groundwater levels and peatland CO₂ fluxes, the title could potentially be framed more explicitly around this central research question. A question-based or finding-based title could also help make the main objective and scientific motivation more apparent to the reader. At the same time, land use and management practices are also important themes throughout the manuscript. I therefore suggest that the authors consider which aspect (groundwater level, land use/management, or their interaction) is intended to be the primary focus of the paper and ensure that the title reflects this hierarchy.
2) Abstract: I think the general structure of the abstract is good. However, it may benefit from a stronger focus on the central research question and main take-home message. I found the key message of the ERF analysis somewhat difficult to discern, as several different findings are presented. First, the focus is on the relationship between carbon fluxes and groundwater level and the effect of adding information about the clay layer. Later, the focus shifts more towards land management. The statement that the ERF sensitivities “align with previously published ERFs” is not entirely clear to me, and it would be helpful to clarify what this comparison with previous studies demonstrates and how it relates to the main finding, so that the abstract is more self-contained.
More generally, I think the abstract would be strengthened by making the hierarchy of the results clearer, particularly the relative importance of groundwater level, land use, and the clay layer, and what these findings imply for groundwater-management strategies.3) Introduction: I like the content provided in the introduction, as it gives the reader the relevant background on peatlands and their CO₂ emissions. However, similar to the abstract, the introduction could benefit from a clearer focus. There is a substantial amount of knowledge and several relevant ideas presented, but I think the narrative would be stronger if the introduction more explicitly distinguished between (1) the current state of knowledge, (2) the specific knowledge gap, and (3) what new knowledge this study aims to provide.
4) Introduction – research questions: I think the research questions are well described, and the expected outcomes are also clearly presented. To make it easier for the reader to identify the overarching aim of the study, I suggest adding a sentence that brings the individual research questions together into a more general aim or purpose.5) Methodology: I think this section is generally clearly written. As a general suggestion, I would propose linking the methodological steps more explicitly to the overall aim of the study, so that the rationale for each analytical step is clearer throughout.
Regarding the machine-learning analysis, I feel that some methodological information is missing. First, how were the hyperparameters optimised, and which optimisation algorithm and objective function were used? Regarding the estimation of epistemic uncertainty, I also wonder why the authors chose 64 repeated model runs. Was the convergence or stability of the uncertainty estimate assessed? For example, does the estimated uncertainty stabilise after 20, 40, 64, or 100 runs? Without such an assessment or methodological justification, the choice of 64 runs is difficult to evaluate.
Additionally, the authors refer to the residual error as “aleatoric uncertainty”. Please clarify this terminology and discuss to what extent the residuals may also contain model misspecification, measurement error, and unresolved temporal or environmental variability. In other words, please clarify why the residual error is considered to represent predominantly irreducible variability rather than a combination of different sources of uncertainty.
For model performance evaluation, I would also expect one or more quantitative metrics, such as RMSE, MAE, and/or R², to be reported alongside the distributional and residual analyses. Looking only at the joint distributions and marginal densities may not sufficiently demonstrate how well the model reproduces the observed time series, particularly under specific conditions (e.g. day/night for sub-daily data, seasonal transitions, or extreme fluxes).
Finally, the use of LOOCV for model evaluation and uncertainty estimation may be problematic if individual observations are left out while temporally adjacent observations remain in the training set, given the strong temporal autocorrelation in flux data. Pointwise cross-validation may overestimate predictive performance under these circumstances, which could in turn lead to underestimated gap-filling uncertainty. I therefore suggest that the authors demonstrate that the reported model performance and uncertainty estimates are robust to temporally blocked validation, ideally using block lengths or gap structures that are representative of the actual missing-data periods.
6) Results: In the Results section, I suggest that the authors add a sentence at the beginning of each subsection explicitly linking the results to the aim of the study. This would help readers follow the overall narrative of the paper and more clearly understand why the particular results are being presented.
Also, several results are only shown in the Discussion section. I understand the intention, but if the authors intend to follow a classical Results and Discussion structure, I suggest that all results, including the corresponding figures and tables, are presented in the Results section rather than being introduced for the first time in the Discussion.
7) Water-level influence on carbon fluxes: While reading the manuscript, I was somewhat confused about the main message regarding the influence of groundwater level on carbon fluxes. In some parts of the Results, groundwater level is presented as a relatively weak predictor of annual NEE/NECB, with limited explanatory power (e.g. R² = −0.08 and R² = 0.26). In Section 3.6, however, the analysis of interannual changes in NECB and groundwater level shows a stronger relationship (R² = 0.37). Elsewhere in the manuscript, the authors describe groundwater level as showing clear and consistent relationships with carbon exchange. For example, the Conclusions state that “Groundwater level consistently emerged as a determinant of carbon exchange across sites and years.” I think the authors should clarify how these findings should be reconciled and ensure that the wording of the conclusions is consistent with the strength of the relationships demonstrated by the different analyses.
In particular, I suggest distinguishing more clearly between groundwater level being an important mechanistic or ecological driver of carbon exchange and groundwater level being a strong predictor of annual carbon balances in the present dataset. These are not necessarily equivalent: a variable can have a meaningful ecological effect or sensitivity while still explaining only a relatively modest proportion of the variation in annual NECB. If this is the interpretation intended by the authors, I suggest making this distinction explicit throughout the manuscript.
It would also be helpful to explain why changes in groundwater level appear to be more informative for explaining changes in NECB than mean groundwater level is for explaining annual NECB. Are these analyses intended to capture different aspects of the groundwater–carbon relationship? Making this distinction explicit would help the reader understand the apparently different strengths of the relationships reported throughout the manuscript.
Finally, regarding the comparison with previous literature, I think it would be useful to go beyond comparing the direction or slope of the reported relationships. Where possible, could the authors also compare the predictive or explanatory performance reported in previous studies and discuss why the groundwater–carbon relationship appears less pronounced in the present dataset? Differences in temporal or spatial scale, land management, vegetation, hydrological conditions, site selection, data processing, or the aggregation of fluxes into annual balances could potentially contribute to these differences.
8) Relationship to previous work and novelty: I think the manuscript needs to be more explicit about how it differs from, and advances beyond, previous work from the same research context. In particular, I see considerable conceptual overlap with the recent study on the groundwater–CO₂ relationship in Dutch peatlands (van der Poel et al., 2025), as well as some overlap with the study of methane emissions from Dutch peatlands (Buzacott et al., 2024).
Given the overlap in study region, measurements, authorship and, potentially, underlying datasets, I think the authors should provide a clear explanation of the specific scientific contribution of the present manuscript relative to these studies. In particular, it would be helpful to state explicitly what is novel in the present analysis—for example, in terms of research question, temporal or spatial scale, land-use comparison, analytical approach, variables considered, or interpretation of the groundwater–carbon relationship. This would also help clarify how the present study builds upon rather than duplicates previous work.Detailed Comments:
1) In line 33, the authors refer to the Dutch climate agreement policy document (“Klimaatakkoord”), which they cite as follows: Ministerie van Economische Zaken en Klimaat, 2019 - Ministerie van Economische Zaken en Klimaat: Klimaatakkoord, Government report, 2019. I suggest to adjust the citation to make it more complete and easier to find. I would add a link to the official document in English to make it easier to access and understand for the international audience: https://english.rvo.nl/sites/default/files/2020/07/National%20Climate%20Agreement%20The%20Netherlands%20-%20English.pdf. Also, what about adding information about the EU restoration law here and how this relates to rewetting etc.? This is however just a suggestion, not a “need to”.
2) Abstract, line 8: “Net Ecosystem Exchange (NEE) of CO2 demonstrated pronounced seasonal dynamics across different land-use types.” I would suggest to adjust this sentence and make the land use types the subject of the sentence. Different land-use types showed pronounced ...”
3) Line 59/60: “The causal diagram (Fig. 2) illustrates the interconnected controls of climate drivers, hydrology, and soil properties on peat oxidation in Dutch peatlands.” Is this the case for Dutch peatlands alone? Or is this working for peatlands more generally? If it is super specific for the Netherlands, would be worthwhile mentioning why this is the case. If it’s more generally applicable, I would point that out, because then a wider audience could use it.
4) Line 66: “This complexity makes it inherently difficult to formulate a general ERF,..” Is the trouble not also that one would need to have a lot of data? Considering that one needs to measure a lot of different elements to really get this interactions all accurately in an ERF?
5) Line 76-77: “As a result, ERFs inevitably smooth over much of the temporal variability and spatial heterogeneity that shape CO2 emissions, contributing to the observed non-linearity and site-specificity in groundwater–carbon relationships.” Is this really true? While ERFs based on annual or mean groundwater levels inevitably smooth temporal variability and spatial heterogeneity, it is not clear to me that this smoothing necessarily produces non-linearity.
6) Figure 3: In the caption, not all therms are explained (Fmanure etc. are not in there) I would make sure that all are explained.7) Line 133 – 136: “The eddy covariance (EC) monitoring network, including detailed documentation of data processing and harmonization, is described in a companion data paper currently in preparation. The machine-learning modelling framework is presented in a separate methodological manuscript, also in preparation.” Maybe you can add this to the Code and data availability section. I think the fact that data and code is not available makes the study less reproducible (at least during the review process and also during the period when readers do not have access to the code and data), but I hope tihs data will become available in due time.
8) Line 165 – section: Here, it would be probably good to mention if the same sensors were used for the mobile tower setup.
9) Figure 4: Although I generally like the figure, I suggest making the map somewhat larger and the station symbols slightly smaller, as this would make it easier to identify the locations of the EC towers. If colour is used to distinguish land-use types, I would also suggest using a consistent colour scheme for the same land-use categories wherever possible. At present, the large number of colours makes it difficult to obtain an overview of the spatial distribution of the different land-use types.
In addition, the current labels are not always very informative without referring back to the text. For example, labels such as “Fen meadow clay thin layer” and station codes such as “LAW_ICOS” do not immediately tell the reader which site is being referred to or what its main characteristics are. I suggest using a simple site identifier that can be directly linked to Table A, and explicitly providing this link in the figure caption. Alternatively, the stations could be numbered and the corresponding numbers linked to Table A. This would make the figure easier to interpret independently of the main text.
10) Line 299. 200: “Gaps greater than 60 days in length and gaps resulting from the mobile measuring strategy were not filled with the MDS algorithm; beyond this threshold, conditions at the time of the gap are too far removed from those of available look-up entries to yield reliable estimates, and the ML model described below is used instead”. I would adjust this sentence as follows (1) removed is probably not the right wording here, adjust this and (2) shorten the sentence or rewrite it into two sentence to make it easier to follow for the reader.
11) Line 203 – 204: “Missing meteorological data was filled with data from nearby NOBV stations and KNMI (Dutch meteorological institute) stations.” Here, it would be great if the authors could give some information on how close these stations are to the EC measurement towers.Table 1: Change “Depthof to “Depth of”.
12) Line 225 – 226: “Each dataset was first filtered for data quality and screened for outliers or inconsistencies”. Maybe the authors could add a sentence on the way this was done or refer to publications that have used the same filtering etc.
13) Figure 5: I like the graph, I think it gives a good overview of the workflow. I just have some small remarks to slightly improve the graph. In step 2 – at one arrow there is written “repeated 64 times”. It looks like the “repeated is written with a different font and it’s more tricky to read. Generally, in multiple cases the text is not really fitting in the boxes. I would propose to adjust the size of the boxes, so that it is easier to read it. In step 2, there is written “Partition in 3 weeks contiguous blocks for for LOO”. Should this not read “Partition in 3 weeks continuous blocks for LOGO”? In the text in step 2, I found several LOO, either this is an issue with the way I display the .pdf (Adobe) or the G is not printed here appropriately (or it should read LOO and I have not understood this).
14) Figure 6: I wonder whether the figure could be adjusted to make the main message of Section 3.1 easier to identify. At present, the large number of boxplots across the different 3-month periods and classes makes some of the broader patterns, such as differences in peak seasonality, relatively difficult to distinguish. One option could be to merge some of the classes in the main figure to highlight the broader patterns, while presenting the variability among the individual classes separately if this level of detail is important for the analysis. For example, one panel could show the main patterns using broader land-use categories, with an additional panel showing the individual classes. This might make the figure easier to interpret while retaining the information on differences between specific classes.
15) Line 197 – 198: “Sites with similar mean groundwater levels displayed markedly different NECB values, underscoring the influence of strong site-specific controls.” Here, I assume that the authors are referring to Figure 7. If so, please consider explicitly referring to the spread of NECB values at similar groundwater levels, for example at groundwater levels between approximately −60 and −40 cm in Fig. 7A.
16) Line 198 – 199: “Several pronounced deviations corresponded to sites underlain by a marine clay layer ... ”. Maybe explain here in an extra sentence why these deviations refer to the clay layer depth? Why this is so important here? I think this is important, because it makes it probably easier to understand for the reader without reading the paper by Nijman.17 Figure 8: I wonder whether the presentation could be adjusted to make the distinction between net carbon sources and sinks more explicit. While the figure provides information on NECB and its components, it is not immediately clear from the current presentation whether and where the peatland systems shift from a net source to a net sink of carbon, or vice versa.
This seems particularly important because such source-to-sink transitions are potentially among the most relevant outcomes of the study from a management and policy perspective. Relative changes are useful for evaluating the effects of different management practices, but the absolute magnitude and sign of NECB are also important for interpreting the actual implications of these changes. I therefore wonder whether the authors could make the zero line and/or the sign of NECB more visually explicit, or alternatively refer more clearly to another figure or analysis showing the absolute NECB values. This would help the reader interpret both the magnitude of the management effect and whether it results in a change in the overall carbon balance.
Citation: https://doi.org/10.5194/egusphere-2026-3151-RC2 -
AC3: 'Reply on RC2', Laurent Bataille, 26 Sep 2026
We thank Dr. Wiekenkamp for a constructive and detailed review. The comments have helped us clarify the paper’s central message and its structure. We respond to the general comments first, then the detailed ones.
Erratum (Ilperveld, ILP_PT): During the discussion phase, we identified that an incorrect input data file had been used for the Ilperveld site. We will correct the ILP_PT fluxes and all derived quantities (including its seasonal NEE in Fig. 6, its points in Figs 7 and 9, its annual budgets in Table C1, and the fitted slopes and R² values) in the revised manuscript. Ilperveld represents 2 of approximately 41 site-years. We anticipate that the ERF slopes, their confidence intervals, and all qualitative conclusions will remain essentially unchanged whether the site is corrected or excluded, and we will confirm this with a leave-one-site-out analysis in the revised manuscript. We thank the referee for prompting this verification.
General comments
(1) Title. We agree it is long and mixes two foci. We will shorten it to foreground the central finding and its hierarchy, for example: “Groundwater level as a predictor of annual CO₂ balances across Dutch peatlands: land management and clay cover are as important.” The European comparison will be introduced in the abstract rather than the title.
(2) Abstract. We will restructure the abstract around a clearer hierarchy and a single take-home message: mean-annual surface groundwater level is a weak predictor of annual NECB; referencing the exposed peat below the clay layer improves it; land management and the presence of a clay cap are as important as water level, because they determine how much peat is exposed for a given water level. We will replace the unclear phrase “align with previously published ERFs” with a concrete statement: our inferred slopes fall within the range of published values, and we will give bootstrap confidence intervals showing which published slopes they can and cannot be distinguished from (see our reply to Referee #1, L351). We will make the land-use types the subject of the seasonal-NEE sentence (detailed comment 2).
(3) Introduction. We will restructure the introduction to distinguish (i) the current state of knowledge, (ii) the specific knowledge gap, and (iii) the contributions of this study, and will lead with the exposed-organic-matter framing (see general comment 7).
(4) Overarching aim. We will add a sentence before the numbered research questions to state the overall aim: to determine how strongly, and through which pathways, groundwater level governs annual CO₂ balances across the full range of Dutch managed-peatland land uses, and how this compares with published ERFs.
(5) Methodology. We will add brief rationale clauses linking each methodological step to the aim. On the machine-learning analysis: this paper is not a gap-filling methodology study. The gap-filler is a tool used to produce the budgets, and its full development and benchmarking belong to a dedicated companion manuscript. We will therefore keep the following additions concise and, where appropriate, point to that manuscript.
(a) Hyperparameters were tuned by grid search within each cross-validation training fold. We will state the search space and objective function in one sentence and refer to the companion manuscript for detail.
(b) Number of repetitions. We chose 64 seeds as a pragmatic ensemble size and will add a short convergence check showing that the ensemble spread has stabilized by 64. The seeds vary only the model’s internal randomness on the same training data, whereas the cross-validation blocks are fixed and not seeded. The across-seed spread is therefore the epistemic (model) uncertainty, and the held-out residual is the residual term of (c). At the annual scale, the epistemic term dominates: residual errors largely average out under annual aggregation, whereas the spread across seeds does not, so the residual term is negligible in the annual totals. We will state how both components propagate to the annual totals, including how we treat residual autocorrelation.
(c) “Aleatoric” uncertainty. We agree the term overstates irreducibility. We will relabel it “residual (cross-validation) uncertainty” and state that it conflates measurement noise, model misspecification and unresolved environmental variability.
(d) Quantitative metrics. Appendix B already reports NSE, RMSE and bias. We will add MAE and R² and a daily-aggregated value (which removes the diurnal correlation), extending the existing panel rather than adding a new methods section.
(e) Blocked cross-validation. We agree pointwise leave-one-out would be problematic under temporal autocorrelation, which is why our cross-validation is already blocked: it is leave-one-group-out with three-week contiguous blocks matching the mobile-EC rotation, chosen so that withheld periods resemble the actual missing-data periods. Some instances in the text and Fig. 5 read “LOOCV/LOO”; we will correct these to “LOGO-CV” throughout. Blocked validation is therefore already the method used. We will describe the block structure explicitly and report the metrics under (d) on these blocked folds.
(6) Results structure. We will add an opening sentence to each results subsection linking it to the aim, and move all results currently introduced in the Discussion (Table 2 and Figures 10–12) into a Results subsection (“Comparison with published ERFs”), keeping the Discussion for interpretation.
(7) Reconciling the water-level message. We agree this needs to be made coherent, and we adopt the referee’s framing. We will distinguish groundwater level as a mechanistic driver from its predictive power for annual site-year NECB. The apparently different strengths reflect this distinction. The cross-site regression (out-of-sample R² ≈ −0.08 for surface water level, negative because the surface fit predicts annual NECB slightly worse than the sample mean out of sample, and ≈ 0.26 clay-referenced) mixes the water-level sensitivity with large between-site baseline differences (clay, stratigraphy, land use, management). The within-site interannual analysis (R² ≈ 0.37, a fit statistic, whereas the cross-site values are cross-validated; we will report all three on the same basis) removes the fixed-site baseline and isolates sensitivity to hydrological change, which is why it is more coherent. On explanatory power in previous work: Evans et al. (2021) obtained R² = 0.90 using “effective water table depth” (an exposed-organic-matter metric) at the site-mean level across a wide bog-to-cropland gradient, a range that inflates R², and Aben et al. (2024) found that annual exposed carbon explained more variance than water-table depth. The two effective-depth metrics reflect the same principle. Evans’ effective depth is capped at the peat base (min(WTD, peat depth)), so a water table below the mineral horizon underlying the peat adds no exposed peat; our clay-referenced depth starts at the base of the clay cap, so a water table within the mineral horizon overlying the peat adds no exposed peat either. Our results are consistent with both: referencing to exposed peat (our clay metric) improves the relationship, and the lower absolute R² here reflects our site-year resolution, the narrow managed-only range, EC footprint heterogeneity and the management terms, not a data-quality gap. We will state this explicitly and make the Conclusions consistent (the sentence “Groundwater level consistently emerged as a determinant …” will be reworded to the driver-versus-predictor formulation).
(8) Relationship to previous work and novelty. We will add an explicit statement of novelty. The present study is complementary to, and by some of the same authors as, the works noted. Buzacott et al. (2024) address CH₄, not the CO₂/NECB budgets analyzed here. Van der Poel et al. (2025) build a generalizable, large-scale flux–groundwater response surface at sub-annual (half-hourly) resolution from airborne and ground EC using machine learning; to that end, they use modeled water levels and do not resolve the annual carbon budget or its management terms. The present paper uses measured, site-specific water levels and site-specific gap-filling models to construct annual NECB budgets that include the management and harvest terms (Eq. 1) required for a carbon balance, introduces and tests the clay-referenced groundwater metric, contrasts cross-site with interannual sensitivity, and benchmarks the results against European ERFs. The two studies are complementary: a spatially generalizable response model versus site-resolved annual carbon accounting. We will make this distinction explicit in the Introduction and Discussion.
Detailed comments
(1) We will complete the Klimaatakkoord citation with the official English-language URL.
(2) We will make land-use types the subject: “Different land-use types showed pronounced seasonal dynamics in NEE …” (implemented in the revised abstract).
(3) Fig. 2 scope. The mechanisms shown (water level, air-filled pore space, oxygen, degradation) are general to drained temperate peatlands; the clay-deck pathway is characteristically Dutch. We will state this so a wider audience can use the diagram.
(4) L66. We agree. Formulating a general ERF is also data-limited, since resolving all interacting controls would require simultaneous, co-located measurements that are rarely available together. We will add this.
(5) L76–77. We agree that smoothing does not, by itself, produce non-linearity. We will reword this to say that averaging temporal and spatial variability contributes to the apparent scatter and site-specificity of groundwater–carbon relationships, and remove the claim that it produces non-linearity.
(6) Fig. 3 caption. We will define all terms (including F_manure) in the caption (also addressed for Referee #1).
(7) L133–136. We will remove these two “in preparation” sentences from the Introduction and give precise pointers in the Code and Data Availability section instead. The annual budgets are tabulated in full (Table C1); see our response to Referee #1 on reproducibility.
(8) L165. Yes. The mobile towers used the same sensor models as the fixed towers (Gill or Metek 3-D sonic anemometer; LI-COR LI-7500DS open-path CO₂/H₂O analyser, LI-7500RS at a few fixed sites; LI-COR LI-7700 CH₄), differing mainly in mast height and the shared-meteorology rotation design. The one exception is the permanent ICOS-type tower at Langeweide (LAW), which uses a closed-path analyser. The instrumentation is otherwise homogeneous across the network. We will state this in the Methods.
(9) Fig. 4. We will redraw the map larger with smaller station symbols, the same land-use colour scheme as the results figures (Figs 6–9) so that map and plots read together, and simple site identifiers cross-referenced to Appendix A in the caption.
(10) L199–201. We will rewrite this as two shorter sentences and replace “removed”: gaps longer than 60 days, and gaps from the mobile rotation, are not filled with MDS; beyond ~60 days the look-up conditions differ too much from the gap, so the ML model is used instead.
(11) L203–204. We will add the typical distances of the substitute stations to the EC towers. The NOBV network is dense and substitution was mostly between sites of the network rather than from KNMI stations, so these distances are generally short. We will tabulate them in the Methods.
(Table 1) We will correct “Depthof” to “Depth of.”
(12) L225–226. We will add a brief description of the filtering and outlier screening (physical-range limits, flag-based rejection, footprint filter, cross-referencing Section 2.2.1) or a reference to the harmonised NOBV procedure.
(13) Fig. 5. We will redraw the figure so the text fits the boxes, fix the “for for” typo, and use “LOGO-CV” consistently (the “G” was dropped in several places; the referee is correct that it should read LOGO).
(14) Fig. 6. We will redesign the figure into two panels, one showing the main seasonal patterns by broad land-use category and one showing the individual classes, so the broad patterns are easier to read while the detail is retained.
(15) L297–298 (the referee’s L197–198). We will refer explicitly to the spread of NECB values at similar groundwater levels (for example between approximately −60 and −40 cm in Fig. 7, left panel) to make the “site-specific controls” statement concrete.
(16) L298–299 (the referee’s L198–199). We will add a sentence explaining why the clay layer matters: a marine clay cap over peat raises the effective anoxic boundary and limits oxygen ingress into the underlying peat, reducing the exposure of peat organic matter to oxidation for a given surface water table (Nijman et al., 2024). We will also note that the organic-matter content of the clay itself varies, which adds scatter and is probably why clay-referencing improves but does not fully explain NECB.
(17) Fig. 8. We will make the source/sink distinction explicit: an emphasised zero line and clear shading or annotation of net source versus net sink, with the sign convention stated in the caption. Source-to-sink transitions are among the most policy-relevant outcomes, and we will reference the absolute NECB values (Table C1).
Citation: https://doi.org/10.5194/egusphere-2026-3151-AC3
-
AC3: 'Reply on RC2', Laurent Bataille, 26 Sep 2026
Viewed
| HTML | XML | Total | BibTeX | EndNote | |
|---|---|---|---|---|---|
| 237 | 125 | 38 | 400 | 29 | 30 |
- HTML: 237
- PDF: 125
- XML: 38
- Total: 400
- BibTeX: 29
- EndNote: 30
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
Summary: The manuscript titled “Groundwater-Level Response of Annual CO2 Budgets from a Dutch Eddy Covariance Network: Comparison with European Temperate Peatlands across Land Use and Sites” uses a large set of Eddy Covariance measurements, ground water level measurements and additional site information to relate net ecosystem CO2 exchange with hydrological conditions of peatlands in the Netherlands. The manuscript thus addresses an important topic in large scale emission modelling, land use transformation, emission reduction and agricultural policy. Herewith, the manuscript fits well to the aims and scope of EGUsphere.
The main findings of this manuscript are, that water level is only a weak predictor for peatland emissions, strongly impacted by site-specific conditions like land use and clay cover. Impacts of specific water management measures were characterized as ambiguous.
General comments:
The main concern about this manuscript is, that it is based entirely on large, but yet unpublished data sets, unpublished methods for green house gas measurements and gap filling routines, making it impossible to evaluate the quality of the finding presented. In some data sets, biases are introduced by the data acquisition and processing (dry dip wells, completely omitted incorporation of grazing into CO2 balances, although 10 sites out of 21 were dairy pastures). Sites were further misclassified by ignoring the actual state of the hydrologic conditions or vegetation cover. Some of the presented findings cannot be backed up by the data presented. Finally, the major outcome, that water level in peatland soils is only a weak predictor for its CO2 balance, is a strong claim and contradicting current process understanding in the vast universe of literature on this topic. If true, this requires a comprehensive discussion and plausible explanations, which the authors omit. Please see the detailed comments below.
Specific comments:
L 15: The correlation to summer water levels were nowhere shown in the manuscript nor discussed.
L 34: The most obvious measure, restoration, is omitted from the discussion. It would be valuable for the reader if the national context of restoration, e.g. in conjunction with the upcoming demands of the EU restoration law, would be discussed as well. It would be valuable to know in the context of this paper, if the Dutch government is not taking into account this land use for peatlands in its national strategies.
L 42: Here the authors state, that ground water management plays a crucial role in peat degradation and emissions, which they clearly contradict in their findings. Please clarify that.
L 50: “…, whereby capturing emissions caused by deeper soil drainage is the primary motivation for introducing this non-linearity.” This sentence is not clear to me, as you state in the next sentence, when using non-linear functions, emissions from deeper soil layers are smaller (since not the same amount of oxygen can reach down there as a result of diffusion through the reactive porous media, compared to upper layers). Please make that clearer.
Fig 1: It is not explained how you calculated “effective linear slope”, is it the arithmetic mean slope between the given intervals? If so, state clearly.
Fig 3: The sentence “NEE measured by eddy-covariance integrates plant-driven fluxes…” omits a closing bracket, leaving unclear which fluxes are considered plant- or peat-based, in particular short-term microbial respiration. Further, the explanation omits peat based fluxes from all upper layers, e.g. the root zone, since only peat oxidation “from deeper degraded pools” is mentioned, leaving this description unclear if not highly misleading. According to this figure, no long term peat oxidation occurs in the root zone, which is incorrect.
L 90: Lateral exports were estimated only at fully inundated sites and set elsewhere to zero. Although it is desired to use all available data at hand, this introduces an artificial bias when comparing NEC balances among sites, rendering these comparisons corrupt. Please treat lateral fluxes the same way for the purpose of site comparisons. Further, Appendix F details that the impact of grazing on NECB was fully ignored in this study, although 10 sites out of 21 were dairy pastures. This further adds to erroneous site comparisons.
L 96: The no-accumulation assumption of pastures is valid here for annual, and even more, multi-annual balances and in line with international standards, e.g. IPCC guidelines. These could be cited here to back-up this assumptions.
L 101: “The addition of manure might imply the addition of carbon to the system..” The addition of manure is always an addition of carbon to the system.
L 132: The entire data analysis is based on large, but yet unpublished data and the data (e.g. flux and water level time series) is not at all shown or discussed in this paper. Thus, the quality of the underlying data remains completely unknown. This strongly violates established scientific practice and the standards of EGUsphere, giving rise to possibly tremendous misinterpretation of the data.
L 140 / 165 / 200: The new approach of using intermittend mobile eddy towers is not an established scientific strategy for conducting tower-based flux observations, due to the resulting, extremely large data gaps. This method was not published or validated befor this manuscript, similar to the newly proposed, ML-based gap filling method. Thus, this manuscript, similar to the underlying data sets, relies on unpublished methods, questioning the quality of the outcomes.
L 249: Water levels exceeding this range (0 to 1-1.5 m) were recorded at the installation depth limit. This means, that whenever the water level fell below this value, recorded values were set to e.g. 1 m. This of course introduces a bias in the annual mean water levels (towards higher water levels). Were dip wells deepened after data retrieval discovered this incorrect installation to reduce this type of bias? Why were these times not filled with more realistic values, e.g. by some ML algorithm as the authors apparently are used to these types of methods? This artificial bias could substantially lower the correlation strenght to emissions and need to be elaborated, since this is one of the main outcomes of this manuscript, and also making non-linear behaviour more likely in fig 7, altering the NECB=0 intersection.
L 280: The authors state, that the vegetation composition of the Sphagnum site (ILP_PT) was in the early state of establishment an resembles more of an extensive grassland. Attributing the flux results to Sphagnum paludiculture is thereby wrong and incites erronous interpretation by the reader. The authors should either name the site properly reflecting its conditions, e.g. semi-natural grassland, or remove it from the analysis. The authors need to make sure to use an appropriate classification and naming as well for all other sites. See also caption of Fig 6: “Sphagnum-based paludiculture is comparable to extensive grasslands”, this a strongly misleading and false statement. See also Fig 7, which shows a wet pasture, but with an annual mean water level of -50 cm?
Fig 6: Authors attribute the reduced summer emission balances at sites with a clay cover to reduced peat oxidation. This claim is nowhere backed up by, e.g. biomass data, since a reduced biomass growth on these sites could also be a cause with the same impact on the seasonal balance, bearing the strong risk of mixing fossil peat emissions with recent biomass turn-over.
Fig 7: How do the authors explain net uptake or near to neutral emissions at deeply drained sites, since no other synthesis publications like Evans et al., (2021) or Tiemeyer et al., (2020) show this behaviour. This strongly contradicts current process understanding.
L 294: As it is also common in the field (see Table 2), was there ever a non-linear function tested against the presented data set, e.g. parabolic or compertz, to see if this dependency would yield a better representation of the data than a linear one and thus increasing the likelihood of a more realistic process understanding? Otherwise, the conclusion of weak correlation of NECB to WLEV could be a matter of choice instead of inference.
L 311: Section 3.4, set out to display outcomes with relation to Land-use and management, consists of just two sentences and seems thus, fairly incomplete. Maybe some text went missing during the writing?
L 316: Section 3.5 is set out to display comparisons between pairs of control and treatment sites. However, may be because the entire section is just two sentences long, it omits completely to explain the different treatments, e.g. what is the difference between the ALB_RF and the ALB_MS sites. I could not find it in the entire manuscript nor the appendix tables. This should be part of the Methods section, however, the site description part there is also only 4 sentences long. Conclusively, with out the necessary information the entire control-treatment comparison remains pointless.
L 333: This important claim cannot be backed up by data in this paper, since the observed annual mean water levels often violates the given land use classification (see comment to Fig 7.) and the range of water level changes in Fig 9 (~25 cm) is way to small to conclude findings on saturation effects (in the wet or dry range) given the actual range observed (~80 cm, compare Fig 7).
L 351: Have the authors taken into account the uncertainty behind all slope values given, when comparing their outcome with previous studies? If done so, is a significant distinction even possible or are the majority of slopes so close, that they can not be distinguished from one another?
Table 2 and Fig 10, 11, 12: It is unusual to present data outside the result section (like here, in the discussion). This way, there is no place to thoroughly describe these findings in detail there, and the necessary background for the reader is missing at this point. Please move Table 2 and Figure 10, 11, 12 into the results section.
L 388: Here you state, that you have fitted various slopes, apparently for different land use types, however, in the entire manuscript, there is just the fit over all sites given (Fig 7.). Were these additional fits not included into the manuscript? Why have these never mentioned before?
L 398: The authors derive the conclusion, that type of measurement (e.g. Eddy Covariance vs. chamber) is likely to contribute to observed differences in data sets. This conclusion cannot be derived, since reported EF, as they state themselves, strongly depend on various site conditions as well as the specific site selection included in each data set. Given this variability, systematic differences between measurement techniques cannot be derived. Which the authors state themselves further down in L 405.
L 412: The authors speculate, that biomass removal impacts NECB, in this case of restored and paludiculture systems. From Eq. 1 and Fig 8, this should rather be a derived outcome of this paper (the authors take biomass export into account), than a guess. Why where the authors unable to derive that conclusion based on their data?
L 426: See comment to L333. Additionally, the authors showed differences in observed water levels among the different country specific studies. Consequently, in each country, e.g. “deeply drained” refers to different water levels. To guarantee a comparability, the authors need to use absolute values instead of inconclusive terms.
L 427: See comment on L 412.
L 445: The authors present the water level as a weak predictor for NECB, omitting that they characterized it as the first-order descriptor (L 363), omitting a better predictor available for large scale modelling or mentioning the vast universe of literature identifying water level as the main driver for emissions from peatlands. This appears to be a biased opinion rather than an objective assessment.
Technical corrections: None