the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
Soil-to-stream export of dissolved organic carbon and its functional composition can be predicted using only widely available data
Abstract. Dissolved organic carbon (DOC) exports from soil to aquatic systems are significant components of the global carbon cycle and are of concern to the water industry owing to the role of DOC in the formation of disinfection byproducts. Modelling the processing of DOC as it flows from soil to the ocean requires distinction between compounds that are relatively more aromatic, coloured, ultraviolet-sorbent and photolabile, but resistant to microbial breakdown (T1), from less photo-reactive but more microbially labile compounds (T2). We assessed whether mean annual DOC concentration and its T1 fraction (pT1) can be predicted from data that are widely available, without in-situ measurements, to enable applications within Earth System Models. Using only spatially resolved data that are available at global scale, we estimated DOC concentration and pT1 for a range of headwater catchments in the United Kingdom. The best fitted models predicted DOC concentration (Nash-Sutcliffe efficiency, NSE = 0.71) and pT1 (NSE = 0.57) from land use (arable and wetland), soil type (peat or mineral), precipitation chemistry (ionic strength) and areal runoff. Application at national scale demonstrated expected spatial patterns in the distribution of DOC concentration and pT1, with high concentrations in peatlands in the north of the country, and highest pT1 in northern and western areas with high peat cover and surface runoff generation. The approach described here could be applied to other regions, transferring understanding of land-to-water DOC export from data-rich to data-poor areas.
- Preprint
(1152 KB) - Metadata XML
- BibTeX
- EndNote
Status: open (until 18 Oct 2026)
- RC1: 'Comment on egusphere-2026-3786', Anonymous Referee #1, 03 Sep 2026 reply
Viewed
| HTML | XML | Total | BibTeX | EndNote | |
|---|---|---|---|---|---|
| 213 | 86 | 24 | 323 | 20 | 20 |
- HTML: 213
- PDF: 86
- XML: 24
- Total: 323
- BibTeX: 20
- EndNote: 20
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
Review of egusphere-2026-3786
Title: Soil-to-stream export of dissolved organic carbon and its functional composition can be predicted using only widely available data, by Sawicka et al.
Sawicka et al. investigate whether dissolved organic carbon (DOC) concentration and “functional composition” in headwater streams can be predicted from widely available spatial data. They compiled data from 150 catchments across the UK, with areas up to 100 km2, and related annual median DOC concentration and SUVA254 to catchment characteristics including land-cover proportions, climate, runoff, and atmospheric deposition and precipitation chemistry using multiple linear regression. The selected model for DOC concentration identified peat cover and annual runoff as the strongest predictors, whereas SUVA254 was primarily related to peat and arable land cover. The authors report NS efficiencies of 0.71 for DOC concentration and 0.57 for the SUVA254-derived “pT1”, which represents the proportion of the UV-absorbing, photolabile T1 fraction within the UniDOM framework. They subsequently applied the fitted relationships across UK at 1-km resolution and argue that the approach could provide spatially distributed estimates of terrestrial DOC inputs and composition for land-to-ocean and Earth System Models, particularly in regions where direct observations are unavailable.
Unfortunately, my overall assessment is that the manuscript does not currently meet the standards of quality and scientific rigour required for publication. Below I provide a detailed account of the main reasons for this recommendation. The list is not intended to be exhaustive, nor to suggest that every individual issue would by itself justify rejection. Rather, I include these points to substantiate my overall assessment as clearly and transparently as possible, as the concerns are numerous and, in several cases, affect the conceptual framing, methodological robustness, and strength of the conclusions.
[01] The novelty and knowledge gap are not convincingly established. A substantial literature, some of which is cited in the manuscript itself, has already related catchment characteristics such as peat extent, land cover, hydrology and climate to DOC concentrations and DOM composition in temperate and boreal systems. The main results here, i.e., that peat cover is strongly associated with higher DOC concentrations and higher SUVA254, therefore largely confirm relationships that are already well established from my point of view. If the novelty lies instead somewhere else, this needs to be demonstrated much more clearly. At present, I find the scientific advance rather limited.
[02] There is a fundamental mismatch between the stated objective and what is actually modelled. The manuscript repeatedly refers to predicting soil-to-stream DOC “fluxes” or “export”, including in the context of improving land-to-ocean carbon budgets and Earth System Models. However, the response variable is annual median DOC concentration. Including runoff as an explanatory variable does not turn concentration into a mass flux. This distinction is important, particularly because the principal motivation of the study is framed around quantifying lateral carbon transport.
[03] In relation to that, the claims of broad or global transferability go beyond what is demonstrated by the analysis and the evidence presented. The models are developed entirely from UK data and, as far as I understand, are not independently tested outside the calibration dataset. Applying the fitted relationships spatially across UK demonstrates extrapolation of the models, but does not demonstrate that they will perform in data-poor regions elsewhere, where relationships among soil types, climate, hydrology, and DOM may differ substantially.
[04] The national-scale extrapolation itself is not sufficiently described and I can see potential issues. The statistical models are fitted using catchment-scale predictors such as the proportion of peat or arable land within each contributing catchment, yet the Methods only state that the models were subsequently applied on a 1x1 km grid. It is unclear to me whether predictor values for each grid cell represent the cell itself or the entire upstream contributing area. If the former, the model is being applied at a different spatial support from that at which it was developed. In either case, the procedure used to generate Figure 5 would need to be described more thoroughly for the analysis to be evaluated with more confidence.
[05] I am also not convinced that the extensive UniDOM/T1-T2 framing adds what the manuscript implies it does. The DOC-quality variable actually measured and modelled is SUVA254, a widely used proxy for DOM aromaticity. pT1 is then calculated deterministically from predicted SUVA254 using an existing empirical relationship. The added value of reframing the SUVA254 analysis in terms of pT1 therefore needs much stronger justification. This also contributes to considerable inconsistency throughout the manuscript regarding whether SUVA254 or pT1 is the response being modelled.
[06] DOC and DOM are also used inconsistently. The Introduction moves between DOM and DOC without initially defining their relationship, despite the distinction being important here since DOC is quantified as a concentration, whereas SUVA254 is being used to characterise properties of the DOM pool. The terminology should be conceptually clear and consistent throughout.
[07] The rationale for restricting the analysis to “headwater” catchments relies too strongly on the assumption that autochthonous production and aquatic processing can be neglected. Headwater status alone does not ensure short residence times, low nutrient inputs, or negligible autochthonous DOM production. Regardless, this defence of the use of headwaters is then difficult to reconcile when the selected catchments have areas up to 100 km2. There is no universal catchment-size threshold for defining a headwater, but most readers would understand the upper end of the range outside that definition. Furthermore, some study catchments also contain substantial agricultural land and are elsewhere described as “lowland catchments” (L. 146).
[08] The temporal aggregation of the observations needs further details. The six datasets have very different sampling frequencies and record lengths, ranging from approximately one year to more than a decade and from weekly to two-monthly sampling, yet they ultimately provide one annual-median response for each site. It is not clear how multi-year records were reduced to these values, how unequal sampling frequencies and years were handled, or how the corresponding annual predictor variables were matched to them. This is important because the analysis is interpreted primarily as spatial despite combining observations collected over markedly different periods.
[09] The statistical model-development procedure also needs to be described in more detail. It is unclear exactly which variables from Table 2 entered the candidate model set, which transformations were applied to which variables, or what the starting models were. Given the relatively large set of correlated candidate predictors, these choices are important for evaluating model stability and interpretation.
[10] More importantly, I cannot identify an independent or cross-validated assessment of “predictive performance”. NSE appears to have been calculated using the same 150 sites used for predictor selection and model fitting. In that case, the reported NSE values represent apparent goodness of fit rather than predictive performance. This is an issue for a manuscript which main ambition is prediction and transferability to unobserved regions. Figure 4 also shows some structured differences among datasets, while no meaningful residual diagnostics or assessment of model assumptions are presented.
[11] Several sampling sites also appear to belong to nested catchments (?). If so, the observations are not statistically independent because nested sites share parts of their contributing area and environmental characteristics. This should be explicitly assessed and, where relevant, accounted for in the statistical analysis.
[12] The calculation of runoff is particularly surprising. The manuscript states that runoff was estimated by subtracting “estimated evaporative loss” from precipitation, but only potential evapotranspiration is described among the meteorological datasets. It is therefore unclear what quantity was actually used. Estimating actual catchment evapotranspiration, and consequently runoff, is generally very far from trivial, particularly across 150 catchments and multiple years. Because annual runoff emerges as one of the strongest predictors of DOC concentration, this methodological step needs to be fully documented and justified.
[13] I am not convinced about the interpretation of the negative wetland-moorland coefficient. If I understand it correctly, peat cover is effectively a subset of the broader wetland-moorland category and the two predictors are strongly correlated. After including both in the same regression, a negative conditional coefficient for wetland-moorland cover could simply result from the partitioning of overlapping predictors as a statistical artefact. So, interpret it as evidence that “non-peat wetlands” might reduce DOC is far-fetched unless that quantity can be defined and modelled separately.
[14] Much of the discussion of peat, agricultural land, runoff and ionic strength reiterates mechanisms that are already well established in the DOC literature. Given the limited novelty of the empirical relationships themselves, I would have rather expected more evaluation of model uncertainty, failure modes, spatial transferability and what new process understanding is obtained from this analysis.
[15] I consider the comparison with process-based models in L. 356-358 not entirely appropriate. That the relationships shown in this manuscript can be applied “without site-specific calibration” imply an advantage over mechanistic models, but the two approaches generally operate at very different temporal scales and are designed to answer different questions. From my point of view, this comparison therefore conflates model purposes rather than demonstrating an equivalent approach with reduced data requirements.
[16] The Introduction could be better structured, with some rationale repeated and the study objective introduced before the specific knowledge gap has been established. The literature used for some of the broader framing is also relatively dated and does not represent the more recent work on terrestrial-aquatic carbon transfer and CO2 emissions.
[17] The origin of the SEPA dataset would need clarification. Table A3 states that SEPA “lost all their data and documentation” during a cyberattack. Yet, SEPA contributes 87 of the 150 study sites and by far the largest number of observations. There may of course be a straightforward explanation (e.g., copies of the data may have been obtained before the attack), but this should be stated explicitly, particularly given the importance of this dataset to the analysis.
[18] Several parts of the Results report mechanistic interpretations that would be more appropriately placed in the Discussion.
[19] The appendix tables are first referred to in the manuscript in the order A3, A1, and A2. They should be renumbered or reorganised so that their first appearance follows sequential order.
[20] Finally, “United Kingdom” and “Great Britain” are used interchangeably in places. The terminology should be standardised in this regard.