the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
UK-Flow15-QC: A quality control framework for better river flow data in hydrological research
Abstract. The significant increase in computing power over the past 70 years has progressively enabled the use of extensive datasets for hydrological modelling. The colossal scale of these datasets, i.e., over one million timesteps per station for a 30-year record at 15-min resolution, makes implementing effective quality control (QC) particularly challenging. In this study, we present a national-scale, open-source quality-control framework tailored for the UK’s 15-minute river flow dataset, UK-Flow15, which is described in Part 1 of this paper series. The framework combines manual visual inspection of anomalies with automated detection of statistical artefacts, incorporating both established and novel procedures. In particular, we introduce methods to evaluate high-flow events by comparing them with rainfall records and flow observations from neighbouring catchments. Application of the framework within a UK dataset reveals that while many stations maintain generally reliable records, over 20 % exhibit visually identifiable issues such as truncations, discontinuities, or missing data. Automated checks indicate that most (78 %) stations contain at least isolated segments of suspicious behaviour. Our high-flow event validation procedures confirm most peak flows, but also flag a small proportion of events as potentially spurious due to a lack of consistency with nearby flow (10.5 %) or rainfall (14.5 %) support. We further demonstrate that data quality has a measurable impact on hydrological modelling, with catchments containing flagged anomalies producing the least reliable simulations in terms of NSE and High-Flow Bias. By making flagged data and metadata openly accessible, the framework enables users to make informed decisions about data suitability. This work highlights the critical importance of rigorous QC in sub-daily hydrology and provides a scalable tool to support the development of more reliable, high-resolution hydrological data.
- Preprint
(1331 KB) - Metadata XML
- BibTeX
- EndNote
Status: final response (author comments only)
- RC1: 'Comment on egusphere-2026-277', Anonymous Referee #1, 07 Jul 2026
-
RC2: 'Comment on egusphere-2026-277', Anonymous Referee #2, 27 Jul 2026
General comment
The paper identifies a real problem with these large datasets and develops a framework to automate QC of flow dataset. They develop a set of tests and outlier detection thresholds, including tests in comparing outlier flags to climate data and neighbouring watersheds flow rates within a temporal window. However, as is this method isn't transferable enough to be published as a general framework. As a result, I recommend this paper be rejected and reconsidered after major revisions.
For instance, the selection of a 5x threshold for spike detection was based on how many stations get flagged (sensitivity analysis approach), which is not the same as detection accuracy. If we were to use a highly “polluted” dataset with many erroneous flows, this approach would break down as these results are self-referential and impact the threshold selection. Furthermore, there does not appear to be a validation step where authors identify false/positives.
Similarly, in the hydrometric area comparison thresholds, the "0.5 annual likelihood" and the "20x more likely to happen" thresholds are arbitrary. The authors do not state a statistical rationale for their choices in this section.
In general, the authors suggest that the paper provides a scalable tool for new datasets and regions, however, the thresholds are tuned to one national dataset's flagging behavior. To improve this method, I encourage the authors to separate statistical thresholds from purely heuristic thresholds, and provide justification and validation for each threshold.
The selection of statistical thresholds are mostly unjustified; however, the definition of what is considered an outlier is dependent on the end-user’s needs and objective. The statistical thresholds selected should have their purpose and drawbacks clearly stated. For flexibility, the framework, especially if/when published as a package, should offer different approaches so the user can decide which threshold suits their needs. The paper could thus offer comparison of outlier detection between different approaches/thresholds.
The heuristic thresholds need to be justified clearly and and validated robustly. The framework would be improved if these thresholds were validated either against synthetic data or a set of manually-labeled subset of timeseries. Furthermore, for broader applicability, the authors have to have a method for users to select these heuristic thresholds themselves as configurable defaults in the user-friendly package.
Line by line comments
Line 51: “Adaptation to extreme flows, however, has lagged, partially due to the scarcity of large‐scale sub‐daily flow datasets.” Is this true in the context of this paper? Certainly in some areas this could be argued, however, I question the relevancy of this claim for the UK, which does have relatively abundant high-frequency data (as the paper demonstrates). Secondly, the US has a sizable network of high frequency flow measurements, but places still lag behind in flood mitigation because of other reasons (more so political, likely). The authors’ use of “partially” does qualify the statement, however, I’m still not sure whether the lack of high frequency sensors is a significant reason why we lack adaptations. Consider adding a citation and using more precise language. See Rozemeijer et al. (2025) and related work.
Lines 59-65: Authors state that national datasets lack standardization, but then follow up a statement with two examples of standardization within the country, which is often the scope of where policy decisions are made. This argument is in contradiction with the authors’ work as the framework is developed using a national-scale dataset, and not a global dataset where concerns regarding mismatch of timestep are meaningful. I agree that accessibility is an issue and that standardization is an issue, but I think it’s deeper than standardization of timing across national dataset (which appears to be what the authors are claiming). Perhaps the lack of standardization has more to do with lack of harmonised, quality-controlled, analysis-ready large-sample sub-daily flow datasets? I believe Rozemeijer et al. (2025) also touches on these issues as well.
Lines 74-78: Authors should consider fleshing out this section some more with specifics regarding the methods as it is the foundation of the paper. There are also many papers that discuss outlier detection and related QAQC approaches for high-frequency data that should be summarized and cited here.
Line 142: Add oxford comma after “type”, for clarity.
Citation: https://doi.org/10.5194/egusphere-2026-277-RC2
Model code and software
UK-Flow15-QC F. Fileni https://github.com/felipef93/UK-Flow15-QC
Viewed
| HTML | XML | Total | BibTeX | EndNote | |
|---|---|---|---|---|---|
| 338 | 176 | 36 | 550 | 26 | 36 |
- HTML: 338
- PDF: 176
- XML: 36
- Total: 550
- BibTeX: 26
- EndNote: 36
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
Review of “UK-Flow15-QC: A quality control framework for better river flow data in hydrological research” by Fileni et al. The paper describes a comprehensive and reproducible quality control framework applied to UK sub-daily time series of river flow data. The paper is well written and clear in its scope. However, I do not think that HESS is the right journal, therefore I recommend that the paper is rejected and encourage the authors to publish in a journal that would be better suited.
The methodology is very specific to UK, and there is no discussion on how this can be extended to other regions, therefore it fails on the criteria of being of interest for a wider audience. The dataset is valuable and care should always be taken when using observed data in modelling exercises, but there is no application to show the magnitude of error it can cause. For example, an exercise of calibration using an identified flawed dataset with a synthetic “perfect” data set could illustrate the impact of data quality on hydrological modelling. Including such results in the paper would increase the applicability of the results in other studies and. As the paper stands now it is a good technical report, but I do not think it is suitable for the journal.
Minor comments