Predicting streamflow drought in the conterminous United States using machine learning and a donor-gage approach, 1982–2020
Abstract. Drought is a highly consequential natural disaster that may likely increase in both severity and extent across the conterminous United States (CONUS). The mechanisms affecting the propagation of drought from the atmosphere to streamflow are complex and interactive, making the prediction of streamflow drought difficult with current modeling approaches. Machine learning is an emerging tool in the field of hydrology that may be well-suited to prediction of streamflow drought across large and topographically diverse areas. Here, we train and analyze 3,198 random forest models at U.S. Geological Survey streamgages to understand common meteorological drivers of streamflow drought and to define physiographic characteristics of basins sensitive to these drivers. We also develop a novel dynamic regionalization approach using donor gages to predict daily streamflow drought at pseudo-ungaged locations. For this study, CONUS was divided into nine regions for ease of reference in describing results. Our results show that teleconnections, temperature, evaporative demand, and snow-water equivalent are important drivers of streamflow drought in the West, Southwest, and Northern Rocky Mountains (Northern Rockies) regions of the United States, and precipitation and soil moisture are primary drivers of streamflow drought in the Northeast, Southeast, and the Northwest regions. Prediction using dynamic regionalization shows comparable performance to at-site models.
- Preprint
(1994 KB) - Metadata XML
-
Supplement
(380 KB) - BibTeX
- EndNote
Status: final response (author comments only)
-
RC1: 'Comment on egusphere-2025-6064', Fedor Scholz, 28 Jul 2026
- AC1: 'Reply on RC1', Aaron Heldmyer, 19 Aug 2026
-
RC2: 'Comment on egusphere-2025-6064', Anonymous Referee #2, 17 Aug 2026
Predicting streamflow drought in the conterminous United States using machine learning and a donor-gage approach, 1982-2020
General
The paper covers the prediction of streamflow drought with random forest approaches and a new dynamic regionalization method for ungauged catchments. The study site covers most of United States and over 3000 sites with different topography and hydrogeological conditions. The goals are to identify the most common drivers of streamflow drought, find the defining basin-characteristics of the most impacted by the drivers, and investigate the use of the new method for ungagged basins.
The paper is generally well-structured and well written. And the subject and approach are very relevant for the readers of HESS and the scientific field. Figures are well presented and easy to decode. The method section is detailed and well-described. There are, however, some points in the comments below that should be addressed prior to publication. Most important are:
- Lack of background meteorological and hydrological information in the study description
- The impact on the very broad definition of droughts applied
Comments
L28: if space allows, the abstract could benefit from a concluding remark on potential of dynamic regionalization in other studies
L38: consider moving “have all resulted in economic losses totalling billions of dollars” to L35 after “modern droughts”
L41: could you add a reference for this statement
L55: “traditional approaches” such as?
L61: remove “potentially”
L78: the study site description is very short, and main talks about the data from GAGES-II. Even though you cover an immense area, it would be good to add a little more information on precipitation, streamflow and temperature patterns in general terms.
Figure1. it seems a little odd to divide the area by administrative boundaries and not hydrological/geological or climatological zones. Is there a specific reason for this approach?
L119: this means that a dry wet-season, would be defined as a drought. In your opinion is this a relevant measure for streamflow drought? Would it not be more relevant to select a threshold based on the dry season? And what about the different predictors of droughts would you not expect these to be different depending on the season they occur. E.g., droughts during low flow season may be very impacted by lower-than normal groundwater baseflow or high PET, while spring droughts could be more related to too little meltwater from mountain ranges or precipitation deficits. The impact of the choice of defining droughts in the way done here in relation to the identified predictors should at least be part of the discussion.
L120: is there trends in dataset? There must be other studies on this
L126: ” Teleconnections data ” would be good to add examples of these in () in the text
L144: what about groundwater or connectivity to groundwater, are there any predictors covering this aspect?
Figure 2: agree with reviewer 1. good figure illustrating the workflow, but the location of the figure and the reference to it in L197, seem very strange and as an afterthought. More introduction to the figure should be added.
L325: extra space should be removed after “range”
L375: very good and informative to include specific examples
L402: “This is likely indicative of hydrologic memory effects that influence the persistence of drought through sustained dry periods, acting as a critical buffer to transient meteorological anomalies.” I don’t understand this. Should it be the case for low aridity areas more than for high?
L417: please add numbers to guide the reader: “1. longer duration, less frequent, and less intense droughts; 2. moderate duration, moderately frequent, and moderately intense droughts; and 3. shorter duration, more frequent, and more intense droughts”
L509: this sentence is very long, consider rephrasing
L513: I don’t understand this sentence
L530: maybe add here that it makes impossible for the user to understand and evaluate the model performance
L542: a donor-based?
L543: new line shift before “We conclude”
Discussion:
Reflections on the impacts of drought definition and the results obtained in this study should be added.
Groundwater is not mentioned a single time in the paper, apart from the mention of baseflow-driven basins and the baseflow-index. How information on groundwater or the geological setting of the basins is included in this analysis is not clear to me. Could you add some reflections on this.
Would be interesting to add some reflections on the future implications of your findings in relation to predicted climate change in the region. What is expected in the region and would we expect the drivers of drought to shift, and how?
Citation: https://doi.org/10.5194/egusphere-2025-6064-RC2 - AC2: 'Reply on RC2', Aaron Heldmyer, 19 Aug 2026
Viewed
| HTML | XML | Total | Supplement | BibTeX | EndNote | |
|---|---|---|---|---|---|---|
| 1,378 | 927 | 110 | 2,415 | 262 | 90 | 106 |
- HTML: 1,378
- PDF: 927
- XML: 110
- Total: 2,415
- Supplement: 262
- BibTeX: 90
- EndNote: 106
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
This work analyzes and predicts streamflow drought in the CONUS via machine learning and statistics. Random forests are trained to predict droughts from meteorological forcings. Gini impurity values are computed to investigate variable importance at the regional scale. Principal component analysis (PCA) on the variable importances allows the investigation of large-scale drivers of drought across the CONUS. Linear regression is used to predict principal components (PCs) from basin characteristics in ungaged basins. Similarity between PCs is used to select donor gages from which forcings are translated to the ungaged basin and converted by the random forest into drought prediction.
The study is well executed and very well documented. The motivation and objectives are clear from the beginning. The methodology is explained in detail and likely allows for reproduction. The dicussion nicely picks up on the research objectives, offers explanations for surprising findings, and draws elaborate connections to existing literature.
Specific comments (most of them are suggestions and therefore optional):
- l 23-24: Sentence about dividing CONUS into nine regions could be removed
- l 61-62: Indeed, ML requires large amounts of data processing during training. Once a model is trained, though, costs are often amortized
- l 62-63: Sentence about interpretability feels out of place here, though interpretability is important, so I would put it somewhere else
- l 66-68: Should conflicting training signals from different basins be called noise? Also, isn't the main advantage of the donor-based approach that it allows PUB?
- l 81: Optionally replace "faced across the nation" with "present across the country" or similar
- l 88: Did you mean non-uniform?
- ll 87-91: I feel a little lost with this information and I don't think it is picked up again. Are you trying to say that the results should be taken with a grain of salt?
- ll 105-107: Once you understand that sentence, it is clear, but it took me a while. Maybe the formula could be extended a bit to make it easier
- l 131-132: Comment: Also, if you would include antecedent streamflow, the task would be much easier since droughts are defined based on streamflow, and the model likely wouldn't learn the connections you're after
- Table 1: I would add horizontal lines to delimit different sources and references. Also, it would be nice if the table wasn't split across two pages
- ll 150-152: What exactly is meant with robustness of the method, especially considering that robustness to redundancies and correlations are separately mentioned right after?
- ll 156-162: This could be shortened a bit
- l 169: Did you play around with the node size parameter and found this value empirically or is this a common choice?
- l 180 and l 176 are the same formula
- ll 191-192: But didn't you explicitly reweight the data so that the classes are balanced? Most likely I'm missing something here, so maybe you could elaborate a bit
- ll 198-201: It was not directly obvious to me how the enumeration relates to the figure. I think adding the paragraph numbers to the enumeration sentence would already help a lot
- Figure 2: This figure contains more than the "development and evaluation of dynamic regionalization method", which is the name of this subsection. It contains all the study's research questions! So I would suggest to reference it not only in that subsection but more generally. Also, the figure is almost completely linear, so I think simply plotting it from top to bottom would be more intuitive than the current layout, but that might be a personal preference
- ll 248-250: These sentences seem redundant to me
- l 256: Similarly to node size, I'm wondering whether the 1% were found empirically
- l 259: I understand that you found 0.35 empirically here, but wouldn't one expect it to be 0.2 given you drought definition? On the other hand, you reweighed your training data, so one might even expect a threshold of 0.5. Interestingly, 0.35 is right in between. Is there something to it, or did I miss something?
- Figure 4: First (and only) occurrence of "Julian Day" is in this caption, maybe introduce it earlier. Also, how does it relate to "Decimal Date"?
- ll 313-314: You explain what loadings are here, but you already used the term in l 229
- Figure 7: If I understand correctly, sometimes, prediction works better with the donor approach than with an at-site models. I think this should be discussed, what could be the reason for this?
- ll 546-550: The conclusion is mostly on a rather high level (which is good), so I wouldn't explicitly mention PC1 etc. here, because those terms are quite technical. Maybe there is a way to phrase it more generally
Technical corrections:
- l 293: "Decimal Date" -> "Decimal date"
- l 350: Remove closing parenthesis
- l 409: "Date" -> "date"
- l 596: doi should be "https://doi.org/10.1175/1520-0493(1987)115<1083:CSAPOL>2.0.CO;2"
- l 603: doi should be "https://doi.org/10.1175/1520-0493(1969)097<0163:ATFTEP>2.3.CO;2"
- l 635: doi should be "https://doi.org/10.1175/1520-0477(1998)079<2715:IMOET>2.0.CO;2"
All in all, this is a very good study in my opinion. After some minor revisions, which mostly concern intelligibility, it is ready for publication in HESS.