Towards a universal hydrologic metric for predicting rainfall-triggered landslide timing
Abstract. Metrics that approximate hillslope hydrologic response to rainfall are fundamental for informing landslide risk reduction efforts, such as early warning systems and hazard models. Notwithstanding the numerous publications using different wetness metrics that underpin and largely control the accuracy of landslide risk reduction products, a robust comparison of a broad array of different wetness metrics for regional landslide analysis is currently lacking in the literature. In this study, we statistically compare common wetness metrics for predicting the temporal occurrence of rainfall-triggered landslides at regional scales (> 1000 km2) using a landslide inventory covering the contiguous United States. We find that representations of hillslope wetting and drainage using parsimonious leaky bucket models that only require rainfall input and an estimated drainage factor can identify landslide-triggering hydrologic conditions across disparate ecological regions more accurately than unmodified precipitation metrics or more complex hydrological models. Due to the proliferation of global precipitation datasets and the limited input data needed for the parsimonious leaky bucket model, this model could be used to improve tools for landslide risk reduction.
- Preprint
(2217 KB) - Metadata XML
-
Supplement
(2059 KB) - BibTeX
- EndNote
Status: open (until 07 Sep 2026)
- RC1: 'Comment on egusphere-2026-3513', Anonymous Referee #1, 14 Jul 2026 reply
-
RC2: 'Comment on egusphere-2026-3513', Anonymous Referee #2, 06 Aug 2026
reply
-
General Summary: This manuscript presents a comprehensive and ambitious evaluation of 15 hydrologic metrics for predicting the timing of rainfall-triggered landslides across the contiguous United States (CONUS). The study's use of Bayesian statistical modeling and ecoregion-level analysis provides robust evidence that parsimonious "leaky bucket" models (specifically AWI with kd = 0.01) and Transient Rainfall Infiltration (TRI) models outperform more complex Land Surface Models (LSMs) in identifying landslide-triggering conditions. This finding is highly significant for the development of regional early warning systems and hazard models. The methodology is sound, and the results are well-supported. However, minor clarifications regarding terminology, data resolution, and specific disaster mechanisms are required to enhance the manuscript’s clarity and reach.
-
Specific Comments
A. Definition and Scope of "Landslides" The study uses the term "landslide" as a broad category that includes various movement types, such as debris flows. Notably, the southern California inventory consists of approximately 30% post-fire debris flows.
- Point for Revision: While the authors cite the Varnes (1978) classification, they acknowledge that specific movement types were not consistently reported in the inventory. The authors should more explicitly discuss the implications of applying the same infiltration-based metrics to both soil-slips (infiltration-triggered) and debris flows (often runoff-triggered). Additionally, it would be beneficial to state clearly that slow-moving landslides, where timing is poorly constrained, were excluded to focus on "rapid failure" events.
B. Spatial and Temporal Resolution Correspondence The study utilizes precipitation and soil data at a spatial resolution of ~800m and a temporal resolution of 1 hour.
- Point for Revision: It is well-documented that debris flows, particularly in burned areas, are often triggered by high-intensity rainfall over 5 to 15-minute intervals. Since the metrics analyzed have a maximum resolution of 1 hour, there is an inherent limitation in capturing these short-duration peaks. Please add a brief discussion on how this "data coarseness" serves as a technical bottleneck (currently unfeasible at regional scales) for predicting small-scale or rapid-onset disasters.
C. Missing Attributes: Scale and Volume
- Point for Revision: The authors note that size, depth, and volume were not analyzed due to inconsistent reporting in the USGS database. However, the relationship between landform resolution (800m grid) and disaster scale is important; extremely small failures might be "noise" relative to such a large grid. Please include a sentence in the Discussion regarding how the lack of magnitude data might limit the evaluation of metric sensitivity for different landslide scales.
D. Performance Equivalence in Specific Regions (e.g., Mediterranean California) The results show that many metrics perform equivalently well in the Mediterranean California (MC) ecoregion.
- Point for Revision: This suggests a "ceiling effect" or "local abnormality" where landslide-triggering events are such extreme outliers in that climate that even simple metrics can identify them with high accuracy. Please provide a more detailed physical or statistical explanation for this regional performance overlap.
E. Post-fire Debris Flow Dynamics
- Point for Revision: Post-fire debris flows involve fundamentally different mechanisms, such as soil hydrophobicity and enhanced surface runoff, rather than the infiltration-drainage balance modeled by AWI or TRI. The authors should clarify that the lower performance of metrics in southern California is likely due to these unique properties of burned areas. This would further support the argument for why a "universal" metric remains elusive while validating the AWI model as a "first-order estimate" for regional unburned terrain.
3.Conclusion
The manuscript is well-written and provides a valuable contribution to the field. Addressing the points above will clarify the boundary conditions of the proposed models and strengthen the manuscript's conclusions. I look forward to seeing the revised version.
Citation: https://doi.org/10.5194/egusphere-2026-3513-RC2 -
Viewed
| HTML | XML | Total | Supplement | BibTeX | EndNote | |
|---|---|---|---|---|---|---|
| 136 | 50 | 12 | 198 | 22 | 11 | 6 |
- HTML: 136
- PDF: 50
- XML: 12
- Total: 198
- Supplement: 22
- BibTeX: 11
- EndNote: 6
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
This manuscript compares multiple wetness metrics for identifying the timing of rainfall-triggered landslides across the contiguous United States and concludes that a simple antecedent wetness index may provide broadly transferable predictive capability. However, I do not believe that the underlying landslide data or the evaluation framework can support this conclusion. The inventory combines heterogeneous sources and mapping procedures, converts polygon inventories to centroid points, accepts failure times with only daily precision, and identifies rainfall-triggered cases largely by excluding earthquake inventories and requiring non-zero precipitation on the reported failure day. This procedure does not establish rainfall causality, while essential attributes such as landslide type, size, depth, and movement mechanism are unavailable or inconsistent. The inclusion of post-fire debris flows and other mechanistically distinct phenomena further compromises sample homogeneity. More importantly, the percentile-based evaluation relies on the unverifiable assumption that the incomplete inventory is unbiased with respect to rainfall extremity, although the manuscript itself acknowledges substantial reporting biases toward particular regions and widespread landslide events. Multiple landslides generated by the same storm and the repeated metric values calculated for each landslide also introduce strong dependence and pseudoreplication, yet the Bayesian model does not appear to include landslide-, storm-, inventory-, or region-level dependence structures. No independent temporal, spatial, or event-based validation is presented, and the analysis therefore represents a retrospective ranking at known landslide locations rather than a robust assessment of predictive skill or false-alarm performance. In addition, the statistical formulation in the supplement contains serious inconsistencies: Equation S6 reverses the percentile scale as written, Equation S7 uses a standard deviation where a variance is required for beta-distribution parameterization, and the upper bound in Equation S15 reduces to zero. These problems make the analysis difficult to reproduce and cast doubt on the reported statistical comparisons. Because the principal conclusions depend directly on uncertain event labels, non-independent samples, an inadequately validated performance measure, and an internally inconsistent statistical model, these are foundational issues that cannot be addressed through routine revision. I therefore recommend rejection.