Comparing and validating modelled snow stability metrics
Abstract. Recent developments in snow stability modelling demonstrated the great potential of numerical snow cover modelling for avalanche forecasting. These recently developed metrics provided promising results, possibly even superseding traditional stability indices. To further validate these results, we compared the temporal evolution of various stability metrics to a unique dataset of measurements, including snow stratigraphy, snow microstructure, and stability obtained at the high Alpine study site Steintälli above Davos (Eastern Swiss Alps) during the winter 2015/16. There, we measured the shear strength of a prominent weak layer of depth hoar crystals using the shear frame and assessed the propagation propensity of this weak layer with propagation saw tests. Concurrently, we characterized snow microstructure with the snow micro-penetrometer (SMP). At the study site, an automated weather station is located, providing the data to run the numerical snow cover model SNOWPACK so that modelled stability metrics can be derived. Field measurements showed that weak layer strength and toughness were initially low but increased over time, in parallel with the increasing slab load; all three parameters were correlated. While field observations focused on the most prominent weak layer, modelled stability metrics were evaluated for the persistent weak layers identified by the respective modelling approaches. These generally corresponded well to the layers visually identified in simulated snow stratigraphy. Data-driven and process-based stability metrics exhibited similar temporal evolution and comparable skill in relation to local avalanche activity, achieving overall accuracies of 60–70 %. This site-specific validation demonstrates that both types of numerical stability metrics provide valuable support for avalanche forecasting, although differences in their behaviour highlight the need for metric-specific interpretation. Hence, a multi-model approach may be advantageous for operational forecasting.
General Comments:
The paper compares several snow stability metrics derived from snowpack modeling against a rich dataset of field measurements collected at the Steintälli study site during the winter 2015-16. The stability metrics may be classed as process-based (skier stability index SK38, critical cut length rc) or data-driven (machine learning snow instability metric Punstable). The paper discusses the challenges of determining which weak layer in simulated stratigraphy output is the most critical and therefore the most likely to cause avalanche activity. The authors consider three approaches to selecting most critical weak layer: maximum Punstable, relative threshold analysis (RTA), and persistent avalanche problem (pAp) approach. In total, six methods that perform the dual tasks of selecting the critical weak layer and quantifying its stability were evaluated, along with the non-modelled, three-day new snow (HN3d) benchmark.
The field measurement dataset is unusually detailed, comprising shear frame measurements, stability tests (PSTs, ECTs, CTs), detailed snow profiles, automated weather station data, and snow micro-penetrometer profiles. This, combined with how the season evolved to develop several persistent weak layers and several significant avalanche cycles, makes it an extremely valuable dataset, well suited to the task of comparing modeled stability metrics.
The manuscript is well written and the analysis is sound. The authors discuss limitations (training data/validation data overlap, and scale mismatch between point observations and regional danger ratings) and the results generally support the interpretations and conclusions. I recommend publication after the following comments are addressed.
Specific Comments:
The overall accuracies of the stability metrics reported in Table 5 suggest a range between 53% and 73% (PC of 0.53 to 0.73). Can you explain why the overall accuracy range suggested in the abstract is stated as 60-70%?
Several of the metrics have performance skill scores that are very similar, yet no formal confidence intervals are reported. Despite this, there are several places in the manuscript where the relative ranking of the metrics is discussed, admittedly with cautious language: e.g. L. 495 “Punstable performed slightly better than the traditional skier stability index RTA-sk38”. How confident are you in their ranking? How confident would you be that the relative ranking would change if applied to a different area, or a different set of conditions in the same area?
Further to the similarity of some of the metrics, the authors explicitly mention (L. 418) that the strong negative correlation between Punstable and pAp-rc is expected as rc is one of the 6 features considered to estimate Punstable. Does this erode the statement in the conclusion that a multi-model approach may be advantageous for operational forecasting? Are there additional stability metrics available or being developed that offer a truly independent assessment from existing approaches?
The authors are honest about the overlap between the training dataset and the validation dataset (L. 453), where they state that the training dataset for Punstable contained 11 profiles from Davos from 2015/16. Ideally the impact should be characterized by reporting whether excluding the 11 overlapping profiles from the training dataset changes the results. This feels especially pertinent in light of the discussion above, that relatively small differences in performance skill scores are being discussed and characterized. If such reanalysis is not possible, it would be better to highlight this limitation in a stronger way, perhaps by repeating it in the conclusion.
What is the rationale for correcting the avalanche danger level (L. 190)? If the uncorrected danger level was what was available to the public when the regional forecast was issued, isn’t that perhaps a more meaningful comparison, as it recognizes the limitations of human forecasting. As presented, the danger level has been altered to improve correlation with the AAI. Both are presented in Figure 4, ostensibly to offer comparisons to two different measures of avalanche danger. However, these are no longer independent, as one has been corrected against the other. I note that the caption describing the avalanche danger level (“Avalanche danger level for the region of Davos”) does not indicate that a corrected version of the danger level has been used. At a minimum, I would like to see the caption of Figure 4 updated to indicate that the corrected avalanche danger level is used.
PON (L. 334) is not defined (presumably the probability of non-events).
Figure 4: There is a small orange spot on Pane (c) in mid-March that does not have a corresponding orange shaded region on the lower half of Pane (a). I suggest adding a small shaded region on the Punstable region of Pane (a).
Technical Corrections: