the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
Probabilistic tsunami hazard assessment at Stromboli volcano: 1. Review of historical sources and expert elicitation findings
Abstract. Active volcanic islands, such as Stromboli in southern Italy, are sites where tsunamis generated by volcanic activity could be frequent and potentially destructive. Stromboli Island has experienced several landslides over the past decades, some of which have generated destructive tsunamis. This paper is part of a broader project aimed at developing a Probabilistic Tsunami Hazard assessment (PTHA) for Stromboli. We present here a review of historical tsunamis sourced from Stromboli, their correlation with explosive activity, and the results of an expert elicitation on tsunamigenic landslides.
In our review of historical tsunamis, we identified 16 events from 1879 to 2024, grouped into three classes based on the degree of inundation observed in the village of Stromboli, ranging from few ten to few hundred of meters inundation distances, the latter comparable to the widely documented December 2002 tsunami event. Four historical tsunamis (in 1879, 1921, 1924, and 1959 CE) have been critically discussed for the first time. Over the past 150 years, ~69 % = 11/16 of the catalogued tsunamis, with 90 % confidence [45 %, 87 %], were associated with paroxysms, while only ~27 % of historical paroxysms were associated with catalogued tsunamis, confidence [16 %, 40 %]. Similar conditional probabilities and uncertainty intervals were estimated from 1916 to 2025, excluding the tsunamis without significant inundation. The expert elicitation was divided in three parts. Part I provided estimates and uncertainty quantification of the number of tsunamigenic landslides at Stromboli (with volumes ≥ 1 × 106 m3) in the past and of those expected in the next 50 years. Part II focused on the probabilities of different triggering mechanisms in the next 50 years. Part III quantified the probabilities of different tsunamigenic landslides (volume and positions) along the Sciara del Fuoco in the next 50 years. Results of the expert elicitation indicate that return periods of tsunamigenic landslides at Stromboli in the next 50 years have median values in the order of 10–12 years (with uncertainty from 3 to 50 years), and a median probability of their occurrence along the Sciara del Fuoco of either 82 or 86 % (depending on the weighting scheme used in the elicitation). Results of part III indicate slightly higher median probabilities for landslides occurring at elevation 300 to 700 m a.s.l. along the Sciara del Fuoco with volumes 1–5 × 106 m3 as compared to other elevation and volume ranges.
- Preprint
(2433 KB) - Metadata XML
-
Supplement
(17545 KB) - BibTeX
- EndNote
Status: final response (author comments only)
- RC1: 'Comment on egusphere-2026-1310', Anonymous Referee #1, 05 Jun 2026
-
RC2: 'Comment on egusphere-2026-1310', Anonymous Referee #2, 24 Jul 2026
First of all, I apologize for the delay in completing this review. I incorrectly assumed that it was not needed, and I only received a notification regarding it very recently.
This manuscript employs established methodologies to derive crucial parameters intended as inputs for a companion paper. It also presents compelling results and several valuable points for discussion. In my opinion, it is worth publishing in NHESS, but there are several concerns that need to be thoroughly addressed before publication.
Main Concerns
1. Catalog Completeness & Probability Estimates: All probabilities quantified from the catalog of tsunami can be highly misleading without a robust discussion on record completeness. Catalog incompleteness can lead to poorly constrained probabilities and may contribute far more to overall uncertainty than the formal uncertainties derived from Clopper and Pearson (1934), which assume complete records. Furthermore, the expert elicitation results themselves seem to imply some incompleteness (questions TQ1 and TQ2 would be redundant). If this is indeed the case, raw probability estimates—particularly in the abstract—should be de-emphasized in favor of simple counts and percentages.2. Conditional Probabilities Evaluation: Similar concerns apply to the estimation of conditional probabilities, even if they emerge from different reasons. The rationale and justification behind the counting approach in Section 3.1 are, in my opinion, weak. Specifically, the comparison between volcanic explosions, their annual rates, and historical tsunamis is poorly described in terms of both methodology and underlying logic. Several things are not fully clear, like for example how the associations between individual events were defined, from where the defined counts exactly come from, how delay/anticipation of tsunamis is dealt with, why the association between individual events is defined with a plot reporting a smoothed function of explosion rates instead of individual explosions, etc. More details are reported below, in the detailed comments. Additionally, as above, also these calculations are susceptible to completeness issues for both tsunami and explosion records. Therefore, I think that the underlying quantification methodology must be clarified, and the probability claims should be scaled back (e.g., in the abstract) in favor of simple counts and percentages.
3. Context and Methodological Rationale for Expert Elicitation: Although expert elicitation is central to this study, the chosen framework is not properly contextualized within the broader elicitation literature. The manuscript lacks a background introduction to expert elicitation methods, particularly those applied in volcanology and for tsunamis in the past, and fails to root the selected approach into this background. Furthermore, if the elicitation sessions were conducted online, the authors must discuss the potential disturbance of a performance-based weighting method in the context of modern internet tools (e.g. AI), given that several seed questions could be in principle easily resolved using online AI resources.
4. Logic Tree Logic & MECE Assumptions: Regarding the structuring of questions, it is unclear whether the branches of the trees used to define elicitation questions are assumed to be Mutually Exclusive and Collectively Exhaustive (MECE), though the presented results seem to imply this assumption. For several levels of the trees, the MECE property appears questionable and requires more justification. Is this potential non-exhaustive formulation of questions problematic for the results? In addition, if MECE is assumed, it is also not clear to me how this MECE assumption was dealt with in analyzing the elicitation results. Since reported probabilities seem to sum to 1 across different percentiles, was this normalization explicitly enforced during the elicitation, or were questions explicitly framed to require MECE inputs? More details and discussion about this are required.
5. Missing Critical Information from the Main Text: Essential details are currently omitted or buried in the supplementary material. The main text must explicitly state when the two elicitations took place and whether those sessions were held online or in person. Furthermore, a expert weights distribution plot and a full summary table of all questions (including seed and target questions) must be integrated directly into the main text for immediate visibility. In addition, a snapshot of the entire questionnaire proposed to experts (as visible to them in the first place) should be placed as supplementary file.
6. Discussion of Elicitation Results & Temporal Evolution: The discussion surrounding Part I of the elicitation is underdeveloped given its importance and the striking results. Further details are needed, including potential biases introduced by the timing of the sessions or by specific topics discussed during group meetings (also providing meeting notes in the supplementary materials would be very helpful in this). Additionally, a dedicated comparison between Elicitations 1 and 2 is required to spot specific shifts in expert judgment, connecting them to time history of events that occurred between the two rounds.
All these points are further detailed in the following detailed comments.
Detailed Comments
Line 155: Why is "greater than" not considered here?Line 162 and followings: There is no completeness analysis provided, making these probabilities highly questionable. A formal discussion on catalog completeness must be introduced; without it, these numerical values lack physical meaning, and simple counts should be considered.
Line 173 and followings: This section is very short and does not describe the context. The manuscript lacks an introduction to expert elicitation backgrounds and the rationale for the selected approach. Please summarize standard methodologies used in volcanology and tsunami science in the past, discussing the Classical Model in this and in the wider context, and clarify the rational of choosing this method for this analysis.
Line 190: Are performance-based elicitation protocols compatible with online delivery in the era of accessible AI tools? Please discuss potential biases, such as artificial expert convergence due to AI use more than to expert's competence or ability in evaluating personal uncertainty.
Line 205: How were these combined? Directly by the experts, or using a formal mathematical aggregation method?
Line 218: As noted for Line 190, were these sessions in-person or online? Did all experts participate in every meeting, and is a full list of participants available for all meetings? Please, add these important details. Were notes taken during plenary meetings regarding key discussions, and on what exact dates did these meetings occur? If so, it would be greatly helpful to have them in the supplementary material, along with the questionnaires as they have been presented to the experts.
Line 248: Regarding the logic tree: what is its structural relationship to standard logic trees adopted in seismic and volcanic hazard assessments? Are these event trees or logic trees? Maybe it is less confusing to call simply tree.
Line 248: Do the branch probabilities strictly sum to 1, and are they defined as Mutually Exclusive and Collectively Exhaustive (MECE)? This requires explicit discussion, as percentage representations in Figure 2 seem to imply a MECE setup, but the individual levels seem not fully designed to be MECE.
Lines 250–258: At Level 2, why must a tsunami necessarily be triggered by something? Is the probability of no trigger or of an unlisted trigger mechanism exactly zero? On what basis can a zero probability be assigned without expert input? Similarly, at Level 3-2 and Level 4, are there no alternative mechanisms? Why was an "other" category omitted to guarantee collective exhaustiveness? While historical records may be dominated by these mechanisms, low-probability alternatives could be critical for 50-year long hazard forecasts.
Line 255 (Level 3-1): Please clarify what is meant by an "exogenous volcanic trigger." Cross-referencing the discussion in Level 4 would help clarify this term.
Lines 266–274: "examples of volcanic settings" mean that there are also examples in which there are no trigger? Please, clarify.
Lines 270–274: Please include also the 1980s tsunami at Vulcano Island, given its location within the same archipelago.
Line 275: The explanation for Level 3 – Case 1 appears to be missing.
Line 280: While I understand the reasoning behind evaluating individual factors, specific parameter combinations may exert a stronger control on hazard than any isolated trigger. Please provide further justification for this simplification. Was it possible to capture combined factor effects through the questionnaire structure or free-comment fields? Did you have any specific indication from experts?
Figure 2: Do these percentages sum to 100%? Was this elicited via a single question (calculating the complement) or as separate queries?
Line 292: Why are depths greater than 700 m excluded? Are they considered negligible in terms of tsunami generation?
Line 301: At this point in the text, it is unclear whether volume is treated independently of water depth or subaerial/submarine setting. This only becomes clear in Figure 8; please clarify it earlier here.
Line 335: Please clarify the phrase "discussed here in detail for the first time." Does this mean these events are introduced for the first time, or that they were already discussed and new analytical details are being presented here?
Lines 359–361 & Figure 5: The application of kernel density estimation here is confusing to me. Does a single eruption imply tsunami occurrence within a 1-year window before or after an explosion (as modelled with the Gaussian kernel bandwidth)? If the explosion works as a trigger, the link should be 1-1, the events must be almost simultaneous, and the tsunami must follow. Why do you prefer this method? Furthermore, why compare annual explosion rates directly against discrete tsunami occurrences? Clarify whether individual explosions or background annual rates govern the process; the text implies individual explosion triggers, whereas this approach and Figure 5 suggests a rate-driven process.
Lines 365–368: Counts such as "11 out of 16" are not readily apparent from Figure 5, which displays rates rather than discrete events. Was a specific rate threshold or temporal proximity window applied? Or a threshold in smoothed annual rate of explosions? If a tsunami precedes an explosion in the record, is it counted identically?
Figure 5: The caption is unclear regarding the line styles (e.g., whether the solid blue line equals the red line). Additionally, the solid blue and solid black lines are visually indistinguishable to me.
Line 365: Which probabilities are being referenced here? If referring to the percentages calculated at Line 365, why compare them directly when one sample considers only major tsunamis (Class 2–3) and the other includes all recorded tsunamis?
Lines 376–384: Please clarify the conditioning window used (e.g., a paroxysm conditional on a tsunami occurring within a 6-month window before or after).
Line 385: What level of completeness is assumed across the dataset (major explosions, paroxysms, and tsunami classes)? Why does the evaluated timeframe extend longer (150 years vs. 110 years) when including smaller, easily missed tsunamis? Could completeness issues dominate parameter uncertainty more than the formal binomial confidence intervals (Clopper & Pearson, 1934)?
Line 391: The choice of cutoff value and its sensitivity impact are not discussed.
Line 392: Performance weights are only provided in the supplementary material. A weight distribution plot should be added to the main text. Note that if only 6 out of 21 experts scored above the cutoff, this represents <30% of the panel, not 40%.
Line 393: Please clarify what is meant by "seed realization" (e.g., expert responses to calibration questions). Explain the role of "balanced" formulations and alternative choices to make this section self-contained.
Lines 396–401: The fact that 18 out of 21 experts revised their answers is a significant finding. Did this shift also occur for seed questions, and how substantial were these revisions? Were changes driven by group discussions or new observational data? Please provide metrics on the shift between rounds.
Line 451: The bimodal distribution in TQ15 appears only for the Classical Model. Does this reflect disagreement across the whole group, or specifically among high-weight experts?
Lines 452–453: Please clarify what is meant by "experts assigned nearly all probability (over 80% median value)." Why does this indicate dominance? In Figure 7, lava accumulation appears to show equal or higher probability.
Figure 7: This figure shows decision-maker (DM) weight distributions rather than the logic tree structure.
Lines 569–571: The elicitation results show noticeable anchoring to recent observational data (TQ2), which is projected backward historically (TQ1) but not forward over the 50-year forecast window. What drove this discrepancy? Did experts explicitly discuss catalog completeness between 1879–2024? Additionally, regarding Lines 572–575, assuming a 145-year average rate matches a ~1200-year history but fails to apply to the next 50 years is a strong assertion. Was there evidence discussed showing that recent activity (1995–2025) is anomalous enough to alter short-term future rates? How did timing and recent volcanic crises influence these responses between Elicitations 1 (2022) and 2 (2024)?
Line 573: Clarify whether activity during 1995–2025 was higher or lower than the baseline average.
Line 574: Please clarify the usage of "reasonable" vs. "likely." Are you describing the physical validity of the rate change, or the likelihood of expert cognitive bias?
Lines 599–600: This point is critical and links back to the decision to omit parameter combinations (as noted for Line 280).
Line 599: As with Line 574, clarify the precision of the term "reasonably."
Lines 606–607: This connects directly to the discussion on TQ7 in Line 579.
Lines 630–635: Since the question used an open-ended interval (>30), the only constraint lies in the power-law computations. Why focus specifically on the 30–100 interval there, rather than a broader range? Does this choice alter the outcomes?
Lines 668–686: This hazard quantification discussion belongs more appropriately in the companion paper (de' Micheli Vitturi et al.), as hazard results are not presented here. I suggest retaining only the details regarding input parameters and the final concluding sentence, deleting from here all the considerations regarding the hazard results
Citation: https://doi.org/10.5194/egusphere-2026-1310-RC2
Viewed
| HTML | XML | Total | Supplement | BibTeX | EndNote | |
|---|---|---|---|---|---|---|
| 352 | 234 | 33 | 619 | 38 | 26 | 28 |
- HTML: 352
- PDF: 234
- XML: 33
- Total: 619
- Supplement: 38
- BibTeX: 26
- EndNote: 28
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
Please see the attached report