the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
Improving Europe-wide windstorm damage modeling using insurance loss data
Abstract. Winter windstorms are among Europe’s deadliest and most damaging natural hazards. In a changing climate, reliably estimating and projecting their impacts is essential for effective risk management. Such risk is often modeled as the intersection between hazard, exposure, and vulnerability, which are linked through functional relationships known as vulnerability curves (or damage models). The Schwierz et al. (2010) damage model is a widely used open-source standard for European windstorms, but its original calibration is based on a limited set of historical UK storms. In this study, this model is calibrated against recent loss data from the PERILS database (1999–2024) within the CLIMADA open-source framework, across 12 European countries. Calibration is conducted independently using two cost functions: Root Mean Squared Error (RMSE) and Root Mean Squared Logarithmic Error (RMSLE), enabling a systematic comparison of their influence on the optimised damage parameter and derived risk metrics. The default model is found to systematically underestimate losses across Europe, and a single pan-European model cannot capture the distinct vulnerability profiles of individual countries. By calibrating the model against PERILS losses, a new set of country-specific damage functions is developed, which reflect spatial heterogeneities in vulnerability. In addition, the analysis demonstrates that the chosen loss function (RMSE versus RMSLE) fundamentally shapes the calibrated curves and the resulting risk profile, underscoring that calibration metric is itself a key modeling decision. The results offer practical guidance for calibrating damage models and support more rigorous climate risk assessment.
Competing interests: At least one of the (co-)authors is a member of the editorial board of Natural Hazards and Earth System Sciences.
Publisher's note: Copernicus Publications remains neutral with regard to jurisdictional claims made in the text, published maps, institutional affiliations, or any other geographical representation in this paper. While Copernicus Publications makes every effort to include appropriate place names, the final responsibility lies with the authors. Views expressed in the text are those of the authors and do not necessarily reflect the views of the publisher.- Preprint
(4687 KB) - Metadata XML
- BibTeX
- EndNote
Status: open (until 21 Aug 2026)
- RC1: 'Comment on egusphere-2026-3099', Anonymous Referee #1, 27 Jul 2026 reply
-
RC2: 'Comment on egusphere-2026-3099', Anonymous Referee #2, 29 Jul 2026
reply
The manuscript recalibrates the Schwierz et al. (2010) windstorm vulnerability function within CLIMADA, against PERILS insured-loss data (1999-2024) for 12 countries, separately under RMSE and RMSLE cost functions, with C3S-EWS gust footprints as hazard and LitPop as exposure. It reports pan-European and country-specific calibrated parameters, event-level scatter, loss-frequency curves, and Unsequa-based uncertainty distributions of normalized average annual impact (AAI), and concludes that (i) the default model systematically underestimates losses across Europe, (ii) a single pan-European model cannot capture country-level vulnerability differences, and (iii) the choice of cost function fundamentally shapes the calibrated model and derived risk metrics.
The aim of the study, a transparent recalibration of a windstorm vulnerability function under two cost functions, is in line with the scope of NHESS and worthwhile for the community. However, my review identified the following issues and limitations, and I therefore recommend a major revision.
Beyond the specific comments below, I want to be transparent that, in my assessment, the central product of this study, the pan-European and country-specific calibrated damage models, will remain highly uncertain even if all comments are fully addressed. The calibration rests on only 3-15 events per country (103 event-country observations in total), drawn from a censored and temporally inhomogeneous loss record; the loss data are aggregated to country-event level, so each fitted multiplier inseparably absorbs hazard biases, exposure-valuation error, and insurance penetration and market coverage alongside actual vulnerability, and is therefore strictly meaningful only within the exact C3S-EWS + LitPop + insured-loss configuration used here; and the fixed curve shape, which saturates at 60 m/s, leaves the damaging tail essentially unconstrained by the data. These are limitations of the data class rather than of the authors’ execution, but they cap the reliability that this type of product can attain.
In that light, the current frontier in this field lies in the resolution, quality and feature extend of the underlying data rather than in recalibration of aggregated curves or in model sophistication. For example, using approximately 9.4 million policy-level records for about one million residential buildings across nine storms in the Netherlands, van Ederen et al. (2026) show that high-resolution data, not model flexibility, drive accurate predictions of European winter storm damage, reducing storm-level prediction errors by about 90% relative to generic European benchmark models (https://doi.org/10.5281/zenodo.20556499). I would encourage the authors to position their contribution relative to such data-driven developments, and to temper the stated reliability and applicability of the recalibrated models accordingly.
Major comments
M1. MDD is not used consistently with CLIMADA
In CLIMADA the damage ratio applied to exposure is the mean damage ratio (MDR), which equals MDD × PAA, where MDD is the mean damage degree of affected assets and PAA is the percentage of assets affected: Impact Functions — CLIMADA 6.1.0 documentation It appears that the paper uses MDD as if it were the MDR, and does not elaborate on PAA or MDR. The manuscript does not provide enough detail to decide if PAA was also scaled, or a scaled MDD was used directly to calculate losses. Recalibrating only the MDD remains valid, but the implementation should be stated unambiguously, because the calibrated MDD is only meaningful in the configuration in which it was fitted: if the calibration chain applied the default PAA (as the standard CLIMADA impact calculation does), the calibrated MDD must also be applied in combination with that PAA, whereas if losses were computed from the MDD alone, subsequently combining the calibrated MDD with the PAA would underestimate damage (by up to an order of magnitude below 60 m/s, where PAA < 1). Please describe the implemented damage model consistently with the CLIMADA terminology and explicitly state if PAA is left out.
Also, the manuscript’s description of the CLIMADA MDD is wrong on several occasions, for instance:
“Negligible damage below approximately 15 m/s” (L127): the function is zero below 20 m/s.
“Reaches approximately 50% of maximum damage at moderate gust speeds (35-40 m/s)” (L128): MDD reaches ~10% of its maximum at 35 m/s and ~20% at 40 m/s; 50% is reached at ~50 m/s.
“The transition region between 30-50 m/s where the slope is steepest” (L130-131): the steepest MDD segment is 55-60 m/s.
“The model’s default state has MDDmax of 0.0387”: Eq. (1) (≈L141) states MDDmaxdefault = 0.0373. With the correct default, the pan-European RMSE calibration (0.0379) raises the default amplitude by ≈1.6%, with the erroneous 0.0387, the same result reads as a slight decrease.
M2. Figs. 3m-n and Fig. 6 appear mutually inconsistent under the stated methodology
Because event losses are exactly linear in θ, and the Unsequa perturbations are identical in distribution for the RMSE- and RMSLE-calibrated runs, shouldn’t the ratio of the two ensemble means of normalized AAI for a given country equal the ratio of the two calibrated parameters (up to Monte-Carlo error)? For instance for Belgium: θ ratio RMSLE/RMSE (Fig. 3) ~ 0.16/0.02 = 8 while AAI ratio RMSLE/RMSE (Fig. 6) ~2.73/0.6 = 4.55 which is a ~ +80% mismatch. The likeliest explanations are that Fig. 6 was produced with a different (earlier, shrunken, or otherwise transformed) parameter set than displayed in Fig. 3, or that its normalization differs from the caption. Please recompute or clarify.
In general, it would also help interpreting the results to complement the figures (e.g. Figs. 2 and 3) with tables reporting their numerical values and, for instance, % differences between modelled and observed damage.
M3. The simplified curve in Schwierz et al. (2010) was derived primarily from insured residential-sector damage but is applied to other asset types as well
The manuscript applies this model to LitPop exposure, such that it effectively transfers a residential insured-property wind-damage relationship to all kinds of exposure, such as commercial, industrial or infrastructure assets. Therefore, the fitted (country) multiplier will (also) absorb sectoral composition. Sector-specific damage models would be preferable to control for their vulnerability characteristics. Hence, the recalibrated damage models are likely too aggregated to successfully transfer to another period or portfolio whose residential/commercial mix differs. The paper should be clear about this limitation.
M4. The headline claim “the default model systematically underestimates losses across Europe” is not fully supported and exposure assumption dependent
Pan-European RMSE calibration indicates that there is no pan-European underestimation to correct. In 6 of 12 countries (BEL 0.020, NLD 0.019, NOR 0.021, GBR 0.030, IRL 0.031, LUX 0.034, vs default 0.0373), the country-level RMSE calibration moves the parameter below the default: for these countries the squared-error view says the default over-predicts.
The abstract, L249-250, L272-274, and the first sentence of the Conclusions should be reformulated to make the claim loss-range- and country-specific.
“The default model underestimates losses” conclusion is conditional on the LitPop exposure: assuming higher exposure values would result in a lower recalibration, which could reach up to the point that the default model overestimates losses. Hence, this is not an inherent flaw of the Schwierz et al. (2010) model, but rather a result of the exposure (and hazard) input assumptions. Readers should (for that reason) not be discouraged from using the Schwierz et al. (2010) model. The Schwierz et al. (2010) model relied on insured values at risk, which differs from the LitPop exposure proxy. The paper should be more clear about the dependence on the exposure assumption.
M5. The LOOCV analysis cannot support the conclusions drawn from it
Section 2.4 reports a single statistic (correlation 0.97 between LOO predictions and in-sample fits) and concludes that the calibration is “robust and not driven by individual extreme events” and that “any calibrated MDDmax value is a generalized parameter suitable for future event prediction”. Predictive skill is asserted against the wrong reference, as correlating LOO predictions with in-sample predictions says nothing about agreement with observed losses.
M6. The AAI evaluation appears in-sample, which would not support prediction accuracy conclusions
AAInorm is computed over the same 1999-2024 events used to fit θ. Statements such as RMSE giving “a stable estimate of the expected annual loss” (DEU 0.98, CHE ~1.06, AUT ~0.95) present in-sample agreement as accuracy. Out-of-sample AAI assessment would make these statements meaningful.
M7. The practical guidance is contradicted by the manuscript’s own results
The Discussion recommends RMSLE calibration for “applications focused on pricing attritional losses or estimating expected annual damages” (L335-337). But Fig. 6 shows RMSLE-calibrated models overestimate AAI in 11 of 12 countries, by factors of 1.3-3.0 (BEL 2.73, NLD 3.01, NOR 2.94, LUX 2.28, DEU 2.11…), while the same section elsewhere concedes RMSLE “frequently overestimates in aggregate loss assessments” (L300-301). Recommending the systematically AAI-overestimating calibration for expected-annual-damage estimation is a direct internal contradiction. RMSLE is defensible for event-level relative accuracy in the attritional range (Fig. 5 supports this); it is not defensible for AAI.
M8. The EM-DAT comparison conflates loss definitions
PERILS records insured losses; EM-DAT records total economic damages (and its entries are not adjusted here in any stated way). Total economic windstorm losses are often higher than insured losses, which may explain why an EM-DAT-calibrated amplitude is about 2-3× that of the PERILS. The manuscript instead frames the divergence as “uncertainty introduced by a much smaller sample size” (L220-221) and tests it against PERILS subsampling variability. The difference in loss definitions are never mentioned, leaving the reader with the impression that EM-DAT is merely noisy rather than measuring a different quantity.
M9. Catalog inhomogeneity and indexation handling are unaddressed
Per Section 2.1, PERILS contains all qualifying events only since 2009; 1999-2008 is represented by exactly five retrospectively added major storms, and the qualification threshold changed from EUR 200M to 300/500M in September 2022. The 1999-2024 record is therefore censored non-uniformly in time: the first decade contains only the largest storms. This biases the empirical “occurrence frequency” axis of Fig. 5 (return periods of moderate events are overstated), deflates AAIrep relative to a true above-threshold AAL, and weights the calibration toward whichever epochs a country’s events happen to fall in. The Discussion acknowledges threshold censoring generically (L341-343) but not the 1999-2008 inhomogeneity. At minimum this should be analyzed (e.g., recalibrating on 2009-2024 only as a sensitivity test).
The manuscript never states whether PERILS losses are used as originally reported (nominal, at-event-date market conditions) or as-if-adjusted to present-day exposure, nor which LitPop vintage is used. This must be specified, and if unadjusted losses were used, corrected.
M10. Interpreting calibrated amplitudes as “vulnerability profiles” / “the underlying wind-damage relationship” overstates what θ can identify
θ is estimated from insured losses against a total-asset exposure (LitPop). It therefore absorbs, inseparably: physical vulnerability, insurance penetration and PERILS market-coverage share (which vary strongly across these 12 countries) and systematic hazard-intensity errors. The Discussion’s single clause “market or claims conditions” (L314) is not enough to license the abstract’s “reflect spatial heterogeneities in vulnerability”, L247-248’s “reflects the underlying wind-damage relationship observed in each country”, or the Conclusions’ framing. Since only the amplitude is calibrated with a fixed curve shape, genuine cross-country differences in threshold or steepness cannot be accounted for. Relatedly, no per-country parameter uncertainty is ever shown (the subsampling band exists only pan-European), with n = 3-15 events the sampling uncertainty is very high.
M11. Descriptive statistics are missing
The manuscript reports no descriptive statistics of the calibration data. Without these, readers cannot judge sample representativeness, the leverage of individual events on the calibrated parameters, or the severity range over which the models are actually constrained.
Minor comments
- Figure cross-references are systematically wrong in Sections 3-4. L289-292 cite Fig. 3e/3i/3q/3u for values that are in Fig. 6e/6i/6q/6u (Fig. 3 has no panels q-u); L310 cite Fig. 5 for the parameter panels of Fig. 3; L275-276 cite Fig. 3c/3d/3h within a Fig. 5 discussion (the panel letters coincide, but the figure number should be 5); L280-281 similarly cite Fig. 3a/3j/3k/3l for exceedance-curve behavior shown in Fig. 5.
- Transposed values. L288-289 quote RMSE mean AAI “0.98, 0.95, and 1.06” for Germany, Switzerland, Austria; Fig. 6 gives DEU 0.98, CHE 1.06, AUT 0.95 (CHE and AUT swapped).
- Norway misclassified. L244 claims IRL, GBR, NOR have RMSLE MDDmax < 0.04 (“substantially flatter”); Fig. 3m shows NOR at ≈0.044-0.045, above both 0.04 and the default (0.0373), and the Fig. 4 RMSLE map correspondingly shows Norway positive (≈ +18%). Only IRL (0.013) and GBR (0.034) are below default.
- “British Isles show small negative changes” (L263, RMSLE map): Great Britain yes (≈ −9%); Ireland no (≈ −65%, the darkest negative cell on the map).
- “a 25-year catalog” (L271) vs “1999-2024, 26 years” (Figs. 5 and 6 captions): inconsistent; 1999-2024 inclusive is 26 years.
- 5 x-axis is mislabeled. The axis labeled “Occurence Frequency (1/years)” increases toward rarer, larger events and spans ≈26/n to ≈26; it is the empirical return period in years, not a frequency in 1/years (also: “Occurrence” is misspelled). The caption and L270-272 should be corrected accordingly.
- Prior article. CLIMADA itself deploys the damage model from Welker et al. (2021), a multiplicative RMSE recalibration of the same Schwierz function against Swiss insurance data, methodologically the direct predecessor of this study. Only the 2020 discussion version is cited, in passing. The novelty framing (multi-country PERILS calibration; RMSE-vs-RMSLE comparison) would be strengthened, not weakened, by engaging with it.
- ±20% vulnerability-uncertainty attribution (L203-205): please point to where Prahl et al. (2012) and Severino et al. (2024) establish ±20% as a bound for this parameter, and specify what exactly is “consistent with” Severino et al. (2024) at L142-143 (their calibration strategy differs in implementation details).
- Koks and Haer (2020) did not employ the Schwierz et al. (2010) damage model (L52-53)
- Typographical: “Infact” (L25); “optimzation” (L309); missing space “(>60 m/s)(see Fig. 2a)” (L129); “formal analysis,visualization”
Citation: https://doi.org/10.5194/egusphere-2026-3099-RC2
Viewed
| HTML | XML | Total | BibTeX | EndNote | |
|---|---|---|---|---|---|
| 179 | 71 | 13 | 263 | 9 | 13 |
- HTML: 179
- PDF: 71
- XML: 13
- Total: 263
- BibTeX: 9
- EndNote: 13
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
Summary
The author team recalibrate a widely used damage model for European wind storms using recent loss data from the PERILS database (1999–2024). Thereby the used separate calibration on the country level of 12 European countries. Calibration uses two cost functions: Root Mean Squared Error (RMSE) and Root Mean Squared Logarithmic Error (RMSLE). The default model systematically underestimates losses across Europe, and a single pan-European model cannot capture the distinct vulnerability profiles of individual countries. Furthermore, the selected cost functions shape the calibrated curves and the resulting risk profile, suggesting that cost function is itself a key modeling decision.
General
The author team presents a newly calibrated damage model for European storms and highlights the importance of regionalization and dependence on the selection of the calibration metric. Clearly, this is new and perfectly fits the objective of NHESS. The manuscript is well structured and well written, though some of the figures could be improved. I recommend minor revisions for possible publication in NHESS.
Comments
L13: I think the authors mean “cost function” here.
L14: I suggest to change “calibration metric” to “cost function” here to make it clear.
L25-26: Suggestion “In recent years, these regions have witnessed multiple billion-euro catastrophe events, including the February 2022 European…”
L36-38: The entire sentence must be reviewed. Note I normally reinsurance contracts run max 2 years, so climate change plays not a big role in business.
L59: “Therefore, robust calibration and validation is necessary, …”
L60: Which parameterizations are meant here, maybe avoid using this wording as people might think about parameterizations in dynamical models, I guess you mean statistical models here, such as the damage model used.
L93: The publications at the end of the sentence are misleading as they suggest that they refer to the dataset, but they only describe the storms, so please cite the dataset here.
L101-102: How is inflation handled in these data sets?
L110: This sentence can be removed as it is described in 2.3
L133: Please name the subsection “Calibration”, this is sufficient
Section 2.3: The number of events n is rather small and should be mentioned in this section.
L189-191: The sentence is not clear. It is also not clear why any calibrated MDD_max value is a generalized parameter. You need to give an explanation.
L204: “who estimated”
Results: Maybe this is the only more major comment. I think the authors need to discuss the effect of the rather small number of events on the calibrations much more. It is mentioned in the last section, but I think the reader needs in the results section information about which region the calibration is trustworthy and in which it is at the limit.
L234: “pulling down modeled damages” this is not visible in Fig 2c, it looks identical.
Fig. 3: Too small labels and axes labels. Maybe split the figure into two (m,n as separate figure).
Fig 5: Too small labels and axes labels. You also change the use of the colors, In Fig 3 orange is used for RMSLE, here for RMSE, please adjust.