the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
Improving Europe-wide windstorm damage modeling using insurance loss data
Abstract. Winter windstorms are among Europe’s deadliest and most damaging natural hazards. In a changing climate, reliably estimating and projecting their impacts is essential for effective risk management. Such risk is often modeled as the intersection between hazard, exposure, and vulnerability, which are linked through functional relationships known as vulnerability curves (or damage models). The Schwierz et al. (2010) damage model is a widely used open-source standard for European windstorms, but its original calibration is based on a limited set of historical UK storms. In this study, this model is calibrated against recent loss data from the PERILS database (1999–2024) within the CLIMADA open-source framework, across 12 European countries. Calibration is conducted independently using two cost functions: Root Mean Squared Error (RMSE) and Root Mean Squared Logarithmic Error (RMSLE), enabling a systematic comparison of their influence on the optimised damage parameter and derived risk metrics. The default model is found to systematically underestimate losses across Europe, and a single pan-European model cannot capture the distinct vulnerability profiles of individual countries. By calibrating the model against PERILS losses, a new set of country-specific damage functions is developed, which reflect spatial heterogeneities in vulnerability. In addition, the analysis demonstrates that the chosen loss function (RMSE versus RMSLE) fundamentally shapes the calibrated curves and the resulting risk profile, underscoring that calibration metric is itself a key modeling decision. The results offer practical guidance for calibrating damage models and support more rigorous climate risk assessment.
Competing interests: At least one of the (co-)authors is a member of the editorial board of Natural Hazards and Earth System Sciences.
Publisher's note: Copernicus Publications remains neutral with regard to jurisdictional claims made in the text, published maps, institutional affiliations, or any other geographical representation in this paper. While Copernicus Publications makes every effort to include appropriate place names, the final responsibility lies with the authors. Views expressed in the text are those of the authors and do not necessarily reflect the views of the publisher.- Preprint
(4687 KB) - Metadata XML
- BibTeX
- EndNote
Status: final response (author comments only)
-
RC1: 'Comment on egusphere-2026-3099', Anonymous Referee #1, 27 Jul 2026
-
AC1: 'Reply on RC1', Aditya Narayan Mishra, 23 Sep 2026
Summary
The author team recalibrate a widely used damage model for European wind storms using recent loss data from the PERILS database (1999–2024). Thereby the used separate calibration on the country level of 12 European countries. Calibration uses two cost functions: Root Mean Squared Error (RMSE) and Root Mean Squared Logarithmic Error (RMSLE). The default model systematically underestimates losses across Europe, and a single pan-European model cannot capture the distinct vulnerability profiles of individual countries. Furthermore, the selected cost functions shape the calibrated curves and the resulting risk profile, suggesting that cost function is itself a key modeling decision.
General
The author team presents a newly calibrated damage model for European storms and highlights the importance of regionalization and dependence on the selection of the calibration metric. Clearly, this is new and perfectly fits the objective of NHESS. The manuscript is well structured and well written, though some of the figures could be improved. I recommend minor revisions for possible publication in NHESS.
We thank Referee #1 for carefully reading the manuscript and for their constructive comments. We appreciate the positive overall assessment and address each point below.
Comments
L13: I think the authors mean "cost function" here.
We agree and propose to replace "loss function" with "cost function" at L13 for terminological consistency with the rest of the manuscript.
L14: I suggest to change "calibration metric" to "cost function" here to make it clear.
We agree here as well and propose to change "calibration metric" to "cost function" at L14.
L25–26: Suggestion "In recent years, these regions have witnessed multiple billion-euro catastrophe events, including the February 2022 European..."
We will revise the sentence as suggested by the referee.
L36–38: The entire sentence must be reviewed. Note I normally reinsurance contracts run max 2 years, so climate change plays not a big role in business.
We agree with the referee. The sentence implicitly frames climate change as a direct driver of near-term insurance pricing decisions, which overstates its relevance given that reinsurance contracts are typically renewed annually or biennially and are priced primarily on recent loss experience rather than long-horizon climate trends. We propose to rephrase the sentence: “As multifaceted impacts of climate change compound on top of these projection uncertainties, robust risk quantification in the form of reliable estimation and projection of windstorm damage becomes imperative to accurately inform risk management strategies and climate adaptation planning (Pinto et al., 2007; Donat et al., 2011; Pinto et al., 2019).”
L59: "Therefore, robust calibration and validation is necessary, ..."
We will revise the sentence as advised.
L60: Which parameterizations are meant here, maybe avoid using this wording as people might think about parameterizations in dynamical models, I guess you mean statistical models here, such as the damage model used.
We agree that the word "parameterizations" here could be misread as referring to physical/dynamical model parameterizations. We propose to replace the word ‘parameterize’ with ‘calibration’ to correct the ambiguity flagged by the referee.
L93: The publications at the end of the sentence are misleading as they suggest that they refer to the dataset, but they only describe the storms, so please cite the dataset here.
We thank the referee for pointing this out. We agree that Ulbrich et al. (2001) and Fink et al. (2009) are meteorological/synoptic descriptions of the named storms, not sources for the PERILS loss data itself. We propose to replace these with a direct citation to the 2011 PERILS database update that added these specific storms, while retaining the meteorological papers as supplementary references. Rephrase - “In 2011, the dataset was retrospectively expanded to include loss estimates for five major European storm events from 1999 onwards: Anatol, Lothar, Martin, Jeanett, and Kyrill (PERILS AG, 2011; see also Ulbrich et al., 2001; Fink et al., 2009 for storm descriptions).”
L101–102: How is inflation handled in these data sets?
We thank the referee for raising this point. Reported losses are provided in nominal (as-reported) currency values for each event's occurrence year. To ensure consistency with our exposure data, which is anchored to reference year 2018, we inflation-adjust all reported losses to 2018 equivalent values using country-specific CPI prior to calibration. We will add this information in the revised manuscript.
L110: This sentence can be removed as it is described in 2.3.
We agree and propose to remove the sentence as advised.
L133: Please name the subsection "Calibration", this is sufficient.
We propose to shorten the subsection heading to "Calibration."
Section 2.3: The number of events n is rather small and should be mentioned in this section.
We agree that this is a limiting factor and must be stated within this section. We propose to add this information in the revised manuscript.
189–191: The sentence is not clear. It is also not clear why any calibrated MDD_max value is a generalized parameter. You need to give an explanation.
We thank the referee for pointing this out. Based on Referee #2’s recommendations, we have conducted a fresh analysis of LOO. Please see our response to comment M5 of Referee #2.
L204: "who estimated"
Based on Referee #2 minor comment number 8, we have updated this sentence.
Results: Maybe this is the only more major comment. I think the authors need to discuss the effect of the rather small number of events on the calibrations much more. It is mentioned in the last section, but I think the reader needs in the results section information about which region the calibration is trustworthy and in which it is at the limit.
We agree with the referee that this deserves more attention in the Results section itself. We propose to mention the limitations of a small number of events available for certain countries in the Results section. In particular, which countries' results should be interpreted more cautiously due to a limited sample size.
L234: "pulling down modeled damages" this is not visible in Fig 2c, it looks identical.
Thank you for rightly pointing out this mistake in the phrasing. We propose rewriting the sentence in the revised manuscript to provide an accurate description of Fig 2c.
Fig. 3: Too small labels and axes labels. Maybe split the figure into two (m, n as separate figure).
We agree with this assessment and propose increasing the label font and axes labels. We propose keeping the figure intact and improving the overall readability by utilising the extra space between scatter plots.
Fig. 5: Too small labels and axes labels. You also change the use of the colors, In Fig 3 orange is used for RMSLE, here for RMSE, please adjust.
We thank the reviewer for rightly pointing this out. We propose increasing the label font and standardising the colour scheme across all figures throughout the manuscript.
Citation: https://doi.org/10.5194/egusphere-2026-3099-AC1
-
AC1: 'Reply on RC1', Aditya Narayan Mishra, 23 Sep 2026
-
RC2: 'Comment on egusphere-2026-3099', Anonymous Referee #2, 29 Jul 2026
The manuscript recalibrates the Schwierz et al. (2010) windstorm vulnerability function within CLIMADA, against PERILS insured-loss data (1999-2024) for 12 countries, separately under RMSE and RMSLE cost functions, with C3S-EWS gust footprints as hazard and LitPop as exposure. It reports pan-European and country-specific calibrated parameters, event-level scatter, loss-frequency curves, and Unsequa-based uncertainty distributions of normalized average annual impact (AAI), and concludes that (i) the default model systematically underestimates losses across Europe, (ii) a single pan-European model cannot capture country-level vulnerability differences, and (iii) the choice of cost function fundamentally shapes the calibrated model and derived risk metrics.
The aim of the study, a transparent recalibration of a windstorm vulnerability function under two cost functions, is in line with the scope of NHESS and worthwhile for the community. However, my review identified the following issues and limitations, and I therefore recommend a major revision.
Beyond the specific comments below, I want to be transparent that, in my assessment, the central product of this study, the pan-European and country-specific calibrated damage models, will remain highly uncertain even if all comments are fully addressed. The calibration rests on only 3-15 events per country (103 event-country observations in total), drawn from a censored and temporally inhomogeneous loss record; the loss data are aggregated to country-event level, so each fitted multiplier inseparably absorbs hazard biases, exposure-valuation error, and insurance penetration and market coverage alongside actual vulnerability, and is therefore strictly meaningful only within the exact C3S-EWS + LitPop + insured-loss configuration used here; and the fixed curve shape, which saturates at 60 m/s, leaves the damaging tail essentially unconstrained by the data. These are limitations of the data class rather than of the authors’ execution, but they cap the reliability that this type of product can attain.
In that light, the current frontier in this field lies in the resolution, quality and feature extend of the underlying data rather than in recalibration of aggregated curves or in model sophistication. For example, using approximately 9.4 million policy-level records for about one million residential buildings across nine storms in the Netherlands, van Ederen et al. (2026) show that high-resolution data, not model flexibility, drive accurate predictions of European winter storm damage, reducing storm-level prediction errors by about 90% relative to generic European benchmark models (https://doi.org/10.5281/zenodo.20556499). I would encourage the authors to position their contribution relative to such data-driven developments, and to temper the stated reliability and applicability of the recalibrated models accordingly.
Major comments
M1. MDD is not used consistently with CLIMADA
In CLIMADA the damage ratio applied to exposure is the mean damage ratio (MDR), which equals MDD × PAA, where MDD is the mean damage degree of affected assets and PAA is the percentage of assets affected: Impact Functions — CLIMADA 6.1.0 documentation It appears that the paper uses MDD as if it were the MDR, and does not elaborate on PAA or MDR. The manuscript does not provide enough detail to decide if PAA was also scaled, or a scaled MDD was used directly to calculate losses. Recalibrating only the MDD remains valid, but the implementation should be stated unambiguously, because the calibrated MDD is only meaningful in the configuration in which it was fitted: if the calibration chain applied the default PAA (as the standard CLIMADA impact calculation does), the calibrated MDD must also be applied in combination with that PAA, whereas if losses were computed from the MDD alone, subsequently combining the calibrated MDD with the PAA would underestimate damage (by up to an order of magnitude below 60 m/s, where PAA < 1). Please describe the implemented damage model consistently with the CLIMADA terminology and explicitly state if PAA is left out.
Also, the manuscript’s description of the CLIMADA MDD is wrong on several occasions, for instance:
“Negligible damage below approximately 15 m/s” (L127): the function is zero below 20 m/s.
“Reaches approximately 50% of maximum damage at moderate gust speeds (35-40 m/s)” (L128): MDD reaches ~10% of its maximum at 35 m/s and ~20% at 40 m/s; 50% is reached at ~50 m/s.
“The transition region between 30-50 m/s where the slope is steepest” (L130-131): the steepest MDD segment is 55-60 m/s.
“The model’s default state has MDDmax of 0.0387”: Eq. (1) (≈L141) states MDDmaxdefault = 0.0373. With the correct default, the pan-European RMSE calibration (0.0379) raises the default amplitude by ≈1.6%, with the erroneous 0.0387, the same result reads as a slight decrease.
M2. Figs. 3m-n and Fig. 6 appear mutually inconsistent under the stated methodology
Because event losses are exactly linear in θ, and the Unsequa perturbations are identical in distribution for the RMSE- and RMSLE-calibrated runs, shouldn’t the ratio of the two ensemble means of normalized AAI for a given country equal the ratio of the two calibrated parameters (up to Monte-Carlo error)? For instance for Belgium: θ ratio RMSLE/RMSE (Fig. 3) ~ 0.16/0.02 = 8 while AAI ratio RMSLE/RMSE (Fig. 6) ~2.73/0.6 = 4.55 which is a ~ +80% mismatch. The likeliest explanations are that Fig. 6 was produced with a different (earlier, shrunken, or otherwise transformed) parameter set than displayed in Fig. 3, or that its normalization differs from the caption. Please recompute or clarify.
In general, it would also help interpreting the results to complement the figures (e.g. Figs. 2 and 3) with tables reporting their numerical values and, for instance, % differences between modelled and observed damage.
M3. The simplified curve in Schwierz et al. (2010) was derived primarily from insured residential-sector damage but is applied to other asset types as well
The manuscript applies this model to LitPop exposure, such that it effectively transfers a residential insured-property wind-damage relationship to all kinds of exposure, such as commercial, industrial or infrastructure assets. Therefore, the fitted (country) multiplier will (also) absorb sectoral composition. Sector-specific damage models would be preferable to control for their vulnerability characteristics. Hence, the recalibrated damage models are likely too aggregated to successfully transfer to another period or portfolio whose residential/commercial mix differs. The paper should be clear about this limitation.
M4. The headline claim “the default model systematically underestimates losses across Europe” is not fully supported and exposure assumption dependent
Pan-European RMSE calibration indicates that there is no pan-European underestimation to correct. In 6 of 12 countries (BEL 0.020, NLD 0.019, NOR 0.021, GBR 0.030, IRL 0.031, LUX 0.034, vs default 0.0373), the country-level RMSE calibration moves the parameter below the default: for these countries the squared-error view says the default over-predicts.
The abstract, L249-250, L272-274, and the first sentence of the Conclusions should be reformulated to make the claim loss-range- and country-specific.
“The default model underestimates losses” conclusion is conditional on the LitPop exposure: assuming higher exposure values would result in a lower recalibration, which could reach up to the point that the default model overestimates losses. Hence, this is not an inherent flaw of the Schwierz et al. (2010) model, but rather a result of the exposure (and hazard) input assumptions. Readers should (for that reason) not be discouraged from using the Schwierz et al. (2010) model. The Schwierz et al. (2010) model relied on insured values at risk, which differs from the LitPop exposure proxy. The paper should be more clear about the dependence on the exposure assumption.
M5. The LOOCV analysis cannot support the conclusions drawn from it
Section 2.4 reports a single statistic (correlation 0.97 between LOO predictions and in-sample fits) and concludes that the calibration is “robust and not driven by individual extreme events” and that “any calibrated MDDmax value is a generalized parameter suitable for future event prediction”. Predictive skill is asserted against the wrong reference, as correlating LOO predictions with in-sample predictions says nothing about agreement with observed losses.
M6. The AAI evaluation appears in-sample, which would not support prediction accuracy conclusions
AAInorm is computed over the same 1999-2024 events used to fit θ. Statements such as RMSE giving “a stable estimate of the expected annual loss” (DEU 0.98, CHE ~1.06, AUT ~0.95) present in-sample agreement as accuracy. Out-of-sample AAI assessment would make these statements meaningful.
M7. The practical guidance is contradicted by the manuscript’s own results
The Discussion recommends RMSLE calibration for “applications focused on pricing attritional losses or estimating expected annual damages” (L335-337). But Fig. 6 shows RMSLE-calibrated models overestimate AAI in 11 of 12 countries, by factors of 1.3-3.0 (BEL 2.73, NLD 3.01, NOR 2.94, LUX 2.28, DEU 2.11…), while the same section elsewhere concedes RMSLE “frequently overestimates in aggregate loss assessments” (L300-301). Recommending the systematically AAI-overestimating calibration for expected-annual-damage estimation is a direct internal contradiction. RMSLE is defensible for event-level relative accuracy in the attritional range (Fig. 5 supports this); it is not defensible for AAI.
M8. The EM-DAT comparison conflates loss definitions
PERILS records insured losses; EM-DAT records total economic damages (and its entries are not adjusted here in any stated way). Total economic windstorm losses are often higher than insured losses, which may explain why an EM-DAT-calibrated amplitude is about 2-3× that of the PERILS. The manuscript instead frames the divergence as “uncertainty introduced by a much smaller sample size” (L220-221) and tests it against PERILS subsampling variability. The difference in loss definitions are never mentioned, leaving the reader with the impression that EM-DAT is merely noisy rather than measuring a different quantity.
M9. Catalog inhomogeneity and indexation handling are unaddressed
Per Section 2.1, PERILS contains all qualifying events only since 2009; 1999-2008 is represented by exactly five retrospectively added major storms, and the qualification threshold changed from EUR 200M to 300/500M in September 2022. The 1999-2024 record is therefore censored non-uniformly in time: the first decade contains only the largest storms. This biases the empirical “occurrence frequency” axis of Fig. 5 (return periods of moderate events are overstated), deflates AAIrep relative to a true above-threshold AAL, and weights the calibration toward whichever epochs a country’s events happen to fall in. The Discussion acknowledges threshold censoring generically (L341-343) but not the 1999-2008 inhomogeneity. At minimum this should be analyzed (e.g., recalibrating on 2009-2024 only as a sensitivity test).
The manuscript never states whether PERILS losses are used as originally reported (nominal, at-event-date market conditions) or as-if-adjusted to present-day exposure, nor which LitPop vintage is used. This must be specified, and if unadjusted losses were used, corrected.
M10. Interpreting calibrated amplitudes as “vulnerability profiles” / “the underlying wind-damage relationship” overstates what θ can identify
θ is estimated from insured losses against a total-asset exposure (LitPop). It therefore absorbs, inseparably: physical vulnerability, insurance penetration and PERILS market-coverage share (which vary strongly across these 12 countries) and systematic hazard-intensity errors. The Discussion’s single clause “market or claims conditions” (L314) is not enough to license the abstract’s “reflect spatial heterogeneities in vulnerability”, L247-248’s “reflects the underlying wind-damage relationship observed in each country”, or the Conclusions’ framing. Since only the amplitude is calibrated with a fixed curve shape, genuine cross-country differences in threshold or steepness cannot be accounted for. Relatedly, no per-country parameter uncertainty is ever shown (the subsampling band exists only pan-European), with n = 3-15 events the sampling uncertainty is very high.
M11. Descriptive statistics are missing
The manuscript reports no descriptive statistics of the calibration data. Without these, readers cannot judge sample representativeness, the leverage of individual events on the calibrated parameters, or the severity range over which the models are actually constrained.
Minor comments
- Figure cross-references are systematically wrong in Sections 3-4. L289-292 cite Fig. 3e/3i/3q/3u for values that are in Fig. 6e/6i/6q/6u (Fig. 3 has no panels q-u); L310 cite Fig. 5 for the parameter panels of Fig. 3; L275-276 cite Fig. 3c/3d/3h within a Fig. 5 discussion (the panel letters coincide, but the figure number should be 5); L280-281 similarly cite Fig. 3a/3j/3k/3l for exceedance-curve behavior shown in Fig. 5.
- Transposed values. L288-289 quote RMSE mean AAI “0.98, 0.95, and 1.06” for Germany, Switzerland, Austria; Fig. 6 gives DEU 0.98, CHE 1.06, AUT 0.95 (CHE and AUT swapped).
- Norway misclassified. L244 claims IRL, GBR, NOR have RMSLE MDDmax < 0.04 (“substantially flatter”); Fig. 3m shows NOR at ≈0.044-0.045, above both 0.04 and the default (0.0373), and the Fig. 4 RMSLE map correspondingly shows Norway positive (≈ +18%). Only IRL (0.013) and GBR (0.034) are below default.
- “British Isles show small negative changes” (L263, RMSLE map): Great Britain yes (≈ −9%); Ireland no (≈ −65%, the darkest negative cell on the map).
- “a 25-year catalog” (L271) vs “1999-2024, 26 years” (Figs. 5 and 6 captions): inconsistent; 1999-2024 inclusive is 26 years.
- 5 x-axis is mislabeled. The axis labeled “Occurence Frequency (1/years)” increases toward rarer, larger events and spans ≈26/n to ≈26; it is the empirical return period in years, not a frequency in 1/years (also: “Occurrence” is misspelled). The caption and L270-272 should be corrected accordingly.
- Prior article. CLIMADA itself deploys the damage model from Welker et al. (2021), a multiplicative RMSE recalibration of the same Schwierz function against Swiss insurance data, methodologically the direct predecessor of this study. Only the 2020 discussion version is cited, in passing. The novelty framing (multi-country PERILS calibration; RMSE-vs-RMSLE comparison) would be strengthened, not weakened, by engaging with it.
- ±20% vulnerability-uncertainty attribution (L203-205): please point to where Prahl et al. (2012) and Severino et al. (2024) establish ±20% as a bound for this parameter, and specify what exactly is “consistent with” Severino et al. (2024) at L142-143 (their calibration strategy differs in implementation details).
- Koks and Haer (2020) did not employ the Schwierz et al. (2010) damage model (L52-53)
- Typographical: “Infact” (L25); “optimzation” (L309); missing space “(>60 m/s)(see Fig. 2a)” (L129); “formal analysis,visualization”
Citation: https://doi.org/10.5194/egusphere-2026-3099-RC2 -
AC3: 'Reply on RC2', Aditya Narayan Mishra, 23 Sep 2026
The manuscript recalibrates the Schwierz et al. (2010) windstorm vulnerability function within CLIMADA, against PERILS insured-loss data (1999-2024) for 12 countries, separately under RMSE and RMSLE cost functions, with C3S-EWS gust footprints as hazard and LitPop as exposure. It reports pan-European and country-specific calibrated parameters, event-level scatter, loss-frequency curves, and Unsequa-based uncertainty distributions of normalized average annual impact (AAI), and concludes that (i) the default model systematically underestimates losses across Europe, (ii) a single pan-European model cannot capture country-level vulnerability differences, and (iii) the choice of cost function fundamentally shapes the calibrated model and derived risk metrics.
The aim of the study, a transparent recalibration of a windstorm vulnerability function under two cost functions, is in line with the scope of NHESS and worthwhile for the community. However, my review identified the following issues and limitations, and I therefore recommend a major revision.
We thank Referee #2 for their careful reading of the manuscript and the thorough assessment of our work. We genuinely believe that the Referee’s valuable comments, suggestions and corrections will be very helpful in improving the clarity and rigour of the study. We address each of their concerns in detail below.
Beyond the specific comments below, I want to be transparent that, in my assessment, the central product of this study, the pan-European and country-specific calibrated damage models, will remain highly uncertain even if all comments are fully addressed. The calibration rests on only 3-15 events per country (103 event-country observations in total), drawn from a censored and temporally inhomogeneous loss record; the loss data are aggregated to country-event level, so each fitted multiplier inseparably absorbs hazard biases, exposure-valuation error, and insurance penetration and market coverage alongside actual vulnerability, and is therefore strictly meaningful only within the exact C3S-EWS + LitPop + insured-loss configuration used here; and the fixed curve shape, which saturates at 60 m/s, leaves the damaging tail essentially unconstrained by the data. These are limitations of the data class rather than of the authors’ execution, but they cap the reliability that this type of product can attain.
In that light, the current frontier in this field lies in the resolution, quality and feature extend of the underlying data rather than in recalibration of aggregated curves or in model sophistication. For example, using approximately 9.4 million policy-level records for about one million residential buildings across nine storms in the Netherlands, van Ederen et al. (2026) show that high-resolution data, not model flexibility, drive accurate predictions of European winter storm damage, reducing storm-level prediction errors by about 90% relative to generic European benchmark models (https://doi.org/10.5281/zenodo.20556499). I would encourage the authors to position their contribution relative to such data-driven developments, and to temper the stated reliability and applicability of the recalibrated models accordingly.
We agree with the referee's main point that the reliability ceiling of any country-event-aggregated calibration is fundamentally bounded by the resolution and volume of the underlying loss data, and no methodological refinement can substitute for richer, disaggregated ground-truth data. Importantly, the purpose of this study is not to develop a new damage-model functional form or to claim that the fixed Schwierz et al. (2010) curve shape, including its saturation at high wind speeds, is universally optimal. Rather, we deliberately retain this widely used and established formulation in order to assess how far its performance can be improved through transparent recalibration against a consistent, high-quality insurance-loss dataset, and to isolate the consequences of choosing RMSE versus RMSLE as the cost function. This provides a directly interpretable update of an existing model that remains widely applied in European climate-risk studies.We thank the referee for highlighting the preprint of the study from van Ederen et al. (2026). The fact that the type of data used by Ederen et al. (2026) is seldom available at continental scale, underscores that in the current data-constrained setting the calibration process remains important. It is thus still relevant to ask what can be learned from the industry-aggregated data currently available consistently across 12 European countries, even when high-resolution national data may be available for some of them. We propose to add a discussion whereby we explicitly identify high-resolution, harmonized multi-country loss and exposure data as one of the key directions for future improvement, and will revise the scope and applicability of our recalibrated models more precisely, consistent with the limitations outlined above.
M1. MDD is not used consistently with CLIMADA
In CLIMADA the damage ratio applied to exposure is the mean damage ratio (MDR), which equals MDD × PAA, where MDD is the mean damage degree of affected assets and PAA is the percentage of assets affected: Impact Functions — CLIMADA 6.1.0 documentation It appears that the paper uses MDD as if it were the MDR, and does not elaborate on PAA or MDR. The manuscript does not provide enough detail to decide if PAA was also scaled, or a scaled MDD was used directly to calculate losses. Recalibrating only the MDD remains valid, but the implementation should be stated unambiguously, because the calibrated MDD is only meaningful in the configuration in which it was fitted: if the calibration chain applied the default PAA (as the standard CLIMADA impact calculation does), the calibrated MDD must also be applied in combination with that PAA, whereas if losses were computed from the MDD alone, subsequently combining the calibrated MDD with the PAA would underestimate damage (by up to an order of magnitude below 60 m/s, where PAA < 1). Please describe the implemented damage model consistently with the CLIMADA terminology and explicitly state if PAA is left out.
We thank the referee for this comment and agree that the manuscript did not elaborate on the PAA component of the damage model. We clarify that we implemented the Schwierz et al. (2010) model as coded in CLIMADA, where losses (both default and calibrated) are computed using a combination of MDD and PAA. We recalibrated only the MDD component and applied it in combination with the default PAA. We propose to add this information in the revised manuscript.
Also, the manuscript’s description of the CLIMADA MDD is wrong on several occasions, for instance:
“Negligible damage below approximately 15 m/s” (L127): the function is zero below 20 m/s.
“Reaches approximately 50% of maximum damage at moderate gust speeds (35-40 m/s)” (L128): MDD reaches ~10% of its maximum at 35 m/s and ~20% at 40 m/s; 50% is reached at ~50 m/s.
“The transition region between 30-50 m/s where the slope is steepest” (L130-131): the steepest MDD segment is 55-60 m/s.
We thank the referee for their careful observation. We propose to make the necessary corrections in the referred sentences.
“The model’s default state has MDDmax of 0.0387”: Eq. (1) (≈L141) states MDDmaxdefault = 0.0373. With the correct default, the pan-European RMSE calibration (0.0379) raises the default amplitude by ≈1.6%, with the erroneous 0.0387, the same result reads as a slight decrease.
We thank the referee again for pointing this out. We will correct this in L216, and specify the correct value which is 0.0387.
M2. Figs. 3m-n and Fig. 6 appear mutually inconsistent under the stated methodology
Because event losses are exactly linear in θ, and the Unsequa perturbations are identical in distribution for the RMSE- and RMSLE-calibrated runs, shouldn’t the ratio of the two ensemble means of normalized AAI for a given country equal the ratio of the two calibrated parameters (up to Monte-Carlo error)? For instance for Belgium: θ ratio RMSLE/RMSE (Fig. 3) ~ 0.16/0.02 = 8 while AAI ratio RMSLE/RMSE (Fig. 6) ~2.73/0.6 = 4.55 which is a ~ +80% mismatch. The likeliest explanations are that Fig. 6 was produced with a different (earlier, shrunken, or otherwise transformed) parameter set than displayed in Fig. 3, or that its normalization differs from the caption. Please recompute or clarify.
*In general, it would also help interpreting the results to complement the figures (e.g. Figs. 2 and 3) with tables reporting their numerical values and, for instance, % differences between modelled and observed damage.
We thank the referee for pointing out this inconsistency. Upon investigating the code used to generate Figure 6, we identified a bug responsible for the discrepancy. This has now been corrected, and we propose to include the updated Figure 6 in the revised manuscript (attached). Kindly note that we’ve revised the colour palette to ensure colourblind-friendly visualisation and to maintain consistent colour coding across all figures in the manuscript.
This resolves the discrepancy mentioned by the referee. Additionally, we agree with the referee that a table reporting numerical values and/or percentage differences between modelled and reported damage could aid interpretation of the results. However, such a table poses a risk of recovering the underlying proprietary PERILS loss values, which our data license prohibits disclosing.
M3. The simplified curve in Schwierz et al. (2010) was derived primarily from insured residential-sector damage but is applied to other asset types as well
The manuscript applies this model to LitPop exposure, such that it effectively transfers a residential insured-property wind-damage relationship to all kinds of exposure, such as commercial, industrial or infrastructure assets. Therefore, the fitted (country) multiplier will (also) absorb sectoral composition. Sector-specific damage models would be preferable to control for their vulnerability characteristics. Hence, the recalibrated damage models are likely too aggregated to successfully transfer to another period or portfolio whose residential/commercial mix differs. The paper should be clear about this limitation.
We thank the referee for raising this concern. The comment relates to a similar concern raised by reviewer #3. We calibrate the model specifically for building damages across the residential and commercial sectors, consistent with how the original model has itself been applied in the study. Schwierz et al. (2010) note that, though the model is derived from the residential sector in their simplified formulation, it is used for all risk types - residential, commercial, industrial, and agricultural. This same practice is reflected in subsequent CLIMADA-based studies, including Röösli et al. (2021) and Severino et al. (2024), both of which apply the model to multi-sector building losses. We agree with the reviewer, however, that this approach has a limitation. Because the multiplier is a single scalar, it can only correct the average level of damage for the specific storm set and sectoral mix present in our calibration window. Consequently, if the model is applied to a future period or portfolio whose sectoral composition differs materially from that of the calibration sample, the fitted multiplier may not transfer well, as it implicitly encodes the sectoral mix of the calibration data alongside the true hazard-vulnerability relationship. We propose to add this limitation in the discussion section.
M4. The headline claim “the default model systematically underestimates losses across Europe” is not fully supported and exposure assumption dependent
Pan-European RMSE calibration indicates that there is no pan-European underestimation to correct. In 6 of 12 countries (BEL 0.020, NLD 0.019, NOR 0.021, GBR 0.030, IRL 0.031, LUX 0.034, vs default 0.0373), the country-level RMSE calibration moves the parameter below the default: for these countries the squared-error view says the default over-predicts.
The abstract, L249-250, L272-274, and the first sentence of the Conclusions should be reformulated to make the claim loss-range- and country-specific.
“The default model underestimates losses” conclusion is conditional on the LitPop exposure: assuming higher exposure values would result in a lower recalibration, which could reach up to the point that the default model overestimates losses. Hence, this is not an inherent flaw of the Schwierz et al. (2010) model, but rather a result of the exposure (and hazard) input assumptions. Readers should (for that reason) not be discouraged from using the Schwierz et al. (2010) model. The Schwierz et al. (2010) model relied on insured values at risk, which differs from the LitPop exposure proxy. The paper should be more clear about the dependence on the exposure assumption.
We thank the referee for this important comment. We agree that the calibrated MDDmax values cannot be interpreted as estimates of physical building vulnerability alone. In the present modelling chain, modeled impact is jointly determined by the C3S-EWS wind footprints, the LitPop exposure representation, and the impact function (damage model). Consequently, the calibration parameter necessarily absorbs residual differences associated with: (i) biases or uncertainty in modeled wind gusts; (ii) uncertainty in the magnitude and spatial distribution of LitPop asset values; (iii) differences between total economic asset values represented by LitPop and the insured values at risk underlying PERILS losses; and (iv) country-specific differences in insurance penetration, claims practices, policy conditions, and market coverage. A different assumed exposure value would, all else being equal, require a different fitted MDDmax to reproduce the same reported insured loss. We will make this identifiability limitation explicit.
At the same time, we emphasise that this does not make the exercise arbitrary or invalidate the comparison. The purpose of our study is not to establish that the Schwierz et al. (2010) model is intrinsically deficient, nor to claim that our fitted functions universally replace it. Rather, we evaluate how this widely used open-source model performs and how it can be recalibrated within a specified, internally consistent, Europe-wide application configuration: C3S-EWS hazard footprints, LitPop exposure, and PERILS insured-loss data. The fitted parameter should therefore be interpreted as effective for this specific hazard-exposure-loss configuration, rather than as universal corrections to the original Schwierz et al. (2010) model.
We agree with the referee that readers should not be discouraged from using the original Schwierz et al. (2010) model. Our revised interpretation will be that the results demonstrate the importance of evaluating the transferability of this default model for the specific exposure dataset, hazard dataset, and cost function under consideration.
M5. The LOOCV analysis cannot support the conclusions drawn from it
Section 2.4 reports a single statistic (correlation 0.97 between LOO predictions and in-sample fits) and concludes that the calibration is “robust and not driven by individual extreme events” and that “any calibrated MDDmax value is a generalized parameter suitable for future event prediction”. Predictive skill is asserted against the wrong reference, as correlating LOO predictions with in-sample predictions says nothing about agreement with observed losses.
We thank the referee for this important comment, and we agree with the critique. We agree that the previously reported correlation of 0.97 between leave-one-out predictions and in-sample fitted losses is not the appropriate metric for establishing out-of-sample predictive performance. We have therefore conducted a fresh analysis and propose revising the manuscript to report a direct comparison between LOO predictions and the reported losses of the held-out country-event records.
We find that for the RMSLE-calibrated model, the Pearson correlation between log-transformed observed and LOO-predicted losses is 0.804 and for the RMSE-calibrated model it is 0.807. This result indicates that both calibration approaches retain a meaningful association with the magnitude of held-out losses, while also exhibiting substantial event-level residual uncertainty, as expected for winter-storm loss modelling across heterogeneous countries and events.
We also propose to revise the interpretation of parameter stability. Rather than inferring general suitability for future-event prediction, we now state that the fitted parameter is stable under omission of individual observations, i.e., over the 103 LOO folds, the RMSLE-calibrated MDDmax ranged from 0.0690 to 0.0730 and the RMSE-calibrated MDDmax from 0.0298 to 0.0350. We propose to remove the earlier statement that any calibrated MDDmax is necessarily suitable for future-event prediction and now describe the LOO results as evidence of out-of-sample consistency subject to the inherent uncertainty of event-level insured-loss modelling.
M6. The AAI evaluation appears in-sample, which would not support prediction accuracy conclusions
AAInorm is computed over the same 1999-2024 events used to fit θ. Statements such as RMSE giving “a stable estimate of the expected annual loss” (DEU 0.98, CHE ~1.06, AUT ~0.95) present in-sample agreement as accuracy. Out-of-sample AAI assessment would make these statements meaningful.
We acknowledge the concern raised by the referee. We would like to clarify that our intent with the UNSEQUA/AAI experiment (Section 2.4) was to specifically propagate realistic input uncertainties in hazard, exposure, and the vulnerability parameter through the calibrated model and to compare, in relative terms, how the RMSE and RMSLE calibrated models each respond to that uncertainty. This would be an analysis for internal consistency and bias-under-uncertainty, not a test of predictive skill on unseen events or years. We agree that a genuine assessment of predictive accuracy for AAI would require an out-of-sample framework, which is beyond the scope of the current dataset given the limited number of events per country. We propose to revise statements that appear to communicate an evaluation of "accuracy" using in-sample data.
M7. The practical guidance is contradicted by the manuscript’s own results
The Discussion recommends RMSLE calibration for “applications focused on pricing attritional losses or estimating expected annual damages” (L335-337). But Fig. 6 shows RMSLE-calibrated models overestimate AAI in 11 of 12 countries, by factors of 1.3-3.0 (BEL 2.73, NLD 3.01, NOR 2.94, LUX 2.28, DEU 2.11…), while the same section elsewhere concedes RMSLE “frequently overestimates in aggregate loss assessments” (L300-301). Recommending the systematically AAI-overestimating calibration for expected-annual-damage estimation is a direct internal contradiction. RMSLE is defensible for event-level relative accuracy in the attritional range (Fig. 5 supports this); it is not defensible for AAI.
We note that the specific overestimation factors cited in the referee's comment correspond to an earlier version of Fig. 6; we have since recalculated this figure based on the referee’s comment M2, and the updated per-country values differ from those originally reported. However, referee's underlying critique here still remains fully valid in our opinion and we propose to correct the recommendation in the revised manuscript to align with the evidence presented across the two figures: RMSLE calibration recommended specifically for applications requiring accurate event-level, relative/attritional-loss reproduction across the full frequency spectrum (Fig. 5), while RMSE calibration is recommended for applications requiring robust, aggregate expected-annual-damage (AAI) estimates, given its narrower and less-biased AAI distributions (Fig. 6). We thank the referee for prompting this correction, which improves the practical guidance offered by the study.
M8. The EM-DAT comparison conflates loss definitions
PERILS records insured losses; EM-DAT records total economic damages (and its entries are not adjusted here in any stated way). Total economic windstorm losses are often higher than insured losses, which may explain why an EM-DAT-calibrated amplitude is about 2-3× that of the PERILS. The manuscript instead frames the divergence as “uncertainty introduced by a much smaller sample size” (L220-221) and tests it against PERILS subsampling variability. The difference in loss definitions are never mentioned, leaving the reader with the impression that EM-DAT is merely noisy rather than measuring a different quantity.
We thank the referee for raising this point. To clarify, EM-DAT loss figures used in this study are drawn specifically from EM-DAT's insured-loss field, not its total economic damage field, and were adjusted for inflation prior to comparison with PERILS. The two databases are therefore intended to measure the same underlying quantity (insured losses), and the observed divergence is not attributable to a mismatch between insured and total economic loss definitions. However, we recognise that neither the specific EM-DAT field used nor the inflation adjustment applied was stated in the manuscript, and we agree this omission left room for the loss-definition explanation the referee proposed. We propose to revise the Data and Methods section to explicitly state the EM-DAT field used and the inflation-adjustment procedure. The offset is more plausibly explained, at least in part, by structural differences between the two databases, including EM-DAT's reliance on secondary and media-derived reporting for some entries versus PERILS' primary insurer-sourced data, and differing event-inclusion thresholds and matching criteria between the two sources, rather than sampling noise alone. We propose to add these relevant details to the revised manuscript.
M9. Catalog inhomogeneity and indexation handling are unaddressed
Per Section 2.1, PERILS contains all qualifying events only since 2009; 1999-2008 is represented by exactly five retrospectively added major storms, and the qualification threshold changed from EUR 200M to 300/500M in September 2022. The 1999-2024 record is therefore censored non-uniformly in time: the first decade contains only the largest storms. This biases the empirical “occurrence frequency” axis of Fig. 5 (return periods of moderate events are overstated), deflates AAIrep relative to a true above-threshold AAL, and weights the calibration toward whichever epochs a country’s events happen to fall in. The Discussion acknowledges threshold censoring generically (L341-343) but not the 1999-2008 inhomogeneity. At minimum this should be analyzed (e.g., recalibrating on 2009-2024 only as a sensitivity test).
The manuscript never states whether PERILS losses are used as originally reported (nominal, at-event-date market conditions) or as-if-adjusted to present-day exposure, nor which LitPop vintage is used. This must be specified, and if unadjusted losses were used, corrected.
We thank the referee for this detailed comment and acknowledge their concern. We clarify, however, that the intended role of Fig. 5 is not to infer the climatological occurrence frequency or an unbiased actuarial return period for European windstorm losses. For each country and loss series, events are ordered by loss magnitude and assigned a catalog-based empirical exceedance frequency, calculated from the number of available catalog events with loss equal to or exceeding the plotted loss. Its primary purpose is comparative: it shows, on the same country-specific event set, how the default, RMSE-calibrated, and RMSLE-calibrated models reproduce the observed PERILS loss distribution across ranked loss magnitudes, and where the two calibration objectives converge or diverge relative to the observed losses. The conclusion drawn from Fig. 5 therefore concerns the relative behaviour of the calibration objectives within the assembled catalog, rather than a claim about the absolute real-world recurrence of a particular loss level. Nevertheless, we agree that the temporal inhomogeneity means that the catalog-based exceedance frequencies, especially for lower-to-moderate losses, should not be interpreted as homogeneous estimates of underlying occurrence rates. We propose to address this in the discussion section of the revised manuscript.
We retain the full matched catalog in the primary analysis because windstorm losses are scarce at the country-event level. The 103 observations used across Europe represent country–storm observations rather than independent Europe-wide storms: a single large storm may contribute records for several affected countries, each with a distinct country-level insured loss, exposure representation, and hazard footprint. The five retrospectively included major storms are therefore important sources of information on damaging historical events across multiple country calibrations. Excluding them would simultaneously remove several country-event observations and, for countries with already sparse samples, would materially reduce the information available to constrain the fitted parameter. Finally, we also clarify that reported loss are provided in nominal (as-reported) currency values for each event's occurrence year. To ensure consistency with our exposure data, which is anchored to reference year 2018, we inflation-adjust all reported losses to 2018 equivalent values using country-specific CPI prior to calibration. We propose to add this information in the revised manuscript.
M10. Interpreting calibrated amplitudes as “vulnerability profiles” / “the underlying wind-damage relationship” overstates what θ can identify
θ is estimated from insured losses against a total-asset exposure (LitPop). It therefore absorbs, inseparably: physical vulnerability, insurance penetration and PERILS market-coverage share (which vary strongly across these 12 countries) and systematic hazard-intensity errors. The Discussion’s single clause “market or claims conditions” (L314) is not enough to license the abstract’s “reflect spatial heterogeneities in vulnerability”, L247-248’s “reflects the underlying wind-damage relationship observed in each country”, or the Conclusions’ framing. Since only the amplitude is calibrated with a fixed curve shape, genuine cross-country differences in threshold or steepness cannot be accounted for. Relatedly, no per-country parameter uncertainty is ever shown (the subsampling band exists only pan-European), with n = 3-15 events the sampling uncertainty is very high.
We thank the reviewer for raising this point. In our framework, θ is an effective calibration parameter describing the amplitude required to reconcile the fixed model shape, C3S-EWS hazard fields, LitPop exposure, and PERILS industry insured losses. Consequently, country-to-country differences in θ may reflect not only physical vulnerability but also differences between insured and total-asset exposure and uncertainties in the hazard and exposure components. We would, however, clarify that PERILS quality-controls and aggregates insurer-contributed data and extrapolates the resulting losses to the industry or market level; therefore, the calibrated parameter is not directly proportional to the undisclosed share of the market represented by the contributing insurers. Nevertheless, differences in insurance penetration and in the correspondence between PERILS insured exposure and LitPop total-asset exposure may influence the resulting calibration. We also clarify that our calibration deliberately does not alter the functional form due to limited number of events available for calibration. We will ensure that these clarifications are explicit in the revised manuscript.
M11. Descriptive statistics are missing
The manuscript reports no descriptive statistics of the calibration data. Without these, readers cannot judge sample representativeness, the leverage of individual events on the calibrated parameters, or the severity range over which the models are actually constrained.
We thank the referee for this comment and agree that descriptive statistics help readers judge sample representativeness, event leverage, and the severity range over which the models are constrained. We can report descriptive statistics for the hazard intensity (wind gust speed); however, we are unable to report descriptive statistics for the loss data themselves due to our PERILS data-license agreement.
To this end, we propose adding a supplementary table that reports, for each country, the number of matched calibration events (n) and descriptive statistics for the corresponding wind gust intensities (mean, standard deviation, minimum, 25th percentile, median, 75th percentile, and maximum). This directly conveys the sample size and severity range over which each country-specific model was calibrated, and allows readers to assess the relative leverage of extreme events. We note that these hazard-based statistics describe the range and concentration of the matched, thresholded event sample used for calibration; they do not, by themselves, establish representativeness of the full windstorm-loss distribution, particularly for events falling below the PERILS reporting threshold (L341–343).
Minor comments
1. Figure cross-references are systematically wrong in Sections 3-4. L289-292 cite Fig. 3e/3i/3q/3u for values that are in Fig. 6e/6i/6q/6u (Fig. 3 has no panels q-u); L310 cite Fig. 5 for the parameter panels of Fig. 3; L275-276 cite Fig. 3c/3d/3h within a Fig. 5 discussion (the panel letters coincide, but the figure number should be 5); L280-281 similarly cite Fig. 3a/3j/3k/3l for exceedance-curve behavior shown in Fig. 5.
We thank the referee for this careful check. We will correct all of these cross-references in the revised manuscript
2.Transposed values. L288-289 quote RMSE mean AAI “0.98, 0.95, and 1.06” for Germany, Switzerland, Austria; Fig. 6 gives DEU 0.98, CHE 1.06, AUT 0.95 (CHE and AUT swapped).
We will correct this in the revised manuscript.
3. Norway misclassified. L244 claims IRL, GBR, NOR have RMSLE MDDmax < 0.04 (“substantially flatter”); Fig. 3m shows NOR at ≈0.044-0.045, above both 0.04 and the default (0.0373), and the Fig. 4 RMSLE map correspondingly shows Norway positive (≈ +18%). Only IRL (0.013) and GBR (0.034) are below default.
We will correct this in the revised manuscript.
4.“British Isles show small negative changes” (L263, RMSLE map): Great Britain yes (≈ −9%); Ireland no (≈ −65%, the darkest negative cell on the map).
We thank the referee for pointing this out. We will use the appropriate country name in the revised manuscript
5. “a 25-year catalog” (L271) vs “1999-2024, 26 years” (Figs. 5 and 6 captions): inconsistent; 1999-2024 inclusive is 26 years.
We will correct this in the revised manuscript
6. 5 x-axis is mislabeled. The axis labeled “Occurence Frequency (1/years)” increases toward rarer, larger events and spans ≈26/n to ≈26; it is the empirical return period in years, not a frequency in 1/years (also: “Occurrence” is misspelled). The caption and L270-272 should be corrected accordingly.
We thank the referee for this careful comment. We propose to rename the label as “Empirical exceedance interval (years)” in the revised manuscript.
7.Prior article. CLIMADA itself deploys the damage model from Welker et al. (2021), a multiplicative RMSE recalibration of the same Schwierz function against Swiss insurance data, methodologically the direct predecessor of this study. Only the 2020 discussion version is cited, in passing. The novelty framing (multi-country PERILS calibration; RMSE-vs-RMSLE comparison) would be strengthened, not weakened, by engaging with it.
We thank the reviewer for this important observation. On checking, we agree that our reference to Welker et al. (2021) was inadvertently pointing to the 2020 discussion-stage manuscript due to a broken/outdated Google Scholar link. We will update this citation to the correct version in the revised manuscript. We will also ensure that Welker et al. (2021)'s contribution is discussed in the revised manuscript within the broader context of model calibration.8. ±20% vulnerability-uncertainty attribution (L203-205): please point to where Prahl et al. (2012) and Severino et al. (2024) establish ±20% as a bound for this parameter, and specify what exactly is “consistent with” Severino et al. (2024) at L142-143 (their calibration strategy differs in implementation details).
We thank the reviewer for identifying this imprecise attribution. Upon re-examination, we agree that Prahl et al. (2012) does not establish ±20% as an uncertainty bound for the MDD parameter, however, Severino et al. (2024) uses ±20% as an assumed range in its uncertainty-analysis design, applying it to the output-MDD scaling factors. This range is presented more generally as a reasonable estimate of uncertainty in European winter-storm damage modelling. We consider ±20% a reasonable modelling-error estimate, and we will ensure that the sentence is revised appropriately to reflect this.
We agree that Severino et al. 2024’s calibration strategy differs in implementation details to ours. The consistency is limited to the fact that both studies recalibrate the model multiplicatively rescaling its MDD output while preserving its functional shape. We will revise the sentence accordingly to make this clear.9. Koks and Haer (2020) did not employ the Schwierz et al. (2010) damage model (L52-53)
We propose removing this citation from the sentence.
10.Typographical: “Infact” (L25); “optimzation” (L309); missing space “(>60 m/s)(see Fig. 2a)” (L129); “formal analysis,visualization”
We thank the referee for pointing these out. We propose to correct these in the revised manuscript.
Citation: https://doi.org/10.5194/egusphere-2026-3099-AC3 - AC4: 'Reply on RC2', Aditya Narayan Mishra, 23 Sep 2026
-
RC3: 'Comment on egusphere-2026-3099', Anonymous Referee #3, 04 Aug 2026
The study calibrates a loss model for winter storms using the CLIMADA framework and data from the PERILS data set from 1999 to 2024. Within CLIMADA, the Schwierz et al. (2010) vulnerability function is used as default. Using two different cost functions for optimization, a parameter for the vulnerability function is tuned to fit the PERILS data. The authors conclude that the default vulnerability function systematically underestimates losses for historical events across Europe, and that a single pan-European model is inadequate to capture the distinct vulnerability profiles inherent to each country.
The manuscript is well written and nicely fits to the scope of NHESS. There are a couple of comments and minor issues, therefore I suggest minor revision before publishing.
General
The PERILS data set includes a moderate number of events if it comes to country level. I don’t suggest to use a different data set, but for the reader it has to be very clear to know about the limits of the study with respect to uncertainty of the calibration approach in individual countries.
The framework loss=hazard X exposure X vulnerabilities assumes that the damage on the left hand side can be projected by the exposed asset on the right hand side. Schwierz et al. (2010) simplified their model with a reduction to the residential sector. What does PERILS include exactly (residential, commercial, agriculture)? This is important when calibrating the left hand side. Your right hand side uses LitPop to disaggregate national asset values to a grid. Which macroeconomic indicator is used for the national asset value? Does that fit with PERILS? Is LitPop the an adequate choice for that disaggregation wrt the sectors in PERILS (residential, commercial)? Is your result sensitive on the choice of the LitPop weighting parameters? Your conclusion that the a Pan-European is inadequate due to different vulnerability curves on country level, is only valid, if you don’t include a country dependent bias in your framework (loss=hazard X exposure X vulnerabilities). Can you clarify these questions in the data section to help the reader to follow your arguments in the conclusion.
How is inflation covered in the PERILS data set? Does it influence the calibration process?
Minor comments
L149: calibration is done for each country separately. This leads to a 12 (country)x2 (cost-function) calibration matrix (L177) which makes sense. But calibration is conducted for the whole Europe first. What does that mean in the context of the 12x2 calibration matrix? Do you calculate two matrices: an European one with 1(Europe)x2(cost-function) parameters and another one with 12x2 parameters? Or do you calculate the European parameters out of the 12x2 matrix?
L214: how is MDD calculated here? For whole Europe? This refers to my question for L177.
L216: the default value is almost identical to the PERILS calibration with RMSE. What does that mean? The Schwierz function is already well calibrated but the whole method is sensitive to the cost-function?
L257 : cost function?
L308 : cost function? Please clarify at several points of the manuscript whether cost-function should be used instead of loss function.
L267 and L315: You are arguing that one vulnerability curve is inadequate. The loss model of Klawa and Ulbrich (2003) which is also mentioned by the authors, uses a local percentile approach to take local adaptation strategies into account. Integrated measures, like the storm severity index try to overcome this issue with the same local percentile approach (Leckebusch et al., 2008). Severino et al. (2024) use both Schwierz and Klawa/Ulbrich within the CLIMADA framework. Why don’t you use Klawa and Ulbrich (2003) to build a much more general pan-European model in your calibration framework?
Citation: https://doi.org/10.5194/egusphere-2026-3099-RC3 -
AC2: 'Reply on RC3', Aditya Narayan Mishra, 23 Sep 2026
The study calibrates a loss model for winter storms using the CLIMADA framework and data from the PERILS data set from 1999 to 2024. Within CLIMADA, the Schwierz et al. (2010) vulnerability function is used as default. Using two different cost functions for optimization, a parameter for the vulnerability function is tuned to fit the PERILS data. The authors conclude that the default vulnerability function systematically underestimates losses for historical events across Europe, and that a single pan-European model is inadequate to capture the distinct vulnerability profiles inherent to each country.
The manuscript is well written and nicely fits to the scope of NHESS. There are a couple of comments and minor issues, therefore I suggest minor revision before publishing.
We thank Referee #3 for their positive assessment of the manuscript. The referee’s comments and concerns are addressed below.
General
The PERILS data set includes a moderate number of events if it comes to country level. I don’t suggest to use a different data set, but for the reader it has to be very clear to know about the limits of the study with respect to uncertainty of the calibration approach in individual countries.
We agree with the referee, and note that this concern was also raised independently by Referee #1 and #2. We propose to make this limitation explicit in the Results section itself.
The framework loss = hazard × exposure × vulnerability assumes that the damage can be projected by the exposed asset. Schwierz et al. (2010) simplified their model with a reduction to the residential sector. What does PERILS include exactly (residential, commercial, agriculture)? This is important when calibrating the left-hand side.
We thank the referee for raising this concern. We calibrate the model specifically for building damages across the residential and commercial sectors, consistent with how the original model has itself been applied. Schwierz et al. (2010) note that, though the model is derived from the residential sector in their simplified formulation, it is used for all risk types - residential, commercial, industrial, and agricultural. This same practice is reflected in subsequent CLIMADA-based studies, including Röösli et al. (2021) and Severino et al. (2024), both of which apply the model to multi-sector building losses. We will ensure that this is clearly stated in the Methods section.
Your right-hand side uses LitPop. Which macroeconomic indicator is used for the national asset value? Does that fit with PERILS? Is LitPop an adequate choice for disaggregation with respect to the sectors in PERILS? Is your result sensitive to the choice of LitPop weighting parameters?
We use produced capital as the macroeconomic indicator for national asset value in LitPop, which represents a country's total stock of built capital from national accounts. We acknowledge that this does not map directly onto PERILS, since PERILS reports insured sums whereas produced capital is a broader total covering all built capital in the economy. However, we consider this an acceptable approximation because the calibration fits the damage model directly to reported PERILS losses, which implicitly absorbs a share of this exposure-level discrepancy into the fitted vulnerability parameter (Röösli et al. 2021), though we acknowledge it is not a perfect match. LitPop is adequate for representing the overall spatial distribution of total assets, which is why it is the standard exposure choice in CLIMADA windstorm studies. It distributes its national total spatially using nightlight intensity and population density but produces a single undifferentiated asset-value grid with no sector attribution. Since in our study we do not use PERILS' sector breakdown separately either, both the exposure and the loss side of our calibration are treated at the same undifferentiated level, keeping the two consistent with each other. Finally, on sensitivity to the LitPop weighting parameters, we use the default exponents (m=1, n=1), which weight nightlight intensity and population density equally and which Eberenz et al. (2020) identify as the best-performing parameterization for spatially distributing national asset value based on validation against subnational GDP data. The same setting has been used in the literature (Severino et al. 2024), however, we do acknowledge that the results are sensitive to the weighting.
Your conclusion that a pan-European model is inadequate due to different vulnerability curves at the country level is only valid if you don't include a country-dependent bias in your framework. Can you clarify these questions in the Data section to help the reader follow your argument in the conclusion?
We thank the reviewer for this insightful comment. We agree that a single pan-European model could be reconciled with country-level discrepancies by introducing a country-dependent bias-correction term applied to the model output. However, this raises a more fundamental question: what advantage does a single pan-European model retain if it needs a country-specific bias term to perform adequately? In our framework, since calibration scales the model multiplicatively through MDDmax, a country-dependent bias correction is mathematically equivalent to fitting country-specific curves directly. In practice, "single model with bias terms" and "country-specific models" converge to the same outcome. We recognize that this is, in part, a conceptual distinction rather than a purely empirical one, and we do not wish to overstate our claim. We therefore propose to revise our conclusion to state more precisely that a single, uniform pan-European model - one applied without any form of country-level adjustment, is inadequate to capture the observed spatial heterogeneity in windstorm losses across Europe. We further note that a single model could, in principle, be preserved through the introduction of country-specific bias terms, if maintaining one shared underlying model is important for a given application.
How is inflation covered in the PERILS dataset? Does it influence the calibration process?
We thank the referee for raising this point. Reported losses are provided in nominal (as-reported) currency values for each event's occurrence year. To ensure consistency with our exposure data, which is anchored to reference year 2018, we inflation-adjust all reported losses to 2018 equivalent values using country-specific CPI prior to calibration. We will add this information in the revised manuscript.
Minor Comments
L149: Calibration is done for each country separately, leading to a 12×2 matrix (L177), which makes sense. But calibration is conducted for the whole of Europe first. What does that mean in the context of the 12×2 matrix? Do you calculate two matrices: a European one (1×2) and another (12×2)? Or do you calculate the European parameters out of the 12×2 matrix?
We compute two independent calibration matrices. First, a single pan-European calibration is performed using all matched events across all 12 countries pooled together as one training set, for each cost function (RMSE and RMSLE), yielding a 1×2 pan-European parameter set presented in Fig. 2 (a–c). Second, and entirely separately, calibration is repeated per country, using only that country's matched events as the training set for each cost function, yielding the 12×2 matrix presented in Fig. 3 (m–n). The pan-European parameters are not derived from (e.g., averaged from) the 12×2 matrix; they come from an independent pooled-data optimization. We propose to make this clearer in the revised manuscript.
L214: How is MDD calculated here? For whole Europe? This refers to my question for L177.
Yes, here we are referring to pan-European calibration. We propose to make this explicit.
L216: The default value is almost identical to the PERILS calibration with RMSE. What does that mean? The Schwierz function is already well calibrated but the whole method is sensitive to the cost function?
We acknowledge the confusion here. The near-equivalence between the default and RMSE-calibrated pan-European parameter does not indicate that the default Schwierz et al. (2010) model is already well calibrated overall; rather, it reflects the fact that RMSE, by penalizing squared absolute error, is dominated by the handful of largest-magnitude events in the catalogue (Fig. 2c). The default model overestimates several of these largest events while underestimating the numerous mid-range events, these two opposing corrections partially cancel out, resulting in a small net change from the default state. This is precisely why the choice of cost function is important, as argued in the article.
L257: cost function?
We will make the necessary revision.
L308: cost function? Please clarify at several points of the manuscript whether cost function should be used instead of loss function.
We thank the referee for pointing this out. We will ensure that we use consistent terminology throughout the manuscript.
L267 and L315: You argue that one vulnerability curve is inadequate. Klawa and Ulbrich (2003), which you also mention, uses a local percentile approach for local adaptation. Severino et al. (2024) use both Schwierz and Klawa/Ulbrich within CLIMADA. Why don't you use Klawa and Ulbrich (2003) to build a more general pan-European model in your calibration framework?
The referee correctly notes that the Klawa and Ulbrich (2003) local-percentile (V/V98) approach builds in a form of local climatological adaptation by normalizing wind intensity relative to each location's own historical wind climate. This could be useful in reducing some of the country-level vulnerability heterogeneity. However, our choice to focus exclusively on recalibrating the Schwierz et al. (2010) model was motivated by our aim to isolate and demonstrate the effect of cost function choice (RMSE vs. RMSLE) as a modeling decision, independent of a change in functional form, which would have introduced an additional confounding degree of freedom. Also to mention that the Schwierz et al. (2010) model has a widespread use in literature and has an open-source status within CLIMADA - which is important to our study. We agree that the Klawa and Ulbrich approach is a useful alternative and see it as a promising direction for future work comparing functional forms alongside cost-function choice.
Citation: https://doi.org/10.5194/egusphere-2026-3099-AC2
-
AC2: 'Reply on RC3', Aditya Narayan Mishra, 23 Sep 2026
Viewed
| HTML | XML | Total | BibTeX | EndNote | |
|---|---|---|---|---|---|
| 329 | 154 | 43 | 526 | 25 | 25 |
- HTML: 329
- PDF: 154
- XML: 43
- Total: 526
- BibTeX: 25
- EndNote: 25
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
Summary
The author team recalibrate a widely used damage model for European wind storms using recent loss data from the PERILS database (1999–2024). Thereby the used separate calibration on the country level of 12 European countries. Calibration uses two cost functions: Root Mean Squared Error (RMSE) and Root Mean Squared Logarithmic Error (RMSLE). The default model systematically underestimates losses across Europe, and a single pan-European model cannot capture the distinct vulnerability profiles of individual countries. Furthermore, the selected cost functions shape the calibrated curves and the resulting risk profile, suggesting that cost function is itself a key modeling decision.
General
The author team presents a newly calibrated damage model for European storms and highlights the importance of regionalization and dependence on the selection of the calibration metric. Clearly, this is new and perfectly fits the objective of NHESS. The manuscript is well structured and well written, though some of the figures could be improved. I recommend minor revisions for possible publication in NHESS.
Comments
L13: I think the authors mean “cost function” here.
L14: I suggest to change “calibration metric” to “cost function” here to make it clear.
L25-26: Suggestion “In recent years, these regions have witnessed multiple billion-euro catastrophe events, including the February 2022 European…”
L36-38: The entire sentence must be reviewed. Note I normally reinsurance contracts run max 2 years, so climate change plays not a big role in business.
L59: “Therefore, robust calibration and validation is necessary, …”
L60: Which parameterizations are meant here, maybe avoid using this wording as people might think about parameterizations in dynamical models, I guess you mean statistical models here, such as the damage model used.
L93: The publications at the end of the sentence are misleading as they suggest that they refer to the dataset, but they only describe the storms, so please cite the dataset here.
L101-102: How is inflation handled in these data sets?
L110: This sentence can be removed as it is described in 2.3
L133: Please name the subsection “Calibration”, this is sufficient
Section 2.3: The number of events n is rather small and should be mentioned in this section.
L189-191: The sentence is not clear. It is also not clear why any calibrated MDD_max value is a generalized parameter. You need to give an explanation.
L204: “who estimated”
Results: Maybe this is the only more major comment. I think the authors need to discuss the effect of the rather small number of events on the calibrations much more. It is mentioned in the last section, but I think the reader needs in the results section information about which region the calibration is trustworthy and in which it is at the limit.
L234: “pulling down modeled damages” this is not visible in Fig 2c, it looks identical.
Fig. 3: Too small labels and axes labels. Maybe split the figure into two (m,n as separate figure).
Fig 5: Too small labels and axes labels. You also change the use of the colors, In Fig 3 orange is used for RMSLE, here for RMSE, please adjust.