Rapid flood damage estimation tool for urban pluvial floods with scarce data
Abstract. Reliable and rapid flood damage estimation is crucial for both disaster risk reduction and crisis management. Yet, existing models primarily focus on riverine floods, neglecting urban pluvial floods – a substantial gap, as heavy rainfall can lead to flooding virtually anywhere. Here, we present FlooDEsT, a new machine learning (ML)-based tool to rapidly estimate building-level damage from urban pluvial flooding with four key improvements, compared to traditional models. First, the model was trained specifically on damage data from urban pluvial flood events. Second, the tool utilises XGBoost, a ML technique capable of capturing complex non-linear data relationships. Third, the tool efficiently utilises geographical information only as necessary, reducing pre-processing time. Fourth, to address the common challenge of missing data, the tool uses smart random sampling techniques to impute building-level features that are representative to buildings affected by this flood pathway, reducing exposure bias. The tool’s computational performance was evaluated in two German case studies, involving about 2300 and 440,000 buildings. The tool provided damage estimates in respectively 2.1 to 5.6 seconds per thousand buildings, representing a 2.7- to 6.6-fold improvement in speed over a baseline approach. FlooDEsT tackles critical gaps in damage modelling, offering valuable support for disaster preparedness and response.
This work aiming presents FlooDEsT, an XGBoost-based tool for rapid damage estimation from pluvial floods. The training data is based on household survey data from Germany and validated on independent datasets from Amsterdam and Karlsruhe. It addresses an important innovation for under-modelled hazard pathway (pluvial floods). Yet, major revision and clarification is needed. Please see my comments and suggestions below.
Introduction
1. The motivation that pluvial floods are under-represented in damage modelling is clearly stated. However, why impacts are not systematically documented (line 21) required additional information that further supports this as the gap.
2. Consider reviewing the damage models and present them in a more structured manner. For example, distinguishing them between type of model, and summarising key limitations and constraints before proposing multivariable models. Additionally, it is not clear how multivariable modelling improve damage estimations. In Line 50 you introduce this as a solution to the deficiencies in other models. Provide some references to this.
3. While the four improvements are clear, consider clarifying how FlooDEsT differs from earlier work by Bronstert et al., Lindenlaub et al., and Dobkowitz et al. beyond the extended dataset and improved computational performance. Specify the novelty of the work. Consider explicitly stating this in the text as well, as it would help to articulate the main scientific contribution.
Data and methods
4. Have you considered using other tree-based methods, such as random forest or lightGBM? How is XGBoost particularly suitable for your method?
5. People tend to forget and/or underestimate the level of impacts for diverse reasons and adding a temporal aspect could increase this effect. You have validation points from 10 years after the first points, and 14 years after the training set. Are you accounting for this or handling this in any way? If so, how? What happens when households do not repair/replace in that span of time?
6. The training and test sets split is described, but details on the splitting procedure are missing. Please specify whether the split was random, stratified, or followed another method.
7. In your methods you are working with data exclusively from Germany, and have external data points from Amsterdam for the validation. This could result in more localised/regional differences inherent to the affected households after pluvial flooding events. Have you considered this as a limitation for the work? This already led to removing entries (see specific comments below).
8. The predictor set points need further clarification:
9. The filtering of Amsterdam data by removing entries with water entry via windows, roofs or downpipes should be described more explicitly by clarifying how many points were removed. Is this introducing any selection bias?
10. The estimation of building value with linear mixed-effect models is a critical step to obtain absolute losses. Its performance is not reported. Please provide performance metrics for the building value model. Discuss how uncertainty propagates to relative loss and damage class assignment.
11. In Line 139 you mention a case was dismissed because the building value was "much lower that reported losses". Could you explicitly provide information as to which case that was? Please specify the used criterion. Are there other points close to the threshold? Was this done before or after fitting the model? Additional explanation is needed for the Amsterdam case. Building values there would be based and modelled on the German insurance market standards.
12. To improve clarity and robustness of this section, some gaps need to be filled.
13. In your sensitivity analysis, have you considered testing scenarios with multiple missing variables simultaneously? This could provide a more realistic setting in operationalising realistic data gaps. Report not only RMSE but changes in damage class under these scenarios.
Results and Discussion
14. The independent validation datasets are valuable, but with a small total sample size. The manuscript acknowledges this limitation, but the implications for the robustness are not fully explored. Discuss how sensitive RMSE and the damage classes are to individual points, given the small sample size. Confidence intervals or other uncertainty estimates could be a valuable addition.
15. The model fails to predict "very high" damage cases in the validation set. This is a critical limitation for a tool intended for flood damage estimation, where extreme losses are key.
16. You provide references mostly for the results in the calibration performance. Please provide additional references in the following subsections of your results. These provide a more in-depth understanding and discussion of your approach, strengths and weaknesses.
17. Are there other ways to optimise the procedure? How can reproducibility be ensured? Please consider including a short paragraph on the main text.
18. Figure 10 is a very useful way of showing visually the estimated damage classes. Consider adding a brief interpretation of patterns if known, to connect the damage outputs to the drivers. It could be connected to what point 15 mentions.
Conclusions
19. Please consider adding recommendations for practitioners, including testing under different infrastructure, etc.
General remarks:
Consider reviewing the language and other minor details. Examples include: