the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
An XBeach-informed machine-learning surrogate for rapid coastal flood prediction
Abstract. Rapid assessment of coastal inundation along the Western Black Sea coast requires computationally efficient methods that can screen multiple compound storm scenarios while preserving a clear link to the controlling physical drivers. This study develops and evaluates an XBeach-informed machine-learning surrogate for coastal inundation and threshold-based flood-extent screening. The framework combines high-frequency sea-level observations from the Maritime Hydrographic Directorate monitoring network, Copernicus Marine wave and Black Sea physical products, ERA5 atmospheric forcing, Global Runoff Data Centre Danube discharge, bathymetric and coastal-elevation descriptors, and an intermediate cross-shore physical-response layer. The modelling dataset consists of 5,000 Monte Carlo forcing scenarios, 22 forcing and derived predictors, and a 10 × 10 spatial target grids, corresponding to 100 inundation-depth nodes per scenario. Four tree-based residual models, namely Random Forest, Gradient Boosting, Extra Trees, and Histogram Gradient Boosting, are combined using validation-derived weights.
The conservative held-out node-level evaluation shows moderate quantitative skill for continuous inundation-depth prediction. The ensemble reaches R² = 0.409, RMSE = 1.181 m, MAE = 0.678 m, NSE = 0.409, KGE = 0.235, Pearson r = 0.914, and Spearman ρ = 0.923. In contrast, threshold-based flood-extent classification above the 0.30 m operational inundation threshold is substantially stronger, with F1 = 0.995, AUC-ROC = 0.9998, Matthews correlation coefficient = 0.989, and Cohen’s κ = 0.989. Split conformal prediction intervals provide near-nominal 90 % marginal coverage, with empirical coverage of 0.911, but the mean 90 % interval width is large, 3.25 m, indicating limited sharpness for local quantitative depth estimation. Extreme-event diagnostics show systematic underprediction of upper-tail inundation depths, with a mean bias of approximately −2.59 m for events above the 90th percentile.
The present configuration is therefore most appropriate for rapid flood-extent screening, scenario ranking, and identification of cases where more detailed process-based simulations are required. It should not yet be interpreted as a stand-alone operational predictor of maximum inundation depth. The study demonstrates how in situ observations, Copernicus Marine products, physical-response modelling, machine-learning emulation, and uncertainty diagnostics can be combined into a transparent coastal-hazard screening workflow for the Western Black Sea coast.
- Preprint
(1655 KB) - Metadata XML
- BibTeX
- EndNote
Status: open (until 09 Sep 2026)
-
RC1: 'Comment on egusphere-2026-3628', Anonymous Referee #1, 21 Jul 2026
reply
-
AC1: 'Reply on RC1', Maria Mihailov, 23 Jul 2026
reply
Thank you for this constructive suggestion.
Response to Comment 1: We agree that the previous formulation described several methodological steps separately and therefore masked the broader progression of the study. We have revised the final part of the Introduction by consolidating the objectives into three overarching themes: (1) data and framework development; (2) model development and validation; and (3) process interpretation and practical application. We also reformulated the following paragraph so that continuous inundation-depth prediction, threshold-based flood-extent identification, and prediction-interval performance are presented as complementary dimensions of model evaluation rather than as a second list of objectives. We hope that this revision reduces repetition and clarifies the connection between the study design, model assessment, physical interpretation, and intended coastal-hazard application.
“The objective of this paper is to develop and evaluate a physically informed machine-learning surrogate for rapid coastal-inundation prediction and flood-extent screening along the Western Black Sea coast. The study is structured around three main objectives: (1) data and framework development, through the integration of in situ sea-level observations, Copernicus Marine products, ERA5 atmospheric forcing (Hersbach et al., 2020, 2023), Danube discharge (GRDC, 2026), and coastal-morphology descriptors within an XBeach-informed reference framework (Roelvink et al., 2009) and a Monte Carlo scenario space; (2) model development and validation, through the training and comparison of tree-based residual surrogates for spatial inundation-depth prediction and threshold-based flood-extent identification, together with a comprehensive assessment of predictive skill, residual behaviour, extreme-event performance, and uncertainty; and (3) process interpretation and practical application, through the identification of the dominant physical controls on the surrogate response and the evaluation of its suitability and limitations for rapid coastal-flood screening, scenario prioritisation, and Copernicus downstream decision support.”
Response to Comment 2: The Code and data availability section has been expanded to identify the publicly accessible datasets used in the study, together with their data providers, repositories, product identifiers, and persistent DOI or portal addresses. The revised statement now includes the FOCCUS/MHD sea-level observations, the Copernicus Marine Black Sea wave, physical-state, and sea-surface-temperature products, ERA5 atmospheric forcing, GRDC Danube discharge records, and EMODnet Digital Bathymetry. We also clarify the applicable GRDC data-sharing restrictions. The processed forcing matrix, Monte Carlo scenarios, XBeach-informed target tables, validation outputs, trained model objects, environment information, source code, and figure-generation scripts will be deposited in a DOI-bearing Zenodo repository associated with the final publication. Until this repository is released, the authors’ processed datasets, code, and algorithms are available from the corresponding authors upon reasonable request, subject to third-party licensing and redistribution restrictions. The planned repository metadata will document product identifiers, versions, retrieval dates, processing steps, and software dependencies to support reproducibility and FAIR reuse.
“All input datasets used in this study are publicly available online. Sea-level elevation and atmospheric-pressure observations from the Romanian Black Sea stations at Constanța, Gura Portiței, Mangalia, and Midia were provided by the Maritime Hydrographic Directorate and are distributed through Zenodo (Dutu et al., 2026, https://doi.org/10.5281/zenodo.19133361). Black Sea wave data were obtained from the Copernicus Marine Service Black Sea Waves Analysis and Forecast product (BS-Waves, EAS5; Staneva et al., 2022, https://doi.org/10.25423/CMCC/BLKSEA_ANALYSISFORECAST_WAV_007_003_EAS5). Black Sea physical-state data were obtained from the Copernicus Marine Service Black Sea Physics Analysis and Forecast product (BLK-PHY, EAS7; Jansen et al., 2024, https://doi.org/10.48670/mds-00355). High-resolution Level-4 sea-surface-temperature data were obtained from the Copernicus Marine Service (Buongiorno Nardelli et al., 2010, 2013, https://doi.org/10.48670/moi-00159). Hourly ERA5 atmospheric data on single levels, including the 10 m wind components and mean sea-level pressure, were obtained from the Copernicus Climate Change Service Climate Data Store (Hersbach et al., 2020, 2023, https://doi.org/10.24381/cds.bd0915c6). Daily Danube River discharge data were obtained from the Global Runoff Data Centre, Federal Institute of Hydrology, Koblenz, Germany (GRDC, 2026, https://portal.grdc.bafg.de/, last access: 20 June 2026). Bathymetric and coastal-elevation data were obtained from the EMODnet Bathymetry Consortium (2026, https://doi.org/10.12770/bb6a87dd-e579-4036-abe1-e649cea9881a).
The computational workflow was implemented in Python using a notebook-based workflow developed in the Jupyter environment (Kluyver Thomas et al., 2016) and managed through the Anaconda Python distribution. The computational workflow was implemented in Python and includes data alignment, Monte Carlo scenario generation, construction of the XBeach-informed physical-response layer, residual-surrogate training, model validation, uncertainty analysis, and figure production. The processed model-input matrices, Monte Carlo scenarios, spatial target data, trained model objects, validation outputs, and scripts used to generate the results presented in this study will be made available in a DOI-bearing repository accompanying the final publication, subject to third-party data-licence constraints.
All metrics, figures and tables reported in this manuscript are based on the aligned observational and gridded forcing dataset described in Sects. 2 and 3.”
Response to Comment 3: Additional methodological references have been introduced in Sections 3.2, 3.4, and 3.5. These references support the use of block-bootstrap sampling, Earth Mover’s Distance, the selected tree-ensemble algorithms, permutation-based feature importance, the NSE and KGE performance criteria, the MCC, Cohen’s kappa and ROC-AUC classification metrics, and split conformal prediction intervals. We have also clarified that the 0.30 m inundation threshold is a study-specific operational screening threshold and should not be interpreted as a universal flood-damage threshold.
“ 3.2 Monte Carlo scenario generation
A Monte Carlo simulation (MCS) framework was used to generate 5,000 forcing scenarios, providing broad coverage of the multivariate forcing space while keeping the resulting 500,000 node-level samples computationally tractable for repeated model training and validation. The scenario matrix includes both sampled variables and process-derived predictors, including wind speed and direction, significant wave height, peak wave period, wave direction, sea-level elevation, mean sea-level pressure, Danube discharge, sea-surface temperature, storm duration, inverse-barometer contribution, wave-power-related terms, and interaction terms between sea-level anomaly and wave forcing. Where the dependence structure between variables was relevant, empirical or block-bootstrap sampling was used to retain part of the observed co-variability, following the general block-resampling approach of Künsch (1989). The consistency of the sampled variables with the observed or aligned forcing data was evaluated using distributional comparisons, including histograms and kernel-density estimates, together with Earth Mover’s Distance (EMD) as a quantitative measure of marginal distributional agreement (Rubner et al., 2000; Table 3; Figure 4).
(……)
3.4 Hybrid residual surrogate
The final surrogate was formulated as a residual correction to the XBeach-informed physical baseline. In this formulation, the physical layer first provides a baseline estimate of inundation depth, and the machine-learning component then predicts the residual term, defined as the difference between the target inundation response and the baseline estimate. This structure preserves the physical interpretability of the model because the baseline response remains explicitly visible, while the residual model accounts for the part of the inundation signal that is not captured by the simplified physical-response layer.
Four tree-based regression models were trained: Random Forest (RF; Breiman, 2001), Gradient Boosting (GB) and Histogram Gradient Boosting (HGB), following the gradient-boosting framework of Friedman (2001), and Extra Trees (ET; Geurts et al., 2006). The individual model predictions were combined into a validation-weighted ensemble, using weights derived from their validation performance (Table 4). The resulting weights are nearly equal, indicating that the four models contributed similarly to the ensemble prediction. This ensemble formulation reduces dependence on a single algorithm and provides a more stable residual correction across the Monte Carlo scenario space.
The residual-surrogate design also supports model interpretation. When the predicted residual is small, the XBeach-informed baseline explains most of the simulated inundation response. When the residual correction is large, feature-importance diagnostics can identify which forcing variables or morphology-related descriptors contribute most strongly to the correction. Permutation importance was calculated on held-out data by shuffling one predictor at a time and measuring the resulting increase in residual RMSE, consistent with permutation-based variable-importance analysis for tree ensembles (Breiman, 2001).
3.5 Validation metrics and uncertainty
Model performance was evaluated using complementary regression, classification, residual, and uncertainty diagnostics. Continuous inundation-depth predictions were assessed using the coefficient of determination (R²), root mean square error (RMSE), mean absolute error (MAE), mean bias, Nash–Sutcliffe efficiency (NSE; Nash and Sutcliffe, 1970), Kling–Gupta efficiency (KGE; Gupta et al., 2009), Pearson correlation, and Spearman rank correlation. These metrics jointly describe error magnitude, systematic overprediction or underprediction, linear and rank-based association, and model skill relative to reference predictions.
Threshold-based flood extent was evaluated by treating each grid node as flooded or non-flooded according to an operational inundation threshold of 0.30 m. This value is used here as a study-specific screening threshold and should not be interpreted as a universal damage threshold; comparable threshold-based evaluations have been used for coastal urban flood surrogates (Zahura et al., 2020). Binary classification skill was quantified using accuracy, precision, recall, F1 score, Matthews correlation coefficient (MCC; Matthews, 1975), Cohen’s kappa (Cohen, 1960), area under the receiver operating characteristic curve (AUC-ROC; Hanley and McNeil, 1982), and average precision. This separate classification assessment is necessary because identifying flooded nodes and reproducing continuous inundation depth represent different levels of predictive difficulty.
Residual behaviour was analysed using residual histograms, fitted normal-density curves, residual-versus-predicted plots, and quantile–quantile (Q–Q) diagnostics. These diagnostics were used to identify systematic bias, non-normal error structure, heteroscedasticity, and tail errors associated with the largest inundation values.
Predictive uncertainty was quantified using split conformal prediction intervals following the distribution-free regression framework described by Lei et al. (2018). Under exchangeability between calibration and test data, split conformal intervals provide finite-sample marginal coverage without requiring a parametric residual distribution. Their practical value for coastal-hazard interpretation nevertheless depends on interval sharpness, calibration across the inundation-depth range, and robustness under distribution shifts, particularly for rare or extreme flood scenarios.”
Sections 5.1–5.3 have been expanded to compare the present results with previous surrogate-modelling studies. Section 5.1 now discusses the commonly reported contrast between high skill in identifying flooded locations and lower skill in reproducing continuous or maximum inundation depth. The revised text also explains why direct numerical comparison among studies is not appropriate when the forcing ranges, reference models, spatial resolutions, morphologies, threshold definitions, and validation strategies differ. Section 5.2 relates the present framework to comparable rapid flood-mapping and early-warning applications and clarifies its intended role as a computational screening and prioritisation layer rather than a replacement for locally calibrated process-based models. Section 5.3 now supports the interpretation of upper-tail underprediction with studies reporting similar limitations arising from rare-event scarcity, incomplete predictor representation, morphology-dependent generalisation, and extrapolation beyond the well-sampled training domain. We have also added specific future-development priorities, including targeted sampling of severe compound events, event-based XBeach calibration, richer morphology descriptors, independent impact validation, and tail-aware uncertainty calibration.
“5.1 Scientific interpretation of the model skill pattern
The main scientific result is not represented by a single performance metric, but by the contrast between threshold-based flood-extent classification and continuous inundation-depth prediction. The surrogate identifies flooded and non-flooded grid nodes with high skill above the 0.30 m operational inundation threshold, but it remains less accurate in reproducing the magnitude of the largest inundation depths. Comparable hybrid hydraulic–machine-learning studies have reported the same qualitative hierarchy of predictive difficulty. Hosseiny et al. (2020) obtained very high wet–dry classification accuracy while the corresponding depth regression remained less exact, and Zahura et al. (2020, 2022) evaluated thresholded inundation extent separately from continuous water depth in physics-based coastal urban surrogates. The present combination of F1 = 0.995 and R² = 0.409 is therefore consistent with the broader observation that delineating inundated locations is generally less demanding than estimating the exact depth at each location. Direct numerical comparison is not appropriate, however, because the studies differ in forcing ranges, reference models, spatial resolution, morphology, threshold definition, and validation design.
The observed skill pattern indicates that the surrogate learns the dominant inundation-response structure and preserves much of the event ranking, but does not yet fully resolve the most severe inundation conditions. High surrogate skill has been reported when models are trained on extensive, high-fidelity and site-specific simulation databases that explicitly represent local geometry and boundary conditions (Ayyad et al., 2022; Ma et al., 2022; Ahmadi Gharehtoragh and Johnson, 2024). Relative to those configurations, the present model uses a deliberately simplified XBeach-informed baseline and idealised spatial descriptors over a morphologically heterogeneous coast. Its strong Pearson and Spearman correlations therefore indicate useful preservation of scenario ordering and spatial response, whereas the moderate R² and systematic upper-tail bias show that response magnitude is compressed. The likely causes include limited representation of rare severe events, the simplified physical baseline, the idealised 10 × 10 grid, and the absence of richer local descriptors such as dune-crest elevation, beach width, backshore slope, harbour exposure, lagoon connectivity, and surface roughness. Consequently, the model is better suited to rapid flood-extent screening and scenario prioritisation than to stand-alone prediction of maximum inundation depth.
5.2 Implications for coastal early warning
The present surrogate is best interpreted as a rapid screening and prioritisation tool rather than as a deterministic predictor of absolute inundation depth. This role is consistent with published surrogate applications developed to accelerate coastal and urban flood assessment. Rapid storm-surge mapping systems and machine-learning emulators have been used to reduce the computational burden of large scenario ensembles and to support faster-than-full-physics evaluation (Yang et al., 2020; Zahura et al., 2020; Ayyad et al., 2022). In a coastal urban setting, Zahura and Goodall (2022) showed that a Random-Forest surrogate could reproduce tidal and pluvial flood patterns within seconds and could be compared with independent crowdsourced flood reports. These studies support the use of surrogates as a computational triage layer that identifies potentially affected locations and prioritises the scenarios requiring detailed hydrodynamic simulation.
This type of framework is relevant for Copernicus downstream coastal services because it links in situ observations, gridded marine and atmospheric products, physical-response modelling, and uncertainty diagnostics within a computationally efficient workflow. Nevertheless, the comparison with previous applications also clarifies the present boundary of use. Operational or near-real-time surrogate systems generally depend on a locally calibrated high-fidelity reference model, representative training events, and independent impact or inundation observations (Yang et al., 2020; Zahura et al., 2020, 2022). In the present study, the XBeach-informed layer is simplified and the storm-sequence diagnostic is an internal consistency check rather than independent validation. The model can therefore provide rapid guidance for scenario exploration, flood-extent screening, and selection of cases for further simulation, but operational quantitative depth forecasting would require locally calibrated XBeach runs, richer morphology, and independent flood, overtopping, or impact observations.
5.3 Why the extreme tail remains difficult
The reduced skill for upper-tail inundation depths is likely caused by a combination of sampling, physical-representation, and spatial-description limitations, and is consistent with a recognised difficulty in statistical and machine-learning representations of extremes. Harter et al. (2024) showed that storm-surge reconstructions commonly underestimate the largest events because extremes are rare, relevant drivers may be omitted, and atmospheric forcing can itself be biased. Ayyad et al. (2022) similarly noted that limited training samples constrain prediction of low-probability, high-consequence surges, while Ahmadi Gharehtoragh and Johnson (2024) found larger surrogate errors for lower annual exceedance probabilities and the most extreme landscape configuration. In the present dataset, severe inundation cases form only a small fraction of the training space, and the concentration of target values near the imposed upper bound of approximately 4 m creates an additional plateau that limits discrimination among the largest responses.
Second, the XBeach-informed physical baseline is intentionally simplified and does not resolve all local processes that may control extreme inundation, including wave-overtopping pathways, backshore storage, small topographic barriers, harbour geometry, lagoon connectivity, and spatial variations in roughness. The resulting pattern—high skill for the inundation footprint but underestimation of large-event depths—has also been reported in a different coastal-inundation context by Ragu Ramalingam et al. (2025), whose machine-learning surrogate reproduced inundation maps more reliably than maximum depths for the largest test events. This comparison does not imply equivalence between storm-surge and tsunami processes, but it illustrates a general surrogate-modelling limitation when the training database or reduced physical representation does not fully span the most severe onshore response.
Third, the spatial grid and morphology descriptors remain idealised. They provide a consistent framework for rapid screening, but they do not fully distinguish between dune-backed beaches, engineered harbour areas, low-lying lagoonal sectors, cliff-backed coasts, and urbanised backshore environments. Studies that explicitly include changing landscape and boundary-condition descriptors show that morphology information can materially improve transfer across scenarios, while the most extreme or least represented landscape states remain more difficult to predict (Ahmadi Gharehtoragh and Johnson, 2024). The simplified geometry used here therefore helps explain why the surrogate captures flood occurrence more reliably than the exact magnitude of upper-tail inundation depths.
Future development should therefore prioritise event-based XBeach calibration at representative coastal sites, targeted sampling of severe compound wave–surge conditions, removal or reformulation of the imposed upper-depth plateau, richer morphology descriptors, and independent validation against observed flood, overtopping, or impact data. Tail-aware training strategies, quantile-oriented objectives, and conformal calibration stratified by depth range or morphology class should also be evaluated, but only after the physical baseline and extreme-event training sample have been strengthened. These steps are required before the framework can support quantitative operational prediction of maximum inundation depth.
These limitations do not reduce the value of the present surrogate as a rapid screening tool; rather, they define the boundary between exploratory flood-extent assessment and operational quantitative inundation-depth forecasting. Within that boundary, the framework provides a transparent means of combining observed and gridded forcing, a physics-informed baseline, machine-learning correction, and explicit uncertainty diagnostics.”
Citation: https://doi.org/10.5194/egusphere-2026-3628-AC1
-
AC1: 'Reply on RC1', Maria Mihailov, 23 Jul 2026
reply
Viewed
| HTML | XML | Total | BibTeX | EndNote | |
|---|---|---|---|---|---|
| 102 | 38 | 14 | 154 | 9 | 11 |
- HTML: 102
- PDF: 38
- XML: 14
- Total: 154
- BibTeX: 9
- EndNote: 11
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
The manuscript provides a robust and comprehensive description of the materials and methods. However, several aspects could be improved to enhance clarity and strengthen the scientific presentation.
In the Introduction, describing 7 specific objectives is unnecessarily detailed and makes the section appear repetitive, as it lists individual analytical steps rather than the overarching research goals. Would be positive to restructuring them into three broader objectives following a clearer logical progression: (1) data and framework development; (2) model development and validation; and (3) process interpretation and practical application.
The Methods section is comprehensive and well documented. However, publicly available datasets should also be listed in the Data Availability Statement, including their repositories, persistent links (e.g., DOI or URL), and data providers. In addition, datasets, source code, or algorithms that are available upon reasonable request could also be mentioned to further strengthen compliance with the FAIR principles. Adopting these recommendations would improve the transparency, reproducibility, and overall impact of the study. Furthermore, several methodological choices and parameter settings are presented without supporting references. Although many of these statistical approaches are well established, citing the original or standard methodological references would better justify their selection and increase the scientific credibility of the adopted framework.
The Discussion provides a clear interpretation of the results but would benefit from a stronger comparison with previous studies. While the Introduction section is well referenced, Section 5.1 would be strengthened by discussing similar studies and demonstrating that the observed model-skill patterns are consistent with findings reported in the literature. Likewise, Section 5.2 highlights the potential and limitations of surrogate models for early-warning systems but does not include references to comparable applications using similar methodologies. Finally, the discussion in Section 5.3 regarding the reduced performance for extreme-tail events should also be supported by references to studies reporting similar limitations, as this is a common challenge in surrogate modelling and extreme-event prediction.
Overall, the study successfully addresses its stated objectives and presents a framework with significant practical relevance. The proposed methodology has considerable potential for regional coastal-flood hazard assessment and provides a valuable reference for the development of operational early-warning systems. In particular, the framework demonstrates an appropriate balance between computational efficiency, physical realism, and transparent communication of uncertainties, making it well suited to support coastal management and decision-making for compound coastal flooding.