the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
An XBeach-informed machine-learning surrogate for rapid coastal flood prediction
Abstract. Rapid assessment of coastal inundation along the Western Black Sea coast requires computationally efficient methods that can screen multiple compound storm scenarios while preserving a clear link to the controlling physical drivers. This study develops and evaluates an XBeach-informed machine-learning surrogate for coastal inundation and threshold-based flood-extent screening. The framework combines high-frequency sea-level observations from the Maritime Hydrographic Directorate monitoring network, Copernicus Marine wave and Black Sea physical products, ERA5 atmospheric forcing, Global Runoff Data Centre Danube discharge, bathymetric and coastal-elevation descriptors, and an intermediate cross-shore physical-response layer. The modelling dataset consists of 5,000 Monte Carlo forcing scenarios, 22 forcing and derived predictors, and a 10 × 10 spatial target grids, corresponding to 100 inundation-depth nodes per scenario. Four tree-based residual models, namely Random Forest, Gradient Boosting, Extra Trees, and Histogram Gradient Boosting, are combined using validation-derived weights.
The conservative held-out node-level evaluation shows moderate quantitative skill for continuous inundation-depth prediction. The ensemble reaches R² = 0.409, RMSE = 1.181 m, MAE = 0.678 m, NSE = 0.409, KGE = 0.235, Pearson r = 0.914, and Spearman ρ = 0.923. In contrast, threshold-based flood-extent classification above the 0.30 m operational inundation threshold is substantially stronger, with F1 = 0.995, AUC-ROC = 0.9998, Matthews correlation coefficient = 0.989, and Cohen’s κ = 0.989. Split conformal prediction intervals provide near-nominal 90 % marginal coverage, with empirical coverage of 0.911, but the mean 90 % interval width is large, 3.25 m, indicating limited sharpness for local quantitative depth estimation. Extreme-event diagnostics show systematic underprediction of upper-tail inundation depths, with a mean bias of approximately −2.59 m for events above the 90th percentile.
The present configuration is therefore most appropriate for rapid flood-extent screening, scenario ranking, and identification of cases where more detailed process-based simulations are required. It should not yet be interpreted as a stand-alone operational predictor of maximum inundation depth. The study demonstrates how in situ observations, Copernicus Marine products, physical-response modelling, machine-learning emulation, and uncertainty diagnostics can be combined into a transparent coastal-hazard screening workflow for the Western Black Sea coast.
- Preprint
(1655 KB) - Metadata XML
- BibTeX
- EndNote
Status: open (until 09 Sep 2026)
-
RC1: 'Comment on egusphere-2026-3628', Anonymous Referee #1, 21 Jul 2026
reply
-
AC1: 'Reply on RC1', Maria Mihailov, 23 Jul 2026
reply
Thank you for this constructive suggestion.
Response to Comment 1: We agree that the previous formulation described several methodological steps separately and therefore masked the broader progression of the study. We have revised the final part of the Introduction by consolidating the objectives into three overarching themes: (1) data and framework development; (2) model development and validation; and (3) process interpretation and practical application. We also reformulated the following paragraph so that continuous inundation-depth prediction, threshold-based flood-extent identification, and prediction-interval performance are presented as complementary dimensions of model evaluation rather than as a second list of objectives. We hope that this revision reduces repetition and clarifies the connection between the study design, model assessment, physical interpretation, and intended coastal-hazard application.
“The objective of this paper is to develop and evaluate a physically informed machine-learning surrogate for rapid coastal-inundation prediction and flood-extent screening along the Western Black Sea coast. The study is structured around three main objectives: (1) data and framework development, through the integration of in situ sea-level observations, Copernicus Marine products, ERA5 atmospheric forcing (Hersbach et al., 2020, 2023), Danube discharge (GRDC, 2026), and coastal-morphology descriptors within an XBeach-informed reference framework (Roelvink et al., 2009) and a Monte Carlo scenario space; (2) model development and validation, through the training and comparison of tree-based residual surrogates for spatial inundation-depth prediction and threshold-based flood-extent identification, together with a comprehensive assessment of predictive skill, residual behaviour, extreme-event performance, and uncertainty; and (3) process interpretation and practical application, through the identification of the dominant physical controls on the surrogate response and the evaluation of its suitability and limitations for rapid coastal-flood screening, scenario prioritisation, and Copernicus downstream decision support.”
Response to Comment 2: The Code and data availability section has been expanded to identify the publicly accessible datasets used in the study, together with their data providers, repositories, product identifiers, and persistent DOI or portal addresses. The revised statement now includes the FOCCUS/MHD sea-level observations, the Copernicus Marine Black Sea wave, physical-state, and sea-surface-temperature products, ERA5 atmospheric forcing, GRDC Danube discharge records, and EMODnet Digital Bathymetry. We also clarify the applicable GRDC data-sharing restrictions. The processed forcing matrix, Monte Carlo scenarios, XBeach-informed target tables, validation outputs, trained model objects, environment information, source code, and figure-generation scripts will be deposited in a DOI-bearing Zenodo repository associated with the final publication. Until this repository is released, the authors’ processed datasets, code, and algorithms are available from the corresponding authors upon reasonable request, subject to third-party licensing and redistribution restrictions. The planned repository metadata will document product identifiers, versions, retrieval dates, processing steps, and software dependencies to support reproducibility and FAIR reuse.
“All input datasets used in this study are publicly available online. Sea-level elevation and atmospheric-pressure observations from the Romanian Black Sea stations at Constanța, Gura Portiței, Mangalia, and Midia were provided by the Maritime Hydrographic Directorate and are distributed through Zenodo (Dutu et al., 2026, https://doi.org/10.5281/zenodo.19133361). Black Sea wave data were obtained from the Copernicus Marine Service Black Sea Waves Analysis and Forecast product (BS-Waves, EAS5; Staneva et al., 2022, https://doi.org/10.25423/CMCC/BLKSEA_ANALYSISFORECAST_WAV_007_003_EAS5). Black Sea physical-state data were obtained from the Copernicus Marine Service Black Sea Physics Analysis and Forecast product (BLK-PHY, EAS7; Jansen et al., 2024, https://doi.org/10.48670/mds-00355). High-resolution Level-4 sea-surface-temperature data were obtained from the Copernicus Marine Service (Buongiorno Nardelli et al., 2010, 2013, https://doi.org/10.48670/moi-00159). Hourly ERA5 atmospheric data on single levels, including the 10 m wind components and mean sea-level pressure, were obtained from the Copernicus Climate Change Service Climate Data Store (Hersbach et al., 2020, 2023, https://doi.org/10.24381/cds.bd0915c6). Daily Danube River discharge data were obtained from the Global Runoff Data Centre, Federal Institute of Hydrology, Koblenz, Germany (GRDC, 2026, https://portal.grdc.bafg.de/, last access: 20 June 2026). Bathymetric and coastal-elevation data were obtained from the EMODnet Bathymetry Consortium (2026, https://doi.org/10.12770/bb6a87dd-e579-4036-abe1-e649cea9881a).
The computational workflow was implemented in Python using a notebook-based workflow developed in the Jupyter environment (Kluyver Thomas et al., 2016) and managed through the Anaconda Python distribution. The computational workflow was implemented in Python and includes data alignment, Monte Carlo scenario generation, construction of the XBeach-informed physical-response layer, residual-surrogate training, model validation, uncertainty analysis, and figure production. The processed model-input matrices, Monte Carlo scenarios, spatial target data, trained model objects, validation outputs, and scripts used to generate the results presented in this study will be made available in a DOI-bearing repository accompanying the final publication, subject to third-party data-licence constraints.
All metrics, figures and tables reported in this manuscript are based on the aligned observational and gridded forcing dataset described in Sects. 2 and 3.”
Response to Comment 3: Additional methodological references have been introduced in Sections 3.2, 3.4, and 3.5. These references support the use of block-bootstrap sampling, Earth Mover’s Distance, the selected tree-ensemble algorithms, permutation-based feature importance, the NSE and KGE performance criteria, the MCC, Cohen’s kappa and ROC-AUC classification metrics, and split conformal prediction intervals. We have also clarified that the 0.30 m inundation threshold is a study-specific operational screening threshold and should not be interpreted as a universal flood-damage threshold.
“ 3.2 Monte Carlo scenario generation
A Monte Carlo simulation (MCS) framework was used to generate 5,000 forcing scenarios, providing broad coverage of the multivariate forcing space while keeping the resulting 500,000 node-level samples computationally tractable for repeated model training and validation. The scenario matrix includes both sampled variables and process-derived predictors, including wind speed and direction, significant wave height, peak wave period, wave direction, sea-level elevation, mean sea-level pressure, Danube discharge, sea-surface temperature, storm duration, inverse-barometer contribution, wave-power-related terms, and interaction terms between sea-level anomaly and wave forcing. Where the dependence structure between variables was relevant, empirical or block-bootstrap sampling was used to retain part of the observed co-variability, following the general block-resampling approach of Künsch (1989). The consistency of the sampled variables with the observed or aligned forcing data was evaluated using distributional comparisons, including histograms and kernel-density estimates, together with Earth Mover’s Distance (EMD) as a quantitative measure of marginal distributional agreement (Rubner et al., 2000; Table 3; Figure 4).
(……)
3.4 Hybrid residual surrogate
The final surrogate was formulated as a residual correction to the XBeach-informed physical baseline. In this formulation, the physical layer first provides a baseline estimate of inundation depth, and the machine-learning component then predicts the residual term, defined as the difference between the target inundation response and the baseline estimate. This structure preserves the physical interpretability of the model because the baseline response remains explicitly visible, while the residual model accounts for the part of the inundation signal that is not captured by the simplified physical-response layer.
Four tree-based regression models were trained: Random Forest (RF; Breiman, 2001), Gradient Boosting (GB) and Histogram Gradient Boosting (HGB), following the gradient-boosting framework of Friedman (2001), and Extra Trees (ET; Geurts et al., 2006). The individual model predictions were combined into a validation-weighted ensemble, using weights derived from their validation performance (Table 4). The resulting weights are nearly equal, indicating that the four models contributed similarly to the ensemble prediction. This ensemble formulation reduces dependence on a single algorithm and provides a more stable residual correction across the Monte Carlo scenario space.
The residual-surrogate design also supports model interpretation. When the predicted residual is small, the XBeach-informed baseline explains most of the simulated inundation response. When the residual correction is large, feature-importance diagnostics can identify which forcing variables or morphology-related descriptors contribute most strongly to the correction. Permutation importance was calculated on held-out data by shuffling one predictor at a time and measuring the resulting increase in residual RMSE, consistent with permutation-based variable-importance analysis for tree ensembles (Breiman, 2001).
3.5 Validation metrics and uncertainty
Model performance was evaluated using complementary regression, classification, residual, and uncertainty diagnostics. Continuous inundation-depth predictions were assessed using the coefficient of determination (R²), root mean square error (RMSE), mean absolute error (MAE), mean bias, Nash–Sutcliffe efficiency (NSE; Nash and Sutcliffe, 1970), Kling–Gupta efficiency (KGE; Gupta et al., 2009), Pearson correlation, and Spearman rank correlation. These metrics jointly describe error magnitude, systematic overprediction or underprediction, linear and rank-based association, and model skill relative to reference predictions.
Threshold-based flood extent was evaluated by treating each grid node as flooded or non-flooded according to an operational inundation threshold of 0.30 m. This value is used here as a study-specific screening threshold and should not be interpreted as a universal damage threshold; comparable threshold-based evaluations have been used for coastal urban flood surrogates (Zahura et al., 2020). Binary classification skill was quantified using accuracy, precision, recall, F1 score, Matthews correlation coefficient (MCC; Matthews, 1975), Cohen’s kappa (Cohen, 1960), area under the receiver operating characteristic curve (AUC-ROC; Hanley and McNeil, 1982), and average precision. This separate classification assessment is necessary because identifying flooded nodes and reproducing continuous inundation depth represent different levels of predictive difficulty.
Residual behaviour was analysed using residual histograms, fitted normal-density curves, residual-versus-predicted plots, and quantile–quantile (Q–Q) diagnostics. These diagnostics were used to identify systematic bias, non-normal error structure, heteroscedasticity, and tail errors associated with the largest inundation values.
Predictive uncertainty was quantified using split conformal prediction intervals following the distribution-free regression framework described by Lei et al. (2018). Under exchangeability between calibration and test data, split conformal intervals provide finite-sample marginal coverage without requiring a parametric residual distribution. Their practical value for coastal-hazard interpretation nevertheless depends on interval sharpness, calibration across the inundation-depth range, and robustness under distribution shifts, particularly for rare or extreme flood scenarios.”
Sections 5.1–5.3 have been expanded to compare the present results with previous surrogate-modelling studies. Section 5.1 now discusses the commonly reported contrast between high skill in identifying flooded locations and lower skill in reproducing continuous or maximum inundation depth. The revised text also explains why direct numerical comparison among studies is not appropriate when the forcing ranges, reference models, spatial resolutions, morphologies, threshold definitions, and validation strategies differ. Section 5.2 relates the present framework to comparable rapid flood-mapping and early-warning applications and clarifies its intended role as a computational screening and prioritisation layer rather than a replacement for locally calibrated process-based models. Section 5.3 now supports the interpretation of upper-tail underprediction with studies reporting similar limitations arising from rare-event scarcity, incomplete predictor representation, morphology-dependent generalisation, and extrapolation beyond the well-sampled training domain. We have also added specific future-development priorities, including targeted sampling of severe compound events, event-based XBeach calibration, richer morphology descriptors, independent impact validation, and tail-aware uncertainty calibration.
“5.1 Scientific interpretation of the model skill pattern
The main scientific result is not represented by a single performance metric, but by the contrast between threshold-based flood-extent classification and continuous inundation-depth prediction. The surrogate identifies flooded and non-flooded grid nodes with high skill above the 0.30 m operational inundation threshold, but it remains less accurate in reproducing the magnitude of the largest inundation depths. Comparable hybrid hydraulic–machine-learning studies have reported the same qualitative hierarchy of predictive difficulty. Hosseiny et al. (2020) obtained very high wet–dry classification accuracy while the corresponding depth regression remained less exact, and Zahura et al. (2020, 2022) evaluated thresholded inundation extent separately from continuous water depth in physics-based coastal urban surrogates. The present combination of F1 = 0.995 and R² = 0.409 is therefore consistent with the broader observation that delineating inundated locations is generally less demanding than estimating the exact depth at each location. Direct numerical comparison is not appropriate, however, because the studies differ in forcing ranges, reference models, spatial resolution, morphology, threshold definition, and validation design.
The observed skill pattern indicates that the surrogate learns the dominant inundation-response structure and preserves much of the event ranking, but does not yet fully resolve the most severe inundation conditions. High surrogate skill has been reported when models are trained on extensive, high-fidelity and site-specific simulation databases that explicitly represent local geometry and boundary conditions (Ayyad et al., 2022; Ma et al., 2022; Ahmadi Gharehtoragh and Johnson, 2024). Relative to those configurations, the present model uses a deliberately simplified XBeach-informed baseline and idealised spatial descriptors over a morphologically heterogeneous coast. Its strong Pearson and Spearman correlations therefore indicate useful preservation of scenario ordering and spatial response, whereas the moderate R² and systematic upper-tail bias show that response magnitude is compressed. The likely causes include limited representation of rare severe events, the simplified physical baseline, the idealised 10 × 10 grid, and the absence of richer local descriptors such as dune-crest elevation, beach width, backshore slope, harbour exposure, lagoon connectivity, and surface roughness. Consequently, the model is better suited to rapid flood-extent screening and scenario prioritisation than to stand-alone prediction of maximum inundation depth.
5.2 Implications for coastal early warning
The present surrogate is best interpreted as a rapid screening and prioritisation tool rather than as a deterministic predictor of absolute inundation depth. This role is consistent with published surrogate applications developed to accelerate coastal and urban flood assessment. Rapid storm-surge mapping systems and machine-learning emulators have been used to reduce the computational burden of large scenario ensembles and to support faster-than-full-physics evaluation (Yang et al., 2020; Zahura et al., 2020; Ayyad et al., 2022). In a coastal urban setting, Zahura and Goodall (2022) showed that a Random-Forest surrogate could reproduce tidal and pluvial flood patterns within seconds and could be compared with independent crowdsourced flood reports. These studies support the use of surrogates as a computational triage layer that identifies potentially affected locations and prioritises the scenarios requiring detailed hydrodynamic simulation.
This type of framework is relevant for Copernicus downstream coastal services because it links in situ observations, gridded marine and atmospheric products, physical-response modelling, and uncertainty diagnostics within a computationally efficient workflow. Nevertheless, the comparison with previous applications also clarifies the present boundary of use. Operational or near-real-time surrogate systems generally depend on a locally calibrated high-fidelity reference model, representative training events, and independent impact or inundation observations (Yang et al., 2020; Zahura et al., 2020, 2022). In the present study, the XBeach-informed layer is simplified and the storm-sequence diagnostic is an internal consistency check rather than independent validation. The model can therefore provide rapid guidance for scenario exploration, flood-extent screening, and selection of cases for further simulation, but operational quantitative depth forecasting would require locally calibrated XBeach runs, richer morphology, and independent flood, overtopping, or impact observations.
5.3 Why the extreme tail remains difficult
The reduced skill for upper-tail inundation depths is likely caused by a combination of sampling, physical-representation, and spatial-description limitations, and is consistent with a recognised difficulty in statistical and machine-learning representations of extremes. Harter et al. (2024) showed that storm-surge reconstructions commonly underestimate the largest events because extremes are rare, relevant drivers may be omitted, and atmospheric forcing can itself be biased. Ayyad et al. (2022) similarly noted that limited training samples constrain prediction of low-probability, high-consequence surges, while Ahmadi Gharehtoragh and Johnson (2024) found larger surrogate errors for lower annual exceedance probabilities and the most extreme landscape configuration. In the present dataset, severe inundation cases form only a small fraction of the training space, and the concentration of target values near the imposed upper bound of approximately 4 m creates an additional plateau that limits discrimination among the largest responses.
Second, the XBeach-informed physical baseline is intentionally simplified and does not resolve all local processes that may control extreme inundation, including wave-overtopping pathways, backshore storage, small topographic barriers, harbour geometry, lagoon connectivity, and spatial variations in roughness. The resulting pattern—high skill for the inundation footprint but underestimation of large-event depths—has also been reported in a different coastal-inundation context by Ragu Ramalingam et al. (2025), whose machine-learning surrogate reproduced inundation maps more reliably than maximum depths for the largest test events. This comparison does not imply equivalence between storm-surge and tsunami processes, but it illustrates a general surrogate-modelling limitation when the training database or reduced physical representation does not fully span the most severe onshore response.
Third, the spatial grid and morphology descriptors remain idealised. They provide a consistent framework for rapid screening, but they do not fully distinguish between dune-backed beaches, engineered harbour areas, low-lying lagoonal sectors, cliff-backed coasts, and urbanised backshore environments. Studies that explicitly include changing landscape and boundary-condition descriptors show that morphology information can materially improve transfer across scenarios, while the most extreme or least represented landscape states remain more difficult to predict (Ahmadi Gharehtoragh and Johnson, 2024). The simplified geometry used here therefore helps explain why the surrogate captures flood occurrence more reliably than the exact magnitude of upper-tail inundation depths.
Future development should therefore prioritise event-based XBeach calibration at representative coastal sites, targeted sampling of severe compound wave–surge conditions, removal or reformulation of the imposed upper-depth plateau, richer morphology descriptors, and independent validation against observed flood, overtopping, or impact data. Tail-aware training strategies, quantile-oriented objectives, and conformal calibration stratified by depth range or morphology class should also be evaluated, but only after the physical baseline and extreme-event training sample have been strengthened. These steps are required before the framework can support quantitative operational prediction of maximum inundation depth.
These limitations do not reduce the value of the present surrogate as a rapid screening tool; rather, they define the boundary between exploratory flood-extent assessment and operational quantitative inundation-depth forecasting. Within that boundary, the framework provides a transparent means of combining observed and gridded forcing, a physics-informed baseline, machine-learning correction, and explicit uncertainty diagnostics.”
Citation: https://doi.org/10.5194/egusphere-2026-3628-AC1
-
AC1: 'Reply on RC1', Maria Mihailov, 23 Jul 2026
reply
-
RC2: 'Comment on egusphere-2026-3628', Anonymous Referee #2, 06 Aug 2026
reply
In this paper, entitled “An XBeach-Informed Machine-Learning Surrogate for Rapid Coastal Flood Prediction”, the authors develop a machine-learning surrogate model for coastal flood forecasting. Although the topic is interesting, the reviewer believes that the manuscript contains several issues that preclude its publication in its current form. My comments and suggestions are reported below.
My major concern relates to the fact that the methodology for predicting coastal flooding has been developed and validated using tide-gauge observations of sea level rather than direct inundation data. Tide gauges measure sea level near the coast and, because they are typically installed within harbors or near coastal structures, they cannot adequately represent coastal flooding processes. In particular, they do not account for wave-induced effects such as wave setup and runup, and their records may be influenced by harbor resonance. This appears to be the case in the present study, where the tide gauges at Constanța, Mangalia, and Midia are located within well-protected harbors, while the coordinates reported for Gura Portiței indicate an offshore location. Indeed, Figures 5 and 7 suggest that waves have a negligible influence on the estimated inundation depth. Moreover, the selection of the flooding threshold should be related to the specific characteristics of each coastal segment. In the present work, however, a single threshold value has been adopted for all investigated coastal sectors. This also raises the question of how sensitive the flood predictions are to the choice of this threshold. For these reasons, although the proposed modelling framework may be conceptually sound, it cannot be properly applied to and validated for the western Black Sea coast using the available data. I therefore strongly encourage the authors to seek and incorporate direct inundation observations. In my opinion, without such data, the methodology cannot be adequately tested and the manuscript is not suitable for publication.
Secondly, the manuscript lacks a clear and comprehensive description of the methodology. This aspect should be substantially improved to enable readers to fully understand the proposed approach. In particular: i) the Methods section should be reorganized to follow the workflow illustrated in Fig. 3 (observations, XBeach model, Monte Carlo scenarios, surrogate model, etc.); ii) it is not clear whether the Monte Carlo scenarios were generated using all sea-level observations from the different tide gauges collectively, or whether separate scenarios were developed for each coastal sector; iii) the description of the XBeach model should also provide details of the simulation setup, including the model domain, boundary and forcing conditions, spatial resolution, time step, and the representation of coastal and harbor structures; iv) the characteristics of the 10 × 10 spatial target grid used by the surrogate model should be reported, including its extent, resolution, and the bathymetric and topographic datasets employed; v) a clear description on how the dataset have been used for training and testing the surrogate models is missing; vi) it is unclear how the 0D tide-gauge observations, the 1D XBeach profiles, and the 2D surrogate-model grid are integrated.
Thirdly, the authors should pay greater attention to the use of appropriate technical terminology. Below are some examples of terms that are either unclear or potentially misleading:
-
Coastal flood depth, inundation depth, inundation field, inundation response, spatial inundation. In my opinion, the term inundation is not appropriate in this context, since both the observations and the modelling framework appear to represent the (total) sea level rather than actual land flooding. A clear definition of inundation depth should be clearly stated at the beginning of the Methods section.
-
XBeach-informed physical baseline, physics baseline. These terms appear to refer simply to the outputs generated by the XBeach model. I therefore suggest replacing them with a more straightforward expression such as XBeach outputs or XBeach simulations. The terms informed, physical, and baseline add unnecessary complexity and may confuse the reader.
-
Storm surge, storm-surge setup, sea-level anomaly. Tide gauges record the total sea level and additional processing is required to separate and quantify the different components of sea-level variability (e.g., storm surge, sea-level anomaly, astronomical tide). The manuscript should clearly explain how these quantities were derived and used.
-
Storm sequence: It is unclear what is meant by this term. The authors may be referring to storm duration, a storm event, or a storm case. Please either use a more precise term or clearly define its meaning in the manuscript.
Lastly, the results indicate that the different surrogate models exhibit very similar performance. I therefore suggest presenting only the ensemble mean results in the main manuscript and moving the complete set of performance statistics to the Supplementary Material. This would improve the readability of the paper without affecting the interpretation of the results. Furthermore, the authors should consider that these findings may be strongly influenced by the methodological limitations highlighted in the above concerns. Therefore, the reported model performance should be interpreted with caution until the proposed framework is validated against appropriate inundation observations.
Additional comments:
-
As it stands, the abstract reads more like that of a paper intended for a national or regional journal. Moreover, it contains unnecessary details on the datasets and statistical results, including unexplained acronyms. The abstract should be streamlined to emphasize the study's objectives, methodology, key findings, and significance.
-
Lines 113-126: details about the datasets, model performance metrics and predictors should not be reported here.
-
Lines 128-129: the flood extent, expressed as grid nodes exceeding an operational inundation, does not implies continuity of the flooding area. Please discuss the limitation of this approach.
-
exceeding an operational inundation threshold
-
Lines 141-144: is this a result of the present study? If so, it should be moved to the Results section. Otherwise, an appropriate reference should be provided.
-
Line 154: main Romanian Western coastal sectors.
-
Table 1: include the temporal extension and the number of valid data for each station.
-
Figure 2: This figure suggests that the tide gauges use different vertical reference datums, while Table 1 reports very similar mean SLEV values. Please clarify whether the data were adjusted to a common datum before the analysis.
-
Line 175: what do you mean with related hydrodynamic variables. You should only mention the variables that have been used in the present study.
-
Line 182: provide the reference for the topo-bathymetric dataset.
-
Table 2: many variables (SLEV, ATMS, Hs, Tp, sea-surface state, SST, MSLP) and roles (storm surge, erosion-related forcing, plume related forcing, …) need to better defined.
-
Line 203: define “target inundation response”.
-
Lines 214-216: the variables storm duration, inverse-barometer contribution, wave-power-related terms, and interaction terms between sea-level anomaly and wave forcing must be clearly defined.
-
Table 3: The mean values obtained from the Monte Carlo simulations (MCS) are significantly higher than the observed means. This discrepancy is particularly evident for sea level, for which the simulated mean exceeds the observed mean by more than 40%. The authors should provide an explanation for this difference and discuss its potential implications for the modelling results.
-
Lines 229-230: define “design-of-experiments forcing combinations”.
-
Line 230: why sea-level anomaly?
-
Line 231: specify the topo-bathymetric dataset used for implementing the XBeach cross-sections.
-
Lines 234-236: I do not agree with this approach. Each component of the modelling system must be calibrated and verified to obtain a reliable prediction system.
-
Lines 236-238: this part should be moved to section 3.4.
-
Figure 5: The left panel clearly shows that the XBeach profiles provide an unrealistic representation of the offshore/submerged part of the coastal profile, including a vertical wall exceeding 5 m in height. The authors should explain the origin of this feature and discuss its potential effects on the model results.
-
Line 291: from where do you get observations of the inundation depth at grid-node level?
-
Lines 296, tables 5 and 6, and somewhere else: always specify “mean” when referring to average value of the surrogate models.
-
Table 5: why the physics baseline has negative R2?
-
Figure 6: What’s the Observed Coastal Flood Depth? Tide gauge measure sea level not flood depth.
-
Captions of figures 6, 7, 8 and 9: do not report comments in the captions.
-
Section4.5: explain how the permutation-importance diagnostics have been computed.
-
Line 353: clarify how the interaction between sea-level anomaly and significant wave height have been defined.
-
Line 356: no compound sea-level–wave interaction effects can be reliably assessed in this study, as the observations were collected from tide gauges located within harbors, and the XBeach simulations include an unrealistic 5 m vertical wall structure.
-
It is not necessary to include both Figure 7 and Table 8 as they report the same information.
-
Section 4.7: from where these other stations come from?
-
Figure 9b: do not mix observation with XBeach model output.
Citation: https://doi.org/10.5194/egusphere-2026-3628-RC2 -
-
RC3: 'Comment on egusphere-2026-3628', Anonymous Referee #3, 14 Aug 2026
reply
Recommendation: major revision
General comment:
The manuscript develops a tree-ensemble surrogate for rapid coastal-inundation screening along the Western (Romanian) Black Sea coast, trained on a Monte Carlo forcing ensemble and referenced to an intermediate “XBeach-informed” physical-response layer. The topic is timely and operationally relevant: reducing the computational cost of process-based inundation modelling while keeping a traceable link to the controlling physical drivers is a real need for coastal early warning and for Copernicus downstream services. The assembly of a coherent forcing chain for this specific coastal sector — high-frequency MHD tide-gauge records, Copernicus Marine Black Sea wave and physical products, Copernicus L4 SST, ERA5 atmospheric fields and GRDC Danube discharge — is a genuine contribution, and to my knowledge no comparable multi-driver surrogate has been published for the Western Black Sea.
I also want to acknowledge the diagnostic honesty of the manuscript. The authors do not hide the weak performance in the upper tail; they deliberately separate threshold-based flood-extent classification from continuous depth regression, they report conformal prediction intervals including their poor sharpness, and they provide residual, Q–Q, permutation-importance, site-specific and seasonal diagnostics. This level of transparency is above average for surrogate-modelling papers, and the resulting discussion (Sects. 5.1 and 5.3) is thoughtful.
That said, in its present form the manuscript is not fulfilling its central claims, for the following reasons:
- The manuscript refers to “the target inundation response” (e.g. L246, L285–L291) but nowhere states how the 5,000 × 100 target matrix was actually computed. This is the single most important omission, because every metric in the paper is computed against this quantity. Another misleading point is that the XBeach-informed response layer of Fig. 5c gives a maximum flood depth of 0.90 m across the entire design-of-experiments envelope, including a sea-level anomaly approaching 1 m combined with Hs above 5 m. Consistently with this, Fig. 5b gives a resulting depth of 0.1–0.8 m, Fig. 9b a peak inundation envelope below 0.9 m, and Fig. A2 a flood-depth proxy below 0.8 m with R2 % run-up below 1.6 m. Table 7, by contrast, reports that the 90th percentile of the target inundation depth is 3.966 m and that the distribution saturates at exactly 4.000 m. Please describe better or fix this.
- XBeach appears in the title and names the central methodological component, yet the manuscript gives no model version, no computational domain or grid resolution, no hydrodynamic or morphodynamic parameter settings, no boundary-condition specification, no number of simulations performed, and no calibration or validation of the XBeach results against observations. As it is, a reader cannot reproduce the physical baseline, and the “XBeach-informed” framing of the title is not yet demonstrated, thus please provide xbeach model details or a literature reference to support the data.
- Some reported diagnostics contradict each other: Table 5 gives the physics baseline a Pearson r of 0.026 and an NSE of −0.479, and Table 6 gives it MCC = −0.0028 — i.e. no skill whatsoever — yet Table 8 and Fig. 7 identify physics_baseline_m as by far the dominant predictor of the residual. A predictor that carries essentially no information about the target cannot simultaneously be the leading predictor. Further inconsistencies are listed in Sect. 2 below (Fig. 6a annotations vs. Table 5; Table C1 has several values repeated with no particular value for the manuscript; the 0.95 and 0.99 rows of Table 7).
- Beyond these three, two further points seem to me important enough to raise at this level. The surrogate takes sea level as a prescribed input rather than predicting it, so all reported skill is conditional on perfect knowledge of the forcing and is an optimistic bound on what the framework would achieve in forecast mode (M10). And the training loss function — the most likely proximate cause of the upper-tail bias that the manuscript identifies as its own main limitation — is never stated or discussed, nor is any tail-sensitive verification metric used for model selection (M11, S1).
In addition, an evaluation comparing against any observed flooding, overtopping or damage record would greatly improve the manuscript, so achieving the objective indicated in the abstract and title which promise “spatial flood prediction”. Although the 10 × 10 target grid is described as idealised and is never geo-referenced, given a resolution, or linked to the coastline mapped in Fig. 1. The claim of computational speed-up, which one of the motivation is never quantified: no timings are reported for either XBeach or the surrogate. In some cases, the manuscript is repetitive, with the “classification good / regression poor” message restated in the abstract, Sect. 1, Sects. 4.2 and 4.3, Sects. 5.1 and 5.2 and the conclusions, so it could be shortened appreciably without any loss of content.
I want to close this section by repeating that the underlying idea is interesting and that the regional dataset is worth publishing, but the manuscript needs a major revision.
Major points to address:
- Please state what the 5,000 × 100 target matrix contains and how it was generated. Specifically: is each target value the maximum inundation depth extracted from an actual XBeach 2D simulation for that forcing scenario? What it the location, the timing,the model configuration? What are the forcing used?
- Please check the range of the target with the range of the intermediate response layer. As set out in the general comments, Fig. 5b–c, Fig. 9b and Fig. A2 all bound flood depth below about 0.9 m over the full forcing envelope, while Table 7 is about 4 m.
- Table 7 has some mistakes. The test set contains 115,200 nodes — this is the sum of both Table 10 and Table 11, which is reassuring — so the 0.90, 0.95 and 0.99 quantiles should select 11,520, 5,760 and 1,152 nodes respectively. The 0.90 row is consistent (n = 11,520), but the 0.95 and 0.99 rows both report n = 6,034 with an identical threshold of 4.0 m. As stated by the authors, there is a clip of the data, thus the latter 2 percentiles have no particular relevance. However is not clear where 4m comes from and what is the variable interested.
- The consequence is that the headline “extreme-event underprediction of −2.59 m” has 2 points to focus on: this is referred to inundation depth, which should be better detailed in the manuscript. In this version is not exactly well defined and clear. However, it should be not a physical finding about compound storm response, but it is the arithmetic difference between two saturation levels, one imposed on the target and one learned by the model.
- Please check the standard deviation in Table 1 which seems are not aligned with Figure 2: the std is very similar between the 4 sites, but the timeseries, in particular Gura seems to be quite different
- Figure 2 shows that the four stations cover very different periods, with large gaps and different ranges. The Table C2 reports 35,064 valid records for every variable, which is exactly four years of hourly data. Please state the common reference period, how gaps were treated, how much inputs were used for the analysis.
- Table 10 and Sect. 4.7 report site-specific metrics for Constanța, Eforie Nord, Gura Portiței, Mamaia, Mangalia, Midia and Vama Veche, and the text singles out Mamaia and Midia. However, Figs. 1, 2 and 8, and Table 1 describe only four monitoring stations. Where do the profiles, forcing and targets for Eforie Nord, Mamaia and Vama Veche come from?
- The held-out evaluation throughout was never defined. Please state explicitly whether the split is by scenario (grouped) or by row, give the split fractions, and confirm that the reported test set of 115,200 nodes contains no scenario also seen in training.
- All four models are tree ensembles that, in their standard implementations, minimise a squared-error criterion, and their predictions are averages over the training samples falling in the terminal leaves. Such models cannot extrapolate beyond the training range and shrink towards the conditional mean precisely where the target density is lowest. Yet the manuscript nowhere states which loss or split criterion was used for each model, and never considers alternatives. Given that upper-tail underprediction is presented as the limitation of the study, this deserves a dedicated subsection rather than no mention at all.
Concretely, I suggest the authors (a) state the loss and split criterion for each of RF, GB, ET and HGB, together with all hyperparameters and how they were tuned (currently absent entirely); (b) test at least one tail-oriented alternative — quantile regression, an asymmetric or extreme-weighted loss, sample weighting proportional to target exceedance, or a two-stage design with a dedicated upper-tail regressor above the 90th percentile; and (c) add at least one verification metric that is not dominated by the bulk of the distribution.
There is a coherent and rapidly growing literature on exactly this problem in the surge and flood emulation context that the manuscript does not engage with, although it already cites two papers from the same strand (Longo et al., 2026; Wilkinson et al., 2026). Hermans et al. (2025) show that densifying the loss towards extremes substantially improves neural-network estimates of extreme surges in Europe. Campos-Caba et al. (2026) benchmark emulators ranging from multivariate linear regression to LSTM networks against a 50 m coastal-resolution hydrodynamic model in the northern Adriatic and find that the choice of loss function matters more than model complexity: a linear model trained with an extreme-oriented loss (the corrected mean absolute deviation squared, MADc²) matches an LSTM at orders of magnitude lower computational cost — a result with obvious implications for the four-model ensemble of Sect. 3.4. Campos-Caba et al. (2024) address the companion question of which error indicators are capable of measuring skill on extremes at all.
- For a single dataset evaluated against the mean of the reference values, the coefficient of determination in its NSE form and Nash–Sutcliffe efficiency are identical by definition — which is why the abstract reports R² = 0.409 and NSE = 0.409, and why the two columns of Table 5 are identical. Please drop one, or state clearly that R² is the squared Pearson correlation (in which case it should equal 0.914² = 0.836, not 0.409) and give the definition used. More broadly: every regression metric in Table 5 — R², RMSE, MAE, NSE, KGE, r, ρ — is dominated by the bulk of the distribution, and none is sensitive to the upper tail, which is where the model actually fails. The tail behaviour appears only in Table 7, and not used for any model decision. I would recommend adding at least one extreme-oriented indicator to Table 5, for example a peak-over-threshold weighted error, a conditional bias above a fixed quantile, or the corrected mean absolute deviation of Campos-Caba et al. (2024), and using it in model selection alongside RMSE.
- The four “validation-derived” weights are 0.2499, 0.2501, 0.2499 and 0.2501, i.e. indistinguishable from a plain arithmetic mean, and the ensemble does not outperform the best individual model (ensemble R² = 0.4087 vs. HGB R² = 0.4091; ensemble F1 = 0.9948 vs. HGB F1 = 0.9963). Please state how the weights were derived and over what quantity they were optimised, and either demonstrate a benefit or simplify to a single model or a plain average. As written, the ensemble is presented as a methodological feature (L250–L255) that the results do not support.
- Please add a table listing all 22 forcing and derived predictors, with units, source dataset and, for derived quantities, the exact formula. This is essential for reproducibility
- Speed is the entire motivation for the work (“rapid” appears in the title, abstract and conclusions) but no timings are given anywhere. Please report the wall-clock cost of one XBeach simulation on stated hardware, the total cost of building the training set, the surrogate training time, and the inference time per scenario, so that the claimed speed-up can be assessed quantitatively.
Figure 1: the map resolution is too low to be legible — station labels, coastline and scale bar are hard to read at print size. Please supply a higher-resolution version, and consider adding bathymetric contours and the seven evaluation sites of Table 10.
Every Copernicus Marine product has a persistent DOI or product identifier. Please add them for all products used, in Table 2 and in the reference list, together with product version and the temporal coverage actually downloaded.
L174: add the full reference and product identifier for the Copernicus Marine Black Sea Physics Analysis and Forecast product.
L176–L177: add the reference and product identifier for the Copernicus Marine Black Sea high-resolution Level-4 SST product. The two Buongiorno Nardelli papers currently cited at L106 describe the methodology rather than the operational product, and both are needed.
L182: the bathymetric and coastal-elevation data need a proper description and citation — source, horizontal resolution, vertical datum, year, and how EMODnet bathymetry was merged with the topographic data. The EMODnet entry in the reference list has no URL, DOI or access date.
Figure 2: the four panels use different y-axis ranges, which makes visual comparison misleading (Gura Portiței spans −0.2 to +0.1 m, the others −0.5 to +0.5 m or wider). Please use a common axis or explain why not. The image resolution should be increased. The green color in Gura Portiței is very difficult to read.
Figure 9a: the horizontal axis is labelled “Operational Radar OTT RLS Tide-Gauge Stations Network”, refers to an instrument type which is introduced here for the first time, it should be included in the text in the observation sdescription.
The threshold is described as “operational” (L269, L314) please add a reference or rationale.
Figure B2: please improve resolution
Units are mixed between text, tables and figures: some Tables use centimetres, Table 8 uses metres, and the text uses metres throughout. Please use metres consistently, and state the unit in every axis label and column header.
Table 6: MCC and Cohen’s κ agree to four decimals in every row, so reporting both adds no information. Two would be enough
Table 7: No needs of 4 decimals in the quantiles
Sections 4.2, 4.3, 5.1, 5.2 and 6 repeat the same interpretation of the classification/regression contrast. Consolidating it into Sect. 5.1 would shorten the manuscript appreciably and make the argument stronger.
Citation: https://doi.org/10.5194/egusphere-2026-3628-RC3
Viewed
| HTML | XML | Total | BibTeX | EndNote | |
|---|---|---|---|---|---|
| 114 | 41 | 15 | 170 | 10 | 11 |
- HTML: 114
- PDF: 41
- XML: 15
- Total: 170
- BibTeX: 10
- EndNote: 11
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
The manuscript provides a robust and comprehensive description of the materials and methods. However, several aspects could be improved to enhance clarity and strengthen the scientific presentation.
In the Introduction, describing 7 specific objectives is unnecessarily detailed and makes the section appear repetitive, as it lists individual analytical steps rather than the overarching research goals. Would be positive to restructuring them into three broader objectives following a clearer logical progression: (1) data and framework development; (2) model development and validation; and (3) process interpretation and practical application.
The Methods section is comprehensive and well documented. However, publicly available datasets should also be listed in the Data Availability Statement, including their repositories, persistent links (e.g., DOI or URL), and data providers. In addition, datasets, source code, or algorithms that are available upon reasonable request could also be mentioned to further strengthen compliance with the FAIR principles. Adopting these recommendations would improve the transparency, reproducibility, and overall impact of the study. Furthermore, several methodological choices and parameter settings are presented without supporting references. Although many of these statistical approaches are well established, citing the original or standard methodological references would better justify their selection and increase the scientific credibility of the adopted framework.
The Discussion provides a clear interpretation of the results but would benefit from a stronger comparison with previous studies. While the Introduction section is well referenced, Section 5.1 would be strengthened by discussing similar studies and demonstrating that the observed model-skill patterns are consistent with findings reported in the literature. Likewise, Section 5.2 highlights the potential and limitations of surrogate models for early-warning systems but does not include references to comparable applications using similar methodologies. Finally, the discussion in Section 5.3 regarding the reduced performance for extreme-tail events should also be supported by references to studies reporting similar limitations, as this is a common challenge in surrogate modelling and extreme-event prediction.
Overall, the study successfully addresses its stated objectives and presents a framework with significant practical relevance. The proposed methodology has considerable potential for regional coastal-flood hazard assessment and provides a valuable reference for the development of operational early-warning systems. In particular, the framework demonstrates an appropriate balance between computational efficiency, physical realism, and transparent communication of uncertainties, making it well suited to support coastal management and decision-making for compound coastal flooding.