the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
An XBeach-informed machine-learning surrogate for rapid coastal flood prediction
Abstract. Rapid assessment of coastal inundation along the Western Black Sea coast requires computationally efficient methods that can screen multiple compound storm scenarios while preserving a clear link to the controlling physical drivers. This study develops and evaluates an XBeach-informed machine-learning surrogate for coastal inundation and threshold-based flood-extent screening. The framework combines high-frequency sea-level observations from the Maritime Hydrographic Directorate monitoring network, Copernicus Marine wave and Black Sea physical products, ERA5 atmospheric forcing, Global Runoff Data Centre Danube discharge, bathymetric and coastal-elevation descriptors, and an intermediate cross-shore physical-response layer. The modelling dataset consists of 5,000 Monte Carlo forcing scenarios, 22 forcing and derived predictors, and a 10 × 10 spatial target grids, corresponding to 100 inundation-depth nodes per scenario. Four tree-based residual models, namely Random Forest, Gradient Boosting, Extra Trees, and Histogram Gradient Boosting, are combined using validation-derived weights.
The conservative held-out node-level evaluation shows moderate quantitative skill for continuous inundation-depth prediction. The ensemble reaches R² = 0.409, RMSE = 1.181 m, MAE = 0.678 m, NSE = 0.409, KGE = 0.235, Pearson r = 0.914, and Spearman ρ = 0.923. In contrast, threshold-based flood-extent classification above the 0.30 m operational inundation threshold is substantially stronger, with F1 = 0.995, AUC-ROC = 0.9998, Matthews correlation coefficient = 0.989, and Cohen’s κ = 0.989. Split conformal prediction intervals provide near-nominal 90 % marginal coverage, with empirical coverage of 0.911, but the mean 90 % interval width is large, 3.25 m, indicating limited sharpness for local quantitative depth estimation. Extreme-event diagnostics show systematic underprediction of upper-tail inundation depths, with a mean bias of approximately −2.59 m for events above the 90th percentile.
The present configuration is therefore most appropriate for rapid flood-extent screening, scenario ranking, and identification of cases where more detailed process-based simulations are required. It should not yet be interpreted as a stand-alone operational predictor of maximum inundation depth. The study demonstrates how in situ observations, Copernicus Marine products, physical-response modelling, machine-learning emulation, and uncertainty diagnostics can be combined into a transparent coastal-hazard screening workflow for the Western Black Sea coast.
- Preprint
(1655 KB) - Metadata XML
- BibTeX
- EndNote
Status: open (until 09 Sep 2026)
-
RC1: 'Comment on egusphere-2026-3628', Anonymous Referee #1, 21 Jul 2026
reply
-
AC1: 'Reply on RC1', Maria Mihailov, 23 Jul 2026
reply
Thank you for this constructive suggestion.
Response to Comment 1: We agree that the previous formulation described several methodological steps separately and therefore masked the broader progression of the study. We have revised the final part of the Introduction by consolidating the objectives into three overarching themes: (1) data and framework development; (2) model development and validation; and (3) process interpretation and practical application. We also reformulated the following paragraph so that continuous inundation-depth prediction, threshold-based flood-extent identification, and prediction-interval performance are presented as complementary dimensions of model evaluation rather than as a second list of objectives. We hope that this revision reduces repetition and clarifies the connection between the study design, model assessment, physical interpretation, and intended coastal-hazard application.
“The objective of this paper is to develop and evaluate a physically informed machine-learning surrogate for rapid coastal-inundation prediction and flood-extent screening along the Western Black Sea coast. The study is structured around three main objectives: (1) data and framework development, through the integration of in situ sea-level observations, Copernicus Marine products, ERA5 atmospheric forcing (Hersbach et al., 2020, 2023), Danube discharge (GRDC, 2026), and coastal-morphology descriptors within an XBeach-informed reference framework (Roelvink et al., 2009) and a Monte Carlo scenario space; (2) model development and validation, through the training and comparison of tree-based residual surrogates for spatial inundation-depth prediction and threshold-based flood-extent identification, together with a comprehensive assessment of predictive skill, residual behaviour, extreme-event performance, and uncertainty; and (3) process interpretation and practical application, through the identification of the dominant physical controls on the surrogate response and the evaluation of its suitability and limitations for rapid coastal-flood screening, scenario prioritisation, and Copernicus downstream decision support.”
Response to Comment 2: The Code and data availability section has been expanded to identify the publicly accessible datasets used in the study, together with their data providers, repositories, product identifiers, and persistent DOI or portal addresses. The revised statement now includes the FOCCUS/MHD sea-level observations, the Copernicus Marine Black Sea wave, physical-state, and sea-surface-temperature products, ERA5 atmospheric forcing, GRDC Danube discharge records, and EMODnet Digital Bathymetry. We also clarify the applicable GRDC data-sharing restrictions. The processed forcing matrix, Monte Carlo scenarios, XBeach-informed target tables, validation outputs, trained model objects, environment information, source code, and figure-generation scripts will be deposited in a DOI-bearing Zenodo repository associated with the final publication. Until this repository is released, the authors’ processed datasets, code, and algorithms are available from the corresponding authors upon reasonable request, subject to third-party licensing and redistribution restrictions. The planned repository metadata will document product identifiers, versions, retrieval dates, processing steps, and software dependencies to support reproducibility and FAIR reuse.
“All input datasets used in this study are publicly available online. Sea-level elevation and atmospheric-pressure observations from the Romanian Black Sea stations at Constanța, Gura Portiței, Mangalia, and Midia were provided by the Maritime Hydrographic Directorate and are distributed through Zenodo (Dutu et al., 2026, https://doi.org/10.5281/zenodo.19133361). Black Sea wave data were obtained from the Copernicus Marine Service Black Sea Waves Analysis and Forecast product (BS-Waves, EAS5; Staneva et al., 2022, https://doi.org/10.25423/CMCC/BLKSEA_ANALYSISFORECAST_WAV_007_003_EAS5). Black Sea physical-state data were obtained from the Copernicus Marine Service Black Sea Physics Analysis and Forecast product (BLK-PHY, EAS7; Jansen et al., 2024, https://doi.org/10.48670/mds-00355). High-resolution Level-4 sea-surface-temperature data were obtained from the Copernicus Marine Service (Buongiorno Nardelli et al., 2010, 2013, https://doi.org/10.48670/moi-00159). Hourly ERA5 atmospheric data on single levels, including the 10 m wind components and mean sea-level pressure, were obtained from the Copernicus Climate Change Service Climate Data Store (Hersbach et al., 2020, 2023, https://doi.org/10.24381/cds.bd0915c6). Daily Danube River discharge data were obtained from the Global Runoff Data Centre, Federal Institute of Hydrology, Koblenz, Germany (GRDC, 2026, https://portal.grdc.bafg.de/, last access: 20 June 2026). Bathymetric and coastal-elevation data were obtained from the EMODnet Bathymetry Consortium (2026, https://doi.org/10.12770/bb6a87dd-e579-4036-abe1-e649cea9881a).
The computational workflow was implemented in Python using a notebook-based workflow developed in the Jupyter environment (Kluyver Thomas et al., 2016) and managed through the Anaconda Python distribution. The computational workflow was implemented in Python and includes data alignment, Monte Carlo scenario generation, construction of the XBeach-informed physical-response layer, residual-surrogate training, model validation, uncertainty analysis, and figure production. The processed model-input matrices, Monte Carlo scenarios, spatial target data, trained model objects, validation outputs, and scripts used to generate the results presented in this study will be made available in a DOI-bearing repository accompanying the final publication, subject to third-party data-licence constraints.
All metrics, figures and tables reported in this manuscript are based on the aligned observational and gridded forcing dataset described in Sects. 2 and 3.”
Response to Comment 3: Additional methodological references have been introduced in Sections 3.2, 3.4, and 3.5. These references support the use of block-bootstrap sampling, Earth Mover’s Distance, the selected tree-ensemble algorithms, permutation-based feature importance, the NSE and KGE performance criteria, the MCC, Cohen’s kappa and ROC-AUC classification metrics, and split conformal prediction intervals. We have also clarified that the 0.30 m inundation threshold is a study-specific operational screening threshold and should not be interpreted as a universal flood-damage threshold.
“ 3.2 Monte Carlo scenario generation
A Monte Carlo simulation (MCS) framework was used to generate 5,000 forcing scenarios, providing broad coverage of the multivariate forcing space while keeping the resulting 500,000 node-level samples computationally tractable for repeated model training and validation. The scenario matrix includes both sampled variables and process-derived predictors, including wind speed and direction, significant wave height, peak wave period, wave direction, sea-level elevation, mean sea-level pressure, Danube discharge, sea-surface temperature, storm duration, inverse-barometer contribution, wave-power-related terms, and interaction terms between sea-level anomaly and wave forcing. Where the dependence structure between variables was relevant, empirical or block-bootstrap sampling was used to retain part of the observed co-variability, following the general block-resampling approach of Künsch (1989). The consistency of the sampled variables with the observed or aligned forcing data was evaluated using distributional comparisons, including histograms and kernel-density estimates, together with Earth Mover’s Distance (EMD) as a quantitative measure of marginal distributional agreement (Rubner et al., 2000; Table 3; Figure 4).
(……)
3.4 Hybrid residual surrogate
The final surrogate was formulated as a residual correction to the XBeach-informed physical baseline. In this formulation, the physical layer first provides a baseline estimate of inundation depth, and the machine-learning component then predicts the residual term, defined as the difference between the target inundation response and the baseline estimate. This structure preserves the physical interpretability of the model because the baseline response remains explicitly visible, while the residual model accounts for the part of the inundation signal that is not captured by the simplified physical-response layer.
Four tree-based regression models were trained: Random Forest (RF; Breiman, 2001), Gradient Boosting (GB) and Histogram Gradient Boosting (HGB), following the gradient-boosting framework of Friedman (2001), and Extra Trees (ET; Geurts et al., 2006). The individual model predictions were combined into a validation-weighted ensemble, using weights derived from their validation performance (Table 4). The resulting weights are nearly equal, indicating that the four models contributed similarly to the ensemble prediction. This ensemble formulation reduces dependence on a single algorithm and provides a more stable residual correction across the Monte Carlo scenario space.
The residual-surrogate design also supports model interpretation. When the predicted residual is small, the XBeach-informed baseline explains most of the simulated inundation response. When the residual correction is large, feature-importance diagnostics can identify which forcing variables or morphology-related descriptors contribute most strongly to the correction. Permutation importance was calculated on held-out data by shuffling one predictor at a time and measuring the resulting increase in residual RMSE, consistent with permutation-based variable-importance analysis for tree ensembles (Breiman, 2001).
3.5 Validation metrics and uncertainty
Model performance was evaluated using complementary regression, classification, residual, and uncertainty diagnostics. Continuous inundation-depth predictions were assessed using the coefficient of determination (R²), root mean square error (RMSE), mean absolute error (MAE), mean bias, Nash–Sutcliffe efficiency (NSE; Nash and Sutcliffe, 1970), Kling–Gupta efficiency (KGE; Gupta et al., 2009), Pearson correlation, and Spearman rank correlation. These metrics jointly describe error magnitude, systematic overprediction or underprediction, linear and rank-based association, and model skill relative to reference predictions.
Threshold-based flood extent was evaluated by treating each grid node as flooded or non-flooded according to an operational inundation threshold of 0.30 m. This value is used here as a study-specific screening threshold and should not be interpreted as a universal damage threshold; comparable threshold-based evaluations have been used for coastal urban flood surrogates (Zahura et al., 2020). Binary classification skill was quantified using accuracy, precision, recall, F1 score, Matthews correlation coefficient (MCC; Matthews, 1975), Cohen’s kappa (Cohen, 1960), area under the receiver operating characteristic curve (AUC-ROC; Hanley and McNeil, 1982), and average precision. This separate classification assessment is necessary because identifying flooded nodes and reproducing continuous inundation depth represent different levels of predictive difficulty.
Residual behaviour was analysed using residual histograms, fitted normal-density curves, residual-versus-predicted plots, and quantile–quantile (Q–Q) diagnostics. These diagnostics were used to identify systematic bias, non-normal error structure, heteroscedasticity, and tail errors associated with the largest inundation values.
Predictive uncertainty was quantified using split conformal prediction intervals following the distribution-free regression framework described by Lei et al. (2018). Under exchangeability between calibration and test data, split conformal intervals provide finite-sample marginal coverage without requiring a parametric residual distribution. Their practical value for coastal-hazard interpretation nevertheless depends on interval sharpness, calibration across the inundation-depth range, and robustness under distribution shifts, particularly for rare or extreme flood scenarios.”
Sections 5.1–5.3 have been expanded to compare the present results with previous surrogate-modelling studies. Section 5.1 now discusses the commonly reported contrast between high skill in identifying flooded locations and lower skill in reproducing continuous or maximum inundation depth. The revised text also explains why direct numerical comparison among studies is not appropriate when the forcing ranges, reference models, spatial resolutions, morphologies, threshold definitions, and validation strategies differ. Section 5.2 relates the present framework to comparable rapid flood-mapping and early-warning applications and clarifies its intended role as a computational screening and prioritisation layer rather than a replacement for locally calibrated process-based models. Section 5.3 now supports the interpretation of upper-tail underprediction with studies reporting similar limitations arising from rare-event scarcity, incomplete predictor representation, morphology-dependent generalisation, and extrapolation beyond the well-sampled training domain. We have also added specific future-development priorities, including targeted sampling of severe compound events, event-based XBeach calibration, richer morphology descriptors, independent impact validation, and tail-aware uncertainty calibration.
“5.1 Scientific interpretation of the model skill pattern
The main scientific result is not represented by a single performance metric, but by the contrast between threshold-based flood-extent classification and continuous inundation-depth prediction. The surrogate identifies flooded and non-flooded grid nodes with high skill above the 0.30 m operational inundation threshold, but it remains less accurate in reproducing the magnitude of the largest inundation depths. Comparable hybrid hydraulic–machine-learning studies have reported the same qualitative hierarchy of predictive difficulty. Hosseiny et al. (2020) obtained very high wet–dry classification accuracy while the corresponding depth regression remained less exact, and Zahura et al. (2020, 2022) evaluated thresholded inundation extent separately from continuous water depth in physics-based coastal urban surrogates. The present combination of F1 = 0.995 and R² = 0.409 is therefore consistent with the broader observation that delineating inundated locations is generally less demanding than estimating the exact depth at each location. Direct numerical comparison is not appropriate, however, because the studies differ in forcing ranges, reference models, spatial resolution, morphology, threshold definition, and validation design.
The observed skill pattern indicates that the surrogate learns the dominant inundation-response structure and preserves much of the event ranking, but does not yet fully resolve the most severe inundation conditions. High surrogate skill has been reported when models are trained on extensive, high-fidelity and site-specific simulation databases that explicitly represent local geometry and boundary conditions (Ayyad et al., 2022; Ma et al., 2022; Ahmadi Gharehtoragh and Johnson, 2024). Relative to those configurations, the present model uses a deliberately simplified XBeach-informed baseline and idealised spatial descriptors over a morphologically heterogeneous coast. Its strong Pearson and Spearman correlations therefore indicate useful preservation of scenario ordering and spatial response, whereas the moderate R² and systematic upper-tail bias show that response magnitude is compressed. The likely causes include limited representation of rare severe events, the simplified physical baseline, the idealised 10 × 10 grid, and the absence of richer local descriptors such as dune-crest elevation, beach width, backshore slope, harbour exposure, lagoon connectivity, and surface roughness. Consequently, the model is better suited to rapid flood-extent screening and scenario prioritisation than to stand-alone prediction of maximum inundation depth.
5.2 Implications for coastal early warning
The present surrogate is best interpreted as a rapid screening and prioritisation tool rather than as a deterministic predictor of absolute inundation depth. This role is consistent with published surrogate applications developed to accelerate coastal and urban flood assessment. Rapid storm-surge mapping systems and machine-learning emulators have been used to reduce the computational burden of large scenario ensembles and to support faster-than-full-physics evaluation (Yang et al., 2020; Zahura et al., 2020; Ayyad et al., 2022). In a coastal urban setting, Zahura and Goodall (2022) showed that a Random-Forest surrogate could reproduce tidal and pluvial flood patterns within seconds and could be compared with independent crowdsourced flood reports. These studies support the use of surrogates as a computational triage layer that identifies potentially affected locations and prioritises the scenarios requiring detailed hydrodynamic simulation.
This type of framework is relevant for Copernicus downstream coastal services because it links in situ observations, gridded marine and atmospheric products, physical-response modelling, and uncertainty diagnostics within a computationally efficient workflow. Nevertheless, the comparison with previous applications also clarifies the present boundary of use. Operational or near-real-time surrogate systems generally depend on a locally calibrated high-fidelity reference model, representative training events, and independent impact or inundation observations (Yang et al., 2020; Zahura et al., 2020, 2022). In the present study, the XBeach-informed layer is simplified and the storm-sequence diagnostic is an internal consistency check rather than independent validation. The model can therefore provide rapid guidance for scenario exploration, flood-extent screening, and selection of cases for further simulation, but operational quantitative depth forecasting would require locally calibrated XBeach runs, richer morphology, and independent flood, overtopping, or impact observations.
5.3 Why the extreme tail remains difficult
The reduced skill for upper-tail inundation depths is likely caused by a combination of sampling, physical-representation, and spatial-description limitations, and is consistent with a recognised difficulty in statistical and machine-learning representations of extremes. Harter et al. (2024) showed that storm-surge reconstructions commonly underestimate the largest events because extremes are rare, relevant drivers may be omitted, and atmospheric forcing can itself be biased. Ayyad et al. (2022) similarly noted that limited training samples constrain prediction of low-probability, high-consequence surges, while Ahmadi Gharehtoragh and Johnson (2024) found larger surrogate errors for lower annual exceedance probabilities and the most extreme landscape configuration. In the present dataset, severe inundation cases form only a small fraction of the training space, and the concentration of target values near the imposed upper bound of approximately 4 m creates an additional plateau that limits discrimination among the largest responses.
Second, the XBeach-informed physical baseline is intentionally simplified and does not resolve all local processes that may control extreme inundation, including wave-overtopping pathways, backshore storage, small topographic barriers, harbour geometry, lagoon connectivity, and spatial variations in roughness. The resulting pattern—high skill for the inundation footprint but underestimation of large-event depths—has also been reported in a different coastal-inundation context by Ragu Ramalingam et al. (2025), whose machine-learning surrogate reproduced inundation maps more reliably than maximum depths for the largest test events. This comparison does not imply equivalence between storm-surge and tsunami processes, but it illustrates a general surrogate-modelling limitation when the training database or reduced physical representation does not fully span the most severe onshore response.
Third, the spatial grid and morphology descriptors remain idealised. They provide a consistent framework for rapid screening, but they do not fully distinguish between dune-backed beaches, engineered harbour areas, low-lying lagoonal sectors, cliff-backed coasts, and urbanised backshore environments. Studies that explicitly include changing landscape and boundary-condition descriptors show that morphology information can materially improve transfer across scenarios, while the most extreme or least represented landscape states remain more difficult to predict (Ahmadi Gharehtoragh and Johnson, 2024). The simplified geometry used here therefore helps explain why the surrogate captures flood occurrence more reliably than the exact magnitude of upper-tail inundation depths.
Future development should therefore prioritise event-based XBeach calibration at representative coastal sites, targeted sampling of severe compound wave–surge conditions, removal or reformulation of the imposed upper-depth plateau, richer morphology descriptors, and independent validation against observed flood, overtopping, or impact data. Tail-aware training strategies, quantile-oriented objectives, and conformal calibration stratified by depth range or morphology class should also be evaluated, but only after the physical baseline and extreme-event training sample have been strengthened. These steps are required before the framework can support quantitative operational prediction of maximum inundation depth.
These limitations do not reduce the value of the present surrogate as a rapid screening tool; rather, they define the boundary between exploratory flood-extent assessment and operational quantitative inundation-depth forecasting. Within that boundary, the framework provides a transparent means of combining observed and gridded forcing, a physics-informed baseline, machine-learning correction, and explicit uncertainty diagnostics.”
Citation: https://doi.org/10.5194/egusphere-2026-3628-AC1
-
AC1: 'Reply on RC1', Maria Mihailov, 23 Jul 2026
reply
-
RC2: 'Comment on egusphere-2026-3628', Anonymous Referee #2, 06 Aug 2026
reply
In this paper, entitled “An XBeach-Informed Machine-Learning Surrogate for Rapid Coastal Flood Prediction”, the authors develop a machine-learning surrogate model for coastal flood forecasting. Although the topic is interesting, the reviewer believes that the manuscript contains several issues that preclude its publication in its current form. My comments and suggestions are reported below.
My major concern relates to the fact that the methodology for predicting coastal flooding has been developed and validated using tide-gauge observations of sea level rather than direct inundation data. Tide gauges measure sea level near the coast and, because they are typically installed within harbors or near coastal structures, they cannot adequately represent coastal flooding processes. In particular, they do not account for wave-induced effects such as wave setup and runup, and their records may be influenced by harbor resonance. This appears to be the case in the present study, where the tide gauges at Constanța, Mangalia, and Midia are located within well-protected harbors, while the coordinates reported for Gura Portiței indicate an offshore location. Indeed, Figures 5 and 7 suggest that waves have a negligible influence on the estimated inundation depth. Moreover, the selection of the flooding threshold should be related to the specific characteristics of each coastal segment. In the present work, however, a single threshold value has been adopted for all investigated coastal sectors. This also raises the question of how sensitive the flood predictions are to the choice of this threshold. For these reasons, although the proposed modelling framework may be conceptually sound, it cannot be properly applied to and validated for the western Black Sea coast using the available data. I therefore strongly encourage the authors to seek and incorporate direct inundation observations. In my opinion, without such data, the methodology cannot be adequately tested and the manuscript is not suitable for publication.
Secondly, the manuscript lacks a clear and comprehensive description of the methodology. This aspect should be substantially improved to enable readers to fully understand the proposed approach. In particular: i) the Methods section should be reorganized to follow the workflow illustrated in Fig. 3 (observations, XBeach model, Monte Carlo scenarios, surrogate model, etc.); ii) it is not clear whether the Monte Carlo scenarios were generated using all sea-level observations from the different tide gauges collectively, or whether separate scenarios were developed for each coastal sector; iii) the description of the XBeach model should also provide details of the simulation setup, including the model domain, boundary and forcing conditions, spatial resolution, time step, and the representation of coastal and harbor structures; iv) the characteristics of the 10 × 10 spatial target grid used by the surrogate model should be reported, including its extent, resolution, and the bathymetric and topographic datasets employed; v) a clear description on how the dataset have been used for training and testing the surrogate models is missing; vi) it is unclear how the 0D tide-gauge observations, the 1D XBeach profiles, and the 2D surrogate-model grid are integrated.
Thirdly, the authors should pay greater attention to the use of appropriate technical terminology. Below are some examples of terms that are either unclear or potentially misleading:
-
Coastal flood depth, inundation depth, inundation field, inundation response, spatial inundation. In my opinion, the term inundation is not appropriate in this context, since both the observations and the modelling framework appear to represent the (total) sea level rather than actual land flooding. A clear definition of inundation depth should be clearly stated at the beginning of the Methods section.
-
XBeach-informed physical baseline, physics baseline. These terms appear to refer simply to the outputs generated by the XBeach model. I therefore suggest replacing them with a more straightforward expression such as XBeach outputs or XBeach simulations. The terms informed, physical, and baseline add unnecessary complexity and may confuse the reader.
-
Storm surge, storm-surge setup, sea-level anomaly. Tide gauges record the total sea level and additional processing is required to separate and quantify the different components of sea-level variability (e.g., storm surge, sea-level anomaly, astronomical tide). The manuscript should clearly explain how these quantities were derived and used.
-
Storm sequence: It is unclear what is meant by this term. The authors may be referring to storm duration, a storm event, or a storm case. Please either use a more precise term or clearly define its meaning in the manuscript.
Lastly, the results indicate that the different surrogate models exhibit very similar performance. I therefore suggest presenting only the ensemble mean results in the main manuscript and moving the complete set of performance statistics to the Supplementary Material. This would improve the readability of the paper without affecting the interpretation of the results. Furthermore, the authors should consider that these findings may be strongly influenced by the methodological limitations highlighted in the above concerns. Therefore, the reported model performance should be interpreted with caution until the proposed framework is validated against appropriate inundation observations.
Additional comments:
-
As it stands, the abstract reads more like that of a paper intended for a national or regional journal. Moreover, it contains unnecessary details on the datasets and statistical results, including unexplained acronyms. The abstract should be streamlined to emphasize the study's objectives, methodology, key findings, and significance.
-
Lines 113-126: details about the datasets, model performance metrics and predictors should not be reported here.
-
Lines 128-129: the flood extent, expressed as grid nodes exceeding an operational inundation, does not implies continuity of the flooding area. Please discuss the limitation of this approach.
-
exceeding an operational inundation threshold
-
Lines 141-144: is this a result of the present study? If so, it should be moved to the Results section. Otherwise, an appropriate reference should be provided.
-
Line 154: main Romanian Western coastal sectors.
-
Table 1: include the temporal extension and the number of valid data for each station.
-
Figure 2: This figure suggests that the tide gauges use different vertical reference datums, while Table 1 reports very similar mean SLEV values. Please clarify whether the data were adjusted to a common datum before the analysis.
-
Line 175: what do you mean with related hydrodynamic variables. You should only mention the variables that have been used in the present study.
-
Line 182: provide the reference for the topo-bathymetric dataset.
-
Table 2: many variables (SLEV, ATMS, Hs, Tp, sea-surface state, SST, MSLP) and roles (storm surge, erosion-related forcing, plume related forcing, …) need to better defined.
-
Line 203: define “target inundation response”.
-
Lines 214-216: the variables storm duration, inverse-barometer contribution, wave-power-related terms, and interaction terms between sea-level anomaly and wave forcing must be clearly defined.
-
Table 3: The mean values obtained from the Monte Carlo simulations (MCS) are significantly higher than the observed means. This discrepancy is particularly evident for sea level, for which the simulated mean exceeds the observed mean by more than 40%. The authors should provide an explanation for this difference and discuss its potential implications for the modelling results.
-
Lines 229-230: define “design-of-experiments forcing combinations”.
-
Line 230: why sea-level anomaly?
-
Line 231: specify the topo-bathymetric dataset used for implementing the XBeach cross-sections.
-
Lines 234-236: I do not agree with this approach. Each component of the modelling system must be calibrated and verified to obtain a reliable prediction system.
-
Lines 236-238: this part should be moved to section 3.4.
-
Figure 5: The left panel clearly shows that the XBeach profiles provide an unrealistic representation of the offshore/submerged part of the coastal profile, including a vertical wall exceeding 5 m in height. The authors should explain the origin of this feature and discuss its potential effects on the model results.
-
Line 291: from where do you get observations of the inundation depth at grid-node level?
-
Lines 296, tables 5 and 6, and somewhere else: always specify “mean” when referring to average value of the surrogate models.
-
Table 5: why the physics baseline has negative R2?
-
Figure 6: What’s the Observed Coastal Flood Depth? Tide gauge measure sea level not flood depth.
-
Captions of figures 6, 7, 8 and 9: do not report comments in the captions.
-
Section4.5: explain how the permutation-importance diagnostics have been computed.
-
Line 353: clarify how the interaction between sea-level anomaly and significant wave height have been defined.
-
Line 356: no compound sea-level–wave interaction effects can be reliably assessed in this study, as the observations were collected from tide gauges located within harbors, and the XBeach simulations include an unrealistic 5 m vertical wall structure.
-
It is not necessary to include both Figure 7 and Table 8 as they report the same information.
-
Section 4.7: from where these other stations come from?
-
Figure 9b: do not mix observation with XBeach model output.
Citation: https://doi.org/10.5194/egusphere-2026-3628-RC2 -
AC2: 'Reply on RC2', Maria Mihailov, 16 Aug 2026
reply
Dear Reviewer,
Thank you very much for your careful and constructive assessment of our manuscript.
We have substantially revised the manuscript in response to your comments. In particular, we have clarified the role and limitations of the tide-gauge observations, reorganized and expanded the methodology, incorporated real-event XBeach simulations for the November 2023 and December 2024 storms, revised the coastal profiles and inundation diagnostics, added threshold-sensitivity and independent-event evaluations, and substantially revised the figures and terminology throughout the manuscript.
We also appreciate your observation regarding the regional framing of the study. We agree that the previous version placed excessive emphasis on the regional datasets and did not sufficiently communicate the broader methodological contribution. In the revised manuscript, we therefore distinguish more clearly between the regional application domain and the scientific scope of the methodology. The Western Black Sea is used as a physically distinctive testbed for a more general problem: the efficient emulation of process-based coastal-flood simulations while retaining event dependence, spatial variability, threshold sensitivity, and an explicit applicability domain. The revised framework, including time-resolved XBeach simulations, independent-event testing, threshold-sensitivity analysis, and event-conditioned Monte Carlo screening, is not intrinsically specific to the Romanian coast and can be transferred to other coastal settings following appropriate local model configuration and forcing.
A detailed point-by-point response to all your comments is provided in the uploaded PDF, with the corresponding changes identified in the revised manuscript.
We sincerely appreciate your comments, which helped us to improve our work.
-
AC3: 'Reply on AC2', Maria Mihailov, 19 Aug 2026
reply
Response to Reviewer 2
We would like to inform the Reviewer that, following a careful re-check of our analysis notebooks, we identified an error in the quantitative results reported in our previous response. Some of the performance statistics cited there — in particular those for the maximum land water depth diagnostic — did not correspond to the actual model outputs. We sincerely apologise for this oversight. This document therefore succeeds our earlier reply: the results below have been recomputed directly from the model code and now reflect the correct, verified values. All substantive changes to the manuscript described previously remain in place; only the reported figures have been corrected.
The paper has been revised in the sections directly affected by the review while retaining its overall scientific structure and regional bordering. The central change is that the paper no longer treats tide-gauge sea level as direct evidence of land inundation. The revised work is outlined explicitly as a machine-learning emulator of fixed-bed XBeach simulations forced by two documented storms. The former synthetic 10 × 10 target, residual physics baseline, and pre-XBeach Monte Carlo construction have been removed. The main quantitative validation is now an independent temporal test: models are trained on November 2023 XBeach states and tested on the complete December 2024 event.
Major comment 1
My major concern relates to the fact that the methodology for predicting coastal flooding has been developed and validated using tide-gauge observations of sea level rather than direct inundation data. Tide gauges measure sea level near the coast and, because they are typically installed within harbors or near coastal structures, they cannot adequately represent coastal flooding processes. In particular, they do not account for wave-induced effects such as wave setup and runup, and their records may be influenced by harbor resonance. This appears to be the case in the present study, where the tide gauges at Constanța, Mangalia, and Midia are located within well-protected harbors, while the coordinates reported for Gura Portiței indicate an offshore location. Indeed, Figures 5 and 7 suggest that waves have a negligible influence on the estimated inundation depth. Moreover, the selection of the flooding threshold should be related to the specific characteristics of each coastal segment. In the present work, however, a single threshold value has been adopted for all investigated coastal sectors. This also raises the question of how sensitive the flood predictions are to the choice of this threshold. For these reasons, although the proposed modelling framework may be conceptually sound, it cannot be properly applied to and validated for the western Black Sea coast using the available data. I therefore strongly encourage the authors to seek and incorporate direct inundation observations. In my opinion, without such data, the methodology cannot be adequately tested and the paper is not suitable for publication.
Response: We agree with the important criticism. The previous formulation overstated what tide-gauge data could validate. We have removed the tide gauges from the land-inundation validation chain and removed all labels such as “observed coastal flood depth”. The revised target is actual land wetting calculated from XBeach zs and bed elevation along the cross-shore profiles. We have not obtained spatially georeferenced field maps of inundation for the two storms, so we do not claim end-to-end observational validation. Instead, the paper is explicitly scoped as an XBeach emulator. To strengthen the physical test, we use 14 documented-event XBeach runs, an entirely independent December 2024 temporal test, event-peak diagnostics, source-aware terrain QA, and wet/dry-threshold sensitivity. The relevant parts of the Abstract, Methods, Results, Discussion, and Conclusions have been revised accordingly.
Revised paper text: “The revised study is explicitly framed as an emulator of fixed-bed XBeach simulations rather than as an observationally validated inundation-forecast system. Tide-gauge records are therefore not used as a proxy for land flooding. Instead, the surrogate targets are extracted directly from shoreline-connected wetting simulated by XBeach for two documented storm events. Because spatially georeferenced field observations of inundation depth or flood extent are not available for these events, the XBeach simulations themselves are not claimed to be validated against direct inundation observations; this limitation defines the present scope of the study.”
Response: The threshold issue is also addressed directly. The former universal 0.30 m classification threshold has been removed from the XBeach land-wetting diagnostic. A 0.05 m numerical wet/dry criterion is used for the surrogate target, and the resulting connected flood extent is recomputed at 0.01, 0.05, 0.10, and 0.20 m for both storms. Figure 6 now shows the sector-specific sensitivity.
Revised paper text: “Sensitivity to the wet/dry criterion was assessed by recomputing shoreline-connected inundation extent at thresholds of 0.01, 0.05, 0.10, and 0.20 m. Averaged across the seven modelling sectors, increasing the threshold from 0.01 to 0.20 m reduces mean event-peak extent by approximately 11.7% for November 2023 and 14.1% for December 2024, with strong site-to-site variation. The 0.05 m value is retained as the numerical wet/dry diagnostic used for surrogate training, while the sensitivity curves are reported to make the threshold dependence explicit.”
Major comment 2
Secondly, the paper lacks a clear and comprehensive description of the methodology. This aspect should be substantially improved to enable readers to fully understand the proposed approach. In particular: i) the Methods section should be reorganized to follow the workflow illustrated in Fig. 3 (observations, XBeach model, Monte Carlo scenarios, surrogate model, etc.); ii) it is not clear whether the Monte Carlo scenarios were generated using all sea-level observations from the different tide gauges collectively, or whether separate scenarios were developed for each coastal sector; iii) the description of the XBeach model should also provide details of the simulation setup, including the model domain, boundary and forcing conditions, spatial resolution, time step, and the representation of coastal and harbor structures; iv) the characteristics of the 10 × 10 spatial target grid used by the surrogate model should be reported, including its extent, resolution, and the bathymetric and topographic datasets employed; v) a clear description on how the dataset have been used for training and testing the surrogate models is missing; vi) it is unclear how the 0D tide-gauge observations, the 1D XBeach profiles, and the 2D surrogate-model grid are integrated.
Response: The Methods section has been clarified and reorganised around the revised workflow (new Fig. 3): documented forcing → final terrain profiles → 14 fixed-bed XBeach simulations → time-resolved ML ensemble → event-conditioned Monte Carlo. The previous 10 × 10 grid has been removed, so there is no longer an unexplained 0D–1D–2D mapping. The surrogate predicts three diagnostics extracted directly from each 1-D XBeach state: maximum land water depth, shoreline-connected flood extent, and number of wet land nodes. We now report the XBeach boundary forcing, grid design, 73 h duration, 600 s output interval, fixed-bed configuration, lack of explicit harbour structures, 0.05 m wet/dry criterion, one-hour spin-up exclusion, predictor definitions, leave-one-sector-out model-development procedure, and fully independent December 2024 test.
Revised paper text: “In this study, inundation depth is defined as the positive XBeach-modelled water-column thickness above the local bed elevation at landward grid nodes, h = max(zs - zb, 0). A land node is considered wet when h is at least 0.05 m. The contiguous flood extent is the landward distance from the shoreline to the last consecutively wet node before the first dry interruption. These quantities are XBeach model diagnostics; they are not tide-gauge sea level and they are not direct field observations of inundation.”
Major comment 3
Thirdly, the authors should pay greater attention to the use of appropriate technical terminology. Below are some examples of terms that are either unclear or potentially misleading:
3.1 Coastal flood depth, inundation depth, inundation field, inundation response, spatial inundation. In my opinion, the term inundation is not appropriate in this context, since both the observations and the modelling framework appear to represent the (total) sea level rather than actual land flooding. A clear definition of inundation depth should be clearly stated at the beginning of the Methods section.
3.2 XBeach-informed physical baseline, physics baseline. These terms appear to refer simply to the outputs generated by the XBeach model. I therefore suggest replacing them with a more straightforward expression such as XBeach outputs or XBeach simulations. The terms informed, physical, and baseline add unnecessary complexity and may confuse the reader.
3.3 Storm surge, storm-surge setup, sea-level anomaly. Tide gauges record the total sea level and additional processing is required to separate and quantify the different components of sea-level variability (e.g., storm surge, sea-level anomaly, astronomical tide). The paper should clearly explain how these quantities were derived and used.
3.4 Storm sequence: It is unclear what is meant by this term. The authors may be referring to storm duration, a storm event, or a storm case. Please either use a more precise term or clearly define its meaning in the paper.
Response: Agreed. Inundation is now defined only from positive XBeach water depth on land, not from tide-gauge sea level. The definition appears at the beginning of Methods.
Revised paper text: “In this study, inundation depth is defined as the positive XBeach-modelled water-column thickness above the local bed elevation at landward grid nodes, h = max(zs - zb, 0). A land node is considered wet when h is at least 0.05 m. The contiguous flood extent is the landward distance from the shoreline to the last consecutively wet node before the first dry interruption. These quantities are XBeach model diagnostics; they are not tide-gauge sea level and they are not direct field observations of inundation.”
Response: Agreed. The terms “XBeach-informed physical baseline” and “physics baseline” have been removed. The final surrogate is trained directly on XBeach outputs; the former residual-baseline formulation and its negative R² are no longer part of the paper.
Response: Agreed. We no longer call the CMEMS boundary series a tide-gauge-derived storm surge. The paper consistently uses “CMEMS sea-level forcing” or “CMEMS zos”. No tide/surge/harbour-resonance decomposition is performed or claimed.
Revised paper text: “The revised paper uses the terms CMEMS sea-level forcing, XBeach simulation/output, documented storm event, maximum land water depth, contiguous flood extent, and wet land nodes. No decomposition of tide-gauge total water level into astronomical tide, storm surge, or harbour-resonance components is used in the surrogate workflow.”
Response: Agreed. “Storm sequence” has been replaced by “documented storm event”, “event time”, or “time-resolved event response”.
Major comment 4
Lastly, the results indicate that the different surrogate models exhibit very similar performance. I therefore suggest presenting only the ensemble mean results in the main paper and moving the complete set of performance statistics to the Supplementary Material. This would improve the readability of the paper without affecting the interpretation of the results. Furthermore, the authors should consider that these findings may be strongly influenced by the methodological limitations highlighted in the above concerns. Therefore, the reported model performance should be interpreted with caution until the proposed framework is validated against appropriate inundation observations.
Response: Agreed. Only the validation-weighted ensemble is discussed quantitatively in the main Results, while the complete RF, ET, GB, and HGB statistics are retained in the Supplementary Material. The updated independent-event evaluation shows a target-dependent skill pattern that is important to report transparently. Maximum land water depth has RMSE = 0.116 m but R² = 0.015, indicating limited reproduction of its time-step-to-time-step variance on the unseen December 2024 event. In contrast, shoreline-connected inundation extent and wet-land-node count remain well reproduced (R² = 0.958 and 0.971; RMSE = 14.21 m and 1.87 nodes, respectively). The sector-wise event-peak diagnostic gives the same qualitative pattern, with R² = 0.279 for maximum land water depth, 0.971 for shoreline-connected extent, and 0.966 for wet-land-node count. We therefore interpret the surrogate primarily as a rapid emulator for shoreline-connected wetting and extent screening, and we do not claim equally strong independent-event skill for exact time-resolved maximum depth. All reported skill statistics quantify fidelity to XBeach outputs rather than validation against direct field observations of inundation.
Revised paper text: “Surrogate skill is evaluated against XBeach outputs, not against tide-gauge observations. The November 2023 event provides 3,031 post-spin-up states for model development, while the December 2024 event is held out in its entirety as an independent temporal test with another 3,031 states. On that independent event, the validation-weighted ensemble gives MAE = 0.102 m, RMSE = 0.116 m, bias = -0.102 m, and R² = 0.015 for maximum land water depth; MAE = 9.25 m, RMSE = 14.21 m, bias = -9.18 m, and R² = 0.958 for shoreline-connected inundation extent; and MAE = 1.43 nodes, RMSE = 1.87 nodes, bias = -1.40 nodes, and R² = 0.971 for the number of wet land nodes. The near-zero R² for maximum land water depth indicates limited skill in reproducing its within-event temporal variance on the unseen storm, despite the comparatively modest absolute RMSE. The extent and wet-node diagnostics generalise substantially better. These statistics quantify emulation fidelity to XBeach and must not be interpreted as observational flood-validation skill.”
Additional comments
- As it stands, the abstract reads more like that of a paper intended for a national or regional journal. Moreover, it contains unnecessary details on the datasets and statistical results, including unexplained acronyms. The abstract should be streamlined to emphasize the study's objectives, methodology, key findings, and significance.
Response: The Abstract has been rewritten and shortened. It now states the objective, revised XBeach-to-ML workflow, independent temporal test, principal ensemble results, and the explicit limitation that direct inundation observations are unavailable. Unnecessary dataset acronyms and the former long list of metrics were removed.
Revised paper text: “A wet/dry-threshold sensitivity analysis and an event-conditioned Monte Carlo experiment are used to quantify diagnostic robustness and the surrogate applicability domain. The revised framework does not use tide-gauge sea level as a proxy for land inundation and does not claim validation against direct flood observations, which are unavailable for the analysed events. It is therefore intended as a rapid emulator and scenario-screening tool for XBeach coastal-flood diagnostics rather than as a stand-alone operational predictor of absolute inundation depth.”
- Lines 113-126: details about the datasets, model performance metrics and predictors should not be reported here.
Response: The former introduction-level details on predictors and performance metrics have been removed. The retained Introduction now provides only the scientific motivation, revised scope, and objectives; implementation details are in Methods and quantitative results are in Sect. 4.
Revised paper text: “The objective is to test whether a machine-learning ensemble can reproduce time-resolved XBeach coastal-flood diagnostics at seven Western Black Sea sectors and then use the trained emulator for rapid, event-conditioned scenario screening. The revised workflow therefore separates (i) documented marine and atmospheric forcing, (ii) sector-specific cross-shore terrain and fixed-bed XBeach simulations, (iii) surrogate training and independent temporal testing, and (iv) event-conditioned Monte Carlo application.”
- Lines 128-129: the flood extent, expressed as grid nodes exceeding an operational inundation, does not implies continuity of the flooding area. Please discuss the limitation of this approach.
Response: Agreed. The former count of threshold-exceeding 2-D grid nodes has been removed. Flood extent is now defined as a shoreline-connected 1-D cross-shore extent, which explicitly requires consecutive wet nodes from the shoreline. Its limitation as a transect metric rather than a 2-D polygon is stated.
Revised paper text: “The revised flood-extent metric explicitly enforces one-dimensional shoreline connectivity along each cross-shore profile: only consecutive wet land nodes beginning at the shoreline are counted. This removes the former ambiguity associated with counting isolated threshold exceedances. It should nevertheless be interpreted as a 1-D transect extent, not as a continuous two-dimensional inundation polygon or an alongshore-connected flooded area.”
- exceeding an operational inundation threshold
Response: The wording has been corrected throughout. The revised paper uses “wet/dry criterion” and “shoreline-connected flood extent” rather than the grammatically incomplete phrase in the previous version.
- Lines 141-144: is this a result of the present study? If so, it should be moved to the Results section. Otherwise, an appropriate reference should be provided.
Response: The sentence in question has been removed from the Introduction. Event-specific findings are now confined to the Results section.
- Line 154: main Romanian Western coastal sectors.
Response: The geographical wording has been standardised to “Romanian Western Black Sea coast” or “Western Black Sea sectors”, depending on context.
- Table 1: include the temporal extension and the number of valid data for each station.
Response: The former tide-gauge statistics table has been removed because tide gauges are no longer used as the inundation-validation target. Revised Table 1 instead documents all seven XBeach sectors, and Methods reports the two event windows and the 433 post-spin-up states per sector per event (3,031 states per event).
- Figure 2: This figure suggests that the tide gauges use different vertical reference datums, while Table 1 reports very similar mean SLEV values. Please clarify whether the data were adjusted to a common datum before the analysis.
Response: The former tide-gauge time-series Fig. 2 has been removed from the validation chain. The revised Fig. 2 shows CMEMS–ERA5 forcing for the two documented storms. We no longer compare local tide-gauge datums as if they were directly interchangeable. The remaining vertical-reference limitation is explicitly stated.
Revised paper text: “The former cross-station comparison of tide-gauge elevations has been removed from the validation chain. Terrain is expressed in a local shoreline-relative model datum with z = 0 at the shoreline, while CMEMS zos provides the time-varying offshore water-level boundary. A geodetic transformation that would place all land-survey elevations and CMEMS sea level in a single independently verified vertical reference is not available for every sector; this is therefore treated as a limitation of absolute depth interpretation rather than concealed through datum pooling.”
- Line 175: what do you mean with related hydrodynamic variables. You should only mention the variables that have been used in the present study.
Response: The vague phrase “related hydrodynamic variables” has been removed. Revised Table 2 lists only variables actually used: Hs, Tp, wave direction, CMEMS zos, ERA5 wind speed/direction, and static terrain sources.
- Line 182: provide the reference for the topo-bathymetric dataset.
Response: Topo-bathymetric sources are now specified explicitly: EMODnet DTM for marine bathymetry, Copernicus DEM GLO-30 for land elevation, shoreline/GNSS anchors for coastline registration, and the Zenodo RTK-GNSS/multibeam dataset for Vama Veche.
Revised paper text: “The profile-generation workflow was revised to remove the artificial offshore vertical-wall artefact identified by the reviewer. The shoreline is imposed at z = 0 using the available GNSS/shoreline anchor, bathymetry is taken from EMODnet (with the Vama Veche nearshore profile supplemented by the RTK-GNSS/multibeam Zenodo dataset), and land elevation is taken from Copernicus DEM GLO-30. Source-aware interpolation is used only across the unresolved nearshore transition. Quality control confirms no offshore re-emergent land nodes in the final source-aware profiles; steep nearshore sectors are retained with explicit visual-QA flags rather than being silently smoothed.”
- Table 2: many variables (SLEV, ATMS, Hs, Tp, sea-surface state, SST, MSLP) and roles (storm surge, erosion-related forcing, plume related forcing, …) need to better defined.
Response: Revised Table 2 defines every variable and its modelling role. Variables from the old framework that are not used by the final surrogate (e.g. SST, Danube discharge, MSLP as a direct surrogate predictor) have been removed from the core Methods.
- Line 203: define “target inundation response”.
Response: The phrase “target inundation response” has been replaced by explicit XBeach-derived targets: maximum land water depth, contiguous shoreline-connected flood extent, and wet-land-node count.
Revised paper text: “In this study, inundation depth is defined as the positive XBeach-modelled water-column thickness above the local bed elevation at landward grid nodes, h = max(zs - zb, 0). A land node is considered wet when h is at least 0.05 m. The contiguous flood extent is the landward distance from the shoreline to the last consecutively wet node before the first dry interruption. These quantities are XBeach model diagnostics; they are not tide-gauge sea level and they are not direct field observations of inundation.”
- Lines 214-216: the variables storm duration, inverse-barometer contribution, wave-power-related terms, and interaction terms between sea-level anomaly and wave forcing must be clearly defined.
Response: The final predictor definitions are now explicit. The previous vague terms were replaced by elapsed event time, Hs²Tp as a wave-power proxy, zos×Hs as the sea-level–wave product, cos(θwind−θwave) as wind– wave alignment, and max[cos(θwave),0] as the onshore-wave factor. The old generic “storm duration” and inverse-barometer terms are not part of the final surrogate feature set.
- Table 3: The mean values obtained from the Monte Carlo simulations (MCS) are significantly higher than the observed means. This discrepancy is particularly evident for sea level, for which the simulated mean exceeds the observed mean by more than 40%. The authors should provide an explanation for this difference and discuss its potential implications for the modelling results.
Response: The former 5,000-scenario pre-XBeach Monte Carlo table and the 40% mean sea-level mismatch have been removed. Monte Carlo is now an event-conditioned post-surrogate application based on the documented forcing states with small perturbations and an explicit applicability-domain filter.
Revised paper text: “The Monte Carlo stage was also reformulated. It no longer generates a synthetic forcing climatology before XBeach. Instead, 10,000 states per sector are sampled conditionally from the two documented-event forcing states with small perturbations, and only samples that remain inside the empirical surrogate-training envelope are used in the principal summaries. Depending on sector, 91.92–97.33% of generated states satisfy this applicability criterion (66,233 of 70,000 states overall). These results describe event-conditioned surrogate uncertainty and response ranges; they are neither a flood climatology nor an observational validation.”
- Lines 229-230: define “design-of-experiments forcing combinations”.
Response: The former “design-of-experiments forcing combinations” concept has been removed from the final workflow.
- Line 230: why sea-level anomaly?
Response: The final workflow does not derive a tide-gauge sea-level anomaly. It uses CMEMS zos directly as the time-varying offshore sea-level boundary and labels it as sea-level forcing.
Revised paper text: “The revised paper uses the terms CMEMS sea-level forcing, XBeach simulation/output, documented storm event, maximum land water depth, contiguous flood extent, and wet land nodes. No decomposition of tide-gauge total water level into astronomical tide, storm surge, or harbour-resonance components is used in the surrogate workflow.”
- Line 231: specify the topo-bathymetric dataset used for implementing the XBeach cross-sections.
Response: The terrain sources used for the XBeach cross-sections are now specified in Sect. 2.2/3.2, Table 2, and the revised profile figure.
Revised paper text: “The profile-generation workflow was revised to remove the artificial offshore vertical-wall artefact identified by the reviewer. The shoreline is imposed at z = 0 using the available GNSS/shoreline anchor, bathymetry is taken from EMODnet (with the Vama Veche nearshore profile supplemented by the RTK-GNSS/multibeam Zenodo dataset), and land elevation is taken from Copernicus DEM GLO-30. Source-aware interpolation is used only across the unresolved nearshore transition. Quality control confirms no offshore re-emergent land nodes in the final source-aware profiles; steep nearshore sectors are retained with explicit visual-QA flags rather than being silently smoothed.”
- Lines 234-236: I do not agree with this approach. Each component of the modelling system must be calibrated and verified to obtain a reliable prediction system.
Response: We agree that a predictive system requires calibration and verification of each component. Because direct spatial inundation observations are unavailable for these two events, we have not represented the fixed-bed XBeach simulations as observationally calibrated. Instead, the paper is explicitly limited to XBeach emulation and screening; direct field calibration is identified as a prerequisite for operational forecasting.
Revised paper text: “The present framework remains intentionally limited. It uses one-dimensional fixed-bed XBeach transects, does not resolve harbour geometry, alongshore exchange, morphodynamic feedbacks, or two-dimensional overtopping pathways, and has not been calibrated against direct field maps of inundation depth or extent. Consequently, the surrogate is appropriate for rapid emulation, scenario ranking, and screening of the XBeach response space, but not yet for stand-alone operational prediction of absolute coastal flood depth.”
- Lines 236-238: this part should be moved to section 3.4.
Response: The old text has been removed during the complete Methods restructuring; the final surrogate formulation is now contained in Sect. 3.4.
- Figure 5: The left panel clearly shows that the XBeach profiles provide an unrealistic representation of the offshore/submerged part of the coastal profile, including a vertical wall exceeding 5 m in height. The authors should explain the origin of this feature and discuss its potential effects on the model results.
Response: This issue prompted a full revision of the profile-construction workflow. The former figure containing the artificial offshore wall has been replaced by source-aware, shoreline-registered profiles. The new QA verifies the marine/land transition, removes offshore re-emergent nodes, and flags steep sectors for visual inspection.
Revised paper text: “The profile-generation workflow was revised to remove the artificial offshore vertical-wall artefact identified by the reviewer. The shoreline is imposed at z = 0 using the available GNSS/shoreline anchor, bathymetry is taken from EMODnet (with the Vama Veche nearshore profile supplemented by the RTK-GNSS/multibeam Zenodo dataset), and land elevation is taken from Copernicus DEM GLO-30. Source-aware interpolation is used only across the unresolved nearshore transition. Quality control confirms no offshore re-emergent land nodes in the final source-aware profiles; steep nearshore sectors are retained with explicit visual-QA flags rather than being silently smoothed.”
- Line 291: from where do you get observations of the inundation depth at grid-node level?
Response: There are no observations of inundation depth at grid-node level. The revised paper now states this explicitly. The former 2-D grid target is removed, and all ML targets are extracted directly from XBeach output.
Revised paper text: “In this study, inundation depth is defined as the positive XBeach-modelled water-column thickness above the local bed elevation at landward grid nodes, h = max(zs - zb, 0). A land node is considered wet when h is at least 0.05 m. The contiguous flood extent is the landward distance from the shoreline to the last consecutively wet node before the first dry interruption. These quantities are XBeach model diagnostics; they are not tide-gauge sea level and they are not direct field observations of inundation.”
- Lines 296, tables 5 and 6, and somewhere else: always specify “mean” when referring to average value of the surrogate models.
Response: Agreed. The main paper now refers to the “validation-weighted ensemble” rather than an ambiguous “average model”. Individual-model statistics are placed in the Supplement.
- Table 5: why the physics baseline has negative R2?
Response: The former physics baseline and its negative R² have been removed. The final model is a direct multi-output surrogate of XBeach simulations rather than a residual correction to a simplified baseline.
- Figure 6: What’s the Observed Coastal Flood Depth? Tide gauge measure sea level not flood depth.
Response: The label “Observed Coastal Flood Depth” has been removed. Revised Figs. 7 and 8 explicitly use “XBeach” on the reference axis/series and “ML ensemble” on the prediction axis/series.
- Captions of figures 6, 7, 8 and 9: do not report comments in the captions.
Response: Figure captions have been rewritten as descriptive captions only; interpretive statements have been moved to the Results text.
- Section4.5: explain how the permutation-importance diagnostics have been computed.
Response: The computation is now described explicitly: on the independent December 2024 event, each predefined predictor group is permuted while the trained ensemble is fixed; the operation is repeated five times; importance is the increase in RMSE relative to the unpermuted prediction.
Revised paper text: “Permutation importance is computed on the independent December 2024 event by randomly permuting each predefined predictor group while leaving the trained ensemble unchanged, repeating the perturbation five times, and reporting the increase in RMSE relative to the unpermuted prediction. The grouped analysis shows that profile geometry and site identity dominate all three outputs, with sea level providing the next consistent contribution. Wave forcing and the explicit sea-level–wave interaction have small additional permutation importance in this two-event experiment; no claim of a robust compound sea-level–wave interaction is therefore made.”
- Line 353: clarify how the interaction between sea-level anomaly and significant wave height have been defined.
Response: The interaction feature is now defined explicitly as the product zos×Hs. Importantly, the revised grouped permutation analysis shows negligible additional importance of this interaction in the current two-event experiment.
Revised paper text: “Permutation importance is computed on the independent December 2024 event by randomly permuting each predefined predictor group while leaving the trained ensemble unchanged, repeating the perturbation five times, and reporting the increase in RMSE relative to the unpermuted prediction. The grouped analysis shows that profile geometry and site identity dominate all three outputs, with sea level providing the next consistent contribution. Wave forcing and the explicit sea-level–wave interaction have small additional permutation importance in this two-event experiment; no claim of a robust compound sea-level–wave interaction is therefore made.”
- Line 356: no compound sea-level–wave interaction effects can be reliably assessed in this study, as the observations were collected from tide gauges located within harbors, and the XBeach simulations include an unrealistic 5 m vertical wall structure.
Response: Agreed. We no longer claim a reliably quantified compound sea-level–wave interaction. The revised profile artefact has been addressed, and the new importance analysis shows that the explicit interaction contributes little beyond geometry, site identity, and sea level. This is stated as a result rather than interpreted as a robust physical interaction.
Revised paper text: “Permutation importance is computed on the independent December 2024 event by randomly permuting each predefined predictor group while leaving the trained ensemble unchanged, repeating the perturbation five times, and reporting the increase in RMSE relative to the unpermuted prediction. The grouped analysis shows that profile geometry and site identity dominate all three outputs, with sea level providing the next consistent contribution. Wave forcing and the explicit sea-level–wave interaction have small additional permutation importance in this two-event experiment; no claim of a robust compound sea-level–wave interaction is therefore made.”
- It is not necessary to include both Figure 7 and Table 8 as they report the same information.
Response: The redundant figure/table pair has been removed. The main paper retains the grouped permutation-importance figure; detailed individual feature values remain supporting material rather than duplicating the figure in the main text.
- Section 4.7: from where these other stations come from?
Response: The seven XBeach modelling sectors are introduced explicitly at the beginning of Sect. 2 and documented in the modelling description/Table 1: Constanța, Mamaia, Midia, Eforie Nord, Mangalia, Vama Veche, and Gura Portiței. Figure 1 is retained as the MHD sea-level station map for geographic and observational context; the four tide-gauge locations shown there are clearly distinguished from the seven XBeach modelling sectors and are not treated as the inundation-validation network.
- Figure 9b: do not mix observation with XBeach model output.
Response: Agreed. The revised validation figures no longer mix “observation” with XBeach output. They compare XBeach reference outputs directly against the ML ensemble, and the text repeatedly states that this is simulator-emulation validation, not observational flood validation.
Revised paper text: “Surrogate skill is evaluated against XBeach outputs, not against tide-gauge observations. The November 2023 event provides 3,031 post-spin-up states for model development, while the December 2024 event is held out in its entirety as an independent temporal test with another 3,031 states. On that independent event, the validation-weighted ensemble gives MAE = 0.102 m, RMSE = 0.116 m, bias = -0.102 m, and R² = 0.015 for maximum land water depth; MAE = 9.25 m, RMSE = 14.21 m, bias = -9.18 m, and R² = 0.958 for shoreline-connected inundation extent; and MAE = 1.43 nodes, RMSE = 1.87 nodes, bias = -1.40 nodes, and R² = 0.971 for the number of wet land nodes. The near-zero R² for maximum land water depth indicates limited skill in reproducing its within-event temporal variance on the unseen storm, despite the comparatively modest absolute RMSE. The extent and wet-node diagnostics generalise substantially better. These statistics quantify emulation fidelity to XBeach and must not be interpreted as observational flood-validation skill.”
Citation: https://doi.org/10.5194/egusphere-2026-3628-AC3
-
AC3: 'Reply on AC2', Maria Mihailov, 19 Aug 2026
reply
-
-
RC3: 'Comment on egusphere-2026-3628', Anonymous Referee #3, 14 Aug 2026
reply
Recommendation: major revision
General comment:
The manuscript develops a tree-ensemble surrogate for rapid coastal-inundation screening along the Western (Romanian) Black Sea coast, trained on a Monte Carlo forcing ensemble and referenced to an intermediate “XBeach-informed” physical-response layer. The topic is timely and operationally relevant: reducing the computational cost of process-based inundation modelling while keeping a traceable link to the controlling physical drivers is a real need for coastal early warning and for Copernicus downstream services. The assembly of a coherent forcing chain for this specific coastal sector — high-frequency MHD tide-gauge records, Copernicus Marine Black Sea wave and physical products, Copernicus L4 SST, ERA5 atmospheric fields and GRDC Danube discharge — is a genuine contribution, and to my knowledge no comparable multi-driver surrogate has been published for the Western Black Sea.
I also want to acknowledge the diagnostic honesty of the manuscript. The authors do not hide the weak performance in the upper tail; they deliberately separate threshold-based flood-extent classification from continuous depth regression, they report conformal prediction intervals including their poor sharpness, and they provide residual, Q–Q, permutation-importance, site-specific and seasonal diagnostics. This level of transparency is above average for surrogate-modelling papers, and the resulting discussion (Sects. 5.1 and 5.3) is thoughtful.
That said, in its present form the manuscript is not fulfilling its central claims, for the following reasons:
- The manuscript refers to “the target inundation response” (e.g. L246, L285–L291) but nowhere states how the 5,000 × 100 target matrix was actually computed. This is the single most important omission, because every metric in the paper is computed against this quantity. Another misleading point is that the XBeach-informed response layer of Fig. 5c gives a maximum flood depth of 0.90 m across the entire design-of-experiments envelope, including a sea-level anomaly approaching 1 m combined with Hs above 5 m. Consistently with this, Fig. 5b gives a resulting depth of 0.1–0.8 m, Fig. 9b a peak inundation envelope below 0.9 m, and Fig. A2 a flood-depth proxy below 0.8 m with R2 % run-up below 1.6 m. Table 7, by contrast, reports that the 90th percentile of the target inundation depth is 3.966 m and that the distribution saturates at exactly 4.000 m. Please describe better or fix this.
- XBeach appears in the title and names the central methodological component, yet the manuscript gives no model version, no computational domain or grid resolution, no hydrodynamic or morphodynamic parameter settings, no boundary-condition specification, no number of simulations performed, and no calibration or validation of the XBeach results against observations. As it is, a reader cannot reproduce the physical baseline, and the “XBeach-informed” framing of the title is not yet demonstrated, thus please provide xbeach model details or a literature reference to support the data.
- Some reported diagnostics contradict each other: Table 5 gives the physics baseline a Pearson r of 0.026 and an NSE of −0.479, and Table 6 gives it MCC = −0.0028 — i.e. no skill whatsoever — yet Table 8 and Fig. 7 identify physics_baseline_m as by far the dominant predictor of the residual. A predictor that carries essentially no information about the target cannot simultaneously be the leading predictor. Further inconsistencies are listed in Sect. 2 below (Fig. 6a annotations vs. Table 5; Table C1 has several values repeated with no particular value for the manuscript; the 0.95 and 0.99 rows of Table 7).
- Beyond these three, two further points seem to me important enough to raise at this level. The surrogate takes sea level as a prescribed input rather than predicting it, so all reported skill is conditional on perfect knowledge of the forcing and is an optimistic bound on what the framework would achieve in forecast mode (M10). And the training loss function — the most likely proximate cause of the upper-tail bias that the manuscript identifies as its own main limitation — is never stated or discussed, nor is any tail-sensitive verification metric used for model selection (M11, S1).
In addition, an evaluation comparing against any observed flooding, overtopping or damage record would greatly improve the manuscript, so achieving the objective indicated in the abstract and title which promise “spatial flood prediction”. Although the 10 × 10 target grid is described as idealised and is never geo-referenced, given a resolution, or linked to the coastline mapped in Fig. 1. The claim of computational speed-up, which one of the motivation is never quantified: no timings are reported for either XBeach or the surrogate. In some cases, the manuscript is repetitive, with the “classification good / regression poor” message restated in the abstract, Sect. 1, Sects. 4.2 and 4.3, Sects. 5.1 and 5.2 and the conclusions, so it could be shortened appreciably without any loss of content.
I want to close this section by repeating that the underlying idea is interesting and that the regional dataset is worth publishing, but the manuscript needs a major revision.
Major points to address:
- Please state what the 5,000 × 100 target matrix contains and how it was generated. Specifically: is each target value the maximum inundation depth extracted from an actual XBeach 2D simulation for that forcing scenario? What it the location, the timing,the model configuration? What are the forcing used?
- Please check the range of the target with the range of the intermediate response layer. As set out in the general comments, Fig. 5b–c, Fig. 9b and Fig. A2 all bound flood depth below about 0.9 m over the full forcing envelope, while Table 7 is about 4 m.
- Table 7 has some mistakes. The test set contains 115,200 nodes — this is the sum of both Table 10 and Table 11, which is reassuring — so the 0.90, 0.95 and 0.99 quantiles should select 11,520, 5,760 and 1,152 nodes respectively. The 0.90 row is consistent (n = 11,520), but the 0.95 and 0.99 rows both report n = 6,034 with an identical threshold of 4.0 m. As stated by the authors, there is a clip of the data, thus the latter 2 percentiles have no particular relevance. However is not clear where 4m comes from and what is the variable interested.
- The consequence is that the headline “extreme-event underprediction of −2.59 m” has 2 points to focus on: this is referred to inundation depth, which should be better detailed in the manuscript. In this version is not exactly well defined and clear. However, it should be not a physical finding about compound storm response, but it is the arithmetic difference between two saturation levels, one imposed on the target and one learned by the model.
- Please check the standard deviation in Table 1 which seems are not aligned with Figure 2: the std is very similar between the 4 sites, but the timeseries, in particular Gura seems to be quite different
- Figure 2 shows that the four stations cover very different periods, with large gaps and different ranges. The Table C2 reports 35,064 valid records for every variable, which is exactly four years of hourly data. Please state the common reference period, how gaps were treated, how much inputs were used for the analysis.
- Table 10 and Sect. 4.7 report site-specific metrics for Constanța, Eforie Nord, Gura Portiței, Mamaia, Mangalia, Midia and Vama Veche, and the text singles out Mamaia and Midia. However, Figs. 1, 2 and 8, and Table 1 describe only four monitoring stations. Where do the profiles, forcing and targets for Eforie Nord, Mamaia and Vama Veche come from?
- The held-out evaluation throughout was never defined. Please state explicitly whether the split is by scenario (grouped) or by row, give the split fractions, and confirm that the reported test set of 115,200 nodes contains no scenario also seen in training.
- All four models are tree ensembles that, in their standard implementations, minimise a squared-error criterion, and their predictions are averages over the training samples falling in the terminal leaves. Such models cannot extrapolate beyond the training range and shrink towards the conditional mean precisely where the target density is lowest. Yet the manuscript nowhere states which loss or split criterion was used for each model, and never considers alternatives. Given that upper-tail underprediction is presented as the limitation of the study, this deserves a dedicated subsection rather than no mention at all.
Concretely, I suggest the authors (a) state the loss and split criterion for each of RF, GB, ET and HGB, together with all hyperparameters and how they were tuned (currently absent entirely); (b) test at least one tail-oriented alternative — quantile regression, an asymmetric or extreme-weighted loss, sample weighting proportional to target exceedance, or a two-stage design with a dedicated upper-tail regressor above the 90th percentile; and (c) add at least one verification metric that is not dominated by the bulk of the distribution.
There is a coherent and rapidly growing literature on exactly this problem in the surge and flood emulation context that the manuscript does not engage with, although it already cites two papers from the same strand (Longo et al., 2026; Wilkinson et al., 2026). Hermans et al. (2025) show that densifying the loss towards extremes substantially improves neural-network estimates of extreme surges in Europe. Campos-Caba et al. (2026) benchmark emulators ranging from multivariate linear regression to LSTM networks against a 50 m coastal-resolution hydrodynamic model in the northern Adriatic and find that the choice of loss function matters more than model complexity: a linear model trained with an extreme-oriented loss (the corrected mean absolute deviation squared, MADc²) matches an LSTM at orders of magnitude lower computational cost — a result with obvious implications for the four-model ensemble of Sect. 3.4. Campos-Caba et al. (2024) address the companion question of which error indicators are capable of measuring skill on extremes at all.
- For a single dataset evaluated against the mean of the reference values, the coefficient of determination in its NSE form and Nash–Sutcliffe efficiency are identical by definition — which is why the abstract reports R² = 0.409 and NSE = 0.409, and why the two columns of Table 5 are identical. Please drop one, or state clearly that R² is the squared Pearson correlation (in which case it should equal 0.914² = 0.836, not 0.409) and give the definition used. More broadly: every regression metric in Table 5 — R², RMSE, MAE, NSE, KGE, r, ρ — is dominated by the bulk of the distribution, and none is sensitive to the upper tail, which is where the model actually fails. The tail behaviour appears only in Table 7, and not used for any model decision. I would recommend adding at least one extreme-oriented indicator to Table 5, for example a peak-over-threshold weighted error, a conditional bias above a fixed quantile, or the corrected mean absolute deviation of Campos-Caba et al. (2024), and using it in model selection alongside RMSE.
- The four “validation-derived” weights are 0.2499, 0.2501, 0.2499 and 0.2501, i.e. indistinguishable from a plain arithmetic mean, and the ensemble does not outperform the best individual model (ensemble R² = 0.4087 vs. HGB R² = 0.4091; ensemble F1 = 0.9948 vs. HGB F1 = 0.9963). Please state how the weights were derived and over what quantity they were optimised, and either demonstrate a benefit or simplify to a single model or a plain average. As written, the ensemble is presented as a methodological feature (L250–L255) that the results do not support.
- Please add a table listing all 22 forcing and derived predictors, with units, source dataset and, for derived quantities, the exact formula. This is essential for reproducibility
- Speed is the entire motivation for the work (“rapid” appears in the title, abstract and conclusions) but no timings are given anywhere. Please report the wall-clock cost of one XBeach simulation on stated hardware, the total cost of building the training set, the surrogate training time, and the inference time per scenario, so that the claimed speed-up can be assessed quantitatively.
Figure 1: the map resolution is too low to be legible — station labels, coastline and scale bar are hard to read at print size. Please supply a higher-resolution version, and consider adding bathymetric contours and the seven evaluation sites of Table 10.
Every Copernicus Marine product has a persistent DOI or product identifier. Please add them for all products used, in Table 2 and in the reference list, together with product version and the temporal coverage actually downloaded.
L174: add the full reference and product identifier for the Copernicus Marine Black Sea Physics Analysis and Forecast product.
L176–L177: add the reference and product identifier for the Copernicus Marine Black Sea high-resolution Level-4 SST product. The two Buongiorno Nardelli papers currently cited at L106 describe the methodology rather than the operational product, and both are needed.
L182: the bathymetric and coastal-elevation data need a proper description and citation — source, horizontal resolution, vertical datum, year, and how EMODnet bathymetry was merged with the topographic data. The EMODnet entry in the reference list has no URL, DOI or access date.
Figure 2: the four panels use different y-axis ranges, which makes visual comparison misleading (Gura Portiței spans −0.2 to +0.1 m, the others −0.5 to +0.5 m or wider). Please use a common axis or explain why not. The image resolution should be increased. The green color in Gura Portiței is very difficult to read.
Figure 9a: the horizontal axis is labelled “Operational Radar OTT RLS Tide-Gauge Stations Network”, refers to an instrument type which is introduced here for the first time, it should be included in the text in the observation sdescription.
The threshold is described as “operational” (L269, L314) please add a reference or rationale.
Figure B2: please improve resolution
Units are mixed between text, tables and figures: some Tables use centimetres, Table 8 uses metres, and the text uses metres throughout. Please use metres consistently, and state the unit in every axis label and column header.
Table 6: MCC and Cohen’s κ agree to four decimals in every row, so reporting both adds no information. Two would be enough
Table 7: No needs of 4 decimals in the quantiles
Sections 4.2, 4.3, 5.1, 5.2 and 6 repeat the same interpretation of the classification/regression contrast. Consolidating it into Sect. 5.1 would shorten the manuscript appreciably and make the argument stronger.
Citation: https://doi.org/10.5194/egusphere-2026-3628-RC3 -
AC4: 'Reply on RC3', Maria Mihailov, 19 Aug 2026
reply
Dear Reviewer,
Thank you very much for your careful and constructive assessment of our paper. Several of the concerns raised in this report overlap closely with issues addressed during the same revision cycle, including comments raised by Reviewer 2. To maintain a single scientifically consistent paper, we have therefore not introduced a second, independent restructuring solely in response to the present report. Instead, the responses below refer to the same targeted scientific revisions already incorporated in the revised paper, with only local clarifications where needed.
The central clarification is that tide-gauge sea level is no longer treated as direct evidence or validation of land inundation. The surrogate is evaluated against time-resolved outputs from fixed-bed XBeach simulations for two documented storm events (18–19 November 2023 and 25–26 December 2024), while the absence of spatially georeferenced field observations of flood extent or inundation depth is stated explicitly as a limitation. The revised analysis also includes source-aware coastal-profile quality control, wet/dry-threshold sensitivity, an independent temporal test on the December 2024 event, and an event-conditioned Monte Carlo application. Model-skill statistics are therefore interpreted as emulation fidelity to XBeach, not as observational flood-validation skill.
Major comment 1
My major concern relates to the fact that the methodology for predicting coastal flooding has been developed and validated using tide-gauge observations of sea level rather than direct inundation data. Tide gauges measure sea level near the coast and, because they are typically installed within harbors or near coastal structures, they cannot adequately represent coastal flooding processes. In particular, they do not account for wave-induced effects such as wave setup and runup, and their records may be influenced by harbor resonance. This appears to be the case in the present study, where the tide gauges at Constanța, Mangalia, and Midia are located within well-protected harbors, while the coordinates reported for Gura Portiței indicate an offshore location. Indeed, Figures 5 and 7 suggest that waves have a negligible influence on the estimated inundation depth. Moreover, the selection of the flooding threshold should be related to the specific characteristics of each coastal segment. In the present work, however, a single threshold value has been adopted for all investigated coastal sectors. This also raises the question of how sensitive the flood predictions are to the choice of this threshold. For these reasons, although the proposed modelling framework may be conceptually sound, it cannot be properly applied to and validated for the western Black Sea coast using the available data. I therefore strongly encourage the authors to seek and incorporate direct inundation observations. In my opinion, without such data, the methodology cannot be adequately tested and the paper is not suitable for publication.
Response: We agree with the core scientific concern and have revised the interpretation accordingly. Tide-gauge observations are retained only for observational and geographic context and are not used as direct measurements of land inundation. The machine-learning targets are extracted from XBeach land-wetting diagnostics, calculated from modelled water-surface elevation and bed elevation along the cross-shore profiles. Because spatially georeferenced field maps of inundation depth or flood extent are not available for the two analysed storms, we do not claim end-to-end observational validation. The study is explicitly scoped as an emulator of fixed-bed XBeach simulations and as a rapid scenario-screening framework. This same clarification was already introduced in response to Reviewer 2, and no additional restructuring is required here.
Revised paper text: “The revised study is explicitly framed as an emulator of fixed-bed XBeach simulations rather than as an observationally validated inundation-forecast system. Tide-gauge records are therefore not used as a proxy for land flooding. Instead, the surrogate targets are extracted directly from shoreline-connected wetting simulated by XBeach for two documented storm events. Because spatially georeferenced field observations of inundation depth or flood extent are not available for these events, the XBeach simulations themselves are not claimed to be validated against direct inundation observations; this limitation defines the present scope of the study.”
Response: The threshold issue is also addressed explicitly. The former universal 0.30 m flood-classification threshold is not used as the XBeach wet/dry diagnostic. A 0.05 m numerical wet/dry criterion is used for the surrogate targets, and shoreline-connected flood extent is recomputed at 0.01, 0.05, 0.10, and 0.20 m for both documented events to quantify site-specific sensitivity. The 0.05 m value is therefore treated as a reproducible numerical diagnostic, not as a universal physical damage threshold.
Revised paper text: “Sensitivity to the wet/dry criterion was assessed by recomputing shoreline-connected inundation extent at thresholds of 0.01, 0.05, 0.10, and 0.20 m. Averaged across the seven modelling sectors, increasing the threshold from 0.01 to 0.20 m reduces mean event-peak extent by approximately 11.7% for November 2023 and 14.1% for December 2024, with strong site-to-site variation. The 0.05 m value is retained as the numerical wet/dry diagnostic used for surrogate training, while the sensitivity curves are reported to make the threshold dependence explicit.”
Major comment 2
Secondly, the paper lacks a clear and comprehensive description of the methodology. This aspect should be substantially improved to enable readers to fully understand the proposed approach. In particular: i) the Methods section should be reorganized to follow the workflow illustrated in Fig. 3 (observations, XBeach model, Monte Carlo scenarios, surrogate model, etc.); ii) it is not clear whether the Monte Carlo scenarios were generated using all sea-level observations from the different tide gauges collectively, or whether separate scenarios were developed for each coastal sector; iii) the description of the XBeach model should also provide details of the simulation setup, including the model domain, boundary and forcing conditions, spatial resolution, time step, and the representation of coastal and harbor structures; iv) the characteristics of the 10 × 10 spatial target grid used by the surrogate model should be reported, including its extent, resolution, and the bathymetric and topographic datasets employed; v) a clear description on how the dataset have been used for training and testing the surrogate models is missing; vi) it is unclear how the 0D tide-gauge observations, the 1D XBeach profiles, and the 2D surrogate-model grid are integrated.
Response: The Methods have been clarified around the physically consistent sequence used in the revised analysis: documented forcing → sector-specific terrain profiles → 14 fixed-bed XBeach simulations → time-resolved machine-learning ensemble → event-conditioned Monte Carlo application. The former synthetic 10 × 10 target is not retained in the final surrogate workflow; consequently, the previous ambiguous 0D–1D–2D mapping is removed. The surrogate predicts three diagnostics extracted directly from each one-dimensional XBeach state: maximum land water depth, shoreline-connected flood extent, and number of wet land nodes. The XBeach configuration, forcing, non-uniform cross-shore grids, 73 h simulation duration, 600 s output interval, fixed-bed formulation, lack of explicit harbour structures, 0.05 m wet/dry diagnostic, one-hour spin-up exclusion, predictor definitions, model-development procedure, and independent December 2024 temporal test are now stated explicitly. These methodological clarifications are the same revision set already described in our response to Reviewer 2.
Revised paper text: “In this study, inundation depth is defined as the positive XBeach-modelled water-column thickness above the local bed elevation at landward grid nodes, h = max(zs - zb, 0). A land node is considered wet when h is at least 0.05 m. The contiguous flood extent is the landward distance from the shoreline to the last consecutively wet node before the first dry interruption. These quantities are XBeach model diagnostics; they are not tide-gauge sea level and they are not direct field observations of inundation.”
Major comment 3 — Technical terminology
3.1 Coastal flood depth, inundation depth, inundation field, inundation response, spatial inundation. In my opinion, the term inundation is not appropriate in this context, since both the observations and the modelling framework appear to represent the (total) sea level rather than actual land flooding. A clear definition of inundation depth should be clearly stated at the beginning of the Methods section.
Response: Agreed. Inundation is now defined only from positive XBeach-modelled water-column thickness over land and is not inferred from tide-gauge sea level. The definition is stated at the beginning of the Methods section.
Revised paper text: “In this study, inundation depth is defined as the positive XBeach-modelled water-column thickness above the local bed elevation at landward grid nodes, h = max(zs - zb, 0). A land node is considered wet when h is at least 0.05 m. The contiguous flood extent is the landward distance from the shoreline to the last consecutively wet node before the first dry interruption. These quantities are XBeach model diagnostics; they are not tide-gauge sea level and they are not direct field observations of inundation.”
3.2 XBeach-informed physical baseline, physics baseline. These terms appear to refer simply to the outputs generated by the XBeach model. I therefore suggest replacing them with a more straightforward expression such as XBeach outputs or XBeach simulations. The terms informed, physical, and baseline add unnecessary complexity and may confuse the reader.
Response: Agreed. The terms “XBeach-informed physical baseline” and “physics baseline” are not used in the final formulation. The surrogate is trained directly on XBeach outputs, and the previous residual-baseline formulation is not retained.
3.3 Storm surge, storm-surge setup, sea-level anomaly. Tide gauges record the total sea level and additional processing is required to separate and quantify the different components of sea-level variability (e.g., storm surge, sea-level anomaly, astronomical tide). The paper should clearly explain how these quantities were derived and used.
Response: Agreed. We do not decompose tide-gauge total water level into astronomical tide, storm surge, or harbour-resonance components for the surrogate workflow. The offshore water-level boundary is described consistently as CMEMS sea-level forcing (CMEMS zos).
Revised paper text: “The revised paper uses the terms CMEMS sea-level forcing, XBeach simulation/output, documented storm event, maximum land water depth, contiguous flood extent, and wet land nodes. No decomposition of tide-gauge total water level into astronomical tide, storm surge, or harbour-resonance components is used in the surrogate workflow.”
3.4 Storm sequence: It is unclear what is meant by this term. The authors may be referring to storm duration, a storm event, or a storm case. Please either use a more precise term or clearly define its meaning in the paper.
Response: Agreed. “Storm sequence” has been replaced by the more precise terms “documented storm event”, “event time”, or “time-resolved event response”, depending on context.
Major comment 4
Lastly, the results indicate that the different surrogate models exhibit very similar performance. I therefore suggest presenting only the ensemble mean results in the main paper and moving the complete set of performance statistics to the Supplementary Material. This would improve the readability of the paper without affecting the interpretation of the results. Furthermore, the authors should consider that these findings may be strongly influenced by the methodological limitations highlighted in the above concerns. Therefore, the reported model performance should be interpreted with caution until the proposed framework is validated against appropriate inundation observations.
Response: Agreed. The main Results focus on the validation-weighted ensemble, while individual RF, ET, GB, and HGB statistics are retained in the Supplementary Material. The updated independent-event evaluation shows a clear target-dependent skill pattern: shoreline-connected inundation extent and wet-land-node count remain well reproduced, whereas maximum land water depth shows limited temporal-variance skill on the unseen December 2024 event. We therefore report the low depth R² explicitly and interpret the surrogate primarily as an extent/wetting screening emulator rather than as a precise predictor of time-resolved maximum depth. The sector-wise event-peak diagnostic is consistent with this interpretation (R² = 0.279, 0.971, and 0.966 for depth, extent, and wet-node count, respectively). All reported statistics quantify fidelity to XBeach outputs and not validation against direct field observations of inundation.
Revised paper text: “Surrogate skill is evaluated against XBeach outputs, not against tide-gauge observations. The November 2023 event provides 3,031 post-spin-up states for model development, while the December 2024 event is held out in its entirety as an independent temporal test with another 3,031 states. On that independent event, the validation-weighted ensemble gives MAE = 0.102 m, RMSE = 0.116 m, bias = -0.102 m, and R² = 0.015 for maximum land water depth; MAE = 9.25 m, RMSE = 14.21 m, bias = -9.18 m, and R² = 0.958 for shoreline-connected inundation extent; and MAE = 1.43 nodes, RMSE = 1.87 nodes, bias = -1.40 nodes, and R² = 0.971 for the number of wet land nodes. The near-zero R² for maximum land water depth indicates limited skill in reproducing its within-event temporal variance on the unseen storm, despite the comparatively modest absolute RMSE. The extent and wet-node diagnostics generalise substantially better. These statistics quantify emulation fidelity to XBeach and must not be interpreted as observational flood-validation skill.”
Additional comments
- As it stands, the abstract reads more like that of a paper intended for a national or regional journal. Moreover, it contains unnecessary details on the datasets and statistical results, including unexplained acronyms. The abstract should be streamlined to emphasize the study's objectives, methodology, key findings, and significance.
Response: The Abstract has been shortened and refocused on the methodological objective, the XBeach-to-ML workflow, the independent temporal test, the principal ensemble findings, transferability of the framework, and the explicit limitation that direct inundation observations are unavailable. Detailed dataset acronyms and the former long list of performance metrics have been removed from the Abstract.
- Lines 113-126: details about the datasets, model performance metrics and predictors should not be reported here.
Response: Agreed. Detailed implementation information is confined to the Methods, and quantitative performance statistics are reported in the Results. The Introduction retains only the scientific motivation, scope, and objectives.
- Lines 128-129: the flood extent, expressed as grid nodes exceeding an operational inundation, does not implies continuity of the flooding area. Please discuss the limitation of this approach.
Response: Agreed. Flood extent is now defined as a shoreline-connected one-dimensional cross-shore extent, requiring consecutive wet nodes from the shoreline. The paper also states explicitly that this is a transect metric, not a two-dimensional flood polygon or an alongshore-connected flooded area.
Revised paper text: “The revised flood-extent metric explicitly enforces one-dimensional shoreline connectivity along each cross-shore profile: only consecutive wet land nodes beginning at the shoreline are counted. This removes the former ambiguity associated with counting isolated threshold exceedances. It should nevertheless be interpreted as a 1-D transect extent, not as a continuous two-dimensional inundation polygon or an alongshore-connected flooded area.”
- exceeding an operational inundation threshold
Response: The wording has been corrected. The revised text uses “wet/dry criterion” and “shoreline-connected flood extent” rather than the incomplete phrase in the previous version.
- Lines 141-144: is this a result of the present study? If so, it should be moved to the Results section. Otherwise, an appropriate reference should be provided.
Response: The event-specific statement has been removed from the Introduction; event-specific findings are reported in the Results section.
- Line 154: main Romanian Western coastal sectors.
Response: The geographical wording has been standardised to “Romanian Western Black Sea coast” or “Western Black Sea sectors”, depending on context.
- Table 1: include the temporal extension and the number of valid data for each station.
Response: In the revised analysis the former tide-gauge statistics table is no longer used as an inundation-validation table. Table 1 documents the seven XBeach sectors and terrain sources, while the Methods report the two storm windows and the number of post-spin-up XBeach states: 433 states per sector per event, i.e. 3,031 states per event.
- Figure 2: This figure suggests that the tide gauges use different vertical reference datums, while Table 1 reports very similar mean SLEV values. Please clarify whether the data were adjusted to a common datum before the analysis.
Response: The former cross-station tide-gauge time-series comparison is not used in the validation chain. The revised Figure 2 presents CMEMS–ERA5 forcing for the two documented events. Terrain is expressed in a local shoreline-relative model datum with z = 0 at the shoreline, and CMEMS zos provides the offshore time-varying water-level boundary. Because a common independently verified geodetic transformation is not available for every sector, this limitation is stated explicitly and absolute depth values are interpreted cautiously.
Revised paper text: “The former cross-station comparison of tide-gauge elevations has been removed from the validation chain. Terrain is expressed in a local shoreline-relative model datum with z = 0 at the shoreline, while CMEMS zos provides the time-varying offshore water-level boundary. A geodetic transformation that would place all land-survey elevations and CMEMS sea level in a single independently verified vertical reference is not available for every sector; this is therefore treated as a limitation of absolute depth interpretation rather than concealed through datum pooling.”
- Line 175: what do you mean with related hydrodynamic variables. You should only mention the variables that have been used in the present study.
Response: Agreed. The vague phrase “related hydrodynamic variables” has been removed. The revised data table lists only variables actually used in the final workflow.
- Line 182: provide the reference for the topo-bathymetric dataset.
Response: The terrain sources are now stated explicitly: EMODnet DTM 2024 for marine bathymetry, Copernicus DEM GLO-30 for land elevation, shoreline/GNSS anchors for coastline registration, and the Zenodo RTK-GNSS/multibeam dataset for Vama Veche.
- Table 2: many variables (SLEV, ATMS, Hs, Tp, sea-surface state, SST, MSLP) and roles (storm surge, erosion-related forcing, plume related forcing, …) need to better defined.
Response: Table 2 has been simplified so that each retained variable and its role are explicit. Variables from the previous framework that are not used in the final surrogate are not retained as core predictors.
- Line 203: define “target inundation response”.
Response: The generic phrase “target inundation response” has been replaced by the explicit XBeach-derived targets: maximum land water depth, contiguous shoreline-connected flood extent, and wet-land-node count.
- Lines 214-216: the variables storm duration, inverse-barometer contribution, wave-power-related terms, and interaction terms between sea-level anomaly and wave forcing must be clearly defined.
Response: The final predictor definitions are explicit: elapsed event time; Hs²Tp as a wave-power proxy; zos×Hs as a sea-level–wave product; cos(θwind−θwave) as wind–wave directional alignment; and max[cos(θwave),0] as an onshore-wave factor. The previous generic “storm duration” and inverse-barometer terms are not retained in the final feature set.
- Table 3: The mean values obtained from the Monte Carlo simulations (MCS) are significantly higher than the observed means. This discrepancy is particularly evident for sea level, for which the simulated mean exceeds the observed mean by more than 40%. The authors should provide an explanation for this difference and discuss its potential implications for the modelling results.
Response: The former pre-XBeach 5,000-scenario Monte Carlo table and its mean sea-level mismatch are not part of the final workflow. Monte Carlo is applied after surrogate training as an event-conditioned perturbation experiment based on the documented forcing states, with an explicit applicability-domain filter. It is therefore not interpreted as a synthetic climatology or observational validation.
Revised paper text: “The Monte Carlo stage was also reformulated. It no longer generates a synthetic forcing climatology before XBeach. Instead, 10,000 states per sector are sampled conditionally from the two documented-event forcing states with small perturbations, and only samples that remain inside the empirical surrogate-training envelope are used in the principal summaries. Depending on sector, 91.92–97.33% of generated states satisfy this applicability criterion (66,233 of 70,000 states overall). These results describe event-conditioned surrogate uncertainty and response ranges; they are neither a flood climatology nor an observational validation.”
- Lines 229-230: define “design-of-experiments forcing combinations”.
Response: The former “design-of-experiments forcing combinations” concept has been removed from the final workflow.
- Line 230: why sea-level anomaly?
Response: The final workflow does not derive a tide-gauge sea-level anomaly. It uses CMEMS zos directly as the time-varying offshore sea-level forcing.
- Line 231: specify the topo-bathymetric dataset used for implementing the XBeach cross-sections.
Response: The terrain sources used to construct the XBeach transects are now specified in the Study area/Data and Methods sections and in the revised data-source table: EMODnet DTM, Copernicus DEM GLO-30, shoreline/GNSS anchors, and the local Vama Veche RTK-GNSS/multibeam dataset.
- Lines 234-236: I do not agree with this approach. Each component of the modelling system must be calibrated and verified to obtain a reliable prediction system.
Response: We agree with the principle. Because direct spatial inundation observations are unavailable for the two analysed events, the fixed-bed XBeach simulations are not presented as observationally calibrated. The present paper is therefore limited to emulator development and screening, and direct field calibration/validation is identified as necessary before operational prediction of absolute inundation depth.
Revised paper text: “The present framework remains intentionally limited. It uses one-dimensional fixed-bed XBeach transects, does not resolve harbour geometry, alongshore exchange, morphodynamic feedbacks, or two-dimensional overtopping pathways, and has not been calibrated against direct field maps of inundation depth or extent. Consequently, the surrogate is appropriate for rapid emulation, scenario ranking, and screening of the XBeach response space, but not yet for stand-alone operational prediction of absolute coastal flood depth.”
- Lines 236-238: this part should be moved to section 3.4.
Response: The relevant methodological material is now located with the surrogate formulation in Sect. 3.4; no additional restructuring beyond the revision already made is necessary.
- Figure 5: The left panel clearly shows that the XBeach profiles provide an unrealistic representation of the offshore/submerged part of the coastal profile, including a vertical wall exceeding 5 m in height. The authors should explain the origin of this feature and discuss its potential effects on the model results.
Response: This issue was addressed through the revised profile-construction and quality-control workflow. The former artificial offshore wall is not retained. Profiles are shoreline-registered and assembled from source-aware marine bathymetry and land topography, with explicit checks for the marine/land transition and offshore re-emergent nodes. Steep sectors are flagged for visual quality assurance rather than being silently smoothed.
Revised paper text: “The profile-generation workflow was revised to remove the artificial offshore vertical-wall artefact identified by the reviewer. The shoreline is imposed at z = 0 using the available GNSS/shoreline anchor, bathymetry is taken from EMODnet (with the Vama Veche nearshore profile supplemented by the RTK-GNSS/multibeam Zenodo dataset), and land elevation is taken from Copernicus DEM GLO-30. Source-aware interpolation is used only across the unresolved nearshore transition. Quality control confirms no offshore re-emergent land nodes in the final source-aware profiles; steep nearshore sectors are retained with explicit visual-QA flags rather than being silently smoothed.”
- Line 291: from where do you get observations of the inundation depth at grid-node level?
Response: There are no field observations of inundation depth at grid-node level. The revised paper states this explicitly. The surrogate targets are extracted from XBeach outputs and are model diagnostics, not observations.
- Lines 296, tables 5 and 6, and somewhere else: always specify “mean” when referring to average value of the surrogate models.
Response: Agreed. The revised paper uses the term “validation-weighted ensemble” and distinguishes ensemble values from individual-model values. Individual-model statistics are provided in the Supplementary Material.
- Table 5: why the physics baseline has negative R2?
Response: The previous physics-baseline/residual formulation and its negative R² are not part of the final analysis. The final model is a direct multi-output surrogate of XBeach simulations.
- Figure 6: What’s the Observed Coastal Flood Depth? Tide gauge measure sea level not flood depth.
Response: Agreed. The label “Observed Coastal Flood Depth” has been removed. Validation figures compare XBeach reference outputs directly with the ML ensemble and do not label tide-gauge measurements as flood depth.
- Captions of figures 6, 7, 8 and 9: do not report comments in the captions.
Response: Agreed. Figure captions are descriptive only; interpretation is kept in the Results and Discussion text.
- Section4.5: explain how the permutation-importance diagnostics have been computed.
Response: The computation is now stated explicitly. On the independent December 2024 event, each predefined predictor group is permuted while the trained ensemble remains fixed; the perturbation is repeated five times, and importance is measured as the increase in RMSE relative to the unpermuted prediction.
Revised paper text: “Permutation importance is computed on the independent December 2024 event by randomly permuting each predefined predictor group while leaving the trained ensemble unchanged, repeating the perturbation five times, and reporting the increase in RMSE relative to the unpermuted prediction. The grouped analysis shows that profile geometry and site identity dominate all three outputs, with sea level providing the next consistent contribution. Wave forcing and the explicit sea-level–wave interaction have small additional permutation importance in this two-event experiment; no claim of a robust compound sea-level–wave interaction is therefore made.”
- Line 353: clarify how the interaction between sea-level anomaly and significant wave height have been defined.
Response: The retained interaction feature is defined explicitly as zos×Hs, where zos is CMEMS offshore sea-level forcing and Hs is significant wave height. We do not refer to this as a tide-gauge-derived sea-level anomaly.
- Line 356: no compound sea-level–wave interaction effects can be reliably assessed in this study, as the observations were collected from tide gauges located within harbors, and the XBeach simulations include an unrealistic 5 m vertical wall structure.
Response: We agree that the present two-event experiment does not support a robust physical claim about compound sea-level–wave interaction. The artificial profile artefact has been removed from the revised terrain workflow, and the grouped permutation analysis shows only small additional importance of the explicit sea-level–wave product. The paper therefore avoids interpreting this feature as a reliably quantified compound interaction.
- It is not necessary to include both Figure 7 and Table 8 as they report the same information.
Response: Agreed. The redundant presentation has been removed; the main paper retains the graphical summary, while detailed values are kept in supporting material where appropriate.
- Section 4.7: from where these other stations come from?
Response: The seven XBeach modelling sectors are introduced explicitly at the beginning of Sect. 2 and documented in the modelling description/Table 1: Constanța, Mamaia, Midia, Eforie Nord, Mangalia, Vama Veche, and Gura Portiței. Figure 1 is retained as the MHD sea-level station map for geographic and observational context; the four tide-gauge locations shown there are clearly distinguished from the seven XBeach modelling sectors and are not treated as the inundation-validation network.
- Figure 9b: do not mix observation with XBeach model output.
Response: Agreed. The revised validation figures compare XBeach outputs directly with the ML ensemble. The paper states repeatedly that this is simulator-emulation validation and not observational flood validation.
Revised paper text: “Surrogate skill is evaluated against XBeach outputs, not against tide-gauge observations. The November 2023 event provides 3,031 post-spin-up states for model development, while the December 2024 event is held out in its entirety as an independent temporal test with another 3,031 states. These statistics quantify emulation fidelity to XBeach and must not be interpreted as observational flood-validation skill.”
Citation: https://doi.org/10.5194/egusphere-2026-3628-AC4
Viewed
| HTML | XML | Total | BibTeX | EndNote | |
|---|---|---|---|---|---|
| 124 | 41 | 16 | 181 | 10 | 11 |
- HTML: 124
- PDF: 41
- XML: 16
- Total: 181
- BibTeX: 10
- EndNote: 11
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
The manuscript provides a robust and comprehensive description of the materials and methods. However, several aspects could be improved to enhance clarity and strengthen the scientific presentation.
In the Introduction, describing 7 specific objectives is unnecessarily detailed and makes the section appear repetitive, as it lists individual analytical steps rather than the overarching research goals. Would be positive to restructuring them into three broader objectives following a clearer logical progression: (1) data and framework development; (2) model development and validation; and (3) process interpretation and practical application.
The Methods section is comprehensive and well documented. However, publicly available datasets should also be listed in the Data Availability Statement, including their repositories, persistent links (e.g., DOI or URL), and data providers. In addition, datasets, source code, or algorithms that are available upon reasonable request could also be mentioned to further strengthen compliance with the FAIR principles. Adopting these recommendations would improve the transparency, reproducibility, and overall impact of the study. Furthermore, several methodological choices and parameter settings are presented without supporting references. Although many of these statistical approaches are well established, citing the original or standard methodological references would better justify their selection and increase the scientific credibility of the adopted framework.
The Discussion provides a clear interpretation of the results but would benefit from a stronger comparison with previous studies. While the Introduction section is well referenced, Section 5.1 would be strengthened by discussing similar studies and demonstrating that the observed model-skill patterns are consistent with findings reported in the literature. Likewise, Section 5.2 highlights the potential and limitations of surrogate models for early-warning systems but does not include references to comparable applications using similar methodologies. Finally, the discussion in Section 5.3 regarding the reduced performance for extreme-tail events should also be supported by references to studies reporting similar limitations, as this is a common challenge in surrogate modelling and extreme-event prediction.
Overall, the study successfully addresses its stated objectives and presents a framework with significant practical relevance. The proposed methodology has considerable potential for regional coastal-flood hazard assessment and provides a valuable reference for the development of operational early-warning systems. In particular, the framework demonstrates an appropriate balance between computational efficiency, physical realism, and transparent communication of uncertainties, making it well suited to support coastal management and decision-making for compound coastal flooding.