the Creative Commons Attribution-NonCommercial 4.0 International License.
the Creative Commons Attribution-NonCommercial 4.0 International License.
Robust Uneven Shift of Extreme Storm Surges Observed in Data Sparse Northeast Indian Ocean Cities
Abstract. Reanalysis-driven storm surge datasets enable extreme analysis, but previous studies translate these time series into extremes using a single statistical model, leaving model-selection uncertainty unquantified. In this study, we analyze ERA5-forced surge residual dataset (1950–2024) from Copernicus Climate Change Service for data-sparse 11 Northeast Indian Ocean (NIO) cities using an ensemble of nonstationary extreme value and Bayesian formulations to estimate return levels, implied return period changes in 2000 relative to 1950 baselines, and trends. Our multi-model ensemble analyses reveal that for inner NIO cities (e.g., South 24 Parganas, Patuakhali, Chittagong, Cox’s Bazar)–near the head of the Bay of Bengal–the 1950’s 50-year surge residual level (RL50) becomes more frequent by 2000 (typically a 33- to 39-year event corresponds to increase in median annual exceedance probability up to 51 %), even though surge residual annual-maxima trends are negative (~ - 1 mm/year). These changes are not uniform across the NIO as several western NIO cities (e.g., Colombo, Chennai) show the opposite tendency, with longer implied return periods (typically a 66- to 104-year event) by 2000. Statistical model choice substantially affects design levels and their interpretation in areas with high storm surges. For example, Gumbel or Bayesian median estimates of RL50 can correspond to roughly a median of RL25 under Frechet tail assumptions for inner NIO cities, highlighting nontrivial structural uncertainty. Finally, we show that bias of High Resolution Model Intercomparison Project (HighResMIP)– ERA5 depends on the statistical model used to estimate return levels, and that an ensemble-of-models evaluation provides a conservative and transparent basis for benchmarking climate-model surge extremes against reanalysis.
Status: open (until 07 Oct 2026)
-
RC1: 'Comment on egusphere-2026-3651', Anonymous Referee #1, 28 Aug 2026
reply
The comment was uploaded in the form of a supplement: https://egusphere.copernicus.org/preprints/2026/egusphere-2026-3651/egusphere-2026-3651-RC1-supplement.pdfReplyCitation: https://doi.org/
10.5194/egusphere-2026-3651-RC1 -
RC2: 'Comment on egusphere-2026-3651', Anonymous Referee #2, 01 Oct 2026
reply
This manuscript tries to propose a framework for providing more robust return level estimations of storm surges by using four nonstationary statistical models, rather than relying on a single extreme value distribution as done traditionally. While the topic is relevant and timely, I am not fully convinced by the proposed approach in its current form. My concerns mainly relate to the suitability of the input data for the study area and the way how the results from the four statistical models are combined and interpreted. Please find my detailed review below. Hope these comments would help the authors improve their framework and the manuscript.
Major comments
1. My main concern relates to the motivation of the study. On many occasions (e.g. the title, L120-127, L162-166, and etc.), the manuscript appears to present the proposed framework as a means of providing robust return period estimates of extreme storm surges in data-sparse regions. However, in my opinion, the input GTSM products already address, at least partially, address this issue by providing continuously long-time data at a 2.5km resolution along the coastline. In this context, it is not entirely clear to me in what way the data used in this study can be considered “sparse,” or what specific data-scarcity problem the proposed framework is intended to overcome.
A more fundamental concern, in my view, is whether the GTSM-ERA5 simulations can adequately represent the extreme storm surges relevant to this study. As the authors acknowledged in the manuscript, storm surges at several locations within the study area are mainly affected by TCs (e.g. L237-239). To my knowledge, ERA5 still struggles to accurately represent the TC-related variables because of its relatively coarse resolution even though it has been improved a lot compared to ERA-Interim. This misrepresentation would subsequently propagate into the hydrodynamic simulations for accurately estimating storm surge levels. Besides, due to the stochastic nature of TCs, very long records (e.g. 10,000 synthetic time series) are generally required to characterize the return periods of TC-driven storm surges (e.g. Bloemendaal et al. (2020), Dullaart et al. (2021), Benito et al. (2023)). And for regions affected by both TCs and ETCs like the study area in this manuscript, a first-order event stratification can also be useful for distinguishing surge populations associated with different storm types and for better estimating the frequencies/return levels (e.g. Maduwantha et al. (2026)).
I therefore think the authors need to more carefully demonstrate that the underlying GTSM–ERA5 dataset is suitable for the proposed extreme value analysis. Alternatively, the proposed framework could be applied in a region (e.g. ETC-region) where ERA5-driven GTSM simulations have been shown to capture the relevant extremes more reliably, or the authors could consider stratifying events according to their generating mechanisms and carrying out the subsequent statistical analyses.
2. If I understood the methodology correctly, the authors applied their framework to two storm surge time series, namely GTSM-ERA5 and GTSM-EC-Earth3P-HR. However, this clarification appears late in the Results section (L325-329). I suggest making this explicit much earlier, e.g. in the abstract, introduction, and the method. More generally, the authors should carefully review every mention of the five HighResMIP climate models. Although five models may be available in the original dataset, the analysis presented here appears to rely on only one of them because EC-Earth3P-HR is the only model providing the required continuous record over the study period. The terminology throughout the manuscript should clearly distinguish between the broader HighResMIP dataset and the single GCM actually used in the analysis.
3. Related to Comment 2, I am not fully convinced that the inclusion of HighResMIP strengthens the analysis in its current form, particularly because only one GCM is ultimately used. The authors should clarify what additional scientific question the EC-Earth3P-HR analysis addresses beyond the ERA5-based analysis.
This is particularly important given the uncertainties in GCM-driven storm-surge simulations. For example, in Fig. 2 of Muis et al. (2023), the ensemble mean storm surges show high deviations (>0.4m) and biases (>40%) in the study region of this manuscript. Muis et al. (2023) (and some others) also discuss the importance of bias correction for GCM-driven products because individual climate models can exhibit substantially different biases. I therefore suggest that the authors provide a stronger justification for using EC-Earth3P-HR, evaluate its performance against the ERA5-driven simulations in the study region, and discuss whether bias correction is required before interpreting differences in extreme return levels.
4. Given that one of the motivations of this study is to provide robust frequency/return level estimates and the claims on the uncertainty of the conventional approach, I was surprised that the results from the four statistical models are not more directly evaluated against conventional single-curve frequency analysis approaches. The authors are encouraged to also add the empirical distribution (i.e. weibull’s plotting formula) of the annual maxima between 1950 and 2024 to this comparison.
I would also be interested to know whether the found temporal change in return levels can be identified using a conventional stationary analysis. For example, the authors could divide the record into two periods of approximately equal length (e.g., 1950–1986 and 1987–2024), fit the same stationary extreme-value distribution independently to each period, and compare the resulting return-level curves.
5. Could the authors explain why annual maxima were used? Several studies (e.g., Fischer & Schumann (2016)) have argued that the peaks-over-threshold (POT) approach can make more efficient use of the available information and can provide more robust frequency estimates than annual maxima, particularly for relatively short records. In addition, Wahl et al. (2017) explicitly recommended using a GPD-POT approach with a 99th percentile to assess extreme sea levels at the globe scale.
6. My other concern is on how the results from the four statistical models are combined. As I understand it, posterior samples from the four models are pooled with equal weight to construct the ensemble distribution. Although equal weighting seems an intuitive choice, it implicitly assumes that all four formulations should contribute equally, irrespective of their relative fit or predictive performance.
Besides, the GEV is developed by combining the Gumbel, Fréchet and Weibull families, i.e. the GEV follows the Gumbel when the shape is 0, the Fréchet when the shape is >0, and the Reverse Weibull when the shape is <0. Therefore, The four ensemble members therefore should not necessarily be interpreted as four independent and equally distinct distributional hypotheses.
7. I am also not sure about the city-scale virtual station approach. From an application and practical perspective, the information on the site-specific engineering design thresholds is needed rather than that at the city scale. Besides, the applied spatial model does not capture local nearshore processes, which made me question the actual reliability of this approach. While the manuscript acknowledges some of these limitations, I suggest that the authors more clearly justify why city-scale estimates are preferable than those at individual GTSM output locations. In my opinion, providing station-specific results would also make the analysis more transparent and facilitate the use of the results in other applications.
Minor comments
1. L54: Rephrase ‘because these extremes drive flooding impacts and coastal erosion.’ as ‘because these extremes can drive flooding impacts and cause coastal erosion.’ Not all extreme events necessarily result in flooding or subsequent impacts, as flooding depends on factors such as DEM and local protection structures, while impacts also depend on exposure and vulnerability.
2. L59: Define or give examples of ‘overconfident storm surge hazard estimates’.
3. L71: Further explain systematic regional biases and underestimation of extreme event intensities.
4. L132: The Deltares Global Tide and Surge Model (GTSM version 3.0) ‘was’ used to.
5. L134: The grid resolution of coastal potions is 1.25 km in Europe.
6. L130-134: The GTSM-HighResMIP only covers 1950-2014.
References
Bloemendaal, N., De Moel, H., Muis, S., Haigh, I. D., & Aerts, J. C. (2020). Estimation of global tropical cyclone wind speed probabilities using the STORM dataset. Scientific data, 7(1), 377.
Dullaart, J. C., Muis, S., Bloemendaal, N., Chertova, M. V., Couasnon, A., & Aerts, J. C. (2021). Accounting for tropical cyclones more than doubles the global population exposed to low-probability coastal flooding. Communications Earth & Environment, 2(1), 135.
Benito, I., Aerts, J. C., Eilander, D., Ward, P. J., & Muis, S. (2024). Stochastic coastal flood risk modelling for the east coast of Africa. npj Natural Hazards, 1(1), 10.
Muis, S., Aerts, J. C., Á. Antolínez, J. A., Dullaart, J. C., Duong, T. M., Erikson, L., ... & Yan, K. (2023). Global projections of storm surges using high‐resolution CMIP6 climate models. Earth's Future, 11(9), e2023EF003479.
Fischer, S., & Schumann, A. (2016). Robust flood statistics: comparison of peak over threshold approaches based on monthly maxima and TL-moments. Hydrological Sciences Journal, 61(3), 457-470.Citation: https://doi.org/10.5194/egusphere-2026-3651-RC2
Viewed
Since the preprint corresponding to this journal article was posted outside of Copernicus Publications, the preprint-related metrics are limited to HTML views.
| HTML | XML | Total | BibTeX | EndNote | |
|---|---|---|---|---|---|
| 177 | 0 | 2 | 179 | 0 | 0 |
- HTML: 177
- PDF: 0
- XML: 2
- Total: 179
- BibTeX: 0
- EndNote: 0
Viewed (geographical distribution)
Since the preprint corresponding to this journal article was posted outside of Copernicus Publications, the preprint-related metrics are limited to HTML views.
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1