the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
A New Proposal for Optimizing Maximum Hydrological Events Fitting with Flexible TCEV Distribution
Abstract. Accurate characterization of extreme hydrological events is critical for flood risk assessment and hydraulic engineering design, particularly in the high-value cumulative distribution function (CDF, F(x)) range that governs design extremes. Hydrological records often consist of mixed populations of ordinary and extreme events, leading to a pronounced “dog-leg effect” that limits the applicability of conventional extreme-value distributions such as the Gumbel and Log-Pearson Type III. Although the Two-Component Extreme Value (TCEV) distribution is conceptually well suited to such mixed populations, its practical application is constrained by subjective parameter initialization, uniform weighting schemes that underrepresent right-tail extremes, and evaluation metrics with limited tail sensitivity. In this study, we propose a new fitting method for the TCEV distribution, SR-MWS, which uses piecewise linear fitting for stable initial parameters, right-tail-oriented weighting for extreme events, and a partitioned scoring framework to evaluate global and tail performance. The results of the hydrological dataset indicate that SR-MWS consistently outperforms existing TCEV estimation methods in accuracy and robustness. Further experiments based on simulated data show that this method achieves better global fitting performance while maintaining tail accuracy comparable to the Peaks-Over-Threshold (POT) method, and is significantly better than generalized extreme value (GEV) and Gumbel distributions in capturing extremes. By reducing subjectivity and enhancing robustness, the proposed method provides an automated framework for extreme-event modeling applicable to other mixed-population extreme-value problems.
- Preprint
(3124 KB) - Metadata XML
- BibTeX
- EndNote
Status: final response (author comments only)
-
RC1: 'Comment on egusphere-2026-2664', Anonymous Referee #1, 12 Aug 2026
-
CC1: 'Reply on RC1', Liangyu Ta, 20 Aug 2026
We sincerely thank the reviewer for the careful reading of our manuscript and for providing valuable and constructive comments. We appreciate the reviewer’s concerns regarding the validation framework, evaluation methodology, and interpretation of the proposed SR-MWS method.
We fully recognize that these comments address important aspects of the methodological rigor and applicability of the proposed framework. Below, we provide our preliminary responses and explanations. We appreciate these valuable suggestions and will carefully consider them during the revision process, including additional analyses, clarification of methodological descriptions, and refinement of the interpretation of the proposed framework.
The manuscript proposes a potentially useful modification to TCEV fitting, but its principal conclusions are not supported by the analysis. The problems are fundamental and require a new validation framework.
- SR-MWS is optimized, selected, and evaluated using RMSE against the same empirical CDF. A method designed to minimize this quantity will naturally outperform estimators optimized using likelihood. This demonstrates better in-sample curve fitting, not better prediction of extremes.
Re: We sincerely thank the reviewer for raising this important methodological concern. We agree that when the optimization objective and evaluation criterion are based on the same empirical CDF error measure, the resulting comparison primarily reflects the ability of a method to reproduce the observed sample distribution, rather than directly demonstrating its predictive capability for unobserved extreme events.
The original intention of using RMSE-based evaluation in this study was to assess the improvement of distributional representation, particularly in the upper-tail region of the cumulative distribution function. This is because the main motivation of the proposed SR-MWS framework is to address the insufficient tail representation of conventional TCEV parameter estimation methods, where ordinary events may dominate the fitting process and reduce the accuracy of rare extreme-event characterization.
However, we acknowledge that improved empirical CDF fitting alone cannot fully demonstrate improved estimation of extreme quantiles. In extreme-value analysis, the ultimate objective is not only to reproduce historical observations but also to provide reliable estimates of design values associated with long return periods. Therefore, the reviewer’s distinction between in-sample fitting performance and predictive evaluation is highly valuable.
To address this concern, we intend to further strengthen the validation framework by incorporating return-level-based evaluations in the simulation experiments during the revision process. Since the theoretical generating distributions are known in the simulation framework, the true return levels can be calculated directly and used as references for evaluating the accuracy of estimated extreme quantiles. Specifically, the performance of different methods can be compared based on deviations between estimated and theoretical return levels at selected return periods. RMSE-based evaluation will be retained as an indicator of distributional fitting performance rather than being regarded as the sole evidence of predictive superiority.
For the observed hydrological cases, where the true extreme quantiles are unknown, we will clarify that these datasets are mainly used to evaluate the applicability and practical fitting behavior of different approaches rather than providing direct validation of prediction accuracy.
We appreciate the reviewer’s comment, which helps us better distinguish between distribution fitting performance and predictive capability, and will revise the manuscript accordingly to provide a more comprehensive evaluation of the proposed SR-MWS framework.
- The generating parameters, number of experiments per distribution and sample size, random seeds, parameter ranges, optimization failures, and application of the slope-ratio screening are not reported. Equation (2) produces empirical plotting positions; not the theoretical CDF claimed by the authors.
Re: We agree that the description of the simulation framework in the original manuscript was not sufficiently detailed. Although the simulation experiments were designed to evaluate the performance of SR-MWS under different statistical conditions, some important implementation details, including the parameter settings of the generating distributions, the number of independent simulations, sample-size configurations, and optimization-related information, were not explicitly reported. We acknowledge that these details are essential for ensuring transparency and reproducibility.
The objective of the simulation experiments was to evaluate whether the proposed framework can maintain reliable extreme-value fitting performance under different distributional characteristics. Therefore, simulated datasets were generated from several representative distributions commonly used in hydrological frequency analysis, including TCEV, GEV, lognormal, exponential, and Pareto distributions. We agree that a more comprehensive description of the simulation design is necessary, including the parameter ranges or fixed parameters used for data generation, the number of experiments conducted for each distribution and sample-size scenario, the randomization procedure, and the treatment of potential optimization failures or convergence issues.
Regarding the application of the slope-ratio screening procedure, we appreciate the reviewer’s concern. The slope-ratio criterion was introduced as an applicability diagnosis step to determine whether a dataset exhibits sufficient two-component statistical characteristics before applying the TCEV-based framework. We will further clarify whether and how this screening procedure is applied within the simulation experiments and explain its role in the overall validation framework.
Regarding Equation (2), we appreciate the reviewer for identifying this imprecise terminology. We agree that the Weibull plotting-position formula does not represent the theoretical cumulative distribution function of the underlying population. Instead, it provides empirical plotting positions (or empirical cumulative probabilities) assigned to ordered observations. These empirical values are used to construct probability plots and compare observed samples with fitted theoretical distributions. The theoretical CDF is obtained from the assumed probability distribution model itself, such as the TCEV, GEV, or other extreme-value distributions.
Our intention was to describe the calculation of empirical cumulative probabilities from observed samples for subsequent distribution fitting and comparison, rather than to imply that Equation (2) represents the theoretical CDF. We acknowledge that the original wording was not sufficiently precise and may have caused confusion. Therefore, we will revise the terminology throughout the manuscript to clearly distinguish between empirical plotting positions and theoretical cumulative distribution functions.
We sincerely appreciate the reviewer’s careful reading and constructive suggestion. These comments will help us improve the reproducibility of the simulation framework and enhance the statistical rigor and clarity of the manuscript.
- A small probability error near 1 can produce an enormous error in the corresponding 100- or 200-year event magnitude. The authors must evaluate return-level bias, RMSE, and uncertainty at specific return periods.
Re: We fully agree that, in extreme-value analysis, evaluation in probability space alone may not fully reflect the practical implications of model performance. In particular, small deviations in the upper tail of the cumulative distribution function can result in substantial differences in estimated return levels, especially for long return periods such as 100- and 200-year events.
In the original manuscript, the RMSE-based evaluation was mainly designed to assess the improvement in distributional representation, with particular emphasis on the upper-tail region of the CDF. However, we acknowledge that the ultimate objective of extreme hydrological frequency analysis is to provide reliable estimates of design events associated with specific return periods. Therefore, return-level-based evaluation provides a more direct measure of the practical performance of different methods.
To address this concern, we will further strengthen the validation framework by incorporating return-level evaluations in the simulation experiments. Because the theoretical distributions used for simulation are explicitly defined, the corresponding true return levels can be calculated directly. This allows us to compare the estimated return levels from different methods against the theoretical values and quantify their performance using return-level bias and RMSE at selected return periods.
Furthermore, uncertainty analysis will also be considered in the revised evaluation framework to assess the robustness of return-level estimation. This additional evaluation will help determine not only whether a method can reproduce observed distribution characteristics but also whether it can provide stable estimates of rare extreme events.
For the observed hydrological cases, where the true return levels are unknown, we will clarify that these datasets are primarily used to demonstrate the practical applicability of different approaches rather than to provide direct validation of prediction errors.
- The framework depends on numerous arbitrary decisions: a slope ratio of 1.5, three points per segment, three selected weighting functions, an F= 0.8 threshold, and 60/30/10 scoring weights. None is justified or subjected to sensitivity analysis.
Re: We sincerely thank the reviewer for this valuable comment. We agree that the robustness of a methodological framework depends not only on its formulation but also on the rationality and stability of the parameters involved in its implementation.
We acknowledge that several choices in the current SR-MWS framework, including the slope-ratio threshold, minimum segment length, weighting schemes, tail-region threshold, and scoring weights, require clearer justification. These settings were introduced based on the characteristics of extreme-value analysis and the objective of improving the representation of the upper tail. However, we agree that their influence on the final results should be further examined.
Regarding the slope-ratio criterion, the threshold R=1.5 was introduced as a diagnostic criterion to identify datasets exhibiting a clear change in slope between ordinary and extreme-event regions after the Gumbel transformation. It is not intended as a universal physical threshold, but rather as a practical criterion to avoid applying the TCEV framework to datasets without evident two-component characteristics.
The minimum segment length of three observations was selected because three points were selected as the minimum requirement for estimating a local linear trend while maintaining applicability for relatively short hydrological records. A larger minimum segment length may improve regression stability but could reduce applicability to shorter hydrological records.
The three weighting functions (linear, quadratic, and exponential) were selected to represent different levels of emphasis on the upper tail, ranging from moderate to stronger tail weighting. Together, they provide a simple but representative set of weighting strategies for evaluating the effect of different tail priorities.
The F(x)=0.8 threshold and the 60/30/10 scoring scheme were designed according to the practical importance of extreme-event estimation. The high-probability region was assigned greater importance because it corresponds to longer return periods and is most relevant for hydrological design. Nevertheless, we agree that these choices should not be regarded as universal optimal values.
Therefore, in the revision, we will further clarify the rationale behind these methodological choices and perform sensitivity analyses by varying key thresholds and scoring configurations. The purpose of these analyses will be to examine whether the main conclusions of SR-MWS remain stable under reasonable alternative settings.
- A method receives the maximum score whenever it is the least poor candidate in an interval. Therefore, a score of 100 does not represent a perfect or even acceptable fit. Using the same observations for weight selection and performance evaluation also introduces selection bias.
Re: We sincerely thank the reviewer for this insightful comment regarding the interpretation and potential limitation of the partitioned scoring framework.
We agree that the score generated by the proposed framework should not be interpreted as an absolute goodness-of-fit measure. The purpose of this scoring system is to provide a relative comparison among candidate weighting schemes and to identify the weighting strategy that provides the best performance within the predefined evaluation regions. Therefore, a score of 100 indicates that a weighting scheme achieves the lowest RMSE among the tested candidates in a specific interval, rather than indicating a perfect or universally acceptable fitting performance.
We acknowledge that the original manuscript did not sufficiently clarify this point, which may lead to misunderstanding regarding the meaning of the score. We will revise the description of the scoring framework to emphasize its role as a relative selection criterion rather than an independent accuracy metric.
Regarding the potential selection bias caused by using the same observations for weighting scheme selection and evaluation, we appreciate the reviewer’s concern. The current framework was designed to optimize the representation of the upper-tail region by selecting the most appropriate weighting strategy among several predefined candidates. However, we agree that using the same dataset for both optimization and evaluation may overestimate the performance improvement.
To address this concern, we will further distinguish between the weighting-scheme selection procedure and the final performance evaluation. The selection of weighting functions will be treated as an internal optimization step, while the effectiveness of the resulting SR-MWS framework will be evaluated using additional criteria independent of the selection metric. These criteria will include return-level estimation accuracy (e.g., bias and RMSE) and uncertainty assessment in simulation experiments where the theoretical extreme quantiles are available.
- A dog-leg probability plot does not prove the existence of distinct flood-generating mechanisms. It may result from sampling variability, outliers, nonstationarity, or measurement inconsistencies. No meteorological or hydrological classification is presented.
Re: We agree that a dog-leg pattern observed in an extreme-value probability plot alone cannot be regarded as definitive proof of distinct physical flood-generating mechanisms. Such a statistical pattern may arise from multiple factors, including sampling variability, influential extreme observations, nonstationary hydrological processes, or uncertainties associated with hydrological measurements.
In this study, the dog-leg phenomenon was identified from multiple real-world hydrological datasets, including observed flood and precipitation records from different hydrological stations and basins. These datasets exhibited a consistent right-tail separation in the probability plots, indicating that the corresponding annual maximum series may contain heterogeneous statistical characteristics. The purpose of identifying this phenomenon was not to demonstrate the existence of two independent physical flood-generating mechanisms, but rather to diagnose whether the observed extreme-event population exhibits characteristics that cannot be adequately represented by a single homogeneous distribution.
Specifically, the observed dog-leg behavior suggests that ordinary events and rare extreme events may follow different statistical patterns within the same annual maximum series. This statistical separation provides motivation for considering a two-component extreme-value framework such as TCEV, which is designed to represent mixed extreme-value populations without requiring explicit attribution of each component to a specific physical mechanism.
We acknowledge that the original manuscript may have overemphasized the physical interpretation of the dog-leg phenomenon. In the revised manuscript, we will clarify that the dog-leg structure should be regarded as an indicator of mixed statistical characteristics rather than direct proof of specific meteorological or hydrological generation mechanisms. The underlying causes of such behavior may involve multiple interacting factors, including storm-type variability, catchment response differences, climate variability, and observational uncertainties.
Although this study does not include event-based meteorological classification or detailed hydrological process attribution, the observed right-tail separation in real hydrological datasets provides a valuable statistical basis for evaluating the applicability of mixed-population extreme-value models. Therefore, the objective of the proposed SR-MWS framework is to improve the statistical representation and parameter estimation of extreme hydrological records exhibiting heterogeneous characteristics, rather than to identify the specific physical origins of individual extreme events.
We appreciate the reviewer’s comment, which has helped us refine the interpretation of the dog-leg phenomenon and better define the scope and limitations of the proposed framework.
- The manuscript does not demonstrate that SR-MWS is robust, generalizable, or capable of accurately predicting events outside the historical record. Even the claim that it “consistently outperforms” existing methods is contradicted by Case 5 and by the exclusion of Cases 7 and 8.
Re: We also agree that the original statement that SR-MWS “consistently outperforms” existing methods requires more careful wording. The evaluation of an extreme-value framework should consider not only overall distributional fitting performance but also its capability to characterize upper-tail behavior and estimate rare-event quantiles beyond the range of historical observations. The primary objective of SR-MWS is to improve the representation of the upper-tail behavior of extreme-value distributions, which is directly related to long-return-period hydrological design events. Therefore, the performance evaluation of SR-MWS should consider both general distributional fitting and tail-specific estimation accuracy.
Regarding Case 5, we appreciate the reviewer’s observation. We acknowledge that SR-MWS does not achieve the best performance in every part of the distribution for every dataset, and therefore the original statement that SR-MWS “consistently outperforms” existing methods may be too broad. However, this case does not indicate a failure of the proposed method. The purpose of SR-MWS is not to minimize the global fitting error across the entire probability range, but to improve the characterization of extreme-event regions that are most relevant for hydrological design. In Case 5, although differences may exist among methods when considering the full probability range, SR-MWS shows competitive or improved fitting performance in the upper-tail region (e.g., F(x)>0.9), where rare extreme events and long return periods are represented. This result is consistent with the fundamental motivation of introducing the multi-weighting strategy, namely, reducing the tendency of conventional fitting approaches to be dominated by frequent events while underrepresenting extreme observations.
Regarding Cases 7 and 8, we would like to clarify that these datasets were not excluded because SR-MWS exhibited poor performance. Instead, they were screened out because the observed probability structures did not satisfy the applicability criterion established for the two-component TCEV framework. The purpose of the slope-ratio screening procedure is not to selectively remove unfavorable cases, but to determine whether the statistical characteristics of a dataset provide sufficient evidence for adopting a two-component extreme-value model. When the separation between ordinary and extreme events is not evident, applying TCEV may introduce unnecessary model complexity, and simpler single-component models, such as the Gumbel distribution, may provide a more appropriate representation. This interpretation is also consistent with previous studies indicating that datasets without clear evidence of mixed-population characteristics may be appropriately described by conventional single-component extreme-value distributions.
Therefore, Cases 7 and 8 represent the applicability boundary of the proposed SR-MWS framework rather than unsuccessful applications. The screening procedure is an important component of the proposed methodology because it prevents the inappropriate application of a complex two-component model to datasets that do not exhibit corresponding statistical characteristics.
We also acknowledge that the original manuscript may have overstated the general applicability of SR-MWS. The intended contribution of this study is not to demonstrate that SR-MWS universally replaces existing extreme-value models, but rather to provide an improved statistical fitting framework for hydrological datasets exhibiting heterogeneous extreme-value behavior and enhanced upper-tail characteristics.
In the revised manuscript, we will carefully revise the relevant statements, clarify the applicable conditions and limitations of SR-MWS, and avoid overgeneralized claims regarding its superiority over all existing methods.
We sincerely appreciate the reviewer’s comment, which has helped us better define the scope, advantages, and limitations of the proposed framework.
Citation: https://doi.org/10.5194/egusphere-2026-2664-CC1
-
CC1: 'Reply on RC1', Liangyu Ta, 20 Aug 2026
-
RC2: 'Comment on egusphere-2026-2664', Anonymous Referee #2, 30 Aug 2026
Please find my comments in the attached PDF.
-
AC1: 'Reply on RC2', Chen Yu, 08 Sep 2026
Thank you very much for your careful evaluation of our manuscript and for your constructive comments and suggestions.
We have carefully considered all the comments and prepared a point-by-point response.
Please find our detailed responses in the attached PDF.
-
AC1: 'Reply on RC2', Chen Yu, 08 Sep 2026
Viewed
| HTML | XML | Total | BibTeX | EndNote | |
|---|---|---|---|---|---|
| 177 | 78 | 22 | 277 | 16 | 13 |
- HTML: 177
- PDF: 78
- XML: 22
- Total: 277
- BibTeX: 16
- EndNote: 13
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
The manuscript proposes a potentially useful modification to TCEV fitting, but its principal conclusions are not supported by the analysis. The problems are fundamental and require a new validation framework.