the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
Retrieval of the hail size number distribution from polarimetric radar data using the double-moment normalization
Abstract. Estimating the distribution of hail sizes is essential for assessing potential damage to infrastructure, vehicles, and agriculture. In this study, we introduce a novel technique for retrieving the hail size number distribution (HSND) from polarimetric C-band radar data. Our approach uses a generalized additive model (GAM) to estimate two empirical moments of the HSND, which is then reconstructed using the double-moment normalization technique, exploiting the relative invariance of the normalized HSND. The model is trained using data from the Swiss automatic hailsensor network (August 2018–September 2025) across three hail-prone regions. Hundreds of polarimetric features are extracted from a high-resolution 3D radar composite combining data from the five operational, dual-pol Swiss radars. Among these, the most predictive features selected by the model include the volume of the region where the cross-correlation coefficient ρHV falls below 0.97 and the horizontal reflectivity ZH is above 50 dBZ in a vertical column of 1 km radius, the maximum value of vertical reflectivity ZV in a column of 1 km radius, the integral of ZV in a column of 1 km radius. HSND estimates derived from radar show strong agreement with independent hailsensor measurements. Additional validation is performed using photogrammetric drone surveys and crowd-sourced hail reports. The radar-retrieved HSND closely matches the shape of drone-based HSND for the two events, while overestimating the number of hailstones per diameter bin due to melted hailstones prior to drone observations and to the nature of the training data used (a minimum number of 30 hailstones must have been measured by a sensor for a single event to be used). Radar-based percentile diameters of the retrieved HSND exhibit a slightly higher Pearson correlation and lower bias with crowd-sourced hail reports compared to the MESHS (Maximum Expected Severe Hail Size) product operated by MeteoSwiss. The main advantage of the presented technique is that it enables high-resolution (1 km, 5 min) retrievals of the full HSND and related features, such as kinetic energy, potentially providing valuable insights for real-time hail monitoring, nowcasting and long-term statistical assessments of hail features across Switzerland. The proposed model could be easily adapted to other countries, though the invariance of the normalized HSND outside Switzerland should be verified. Because hail is rare, further ground hail observations remain crucial for refining the HSND retrievals, ensuring a more comprehensive evaluation of the proposed approach, and properly assessing the associated uncertainty.
Competing interests: At least one of the (co-)authors is a member of the editorial board of Atmospheric Measurement Techniques.
Publisher's note: Copernicus Publications remains neutral with regard to jurisdictional claims made in the text, published maps, institutional affiliations, or any other geographical representation in this paper. While Copernicus Publications makes every effort to include appropriate place names, the final responsibility lies with the authors. Views expressed in the text are those of the authors and do not necessarily reflect the views of the publisher.- Preprint
(9655 KB) - Metadata XML
- BibTeX
- EndNote
Status: final response (author comments only)
- RC1: 'Comment on egusphere-2026-594', Anonymous Referee #2, 23 Jul 2026
-
RC2: 'Comment on egusphere-2026-594', Anonymous Referee #4, 23 Jul 2026
Review Report
General Assessment
This manuscript presents a novel methodology for retrieving hail size number distributions (HSND) from polarimetric C-band radar data using generalized additive models (GAMs) trained against automatic hailsensor measurements. The work builds upon prior double-moment normalization theory (Ferrone et al., 2024) and extends it to radar-based retrieval, with validation against drone photogrammetry and crowd-sourced reports. The topic is timely and relevant for AMT, and the authors have assembled an impressive multi-source observational dataset spanning multiple years.
However, the manuscript has several significant methodological and interpretational issues that need to be addressed before publication. The validation approach, while ambitious, contains circularities and inconsistencies that undermine confidence in the claimed performance. The uncertainty quantification is insufficient, and several critical assumptions are not adequately justified or tested.
Major Concerns
- Spatial Extrapolation of Hailsensor Data (Appendix C)
The spatial extrapolation procedure is the cornerstone of the training dataset but raises serious methodological questions:
The assumption that temporal variations at a fixed sensor represent spatial variability within the storm "segment" is untested and likely problematic. Hailfall is notoriously heterogeneous at small scales; assuming stationarity over 5 minutes and extrapolating up to 200 m is a strong assumption without validation.
Lines 317-318 state that "The added value of the spatial extrapolation was found to be limited, with meaningful contributions extending to a maximum distance of roughly 200 m." If the extrapolation adds so little value, why is it included? This suggests that the method is effectively training on point measurements against radar volumes that are ~5 orders of magnitude larger—a fundamental scale mismatch.
The extrapolation creates pseudo-replicates in the training data. Multiple extrapolated positions from the same sensor observations are treated as independent training samples, artificially inflating the effective sample size and likely leading to overconfident predictions.
- Feature Selection and Model Overfitting
Section 3.4.2 describe the feature selection process, which raises serious red flags:
The features were selected based on cross-validated Spearman correlation with the extrapolated moments (which are derived from the same hailsensor data). This creates a risk of information leakage, especially given the limited number of independent events (35 days).
The model uses 4 features to predict two moments from only 35 independent days. With this limited sample size, the risk of overfitting is high, despite the leave-one-event-out cross-validation. The authors acknowledge that "the limited amount of reference data from hailsensors makes robust cross-validation challenging" (lines 491-492), yet proceed with the selected features.
Table A2 shows only marginal improvements in PCC when adding features (e.g., M4 PCC: 0.64 → 0.65 → 0.66 → 0.68). The incremental gain is modest, and the model complexity increases substantially. The authors should report AIC or BIC to justify the additional features.
- Posterior Bias Adjustment (Section 3.4.3)
The bias adjustment procedure (quantile mapping + linear regression) is concerning:
This is essentially post-hoc calibration that uses the same training data to correct the model predictions. While applied in cross-validation, the procedure reduces the interpretability of the GAM's predictive skill.
The adjusted moments show increased RMSE compared to the unadjusted predictions (Table A2: M2 RMSE from 0.31 → 0.34; M4 from 0.41 → 0.44), yet the authors present the adjusted results as superior. This is counterintuitive—if the adjustment increases RMSE, in what sense is it improving performance?
The authors should clearly distinguish between calibration (improving probabilistic properties) and accuracy (reducing error). Quantile mapping improves distributional alignment at the cost of increased point-wise error, which should be discussed transparently.
- Validation Against Drone Data (Section 4.3)
The drone validation contains critical issues:
The melting correction (0.5 mm/min) is arbitrary for the 2022 event, as acknowledged in lines 571-573.
The statement "the HSND obtained via double-moment normalization from hailsensor observations agrees more closely with the radar-based HSND than with the drone-retrieved HSND" (lines 577-578) undermines the value of drone data as independent validation. If the hailsensor (30 m away) matches the radar better than the drone observations, why trust the drone validation?
Figure 7b shows that a 10-minute melting correction gives "great agreement" while the 20-minute correction does not. This is curve-fitting rather than validation, as the melting correction is not independently constrained for this event.
- Validation Against Crowd-Sourced Reports (Section 4.4)
The crowd-sourced validation has a data leakage issue:
The spatial aggregation (3-11 km) and the use of maximum values for radar features compared against percentiles of reports creates an apples-to-oranges comparison. MESHS (also a maximum) is a more appropriate comparator. The authors should justify why maximum radar values should be compared against 75th/90th percentiles of reports.
The improved correlation of HSND percentiles over MESHS (Fig. 8) is not surprising: HSND percentiles are tunable parameters (D99.9, D99.99, etc.) that can be adjusted to match the reports, while MESHS is a fixed operational product. The authors should use a fixed percentile (e.g., always D99.9) and avoid cherry-picking the best performer.
- Non-Rayleigh Scattering Treatment (Section 4.1.1)
The discussion of non-Rayleigh scattering is inadequate:
Lines 495-496 acknowledge that C-band resonance effects occur for hailstones >2 cm, which directly affects the ZV features used in the GAM. The authors state that moment 2 is "generally less sensitive" and moment 4 "is more directly impacted," but provide no quantitative assessment of how resonance affects their specific features.
Given that the hailsensor data include events with hailstones up to 4.2 cm (Fig. 3l), the GAM is being trained on data that may have nonlinear and non-monotonic relationships between radar features and ground truth due to resonance. The monotonic splines used in the GAM may be inappropriate for these conditions.
Minor Issues
- Sample Size and Representativeness
Only 35 days of hailsensor data are used, spanning 2018-2025 (7 years). This is a small sample for a statistical model with 4 features and smooth splines. The authors should report the effective sample size considering the spatial extrapolation (which likely reduces independence).
The training dataset is biased toward intense events (≥30 hailstones per event), which the authors acknowledge. This may explain the systematic overestimation of small hailstones in the drone comparison. However, the authors do not propose a correction or weighting scheme for this selection bias.
- Missing Uncertainty Quantification
No uncertainty estimates are provided for the radar-retrieved HSND. The authors mention that "explicit quantification of model uncertainty" is "not addressed" and should be "explored" (lines 635-638). This is a significant limitation for an operational product and should be addressed before publication.
The authors should at minimum provide prediction intervals for the GAM outputs, perhaps using bootstrapping or quantile regression.
- Physical Interpretability of Selected Features
The selected features (particularly the ρHV < 0.97 and ZH > 50 dBZ volume) are physically plausible (mixed-phase regions with high reflectivity). However, the sum of ZV and maximum ZV features are highly correlated (as noted in Fig. 4), and the physical distinction between them is not clearly explained. The authors should discuss whether these features could be combined or reduced without loss of skill.
The thresholds (ρHV < 0.97, ZH > 50 dBZ, ZV > 30 dBZ) were likely optimized empirically. The authors should explain how these thresholds were chosen and whether they are specific to the Swiss radar calibration or transferable.
- MESHS Comparison
The comparison with MESHS (Section 4.4) is valuable, but MESHS is a single-polarization product. Comparing a dual-polarization model against a single-polarization baseline is not a fair test; the comparison should be against state-of-the-art dual-polarization hail products.
The authors should also compare against POH (Probability of Hail) to demonstrate improvement in detection, not just size estimation.
- Temporal Resolution and Advection
The 5-minute temporal resolution of the radar composite is coarse relative to the 20-s grouping of hailsensor data (Appendix C). The authors should discuss whether this temporal mismatch affects the feature-target relationships.
The advection correction (Appendix C) uses TRT storm motion, but hailstones have different fall speeds than the storm cell. The 30-s time offset (Appendix D) found empirically is plausible, but the authors should justify this based on physical fall speeds (e.g., 20-40 m/s fall speed × 30 s ≈ 600-1200 m aloft, which seems large relative to the assumed 200 m extrapolation distance).
- Writing and Clarity
The manuscript is generally well-written but is too long (46 pages with appendices). The extensive appendices (A-E) contain essential methodological details that should be integrated into the main text. The current structure makes it difficult for readers to follow the workflow (Fig. 2 provides an overview, but the details are scattered).
The abstract should more clearly state the limitations of the approach (bias toward intense events, uncertainty quantification missing, small sample size).
Specific Technical Corrections
- Line 370-371: The phrase "the limited amount of data for large hail diameters prevents the GAM from constructing well-adapted spline functions" contradicts the earlier claim that the model "generalizes well." This should be rephrased to acknowledge the limitation more clearly.
- Line 407: The notation "D99.9, D99.99, D99.999, D99.9999" is confusing. For a dataset with 30-100 hailstones per event, the 99.99th percentile corresponds to a probability level that is not statistically meaningful. The authors should use return levels or N-year exceedance estimates instead.
Citation: https://doi.org/10.5194/egusphere-2026-594-RC2
Viewed
| HTML | XML | Total | BibTeX | EndNote | |
|---|---|---|---|---|---|
| 53 | 50 | 12 | 115 | 9 | 10 |
- HTML: 53
- PDF: 50
- XML: 12
- Total: 115
- BibTeX: 9
- EndNote: 10
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
This paper is based on the dual-polarization radar of Switzerland. It combines the dual-moment normalization theory and the generalized additive model (GAM) to invert the size distribution of ground hail particles (HSND) using the dual-polarization radar data. It integrates multi-source data such as automatic hail sensors, unmanned aerial vehicle photogrammetry, and crowdsourced hail reports for model training and validation. The research idea is clear, the data is detailed, and the methods are complete. It has clear innovations in the field of radar quantitative hail inversion in complex Alpine terrain. Before publication, there are still some issues that need to be addressed.
The manuscript uses the fixed normalized shape function eliminated by Ferrone et al. (2024), but the applicability of this function within the retrieval framework of this study needs further verification. In different radar bands and under different types of thunderstorms, are the parameters of the normalized function fixed, or does there exist variability? If possible, provide some sensitivity verification experiments.
Are the radar data used for training GAM and the drone/crowd-sourced report used for verification completely independent in terms of time and space? The drone observations may be from the same storm as the radar features and sensor data, which could lead to overly optimistic validation results. If there is spatial correlation among multiple radar data from the same event, it can also result in information leakage.
Feature selection was carried out stepwise based on cross-validation using Spearman correlation, but only 4 features were ultimately selected. Among hundreds of candidate features, this stepwise selection method may amplify accidental correlations, especially in cases with limited sample size. If possible, it is recommended to randomly shuffle the target variables and repeat the feature selection process to see if the selected features remain stable.
When comparing the drone data, the melting correction for the radar inversion of HSND was performed. However, this rate (0.5 mm/min) was only derived from one event in 2021 and was directly applied to another event in 2022. The environmental temperature, humidity, and ground conditions of the two events were different, and this assumption may lead to biased estimation.
The GAM model only provides quantitative estimates without any uncertainty intervals (such as uncertainty range, confidence interval, etc.). It is suggested to provide standard deviation or percentile intervals for model predictions.
Five radars are distributed in different areas, with differences in altitude and terrain conditions. When merging the detection data from the five radars, how are the effective detection range of the radars, the radar blind zones, terrain clutter, and the influence of rugged terrain considered, as well as how to unify the characteristic data from different heights?
The hailsensors, drone, and crowdsourced reports show distinct cluster characteristics. They mainly focuses on three centers in this study. Among them, their observations in Locarno Monti is closest to the Monte Lema radar, while the other two clusters are farther away from the radar. It is suggested to conduct some sensitivity tests to discuss the dependence of the model's prediction results on the radar's detection distance.
Extreme hailstones are more destructive in nature, although the samples of extreme hailstones are extremely limited. It is still recommended to verify the performance of the model in Large-diameter hailstones events. Additionally, hailstones and strong convective precipitation share certain similarities in radar characteristics, and it is necessary to verify whether the model can still predict the presence of hailstones in extreme precipitation events.
The introduction section provides a detailed review of existing radar methods and ground observation techniques. However, the presentation of the innovative aspects of this study and the introduction of the specific technical approach are somewhat slow. Please streamline the introduction and emphasize the core innovations.