Process-aware mixture modelling of precipitation-event distributions using event-type clustering
Abstract. Accurate modelling of event-level precipitation distributions is critical for hydrological modelling, flood risk assessment, and climate studies, yet selecting suitable parametric families remains challenging. We propose a process-aware approach that first separates precipitation events into types and then constructs an overall mixture from the resulting type-specific distributions. Using 10-min precipitation records from two climatically contrasting Austrian weather stations (Graz Universität and Dornbirn, 2010–2022), we compare partitioning (k-means) with outcome-guided, model-based clustering implemented as finite mixtures of Gamma regressions (clusterwise distributional regression). Events are characterised by severity, magnitude, duration, intensity, time-to-peak, and lightning occurrence. While partitioning yields interpretable event groups, the outcome-guided approach provides substantially better distributional fit–reducing Kullback-Leibler divergence by 66–73 % relative to a single-Gamma baseline–and improves visual diagnostics. Single-distribution baselines underrepresent process heterogeneity in precipitation; process-aware clusterwise mixtures that account for event-level structure address this gap. The framework is distribution-agnostic and portable, relying on information available at the event scale and lightning data, and can support applications such as hydrological modelling and stochastic weather generation.
This paper proposes a novel, process-aware framework for modeling sub-hourly precipitation distributions (e.g., for stochastic weather generation) by explicitly accounting for the underlying meteorological process heterogeneity rather than relying solely on increasingly complex single-distribution models, though it can potentially be combined with them. The proposed methodology (supplemented by open-source R code) includes the identification and characterization of precipitation events from 10-minute records using severity, magnitude, duration, mean intensity, time-to-peak and lightning occurrence as its features, with lightning serving as a proxy for convection. Two distinct clustering approaches are compared within the study’s framework, the traditional unsupervised partitioning (k-means) and an outcome-guided, clusterwise distributional regression method. Gamma models are fitted both to all non-zero values surpassing a selected wet-interval magnitude threshold and cluster-wise, subsequently aggregating the cluster-specific fits into weighted mixtures. The results show that the proposed framework offers a promising alternative to traditional precipitation modelling, with outcome-guided, process-aware clusterwise mixtures reducing Kullback-Leibler divergence by 66% to 73% relative to a standard single-Gamma baseline across two conducted case studies, outperforming also the k-means alternative.
I believe that the paper represents an important methodological advancement, making a meaningful step forward from the purely descriptive use of clustering for sub-hourly precipitation events. I have the following minor comments, which mainly concern the presentation of the work already conducted.
Minor comments