the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
Evaluation and application of a convolutional neural network for graupel identification in DCMEX deep convective cloud
Abstract. Untangling the ice microphysical interactions within deep convective clouds presents an ongoing issue. Cumulonimbus have implications for localised precipitation and global radiative feedbacks. In situ flight campaign data is informative of these interactions and can consolidate our understanding. Identifying ice particle habits illustrates the evolving cloud on the micro-scale. In particular, the development and growth of graupel continues to be the least understood hydrometeor in numerical models. Consequently, in the ever-evolving machine learning landscape, a multitude of instrument and dataset specific ice habit identification algorithms are becoming commonplace. Here, we complete a key step of independently assessing several generalised and open source algorithms, to better understand their suitability for wider uptake. Evaluation and application of generalised convolutional neural networks (CNN), created by Jaffeux et al. (2025) has been undertaken on unseen two-dimensional stereo (2D-S) and High Volume Particle Spectrometer (HVPS) images from the Deep Convective Microphysics EXperiment (DCMEX). Models were not re-tuned to the dataset. Jaffeux et al.'s global CNN tested with human labelled 2D-S images obtained an accuracy of 72 % and F1 score (harmonic mean of precision and recall) of 70 %. While for HVPS images, the HVPS-specific CNN had an accuracy of 86 % and F1 score of 73 %, which was only marginally better than the global model. Then scaling up CNN application to the whole DCMEX dataset, graupel concentrations were inferred from rimed particle classification. The models constructed by Jaffeux et al. (2025) present an accessible, accurate and adjustable approach for particle identification of optical array probe images.
- Preprint
(1115 KB) - Metadata XML
- BibTeX
- EndNote
Status: open (until 30 Aug 2026)
- RC1: 'Comment on egusphere-2026-3385', Anonymous Referee #1, 24 Jul 2026 reply
-
RC2: 'Referee comment #2 for egusphere-2026-3385', Anonymous Referee #2, 17 Aug 2026
reply
Note: when I tried to upload my comments as a file, there was a bug in which my file was replaced by another referee's comments upon hitting save, so I am posting everything in the text box.
Review of “Evaluation and application of a convolutional neural network for graupel identification in DCMEX deep convective cloud” by Ezri E. Alkilani-Brown, Declan L. Finney, Alan M. Blyth, Paul R. Field, Jonathan Crosier, and Chetan R. Deva
Summary: This work employs convolutional neural networks (CNNs) for ice crystal habit classification, described by Jaffuex et al. (2025), to label in-situ ice crystal images collected from the Two-Dimensional Stereo probe (2D-S) and the High-Resolution Precipitation Spectrometer (HVPS) during the 2022 Deep Convective Microphysics EXperiement (DCMEX). The performances of three CNNs are evaluated: a universal neural network applicable to many optical array probes, and probe-specific networks for the 2D-S and HVPS. Sets of randomly-selected, manually-classified images are used for test sets. About 10 percent of the test images for each probe are manually classified by more than one person, and used to evaluate the uncertainty of manual classification. Performance is quantified using accuracy, precision, recall, and F1 score metrics, with emphasis given to the successful identification of graupel (rimed) particles. Finally, the universal CNN is used for 2D-S images, and the specific CNN for HVPS images, to classify (nearly) all crystals imaged during the DCMEX campaign and quantify the proportions of crystal habits observed in convective clouds.
This study works to fill a gap in the literature, namely in that although numerous algorithms for ice crystal classification have been developed and tested on sets of idealized crystal images, few have been broadly applied to the scale of an entire field campaign. The substantial uncertainty in habit that is seen even between manual classifications highlights the issue that a large portion of crystals collected during field campaigns are ambiguous in habit (and also that the class definitions can be ambiguous), at least when viewed at the resolution of an optical array probe. Likewise, the selection of random crystals for testing, rather than representative crystals, results in a notable decline in performance for the neural networks compared to the studies in which they were developed. Nevertheless, in the context of DCMEX, the neural networks are shown to be useful for distinguishing graupel particles from other particles, a distinction that is important for microphysical understanding.
Recommendation: The subject matter discussed in this manuscript represents a valuable contribution to the literature, in outlining the utility, but also the limitations, of applying CNNs for habit classification across the scale of an entire field campaign. However, it is difficult to determine the main takeaway from certain portions of the manuscript, such that I think reorganization, or more concise writing, is necessary. Additionally, although limitations are discussed, the skewing of the random test dataset toward compact particles is not ideal for evaluation. Given that I think significant changes are required regarding the methodology and organization of the paper, I recommend major revisions.
General comments:
Random sampling of the test set. Random sampling provides a valuable test set for how a CNN will perform for images from this particular field campaign, but likely does not produce results that are broadly applicable to cases with substantially different habit distributions. Ideally, I’d suggest randomly sampling and classifying images until ideally 100 images have been acquired for each class, matching Jaffeux et al. 2025. To more quickly obtain examples of less common habits, I’d suggest randomly sampling from specific time periods when these habits are more frequent, if any such periods are known to exist. Then, to balance the test set, take a subset of 100 images for the overpopulated classes. This could be done with the human-to-human comparison as well, if an additional person could be found to classify many crystals (although this may be prohibitively difficult to do, based on personal experience classifying crystals and the time commitment involved).
As a caveat, the low sample sizes in the current test set for CC in the 2DS, and CBC/HBC for the HVPS suggest to me that there aren’t 100 samples even in the entire dataset, so it would make sense to omit those classes, except for the purpose of incorrect classifications by the CNN. While this balanced evaluation could be covered alongside the current unbalanced one, I think it would be reasonable to present only the balanced evaluation in figures for conciseness. Instead, the human-identified fractional habit distribution could be presented as part of Section 3.3 (DCMEX application), and compared to the CNN results.
Organization. This work appears to have two competing objectives, being the evaluation of CNNs, and the use of CNNs to identify habit in DCMEX. I think that these two objectives should be more cleanly separated for clarity, with each objective having its own dedicated section and associated discussion. By keeping everything together, the manuscript can also be written more concisely. To accomplish this, I’d suggest the following reorganization, although this is far from the only way to do so:
Merge the context of DCMEX and the 2DS/HVPS probes in section 2.1 into the portion of the introduction discussing the same topics.
Title Section 2 something along the lines of “CNN Evaluation”, and include what is currently section 2.2, and sections 3.1 and 3.2. Merge the discussion in what is currently section 4.1 into what is currently section 3.1, and the discussion in what is currently section 4.2 into what is currently section 3.2. Relabel subsections as needed.
Title Section 3 something along the lines of “Graupel Identification in DCMEX” and include what is currently section 2.2.5, everything currently in section 3.3. Merge the discussion in what is currently section 4.3 into the new section 4. Relabel subsections as needed.
The conclusion then becomes Section 4.
I’d prioritize manuscript organization over the above major change, as I think the major findings of this work are an important contribution to the literature even as is, and simply need to be communicated more clearly.
Grammar. The grammatical structure could be improved. While I will not specifically point out everything, I will provide a couple examples. For one, at lines 219-221, the second sentence is a sentence fragment, and should be combined into the first sentence: “However, considering the weighted test scores across all and sub-setted habits, the global CNN performed almost universally better than the 2D-S-specific CNN. Apart from a very slightly higher weighted precision score.” The paragraph in which the above example is found then concludes at line 229 with another sentence fragment, “Respectively a higher weighted precision, recall and F1 scores of 88.60%, 88.30% and 88.28%,” and abruptly ends the section with a statistic without any comment on the importance of the statistic. Also note that the latter issue will be resolved by merging section 3.2 with the discussion in section 4.2, as suggested.
Time series. This is undoubtably a personal preference, but it would be interesting to see time series plots of habit fractions for a couple of flights from the field campaign. I find that time series plots provide a useful visualization of the microphysical variability within clouds that cannot be expressed using cloud-mean quantities.
Quantifying model skill. In a scenario where the total numbers of crystals in each category are the same as in the test set, the F1 scores achieved by the random test in Appendix A provide a quantitative benchmark for evaluating the CNNs. If the CNN cannot exceed the F1 score from random classification, this indicates that the CNN has no skill for that class.
Specific comments:
Line 1: Change to “cumulonimbus is the largest convectively formed cloud, with serious impacts on climate and weather”.
Line 24: Should be “complexities, e.g., secondary ice”. Note that there are other places in the manuscript where commas are missing before and after “e.g.” or “i.e.,” but I will not list them here for brevity.
Line 26: Commenting that “Importantly, the development lifecycle of graupel within these clouds has not been well characterized” gives me the expectation that this issue will be discussed in the manuscript. However, such a characterization is noted to be beyond the scope of this work. While I agree that an in-depth analysis of graupel lifecycle is too much to include here, I think it is important to explain how using a CNN to identify graupel particles will be useful for understanding graupel lifecycle.
Also, the following sentence beginning with “Including the…” is a sentence fragment and should be combined into the quoted sentence.
Paragraph lines 42-55: I’m not sure if the discussion of hail is needed, given that hail is not mentioned later in the manuscript. Further, near the end of the paragraph, it is noted that “a quantitative threshold for when graupel has been created has not been outlined. Instead, relying on a qualitative description…” which makes me expect that such a threshold will be discussed in this work. However, this study also uses a qualitative description, which leads to uncertainty in graupel classification among different human classifiers, as noted later in this work.
For the above reasons, I think this paragraph should be removed, and the previous paragraph should conclude with something along the lines of “Habits are often organized only according to the variance in depositional growth patterns with temperature and supersaturation. However, crystals that undergo aggregation and/or accretion will exhibit habits that cannot be predicted from environmental conditions alone, including limitless aggregating crystal monomers, graupel, and hail. As a result, knowledge of the ambient atmospheric conditions is insufficient to determine crystal habits. Instead, the only surefire way to determine habits is to directly view the crystals through in-situ observations.”
Line 69: It may be worth noting that classification of habit from geometric features has also been done without principal component analysis, as described by Holroyd (1987), and applied generally to field campaign data by Schima et al. (2024). These references are provided below.
Holroyd, E. W. (1987). Some techniques and uses of 2D-C habit classification software for snow particles. Journal of Atmospheric and Oceanic Technology, 4(3), 498–511. https://doi.org/10.1175/1520-0426(1987)004<0498:stauoc>2.0.co;2
Schima, J., McFarquhar, G. M., Delene, D., Heymsfield, A., Bansemer, A., Schnaiter, M., et al. (2024). A multi-probe automated classification of ice crystal habits during the IMPACTS campaign. Journal of Geophysical Research: Atmospheres, 129, e2024JD040895. https://doi.org/10.1029/2024JD040895.
Paragraphs lines 76-90: This is the part of the introduction that really highlights the motivation of this study, which is to evaluate the performance of the Jaffeux et al. 2025 algorithms using untested field campaign data, and then use the classifications to describe ice crystal habits in clouds. I think the importance of this work should be more clearly emphasized. As is noted, virtually all CNNs have been developed on hyper-specific, idealized training images, but the work presented here is a critical test of the ability of the algorithms to adapt to the ambiguity of crystal habits in real clouds.
Lines 111-112: I’ve typically seen these referred to as the horizontal and vertical channels, and I find that labeling them that way helps for my own visualization.
Lines 134-135: It’s probably worth noting that these diameter thresholds were required due to probe resolution constraints. The limitations of these relatively large thresholds should be discussed, as this leaves gaps in the classifications for particles smaller than 300 microns and 1.2-3 microns, which probably constitute a large portion of total crystal mass.
Lines 145: I assume that columns with an aspect ratio greater than 10 were rare enough such that their omission makes negligible difference, but I’d like to see a comment on this, given that the habit statistics could be skewed if they are more frequent.
Lines 149-150: The use of multi-label consensus for some labels, but the label from a single human for others, is a methodological inconsistency. For consistency, I’d recommend sticking to labels from a single person when evaluating the CNNs, given that a large majority of the images were classified by one person. The comparison of human labelers in Section 3.1 effectively quantifies the uncertainty in using classifications performed by one person and thus remains a valuable analysis. In doing this, the need to describe the concept of a group consensus (lines 156-165) would be eliminated, allowing the manuscript to be shortened.
Figure 1: I’m not entirely clear how the pair-wise comparison is performed. How are pairs chosen? My interpretation is that the results could vary depending on which pairs of humans are selected. Please elaborate on the methodology.
Figure 2 and associated discussion: Based on the random consensus results, the accuracy metric is meaningless for evaluation purposes. I’d remove accuracy from all figures, remove accuracy statistics from the text, and give a sentence or two of discussion about why this was done.
Line 220: The better performance of the global CNN for the 2D-S is a surprising and important result. I’d add in the manuscript that this indicates that the specific model was likely overfit to the training data.
Line 233: Include the note that the global CNN doesn’t have an RA class here in the text as justification for merging the classes, rather than, or in addition to, within the description of Figure 4.
As a consideration, it may also make sense to merge the RA and CP classes within the specific HVPS algorithm here, to provide a fair comparison between the general and specific algorithms. As further justification, the classes are already combined for the campaign-wide analysis.
Line 251: For more justification to choose the specific model for the HVPS in this study, it appears to have markedly better performance for the CP class, which comprises the majority of crystals from the campaign.
Lines 263-268: I find the description of the correction factor hard to follow. My interpretation is that the correction factor is equivalent to the fraction of human-labeled CP images classified as HPC in testing, which from Figure 2 I calculated is 18.7%, divided by the total fraction of crystals classified as HPC over one second from the field campaign data. If my interpretation is not correct, it may be good to reword the description.
Additionally, the authors’ fixed correction factor and the one from Jaffeux et al. (2025) are drastically different – why is this the case?
Lines 290-295: Explain why we might expect to see these differences in composition between the probes. I’m not surprised that the HVPS, which images larger crystals, has more classified aggregates, as the aggregation process more commonly occurs with larger crystals, and increases crystal size.
Note: I see that this is explained in the paragraph from lines 430-438. This is a place in which merging the sections would add clarity.
Lines 297-298: What does “outside the scope of CNN classification” mean? Is there a specific minimum threshold concentration needed for the fraction of a given particle type to be considered nonzero?
Figure 6 description: The phrase “Legend percentages display (number) proportions of habit concentration to the DCMEX total ice population…” reads as if these are proportions of the total particle count from the 2D-S and HVPS combined. However, I see that the labeled percentages for the ch0, ch1, and HVPS data are proportions of the ch0, ch1, and HVPS particle counts, respectively. Please clarify.
Also, do the errorbars refer to the standard deviation of the data among all 1-s time periods, or something else?
Tables 1 and 2: Some of the sample sizes are very small. Be cautious about making conclusions from any flights with low sample sizes. I generally follow the rule of thumb that a minimum sample size of 30 is needed to draw any statistical conclusions.
Lines 316-317: It is noted that some pair-wise combinations were significantly higher than random when comparing the two. However, the distribution of habits is even in the random assignment case, whereas CP is heavily favored by both human classifiers in the human comparison, so this comparison is misleading. To create a more fair comparison, the set of labels given by one of the human classifiers could be randomly scrambled to create a random assignment case, ensuring a similar distribution of habits. Ideally, the human comparison would be roughly balanced to begin with, but if this cannot be done, the next best thing is to adjust the random classifications.
Lines 323-325: Did all the human classifiers look over the example images from Jaffeux et al. prior to their classifications? If not, I would expect better agreement if these images were analyzed ahead of time.
Lines 343-344: I wonder if there is a way to quantify this uncertainty from the human comparison data. Would it be possible to construct a margin of uncertainty for the precision, accuracy, and F1 score, knowing that the classifications may be wrong? One idea would be to perform a series of random resamples of the class labels according to the confusion matrices presented in Figure 1, and find the standard deviations of the scores across all resamples.
Figure 7: I like this visualization; it effectively showcases the ambiguity in the habit of real crystals. I’d argue that some images are inherently unclassifiable; at the resolution of the probe, and without any grayscale, sometimes two habits that are distinct in reality can produce the exact same image on an OAP. As a comparison, it may be informative to also show a few cases where both the human and CNN classify a crystal as CP.
Lines 385-389: As noted previously, it would make sense to just exclude the class for small sample size instead of taking the time to explain.
Line 390: “Despite not attaining perfect performance scores” is redundant given its mention later in the paragraph.
Lines 423-426: Figure 6 is very separated from this part of the text; this is one reason why I’ve suggested merging this section with Section 3.3. Also, while I agree with the approach of combining the FA and CP classes, I’m not convinced that there is enough evidence to confidently claim that it is “the correct approach.”
Lines 444-445: For classes with extremely low precision, such as CBC, the size distributions will look like those from other classes, as the overwhelming majority of particles placed in that class will be incorrect classifications. For these classes, I would argue we cannot infer anything at all from the size distribution plots. The most robust result is by far the CP size distribution.
Lines 449-451: Theoretically in-cloud sampling time does relate to the number of particles observed, but this effect is overcome by variations in number concentration and size between cases.
Lines 461-463: While there were issues with CP categorization, I think that this statement is misleading, as it implies classifications for other classes were relatively accurate. Instead, both CNNs performed best for CP crystals per the F1 metric, which should be emphasized here, while classes such as Co were much more poorly represented.
Line 478: I don’t think the images from the 2D-S and HVPS provide sufficient information to determine a degree of riming with any accuracy. Higher-resolution probes with grayscale imaging, like the PHIPS and CDP, are needed.
Context on PHIPS probe if curious: Abdelmonem, A., Järvinen, E., Duft, D., Hirst, E., Vogt, S., Leisner, T., and Schnaiter, M.: PHIPS–HALO: the airborne Particle Habit Imaging and Polar Scattering probe – Part 1: Design and operation, Atmos. Meas. Tech., 9, 3131–3144, https://doi.org/10.5194/amt-9-3131-2016, 2016.
Appendix C: I do think that the accuracy calculation should be kept here, even with my suggestion of removing accuracy from the rest of the manuscript, since it is a frequently used metric.
Citation: https://doi.org/10.5194/egusphere-2026-3385-RC2
Data sets
DCMEX 2D-S and HVPS labelled images for CNN evaluation Ezri Alkilani-Brown, Declan Finney, Alan Blyth, and Paul Field https://doi.org/10.5281/zenodo.20612982
Viewed
| HTML | XML | Total | BibTeX | EndNote | |
|---|---|---|---|---|---|
| 163 | 50 | 18 | 231 | 21 | 21 |
- HTML: 163
- PDF: 50
- XML: 18
- Total: 231
- BibTeX: 21
- EndNote: 21
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
Please see the review attached to this comment.