the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
Deep-learning-based stage-aware classification of Arctic melt ponds and their microtopographic associations from high-resolution UAV imagery
Abstract. Melt ponds are a key component of the summer Arctic sea-ice surface because their formation and evolution strongly affect surface albedo, energy absorption, meltwater redistribution, and sea-ice mass balance. Most previous remote-sensing studies have treated melt ponds as a single surface class, limiting the characterization of heterogeneous pond states during late-summer melt and refreezing. In this study, we developed MP-Unet, a stage-aware semantic segmentation framework for identifying Open, Transitional, and Frozen Melt Ponds from high-resolution unmanned aerial vehicle imagery acquired during the 14th Chinese National Arctic Research Expedition. MP-Unet integrates residual blocks with channel attention, atrous spatial pyramid pooling, attention-gated skip connections, and an auxiliary binary segmentation head. The full model achieved an F1-score of 0.9440 and a mean intersection over union of 0.7466, with class-specific IoU values of 0.5897, 0.7544, and 0.6539 for Open, Transitional, and Frozen Melt Ponds, respectively. Stage-resolved mapping revealed marked spatial heterogeneity among the five observation sites, while Transitional Melt Ponds accounted for approximately 80.3 % of the total pond area in the pooled sample. Pond area–frequency distributions showed a general scale-dependent decline and a sparse large-area tail, although the strength of the fitted scaling relationship varied among sites. Object-level analysis further showed that Frozen Melt Ponds generally had more compact and regular shapes, whereas Open Melt Ponds exhibited broader circularity distributions extending toward lower values. DEM-assisted analysis indicated significant stage-dependent differences in local relative elevation: Transitional Melt Ponds occupied lower local topographic positions than Frozen Melt Ponds, despite the absence of significant differences in distance to the nearest ridge-like feature. These findings demonstrate that stage-aware classification provides information beyond conventional binary melt pond mapping by linking surface-state identification with pond morphology and local microtopographic position. The proposed framework offers a practical basis for fine-scale observations of Arctic sea-ice surface evolution and for the validation and improvement of satellite and numerical melt pond products.
- Preprint
(5478 KB) - Metadata XML
- BibTeX
- EndNote
Status: open (until 21 Sep 2026)
- RC1: 'Comment on egusphere-2026-4430', Anonymous Referee #1, 06 Sep 2026 reply
-
RC2: 'Comment on egusphere-2026-4430', Anonymous Referee #2, 08 Sep 2026
reply
This manuscript proposes a new deep-learning DL method to classify different types of melt ponds and interpret their microtopographic characteristics. I think the description of the DL method is generally clear, and the analyses of the classification results (Sections 4.1–4.3) are well presented. These aspects make the manuscript more suitable for a remote-sensing methodology journal. However, I would recommend a major revision to strengthen the physical interpretation, particularly the topographic analysis in Section 4.4, in order to better fit the scope of TC. The overall quality of the manuscript, including its organization, figure layout, and writing, should also be thoroughly improved.
Major concerns
First, the authors should clearly explain the physical meaning of Open, Transitional, and Frozen Melt Ponds. RGB imagery can certainly be used to classify melt ponds into open, frozen, and transitional stages, but what is the scientific significance of distinguishing among these three classes? What geophysical information can be obtained from this classification?
Without a clear explanation of the physical significance of these classes, the implementation of a U-Net architecture combined with specific encoders and decoders to improve classification accuracy appears to be primarily a methodological contribution and may therefore be more suitable for a remote-sensing methodology journal such as IEEE JSTARS.
Another concern is about the scientific significance of the improvement in classification accuracy. Table 3 demonstrates an improvement in classification accuracy using the proposed DL method. However, what does an improvement of approximately 0.01 in accuracy mean in practice? Is a more complicated DL architecture really necessary compared with a conventional U-Net?
I suggest that the authors examine the prediction maps in detail and identify which pixels, features, or melt-pond conditions are successfully classified by the proposed method but misclassified by the conventional U-Net. The authors should then demonstrate whether these improvements are important for the subsequent physical interpretation. Simply reporting a numerical improvement in accuracy is insufficient if the difference does not have meaningful consequences for the scientific analysis.
I am not sure about the Microtopographic analysis and physical interpretation in the manuscript. Seen from the title, a highlight of this manuscript seems to be the analysis of microtopography retrieved from the UAV imagery. However, there is currently no one figure showing what the retrieved topography actually looks like. Figure 8, for example, could potentially include a DEM or another direct visualization of the surface topography?
I think the microtopographic interpretation (Section 4.4) should be substantially rewritten. In its current form, I do not think this section meets the expected scientific scope and standard. First, the term “microtopography” should be clearly defined in the context of this study. The authors should explain what spatial scales and topographic features are represented by this term. Second, the physical interpretation should be scientifically rigorous and supported by appropriate references:
For example, Lines 564–566 state:
“Transitional Melt Ponds occupied significantly lower local topographic positions than Frozen Melt Ponds, whereas Open Melt Ponds showed intermediate characteristics.”
I find it difficult to understand the physical meaning and implications of this statement. The authors should explain what these differences in topographic position represent physically and why the different melt-pond stages would be expected to occupy these positions.
Similarly, the manuscript states:“Transitional Melt Ponds exhibited a weak but significant positive relationship.”
What exactly is meant by “weak but significant”?
Lines 574–576 state:“Although the interaction between ridge distance and pond stage did not reach conventional statistical significance, the observed stage-specific trends consistently indicate that Transitional Melt Ponds are more sensitive to ridge-related microtopographic variation than the other two stages.”
If the statistical analysis does not support a significant interaction, the authors should be very cautious about making a definitive interpretation based on the apparent trends. This interpretation should either be supported by additional evidence or substantially qualified.
Other points:
A number of statements, methods, and equations require additional references or explanations:
Line 84: When stating that the study area is approximately 90% covered by sea ice, what reference or dataset was used to determine this value? Please provide the relevant data source and/or reference.
Line 198: Please explain these features more clearly and provide equations and appropriate references where necessary.
Eqs 3–18: If these equations were not developed specifically in this study, appropriate references to their original or established sources should be provided.
Eqs 15–18: I may have misunderstood this part, but are all of these topographic features subsequently analyzed and interpreted in the Results and Discussion sections? The connection between the defined features and the subsequent analyses should be made clear.
All figures and figure captions should be carefully revised. Figure captions should be sufficiently self-contained so that readers can understand the main content of each figure without having to search extensively through the main text.
Section 2 is entitled “Research Data,” but it contains subsections describing the study area, data, and processing methods. The structure and section titles should therefore be reconsidered and better organized.
Figure 2: The current figure appears more suitable for a slide or presentation than for a scientific journal article. It should be redesigned to meet journal publication standards.
Figure 3: Please plot the imagery using geographic or projected coordinates. A scale bar and north arrow should also be provided so that readers can understand the spatial context.
Figure 6: Why are the images in the first column displayed in grayscale? Should they not be displayed as RGB images?
Line 208: Why were only 70 images selected for the analysis? According to Table 1, hundreds of images appear to be available. The authors should also provide information about the geographic locations of the selected images and the associated sea-ice conditions. For example, were the images acquired from different types of ice floes or different ice conditions? This information is important for assessing whether the selected samples are representative of the broader dataset and whether the classification results can be generalized.
An abbreviation only needs to be defined at its first occurrence. After that, the abbreviation should be used consistently throughout the manuscript. For example, “melt pond fraction (MPF)” should be defined once and subsequently referred to simply as “MPF.”
Citation: https://doi.org/10.5194/egusphere-2026-4430-RC2
Viewed
| HTML | XML | Total | BibTeX | EndNote | |
|---|---|---|---|---|---|
| 143 | 58 | 21 | 222 | 19 | 19 |
- HTML: 143
- PDF: 58
- XML: 21
- Total: 222
- BibTeX: 19
- EndNote: 19
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
The manuscript presents an interesting high-resolution UAV dataset and a potentially useful machine-learning approach for melt-pond classification. However, I have serious concerns about the physical basis and interpretation of the proposed Open, Transitional, and Frozen Melt Pond classes.
My main concern is that the study proceeds mainly from visually defined RGB (but actually optical) classes to machine-learning classification, physical interpretation to fit the results, rather than establishing the physical states first and then testing whether they can be identified from UAV imagery. Consequently, the manuscript demonstrates that the proposed visual classes can be distinguished, yet does not show that they represent distinct developmental stages of melt ponds.
The manuscript also lacks a sufficiently clear scientific thread connecting machine-learning development to real optical classification, pond morphology, DEM analysis etc... I think the manuscript requires substantial restructuring and clarification of whether its main contribution is a new image-classification method or a physical investigation of melt-pond evolution.
Please see further comments in the supplements.