the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
Observing Seawater Flooding on Arctic First-Year Ice during Pre-Melt Using C- and L-Band Synthetic Aperture Radar
Abstract. Seawater flooding from negative freeboard onto sea ice leads to the formation of a saline slush layer at the bottom of the snowpack, affecting snowpack geophysical properties, ice mass balance, remote sensing retrievals, and on-ice travel safety. Synthetic Aperture Radar (SAR) is sensitive to saline slush, though limited information on this dynamic layer exists. This research assesses the sensitivity of C- and L-band frequency SAR to slush occurrence on Arctic first-year ice near Qikiqtarjuaq (Nunavut, Canada) during pre-melt conditions and evaluates the potential of SAR to estimate flooding. SAR parameters backscatter, logarithm ratio, cross- and co-polarization ratios, coherence, and textural features, are evaluated for their ability to distinguish Flooded from Non-flooded classes and results are tested at a regional scale. The loss of coherence in co-polarization L-band SAR, due to the occurrence of a high-dielectric slush layer, is found to be optimal for flooding detection, with the lower frequency being less influenced by the overlying snow. Backscatter, and its temporal change through the logarithm ratio, consistently discriminates classes when the slush is refreezing or has frozen into snow ice. Regional-scale estimation using L-band coherence is demonstrated as achieving an F1-score of 0.81, with higher F1-scores (up to 0.87) obtained when random forest models combine textural information with coherence. These findings highlight the dynamic nature of seawater flooding, and the potential for SAR to be used in related applications and research studies.
- Preprint
(2916 KB) - Metadata XML
-
Supplement
(2577 KB) - BibTeX
- EndNote
Status: final response (author comments only)
- RC1: 'Comment on egusphere-2026-4311', Anonymous Referee #1, 21 Aug 2026
-
RC2: 'Comment on egusphere-2026-4311', Anonymous Referee #2, 21 Aug 2026
This manuscript presents coincident C- and L-band SAR observations of seawater flooding and slush on landfast first-year ice near Qikiqtarjuaq, supported by in-situ transects, snow pits, and ice thickness measurements collected within days of the satellite overpasses. Datasets combining these frequencies with coincident slush measurements are rare, and this one has clear value beyond the present analysis. The use of repeat-pass L-band coherence as a flooding indicator is, as far as I am aware, new in the sea ice context, and it is a sensible response to the earlier finding that backscatter intensity alone was not sufficient for this problem. The co-design with the community management committee, and the inclusion of a local ice monitor as a co-author, meaningfully improved the ground reference for a feature that is invisible from above and difficult to sample blind.
I have concerns about the link between the evidence presented and the conclusions drawn (major comments), and I think these need to be addressed before the paper is accepted.
Major comments (#):
#1
The central result is that L-band coherence is lower over flooded areas, attributed to a high-dielectric slush layer forming a scattering interface at a new height (Sect. 4.1, Fig. 12). The proposed mechanism is physically reasonable, but I would like to see the alternative explanation addressed more fully than the brief acknowledgement at lines 514–516.
Areas that flood are areas of low or negative freeboard, that is, comparatively thin ice. Thin ice near a landfast boundary is also more susceptible to tidal flexure and small-scale differential motion, both of which reduce coherence over a 16–24 day interval independently of any change in dielectric properties. The observation at lines 486–488 that the flooded zones were the first to open on 9 July is, on its own, consistent with either explanation.
Would the authors consider examining the interferometric phase in addition to the coherence magnitude? If the phase over the flooded zones is spatially flat while the coherence magnitude drops, the dielectric interpretation is considerably strengthened. If instead there are phase ramps or fringe patterns over those zones, that would point toward a motion contribution. This seems to me the most direct discriminator available in the data already in hand, and I suspect it would take relatively little additional processing.
A complementary check would be to stratify the coherence comparison by measured ice thickness, to see whether flooded and non-flooded samples of comparable thickness still separate. If neither is conclusive, I would simply ask that the abstract and conclusions state that the two mechanisms co-vary in this dataset and were not separated.
A smaller point on the same mechanism. A uniform layer at a new height principally changes the interferometric phase rather than causing decorrelation; loss of coherence requires a change in the arrangement of scatterers within the resolution cell. The authors document exactly the relevant condition, namely that slush is patchy at scales below the resolution cell (lines 171–175). Framing the mechanism around sub-resolution patchiness and its evolution, rather than around a uniform interface shift, would make the argument in Sect. 4.1 more robust and is better supported by the authors' own field observations.
#2
Given that coherence carries the main result, could the authors add the following to Sect. 2.4?
I have questions about the coherence estimation window and the effective number of looks. The reported class means fall roughly in the 0.25–0.45 range, where the bias of the sample coherence magnitude is not negligible, so the reader needs these values to judge whether the class separation exceeds the estimator's own uncertainty. A short statement of the expected bias at the reported window size would settle this.
I am curious whether topographic or flat-earth phase was removed, and the co-registration approach. The critical baseline and geometric decorrelation contribution for each pair. I raise this partly because I expect it to help the authors' case: since the critical baseline scales with wavelength, the 480 m perpendicular baseline of orbit 19b is less penalizing at L-band than the comparable figure would be at C-band. Making that explicit would support the temporal interpretation at lines 508–513.
On lines 508–513, the comparison between orbit 19a (16 days, 153 m) and 19b (24 days, 480 m) is used to infer a temporal-baseline dependence, but both baselines differ between the two pairs and there are only two pairs. I would suggest softening this, or supporting it with the geometric decorrelation calculation above.
#3
Section 2.6 describes IDW interpolation of the field slush measurements, generation of 100 random points within each interpolated area, thinning to a 30 m minimum separation, and a 70/30 split stratified by class.
Could the authors confirm whether the split was random or spatially blocked? If it was random, then training and evaluation points are drawn from the same interpolated surfaces at separations well below the scale of the SAR features themselves (the GLCM kernels are 100–135 m per lines 269–270), and the reported ROC-AUC and F1 values would represent something closer to in-sample than out-of-sample performance. With four field sites, a leave-one-site-out scheme would be straightforward and would give a figure that transfers. I would expect the resulting scores to be lower, and I want to say clearly that this would not weaken the paper. A defensible number is more useful to the community than a high one.
Two related requests:
The class labels are themselves interpolated, with reported MAE of 2.2 cm and RMSE of 3.5 cm (lines 314–315), against a class boundary at 5 cm. Could the authors report what fraction of the extended points lie within one interpolation RMSE of the threshold, and show how the results change if those are excluded?
Feature selection by KS statistic and model evaluation currently use the same data. Selecting features on a subset of sites and evaluating on the held-out site would remove this.
#4
The abstract states that L-band coherence is optimal for this application and that the lower frequency is less influenced by the overlying snow (lines 22–25). I do not think either statement is separable from the acquisition configuration in this dataset.
The L-band observations come from a single mission, in one mode, at a single incidence angle and pixel spacing, in quad-pol. The C-band observations span two missions, four modes, incidence angles from 20 to 55 degrees, and dual-pol only. More importantly, C-band coherence is available from one interferometric pair. I do not think a conclusion about C-band coherence for flooding detection can rest on a single pair, and I would suggest restating this as the observation that the one C-band pair available here did not separate the classes, together with a short list of the confounding differences.
The snow influence statement is not tested anywhere in the manuscript. In cold, dry, low-salinity snow, C-band is also close to transparent, so the frequency contrast under these conditions may be smaller than implied. If the authors wish to keep an argument about snow, the salinity profiles they measured in the pits (lines 121–123) are the natural evidence, since a saline basal layer would raise the effective C-band scattering horizon while affecting L-band less. Those measurements are collected but not used anywhere in the analysis, which seems to me a missed opportunity.
The same issue affects Fig. 9a. The mean KS statistic is computed over different orbit subsets for different parameters: coherence exists for three orbits, two of which are the L-band pairs acquired on the latest and most strongly flooded dates, while sigma nought is averaged across all twelve orbits including earlier dates and wide-swath scenes at high incidence. Because this ranking underpins the headline sensitivity result, I would ask that it be restricted to a common subset of orbits and dates, or presented per-orbit only.
#5
Flooded is defined as slush thickness of at least 5 cm and Non-flooded as zero, with the 0–5 cm interval excluded (lines 285–287). From Table 1, mean slush thickness is 2.0, 3.5, 10.1, and 1.0 cm at Sites 1 to 4, so only Site 3 has a mean above the threshold.
Could the authors report the number of Flooded and Non-flooded samples per site, per orbit? If the Flooded class is drawn predominantly from one site, then the reported separations are partly a between-site contrast, and sites differ in snow depth, ice thickness, and date of visit. It would be reassuring to see whether the separation also holds within sites.
Separately, the Table 1 caption states that slush thickness includes the surface snow ice layer where present. This places liquid slush and refrozen snow ice in the same class, whereas Sects. 3.1 and 4.2 argue that the two states differ substantially in their backscatter and log-ratio response. Could the class definition be made consistent with that argument, either by separating the two states or by reporting a sensitivity test with snow-ice records excluded?
#6
Table 2 indicates that between 61 and 96 percent of L-band HV and VH samples fall below the NESZ, and 30 to 37 percent for two of the Sentinel-1 wide-swath orbits. Lines 355–356 then conclude that L-band showed almost no difference between classes in the cross-polarized channels. I would suggest revising this to state that these channels are noise-dominated and therefore uninformative here, since a null result cannot be established from data below the noise floor. Moving the affected panels to the supplement with an explicit note would be one clean way to handle it.
#7
The KS tests are applied to samples spaced 25 m apart, using features derived from kernels of 100–135 m, so the effective sample size is substantially smaller than the nominal one and the associated p-values are difficult to interpret at face value. The tests are also applied across 138 features without correction for multiple comparisons.
The D statistic itself is a perfectly reasonable effect size and I would keep it and the associated figures. My suggestion is simply to present D as an effect size, and to caveat or remove the significance flagging (including the −1.00 sentinel in Fig. 9) rather than to redo the analysis.
#8
The paper is motivated by ice travel safety, which I think is the right motivation and one of its strengths. The random forest is trained on a balanced set, and ROC-AUC is correctly noted to be insensitive to class balance. However, along a given travel corridor the prevalence of thick slush is likely to be low, and at low prevalence precision degrades even for models with good ROC-AUC. The two error types also carry very different costs for someone on the ice.
Because the intended users would rely on this, it would be valuable to report precision and recall at a plausible operating prevalence, and to state the expected false-alarm and missed-detection behaviour. I raise this to make the product more dependable in use rather than to question the application.
Minor comments:
Lines 22–28 (Abstract) "is found to be optimal" reads more strongly than the evidence supports, given two L-band pairs at one location in one season. Similarly, "due to the occurrence of a high-dielectric slush layer" states a causal mechanism that the study associates but, as discussed above, does not isolate. Some qualification would bring the abstract into line with the body.
Lines 93–97 and 595 The study domain is a coastal area within roughly 30 km of the community, informed by four sites. I would suggest "local" or "sub-regional" rather than "regional." Line 595, which pairs "regional scales" with "spatial resolutions of less than 50 m," reads a little awkwardly as a result.
Line 130 The sensors were reinstalled adjacent to their original locations on 2 May. Could the authors state the displacement, and whether the record in Fig. 2c is treated as continuous across the interruption?
Lines 137–141 Millero and Leung (1976) concerns seawater thermodynamics and seems an unusual citation in support of the absence of snow melt. Separately, the Site 1 sensor reading at or above 0 °C from 5 May sits somewhat awkwardly with the statement that melt was absent; a sentence reconciling the two would help.
Lines 156–161 and Table 1 The treatment of snow ice thickness is difficult to follow. The text states it was not considered owing to measurement uncertainty, while the Table 1 caption states that slush thickness includes it. Please make the convention explicit and consistent.
Lines 288–291 Sigma nought is normalized to 28° using a single slope for C-band HH (−0.23 dB/°, R² = 0.67) and for HV (−0.02 dB/°, R² = 0.01). With R² = 0.01 the HV normalization has little basis and might simply be dropped or flagged. More substantively, the angular dependence of flooded surfaces is itself reported in the earlier literature as a discriminator, so applying a single pooled slope across both classes could suppress or introduce class differences. Could the authors derive the slope separately per class, and show the sensitivity of the results to that choice?
Line 291 "Wiebke et al., 2020" appears to use a given name, and I could not find a corresponding entry in the reference list.
Lines 348–349 versus 425–427 and Fig. 6 Line 349 states that the HH class difference was not dependent on incidence angle, while Sect. 3.3 reports increasing separability with incidence angle for GLCM Variance and differences confined below 37° for HV. These appear inconsistent.
Lines 402–404 and 452–462 Treating 1 − coherence as a probability of flooding is convenient, but it is an unscaled decorrelation measure rather than a calibrated probability, so comparing its F1 at a 0.50 threshold against random forest probabilities is not quite like-for-like. Reporting it as a coherence threshold, with the ROC curve and the chosen operating point, would be cleaner and would not weaken the finding.
Lines 465–474 The polynya observations are described in the text as subjective indicators, which I think is right, but the abstract and conclusions read as though they validate the regional predictions. Since polynya formation is also associated with thin and dynamically active ice, I would suggest describing this consistently as a qualitative consistency check.
Lines 501–503 The suggestion regarding upward capillary transport and volume scattering is plausible but appears speculative in this context; marking it as an inference would be clearer.
Lines 505–507 The proposed C-band explanation involving snow volume change between passes is closely related to work already in the reference list on C-band coherence and snow thermodynamics. Engaging with it here would strengthen the interpretation of the single C-band pair.
Line 469 8 May to 11 June is 34 days.
Citation: https://doi.org/10.5194/egusphere-2026-4311-RC2
Viewed
| HTML | XML | Total | Supplement | BibTeX | EndNote | |
|---|---|---|---|---|---|---|
| 121 | 49 | 16 | 186 | 16 | 17 | 15 |
- HTML: 121
- PDF: 49
- XML: 16
- Total: 186
- Supplement: 16
- BibTeX: 17
- EndNote: 15
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
Overall
Thank you to the editor for inviting me to review this manuscript, and thank you to the authors for an interesting paper with a new approach to detecting flooding from remote sensing.
Overall, the paper provides some useful observations of flooding adding to a limited dataset. The SAR approach is interesting, and the paper articulates several parameters that can be explored, focusing on the coherence as a potential parameter to distinguish flooded and non-flooded scenarios, and also providing some results on how snow ice formation may be detectable.
To improve the paper, throughout the text the terminology could be reviewed for clarity when discussing flooding. I have highlighted some instances in the detailed comments. Phrases including ‘surface flooding’, ‘seawater flooding’ and ‘flooded sea ice’ are used, and I would take care to ensure it is clearly articulated that natural flooding occurs from seawater infiltration into the snow atop sea ice (at the snow/ice interface) which is more commonly referred to in the literature generally as ‘flooding’ and instances in the past tense would be for example 'flooded snow' (rather than flooded ice). I would try to be consistent with phrasing throughout.
A clearer explanation could be provided on how the SAR imagery relates to the field data. For example, given the scattered flooded and non-flooded locations at the same sites in the field data, how does this relate to the dates and analysis in the SAR data? I believe this could be explained further in the methodology and linked more clearly in the results. Additionally, several of the figures should be improved to present the data more clearly, with specific notes given in the detailed comments.
I note that my expertise is best aligned with the understanding of flooding, snow ice formation, and field experiment areas of this paper, so the more detailed SAR aspects of the paper would benefit from additional comments from the editor and other reviewer.
Abstract
Introduction
Data and Methods
2.1
2.2
2.3
2.4
2.5
Results
Discussion
4.2