A prior-regularised heteroscedastic ResUNet for fusing passive-microwave sea-ice concentration products
Abstract. Passive microwave (PMW) sea-ice concentration (SIC) products provide pan-Arctic coverage, but their retrieval errors are often elevated near coastlines and within the marginal ice zone (MIZ), where land spillover, mixed pixels, melt ponds, and atmospheric effects complicate retrieval. We present a prior-regularised two-level residual U-Net (ResUNet) that generates a 12.5 km Arctic SIC field by fusing six PMW SIC products: Bremen-ASI, NSIDC-BT, NSIDC-NT, NSIDC-CDR, SICCI-25km, and OSI-450. The model combines a compact encoder-decoder backbone with auxiliary spatial and seasonal encodings and a heteroscedastic Gaussian negative log-likelihood loss. Empirical relationships between PMW SIC error and distance to land or to the ice edge (defined here as the 0.15 SIC contour), together with product-provided uncertainty estimates, are incorporated into the loss function as pixel-wise reliability information. On the held-out test set, the fused product outperforms all six individual PMW SIC products, reducing MAE by about 55 % and RMSE by about 30 % relative to the best-performing PMW product while maintaining near-zero bias. In the most error-prone regions, RMSE is reduced by about 34 % within 20 km of the coast and by about 25 % within 50 km of the ice edge. Independent validation against Landsat-derived SIC gives the lowest MAE and RMSE for the fused product (0.035 and 0.062), corresponding to improvements of about 13 % and 40 % over the best-performing PMW product, respectively. The fused product also has the lowest errors in the most error-prone regions and smaller interannual RMSE variability. The estimated heteroscedastic uncertainty is informative: on the held-out test set it increases consistently with error and reaches its maximum in coastal and ice-edge regions, while the Landsat validation also shows a consistent ordering of errors with uncertainty. Overall, the proposed framework combines complementary PMW SIC products into a single fused product with improved accuracy and reduced errors near coastlines and the ice edge, providing a useful basis for climate applications and near-real-time sea-ice mapping.
Review of
A prior-regularised heteroscedastic ResUNet for fusing passive microwave sea-ice concentration products
by
Fengxin Chen et al.
Summary: This contribution deals with the application of a ResUNet machine learning approach to fuse six different sea ice concentration products based on satellite passive microwave observations with the aim to reduce the uncertainty statistics of the individual products. The study focusses on regions of highly variable sea ice concentrations and/or regions where the known sea ice concentration products tend to have their limitations. The machine learning approach is tuned / trained with a data set of the sea ice cover fraction based on MODIS optical observations. This data set covers years 2000 through 2011 and the summer months May through early September. The authors describe the ResUNet framework well and provide quite a number of details of it. They demonstrate that the training works well and come up with convincing test results. They evaluate the fused sea-ice concentration product, and all input sea ice concentration products separately, against a held-out test case of the MODIS data set used for training and testing and, in addition, against sea ice concentration derived from a product of classified Landsat-5 and 8 optical imagery. Both attempts to evaluate the fused product suggests that this product compares better to the two evaluation data sets than all six input products and that in particular in the (partly hypothesized) main problematic regions the ResUNet based approach is demonstrably more accurate.
This is without doubt an interesting contribution which content should be published. But before this could happen, the authors should have a look at the general comments which summarize some of the comments and concerns I list in my specific comments; the latter also contain additional comments and hints for improvement of the manuscript. I close with a number of editoral comments that the authors also might want to pay attention to.
General Comments:
GC1: A substantial fraction of this manuscript is based on the hypothesis that SIC retrieval errors are particularly large near the coasts. This lacks a properly described basis. It is not sufficient to take the 8-daily MODIS surface fraction data set as a motivation here - simply because that product has the pan-Arctic melt-pond fraction as the main focus and is not particularly well validated near the coasts.
GC2: The manuscript requires a more specific description of the algorithms, of the different filters used, and - in particular - how the resampling to a common grid and the fact that the products may use a different land mask may have influenced your results. Please see the respective specific comments.
GC3: Algorithm development, fusion of derived products and the evaluation of the products requires a sound understanding of the data that are used - particularly of their uncertainty characteristics and their potential limitations. Some data are fit-for-purpose in the sense of what one aims to do, others are not. Using the MODIS product ice edge information as THE reference data set in this regard is certainly not optimal. Your manuscript does not at all reflect this limitation. See also the respective specific comments which all rank around using "distance to the ice edge" as one of the main "problem zones" for passive microwave sea ice concentration retrieval.
GC4: Related to whether a product is fit for which purpose is the fact that the involved products all come with their retrieval uncertainty which - in the best case - is 1% or 0.01 in your notation. When providing your results you must go away from numerical precision and provide results with a resonable precision that matches the accuracy of the products used. Hence, please move a way from 0.0001 or 0.01% precision everywhere in your paper. It is meaningless.
GC5: Several of the figures are not described in an adequate way yet and also contain features that either require amendment of the figure or more text explaining the content.
GC6: Your paper requires a section in which you discuss critically your approach and its limitations, and your results, and put them into context with results published elsewhere. For example, you show results of an intercomparison between your new product and the six input products to a MODIS-based SIC data set and to a Landsat-based SIC data set. These intercomparisons partly overlap with results published elsewhere, e.g. Kern et al., 2020, and Kern et al., 2022 and it would be interesting to see how your results compare to their results. One of the issues that came up repeatedly in the past when evaluating SIC products, is that products are truncated at the lower and upper end of the SIC distribution, resulting in a misestimation of both, accuracy and precision when not taking this truncation of the naturally retrieved SIC values into account. I am also missing a more thoughtful consideration of the different spatial resolutions and the different filters that are used in the input products. In addition, the limitations of the MODIS product in terms of delineating the ice edge - now that you aware of that - should be discussed.
Specific Comments:
Title: You developed and tested your approach for the Arctic Ocean. I'd find it useful to add this to the title.
L32: One overarching question: Why are you focussing on the Arctic Ocean? The sea ice cover in the Southern Ocean has been shown to have a roller-coaster like intra-annual variability since about 2015. So, what is the motivation to stay in the Arctic Ocean?
L49-62 and remainder of the paper: Please make sure that you clearly distinguish between an "algorithm" or "approach" that is used to retrieve or compute SIC from passive microwave observations and "products" which are the result of the application of an algorithm or approach and might involve additional steps such as filtering or, in case of the NSIDC CDR, the combination of the results of two algorithms - or in this case even products.
L63: "coastlines" are not necessarily regions of low-to-intermediate SIC. This is true only during summer. During winter and also during the shoulder seasons, the SIC along the coastlines is often particularly high (and stable), especially when the so-called land-fast sea ice forms. Please check your writing - especially also in the following sentence(s). You might need to go through your manuscript to avoid a false association between SIC retrieval uncertainty sources and the coastlines and/or regions near the coast. --> see also GC1
L65: "even after filtering" --> For improved understanding of what you mean here I recommend to add more text about the different filters that are applied to the various products.
L79/80: It is not sufficiently clear what you refer to by "narrower ice edge regions". Please specify better.
L113: Please note that a new MODIS-based data set is available now which is based on more recent MODIS processing, has daily temporal resolution and spans the years 2000 through 2024, June through August: https://doi.org/10.25592/uhhfdm.18069
L117/118: This data set contains a "low quality" version and a "clear-sky" version of the MPF and OWF data. I hope you used the clear-sky version. Please mention your choice in the text.
Table 1:
Please provide - where possible - the version information of the products used. You did so without knowing for OSI-450 - which now is by the way OSI-450a and OSI-450a1 plus OSI-430a for the extension as ICDR beyond 2020.
Of the NSIDC CDR V4, V5, and V6 exist. Did you consider version V6? If not, why not?
You must provide more information about the NSIDC CDR, how it is created, what the provided uncertainty information is and so forth. It is not clear what you mean by "consensus NT/BT". Please check and correct.
SICCI-25km was an experimental product which found its continuation in the OSI-458 product. Please check out the respective OSI SAF web page. For sure it is NOT merged in any sense with the NT2 product! This is an error that needs to be corrected. It also does NOT use the 89 GHz channels.
I note that you write 19.35 and 37.0 for the frequencies of the NSIDC-NT but you don't use decimal digits for any of the other products.
For some of the products you write that they extent until "present". While this might be true (it depends on the version) the list of sensors is then not correct in some cases because SSMIS data have not been available since September last year. Please check and correct.
Your references for the three NSIDC products also deserve another check because the Cavalieri et al. (1991) paper is quite old and I am almost sure that it would make more sense to go to a more recent publication. In general, I would recommend to decide whether the column named "Literature" is meant to list references of the data product or references of the algorithm. Currently, it seems to be a mixture of both.
I suggest to add in this table which product provides which kind of uncertainty. SICCI-25km and OSI-450 provide a physically based uncertainty, NSIDC-CDR a statistically based uncertainty.
L143: "based on an enhanced NASA Team 2 (NT2)-type retrieval" is wrong and needs to be corrected.
L146/147: How is the resampling of the products to the (mostly) finer grid resolution done? I note that some algorithms come on the same polar stereographic grid but some come on the EASE2.0 grid. This requires a different complexity of the resampling process. It is not sufficiently clear whether (and how) the information content of the finer resolved ASI algorithm product is preserved (or smoothed) and it is not clear how the coarser resolved information of all the other SIC products is "transferred" to the finer grid. The resampling process might change the statistics of the products and as written it is not sufficiently clear to which degree this is taken into account or quantified. --> see GC2
L159/160: "We then make ... products" --> This is a very short description of a potentially more complex process because different products use different land masks AND the pole hole differs between the SIC product and the MODIS product. Please provide more details here. As written this part of the pre-processing is not clear enough. --> see also GC2
"these pixels are excluded from the loss calculation and evaluation metrics" --> I hope these zeros are also not used in the training?
L172: I think this data set has been expanded by data from years 2018 and 2019. Have you considered to include these?
L214: Given the fact that NSIDC-CDR is a combination of NSIDC-NT and NSIDC-BT and, in addition, SICCI-25km and OSI SAF are similar in the algorithm and filtering and only differ in the input sensor, it has to be questioned whether all six PMW SIC products provide complementary information. This is something that falls into GC2 where - at the end - you need to motivate why you have chosen the products you have chosen. This is not sufficiently clear.
L215: PMW SIC errors are largest in the MIZ / along the ice edge - but not necessarily along the coastline. The design of the fusion framework along these two considerations is therefore not optimal. --> see also GC1
L260++
Following up with previous comments into this direction: I seriously doubt that the "distance-to-land error structure" makes sense - at least not without i) giving further explanations how you treat filters and/or how different products differ in their efforts to correct the land-spillover effect, ii) giving further insights in how your resampling to the 12.5 km polarstereographic grid of the MODIS "SIC" data set may have influences SIC estimates near the coastline, iii) giving information about how you "maximized" the land masks, and iv) giving a physically-based explanation why the error near coastlines is supposed to be higher. Any SIC error near the coastline can have multiple sources and one needs to understand them BEFORE trying to use the distance to the coastline as a means to characterize SIC uncertainties. --> see also GC1 & GC2
L260++ #2
The "distance-to-ice-edge" is a reasonable metric. However, when I understand the MODIS "SIC" data set correctly, then that data set had the focus to delineate PRIMARILY melt ponds on sea ice. The classification approach used does ALSO provide information about the open water fraction and, indirectly the (net) ice surface fraction ... but the validation of that product concentrated on the melt-pond fraction. Whether the SIC computed as 1 minus the open water fraction is so super reliable and whether the 15% MODIS SIC isoline is a useful representation of the ice edge has not been checked and was also not in the focus. The new MODIS melt pond fraction data set I mentioned before is a bit better in that respect because the evaluation also covers ice surface fraction and open water fraction - but that data set you did not use. So, in summary, I recommend to be very careful to not over-interpret and over-use the MODIS "SIC" product. If I am not mistaken, the publication by Kern et al. (2020) also deliberately did not focus on the ice edge or the MIZ. --> see also GC3
L275: I note that the non-zero background error level you are reporting here could simply be caused by the fact that your reference data set represents late spring / summer conditions, hence a time period of the year where all six SIC products have their deficiencies. While this is not a problem per se, it should not be overinterpreted and, in particular, this finding should not be taken for granted to be the same for the other months of the year. Here your reference data set clearly has its limits in applicability and I recommend to state clearly that the relatively high RMSE values obtained should not be generalized. See also my comment at the beginning of this sub-section and --> see also GC3.
L310 / Equation 11: I can understand that you set the value of r_i^unc to 1 for those SIC products that do not provide an uncertainty. However, you could follow the example of the NSIDC-CDR product and compute a 3 x 3 grid cell standard deviation for these other products and redo your analysis with this standard deviation. While these cannot compete with the phyically-based uncertainty estimation carried out in SICCI-25km and OSI 450, it is perhaps a more appropriate additional input parameter.
L467-475++: While I can follow your argumentation why you, at the end, selected the two-level ResUNet, I have problems to understand what I see in this table 4. We are talking about RMSE variations at the 3rd decimal digit, i.e. in L471, when transformed into percent RMSE we talk about 13.05% versus 13.16%, so a difference in the RMSE of 1/10 of a percent. I am sceptical that the accuracy of the input SIC products justifies to work and argument with such incremental differences in the results achieved. --> See also GC4.
L480-483: I find this statement interesting because when I compute the sum over the values given in the respective rows I end up with 0.7222 for your favorite two-level ResUNet but with 0.7137 and 0.7142 for the three-level ResUNet and the U-Net++, respectively. But it is possibly nonsense to look at these numbers including the RMSE from the training.
I note that the three-level U-Net++ clearly benefits most from auxiliary decoder heads and PMW derived diagnostics. Do you have an explanation for that?
Table 5: Excuse my question, but for this comparison you used the SIC products on their native grid with their native grid resolution?
L501-503: Even though you did this investigation with a held-out set of the MODIS data, hence those data did not enter the training of the U-Net framework, it is still a comparison to the type of data set that has been used for the training. It is therefore note overly surprizing - to my opinion - that compared to all input SIC products - the new fused product performs best. I dare to say that if you would have used the Bremen-ASI data set for the training you would receive the best performance of the fused product against the held-out part of the Bremen-ASI exactly also for the comparison with the Bremen-ASI SIC.
L505: The RMSE of 0.1364 is a bit worse than the values you showed for this RMSE in Table 4, both for training and validation. Why?
L506: Likewise, for the soft ice edge case, you obtained a better RMSE than both values given for training and validation in Table 4. To me this emphasizes one more time that there is considerable variability in the results and looking into differences in the 3rd digit like you did in the context of Table 4 to select the "correct" approach is not overly helpful. --> see also GC4.
Figure 8: I note that you also show the MAE for open water areas outside the sea ice cover. I doubt that the MODIS "SIC" product is overly reliable here and I suggest that you cut off any comparison (also in these maps) at 15% MODIS SIC. I assume that the values obtained over open water did not enter your evaluation results, for exeample those shown in Table 5?
The fact that the UNet SIC shows a MAE near zero over the open water one more time demonstrates how closely the UNet results are to the MODIS SIC that is used for the training. I don't think this is an overly useful result and just shows how much the UNet results are influenced by the training data.
I suggest that you use the same range over values in both sets of maps: from 0.0 to 0.225, for example. That way the two sets of maps can be compared to each other even a bit better.
You need to be more specific about what the grey and white colors denote. So far you only refer to grey. Also: Why do you have grey AND white at the region without observations at the pole?
In the top right corner of all panels one can see values of the MAE that are located outside of the region where the MODIS data used for training, testing and evaluation have valid SIC values. Where do these MAE values come from?
Figure 9: MODIS and UNet SIC maps show a sea ice concentration over open water that is different from 0. This cannot be correct. Clearly, over open water ResUNet does not provide the correct physical interpretation.
Please note the date of image (a).
The title says "Best clear-sky date" --> What is "best" and how did you decide upon that?
You need to also provide information in the caption what the white areas shown in panel (a) are. I note that grey only denotes the area of missing observations near the pole and land in panels (b) to (h), missing observations seem also to be colored in white in panel (a); this is inconsistent and should be avoided.
The MODIS product does not have valid values south of 60degNorth (see panel a); still ResUNet shows SIC values there which are definitely wrong if they are not at 0% (or 0.0).
L534-537: "Our fused ... spatial structure." --> This is a very qualitative statement. I cannot follow that statement in view of Figure 9. In order to back up your statement I suggest to i) focus on regions where you think that there is less "coastal contamination" and why, and to ii) provide sample maps of the ice edge where you think that ResUNet provides a better degree of detail than the other products.
Figure 10: In light of what I have seen in Figures 8 and 9 this result does not surprise me. It seems clear to me that the ResUNet is very much trained towards the MODIS "SIC"; it therefore shows these small MAE values. In turn, the six SIC products naturally do have a difference relative to the MODIS "SIC" (as shown, e.g. by Kern et al., 2020) simply because all these products have difficulties deriving SIC from satellite passive microwave data during summer conditions. In other words, one could ask the question whether it would not be more useful (and less computationally expensive) to simply use the MODIS "SIC" product.
Apart from that I again would like to make the point that the quality of the MODIS "SIC" product in the sea-ice edge regions is not optimal. And I would like to repeat my comment that so far you could not demonstrate sufficiently well, why coastal regions - physically - might be a difficult region for the passive microwave SIC data. --> see GC1 & GC3.
Figure 12: I cannot follow your interpretation of this figure. To me the mean error (displayes at th y-axis) does not show any dependency to the uncertainty shown at the x-axis. To me it looks as if the error distribution is the same for basically all uncertainty values shown. But perhaps I can also simply not read the legend and annotations properly because the font size is overly small.
It took me a while to figure out that the legend of the left y-axis seems to be the absolute value of the error, no? Please use "Mean absolute error" instead of "|error|" - also in the title and in earlier or following figures. If you think the space is not sufficient use MAE and explain the abbreviation in the caption of the figure.
What does "per hex" refer to?
Figure 13: I suggest that you use the same color (white) for missing values and (grey) for land in all figures that show a map of the quantities derived. Then you can also be more specific with the explanation of what these colors denote.
Comparing these maps with other, similar maps in your paper again suggests that the values shown in the upper right corner are not correct. In addition to non-zero mean uncertainty values (see Figure 8) there is also a cone-shaped region colored in grey for which it is not clear where this comes from. Please correct.
L580: "distributed mainly in coastal regions" --> I am sorry, but Figure 3 clearly shows that at least half of the Landsat images does not overlap with a coast - especially those in the Hudson Bay and in the Beaufort Sea.
Table 6: Here you are comparing two SIC products that were derived from satellite remote sensing. Both of them have - at least - a retrieval uncertainty of 1% - aka 0.01 in your notation. The number of decimals that you have been using is not supported by the accuracy of both products and I therefore suggest that you, at maximum, use three digits - which still would suggest an accuracy one order of magnitude better. So, your RMSE dland<20km column would read: 0.116, 0.173, 0.155, 0.173, 0.148, 0.140, 0.178 - aka 11.6%, 17.3%, 15.5% and so forth. Please also change the text that refers to the numbers given in Table 6 and Figure 14 accordingly. --> See also GC4
Figure 14, panel (d): I suggest to not connect the single points plotted per year because this implies that we are looking at kind of a time series or transect. It also does not fit well with the fact that the Landsat SIC of years 2003-2010 is based on Landsat-5 images, mainly located in the Hudson Bay while the images of the years 2013-2015 are based on Landsat-8 images, located outside of the Hudson Bay in most cases.
It might make sense to somehow indicate the number of Landsat images per year that were used to understand the statistics behind each single data point better.
Figure 14 / L605 & L613-615: Given the comparably small number of Landsat images intersecting with the coast (Fig. 3), I suggest to tell the reader how many Landsat image SIC contribute to the results shown in Figure 14b). I am quite sure that there is a change in the statistics across the x-axis range shown.
Did you observe a difference in the skills between Landsat-5 (mostly Hudson Bay and therefore first-year ice) and Landsat-8 (mostly interior Arctic Ocean and peripheral seas and hence a mixture of first-year and multiyear ice?
Figure 14c: Like for Figure 14b I suggest to tell the reader on how many Landsat image SIC matchups this figure is based to better understand the statistics behind the reults.
L655-658: "Rather than relying ... extensions." --> I cannot follow this last sentence. First of all, the ResUnet approach has been trained with a product that is based on optical data and secondly, the comparison to the held-out test case of the MODIS data used for training clearly demonstrates that the fused product is tuned substantially towards the MODIS product and seems not to be independent.
Editoral Comments / Typos:
L15/16: You mention coastlines and within the marginal ice zone as regions where SIC retrieval might be problematic and then list land-spillover, mixed pixels, melt ponds and atmospheric effects. This list is too broad because 1) land spillover is limited to the coastlines, 2) melt-ponds are a pan-Arctic phenomenon and influence the SIC retrieval everywhere, and 3) atmospheric effects are not just limited to coastlines and the MIZ. A solution could be to only mention the 4 phenomena and get rid of the parts about the regions. --> see also GC1
L38: "grid cell" --> better "known ocean surface area"
L47: "under clear skies ... constrained by cloud cover" --> I guess it is sufficient to mention the problems with clouds once.
L61/62: "However, they ..." --> Please provide a reference for this statement.
L75/76: I guess it would not hurt to also cite the various papers dealing with inter-comparison of various SIC products against various independent data sets, i.e. Kern et al., 2019, 2020, and 2022, all in "The Cryosphere".
L82: "product-provided uncertainty estimates" --> I suggest to add: "where possible"
L157: "The screening yields 156 MODIS SIC scenes" --> I doubt that what you describe in the previous sentence is the reason why you work with 156 MODIS SIC scenes at the end. Since you look into months May through August, I am inclined to say that the 156 is the result of 11 years times the number of 8-day periods with data within May through August minus a few missing 8-day periods. So, you can simply state that you have 156 MODIS scenes available.
Figure 1: While panel (b) can in fact be quadratic because the x- and y-dimensions are similar, this is not possible for panel (a) because the native grid has 608 rows times 896 columns. I strongly recommend to show panel (a) with the original aspect ratio to avoid confusion.
And, by the way, the already mentioned novel MODIS "SIC" product comes on a 531 x 531 grid so that you would not need to crop the MODIS data.
Please provide information about the grey and white areas in panels (a) and (b).
L216/217: "and can ... agreement" --> This part of the sentence is repeated almost one-to-one in the next paragraph and can be deleted here therefore.
L223: Sorry for my ignorance, but what is an "ablation experiment"? Ablation is a term I know from glaciers.
Figure 4: The rectangular box at the top centre of the figure contains something written in a super small font. Consider enlarging this, please.
L226: "base input" --> In the text further up you termed this "core input". Please use the same expression or term for the same meaning.
L250: Again I suggest to add "where available" behind the "product-provided uncertainty".
L499: "within 50km ice edge" --> "within 50 km of the ice edge"
Figure 7 / L511: I am not sure I fully understand the "all samples" case. Shouldn't this be the held-out cases you used for Table 5?
Figure 11: I suggest to add the 1-to-1 line of perfect agreement in all three panels.
Figure 15: I am not sure I understand what you mean by "(%-SIC)" in the annotations of the x- and y-axes of panel (a). It seems you switched to "percent" as the unit for SIC and its uncertainty. Why?
In L630 of the caption you write "median estimated uncertainty" while in panel a) you write "Median Pred sigma". Please ensure you use the same names in the text and in the figures. You also write "within uncertainty bins" --> Where can I find these bins and what are their boundaries?
L633/634: "The monotonic ..." --> I take this sentence in the caption of Figure 15 as an example to suggest to reduce the information content of your captions considerably. You have many captions where you repeart information that is given in the text as well. While figures need to be self-explanatory theys should not include elements of your interpretation or discussion of the results shown. I am aware that this over-arching editoral comment comes quite late; apologies for that.