the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
Super-resolution of Arctic Sea Ice thickness using a conditional diffusion model
Abstract. Small‑scale variability (3–60 km) in Arctic sea‑ice thickness plays a crucial role in sea‑ice predictability and in the climate system. However, these scales are neither directly observed nor adequately represented in climate models. While coarse‑resolution observational products (e.g., CS2SMOS) and some high‑resolution model simulations exist, bridging the scale gap remains challenging.
In this work, we use machine learning to develop a super‑resolution algorithm that reconstructs small‑scale sea‑ice thickness features from low‑resolution input fields. The algorithm is trained on realistic high‑resolution model simulations and is based on diffusion models conditioned on low‑resolution observations. This class of models is inherently probabilistic, enabling the generation of an ensemble of plausible high‑resolution reconstructions from a single coarse‑resolution input.
We apply the method both to model simulation, where high‑resolution ground truth is available, and to the CS2SMOS observational product. We demonstrate that the algorithm produces realistic high‑resolution sea‑ice thickness fields with improved accuracy and provides meaningful uncertainty estimates through the ensemble spread.
- Preprint
(2402 KB) - Metadata XML
- BibTeX
- EndNote
Status: final response (author comments only)
- RC1: 'Comment on egusphere-2026-2318', Anonymous Referee #1, 07 Jun 2026
-
RC2: 'Comment on egusphere-2026-2318', Anonymous Referee #2, 07 Jul 2026
Review of “Super-resolution of Arctic Sea Ice thickness using a conditional diffusion model”
In this manuscript, the authors present a novel method based on a conditional diffusion model to generate an ensemble of high-resolution sea ice thickness (SIT) fields. The model takes as input low-resolution SIT fields along with additional information such as concentration, deformation, and the sea ice mask. To evaluate their method, the authors first focus on a twin experiment, generating low-resolution observations from a high-resolution sea ice model simulation in order to compare the generated outputs against ground truth. In a later stage, they apply their method to real observations from CS2SMOS. Overall, the paper is easy to read and follow, and the method appears sound and promising, with potential for extension to other variables and resolutions. I believe the paper is a valuable contribution to the community and relevant to The Cryosphere. However, I think some questions remain unresolved, and some improvements are needed before the paper can be considered for publication.
Main comments
1- I believe the discussion section could be substantially more elaborated. At present, it focuses mostly on technical neural network parameters and on the split between the deterministic and probabilistic components. However, I think several points raised in the paper merit more in-depth discussion. In particular, the fact that the model is trained on simulated low-resolution observations from a given model (well known for its ability to represent deformation in the ice pack) and then applied to real observations constitutes a strong hypothesis that warrants further discussion. Would the results still hold if a different model were used? I would also be very interested to know how much the pre-processing step used to match the neXtSIM-simulated low-resolution field to observations affects the results.
2- The overall quality of the figures could be improved. All figures appear pixelated, which is particularly problematic for subplots containing many maps, especially Figures 1 and 2. The color scheme in Figure 7 could also be improved, in my opinion, to make it easier to distinguish differences between the CryoSat, ML, and LR PSD curves.
3- I find the question of seasonality quite interesting, and almost puzzling. I understand the explanation offered in terms of increased sea ice thickness, but would it be possible to use a "normalized" metric to further validate this point, for example, by dividing the RMSE by the mean SIT? Additionally, could the use of ECMWF forecasts instead of ERA5 in 2023 be another contributing factor? I think this point deserves a mention somewhere in the manuscript.
4- While the overall structure and organization of the manuscript make sense, I believe there is room for improvement in two respects. First, the naming of the sections could be improved, particularly in Section 3: why is the baseline introduced within the neural network section? The same question applies to the metrics. I would suggest either splitting this section or renaming Section 3 to better reflect its content.
Specific comments
- Sea ice / sea-ice: there are inconsistencies in the use of "sea ice" versus "sea-ice" as an adjective (e.g., "sea ice thickness" in L31 versus "sea-ice thickness" in L66). Please check all instances for consistency.
- L33: The principles of probabilistic approaches could benefit from a clearer introduction, particularly the idea of learning a distribution and then sampling from it.
- L49: The introduction could be updated to include section numbers and a brief mention of the conclusion's content.
- L73: This sentence could be rephrased for clarity, as its current meaning is unclear.
- L123: Equations should be enclosed in parentheses.
- L124: Please clarify whether the mean and standard deviation are computed globally or pixel-wise.
- L131: "Input features" is a familiar term for those accustomed to neural networks, but I would suggest rephrasing this sentence so that readers less familiar with the terminology understand that it refers to the inputs of the NN.
- Algorithms 1 and 2: These describe a standard conditional DDIM algorithm; I suggest moving them to the appendix.
- Section 3.2: The section title is somewhat confusing.
- Section 4.1: The section title could be improved.
- L245: Two instead of three timesteps, why does this apply only to the validation dataset?
- L267: The PSD acronym has already been defined earlier in the text.
- L275: Do you mean "fine-scale bias"?
- Fig. 4: I am puzzled by the differences in PSD across time steps. Is there high variability from one time step to another? Looking at the three PSD plots, the ML PSD appears fairly consistent across all three, whereas it is the neXtSIM PSD that shows differing slopes.
Citation: https://doi.org/10.5194/egusphere-2026-2318-RC2
Viewed
| HTML | XML | Total | BibTeX | EndNote | |
|---|---|---|---|---|---|
| 314 | 97 | 19 | 430 | 42 | 37 |
- HTML: 314
- PDF: 97
- XML: 19
- Total: 430
- BibTeX: 42
- EndNote: 37
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
This manuscript presents a conditional diffusion-model approach for super-resolving Arctic sea-ice thickness from low-resolution observational products. The model is trained using high-resolution neXtSIM simulations and synthetic low-resolution counterparts designed to mimic CS2SMOS and related satellite-derived inputs. The approach is then evaluated both in an idealised model setting, where high-resolution truth is available, and on CS2SMOS data.
Overall, I find the manuscript promising and potentially publishable after revision. The method appears sound, the paper is generally clear, and the results show meaningful improvement over the low-resolution baseline. However, there are some issues that should be addressed before publication, which are shown below.
Main comments:
1:
The baseline comparison is limited. Comparing against the low-resolution input field is necessary, but not sufficient to assess the added value of the proposed diffusion model. The paper would be stronger if the authors added at least one stronger baseline, such as a deterministic neural-network super-resolution model, a directly trained U-Net for the SIT increment (since the diffusion model used here has U-Net as backbone), a simpler stochastic baseline, or an ablation showing the added value of the diffusion formulation. This would help distinguish improvements due to the diffusion model from improvements due simply to supervised learning from neXtSIM.
2:
The construction of the synthetic low-res and high-res data seems to be central to this study, because the main quantitative validation is performed in a way that high-resolution neXtSIM fields are smoothed to create low-res inputs, and the original neXtSIM fields are then used as ground truth for evaluation. The observational application reveals a domain-shift issue. The authors state that spurious coastal features occur because OSI SAF deformation masks were not present in the training data. Since the model is intended for CS2SMOS/OSI SAF inputs, the synthetic training data should mimic their missing-data patterns as well if possible. Why were realistic OSI SAF masks not included during training? Perhaps the authors can clarify this point a bit more and address this in the training/evaluation setup and see if it improves. Also, it is not clear that what kinds of CS2SMOS patterns are absent from the training data, like where they occur, or how strongly they affect the results. The authors should elaborate on this point and provide clearer examples.
Other comments: