the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
Evaluation of subseasonal sea ice changes in the Community Earth System Model version two
Abstract. Forecasting Arctic sea ice is a complex, open problem in polar science, exacerbated by climate change and Arctic amplification. Arctic sea ice dynamics are inherently interconnected to atmosphere and ocean dynamics. This means that coupled climate models are currently the best available tools to model and forecast sea ice. However, climate models consistently underrepresent sea ice processes on subseasonal timescales, contributing to forecasting challenges. In this study, we evaluate the Community Earth System Model Version Two's (CESM2) skill in representing Arctic sea ice by comparing model output to reanalysis and observational data. Doing so establishes CESM2's viability as a tool for studying Arctic sea ice processes and identifies places where model improvements are needed. While sea ice processes in CESM2 have been evaluated on annual and climate timescales, its performance on subseasonal timescales has not. We analyze CESM2's performance in representing subseasonal sea ice variability by comparing its temporal and spatial interannual variability to observations. We also evaluate CESM2's performance at capturing very rapid sea ice loss events (VRILEs). VRILEs are substantial sea ice loss events that occur on the timescale of days. We find that CESM2 does a poor job of representing VRILEs and has less interannual variability than observations.
- Preprint
(18107 KB) - Metadata XML
- BibTeX
- EndNote
Status: open (until 21 Sep 2026)
- RC1: 'Comment on egusphere-2026-3733', Anonymous Referee #1, 13 Aug 2026 reply
-
RC2: 'Comment on egusphere-2026-3733', Anonymous Referee #2, 01 Sep 2026
reply
I appreciate the authors' initiative in undertaking this type of study, as it is critical to understand how we can better use models and forecasts to benefit society, particularly in the polar regions, where we observations are scarce. However, this manuscript needs a substantial overhaul before it can be evaluated for publication. Several claims throughout the text either contradict the study's own findings or rely on citations that don't actually support them, and these issues run from the framing of the results down to the literature review and the methods.
The abstract calls coupled climate models the "best available tools" for sea ice forecasting, based on their structural coupling to the atmosphere and ocean. The study's own results show CESM2 significantly underestimates rapid ice loss events and has less variability than what is actually observed, so the authors need to reconcile this. A model cannot be called the best tool if its output fails this badly against real observations. In my view, these types of models face fundamental challenges in producing skillful sea ice S2S forecasts, and this context should be the baseline against which this type of study is framed. Though I encourage more work to understand how these models can support operations, as stated in the Conclusions, the authors should demonstrate a basic understanding of where the current status of this work stands within the community.
This same lack of rigor shows up in the literature review. Citations are often used to support broad claims about Arctic sea ice, but several of them don't hold up under a direct check of the source material. Some appear to be pulled out of context, and one directly contradicts the paper it's citing. The authors need to fact check their citations and conduct a more comprehensive and relevant literature review.
The paper also never discusses sea ice physics as a possible reason for the model's poor performance. CESM uses the CICE sea ice model with an elastic-viscous-plastic rheology, a computationally efficient version of the older Hibler approach. The authors should discuss whether the dynamics scheme itself, not just resolution, contributes to the problem, and should consider newer approaches, such as the Maxwell-Elasto-Brittle model or others, that may be better suited to the short-term timescales this study is evaluating.
The methods raise similar concerns. The regridding approach used for spatial comparison stretches the coarser model data onto the finer satellite grid, which introduces smearing artifacts in the marginal ice zone, exactly the region this study needs to evaluate precisely. The authors should either test how this affects their results or acknowledge it as a limitation. Additionally, the NSIDC satellite data is treated as ground truth throughout the paper without any acknowledgment of its own known limitations. Passive microwave retrievals have documented uncertainty in the ice concentration range and time of year most relevant to this study, especially during melt season, and this was not mentioned anywhere in the manuscript. Some of the reported gap between the model and observations could be coming from this uncertainty rather than the model alone.
Many of the issues in this review could be addressed by simply acknowledging them as limitations rather than redoing the analysis. However, collectively, that response isn't adequate for this manuscript. If the authors were to add appropriate caveats for every issue identified, there would be little left of the paper's central quantitative claims. An outcome of a study that requires several qualifications to stand is not necessarily a result the reader can act on. Meaningfully addressing the issues identified in this manuscript would require redesigning the analysis, which places this beyond the scope of a standard revision cycle. I therefore recommend rejecting this manuscript, while encouraging the authors to incorporate these recommendations toward a more robust analysis that would offer a more relevant contribution to the community.
Some specific issues are detailed below:
P1, L 14: For a manuscript tracking modern sea ice, citing a "twice the rate" statistic from 2007 and 2012 is quite outdated. Even with the 2017 citation included, this figure may be dated for a manuscript addressing modern sea ice — more recent estimates put Arctic warming well above twice the global rate. Consider updating with a more representative and current citations.
Pg 1, L 18: The citation of Durkalec et al. (2015) here appears to be a misattribution. Durkalec et al. (2015) is a qualitative study focused on the health, safety, and cultural well-being of Inuit community members experiencing sea ice decline. It does not provide data or conclusions regarding a general increase in the Arctic human population due to sea ice loss. Please find a more suitable citation.
Pg 2, L31: grammar error. The citation, DeRepentigny et al. (2020), should all be in parenthesis.
Pg 2 ,L 36: Who is "they?"
Pg 2, L45: This statement "March sea ice variability is driven by radiative feedbacks of clouds and water vapor
(Luo et al., 2017), surface albedo, surface winds, and oceanic heat transport, while September sea ice interannual variability is
primarily driven by surface albedo and fluctuations in temperature advection (Olonscheck et al., 2019)" oversimplifies what drives sea ice variability and does not mention the regional differences and other factors like local wind forcings. The citations used for these types of overarching statements are not used correctly. For example, Luo et al. (2017) is cited to support a general Arctic-wide claim about March variability, but that paper is actually about winter ice decline in the Barents-Kara Sea specifically. And the claim that September variability is "primarily driven by surface albedo" doesn't match Olonscheck et al. (2019) where the citation clearly states "Arctic sea-ice variability is primarily driven by atmospheric temperature fluctuations," not surface albedo. There needs to be more comprehensive literature review for these generalizations and not relying on single studies for broad claims.Pg 3, L63: grammar error: "One"...and then "are" should be "is"
Pg 3, L76-82: From this statement it is uncler to me what the purpose of this study is. The authors used just one model simulation out of many available, picked because it looked "average," not because it gave good results. But testing only one run doesn't tell you if the model actually works well at predicting short-term ice changes — you need to check several runs to know if that's a real pattern or just luck from that one run.If the purpose is to evaluate the model's ability to simulate subseasonal variability, shouldn't multiple individual members be evaluated in order to get a statistically robust sample. If not, the authors should clarify why individual members' daily data, stated as archived in the manuscript, weren't used to test multiple members rather than just one.
Pg 3, L76-77: The manuscript uses the NSIDC Bootstrap product as ground truth without discussing its known retrieval limitations. Passive microwave sea ice concentration algorithms, including Bootstrap, have documented uncertainty in the marginal ice zone (roughly 15-80% SIC, season and region dependant), driven by factors such as thin ice, slush, and liquid water fraction in the snowpack or snowmelt. Separately, retrievals are known to systematically underestimate SIC during melt season, since melt ponds contaminate the reference values the algorithms are calibrated against, with error increasing as SIC decreases. Given that several VRILE cases in this study occur during summer melt conditions (e.g., July 2007, August 2012) and that VRILE identification depends on detecting concentration changes in exactly this intermediate range, this is a relevant source of uncertainty that should be acknowledged. The authors should discuss how retrieval uncertainty in NSIDC data may affect the quantitative comparison to CESM2, particularly the magnitude of the reported VRILE size discrepancy.
Pg 3, L88-92: Sea ice physics are non-linear. The Arctic climate forcing, atmospheric CO₂ concentrations, ocean heat content, and the proportion of thick multi-year ice vs. thin first-year ice (and others) were substantially different in 1960 compared to 2020. Selecting years based on matching sea ice extent does not establish that the underlying variability-generating physics are comparable between these two periods. Since this comparison underlies the paper's interannual variability statistics (PSD, interquartile range, spatial SIC variability), please explain why these statistics should be expected to transfer across two periods with such different background climate/environmental states.
pg 4., L 99-103: Regridding coarse model data onto a finer satellite grid is not a good verification approach. In practice, it introduces smearing artifacts in areas like the marginal ice zone which is critical to capture when looking at subseasonal capabilities. Rather the fine observations should instead be upscaled to the model's resolution. Either the author should try to redo the setup to see how this affects the outcome or acknowledge the practical limitations of this approach.
Pg 4., L 103-107: Two things need clarification. First, since VRILE locations are found using the same gridded data that was regridded earlier (see comment on L99-103), the blurring from that step could shift where the center-of-mass calculation places the event. The authors should quantify how much error this introduces. For example, by comparing VRILE locations derived from the regridded data against locations derived from the model's native grid, and reporting the resulting shift in center-of-mass position. Additionally, percent change in sea ice coverage is used as a proxy for mass, but no formula or example is given, and the phrase is ambiguous. For example, the definition could refer to a percentage-point change in concentration or a relative percent change, which would produce different weightings. Please specify the exact definition used.
Pg 4, L 124-129: The lead times used across the four VRILE case studies range from three to eight days (Table 1), with the manuscript noting that using earlier forecasts "yielded highly unrealistic results" for the 2012 and 2016 cases. In S2S forecasting, lead time directly affects atmospheric predictive skill, so comparing cases initialized at inconsistent lead times makes it difficult to distinguish limitations of the model's physics from limitations of atmospheric predictability at the lead time used. Please either standardize lead time across cases or discuss explicitly how this variation affects the comparability of conclusions drawn from the four cases. Additionally, please clarify the reasoning behind the statement that a 2 July 2007 initialization "does not provide enough time" to evaluate the 3 July event. As written, this is unclear. Is the constraint that the method requires a five-day antecedent SIC change to identify and locate the VRILE, which a one-day lead cannot support, rather than an issue with single-day forecast validity itself?
Pg 5, L131-132: same comment regarding regridding from coarse to fine (see comment for pg 4., L 99-103)
Pg 5, L134-255: Several methodological issues affect most of the results in this section. This is presented in both the individual figures and the overall conclusions drawn throughout Section 3. Rather than repeat these points I have referenced earlier in section 2, they are summarized here:
- Figures 2 through 5 are all derived from a single CESM2 LENS ensemble member, with no evaluation of whether results would differ using other members (see comment on Pg 3, L76-82). This concern is compounded by comparing that member's 1960-2000 period to NSIDC's 1980-2020 observational record, despite differing background climate/environmental conditions between the two periods, without demonstrating the relevance of this comparison (see comment on Pg 3, L88-92). The author should provide a more relevant literature review that characterizes the retrieval uncertainty for the ground-truth data used, particularly in the marginal ice zone and seasons in which this study is presented. It would be good to acknowledge whether melt ponds or wet snow was likely present in regions that may have introduced larger retreival errors.
- The four case studies in Section 3.1 raise a related concern, each using a different forecast lead time (3-8 days), with two cases showing that shorter lead times produced unrealistic results, previously raised as a possibility that the cases shown are not representative (see comment on Pg 4, L124-129).
- Last, the observational uncertainty noted in the comment on Pg 3, L76-77 applies here as well, since the NSIDC data used as the benchmark throughout this section carries its own unaddressed retrieval uncertainty in exactly the concentration range and season most relevant to VRILE detection. These caveats around using passive microwave data as ground truth do not appear to have been addressed anywhere in the study design. The discrepancy between CESM2 and observations shown in this section is clearly substantial, but given all of the above, the paper's central quantitative claim that CESM2 underestimates rapid ice-loss events by roughly an order of magnitude is less certain than presented. It is the precision of the magnitude that is of interest, not the general conclusion that CESM2 underrepresents short-term ice loss, as this is already consistent with prior literature on the model's known limitations. It is recommended that the author provides the model comparison against the observational uncertainty range so it is more clear where the gaps are significant.
Pg 9, L 269 and 270: The abstract already acknowledges that models underrepresent subseasonal sea ice processes, but this statement goes further, suggesting that better representation of synoptic-scale processes would fix the forecasting shortfall. First, is the term "synoptic-scale" processes meant to be "regional or mesoscale" when referring to the detail that is missing? The use of "synoptic-scale" doesn't match the scale identified in the paper's own explanation in Section 3.2 (L251-255). It seems the authors suggest the limiting factor is grid resolution relative to much finer scales, ice floes (~100 m) and low-level jets (~17-23 km), not synoptic scale systems (~1,000+ km). While stated on pg. 8 L255 that this is a parameterization issue and consistent with "representation" being part of the fix, the scale named in section 4 doesn't match the scale the paper's own diagnosis points to. The authors should clarify which scale they mean, since synoptic-scale representation and floe/jet-scale parameterization are different factors that require different solutions to evaluate the model correctly.
Citation: https://doi.org/10.5194/egusphere-2026-3733-RC2
Viewed
| HTML | XML | Total | BibTeX | EndNote | |
|---|---|---|---|---|---|
| 102 | 43 | 17 | 162 | 19 | 11 |
- HTML: 102
- PDF: 43
- XML: 17
- Total: 162
- BibTeX: 19
- EndNote: 11
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
Peer Review of “Evaluation of subseasonal sea ice changes in the Community Earth System Model version two”
This study compares the sea ice evolution from a long-running CESM2 LENS climate forecast to NSIDC ice observations during periods in the simulations when the overall ice extent is comparable to the observations. CESM2 LENS exhibits a strong negative bias in Arctic-wide SIE. Despite this negative bias in CESM2 LENS, shorter reforecasts from a different configuration of CESM2 designed for S2S-scale prediction exhibit reduced ice loss compared to observations in localized regions during rapid episodic ice loss events (termed VRILEs).
The manuscript is well written and logically organized. However, I have two major issues with this study. First, I found it lacks overall coherency because the first part of the analysis evaluates a long-running climate simulation from CESM2 LENS while the second part evaluates shorter, initialized reforecasts using a different configuration of CESM2 for S2S-scale prediction. Sea ice biases in CESM2 LENS and CESM S2S likely differ and arise for different reasons. This makes it difficult to draw overarching conclusions about the ability of CESM2 to make subseasonal sea ice predictions. In addition, the analysis of VRILEs in the S2S reforecasts does not offer physical insight into why the predicted sea ice losses during these events differ from observations or from reforecasts with the CESM2 S2S configuration that uses the WACCM atmospheric model instead of CAM6. Addressing these issues would require redesigning the analysis, which is certainly possible but beyond scope of major revisions. Therefore, I do not recommend accepting the manuscript in its present state.
Decision: Soft Reject
Major Comments:
Specific/Minor Comments:
Line 4: It is unclear what the authors mean by “climate models underrepresent sea ice processes”. Please be more specific about what processes this refers to and what “underrepresent” means here (does it refer to models not melting or growing sea ice fast enough?)
Abstract: A brief overview of results from the evaluations of sea ice variability unrelated to VRILES is missing from the abstract and should be included (i.e., only VRILE analysis results appear in the abstract).
Lines 22-24: The timing of the first ice-free summer should not be used to judge the skill of a given climate model since an ice-free summer has not yet occurred in reality. Projections of the timing of the first ice-free summer are produced from climate models and have uncertainty associated with them. Therefore, a model that has an earlier or later ice-free summer projection than another model (or consensus of models) is not necessarily deficient because we don’t know when the first ice-free summer will actually occur. Thus, I recommend rewording or clarifying the second part of this sentence accordingly.
Lines 31, 172-3, 188, 268: The citations on these lines should appear entirely within parentheses.
Lines 77-79: If there are daily values from each individual ensemble member, shouldn’t it be possible to average all the members together to obtain a daily ensemble mean?
Line 83: I had to read this sentence several times to fully grasp what is meant by the model year in CESM2 LENS not directly relating to the calendar year. I recommend rewording this part to say something like “Since the sea ice extent in CESM2 LENS diverges from the observed sea ice extent, we compare the observed sea ice extent from NSIDC between 1980 and 2020 to earlier periods in the CESM2 forecast when the modeled sea ice extent is more comparable to the observations during this period.” I also felt this sentence should be moved later in this paragraph, to just before the description of the specific periods in the model being compared to NSIDC (~Line 90).
Line 119: I would replace “dynamical cores” here with “atmospheric model configurations” because the differences described here at not actually differences in dynamical cores (dynamical core refers to the numerical solver of the Navier-Stokes equations, not the physics packages or model grid configuration).
Lines 122-123: Do we know why the deeper atmosphere and interactive chemistry in WACCM6 results in better ice forecasts?
Line 127: It’s unclear what is meant by “does not provide enough time to evaluate the VRILE that occurred on 3 July”. Please provide a clearer explanation of what is meant here.
Line 136: Please add the phrase “with differences” before “becoming” on this line.
Section 2: A description of the basic aspects of CESM2 LENS (akin to the description of CESM2 S2S on Lines 112-118) is missing from this section. At minimum, the text should describe when the CESM2 LENS simulation is initialized (I assume 1850?) and what anthropogenic forcing scenario is used.
Line 150 (and elsewhere): Please replace all instances of “inner-quartile range” with the corrected name: “interquartile range”
Lines 166-168/Figure 4: Could the reduced variability of SIC within the marginal ice zone regions in CESM2 compared to NSIDC simply be due to these regions being ice-free in CESM2? This seems plausible given the large negative ice extent bias in CESM2, which would likely result in more ice-free areas at these lower latitudes. I recommend checking what the actual SIC values are in these regions in CESM2. It may also help to also include the 0.15 SIC contour from CESM2 on the panels in Fig. 4 so the reader has a better sense for where the ice edge is in the simulation (not just in the observations).
Line 170: Change “have positive values” to “there are positive values” to make this sentence grammatically correct.
Fig. 5: The caption says CESM2 VRILE sizes are indicated by blues, but when I zoom in on the smaller bars that presumably correspond to CESM2 in this figure, I see black, purple, red, orange and yellow segments, but no blue segments. I recommend either adjusting the color guidance in the caption to better correspond to what is shown in the figure, or changing the figure colors to truly have shades of blue for CESM2.
Fig. 6: This figure should include another row showing the 5-day change in SIC in NSIDC to illustrate what the observed ice loss associated with the VRILE was.
Lines 200-201: It’s not clear to me that the ice biases in a long running CESM2 LENS forecast should have any relationship to biases in short-term forecasts from CESM2 S2S, given the significant differences between the two model configurations (free running climate simulation vs. initialized short forecast). This gets back to my major comment 1 above.
Lines 231-234: Have the authors checked if the areas corresponding to the VRILE are actually open water (ice-free) in CESM at the beginning of the reforecast? The negative differences in Fig. 6d, along with what were known to be low sea ice concentrations in the Beaufort Sea at this time, suggests CESM may not even have sea ice here, which would change the interpretation of the results for this case.
Line 245: Change “unseasonal” to “unseasonable”.