the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
Snow depth retrieval over Pan-Arctic sea ice (2012–2021) using multi-source data and machine learning models
Abstract. Snow depth is a critical climate indicator and a key parameter for Arctic sea ice retrieval. In this study, we retrieve pan-Arctic snow depth from 2012 to 2021 by integrating satellite altimetry, passive microwave brightness temperatures, and multi-source ground/airborne data. We employ four machine learning models—Light Gradient Boosting Machine (LightGBM), Multiple Linear Regression (MLR), Random Forest (RF), and Long Short-Term Memory (LSTM)—to leverage the complementary strengths of altimetry and microwave datasets while evaluating the performance of different machine learning (ML) architectures. Through permutation feature importance analysis, we identified that the 89 GHz polarization ratio has a significantly greater influence on snow depth retrieval over multi-year ice compared to that over first-year ice. Validation against Operation IceBridge and MOSAiC measurements reveals complementary strengths of snow retrieval among the models. The MLR model achieves the highest overall snow depth accuracy (root-means-square-error = 7.19 cm, correlation = 0.67 against OIB), while the LSTM demonstrates minimal mean bias of snow depth between satellite-based and in situ observations (1.98 cm against OIB; 0.30 cm against MOSAiC). All ML models exhibit robust generalization capabilities. Our retrieved snow depth products improved sea ice thickness estimation significantly, reducing bias between satellite-based and a standard climatology-based ice thickness product by nearly an order of magnitude. Our long-term snow products offer users a reliable, high-accuracy dataset for advancing Arctic energy budget modeling and sea ice studies.
Competing interests: At least one of the (co-)authors is a member of the editorial board of The Cryosphere.
Publisher's note: Copernicus Publications remains neutral with regard to jurisdictional claims made in the text, published maps, institutional affiliations, or any other geographical representation in this paper. While Copernicus Publications makes every effort to include appropriate place names, the final responsibility lies with the authors. Views expressed in the text are those of the authors and do not necessarily reflect the views of the publisher.- Preprint
(10173 KB) - Metadata XML
- BibTeX
- EndNote
Status: open (until 18 Oct 2026)
-
RC1: 'Comment on egusphere-2026-2504', Anonymous Referee #1, 16 Aug 2026
reply
-
AC1: 'Reply on RC1', Mengmeng Li, 27 Aug 2026
reply
Thank you for taking your time for the review and providing the helpful comment, please see the attached comment.
-
AC1: 'Reply on RC1', Mengmeng Li, 27 Aug 2026
reply
-
RC2: 'Comment on egusphere-2026-2504', Anonymous Referee #2, 15 Sep 2026
reply
This manuscript uses AMSR2 passive-microwave brightness temperatures, satellite-altimetry-derived snow depth, SnowModel-LG snow density, and OIB and MOSAiC observations to retrieve pan-Arctic sea-ice snow depth. It compares four machine-learning approaches (multiple linear regression (MLR), random forest (RF), LightGBM, and long short-term memory (LSTM)) and further evaluates the effect of the retrieved snow depth on sea-ice-thickness (SIT) estimation. The topic is within the scope of The Cryosphere. The use of multiple data sources, the comparison of different model architectures, and the feature-importance analysis are potentially valuable. I recommend accepting this article after moderate revisions.
Major comments:
1. There appear to be inconsistencies between the table headings, the numerical values, and the descriptions in the text. For example, Table 3 contains an additional “ALL” column, whereas the numerical rows appear to contain seven values corresponding to LGBM, MLR, RF, LSTM, SM, Rost, and MW99.
2. The manuscript interprets the difference between the internal testing results and the independent OIB/MOSAiC results as a “classic indicator of overfitting.” This is possible, but the difference may also result from several other factors, including differences between the ASD and OIB/MOSAiC error structures, spatial and temporal sampling, scale mismatch, and differences in snow-depth distributions. The current experiments do not fully isolate overfitting from dataset shift or reference-product uncertainty.
3. The Results repeatedly interpret high importance of PR(89), PR(06), and other features as direct evidence of specific snow microstructural processes, such as depth-hoar development and peak metamorphism. Permutation importance indicates the predictive contribution of a feature within a particular model and dataset. It does not, by itself, establish a unique physical mechanism, especially in the presence of highly correlated GR and PR variables and possible atmospheric effects at 89 GHz.
4. Please replace causal or overly definitive statements such as “directly tracking,” “demonstrates,” “signifies,” and “providing direct evidence” with more cautious language. The authors should also mention the effect of feature collinearity on permutation-importance rankings.
5. The correlation coefficients against MOSAiC are very low for all products, and the LSTM correlation is reported as 0.13 with a p-value of 0.08. The manuscript correctly notes the point-to-grid scale mismatch, but it should not state that LSTM “most accurately captures the mean snow depth state” solely on the basis of a low bias and MAE, particularly when the differences in MAE among the products are small.
6. The SIT results show a much smaller bias for the machine-learning-based estimates than for FT4 using MW99. This is an important result, but the conclusion that this “confirms” MW99 as the major source of operational SIT error may be too strong. Differences may also result from sea-ice density, snow density, freeboard processing, ice-type classification, matchup selection, and correlations between the snow depth and SIT reference products.
Minor comments:
1. The title should consistently use “pan-Arctic.” such as CryoSat-2, AMSR2, and sea ice thickness should be used consistently throughout the manuscript.
2. The manuscript requires comprehensive language and copy editing. Several language and formatting errors are present, including:“root-means-square-error,” which should be “root-mean-square error”; “horizonal,” which should be “horizontal”; “addictive linear interactions,” which should be “additive linear interactions”; “oceanatmosphere,” which should be “ocean–atmosphere”; “year-around,” which should be “year-round”; and “Cryosat2,” which should be consistently written as “CryoSat-2.”
3. Suggested revisions to key statements in the abstract: Original statement: The 89 GHz polarization ratio has a significantly greater influence on snow-depth retrieval over multi-year ice. Suggested revision: The 89 GHz polarization ratio showed greater predictive importance over multi-year ice than over first-year ice.Original statement: Our long-term snow products offer users a reliable, high-accuracy dataset. Suggested revision: The resulting long-term snow-depth products provide an observationally evaluated dataset for Arctic sea-ice and snow studies, with uncertainties that vary across ice types, seasons, and the extrapolation period.
4. Use consistent model naming and capitalization. Please use one form consistently: LightGBM rather than alternating between LGBM and LightGBM, unless LGBM is explicitly defined as the abbreviation. Similarly, use consistent forms for MLR, RF, LSTM, OIB, MOSAiC, FYI, MYI, and SIT.
5. The manuscript contains several typographical and language problems, including:“generalisability” and “generalization” should be standardized according to the journal’s preferred English style;
6. In Figure 5 caption, “Artic” should be “Arctic”.
7. Line 676, “retrievals..” contains an extra period;
8. Line 490, “The Rost product” should use lower-case “the” when it appears in the middle of a sentence.
9. Line 514, “The value in parentheses represent” should be “The value in parentheses represents”.
10. The captions of Figures 2-7 should provide sufficient methodological information. The figure captions should indicate whether the plotted values are daily or monthly, whether the distributions are based on all collocated samples, and whether the comparisons use the same spatial and temporal mask for all products. Figure 4 should also clarify the meaning of the “point of flawless agreement.”
Viewed
| HTML | XML | Total | BibTeX | EndNote | |
|---|---|---|---|---|---|
| 131 | 69 | 20 | 220 | 14 | 15 |
- HTML: 131
- PDF: 69
- XML: 20
- Total: 220
- BibTeX: 14
- EndNote: 15
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
This manuscript presents a study on snow depth retrieval over Arctic sea ice using multi-source satellite data and four machine learning models. The authors compare MLR, RF, LGBM, and LSTM, evaluate their performance against independent OIB and MOSAiC observations, and demonstrate the practical utility of their retrievals by improving sea ice thickness estimates. The work is methodologically sound, and the validation framework is rigorous. The long-term (2012–2021) snow depth product is an important dataset for Arctic climate studies.
Major Comments
1. The study employs four models (MLR, RF, LGBM, and LSTM) and reveals complex complementary performance—MLR performs best against OIB, while LSTM achieves the lowest bias against MOSAiC. However, the manuscript does not provide explicit guidance on how readers should select among these models for different application scenarios.
2. The MW99 product exhibits very large bias (7.04 cm, Table 3), far exceeding that of all other products. The manuscript notes that MW99 is outdated, but does not elaborate on the physical reasons for its poor performance.
3. The Introduction mentions both Rostosky et al. (2018) and Li et al. (2024) as existing snow depth products, but their roles in this study are different and may confuse readers.
4. The models are trained on data from 2018–2021 (the ASD product) and applied to retrieve snow depth back to 2012. While the temporal extrapolation test (Section 4.3.1) demonstrates generalizability, it remains unclear whether the training period (2018–2021) is representative of the longer period. Did the snow–microwave relationships remain stable across this entire period, or have there been systematic shifts due to changing Arctic conditions?
5. The manuscript's methodological contribution lies in the combination of systematic model intercomparison, rather than simply "applying four ML models to snow depth retrieval." However, the current framing may give readers the impression that this is just another ML-based retrieval study.
6. The authors should consider whether it is necessary to include a discussion of its role in ice thickness retrieval, as this study mainly focuses on snow depth retrieval.
Minor Comments:
1. The abstract states that the validation "reveals complementary strengths of snow retrieval among the models," but does not specify what these complementary strengths are.
2. Line 60, in the sentence introducing data fusion approaches, the term "bridges" should be changed.
3. Line 112, the authors note that 89 GHz is "known" to be affected by atmospheric water vapor but do not cite a reference to support this claim.
4. Table 3 shows seven products compared against OIB, but the layout is dense and could be made more readable.
5. Line 548, "This scale issue affects all products equally". This statement is slightly too strong. While it is true that all products suffer from scale mismatch, the degree of impact may differ. Some models (e.g., LSTM) might be more or less sensitive to the scale mismatch than others.
6. Line 574, "A black line represents the point of flawless agreement with OIB standards". "flawless agreement" is informal.
7. In section 5.1, the uncertainty analysis uses 1000 Monte Carlo iterations for non-linear models. The manuscript does not justify the choice of 1000 iterations.
8. Lines 674-676, the manuscript contains several long, multi-clause sentences that could be broken up for improved readability.
9. Figures 5 and 6 have relatively brief captions that do not fully explain what is being shown.
10. The language still needs further polishing, and consistency should be improved, such as ensuring that symbols are used consistently throughout the manuscript.