Snow depth retrieval over Pan-Arctic sea ice (2012–2021) using multi-source data and machine learning models
Abstract. Snow depth is a critical climate indicator and a key parameter for Arctic sea ice retrieval. In this study, we retrieve pan-Arctic snow depth from 2012 to 2021 by integrating satellite altimetry, passive microwave brightness temperatures, and multi-source ground/airborne data. We employ four machine learning models—Light Gradient Boosting Machine (LightGBM), Multiple Linear Regression (MLR), Random Forest (RF), and Long Short-Term Memory (LSTM)—to leverage the complementary strengths of altimetry and microwave datasets while evaluating the performance of different machine learning (ML) architectures. Through permutation feature importance analysis, we identified that the 89 GHz polarization ratio has a significantly greater influence on snow depth retrieval over multi-year ice compared to that over first-year ice. Validation against Operation IceBridge and MOSAiC measurements reveals complementary strengths of snow retrieval among the models. The MLR model achieves the highest overall snow depth accuracy (root-means-square-error = 7.19 cm, correlation = 0.67 against OIB), while the LSTM demonstrates minimal mean bias of snow depth between satellite-based and in situ observations (1.98 cm against OIB; 0.30 cm against MOSAiC). All ML models exhibit robust generalization capabilities. Our retrieved snow depth products improved sea ice thickness estimation significantly, reducing bias between satellite-based and a standard climatology-based ice thickness product by nearly an order of magnitude. Our long-term snow products offer users a reliable, high-accuracy dataset for advancing Arctic energy budget modeling and sea ice studies.
Competing interests: At least one of the (co-)authors is a member of the editorial board of The Cryosphere.
Publisher's note: Copernicus Publications remains neutral with regard to jurisdictional claims made in the text, published maps, institutional affiliations, or any other geographical representation in this paper. While Copernicus Publications makes every effort to include appropriate place names, the final responsibility lies with the authors. Views expressed in the text are those of the authors and do not necessarily reflect the views of the publisher.
This manuscript presents a study on snow depth retrieval over Arctic sea ice using multi-source satellite data and four machine learning models. The authors compare MLR, RF, LGBM, and LSTM, evaluate their performance against independent OIB and MOSAiC observations, and demonstrate the practical utility of their retrievals by improving sea ice thickness estimates. The work is methodologically sound, and the validation framework is rigorous. The long-term (2012–2021) snow depth product is an important dataset for Arctic climate studies.
Major Comments
1. The study employs four models (MLR, RF, LGBM, and LSTM) and reveals complex complementary performance—MLR performs best against OIB, while LSTM achieves the lowest bias against MOSAiC. However, the manuscript does not provide explicit guidance on how readers should select among these models for different application scenarios.
2. The MW99 product exhibits very large bias (7.04 cm, Table 3), far exceeding that of all other products. The manuscript notes that MW99 is outdated, but does not elaborate on the physical reasons for its poor performance.
3. The Introduction mentions both Rostosky et al. (2018) and Li et al. (2024) as existing snow depth products, but their roles in this study are different and may confuse readers.
4. The models are trained on data from 2018–2021 (the ASD product) and applied to retrieve snow depth back to 2012. While the temporal extrapolation test (Section 4.3.1) demonstrates generalizability, it remains unclear whether the training period (2018–2021) is representative of the longer period. Did the snow–microwave relationships remain stable across this entire period, or have there been systematic shifts due to changing Arctic conditions?
5. The manuscript's methodological contribution lies in the combination of systematic model intercomparison, rather than simply "applying four ML models to snow depth retrieval." However, the current framing may give readers the impression that this is just another ML-based retrieval study.
6. The authors should consider whether it is necessary to include a discussion of its role in ice thickness retrieval, as this study mainly focuses on snow depth retrieval.
Minor Comments:
1. The abstract states that the validation "reveals complementary strengths of snow retrieval among the models," but does not specify what these complementary strengths are.
2. Line 60, in the sentence introducing data fusion approaches, the term "bridges" should be changed.
3. Line 112, the authors note that 89 GHz is "known" to be affected by atmospheric water vapor but do not cite a reference to support this claim.
4. Table 3 shows seven products compared against OIB, but the layout is dense and could be made more readable.
5. Line 548, "This scale issue affects all products equally". This statement is slightly too strong. While it is true that all products suffer from scale mismatch, the degree of impact may differ. Some models (e.g., LSTM) might be more or less sensitive to the scale mismatch than others.
6. Line 574, "A black line represents the point of flawless agreement with OIB standards". "flawless agreement" is informal.
7. In section 5.1, the uncertainty analysis uses 1000 Monte Carlo iterations for non-linear models. The manuscript does not justify the choice of 1000 iterations.
8. Lines 674-676, the manuscript contains several long, multi-clause sentences that could be broken up for improved readability.
9. Figures 5 and 6 have relatively brief captions that do not fully explain what is being shown.
10. The language still needs further polishing, and consistency should be improved, such as ensuring that symbols are used consistently throughout the manuscript.