the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
Snow depth retrieval over Pan-Arctic sea ice (2012–2021) using multi-source data and machine learning models
Abstract. Snow depth is a critical climate indicator and a key parameter for Arctic sea ice retrieval. In this study, we retrieve pan-Arctic snow depth from 2012 to 2021 by integrating satellite altimetry, passive microwave brightness temperatures, and multi-source ground/airborne data. We employ four machine learning models—Light Gradient Boosting Machine (LightGBM), Multiple Linear Regression (MLR), Random Forest (RF), and Long Short-Term Memory (LSTM)—to leverage the complementary strengths of altimetry and microwave datasets while evaluating the performance of different machine learning (ML) architectures. Through permutation feature importance analysis, we identified that the 89 GHz polarization ratio has a significantly greater influence on snow depth retrieval over multi-year ice compared to that over first-year ice. Validation against Operation IceBridge and MOSAiC measurements reveals complementary strengths of snow retrieval among the models. The MLR model achieves the highest overall snow depth accuracy (root-means-square-error = 7.19 cm, correlation = 0.67 against OIB), while the LSTM demonstrates minimal mean bias of snow depth between satellite-based and in situ observations (1.98 cm against OIB; 0.30 cm against MOSAiC). All ML models exhibit robust generalization capabilities. Our retrieved snow depth products improved sea ice thickness estimation significantly, reducing bias between satellite-based and a standard climatology-based ice thickness product by nearly an order of magnitude. Our long-term snow products offer users a reliable, high-accuracy dataset for advancing Arctic energy budget modeling and sea ice studies.
Competing interests: At least one of the (co-)authors is a member of the editorial board of The Cryosphere.
Publisher's note: Copernicus Publications remains neutral with regard to jurisdictional claims made in the text, published maps, institutional affiliations, or any other geographical representation in this paper. While Copernicus Publications makes every effort to include appropriate place names, the final responsibility lies with the authors. Views expressed in the text are those of the authors and do not necessarily reflect the views of the publisher.- Preprint
(10173 KB) - Metadata XML
- BibTeX
- EndNote
Status: open (until 18 Oct 2026)
-
RC1: 'Comment on egusphere-2026-2504', Anonymous Referee #1, 16 Aug 2026
reply
-
AC1: 'Reply on RC1', Mengmeng Li, 27 Aug 2026
reply
Thank you for taking your time for the review and providing the helpful comment, please see the attached comment.
-
AC1: 'Reply on RC1', Mengmeng Li, 27 Aug 2026
reply
-
RC2: 'Comment on egusphere-2026-2504', Anonymous Referee #2, 15 Sep 2026
reply
This manuscript uses AMSR2 passive-microwave brightness temperatures, satellite-altimetry-derived snow depth, SnowModel-LG snow density, and OIB and MOSAiC observations to retrieve pan-Arctic sea-ice snow depth. It compares four machine-learning approaches (multiple linear regression (MLR), random forest (RF), LightGBM, and long short-term memory (LSTM)) and further evaluates the effect of the retrieved snow depth on sea-ice-thickness (SIT) estimation. The topic is within the scope of The Cryosphere. The use of multiple data sources, the comparison of different model architectures, and the feature-importance analysis are potentially valuable. I recommend accepting this article after moderate revisions.
Major comments:
1. There appear to be inconsistencies between the table headings, the numerical values, and the descriptions in the text. For example, Table 3 contains an additional “ALL” column, whereas the numerical rows appear to contain seven values corresponding to LGBM, MLR, RF, LSTM, SM, Rost, and MW99.
2. The manuscript interprets the difference between the internal testing results and the independent OIB/MOSAiC results as a “classic indicator of overfitting.” This is possible, but the difference may also result from several other factors, including differences between the ASD and OIB/MOSAiC error structures, spatial and temporal sampling, scale mismatch, and differences in snow-depth distributions. The current experiments do not fully isolate overfitting from dataset shift or reference-product uncertainty.
3. The Results repeatedly interpret high importance of PR(89), PR(06), and other features as direct evidence of specific snow microstructural processes, such as depth-hoar development and peak metamorphism. Permutation importance indicates the predictive contribution of a feature within a particular model and dataset. It does not, by itself, establish a unique physical mechanism, especially in the presence of highly correlated GR and PR variables and possible atmospheric effects at 89 GHz.
4. Please replace causal or overly definitive statements such as “directly tracking,” “demonstrates,” “signifies,” and “providing direct evidence” with more cautious language. The authors should also mention the effect of feature collinearity on permutation-importance rankings.
5. The correlation coefficients against MOSAiC are very low for all products, and the LSTM correlation is reported as 0.13 with a p-value of 0.08. The manuscript correctly notes the point-to-grid scale mismatch, but it should not state that LSTM “most accurately captures the mean snow depth state” solely on the basis of a low bias and MAE, particularly when the differences in MAE among the products are small.
6. The SIT results show a much smaller bias for the machine-learning-based estimates than for FT4 using MW99. This is an important result, but the conclusion that this “confirms” MW99 as the major source of operational SIT error may be too strong. Differences may also result from sea-ice density, snow density, freeboard processing, ice-type classification, matchup selection, and correlations between the snow depth and SIT reference products.
Minor comments:
1. The title should consistently use “pan-Arctic.” such as CryoSat-2, AMSR2, and sea ice thickness should be used consistently throughout the manuscript.
2. The manuscript requires comprehensive language and copy editing. Several language and formatting errors are present, including:“root-means-square-error,” which should be “root-mean-square error”; “horizonal,” which should be “horizontal”; “addictive linear interactions,” which should be “additive linear interactions”; “oceanatmosphere,” which should be “ocean–atmosphere”; “year-around,” which should be “year-round”; and “Cryosat2,” which should be consistently written as “CryoSat-2.”
3. Suggested revisions to key statements in the abstract: Original statement: The 89 GHz polarization ratio has a significantly greater influence on snow-depth retrieval over multi-year ice. Suggested revision: The 89 GHz polarization ratio showed greater predictive importance over multi-year ice than over first-year ice.Original statement: Our long-term snow products offer users a reliable, high-accuracy dataset. Suggested revision: The resulting long-term snow-depth products provide an observationally evaluated dataset for Arctic sea-ice and snow studies, with uncertainties that vary across ice types, seasons, and the extrapolation period.
4. Use consistent model naming and capitalization. Please use one form consistently: LightGBM rather than alternating between LGBM and LightGBM, unless LGBM is explicitly defined as the abbreviation. Similarly, use consistent forms for MLR, RF, LSTM, OIB, MOSAiC, FYI, MYI, and SIT.
5. The manuscript contains several typographical and language problems, including:“generalisability” and “generalization” should be standardized according to the journal’s preferred English style;
6. In Figure 5 caption, “Artic” should be “Arctic”.
7. Line 676, “retrievals..” contains an extra period;
8. Line 490, “The Rost product” should use lower-case “the” when it appears in the middle of a sentence.
9. Line 514, “The value in parentheses represent” should be “The value in parentheses represents”.
10. The captions of Figures 2-7 should provide sufficient methodological information. The figure captions should indicate whether the plotted values are daily or monthly, whether the distributions are based on all collocated samples, and whether the comparisons use the same spatial and temporal mask for all products. Figure 4 should also clarify the meaning of the “point of flawless agreement.”
-
AC2: 'Reply on RC2', Mengmeng Li, 30 Sep 2026
reply
This manuscript uses AMSR2 passive-microwave brightness temperatures, satellite-altimetry-derived snow depth, SnowModel-LG snow density, and OIB and MOSAiC observations to retrieve pan-Arctic sea-ice snow depth. It compares four machine-learning approaches (multiple linear regression (MLR), random forest (RF), LightGBM, and long short-term memory (LSTM)) and further evaluates the effect of the retrieved snow depth on sea-ice-thickness (SIT) estimation. The topic is within the scope of The Cryosphere. The use of multiple data sources, the comparison of different model architectures, and the feature-importance analysis are potentially valuable. I recommend accepting this article after moderate revisions.
We thank the reviewer for the positive assessment of the manuscript’s scope, multi-source data integration, model comparison, and feature-importance analysis. We also appreciate the recommendation for moderate revision and the constructive comments, which have helped us improve the clarity, consistency, and robustness of the manuscript. In the revised manuscript, we have (i) corrected the inconsistencies in Table 3 and its caption; (ii) replaced the overly specific attribution of the internal-independent validation discrepancy to overfitting with a more cautious discussion of dataset shift, sampling differences, scale mismatch, reference-data uncertainty, and transferability; (iii) revised the interpretation of permutation importance to avoid causal or unique physical claims and added an explicit discussion of feature collinearity; (iv) softened the MOSAiC-based conclusions by acknowledging the low correlations and small MAE differences and by avoiding any definitive ranking of product accuracy; (v) revised the SIT discussion so that the results are interpreted as sensitivity to the snow-depth input rather than confirmation that MW99 is the dominant error source; and (vi) standardized terminology, model naming, abbreviations, language, and figure/table captions throughout the manuscript. Point-by-point responses are provided below.
Major comments:
- There appear to be inconsistencies between the table headings, the numerical values, and the descriptions in the text. For example, Table 3 contains an additional “ALL” column, whereas the numerical rows appear to contain seven values corresponding to LGBM, MLR, RF, LSTM, SM, Rost, and MW99.
Thank you for identifying this inconsistency. We carefully checked Table 3, the associated table caption, and all references to the table in the main text. The “ALL” column was removed, so that the number and order of column headings are now fully consistent with the numerical entries. The revised table now presents results in the following order: LGBM, MLR, RF, LSTM, SM, Rost, and MW99 in the revised manuscript.
- The manuscript interprets the difference between the internal testing results and the independent OIB/MOSAiC results as a “classic indicator of overfitting.” This is possible, but the difference may also result from several other factors, including differences between the ASD and OIB/MOSAiC error structures, spatial and temporal sampling, scale mismatch, and differences in snow-depth distributions. The current experiments do not fully isolate overfitting from dataset shift or reference-product uncertainty.
We agree. Our original wording attributed the performance difference too specifically to overfitting. Although the discrepancy between the internal testing results and the independent OIB/MOSAiC evaluations may be consistent with limited out-of-sample generalization, potentially including overfitting, the available experiments cannot distinguish this explanation from dataset shift and uncertainty in the reference datasets.
We have therefore revised the text to state that the reduced performance in the independent evaluations may reflect a combination of factors, including: (i) differences in the measurement principles and error characteristics of the ASD, OIB, and MOSAiC snow-depth datasets; (ii) differences in spatial and temporal sampling; (iii) point-to-grid and footprint-scale mismatches; (iv) differences in the snow-depth distributions represented in the datasets; and (v) potential overfitting or limited transferability beyond the training domain. We now explicitly state that the present analysis does not isolate the relative contributions of these factors.
We revised the subsection “Synthesis of validation results and implications for model selection” in the Results section to avoid attributing the discrepancy between the internal testing and independent OIB/MOSAiC evaluations solely to overfitting. We also revised the corresponding interpretation in the “Quantification and analysis of retrieval uncertainties” subsection and the Conclusions. The revised text now explicitly discusses dataset shift, differences in reference-data error structures, spatial and temporal sampling, snow-depth distributions, and scale mismatch, and states that the present experiments do not isolate the contribution of overfitting from these factors.
- The Results repeatedly interpret high importance of PR(89), PR(06), and other features as direct evidence of specific snow microstructural processes, such as depth-hoar development and peak metamorphism. Permutation importance indicates the predictive contribution of a feature within a particular model and dataset. It does not, by itself, establish a unique physical mechanism, especially in the presence of highly correlated GR and PR variables and possible atmospheric effects at 89 GHz.
We agree with this important distinction. We have revised the interpretation of permutation importance throughout the manuscript. The revised text now clarifies that permutation importance measures the predictive contribution of a feature within the specific model, training dataset, and evaluation framework. It does not independently establish a unique physical mechanism.
We have removed statements that interpreted the high importance of PR(89), PR(06), or other predictors as direct evidence of depth-hoar development, peak metamorphism, or other specific snow microstructural processes. Instead, these variables are described as statistically informative predictors of snow depth in the present retrieval framework. We also added a statement noting that the importance rankings may reflect overlapping information among correlated GR and PR variables, as well as combined effects of snow properties, surface and ice conditions, and, particularly for the 89 GHz channel, atmospheric influences.
We revised the subsection “Optimizing input parameters for machine learning algorithms” in the Results section, the relevant discussion of feature importance, and the Conclusions. The revised manuscript now interprets permutation importance as model- and dataset-dependent predictive relevance rather than direct evidence of a unique physical mechanism.
- Please replace causal or overly definitive statements such as “directly tracking,” “demonstrates,” “signifies,” and “providing direct evidence” with more cautious language. The authors should also mention the effect of feature collinearity on permutation-importance rankings.
We appreciate this suggestion and have systematically revised the wording throughout the manuscript. Expressions implying direct causality or definitive physical attribution, including “directly tracking,” “demonstrates,” “signifies,” and “providing direct evidence,” have been replaced with more cautious language, such as “is associated with,” “is consistent with,” “suggests,” “may reflect,” “contributes predictive information,” and “is statistically informative in the present model.”
In addition, we added an explicit statement on feature collinearity and its implications for permutation importance. Several GR and PR predictors are derived from overlapping brightness temperature channels and are therefore correlated. Their permutation importance values may reflect shared or redundant predictive information rather than the independent contribution of an individual variable. Consequently, the importance ranking of one predictor may be reduced when correlated predictors remain available to the model, whereas its ranking may increase when those correlated predictors are absent or less informative. We therefore interpret the rankings as model- and dataset-dependent measures of predictive relevance, rather than unique measures of physical relevance or causal influence.
We added a statement on predictor collinearity and the interpretation of permutation importance in the Methods section and revised the relevant feature-importance interpretations in the Results and Conclusions.
- The correlation coefficients against MOSAiC are very low for all products, and the LSTM correlation is reported as 0.13 with a p-value of 0.08. The manuscript correctly notes the point-to-grid scale mismatch, but it should not state that LSTM “most accurately captures the mean snow depth state” solely on the basis of a low bias and MAE, particularly when the differences in MAE among the products are small.
We thank the reviewer for this important comment. We agree that the low correlations with the MOSAiC observations do not support a definitive ranking of product accuracy or a strong statement that LSTM captures the mean snow-depth state most accurately. In particular, the LSTM correlation is low and not statistically significant (R = 0.13, p = 0.08). Although LSTM has the smallest bias (0.30 cm) and a low MAE (5.57 cm), its MAE differs only slightly from those of RF (5.60 cm), SnowModel-LG (5.51 cm), and MLR (5.69 cm). Therefore, these metrics alone are insufficient to establish that LSTM is more accurate than the other products.
We have revised the text in the MOSAiC validation, the synthesis of validation results, and the Conclusions. The revised manuscript now states that LSTM showed the smallest mean bias and an MAE comparable to those of several other products. We further clarify that the weak correlations, together with the point-to-grid scale mismatch and small differences in MAE, prevent a robust ranking of product accuracy based on the MOSAiC comparison alone. The MOSAiC evaluation is therefore interpreted as an assessment of broad agreement in mean error metrics, rather than evidence that any product accurately reproduces point-scale snow-depth variability or local mean conditions.
- The SIT results show a much smaller bias for the machine-learning-based estimates than for FT4 using MW99. This is an important result, but the conclusion that this “confirms” MW99 as the major source of operational SIT error may be too strong. Differences may also result from sea-ice density, snow density, freeboard processing, ice-type classification, matchup selection, and correlations between the snow depth and SIT reference products.
We thank the reviewer for this important comment. We agree that the lower bias of the machine-learning-based SIT estimates relative to FT4 should not be interpreted as confirmation that MW99 is the major source of error in the operational FT4 SIT product. Although the SIT sensitivity experiment was designed to isolate the effect of replacing the MW99 snow-depth input while holding the density assumptions and CryoSat-2 freeboard input constant, the comparison does not account for all sources of uncertainty in operational SIT retrievals.
In particular, uncertainties associated with spatial and temporal matchup procedures, and possible dependencies between the snow-depth and SIT reference products may also affect the observed differences. We have therefore revised the manuscript to state that the results demonstrate a strong sensitivity of SIT bias to the choice of snow-depth input under the assumptions adopted in this study, rather than attributing the FT4 bias uniquely or predominantly to MW99. We also clarify that, although the machine-learning-based SIT estimates substantially reduce mean bias, their RMSE values (0.47–0.51 m) are comparable to that of FT4 (0.48 m). Thus, the experiment primarily indicates an improvement in systematic bias rather than a uniform improvement in all SIT error metrics.
Minor comments:
- The title should consistently use “pan-Arctic.” such as CryoSat-2, AMSR2, and sea ice thickness should be used consistently throughout the manuscript.
We agree. The title and manuscript have been revised to use “pan-Arctic” consistently. We also standardized the terminology and spelling of “CryoSat-2,” “AMSR2,” “sea ice thickness,” and related expressions throughout the manuscript.
- The manuscript requires comprehensive language and copy editing. Several language and formatting errors are present, including:“root-means-square-error,” which should be “root-mean-square error”; “horizonal,” which should be “horizontal”; “addictive linear interactions,” which should be “additive linear interactions”; “oceanatmosphere,” which should be “ocean–atmosphere”; “year-around,” which should be “year-round”; and “Cryosat2,” which should be consistently written as “CryoSat-2.”
We have conducted a comprehensive language and copy edit of the manuscript. The specific terms identified by the reviewer have been corrected.
- Suggested revisions to key statements in the abstract: Original statement: The 89 GHz polarization ratio has a significantly greater influence on snow-depth retrieval over multi-year ice. Suggested revision: The 89 GHz polarization ratio showed greater predictive importance over multi-year ice than over first-year ice.Original statement: Our long-term snow products offer users a reliable, high-accuracy dataset. Suggested revision: The resulting long-term snow-depth products provide an observationally evaluated dataset for Arctic sea-ice and snow studies, with uncertainties that vary across ice types, seasons, and the extrapolation period.
We agree and have revised the abstract to avoid overstating the physical interpretation and reliability of the products.
The original statement:
The 89 GHz polarization ratio has a significantly greater influence on snow-depth retrieval over multi-year ice compared to that over first-year ice.
has been replaced by:
The 89 GHz polarization ratio showed greater predictive importance over multi-year ice than over first-year ice.
The original statement:
Our long-term snow products offer users a reliable, high-accuracy dataset for advancing Arctic energy budget modeling and sea ice studies.
has been replaced by:
The resulting long-term snow-depth products provide an observationally evaluated dataset for Arctic sea-ice and snow studies, with uncertainties that vary across ice types, seasons, and the extrapolation period.
- Use consistent model naming and capitalization. Please use one form consistently: LightGBM rather than alternating between LGBM and LightGBM, unless LGBM is explicitly defined as the abbreviation. Similarly, use consistent forms for MLR, RF, LSTM, OIB, MOSAiC,FYI, MYI, and SIT.
We have standardized model names throughout the manuscript. At the first occurrence, the following definitions are now used:
Light Gradient Boosting Machine (LGBM), Multiple Linear Regression (MLR), Random Forest (RF), and Long Short-Term Memory (LSTM).
We use LGBM consistently thereafter, including in the tables and figure captions. The abbreviation “LightGBM” has been replaced by “LGBM” throughout the manuscript.
The following abbreviations have also been standardized: LGBM, MLR, RF, LSTM, OIB, MOSAiC, FYI, MYI, SIT.
We use “first-year ice (FYI)” and “multi-year ice (MYI)” at their first occurrence and use FYI and MYI consistently thereafter.
- The manuscript contains several typographical and language problems, including:“generalisability” and “generalization” should be standardized according to the journal’s preferred English style;
We have standardized the manuscript to American English and use generalization consistently. Related terms have also been harmonized, including: generalization capability, generalization performance.
- In Figure 5 caption, “Artic” should be “Arctic”.
The caption has been corrected.
- Line 676, “retrievals..” contains an extra period;
The extra period has been removed.
- Line 490, “The Rost product” should use lower-case “the” when it appears in the middle of a sentence.
The sentence has been corrected.
- Line 514, “The value in parentheses represent” should be “The value in parentheses represents”.
The sentence has been corrected.
- The captions of Figures 2–7 should provide sufficient methodological information. The figure captions should indicate whether the plotted values are daily or monthly, whether the distributions are based on all collocated samples, and whether the comparisons use the same spatial and temporal mask for all products. Figure 4 should also clarify the meaning of the “point of flawless agreement.”
We agree. Because two new figures (Figs. 2 and 3) were added in Section 4.1, the original Figures 2–7 have been renumbered as Figures 4–9. The captions of Figures 4–9 have been revised to specify the temporal resolution, the number of collocated samples, the use of common spatial and temporal masks, and the meaning of the reference line in Figure 4.
The revised captions are provided below.
Revised Figure 4 caption
Figure 4: Probability distributions of monthly snow depth from the OIB observations and seven snow depth products: (a) LGBM, (b) MLR, (c) RF, (d) LSTM, (e) SM, (f) Rost, and (g) MW99. The comparison uses 3,167 collocated 25-km grid cell samples from the common spatial and temporal mask for which all products and the OIB observations are available.
Revised Figure 5 caption
Figure 5: Comparison of monthly gridded snow depth estimates from LGBM, MLR, RF, LSTM, SM, and MW99 with MOSAiC snow depth observations from October 2019 to April 2021. The comparison includes 192 collocated samples obtained using the same spatial and temporal mask for all products. The x-axis represents the number of valid grid points. The y-axis represents the corresponding gridded snow depth estimates. The Rost product is excluded because its temporal coverage does not overlap with the MOSAiC period.
Revised Figure 6 caption
Figure 6: Comparison of daily snow depth estimates from LGBM, MLR, RF, and LSTM with daily OIB snow depth observations. Panels (a–d) show the corresponding probability distributions, and panels (e–h) show the differences between the retrieved and OIB snow depths for 5,133 collocated samples. The same spatial and temporal mask was applied to all four products. In panels (e–h), the black line denotes zero retrieval error, corresponding to exact agreement between the estimated and OIB snow depths.
Revised Figure 7 caption
Figure 7: Spatial distributions of monthly mean Arctic snow depth (m) during the 2012–2013 growth season (October–April) derived from the four model retrievals (LGBM, MLR, RF, and LSTM). The maps were generated using the valid sea ice mask (SIC ≥ 80%) and the corresponding ice type classifications for each month.
Revised Figure 8 caption
Figure 8: Daily snow depth retrievals on the 15th of each month during the 2012–2013 growth season from the four models (LGBM, MLR, RF, and LSTM). The maps use the same sea ice mask and ice type classifications as Figure 7.
Revised Figure 9 caption
Figure 9: Distributions of the differences between directly retrieved monthly snow depth and monthly mean snow depth calculated from the corresponding daily retrievals during the 2012–2013 growth season. The comparison is performed for the same valid 25-km grid cells and months for all model. The dashed and solid lines in each violin plot represent the median and mean differences, respectively.
-
AC2: 'Reply on RC2', Mengmeng Li, 30 Sep 2026
reply
-
RC3: 'Comment on egusphere-2026-2504', Anonymous Referee #3, 20 Sep 2026
reply
This study uses four algorithms, including MLR, RF, LightGBM, and LSTM, to retrieve snow depth on Arctic sea ice and further estimate sea ice thickness. However, the authors do not identify an optimal method, and the study appears to produce four separate datasets. The authors should clarify the applicability of each algorithm in the abstract, conclusions, and other relevant sections. Such recommendations should not be based solely on the overall comparison with OIB, but should also consider differences in ice type, season, and temporal resolution. In addition, several figures, tables, and descriptions require further improvement.
Specific comments:
1. Please provide the full name of OIB when it first appears in the abstract.
2. The abstract should be reorganized to describe the datasets, retrieval algorithms, feature selection, temporal coverage of the products, and main conclusions more clearly. At present, the abstract places too much emphasis on the comparison between MLR and OIB. This result only represents model performance for a specific validation dataset and set of evaluation metrics and does not fully reflect the main contribution of the study. Because the model rankings differ between the internal test and independent validation, the general expression "highest accuracy" should be avoided. The performance of the four models and the main characteristics of the four products should be summarized more objectively.
3. Section 4.1 discusses the importance of different frequencies and polarization parameters for different months and ice types, but no corresponding figure or table is provided. Please add a figure or table to support these descriptions and list the final parameters selected as inputs for each model.
4. Please clearly state which datasets are represented by SM, Rost, and MW99 in Lines 471–472.
5. Please further discuss the main differences and advantages of the products developed in this study compared with Rost and other existing products, not only in terms of model validation results but also with respect to spatial and temporal resolution, temporal coverage, seasonal coverage, and other relevant aspects.
6. Lines 531–532 suggest that some models accurately capture the snow depth at the MOSAiC locations based on their relatively small bias and MAE. However, the correlation coefficients of all products are very low. If the low correlations are mainly caused by the scale difference between point measurements and gridded products, the authors should clarify which aspects of product performance can be evaluated using MOSAiC data and weaken the conclusion about "accurately capturing" snow depth variations. In addition, please check the MW99 statistics in Table 4. The reported RMSE of 7.64 cm is smaller than both the absolute bias of 12.29 cm and the MAE of 12.70 cm. This is mathematically impossible, suggesting that some values or labels in the table may be incorrect.
7. Table 2 shows that MLR performs relatively poorly in the internal test but performs well in the OIB validation. In Line 549, the authors explain this difference as possible overfitting of the more complex models. Since overfitting should normally be considered and controlled during model design and training, the authors should explain what measures were used to reduce overfitting and provide more evidence to support this explanation. If RF, LightGBM, and LSTM are clearly affected by overfitting, the generalization and applicability of their products should also be reassessed.
8. The English should be improved throughout the manuscript.
Citation: https://doi.org/10.5194/egusphere-2026-2504-RC3 -
AC3: 'Reply on RC3', Mengmeng Li, 30 Sep 2026
reply
This study uses four algorithms, including MLR, RF, and LSTM, to retrieve snow depth on Arctic sea ice and further estimate sea ice thickness. However, the authors do not identify an optimal method, and the study appears to produce four separate datasets. The authors should clarify the applicability of each algorithm in the abstract, conclusions, and other relevant sections. Such recommendations should not be based solely on the overall comparison with OIB, but should also consider differences in ice type, season, and temporal resolution. In addition, several figures, tables, and descriptions require further improvement.
We thank the reviewer for the constructive and insightful comments. We agree that the four retrieval algorithms should not be treated as having a single universal ranking. Their relative performance depends on the evaluation dataset, spatial and temporal scale, ice type, season, and selected performance metric. We have therefore revised the Abstract, Results, Discussion, and Conclusions to describe the four products as complementary rather than to identify one universally optimal algorithm.
Specifically, we added a model-applicability summary that distinguishes: (i) performance on the ASD-derived internal test set; (ii) agreement with monthly and daily OIB observations, including separate FYI and MYI evaluations and seasonal performance differences; (iii) mean-bias behavior in the MOSAiC comparison; (iv) daily-to-monthly internal consistency; and (v) differences in temporal resolution. We also clarified that the four products are not intended as four independent datasets, but as complementary algorithmic retrievals that characterize structural uncertainty and support application-specific model selection. In addition, we revised statements that previously implied that MLR, LSTM, or any other model was universally superior, clarified the roles and characteristics of the benchmark products, added a feature-importance and final-feature-selection figure, revised the figure and table captions, corrected and rechecked the MOSAiC validation statistics, and conducted comprehensive English editing.
Specific comments:
1. Please provide the full name of OIB when it first appears in the abstract.
Thank you. We have revised the Abstract to define OIB at its first occurrence as Operation IceBridge (OIB).
2. The abstract should be reorganized to describe the datasets, retrieval algorithms, feature selection, temporal coverage of the products, and main conclusions more clearly. At present, the abstract places too much emphasis on the comparison between MLR and OIB. This result only represents model performance for a specific validation dataset and set of evaluation metrics and does not fully reflect the main contribution of the study. Because the model rankings differ between the internal test and independent validation, the general expression "highest accuracy" should be avoided. The performance of the four models and the main characteristics of the four products should be summarized more objectively.
We agree. The Abstract has been reorganized to present the input datasets, retrieval framework, feature-importance analysis, temporal coverage, independent evaluations, and the complementary applicability of the four products. We removed the statement that MLR has the “highest overall accuracy,” because this result depends on the OIB validation dataset and selected metrics. The revised Abstract emphasizes that no single model is optimal in every evaluation setting.
Revised Abstract:
Snow depth is a critical climate indicator and a key parameter for Arctic sea ice retrieval. In this study, we retrieve pan-Arctic snow depth from 2012 to 2021 by integrating satellite altimetry, passive microwave brightness temperatures, and multi-source ground/airborne data. Four machine learning models,Light Gradient Boosting Machine (LGBM), Multiple Linear Regression (MLR), Random Forest (RF), and Long Short-Term Memory (LSTM), are trained using altimeter-derived snow depth (ASD) for 2018–2021 and applied to generate daily and monthly 25-km products. Permutation importance indicates that snow depth and microwave gradient and polarization ratios provide complementary predictive information, with the 89 GHz polarization ratio generally more importance over multi-year ice than over first-year ice. In independent monthly OIB comparisons, MLR shows the lowest RMSE (7.19 cm) and highest correlation (R = 0.67), whereas daily OIB and MOSAiC evaluations show complementary model behavior; MOSAiC correlations are weak for all products and do not support a definitive ranking. The four products therefore provide complementary options rather than a universally optimal retrieval. In a sea ice thickness sensitivity experiment, replacing the Modified Warren 1999 snow depth input with the retrieved snow depth reduces mean thickness bias by nearly an order of magnitude, while RMSE remains comparable among products. The resulting observationally evaluated daily and monthly 25-km snow depth datasets support Arctic sea ice and snow studies, with uncertainties that vary across ice types, seasons, and the extrapolation period.
3. Section 4.1 discusses the importance of different frequencies and polarization parameters for different months and ice types, but no corresponding figure or table is provided. Please add a figure or table to support these descriptions and list the final parameters selected as inputs for each model.
Thank you for this constructive comment. We agree that the importance of different frequencies and polarization parameters should be presented more clearly. In the revised manuscript, we have added two new figures in Section 4.1:
Figure 2: Importance ranking of the final input parameters selected for the FYI snow depth retrieval for each month from October to April. The x-axis lists the final input parameters, and the y-axis shows their importance values. A taller bar indicates greater importance. GR(11/7V) and GR(11/7H) denote the vertically and horizontally polarized gradient ratios, respectively.
Figure 3: Importance ranking of the final input parameters selected for the MYI snow depth retrieval for each month from October to April. The x-axis lists the final input parameters, and the y-axis shows their importance values. A taller bar indicates greater importance. GR(11/7V) and GR(11/7H) denote the vertically and horizontally polarized gradient ratios, respectively.
The factors listed in Figure 2 are the final input parameters selected for the FYI model, and the factors listed in Figure 3 are the final input parameters selected for the MYI model. We have also clarified in Section 4.1 that these parameters were selected according to their importance rankings and are used in the corresponding FYI and MYI models.
4. Please clearly state which datasets are represented by SM, Rost, and MW99 in Lines 471–472.
Thank you. We have defined these abbreviations explicitly at their first occurrence and in the relevant table captions and Results text.
Revised text in Section 4.2:
For comparison, we evaluated three benchmark products: the SnowModel-LG snow-depth output (hereafter SM), the passive-microwave snow-depth product of Rostosky et al. (2018) (hereafter Rost), and the Modified Warren 1999 snow-depth climatology (hereafter MW99).
Revised Table 3 caption:
Table 3. Evaluation of monthly snow depth from LGBM, MLR, RF, LSTM, SnowModel-LG (SM), the Rostosky et al. (2018) passive microwave product (Rost), and the Modified Warren 1999 climatology (MW99) against OIB airborne snow depth observations. (RMSE/bias/MAE units: cm).
5. Please further discuss the main differences and advantages of the products developed in this study compared with Rost and other existing products, not only in terms of model validation results but also with respect to spatial and temporal resolution, temporal coverage, seasonal coverage, and other relevant aspects.
We agree. We added a dedicated comparison subsection to distinguish the products by input data, retrieval approach, spatial and temporal resolution, period of coverage, seasonal availability, ice-type treatment, and intended use. We emphasize that the principal contribution of this study is not only a validation comparison, but also the extension of an altimeter-informed retrieval framework to a continuous daily and monthly pan-Arctic record for the 2012–2021 growth seasons.
6. Lines 531–532 suggest that some models accurately capture the snow depth at the MOSAiC locations based on their relatively small bias and MAE. However, the correlation coefficients of all products are very low. If the low correlations are mainly caused by the scale difference between point measurements and gridded products, the authors should clarify which aspects of product performance can be evaluated using MOSAiC data and weaken the conclusion about "accurately capturing" snow depth variations. In addition, please check the MW99 statistics in Table 4. The reported RMSE of 7.64 cm is smaller than both the absolute bias of 12.29 cm and the MAE of 12.70 cm. This is mathematically impossible, suggesting that some values or labels in the table may be incorrect.
We agree. We revised the MOSAiC evaluation to state explicitly that, because MOSAiC observations are point measurements while the evaluated products are 25-km grid-cell averages, the comparison is mainly informative for systematic differences in the sampled mean state and broad error magnitude. It cannot robustly assess point-scale spatial or temporal variability, and the low correlations prevent a reliable ranking of product accuracy at the MOSAiC locations.
We also thank the reviewer for identifying the inconsistency in the MW99 statistics. The reported MW99 RMSE of 7.64 cm is mathematically inconsistent with the reported absolute bias of 12.29 cm and MAE of 12.70 cm. We rechecked the collocated data, metric calculations, units, and table-transcription process. The corrected values are now reported in Table 4. (The corrected RMSE is 14.47 cm).
7. Table 2 shows that MLR performs relatively poorly in the internal test but performs well in the OIB validation. In Line 549, the authors explain this difference as possible overfitting of the more complex models. Since overfitting should normally be considered and controlled during model design and training, the authors should explain what measures were used to reduce overfitting and provide more evidence to support this explanation. If RF, , and LSTM are clearly affected by overfitting, the generalization and applicability of their products should also be reassessed.
We agree that the available comparisons do not isolate overfitting from dataset shift, reference-product uncertainty, sampling differences, and scale mismatch. We have therefore removed statements that described the internal-versus-independent performance difference as a definitive indicator of overfitting. Instead, we now characterize it as evidence that model rankings are sensitive to the validation setting and that the transferability of the retrievals must be interpreted cautiously.
To reduce overfitting during model development, RF and LGBM hyperparameters were selected using five-fold cross-validation and grid search. RF complexity was constrained through maximum tree depth and the number of candidate predictors used at each split. For LGBM, hyperparameters were selected using cross-validation, with a low learning rate and controlled tree-building settings. The LSTM architecture was restricted to a single layer with 64 neurons, and model selection was based on validation performance. Early stopping, dropout, and validation-loss monitoring were also applied during LSTM training. We have added these details to the Methods section.
However, cross-validation within the ASD-derived training dataset does not fully test transferability to OIB or MOSAiC because these datasets differ in measurement principle, spatial footprint, sampling distribution, and error structure. We therefore revised the Results and Conclusions to state that the four products should be interpreted as complementary, and that the independent evaluations do not provide definitive evidence that RF, LGBM, or LSTM are severely overfitted. These measures, together with the independent evaluations, indicate that the ranking differences are more consistent with validation-setting dependence than with severe overfitting. Therefore, we did not reclassify these products as overfitted; instead, we emphasize their complementary applicability.
8. The English should be improved throughout the manuscript.
We appreciate this comment. The manuscript has undergone comprehensive language editing to improve grammar, sentence structure, terminology, consistency, and readability. We standardized terminology and capitalization, including “pan-Arctic,” “CryoSat-2,” “,” “Operation IceBridge (OIB),” “sea ice thickness (SIT),” “first-year ice (FYI),” and “multi-year ice (MYI).” We also corrected typographical errors, punctuation, hyphenation, and inconsistent American/British English usage.
-
AC3: 'Reply on RC3', Mengmeng Li, 30 Sep 2026
reply
Viewed
| HTML | XML | Total | BibTeX | EndNote | |
|---|---|---|---|---|---|
| 214 | 105 | 45 | 364 | 23 | 26 |
- HTML: 214
- PDF: 105
- XML: 45
- Total: 364
- BibTeX: 23
- EndNote: 26
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
This manuscript presents a study on snow depth retrieval over Arctic sea ice using multi-source satellite data and four machine learning models. The authors compare MLR, RF, LGBM, and LSTM, evaluate their performance against independent OIB and MOSAiC observations, and demonstrate the practical utility of their retrievals by improving sea ice thickness estimates. The work is methodologically sound, and the validation framework is rigorous. The long-term (2012–2021) snow depth product is an important dataset for Arctic climate studies.
Major Comments
1. The study employs four models (MLR, RF, LGBM, and LSTM) and reveals complex complementary performance—MLR performs best against OIB, while LSTM achieves the lowest bias against MOSAiC. However, the manuscript does not provide explicit guidance on how readers should select among these models for different application scenarios.
2. The MW99 product exhibits very large bias (7.04 cm, Table 3), far exceeding that of all other products. The manuscript notes that MW99 is outdated, but does not elaborate on the physical reasons for its poor performance.
3. The Introduction mentions both Rostosky et al. (2018) and Li et al. (2024) as existing snow depth products, but their roles in this study are different and may confuse readers.
4. The models are trained on data from 2018–2021 (the ASD product) and applied to retrieve snow depth back to 2012. While the temporal extrapolation test (Section 4.3.1) demonstrates generalizability, it remains unclear whether the training period (2018–2021) is representative of the longer period. Did the snow–microwave relationships remain stable across this entire period, or have there been systematic shifts due to changing Arctic conditions?
5. The manuscript's methodological contribution lies in the combination of systematic model intercomparison, rather than simply "applying four ML models to snow depth retrieval." However, the current framing may give readers the impression that this is just another ML-based retrieval study.
6. The authors should consider whether it is necessary to include a discussion of its role in ice thickness retrieval, as this study mainly focuses on snow depth retrieval.
Minor Comments:
1. The abstract states that the validation "reveals complementary strengths of snow retrieval among the models," but does not specify what these complementary strengths are.
2. Line 60, in the sentence introducing data fusion approaches, the term "bridges" should be changed.
3. Line 112, the authors note that 89 GHz is "known" to be affected by atmospheric water vapor but do not cite a reference to support this claim.
4. Table 3 shows seven products compared against OIB, but the layout is dense and could be made more readable.
5. Line 548, "This scale issue affects all products equally". This statement is slightly too strong. While it is true that all products suffer from scale mismatch, the degree of impact may differ. Some models (e.g., LSTM) might be more or less sensitive to the scale mismatch than others.
6. Line 574, "A black line represents the point of flawless agreement with OIB standards". "flawless agreement" is informal.
7. In section 5.1, the uncertainty analysis uses 1000 Monte Carlo iterations for non-linear models. The manuscript does not justify the choice of 1000 iterations.
8. Lines 674-676, the manuscript contains several long, multi-clause sentences that could be broken up for improved readability.
9. Figures 5 and 6 have relatively brief captions that do not fully explain what is being shown.
10. The language still needs further polishing, and consistency should be improved, such as ensuring that symbols are used consistently throughout the manuscript.