Random forest predictions of tundra snow density elevate Arctic soil temperatures in CLM5.0
Abstract. Arctic snow exerts a critical control on winter soil temperature and carbon exchange, however representation of its properties in Earth System Models (ESMs) remains simplified. In the Community Land Model v5.0 (CLM5.0), recent updates to snow compaction schemes have led to overly dense tundra snow, increasing snow thermal conductivity and conductive heat loss, producing a persistent cold-soil bias. Here we developed a Random Forest (RF) regression model to derive tundra bulk snow density from meteorological variables, trained on six winter seasons (Sep – May) of Arctic SVS2-Crocus (ASC) simulations constrained by in-situ observations from Trail Valley Creek (TVC), Northwest Territories, Canada. RF-derived snow densities capture the lower bulk densities of Arctic tundra snowpacks, which CLM5.0’s compaction scheme overestimates. Application of the RF-derived snow densities to CLM5.0 reduces snow thermal conductivity and enhances snowpack insulation, decreasing soil temperature RMSE by approximately 2–3 °C relative to field measurements (2017–2023) and increasing future winter soil temperature projections (2016–2100) by 4–7 °C under RCP4.5 and 8.5. Such increases in simulated soil temperature highlight the strong sensitivity of CLM5.0’s soil thermal regime to snow physical properties. The RF model reproduces ASC-simulated density evolution with a mean absolute error of 32 kg m−3 which falls within typical measurement error for volumetric snow density sampling in Arctic tundra (20–40) kg m−3 and matches field measurements more closely than default CLM5.0. Future bulk snow density predictions using the RF model driven by bias-corrected North American-Coordinated Regional Downscaling Experiment (NA-CORDEX) meteorology indicate bulk snow densities 200–450 kg m−3 lower than CLM5.0 and more consistent with tundra conditions.
This study uses a random forest (RF), trained on six winters of Arctic SVS2-Crocus (ASC) simulations at Trail Valley Creek (TVC), to modify snow thermal conductivity in CLM5.0. The reduction in winter soil-temperature bias is encouraging, and the problem is relevant to The Cryosphere. I am less convinced that the RF improves on existing corrections or that it supports the conclusions drawn from the future runs. I recommend major revisions, particularly to the comparisons and validation described below.
Major comments
1. The authors already show that the Sturm conductivity scheme reduces the cold bias. Why is that configuration absent from the comparisons? Default density with Sturm conductivity should be tested against RF density with Jordan conductivity under the same forcing. A seasonal ASC-density climatology and a simple regression would also be useful baselines. With cumulative snowfall as the dominant predictor (L266), it is difficult to tell how much the RF learns beyond the seasonal cycle. Separating seasonal skill from interannual anomalies would help answer this. This is the kind of improvement we would particularly expect under future warming.
2. I would expect to see an analysis using explainable methods to identify what is actually constrained and improved by the observations. The phrase “constrained by in-situ observations” needs qualification. Observations are used to select ASC seasons and ensemble members; ASC then supplies the training labels. Figure 1 tests against the ensemble mean, so the training and evaluation targets also differ. The late-season RF bias relative to ASC happens to bring predictions closer to the measurements, but it is unclear whether this improvement holds across individual years. Please show comparisons on the campaign dates and distinguish errors against ASC from errors against observations. Agreement near peak SWE cannot establish that the early-winter evolution is correct. More generally, we should have interpretable results to understand what happens when the models are integrated. This would benefit the community beyond simply demonstrating better predictions from black-box machine learning models.
3. The authors acknowledge the problem with random partitioning of daily data, but it remains in the tuning procedure. Use leave-one-season-out validation and report the hyperparameters, feature importance and effect of adding training seasons. The exclusion of wind, and the possible role of shortwave radiation as a seasonal proxy, also deserve explanation. A more fundamental issue is Figure 3: NA-CORDEX winters do not reproduce the weather of the corresponding observed years. The RMSE therefore includes forcing mismatch. The present-day evaluation should use observed TVC meteorology for both models, with errors reported by winter and soil depth. This would give readers a clearer understanding of model performance and avoid overstating predictive accuracy based on the full dataset.
4. In the conductivity-only approach, the RF density used to calculate thermal conductivity differs from that implied by CLM5’s simulated snow mass and layer thickness. Please discuss the physical implications of using these different densities and whether the improved soil-temperature agreement could partly reflect compensating errors. How is the RF bulk density applied to individual snow layers (Eqs. 3–7), and how do the two insertion schemes affect snow depth, SWE and snowpack thermal resistance? Energy diagnostics would also help explain why snow layer scaling produces a warm soil-temperature bias (Line 299-300).
5. I remain doubtful about carrying a relationship learned from six selected historical winters through to 2100 before the earlier concerns have been resolved, clearly explained or evaluated. The authors recognise the RF extrapolation problem, but the Jensen–Shannon analysis only describes a shift in the inputs. It does not tell us how large the density or soil-temperature error might become. Figure 5 combines all projection years into one seasonal curve, which obscures the response to warming. Please separate early-, mid- and late-century density and conductivity, test the RF on withheld warm winters, and compare it with ASC under representative future forcing. The uncertainty analysis should include training-season and ASC-label choices, rather than only climate forcing and CLM parameters. The plateau in Figure 6 needs further support for the proposed latent-heat explanation, including temperature profiles, ice/liquid-water and energy diagnostics. These checks are needed before the 4–7°C offset can be interpreted as a more realistic future projection. A winter temperature approaching 0°C at 2 m alone does not establish when or how much permafrost thaws. In general, future projections are more challenging than present-day simulations and require further investigation and evaluation. I agree that using machine learning as an emulator is a good approach, but the projections are somewhat ambitious for the scope of the current manuscript.
6. For the comparison between the RCP4.5 and RCP8.5 scenarios, how much of the apparent scenario difference comes from ensemble composition or initialisation? A comparison using common GCM–RCM pairs and baseline anomalies would help. The “all members have SWE > 5 mm” criterion may also select different winter windows; their lengths should be reported and a common window tested. Please explain how the carbon-parameter choices affect soil-temperature spread. Without carbon-flux results, the claim that previous carbon-loss projections are a conservative lower bound goes beyond this study. In addition, I suggest that the study, including its title and conclusions, should consistently reflect its single-site scope rather than imply a spatial evaluation.
7. The late-season RF bias relative to ASC brings predictions closer to the observations at TVC, but this may reflect compensation for a site-specific ASC bias. Would the same approach still improve agreement with observations at a site where ASC has a different bias, or could it make the predictions worse? Please discuss this possibility and clarify whether the improvement observed at TVC can reasonably be expected at other sites.
Specific comments
Figure 1: The legend, axis labels and units are difficult to read.
L171–175: Please explain the physical rationale for including SW radiation as a predictor of bulk density (it is largely a proxy for day of year at this latitude) and comment on whether it simply encodes seasonality.
L196, L203: "Kullback-Leiber" should be "Kullback-Leibler".
L199: JSD on a 2-D KDE of only two predictors does not capture the joint distribution of all four predictors.
L264, L367 and L479: Clarify whether the 32 kg m⁻³ value refers to MAE or RMSE and define the evaluation target. Why report different metrics instead of using one consistently?
Section 2.2 and Table 1: Specify temporal aggregation, precipitation-phase partitioning, RF-to-CLM time-step mapping and the handling of snow-free conditions. A more detailed description of the data preprocessing and methods would help readers better evaluate your study.
Figure 3: Please check and clarify the depth or scenario assignments in the caption.
Figure 6: Please define the uncertainty bands more clearly.
L214–215: Supply the missing table number.
Appendix A: Correct “RSME” to “RMSE”.
L232–234: The 2 m soil temperature data are unpublished.
L300–302: The 10–24-day delay in spring warming persists across all configurations and is not fully resolved by the density corrections. Please discuss whether melt-energy or snow-cover-fraction processes contribute to this remaining bias.
L310–312: Mean winter air temperature increases from −13 to −7 °C, but accumulated snowfall stays at 59–62 mm; please state whether precipitation phase is recomputed under warming, since constant snowfall with 6 °C warming seems unlikely at TVC.
L412–414: Dutch et al. (2024) is cited for a factor-of-3 overestimate of Keff. I believe this should be Dutch et al. (2022); please check.
L468–475: The claim that existing permafrost carbon estimates are "a conservative lower bound" is not supported by any carbon result in this paper; please either show fluxes or soften the claim.
L500–503 (Appendix B): Please define the formula for the performance score explicitly and give the scores for all 29 seasons in a supplementary table.