SPUN: Deep learning for continuous snow cover fraction retrieval in marginal environments
Abstract. Current operational snow mapping products struggle to accurately map patchy, late-season snow cover in challenging environments like Scotland, hindering hydrological and ecological model validation. We hypothesise that machine learning (ML) could better handle these variable conditions, including dirty and metamorphosed snow, cloud and atmospheric disturbance, varying illumination levels and shading.
To address this, we generated a novel Snow Cover Fraction (SCF) reference dataset for Sentinel-2 (10 m resolution) by aggregating manual and semi-automated "pseudo-labels" derived from 25 cm aerial surveys of Scotland. Using this dataset, we developed a Snow Patch U-Net (SPUN) by adapting a U-Net architecture with a regression head to predict SCF, utilising a combined Dice and L1 loss function. On our Scottish test set, SPUN successfully mapped challenging snow patches – even through cloud gaps and in topographic shadow – outperforming the operational Theia Let-It-Snow (LIS) product with a snow segmentation F1 score of 0.737, a Mean Absolute Error (MAE) of 28.49 %, and an RMSE of 42.16 %.
Furthermore, by integrating pre-trained weights from satellite foundation models, our model demonstrated strong generalisation when applied to an independent global dataset, achieving a macro averages F1 Score of 0.9313 across the three classes (snow, cloud, background) and a snow-only F1 score of 0.8923.
This study demonstrates that ML has the potential to tackle persistent challenges in optical snow retrieval. We suggest that a broader, community-driven effort to create a global, high-resolution training dataset could significantly advance snow mapping capabilities across optical satellite sensors.
This manuscript presents an interesting and potentially valuable contribution to optical snow-cover mapping. The development of a high-resolution reference dataset for a challenging marginal snow environment, together with a deep-learning framework specifically targeting continuous Snow Cover Fraction (SCF), is timely and relevant. I particularly appreciate the effort to move beyond conventional clear-sky validation, to investigate difficult illumination and atmospheric conditions, to make the model and data openly available, and to assess transferability beyond the Scottish training domain. The results are promising and suggest that the proposed approach has potential for improving snow mapping in conditions where traditional NDSI-based methods struggle.
However, in its current form, I think several methodological and interpretative aspects require substantial clarification before the main conclusions can be fully supported. I believe the manuscript could become a strong contribution after these issues are addressed, and my comments below are intended to help strengthen the analysis and better delimit the conclusions supported by the experiments. My main concerns relate to:
Figure 1 does not currently demonstrate the observational difficulties that are most relevant to the study i.e., retrieval of snow in marginal environments. It uses simulated spectral albedo to motivate effects on Sentinel-2/NDSI even though Sentinel-2 measures directional TOA reflectance, and the distinction may be especially relevant under the low-sun and potential complex topographic conditions emphasized in this study. Albedo is essentially hemispherically integrated reflectance, whereas the satellite sees directional (radiance) reflectance for a particular sun–surface–sensor geometry. For NDSI, which is generally calculated from B11 and B3 (and not from B2 as shown in the figure), a low illumination does not necessarily behave as an identical multiplier in both bands. Topographic shadow, atmospheric scattering/path length, snow anisotropy/BRDF, and very low signal can alter the spectral relationship altering the expected behavior of NDSI. Given that the authors reported rapid melt and low solar elevation are central characteristics of the Scottish dataset it would be more informative to quantify these effects instead of difference in grain size, for example by evaluating retrieval behavior as a function of againg, wetting and also of solar zenith angle, topographic illumination and shadow. I like a theoretical background but why not convolve the simulations with your reference Sentinel-2 spectral response functions and actually show the real B3, B11 and NDSI values under different conditions? Furthermore, the definition and representativeness of the assumed 0.1–1 mm "grain size" should be clarified (is it the optical-equivalent radius, effective radius, physical grain radius, diameter, or a value derived from SSA?)
L27: The manuscript repeatedly refers to Sentinel-2 as having a spatial resolution of 10 m. This is imprecise, as the MSI bands have native resolutions of 10, 20, and 60 m. The paper later clarify that all bands are resampled to 10 m. Please distinguish throughout between native spatial resolution and the 10 m model grid. In particular, upsampling the 20/60 m bands does not add spatial information, which should be acknowledged when describing the effective resolution of the SCF product.
L30: The limitation of MODIS for resolving heterogeneous snow cover in complex mountain terrain is well established, but Gascoin et al. (2024) is a review article. More importantly, the statement that snow-cover variability “typically occurs at length scales of less than 100 m” is rather general. The spatial resolution required to represent snow-cover heterogeneity depends on the spatial scales of the underlying snow patterns and should ideally be discussed in the context of spatial sampling principles (e.g. Shannon–Nyquist), rather than through a single generic threshold such as 100 m. In practice, the ability of a sensor to resolve these patterns also depends on its effective spatial response and on the prevalence of mixed pixels. I therefore suggest qualifying this statement and supporting it with primary studies explicitly investigating the scale dependence of snow-cover heterogeneity and mixed pixels (e.g. Selkowitz et al., 2014; Premier et al.). Selkowitz et al. (2014), for example, explicitly demonstrate how the occurrence of mixed snow pixels increases with pixel size from meter to MODIS-scale resolutions.
L39–40: This sentence is central to the motivation of the study, but the link to Figure 1 is not fully clear (see previous comment). The text identifies patchiness, wet snow, and rapid melt–freeze cycling as defining characteristics of marginal Scottish snowpacks, whereas Figure 1 only explores changes in simulated snow albedo associated with grain size and soot contamination. Patchiness is a sub-pixel mixing problem not represented in figure 1, wetness is not represented, and the relationship between the prescribed grain-size range and melt–freeze metamorphism is not explained. That’s why I suggest either revising the interpretation of Figure 1 or expanding the analysis so that it better represents the processes emphasized in the text, particularly wet snow/melt–freeze conditions and illumination effects relevant to the Scottish dataset.
Study region / Figure 2: I found it difficult to navigate between the geographical description of the study region and Figure 2. The text refers to the Isle of Skye, the Nevis range, the Cairngorms, and later Glen Coe, but these locations are not labelled on the map. Please consider adding the main geographic features discussed in the text to Figure 2 and improving the locator inset so that readers unfamiliar with Scotland can immediately understand where the study area lies within Scotland/Great Britain. It could also be useful to indicate, at least approximately, the western versus eastern climatic regions described in Sect. 2.1.
L~96: Please specify how the Sentinel-2 20 m and 60 m bands were resampled to the 10 m model grid (e.g. nearest-neighbour, bilinear, cubic interpolation). This is important for reproducibility and may affect the spectral information and spatial gradients at snow-patch boundaries. It would also be useful to clarify the reference grid/projection used for the resampling.
L98–101: The justification for selecting Level-1C TOA reflectance should be strengthened. First, TOA reflectance is not inherently comparable across multispectral sensors because sensor spectral responses, calibration, acquisition geometry, and atmospheric conditions still differ. If the intended advantage is consistent availability of Level-1 products and independence from sensor-specific atmospheric-correction processors, this should be stated more precisely. Second, the argument that TOA data preserve detail in topographic shadow is plausible, but TOA reflectance also retains atmospheric and illumination effects; please provide supporting references or an empirical L1C–L2A comparison for representative shaded snow patches. A potentially important additional justification is that standard L2A processing may rely on prior scene classification, including snow/cloud probabilities (as Sen2cor for example), creating a possible dependency on an existing snow-detection algorithm. Please identify the relevant L2A processor and clarify whether avoiding this potential circularity motivated the use of L1C data.
L125–133: The use of Cloud Score+ appears reasonable, and the selected cs_cdf threshold of 0.5 lies within the recommended range. However, because CS+ is used to define cloud-contaminated pixels in the reference dataset, the choice has a direct influence on both training and evaluation. I therefore suggest providing more evidence that this configuration is appropriate for the particularly challenging Scottish conditions considered here, ideally through a manual validation on a representative subset and/or a sensitivity analysis to the CS+ threshold. In addition, cs_cdf is a temporally normalized quality/visibility score rather than a strict binary definition of “clear sky”; consequently, the terminology should distinguish truly cloud-free observations from partially affected but potentially usable pixels (e.g. haze, cloud margins, cloud shadow). This distinction seems particularly important given that exploiting such marginal observations is one of the central claims of the study. For example, how would the conditions shown in the first three rows of Fig. D2 be classified according to CS+ and according to the Tier 2a clear-sky criterion?
L146–149: The use of a U-Net with a ResNet50 encoder and continuous output is reasonable for pixel-wise SCF regression, but the architecture used to produce the fractional estimate should be described in greater detail. In particular, please specify the regression head, output normalization/activation, treatment of padded pixels, and how the final 0–100% SCF values are obtained. The combined Dice–regression objective may favor spatial detection and predictions close to 0 or 100%, an effect that the authors themselves note later in the manuscript. Performance stratified by reference SCF, or a predicted-versus-reference SCF analysis, would help demonstrate that the model is retrieving fractional cover rather than primarily performing accurate snow-extent segmentation. Please also clarify how Soft Dice is computed when the reference SCF itself is continuous (or, if it is binarized, specify the threshold). Finally, the manuscript is inconsistent regarding the selected regression loss: the Abstract and Fig. 4 state L1, whereas the hyperparameter optimization identifies Huber as the selected loss. This should be reconciled throughout.
L161–178: The overall U-Net/ResNet50 architecture and use of Sentinel-2 pretrained weights appear reasonable, but several aspects of the implementation require clarification for reproducibility. Please distinguish between the backbone architecture and the different pretraining/initialization strategies (ImageNet, DINO, MoCo, DeCUR, SoftCon), and describe how the 13-band Sentinel-2 input is handled by the pretrained encoder, including band ordering, input normalization/scaling, adaptation of the first layer, and whether encoder weights are fully fine-tuned or partially frozen. There also appears to be an inconsistency regarding class weighting: the text states that a per-pixel snow weighting was used, whereas the selected hyperparameters report “no snow weighting.” The multi-objective Optuna procedure should also explain how the final model was selected from the Pareto-optimal solutions. Finally, the optimized snow decision threshold (0.3997) appears inconsistent with the 0.5 threshold used later for evaluation; please clarify which threshold was used at each stage and why. Additional training details such as number of epochs/stopping criterion, data augmentation, learning-rate scheduling and random-seed treatment would also improve reproducibility.
Experimental design (sec 2.4): The independence of the training, validation, and test datasets should be described more explicitly. The manuscript states that the splits “aim to preserve” spatial separation and that an attempt was made to keep different Sentinel-2 retrievals separate, but it is not clear whether complete Sentinel-2 acquisition scenes/dates are exclusively assigned to a single split. This distinction is important because multiple 1 km image chips extracted from the same Sentinel-2 acquisition share atmospheric conditions, illumination geometry, snow state, and radiometric characteristics. If chips from the same acquisition occur in both training and test sets, the reported test performance may therefore not represent a fully independent evaluation. Given that the complete dataset comprises only 37 Sentinel-2 scenes, this information is particularly important for assessing model generalization. In addition, I suggest being more precise about what is demonstrated by the global validation. The Wang et al. dataset provides categorical snow,cloud,background labels and therefore allows evaluation of the generalization of snow extent detection, but not of the continuous SCF regression itself. Consequently, statements about global generalization should clearly distinguish between global transferability of snow detection and global validation of fractional snow-cover retrieval. Demonstrating the latter would require independent continuous SCF reference data outside Scotland.
Table 1/Table 2: I found the distinction between Tier 1 (“Intrinsic”) and Tier 2b (“Reality”) unclear. Both scenarios appear to use all-sky observations, with the principal difference being that Tier 1 evaluates SPUN and the 10 m baselines at 10 m, whereas Tier 2b aggregates/resamples the products to 20 m to enable comparison with the native 20 m LIS product. If this interpretation is correct, I suggest using more descriptive terminology (e.g. “10 m native-resolution evaluation” and “20 m operational comparison”) and stating explicitly whether exactly the same test scenes/pixels and cloud treatment are used in Tier 1 and Tier 2b. Spatial aggregation can naturally reduce MAE/RMSE because positive and negative sub-pixel errors cancel within the larger 20 m pixel; therefore, this improvement should be interpreted primarily as a scale effect rather than improved retrieval performance. More importantly, please clarify how the 10 m products and reference were aggregated to 20 m. For an area-preserving aggregation over an identical set of pixels, mean bias would normally be expected to remain approximately invariant, whereas SPUN snow-only bias changes from +8.36% to +6.10%.
L365–370: I suggest revising the interpretation that bare rock, water bodies and varying illumination “mimic the spectral signature of snow.” These conditions may produce ambiguous values when the spectral information is reduced to an index such as NDSI, but this does not necessarily imply that their full (or multispectral) signatures are similar to snow. This distinction is particularly relevant here because SPUN uses all 13 Sentinel-2 bands, whereas NDSI uses only two bands. Consequently, the improvement over NDSI cannot necessarily be attributed to the use of spatial context alone; it may also result from the much richer spectral information available to the network. If the authors wish to demonstrate specifically the added value of spatial context, a pixel-wise classifier or regressor using the same spectral inputs would provide a more direct ablation. Comparison with a spectral-unmixing SCF method would also substantially strengthen the evaluation. Otherwise, I suggest moderating this interpretation.
L402 and conclusion: I suggest reconsidering the term “catastrophic overfitting.” The experiment presented here assesses geographic generalization—i.e. whether a model trained on Scottish data transfers to unseen regions—rather than what is usually referred to as catastrophic overfitting in the machine-learning literature. “Overfitting to the Scottish training domain” or “limited out-of-domain generalization” would be more appropriate. Moreover, the Wang et al. dataset provides categorical snow/cloud/background labels, so this experiment demonstrates geographical transfer of snow detection, but does not establish global generalization of the continuous SCF regression. The corresponding conclusions should therefore be moderated accordingly.
L414–416: The modular separation between a snow-specific model and an independent cloud-mask algorithm is an interesting and potentially useful design choice. However, I suggest discussing the corresponding propagation of errors through the two-stage pipeline. The performance of the final product depends jointly on both modules: snow pixels incorrectly masked as cloud cannot subsequently be recovered by SPUN, whereas clouds incorrectly retained by the mask must then be correctly rejected by SPUN. These errors are also unlikely to be statistically independent because both modules respond to the same spectral and illumination conditions. Consequently, replacing CloudScore+ with another cloud algorithm would not simply preserve SPUN performance but would require re-evaluation of the complete pipeline. It would be useful to report the performance of CloudScore+ itself on the global reference data and explicitly quantify how its omission and commission errors propagate into the final snow product.
L455–465: I agree that reproducibility, accessibility and model interpretability are important barriers to wider adoption of ML-based snow products. However, some of the discussion here is quite general and could be more closely connected to the findings of this study. Claims regarding the limited availability of code or training data in the existing literature would benefit from supporting references. I would also distinguish more carefully between explainable AI and physics-informed learning: attribution or perturbation methods can provide insight into the information used by a model, but they do not necessarily demonstrate that the model has learned or represents the underlying physical processes correctly. In this context, it may be particularly useful to discuss the concrete requirements for operationalizing SPUN identified by this study, including cloud-mask dependency, preprocessing, uncertainty characterization and validation of continuous SCF outside the Scottish training domain.
Selkowitz, D.J.; Forster, R.R.; Caldwell, M.K. Prevalence of Pure Versus Mixed Snow Cover Pixels across Spatial Resolutions in Alpine Environments. Remote Sens. 2014, 6, 12478-12508. https://doi.org/10.3390/rs61212478
V. Premier, C. Marin, S. Steger, C. Notarnicola and L. Bruzzone, "A Novel Approach Based on a Hierarchical Multiresolution Analysis of Optical Time Series to Reconstruct the Daily High-Resolution Snow Cover Area," in IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, vol. 14, pp. 9223-9240, 2021, doi: 10.1109/JSTARS.2021.3103585.