the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
A native-grid cover–depth benchmark reveals coupled snow-state biases in CMIP6 over the Tibetan Plateau
Abstract. Snow over the Tibetan Plateau is often widespread in winter but remains shallow over much of the region. This makes snow-covered area and snow depth complementary, not interchangeable, measures of model skill. We develop a monthly native-grid cover–depth benchmark for 2002–2014, evaluating CMIP6 historical simulations and major reanalysis products against MODIS-derived snow-covered area percentage (SNC) and CHE snow depth (SD). The reference fields are transferred to each product’s native grid before bias calculation, so that snow-state errors are diagnosed on the spatial support of the evaluated product. CMIP6 shows a dominant coupled snow-state excess over the Plateau. Median SNC biases reach 37.7 %, 30.2 % and 24.1 % in January, February and March, respectively, while SD bias is also positive but strongly skewed by several large outliers. In the January–March cover–depth bias space, 19 of 21 CMIP6 models overestimate both SNC and SD. The largest coupled biases occur over the western and high-elevation Plateau, indicating a structured cold-season error rather than a domain-wide offset. Reanalysis products show different forms of mismatch: ERA5-Land overestimates both SNC and SD, whereas MERRA-2 combines underestimated SNC with overestimated SD. Forcing diagnostics suggest that cold-biased persistence, snowfall input, snowfall partitioning and snow-process errors all contribute to the model spread. A cover–depth benchmark is therefore needed before Tibetan Plateau snow products are used in cryospheric, hydrological or land–atmosphere studies.
- Preprint
(2097 KB) - Metadata XML
-
Supplement
(290 KB) - BibTeX
- EndNote
Status: open (until 22 Sep 2026)
-
RC1: 'Comment on egusphere-2026-3134', Anonymous Referee #1, 11 Sep 2026
reply
-
AC1: 'Reply on RC1', Yao Xiao, 15 Sep 2026
reply
We thank the referee for the careful and constructive review and for the positive assessment of our study. Please find attached our response to the comments, including how we plan to address the main and specific points in a revised manuscript.
-
AC1: 'Reply on RC1', Yao Xiao, 15 Sep 2026
reply
Viewed
| HTML | XML | Total | Supplement | BibTeX | EndNote | |
|---|---|---|---|---|---|---|
| 127 | 42 | 20 | 189 | 21 | 23 | 18 |
- HTML: 127
- PDF: 42
- XML: 20
- Total: 189
- Supplement: 21
- BibTeX: 23
- EndNote: 18
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
The study presents a monthly, native-grid, paired snow cover and snow depth benchmark for Tibetan plateau from 2002-2014. This benchmark is used to evaluate CMIP6 historical models and 3 common reanalysis products, ERA5-Land, MERRA-2, and JRA-55, in snow cover-depth space. The study seeks to answer which monthly depth-cover state is represented by the evaluated products, how the products differ from that state, and which products are suitable for Tibetan plateau snow state applications. The study finds that all CMIP6 overestimate either snow depth or cover and frequently both, additionally these overestimations are concentrated in high elevation areas and winter months. The study finds that ERA5-Land, MERRA-2, and JRA-55 all misrepresent the snow depth and cover in certain ways, and none provide a consistent correction for CMIP6 models. Overall, I think the paper provides valuable insight on the limitations of commonly used snow products over the Tibetan Plateau, particularly the spatial and temporal distribution of biases. I would like expanded discussion of a few points, please see my general suggestions below and following line recommendations.
------
I would appreciate a clearer explanation of the role of ERA5-Land, MERRA-2, and JRA-55 in this study. From the abstract and introduction, I understood that ERA5-Land, MERRA-2, and JRA-55 were evaluated for usability the same as the CMIP6 models but later ERA5-Land, MERRA-2, and JRA-55 is discussed as to whether they can provide corrections to the CMIP6 models.
I would like some discussion of why all the CMIP6 models have positive snow depth and cover biases. Of the products evaluated only JRA-55 shows widespread underestimation.
From Figures 4, 5c, and 6b the large SD biases seem to be mostly concentrated in the regions with higher SD (fig 2) Could this be a case of depth bias increasing with depth? If you calculate the SD bias as a percentage of SD do you see the same spatial distribution of SD biases?
---------
L71 – “Spatial supports” I'm not familiar with this phrase, do you mean spatial scales?
L76-77 – “CMIP6 models and reanalysis products”, how many and which products? I see that the CMIP6 models are described in Table S1 but I would like more explanation of the CMIP6 models assessed in the text.
L76 - Please define CMIP6 at first mention.
L93-95 – My understanding is that the developed benchmark and the MODIS-derived SNC and CHE SD are the same dataset, however these lines make it sound like they are two separate datasets.
L109-111 – This is a lot of acronyms and the acronyms used are sometimes un-intuitive. This makes the text difficult to follow, I recommend reducing the number of acronyms and making the remaining ones clearer. I would consider SCA (snow covered area), fSC (fractional snow cover), or fSCA (fractional snow covered area) as more standard acronyms for snow-covered area percentage then SNC. I would also consider T_air and PR_snow more common than TAS and PRSN.
L111 – Please define “surface snow amount”. Is this snow water equivalent (SWE)? Or a snow volume estimate?
L114-119 – Please expand on the working subsets division. In Table S1 I see that some models are included in all, none, or a fraction of the TAS, PR, and PRSN subsets and all models are included in the SNC, SD, and SNW subsets therefore I don’t understand what the two subsets you are describing in L114-119 are. From the following section it appears that all the CMIP6 models are evaluated in depth-cover space while a subset of 7 of them have forcing-diagnostic analysis. Please explain why these 7 models are used for forcing-diagnostic analysis.
L134: “The benchmark period is 2002-2014 unless otherwise stated.” Please note the variant benchmark periods in Table 1.
L213-214 – Please explain why these thresholds where chosen.
L292-294 - IPSL-CM6A-LR has over double the SD bias then other CMIP6 models, why does this model behave so poorly?
Fig 5 – Consider plotting the SD biases with the same color scale as how the SNC biases are plotted. MERRA-2 and JRA-55 have much lower SD biases then ERA5 but as currently plotted the bias ranges look similar. This may require extending the color range so that bias distribution is still visible in plots d) and e)
L349-351 – The month range is a little hard to follow here. I think it would be clearer to say that for many models the positive bias begins in Nov-Dec and remains into Apr.
4.3 & 4.4 – Consider moving section 4.3 to after 4.4 as only the CMIP6 models are discussed in section 4.4.
Figure 8b – Please make the arrows larger, they are currently too small to read clearly
Fig 9 – Explain that these 7 CMIP6 models are plotted here because they are in the forcing-signature subset, currently readers need to reference the supplement to know which models are in which subset making this figure confusing. Consider adding an additional column to Table 2 to list which subsets models are included in
5.1 - This section restates much of your introduction, you have already justified the importance of a paired depth-cover benchmark and a native-grid approach. This section can therefore be removed or significantly cut down to focus only on the relevant characteristics of TP snow cover for the discussion.
L459-462 – Please be explicit about how these processes affect simulated snow state. Do you mean that these broader snow-models also overestimate snow in the late season or something else?
L538-541 – Please quantify the common passive-microwave snow-depth error. The text implies that the CMIP6 biases are larger than the passive-microwave snow-depth errors but by how much?