the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
A large-scale evaluation of available subseasonal precipitation forecast products over the contiguous United States
Abstract. Accurate precipitation predictions at the subseasonal timescale (beyond a week but within a season) could benefit a range of human activities, but are highly challenging to achieve. Research efforts have been made through multi-agency and international collaborations, resulting in numerous forecast products such as those included in the Subseasonal Consortium (SubC) and the Subseasonal-to-Seasonal (S2S) Prediction Project. However, a unified and comprehensive evaluation of the full suite of hindcast datasets from these efforts remains limited, partly due to inconsistencies in hindcast frequency and data periods across products. In this study, we employ the full suite of nineteen precipitation hindcast datasets from the SubC and S2S projects over the contiguous United States (CONUS). The hindcast datasets are temporally aggregated into weekly values and are assessed against a reference dataset derived from the Parameter-elevation Regressions on Independent Slopes Model (PRISM). Overall and seasonal evaluations are carried out using statistical metrics including percentage bias (PBIAS), anomaly correlation coefficient (ACC), and continuous ranked probability score (CRPS). Furthermore, we adopt a baseline-referenced skillfulness approach that accounts for differences in hindcast initialization, frequency, and data periods for a relatively fair comparison among the employed hindcast datasets. Our results indicate widespread overestimations in winter and spring across most hindcast datasets, while underestimations are more likely to be observed in summer and autumn. Predictive accuracy generally declines over forecast lead time and remains marginal beyond week three. Notable variations in predictive skill are observed across regions, seasons, and lead times, with no single hindcast dataset consistently outperforming others. In summary, this work provides valuable references for both forecast end-users and model developers, and highlights the need for context-specific selection of available subseasonal forecast products for downstream applications.
- Preprint
(11340 KB) - Metadata XML
- BibTeX
- EndNote
Status: final response (author comments only)
-
RC1: 'Comment on egusphere-2025-5062', Anonymous Referee #1, 01 Jan 2026
- AC1: 'Reply on RC1', Lujun Zhang, 01 Aug 2026
-
CC1: 'Comment on egusphere-2025-5062', Nima Zafarmomen, 08 Jan 2026
This study provides a comprehensive assessment of 19 subseasonal precipitation hindcast datasets from the Subseasonal Consortium (SubC) and the Subseasonal-to-Seasonal (S2S) Prediction Project. The evaluation covers the Contiguous United States (CONUS) and utilizes the PRISM dataset as a high-resolution reference. The paper provides high value to both forecast end-users and model developers. By identifying specific regions, seasons, and lead times where certain models excel, it helps practitioners move beyond "one-size-fits-all" model selection. The systematic use of both deterministic (ACC, PBIAS) and probabilistic (CRPS) metrics ensures a robust evaluation of model performance.
1- The authors correctly note that CRPS is sensitive to precipitation magnitude and should be interpreted alongside ACC. However, the discussion could be slightly strengthened by more explicitly discussing how "dry masks" might impact the perceived skill in arid regions like the Southwest during summer, as briefly mentioned in the discussion of future work2- Â The study re-grids all datasets to a uniform 0.25-degree resolution. While this is a standard approach, it would be beneficial to add a brief comment on whether models with inherently finer native resolutions (like the 1-degree SubC models vs. the 2.5-degree BOM model) showed a consistent advantage in complex terrains due to better representative physics.3- In the results, the authors identify "ACC hotspots" where accuracy remains higher for longer lead times (e.g., the West Coast and Florida Peninsula). It would be helpful to explicitly state in the conclusion if these hotspots are primarily driven by large-scale teleconnections (like ENSO or MJO) which are mentioned in the discussion, to provide a more process-based takeaway for the reader.4- The authors transparently acknowledge the limitations of extending pairwise logic to multi-model comparisons. To add even more value, a sentence could be added to the Discussion recommending a specific "best practice" for future researchers who might have the computational resources to run a fully synchronized multi-model experiment.5- I do strongly recommend the authors consider incorporating a comparison with recent findings in diverse hydroclimatic regions. Specifically, the study 'Analysis of historical global warming impacts on climatological trends for the partially gauged Hirmand river basin based on multiple data products and bias correction methods' demonstrates how similar multi-product evaluations and bias correction techniques perform in data-scarce, partially gauged environments.ÂÂÂCitation: https://doi.org/10.5194/egusphere-2025-5062-CC1 - AC3: 'Reply on CC1', Lujun Zhang, 01 Aug 2026
-
RC2: 'Comment on egusphere-2025-5062', Micha Werner, 05 Jul 2026
Review of Zang et al 2025: A Large-1 scale evaluation of available sub-seasonal precipitation forecast products over the contiguous United States
This manuscript develops and in-depth validation of nineteen sub-seasonal (ensemble) forecasting models across the contiguous United Stated (CONUS). These models are derived from the archives of two initiatives (S2S and SubC) that collated sub-seasonal hindcasts from providers of these modelling systems. Evaluation of the products is developed using standard performance metrics (Anomaly Correlation Coefficient, PBIAS and CRPS) as well as skill using CRPSS. The PRISM dataset, based on interpolated observational data is used here as a reference. The paper addresses the challenge of the disparate frequency as well as forecast start times of these sub-seasonal models, which makes inter-model comparison challenging.
Overall, the results of the study are interesting to the audience of this journal, particularly in the light of the increasing use of extended range (sub-seasonal) forecasts in hydrological streamflow prediction.There are, however, several comments and suggestions to the authors that could help strengthen the scientific merit of this contribution.
The datasets are sourced from two initiatives, the S2S and SubC projects. These two projects established a common dataset to allow comparison of the different models. In establishing this dataset, these projects also applied a number of pre-processing steps. For the models included in the S2S dataset, this included resampling the original model resolution to a common 1.5 degrees resolution. A similar preprocessing may also have been applied to the models participating in the SubC project, where the common spatial resolution of 1 degree was chosen. It may also be that other pre-processing steps were applied to the participating models in each experiment. The authors do not comment on this, but it may be useful to do so. Some of these models may have a much higher resolution, and while it may be that this resampling does not impact forecast skill, it is not entirely clear what influence this has. The authors should to my mind at least acknowledge this resampling in their source datasets. In the preprocessing described in the methods section, the choice is then made to resample all data (forecasts as well as the much finer resolution PRISM data) to the common 0.25 degrees resolution, using nearest neighbour resampling. Â What was the reason for choosing 0.25 degrees, rather than something closer to the resolution of the S2S and SubC datasets? In some of the analysis it is noted that performance metrics vary from cell to cell. The question would then arise to what extent this is an artefact of this up and down sampling, or if it truly reflects the spatial variability of model performance. I think that this at least warrants some discussion.
For the comparison of the forecast models using CPRS, there is no comparison with climatology. As CRPS has dimension (which is noted). It is difficult to interpret the meaning of the values of CRPS provided in for example Figure 6. What would the value have been if the climatology for PRISM had simply been used? I would assume it would be similar to the values found, as these show to become quite stable for most of the models after the 2-3 week lead time. For some values are higher than others (e.g. IAP). What is the implication of this? Is this due to the differences in the comparison (e.g. not the same years used)? To understand the lead time at which there is predictability, the lead time at which skill (CRPSS) goes to zero is considered, using climatology as a reference. CRPSS below zero indicates the forecast is worse than climatology. It would be helpful to understand that for these forecasts. This should at least be discussed.
Related to the above, CRPSS is calculated here using CFSv2 as the reference. Looking at Fig6 this model seems to have a similar performance (spatially as well as in lead time). For the seasonal comparison it appears that CFSv2 is not one of the best performing models at shorter lead times (particularly for ACC). Using this as a reference, as is then done in the results presented in section 4.3 means that the question being asked is not how skilful a model in question is (using e.g. climatology as a reference) but rather if it is better or worse than CFSv2. This makes Figures 16 and 17 quite difficult to interpret. A high value of ACCSS or CRPSS does not necessarily imply a good performance of the model itself. I can understand the use of CFSv2 as a reference to unify the temporal resolution in comparisons, but the interpretation of what the skill scores say about model performance can only be done if the results of CRPS for CFSv2 shown in Fig6 are considered at the same time. It would have been useful to my mind to calculate at leas the CRPSS of CFSv2 using climatology as a reference, which could then help interpret the skill scores of the other models. I understand that this adds quite some scope to the work presented, but also think this adds to the scientific content of the article. Perhaps this was already done? In any case this should be carefully discussed.The results section is quite long, particularly that dealing with the scores for each season. While the maps are well developed, the readability of the paper may increase by making this section a little shorter. As some of the seasons do have similar patterns, it could be an idea to leave only 2 seasons in the main paper and the remainder in the supplementary material. The two remaining are where the patterns are very different (e.g. autumn). The text can of course describe all seasons but focus primarily on the differences between the seasons, using the main one presented as the referenced. That may also raise the interest of this section. This is already done in some sections (e.g. Line 348).
It is in any case important to include high resolution versions of these images in the digital supplementary material, allowing more detailed inspection.
It would be good to expand a little how these sub-seasonal datasets could be used to support hydrological sub-seasonal forecasting. In using such forecasts as a forcing to a hydrological model, it would be commonplace to use a bias-correction method, of which there are several examples in literature. This may be less the case for sub-seasonal model than for seasonal models. Mention of this is made in lines 522-24 as a recommendation, but it is somewhat brief. Would such a bias correction, developed for each model mean that the differences found between the models diminish? Many bias correction methods deal well with systematic biases, but other properties such as ensemble overdispersion or underdispersion are more challenging. There are also studies that show that bias correction can indeed be detrimental to skill. It would add value to the paper to expand this discussion.  Several papers discuss bias correction of precipitation forecasts for hydrological models (e.g. https://doi.org/10.1016/j.jhydrol.2023.129322, several others). Recommendation to extend this work to study the hydrological predictability is given in lines 581 – 587, but the discussion can be extended to my mind.Â
Specific comments:
L74-L83: The authors state that they aim to something, that in fact they actually did! In particular the aim to assess 19 models was achieved, so perhaps reformulate.L160: I can understand the sampling approach, with CFSv2 being the reference as it has the higher (daily) resolution. Please check the number of dashes per week – to mind there should be seven and not six!
L168: complex terrain due
L184: Please check the formula for CRPS. CRPS is the difference (in area) between the cumulative distribution function (CDF) of the forecast and the CDF of the observation, represented by the Heaveside function if the observation is deterministic. This means it integrates the probability values for the predictand value from -Inf to +Inf. In my understanding of the formula as presented here the actual values are integrated (summed) rather than their (cumulative) probabilities. Perhaps this is due to the notation used, which may then be not be consistent with how it is used in PBIAS and ACC. Please check, and also that a correct implementation of the calculation is used.
L255: check spelling of models – here ECMWF is missing the F.Â
L612: In line with the reflection on using CFSv2 as a reference, I believe that this third conclusive point should be revised, as it does not mention the benchmark used when considering a model as skilful.Â
(Please not that this review was provided by the handling editor due to inability to find a second reviewer for this article after inviting over 20 potential referees).
Â
Citation: https://doi.org/10.5194/egusphere-2025-5062-RC2 - AC2: 'Reply on RC2', Lujun Zhang, 01 Aug 2026
Viewed
| HTML | XML | Total | BibTeX | EndNote | |
|---|---|---|---|---|---|
| 1,387 | 1,652 | 139 | 3,178 | 224 | 178 |
- HTML: 1,387
- PDF: 1,652
- XML: 139
- Total: 3,178
- BibTeX: 224
- EndNote: 178
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
This is an exhaustive analysis of the sub seasonal forecasts produced by a number of models that participated in two different model inter comparison campaigns. S2S and SubC. Â While I think this article should be published, it is more of a time capsule of simulations performed nearly a decade ago and experimental design that is somewhat limited by the computational limitations at that time. The number of inter and intra model ensembles is quite small compared to what would be feasible today at these spatial resolutions. In that sense I am not sure how much this work reflects the current model cap abilities that extend to convection resolving scales and much larger ensemble sizes. With that in mind a few additional comments:
a. Why was the choice made to scale everything to 0.25 degree, this is neither native of PRISM or the models. Â I suspect much of the "noise" or striations seen on the western mountainous regions in. figure 4 and 7 for example is probably an artifact of the remapping of the model outputs from 100km to 25 km (which seems to be a direct interpolation without consideration for terrain effects of other features).Â
b. It is remarkable that p[ast two weeks all the model variability of this (I amassing a mix of intra and inter model) variability collapses into approximately the same bias geographically (fig 5 for example). Â This seems to suggest that the deterministic equation setup in these models make this reach a climate 'state' after the initial conditions and no external perturbations (for example as in a climate models). Â Do all these models use the same SSTs and how often do they get updated?Â
c. Have you looked at the inter model variability (for example BOM or CNRM) that have larger ensemble size? Does the bias compared to PRISM for the inter model ensemble remain consistent with each other?Â
d. The CRPS results applied across different models with different initializations may be not appropriate? Â If applied to the same model ensemble it helps understand the model variability and predictability. Across different models it really not very useful as it much impossible to understand what makes these ensembles collapse.Â
e. Analysis of these model predictions under certain initial condition constraints (for example ENSO) could be separated to add additional value to the publication. Â My guess is that while there may be challenges in improving the overall predictability at these time scales, we may be able to find the contained time and spatial domains that would be more predictable under these special initial conditions. Â
f. The link to the ECMWF site for data doesn't work.Â
Â