the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
Fresh Eyes on CMIP Model Biases: Diagnosing Biases Across CMIP Generations and Implications for CMIP7
Abstract. The Coupled Model Intercomparison Project (CMIP) simulates standardized experiments spanning historical, future, and hypothetical conditions, to better understand the Earth system's evolution. Systematic biases remain a persistent limitation of coupled climate models. Within the framework of the CMIP, successive model generations have increased in complexity, resolution, and representation of Earth system processes, yet many long-standing biases remain across atmosphere, ocean, land, and cryosphere components, and their interactions and feedbacks. In this literature review, we identify the characteristics of key systematic biases in coupled models from CMIP6 and earlier CMIP phases to aid future comparisons to CMIP7. We introduce results of a community survey, designed to prioritize diagnostics for the Rapid Evaluation Framework (REF), and use these diagnostics for our review. Biases are shown to be interconnected across the Earth system, via sea surface temperature distributions, ENSO patterns, AMOC´s large scale transport biases, sea ice, the carbon cycle and radiation fluxes among others, with impacts on the Earth system response through metrics such as the Transient Climate Response and Equilibrium Climate sensitivity. The physical mechanisms underlying these biases and strategies to reduce them are presented, including improved parameterizations, increased resolution, and enhanced coupling among Earth system components. This review synthesizes current understandings of systematic biases across the Earth System, providing an assessment of biases in a set of critical diagnostics. In addition we provide a perspective on the advances in physical processes, higher resolution modelling frameworks, and targeted experiments and model evaluation frameworks expected in CMIP7. This review serves as an overview for end users, those new to CMIP data and those looking to connect different aspects of the Earth system.
- Preprint
(3865 KB) - Metadata XML
- BibTeX
- EndNote
Status: open (until 10 Sep 2026)
- RC1: 'Comment on egusphere-2026-3510', Anonymous Referee #1, 24 Aug 2026 reply
Viewed
| HTML | XML | Total | BibTeX | EndNote | |
|---|---|---|---|---|---|
| 334 | 128 | 18 | 480 | 19 | 22 |
- HTML: 334
- PDF: 128
- XML: 18
- Total: 480
- BibTeX: 19
- EndNote: 22
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
Summary: This paper highlights key biases across domains in Earth System Modelling. It is presented in a timely manner as we anticipate the next generation of CMIP model simulations. The paper does not put forth new scientific ideas but is a review of the literature on model biases and uses diagnostics from the CMIP AR7 Fast Track Rapid Evaluation Framework as exemplars. The paper thus serves to identify science gaps for future model development.
Comments:
Major:
1. A reader of this paper, a few years down the line may wonder why there is not a single figure describing any of the biases discussed in the paper. For that reason, it might be a good idea to not only motivate why this paper was written but why it was written this way - the last lines of the introduction give some idea idea but the title doesn't quite do this. The paper as such is a review/synthesis and does not actually diagnose biases across CMIP generations, so it would be good for the title to reflect that.
2. Line 894: Not sure if this sentence was intended to be written as it is now as it doesn’t quite read correctly.
However, the statement about high-ECS models being excluded is incorrect. Please see Chapter 7 (Section 7.5) of IPCC AR6 WG1. In particular note that it says : “Solely based on its ECS or TCR values an individual ESM cannot be ruled out as implausible, though some models with high (greater than 5°C) and low (less than 2°C) ECS are less consistent with past climate change (high confidence).”. The reasons for not using ECS and TCR directly from models in AR6 is also clearly outlined.
It would therefore be incorrect to make the claim that high-ECS models were excluded (as the sentence currently implies) since the IPCC AR6 does not back that up.
See Swaminathan et al, 2024 (https://doi.org/10.1029/2024EF004901) and Appendix C in Dunne et al, 2025 (https://doi.org/10.5194/gmd-18-6671-2025) for a discussion on why ECS alone should not be used as a basis for model sub selection for assessments. The discussion in the TCR section on this is noted but perhaps worth rewriting the conclusion in the ECS section as well.
Minor:
1. In the Abstract, the word “hypothetical” in the first sentence suggests that the scenarios could be imagined and not really true. Suggest using a phrase with scenario in it which is more the scientific terminology and appropriate.
2. If possible, it would be good to not have acronyms or at least acronyms that are not expanded in the Abstract. This would be helpful for readers outside the CMIP community.
3. Line 21: The phrase “dozens of unique” may not be factually correct. Some models share components and if one counts just the unique models, they may not be dozens. Maybe providing some approximate numbers instead would add weight to the claim that there are many models including many independent ones?
4. Line 24: Just a caution that in CMIP6, scenarios do not just include forcings but socio-economic pathways too. This should be made clear.
5. Line 32: Should that be “interactive” carbon cycle and not “active”? I think that might be the term used in the literature. (as opposed to prescribed).
6. Line 35: Suggested change to make it grammatically better - One aim of the CMIP6..was to answer the question - “What are..”
7. Line 38: Note to change the REF paper reference as the paper is now published.
8. Line 40: Should this be a reference to the Data Request paper as this is more a reference to where the 5 domains originated from? This might be https://egusphere.copernicus.org/preprints/2026/egusphere-2026-1641/
9. In Table 1 : Note that the final list of diagnostics in the REF paper has been revised for various reasons. Perhaps putting an asterisk where it says “final” and explaining this was a post-survey summary might clarify the discrepancy between this list and the REF paper list.
10. In Table1, Also, some of the numbering seems off and Snow cover doesn’t have a number. If there is a reason for this, perhaps an explanation can be provided so as to not confuse the reader.
11. Line 81: Suggest changing “these both” -> “Both these”.
12. In Figure, 3 could the authors clarify how the boxes for Key Diagnostics an Other Diagnostic Processes differ? It is not clear how that legend item is used.
13. Line 144: Perhaps not quite correct to call these “inaccurate”, they can be “inadequate” or some other term that the authors can think of. Inaccuracy can be read as models doing something wrong but that’s not quite the case.
14. Line 136: It is not clear how the authors are linking biases in SST to reasons through Figure 3. It would be helpful to call attentions to specific boxes and highlight some of the reasons. Suggest that this should be replicated for other figures as well to help drive home the message from this manuscript.
15. Line 216: Suggest removing the word “inaccurate” as “error compensation” conveys the message here.
16. Line 294: Perhaps the authors could qualify that most CMIP models do not have interactive Greenland Ice sheets. UKESM 1.2 (https://egusphere.copernicus.org/preprints/2025/egusphere-2025-4476/egusphere-2025-4476.pdf) and I think IPSL have interactive Greenlan ice sheets. There doesn't appear to be a documentation or reference paper for IPSL yet though.
17. Line 373: I am not sure how we would characterize or quantify all ENSO teleconnection biases since we may not even be fully aware of all possible teleconnections? Would the authors consider including this aspect of teleconnections as well?
18. Line 717: Suggest changing “culprit” to phenomenon or feature as that would be more appropriate.
19. Section 4.2 : Suggest including a recent publication from the CMIP7 Model Benchmarking Task Team that discusses observational uncertainties in model evaluation :
https://doi.org/10.1175/BAMS-D-25-0079.1