An Emergent Warming-Linked Mode of Cloud Cover in Reanalyses: Systematically Missing in CMIP6 AMIP Simulations
Abstract. Cloud feedback remains the dominant source of uncertainty in climate projections, highlighting the necessity of rigorous cloud-based evaluations of climate models. Current assessments rely predominantly on cloud climatology and responses to internal variability, leaving cloud changes driven by historical warming largely unassessed. Here, we identify an emergent trend mode in total cloud cover (CLT) across multiple reanalysis products that is closely linked to global mean surface temperature. Using this warming-linked mode as the primary benchmark, we evaluate 13 CMIP6 AMIP simulations (1979–2014). While the models adequately capture global warming and internal variability in both temperature and CLT, this CLT trend mode is systematically absent in the simulations. Diagnostic regression reveals that this absence is characterized by a substantial underestimation of the response amplitude and large-scale spatial mismatches. This systematic deficiency points to shared structural limitations in current atmospheric models. Addressing this specific discrepancy offers a targeted pathway to constrain the forced cloud response, thereby reducing cloud feedback uncertainties in future climate projections.
This study utilizes EOF analysis to identify a warming-linked mode and an ENSO-linked mode in near-surface temperature and total cloud cover. By comparing reanalyses and CMIP6 model simulations, the study suggested that AMIP models systematically reproduce ENSO-linked variability mode but fail to capture the warming-linked mode in total cloud cover.
Overall, this work presents a useful technique for separating the forced response from internal variability in total cloud cover, along with metrics for quantitative model evaluation; thus, it has the potential to be a valuable contribution to the field and is worth documenting. That said, there are a few aspects of the manuscript that require further clarification and refinement to strengthen it. My main concerns include the robustness of the benchmark, the limitations of the validation period, and the assumptions underlying the model evaluation framework.
Main comments:
1. Robustness of the reanalysis-derived benchmark
The study identifies the second EOF mode of CLT in ERA5 and interprets as the “forced cloud response”, the validation and evaluation then built on this interpretation. However, it raises a few questions. First, although CLT PC2 has a tight correlation with the GMST timeseries, this does not strictly confirm causality. Is it possible that PC2 is instead related to other low frequency internal variability such as the IPO? In addition, as noted in the manuscript, CLT is s a diagnostic variable in reanalyses, which are still based on data assimilation models and cloud parameterization schemes. Given that there are large discrepancies among reanalyses, including their explained variance and spatial patterns, I wonder whether this “forced” signal is sensitive to model physics and whether this PC2 is indeed a robust signal in the real world. I hope the authors could provide additional evidence to support the use of ERA5 as a “truth benchmark” and/or acknowledge the potential limitations of the framework.
2. Limitations of validating reanalyses against a short satellite record
The study attempts to address the disagreement among reanalyses by comparing them with satellite data, which I think is an essential step. However, the validation has limitations, as there is only one satellite record, and the validation period 2003-2021 is much shorter than and different from the 1979-2014 model evaluation period. Is it possible that the drivers of “forced cloud response” vary between these two periods (e.g., aerosol forcing vs greenhouse gases forcing)? And could the authors clarify whether this would affect the robustness of the validation? In addition, I wonder whether the authors considered using other longer datasets, such as ISCCP or PATMOS-x, to perform the validation or even directly compare with AMIP simulations? I believe these datasets are available from 1983-2018, which is still shorted but overlaps with part of AMIP period.
On a related note, please clarify why the 2003-2021 period was selected. Based on the links in data availability section, ERA5, MERRA2, and MODIS are available to present, R-2 is available from 1979/01-2026/02, CRA appears to be available from 1979 to 2024.
3. Model evaluation using ERA5-derived PC
Since the AMIP models do not produce a warming-related CLT mode, the study regresses AMIP-simulated CLT anomalies onto the ERA5-derived PC2 of CLT. However, this evaluation relies heavily on the assumption that ERA5’s total cloud cover trend (its CLT PC2) is “truth”, and that a perfect model should replicate the same temporal evolution as ERA5. While I understand this is intended to provide a baseline for evaluation, I wonder how much physical insight is gained through this regression and through ranking models based on their similarity to the ERA5 EOF results. For example, what could we learn from the top-ranked models such as IPSL-CM6A-LR versus models like MIROC6 in terms of their performance in simulating total cloud cover? I generally agree with the authors that CMIP6 models systematically lack a warming-linked CLT mode that has been seen in ERA5, but I’m a bit concerned that the value of ranking is less clear. It would be beneficial if the authors could clarify on this point, or perhaps include a discussion on the limitation of evaluating models against ERA5.
4. Tone down the Conclusion
In the discussion section, the paper attributes the model deficiency to the “free-running” nature of atmosphere models, suggesting that a lack of atmospheric anchoring allows systematic errors to accumulate and attenuate forced responses. While the proposed reasoning is plausible to explain the difference between AMIP and reanalyses, the manuscript does not provide a causal link or evidence to support it. Ultimately, climate models must be free-running to predict the future, it seems less insightful to attribute a model deficiency to its lack of observational constraint.
In addition, I found the last paragraph somewhat off. The study is entirely based on AMIP simulations, but it takes a big jump to “critical concerns about fully coupled CMIP6 simulations”. There is no evident in the study to support that coupled models show apparent agreement with observation due to compensating errors. It is also unclear how it raises concerns for ECS and cloud feedback assessments. The two future steps are also somewhat vague. I recommend toning down the conclusion section and consider removing the last paragraph, as it includes a number of assumption and speculative claims and falls outside the scope of this study.
Other Comments:
Text and Figure: