the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
Intercomparison of run-time bias correction methods in LMDZ v6.3
Abstract. Despite progress in physical development and calibration, climate models still exhibit biases with respect to historical observations. As an alternative way to reduce them, run-time bias correction approaches have been developed, which consist in adding empirical tendency adjustment terms to the prognostic equations of some key physical variables. Although their ability to effectively reduce atmospheric circulation biases has been demonstrated, information is still missing regarding which method for estimating the adjustment terms is best suited for a given application. In this study, we implement a set of these methods in the atmospheric general circulation model LMDZ: nudging-based bias correction (the basis approach, a state-dependent version, and an iterative version), and the so-called climatological adaptive bias correction. Applying run-time bias correction on horizontal winds only, using these methods and varying some of their parameters, nine "bias-corrected versions'' of the model are created. They are evaluated using aggregate scores of global mean errors in circulation, temperature, and precipitation, as well as mid-latitude atmospheric variability features. A more regional perspective is also adopted, and a large region covering Europe and the North-Atlantic serves as a case study. It is found that, when evaluated on global aggregate scores, some versions outperform others. We also show that this does not prejudge the outcome on mid-latitude atmospheric variability features or at regional scale. No strict recommendation can be made regarding the optimal methodological choice, and great caution is advised. The choice should be guided by the model user's needs and priorities.
- Preprint
(25910 KB) - Metadata XML
- BibTeX
- EndNote
Status: open (until 08 Aug 2026)
-
RC1: 'Comment on egusphere-2026-1380', John Scinocca, 13 Jul 2026
reply
-
AC1: 'Reply on RC1', Aude Champouillon, 22 Jul 2026
reply
First of all, many thanks for your detailed comments. We would like to use the open discussion as an opportunity to get your insights on a particular point before we start reviewing the paper.
Regarding CABCOR adaptation simulation, the choice of the 1991 year as repeated forcing was indeed unfortunate. We could choose the year 1995, which seems to be free from the Pinatubo eruption's influence and which is not an El Nino year. Is there anything that comes into your mind that would make this year not suitable?
In any case, we wonder about using repeated forcings for the adaptation simulation in the CABCOR methodology. Even a non-volcanic and non-El Nino year is never entirely representative of the targeted period's climate. This means that, no matter the year one chooses, the derived ERBC terms will embody the specificity of the repeated forcing year. This, at least theoretically, might seem unsatisfactory. We wonder if you considered other approaches while developing and evaluating the CABCOR methodology? Could it be suitable for example to apply each year a different forcing, by randomly choosing a year within the 30-year long climatological period used for the adaptation simulation?
We look forwards to your insights on this.
Best regards,
Aude Champouillon
Citation: https://doi.org/10.5194/egusphere-2026-1380-AC1 -
RC2: 'Reply on AC1', John Scinocca, 22 Jul 2026
reply
I would have thought that 1990 might be a more suitable alternative to 1991 as it occurs before Pinatubo and its prescribed forcings lie approximately at the midpoint of the 1981–2020 training period. I had assumed that this would naturally be the year you would select. I should have made that recommendation in my review - apologies for the oversight.
I agree that there is no individual year that can be fully representative of the climatological average over the target period. In that sense, any single repeated year is necessarily an approximation – we just want to ensure that it is not a poor approximation. Alternatively one could use the climatological annual cycle of the forcings over the training period. However, I would expect the impact on the derived ERBC, and hence on the bias-corrected simulations, to be small compared with using a reasonable representative year such as 1990. ENSO shouldn’t be an issue in the repeated annual-cycle simulations, as the climatological annual cycle of sea-surface forcing over the training period should be used in the adaptation.
With respect to your suggestion to apply a different forcing each year selected at random from the range of the adaptation period, I would be concerned that it would introduce artificial interannual variability into the adaptation simulation, which is likely to cause the CABCOR adaptation to take longer to converge.
An alternative approach might be to do interannually varying AMIP simulations over the training period and sequentially cycle the forcings over that period as many times as needed to extend the CABCOR adaptation for convergence. However, to properly sample the forcings the length of the adaptation would need to be an integer multiple of the training period. For the reasons outline in my review, I believe that the implementation of CABCOR in your application converges much more quickly that in my original study. My guess is that, if you were to use the entire 40 year period (1981-2020) for the adaptation, you would likely be able to use just one interannually varying AMIP simulation for the adaptation.
In our own implementation of the CABCOR methodology within a regional climate model driven by reanalysis, we found that adaptation simulations based on AMIP forcing over 1979–2020 performed as well as those based on 1950–2020. This suggests that, at least for a strongly constrained regional model, the adaptation can be performed in one AMIP simulation. Whether the same would hold for a free-running global model is, of course, a separate question that would need to be tested.
Citation: https://doi.org/10.5194/egusphere-2026-1380-RC2
-
RC2: 'Reply on AC1', John Scinocca, 22 Jul 2026
reply
-
AC1: 'Reply on RC1', Aude Champouillon, 22 Jul 2026
reply
Data sets
Intercomparison of run-time bias correction methods in LMDZ v6.3 - LMDZ configuration files, output data and analysis scripts Aude Champouillon https://doi.org/10.5281/zenodo.18955446
Model code and software
Intercomparison of run-time bias correction methods in LMDZ v6.3 - model infrastructure and code Aude Champouillon https://doi.org/10.5281/zenodo.20410139
Viewed
| HTML | XML | Total | BibTeX | EndNote | |
|---|---|---|---|---|---|
| 86 | 34 | 8 | 128 | 9 | 7 |
- HTML: 86
- PDF: 34
- XML: 8
- Total: 128
- BibTeX: 9
- EndNote: 7
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
This study compares several methods for deriving empirical run-time bias correction (ERBC) terms for use during the integration of climate models. The comparison is conducted within the LMDZ atmospheric model by applying the resulting corrections to the zonal and meridional wind components. The objective is to compare the methods themselves and identify characteristics that may generalize beyond this particular model implementation. This is an important and worthwhile objective, as a clearer understanding of the relative strengths and limitations of different ERBC derivation methods would provide guidance for their application in a broader range of modelling systems.
The manuscript is generally well written and presents a substantial number of numerical experiments. However, I have several significant concerns regarding the scope of the manuscript, the conceptual framework used to compare the methods, and the experimental design used to evaluate the CABCOR approach. In my view, these issues limit the conclusions that can be drawn about the methods themselves.
Overall, I recommend major revision. The manuscript requires a reduction in scope by removing the new state-dependent ERBC methodology, which warrants a dedicated methodological paper before being included in a comparative study. The comparison of the remaining methods would be substantially strengthened by first establishing a common conceptual framework for this class of run-time bias corrections and by more carefully examining the influence of the out-of-sample validation strategy on the interpretation of the results. Finally, the current experimental design does not adequately investigate the behaviour of the CABCOR method, as its leading-order control parameter is not explored and the repeated annual cycle simulations on which it is based use a year influenced by the anomalous Mt. Pinatubo eruption. These issues should be addressed before the manuscript can support meaningful comparisons among the different ERBC derivation methods. My detailed comments follow.
John Scinocca
Major Comments
The inclusion of the state-dependent ERBC method (Section 2.2.3 and Appendix A2) raises a fundamental concern regarding the scope of this manuscript. Unlike the other approaches considered, this is a new bias-correction framework in which the corrective term is conditioned on diagnosed weather regimes. However, it receives only a one-paragraph description in the manuscript, and I found the details in Appendix A2 insufficient to fully understand how the methodology operates and how it relates to the other methods discussed in this study. Ultimately, its motivation, theoretical basis, algorithmic formulation, relationship to previous ERBC approaches, expected properties, and sensitivity to key design choices require a much more complete presentation. The authors themselves acknowledge in the Discussion section, “The lack of effect of making correction state dependent with this method may be the sign that our methodology is not adequate and must be adjusted.”
I do not believe this issue can be addressed through a conventional major revision, as doing so would fundamentally change the scope and balance of the manuscript. The state-dependent ERBC method warrants its own dedicated methodological paper before being incorporated into a comparative study. Consequently, I recommend that the authors remove the state-dependent ERBC from the present manuscript and instead develop it as a standalone study. As part of that study, a section could be reserved to evaluate the new methodology using the same diagnostics and evaluation framework employed here, thereby placing it into the context of the established ERBC approaches while allowing the methodology itself to be properly developed and peer reviewed.
With the removal of the state-dependent ERBC methodology, the manuscript becomes conceptually much simpler. All remaining approaches belong to a common class of run-time bias correction in which a cyclostationary, state-independent forcing is added to the right-hand side of the evolution equations of the corrected variables (u and v in this study). The methods differ only in how this corrective forcing (ERBC) is derived.
The manuscript would benefit from first establishing the common properties of this class of bias correction before introducing the individual ERBC derivation methodologies, perhaps by expanding Section 2.2.1, the Introduction, or both. In particular, the discussion should emphasize the following properties of this class of bias corrections:
Establishing these properties at the outset would provide a clearer framework for interpreting many of the later results. For example, from (1) it follows that the control represents a finalized and tuned model, which has climatological biases we would like to correct. It is therefore reasonable to expect that compensating errors exist within the tuned parameterizations to counteract those biases. Combining this with (2) and (3), one can anticipate that directly bias correcting a prognostic variable with a large climatological bias may disrupt those compensating errors as its basic state is altered and so made inconsistent with the tuning of the model physics. Consequently, some aspects of the simulation may degrade as a consequence of detuned physical parameterizations, despite the targeted variable itself improving by construction. This provides a general rationale for why directly bias correcting temperature caused problems in the LMDZ model and explains why other models (and even earlier versions of the LMDZ model) could successfully apply ERBC to T.
This framework also clarifies how the different ERBC methodologies should be evaluated. From (3), the primary comparison should focus on the corrected variables (u and v), since these directly reflect the effectiveness of the different ERBC derivation methods. In contrast, from (4), responses in variables such as temperature and precipitation are indirect consequences of the correction and are therefore more likely to depend on the characteristics of the particular host model rather than on the ERBC methodology itself.
Introducing this conceptual framework before presenting the individual methodologies would substantially improve the organization of the manuscript and help distinguish conclusions that are broadly applicable to the ERBC methodologies from those that are specific to the LMDZ implementation.
The rationale for using out-of-sample validation as the primary basis for comparing the ERBC formulations requires further consideration. The objective of this study is to compare alternative methodologies for deriving the ERBCs, that is, to determine which methodology most effectively reduces the climatological biases of the corrected prognostic variables. This is fundamentally a boundary-value problem in which climatological statistics define both the correction and its evaluation. It therefore differs from an initial-value prediction problem, where split-sample validation is the standard approach because information from the period being predicted cannot be used during model initialization or calibration. It is therefore not obvious that the same validation philosophy is appropriate here.
Furthermore, interpretation of the out-of-sample validation is complicated by anthropogenic climate change, which the authors have identified. Separating the training and validation periods changes the target climatology itself. Consequently, if one ERBC formulation performs better than another during the validation period, it is unclear whether this reflects greater methodological robustness or simply differences between the climatologies of the two periods. Without additional analyses, these competing explanations cannot be distinguished.
This issue is also difficult to reconcile with the manuscript's conclusion that temporal smoothing of the 6-hourly ERBC provides little benefit. If a 20-year training period is already sufficient to estimate a stable ERBC without smoothing, then the principal motivation for withholding an additional 20 years to guard against sampling variability appears substantially weakened. Conversely, if sampling variability is considered sufficiently important to justify split-sample validation, then the conclusion that an unsmoothed 20-year ERBC is adequately converged requires stronger justification.
If the authors retain the out-of-sample validation, then, at a minimum, Fig. 2 should also include results over the training period together with an assessment of the changing climatology. The in-sample validation would establish the relative performance of the ERBC formulations under the conditions for which the corrections were derived. The influence of climate change could then be quantified by comparing the reference climatological annual cycles of the training and validation periods. If RT(d) and RV(d) denote the corresponding reference annual cycles, then B(d) = RV(d) − RT(d) measures the change in the target climatology itself. The magnitude of B(d) could be summarized using the same global annual RMSE metrics employed in Fig. 2. The figure would then contain three complementary panels: (i) training-period validation, (ii) the existing validation-period results, and (iii) the RMSE associated with the change in the reference climatology. This would establish whether the differences among the ERBC formulations are large relative to the change in the target climatology itself.
Depending on the outcome of these analyses, the authors may wish to reconsider whether out-of-sample validation should remain the primary basis for comparing the ERBC formulations. If the changing climatology contributes substantially to the global RMSE, then the out-of-sample comparison may be unnecessarily confounded. In that case, validation over the training period would provide the clearest assessment of the competing methodologies, with the out-of-sample results serving as a complementary robustness analysis. Moreover, if the changing climatology proves important for the globally averaged RMSE, its influence may be impacting the two-dimensional latitude-longitude bias fields presented throughout the manuscript.
The goal of this paper is to compare different methodologies for deriving the cyclostationary correction term (ERBC) appearing on the right-hand side of the evolution equations for u and v, with the hope of drawing conclusions that extend beyond the specific LMDZ model used for this study. It is therefore essential that the leading-order control parameters of each methodology are adequately explored. This has been done for the nudging-based methodologies, where the authors investigate multiple values of the relaxation timescale together with the iterative formulation. However, this has not been done for the CABCOR methodology.
The manuscript implicitly treats the relaxation timescale, denoted by the single symbol tau, as though it represents the same control parameter in both the nudging and CABCOR methodologies. It does not. These two parameters have fundamentally different physical meanings. Following Scinocca and Kharin (2024, hereafter SK24), I will therefore refer to them as tau_R (nudging) and tau_C (CABCOR).
In the nudging methodology, tau_R is the timescale with which the model winds are relaxed toward the single realization of ERA5 meteorology. In contrast, tau_C is the timescale with which the climatological annual cycle of the model is relaxed toward that of ERA5. Unlike nudging, the individual meteorological realizations in the CABCOR adaptation simulation are entirely different from those in the reference. Consequently, there is no physical basis for assuming that tau_C = tau_R. Nevertheless, by considering only a single value of tau, the manuscript implicitly makes this assumption.
This assumption is evident in Fig. 1, where the CABCOR correction terms are substantially weaker than those obtained using the strongest nudging configuration, BC(N1d-it2). A meaningful comparison should allow each method to span a comparable range of correction strengths by varying its principal control parameter. Just as the nudging method is examined using different values of tau_R, the CABCOR method should likewise be examined using different values of tau_C. Based on Figure 1, this likely requires substantially smaller values, such as 12 hours and 6 hours for tau_C, so that the correction tendencies can reach magnitudes comparable to BC(N1d-it2).
Also, the BC(CA1d-y30), BC(CA1d-y70), and BC(CA1d-y110) experiments do not span the behaviour of the CABCOR method. They instead represent different stages in the convergence of a single CABCOR configuration with tau_C = 1 day. The analogy for the nudging method would be to compare ERBCs derived from 2-, 5-, 10-, and 20-year nudged simulations using the same value of tau_R. Such experiments do not span the method. They simply demonstrate convergence for a single parameter choice.
Similarly, CABCOR requires sufficient sampling to converge. In CABCOR, the annually updated correction term is initially derived from only a poor estimate of the bias-corrected model climatology. This can produce unstable behaviour, particularly for smaller values of tau_C, and is the reason why a gradual spin-up, in which tau_C is decreased toward its target value, is required during the early part of the adaptation. Beyond the spin-up, the relaxational nature of the formulation suppresses unstable departures. At the same time, as an increasing number of years contributes to the estimate of the bias-corrected climatology, any remaining instability is shifted to progressively longer timescales. This allows the CABCOR adaptation to keep such departures under control relative to the corresponding BC simulation using ERBCs extracted before convergence. Eventually, the metrics of the corresponding BC simulation converge toward those of the CABCOR adaptation, indicating convergence of the derived ERBC.
An important result of SK24 is that the CABCOR adaptation appears to approach its final metric values well before the corresponding BC simulations converge (see Fig 1 of SK24 – 3d, 2d, 1d). However, the very long convergence times reported in SK24 likely reflect the much more demanding implementation used there, in which corrections were applied throughout the atmospheric column and evaluated from the surface to 10 hPa. In the present study, ERBC is applied only to u and v, only between 850 and 200 hPa, and the evaluation is likewise limited to 200 hPa. The similarity of the BC(CA1d-y30), BC(CA1d-y70), and BC(CA1d-y110) experiments therefore suggests that convergence may be occurring much more rapidly than in SK24. However, the manuscript does not demonstrate this. Instead, it assumes that the convergence characteristics reported by SK24 apply here, which isn’t justified without verification.
I therefore suggest the following revisions:
- Remove material: Remove all but one of the BC(CA1d-y30), BC(CA1d-y70), and BC(CA1d-y110) experiments, since they appear to represent essentially the same converged result for tau_C = 1 day.
- Add material: Investigate several values of tau_C (e.g., 1 day, 12 hours, and 6 hours) so that the CABCOR method spans a range of correction strengths comparable to that of the nudging-based approaches.
- Convergence of the BC simulation and the CABCOR adaptation: The authors should investigate the convergence of the CABCOR adaptation as it appears to be much faster than in SK24 for the reasons mentioned above. Explicit convergence of the BC simulation to the CABCOR adaptation may not be necessary. As discussed above, the stability of the CABCOR adaptation simulation allows it to approach the final metric values more rapidly than the associated BC simulation. Once the adaptation metrics have stabilized, they should be representative of the eventual converged BC results. However, this should be demonstrated for at least one value of tau_C in the present implementation.
The choice of 1991 as the repeated forcing year for the 100-year CABCOR adaptation simulation raises significant concerns. The CABCOR method relies on a repeated annual cycle simulation to define the climatological correction term. The choice of repeated year therefore determines the forcing environment to which the model is repeatedly exposed throughout the adaptation.
The year 1991 is problematic because it includes the eruption of Mt. Pinatubo, one of the largest volcanic aerosol perturbations of the twentieth century. The reference describing the LMDZ6A model states that stratospheric aerosols are prescribed from the CMIP6 forcing data set as monthly and annually varying latitude-height fields of aerosol optical properties that interact radiatively with the model atmosphere. Consequently, the model responds prognostically to the volcanic aerosol forcing associated with the Mt. Pinatubo eruption.
Repeating the 1991 forcings for 100 years therefore repeatedly imposes the Pinatubo volcanic forcing throughout the adaptation. The resulting circulation response has been extensively documented. Enhanced stratospheric aerosol loading warms the tropical lower stratosphere, strengthens the equator-to-pole temperature gradient, alters planetary-wave propagation, strengthens the polar vortex, and produces significant anomalies in the Northern Hemisphere zonal circulation and Arctic Oscillation (Stenchikov et al., 2002). Multi-model studies have likewise demonstrated robust responses of the large-scale circulation and zonal wind fields to the Pinatubo eruption across CMIP-class climate models (Barnes et al., 2016).
These circulation responses involve precisely the variables being bias corrected in the present study. Consequently, the derived CABCOR ERBC is compensating in part for a repeatedly imposed volcanic circulation anomaly rather than solely for the model's climatological wind bias. This could influence both the magnitude and spatial structure of the derived ERBC and thereby compromise comparisons with ERBCs derived using the other methodologies, none of which rely on repeated annual-cycle simulations.
To avoid these issues, the authors should repeat the CABCOR adaptation using a representative non-volcanic forcing year. Alternatively, the repeated annual cycle could be constructed using a climatological volcanic aerosol forcing rather than that associated with the Mt. Pinatubo eruption. Without such a change to the experimental design, the current CABCOR simulations do not provide a reliable basis for evaluating the methodology.
Stenchikov, G., Robock, A., Ramaswamy, V., Schwarzkopf, M. D., Hamilton, K., & Ramachandran, S. (2002). Arctic Oscillation response to the 1991 Mount Pinatubo eruption: Effects of volcanic aerosols and ozone depletion. *Journal of Geophysical Research: Atmospheres*, 107(D24), 4803. DOI: https://doi.org/10.1029/2002JD002090.
Barnes, E. A., Solomon, S., & Polvani, L. M. (2016). Robust Wind and Precipitation Responses to the Mount Pinatubo Eruption, as Simulated in the CMIP5 Models. *Journal of Climate*, 29(13), 4763–4778. DOI: https://doi.org/10.1175/JCLI-D-15-0658.1.