the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
Equation discovery for climate impact: emulating impact models for unexplored climate scenario with interpretable symbolic regression
Abstract. Projected impacts of climate change are assessed with impact models, such as ecological or hydrological models, driven by climate projections. Uncertainties of projected impacts are estimated by driving impact models with a large ensemble of plausible future climate projections. However, most of the time, this is not possible for practical reasons: computing time, data availability. To fix this issue, we propose an approach that links climate projections to impacts with an interpretable equation. First, this equation is discovered based on simulations of the impact model and their corresponding climate projections. Then, we consider that this equation can be used to emulate the impact model for other climate projections. Specifically, the discovered equation maps each year climate indicators, i.e. a list of yearly and seasonally-averaged climate model variables, to a yearly-averaged impact indicator, i.e. a variable computed from the impact model outputs. In our application, the impact indicator is the annual mean Net Primary Production (NPP) of a risk-relevant regional oceanic area located in the North-Western Mediterranean basin. It is computed from the outputs of a biogeochemical model of the Mediterranean Sea, which is driven by climate projections of a coupled regional climate model of the Mediterranean area. In our methodology, we run nine validation schemes each one providing one equation to predict this impact indicator. Our results show that all discovered equations are linear, even though non-linearity is allowed, and that most of them contain four climate indicators that can be interpreted physically: the sea surface temperatures in winter and spring, the sea surface salinity in spring, and the net downward shortwave flux in winter. Based on these four indicators, we fit a linear equation on the historical period (1986–2005) and the scenario RCP8.5 (2006–2099) that reproduces well the trend and the year-to-year correlation of the impact indicator for the scenario RCP4.5 (2006–2099), which was not used for the fit. The predictions of the linear equation however underestimate the interannual variance of NPP. As a perspective, this equation allows us to approximate the impact indicator at a neglected computational cost, i.e. without running the costly biogeochemical impact model, for any regional climate model outputs available.
- Preprint
(1570 KB) - Metadata XML
- BibTeX
- EndNote
Status: open (until 07 Aug 2026)
- RC1: 'Comment on egusphere-2026-1991', Anonymous Referee #1, 30 Jun 2026 reply
-
RC2: 'Comment on egusphere-2026-1991', Anonymous Referee #2, 01 Aug 2026
reply
The manuscript proposes the application of automatic equation discovery, also known as symbolic regression (SR), to derive equations that map annual climate indicators (derived from a climate model) to an impact indicator (computed from an impact model). The paper focuses on the methodology and presents a case study.
For the case study, the authors show that equation discovery is not only able to describe known relationships between some predictors and the predictand, but also identifies new potentially relevant predictors. They also identify weak points of the methodology and propose possible ways to mitigate these weaknesses, which is very important given that the information provided by the methodology is needed for high-stakes decisions.
The paper is well written and the methodology is promising. However, the fact that all discovered equations are linear may potentially limit the practical applicability of the methodology. I therefore consider that the manuscript could be published after major revision. Before this, the authors should address the following comments.
MAJOR COMMENTS
- All discovered equations are linear. This requires explicit justification. In the Western Mediterranean, NPP is regulated by mechanisms that are intrinsically non-linear, both in the predicted field and in the predictor variables. For instance, phytoplankton growth rates, nutrient dynamics, and the photosynthesis-irradiance relationship are not captured by linear equations. Furthermore, the predictor variables exhibit non-linear long-term evolution, both in the historical period and under the RCP scenarios considered. Therefore, the authors should discuss explicitly whether the linearity result reflects a genuine physical property of the system or an artefact of the annual-averaging preprocessing and the complexity penalties inherent to the equation discovery algorithm. I am not impliying that linearity is an algorithmic artifact. Annual averaging acts as a low-pass filter that can smooth non-linear functional responses. Therefore, the authors should explicitly discuss whether the linearity result reflects a genuine emergent property of the aggregated system or an artifact of the preprocessing and/or the algorithm. SR algorithms often aggressively penalise non-linear terms (divisions, exponentials, logarithms) because they increase equation length. If the penalty were relaxed, do non-linear terms (e.g., interactions or quadratic terms) appear, even if later pruned?
- The authors find that the linear equation underestimates interannual NPP variance, which is relevant for the interpretation of the findings. It is therefore necessary to quantify the magnitude of this underestimation relative to the total interannual variance of NPP. Furthermore, the implications for the use of these equations as operational emulators for future NPP projections should be discussed, as interannual variability is critical for economic applications and the ecological systems.
- The systematic emergence of linear equations may reflect a structural bias of the methodology towards linear solutions, or it may reflect that the annual aggregation obscures scale-dependent non-linearities. A potentially valuable extension would be to apply a multi-scale decomposition approach such as Complementary Ensemble Empirical Mode Decomposition (CEEMD; Yeh et al., 2010) to both predictors and NPP prior to equation discovery. This could help isolate relationships that are individually linear at each timescale, even if the aggregate relationship is non-linear, and might better reproduce the interannual variance that the current linear equations underestimate.
MINOR
- Change “As a perspective, this equation allows us to approximate the impact indicator at a neglected computational cost, i.e. without running the costly biogeochemical impact model, for any regional climate model outputs available” to "From this perspective, this equation allows us to approximate the impact indicator at a negligible computational cost, i.e., without running a costly biogeochemical impact model, using any available regional climate model output."
- The absence of MLD from validated equations may reflect its thermodynamic definition, which captures buoyancy-driven stratification rather than the dynamical vertical mixing most directly relevant to nutrient supply. A more physically appropriate predictor for turbulent mixing could be the Ekman layer depth (h_E), which directly characterises wind-driven momentum penetration into the upper ocean. I suggest that the authors explicitly discuss this physical mismatch in the limitations section. If wind stress or wind speed data are readily extractable from the VHR-GCM and RCSM outputs used in this study, testing h_E (or a wind-based proxy) as an additional candidate predictor in a supplementary sensitivity analysis would be a valuable addition. However, if this requires substantial new data extraction and reprocessing, I do not consider it a strict prerequisite for acceptance; a thorough discussion of this limitation and a recommendation for future work would suffice.
Yeh, J.-R., Shieh, J.-S., and Huang, N. E. (2010). Complementary ensemble empirical mode decomposition: A novel noise enhanced data analysis method. Advances in Adaptive Data Analysis, 2(2), 135–156. https://doi.org/10.1142/S1793536910000422
Citation: https://doi.org/10.5194/egusphere-2026-1991-RC2
Viewed
| HTML | XML | Total | BibTeX | EndNote | |
|---|---|---|---|---|---|
| 261 | 65 | 11 | 337 | 23 | 19 |
- HTML: 261
- PDF: 65
- XML: 11
- Total: 337
- BibTeX: 23
- EndNote: 19
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
The paper proposes a symbolic regression (SR) approach to model climate impact for climate scenarios at lower computational cost. The methodology is well presented but although SR allows nonlinear relationships, the final model is a simple linear equation with 4 variables. The work is promising but requires further uncertainty analysis and baseline comparison between models.
Major comments
Minor comments