the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
aiLand v1: Physics-Based Land Surface Emulator with Observational Fine-Tuning
Abstract. A stand-alone emulator of ECMWF's land surface scheme (ecLand) has been developed. This emulator, aiLand, uses a multi-layer perceptron architecture, chosen for its balance of accuracy and efficiency and for its differentiability, which is crucial for integration into data assimilation and parameter estimation systems. In this study, we introduce a two-stage learning framework that leverages both synthetic land surface model simulations and real-world observations. We first pretrain the surrogate on extensive ecLand outputs to capture the core dynamical behaviour of key land surface states, evaluating its accuracy, long-term stability, and transferability across variables, depths, and climates. We then fine-tune the pretrained model on in situ eddy-covariance flux observations for selected diagnostic variables, validating against independent flux-tower sites. The pretrained emulator reproduces ecLand's prognostic soil state with a 90-day RMSE of 1.19 K for surface soil temperature and 0.014 m3 m-3 for surface soil moisture, and remains stable over continuous 4-year autoregressive integrations. A single globally trained model outperforms biome-specialist baselines in cross-biome transfer, with residual errors concentrating in snow-insulated cold biomes where an insulating snowpack decouples the soil from atmospheric forcing. Fine-tuning on FLUXNET observations reduces latent heat flux RMSE by 30 % and sensible heat flux by 20 % at validation sites, while preserving prognostic state integrity and improving the physical consistency of the surface energy budget: per-site Bowen ratio error drops by 42 % and energy balance closure residuals fall from 13.4 W m2m-2 to 3 W m-2 and below. These results demonstrate that combining physics-based pretraining with observation-based fine-tuning provides a flexible pathway for building accurate, stable, and differentiable land surface emulators suitable for data assimilation and coupled modelling applications.
- Preprint
(17661 KB) - Metadata XML
- BibTeX
- EndNote
Status: open (until 07 Sep 2026)
-
CEC1: 'Comment on egusphere-2026-3620 - No compliance with the policy of the journal', Juan Antonio Añel, 07 Aug 2026
reply
-
AC1: 'Reply on CEC1', Nina Raoult, 18 Aug 2026
reply
Dear Prof Añel,
Thank you for your message and for the careful check of our code and data availability statement.
We understand that the concern regarding the training data relates to whether the ECMWF archive satisfies GMD's archive standards on long-term preservation and on preventing unilateral removal, rather than to the persistence of the identifiers themselves. The two identifiers we cite are DataCite DOIs registered by ECMWF (10.21957/0f6t-7f73 and 10.21957/fs25-c406) and resolve to permanent landing pages that specify the exact experiments and the commands needed to retrieve them.
The two training datasets total approximately 1.1 TiB: 1 TiB at n320 resolution and 113 GB at o96. Both are beyond the 100 GB threshold you mention and beyond the capacity of Zenodo or comparable general-purpose repositories, and transferring these volumes is not feasible for us within any reasonable timeframe or cost. We therefore respectfully request that the exception provided for in your comment be applied to these two datasets. We note that equivalent exceptions have been granted for ECMWF-hosted training data in two closely related cases: Moldovan et al. (https://doi.org/10.5194/gmd-19-4703-2026, see the reply at https://doi.org/10.5194/egusphere-2025-4716-AC4), and Wesselkamp et al. (https://doi.org/10.5194/gmd-18-921-2025), the latter concerning training data for an ecLand emulator generated in the same manner as ours.
We will also expand the code and data availability statement to make the provenance of the training data explicit:
"Training data are published at https://doi.org/10.21957/0f6t-7f73 (n320; Raoult et al., 2026a) and https://doi.org/10.21957/fs25-c406 (o96; Raoult et al., 2026b). These datasets are ecLand offline runs, generated using ecLand (Nawab et al., 2025; https://doi.org/10.5281/zenodo.21995581), forced by ERA5 (Hersbach et al., 2020), which is distributed through the Copernicus Climate Data Store (https://cds.climate.copernicus.eu)."
Nawab, A., Bardu, G., Piotrowski, Z., et al. (2025). ecmwf-ifs/ecland: ecLand 2.0.0 (Version 2.0.0) [Computer software]. Zenodo. https://doi.org/10.5281/zenodo.21995581
Finally, thank you for pointing us to the correct persistent identifier for the FLUXNET data. We will cite the FluxDataKit harmonised dataset as https://doi.org/10.5281/zenodo.14808331 in the reference list and in the code and data availability statement of any revised version of the manuscript.
We are happy to provide any further information that would be helpful, and we will of course follow your guidance on the final wording of the availability statement.
With best regards,
Nina Raoult, on behalf of the authors
Citation: https://doi.org/10.5194/egusphere-2026-3620-AC1 -
CEC2: 'Reply on AC1', Juan Antonio Añel, 18 Aug 2026
reply
Dear authors,
Many thanks for the reply. Given the size of the training dataset and the difficulties to store it elsewhere, we can grant you an exception in this case. Therefore, we can consider the new Code and Data Availability section and therefore the current version of your manuscript in compliance with the code and data policy of the journal.
juan A. Añel
Geosci. Model Dev. Executive Editor
Citation: https://doi.org/10.5194/egusphere-2026-3620-CEC2
-
CEC2: 'Reply on AC1', Juan Antonio Añel, 18 Aug 2026
reply
-
AC1: 'Reply on CEC1', Nina Raoult, 18 Aug 2026
reply
-
RC1: 'Comment on egusphere-2026-3620', Anonymous Referee #1, 10 Aug 2026
reply
This work introduces aiLand, an MLP-based emulator of ECMWF’s land surface scheme ecLand. It rigorously evaluates its simulation performance and stability, then demonstrates its utility in two major experiments. The first is a cross-biome extrapolation exercise with unsupervised cluster detection. The second is a two-stage training framework that ablates finetuning approaches, transferring the physics-trained model to Eddy Covariance flux measurements, and evaluates for its skill to correct physical biases. Multiple sub-analyses are conducted, including energy balance closure at station sites.
Thanks, wonderful work and exciting to read! The manuscript is highly polished and in principle ready for publication. Given I was asked for a minor revisions review, I’d ask the authors to address the specific comments below and leave the general comments to their consideration.
General comments:
An operational perspective drives the experimental design with two cool major experiments. Yet, one of them gets buried in the introduction and together they remain slightly fragmented. The broader context that is opened in the introduction remains underdiscussed, restricted to a short sentence in the conclusion. Better integrating them in the discussion could - aside releasing this new model version - make a strong case for pre-training on physical models when forecasting land states under (eco)climatic regime changes.
The finetuned emulator is evaluated for its bias correction, which will be useful as a coupled model component in NWP or AIWP. The introduction and experimental setup (especially experiment 1) also touch on the field of studying transfer learning capability of physically (weak-) constrained ML. The pretrained aiLand together with experiment 1 is in principle a perfect candidate to demonstrate if weak physical priors allow ML prediction under covariate shifts - with implications for upscaled products (see https://arxiv.org/abs/2605.19812). But the physics pre-trained model would need to be compared to an (afap) unconstrained baseline, the closest to which I believe is your scenario S5. As you also deem S5 to yield the best return in generalization skill, I would have liked to see this topic at least discussed, starting from what you’d expect if you trained aiLand on Flux stations directly. You use stations with >20 years of observations, an amount that may be enough for an MLP to learn better than an emulation of a biased target when physical model parameters are constant in time (see Wesselkamp et al. 2024). Perhaps generalization or drift at seasonal lead times if done prognostically will become a problem, and if a learned physical prior reduces it, that would make a strong case for establishing the two-stage training regime (showing how great emulators are, potentially).
The cross-biome experiment gets buried a bit across all sections in the manuscript, and most importantly, it’s goal and implications are not really stated. The methods and results section lack clarity and may profit from a formalization of the analysis that is shown in the results. My assumption about the goal is demonstrating the advantage of a global model, as the experiment is inversely conducted to a spatial cross validation - which in contrast would give an estimate of the true extrapolation error but means training on all but one biome and predict on the hold-out (i.e. Roberts et al. 2016, doi 10.1111/ecog.02881). The approach to manually identify clusters from features that are derived from the training data is a strength of this experiment: it could provide controlled extrapolation regimes, but we don’t see where exactly the clusters are subsets of the global covariate distribution and where not, hence where biome-specialists have to extrapolate. If this was not the goal, it raises the question, why not use established biome classifications (and their naming, see specific comments), or at least validate detected clusters with such (e.g. https://ecoregions.world/ for a qualitative definition based on Dinerstein et al. 2017, https://doi.org/10.1093/biosci/bix014). The identified clusters are meaningful, but their labels blur these advantages and instead highlight limitations.
Specific comments:
L 202: Your probably saying, diagnostic variables are not normalized again before loss computation, I assume they were standardized before
L 362: Is this asymmetry a modelling or a data result? Given that you target two of the diagnostics for finetuning on observational data, the autoregressive error structure of which will be real. I see that the diagnostic prediction is an advantage for gappy data prediction, perhaps avoids unstable rollouts - but was it ever considered to model your diagnostics prognostically, as does the AIFS with e.g. skin temperature?
Section 2.3.2. “data-driven biomes” reads a bit odd, perhaps unsupervised biome detection or alike may work.
Section 2.3.2. Naming conventions are useful for interpretation, but because they are established in existing biom and climate classifications, they are also confusing. E.g. the assignment of large parts of Europe to a subtropical humid cluster, instead of the temperate one. I don’t have a better suggestion though but strictly using their numbered labels throughout
L 250: Why was the period of the training feature space only partially considered for feature computation?
L 253: please specify the validity metrics
Table 3: Reporting e.g. maximum and minimum could better clarify where the models are extrapolating in the cross-biome exercises.
L 275/section 3.1.3/L. 628, … : I suggest using consistent names throughout the manuscript - replace ecosystem regimes with biomes, also because it’s climate-defined clusters
L 284: Is this characterizing short-range or seasonal error growth?
L 285: […] diverge under identical forcing but different initial conditions
Table 4: Do you know why you accumulate more error in deeper layers, where the relative error is smaller?
L 308: Interesting. Is this skewness only showing for stl1?
L 309: Should skewness not be unitless, or did you compute this as the difference between mean and median? then this is large for increments
Figure 7a: Ngl, it takes a while to follow the paragraph that describes these results. Part of it is terminology that is not uniquely defined (i.e. generalization error = transfer skill), or used inconsistently (L. 354: biome naming conventions already mentioned, tropical vs. tropical/subtropical, L. 359 tropical or rainforest –other paragraphs just use the shortages, maybe preferable)
Figure 7b : I don’t understand 7b from the caption description alone: What is the residual penalty under global training in contrast to the penalty recovered by the global model? How are the in- and cross biome evaluations related with the global model? Could you please define and explain these conventions somewhere, e.g. in section 2.3.2 ?
L 360: It could be helpful to visualize the statements in the plot, i.e. highlight the (tropical) rainforest-to-desert pixel and desert cluster.
L 368: This paragraph is most supported by figure 7a left panel and would fit after the one describing the 7a right panel results.
Table 6: I think an aiLand-S0 (Flux) is missing from this table, can be discussed.
Citation: https://doi.org/10.5194/egusphere-2026-3620-RC1 -
RC2: 'Comment on egusphere-2026-3620', Anonymous Referee #2, 14 Aug 2026
reply
Overall, I think this is a well-structured and valuable manuscript. The proposed framework provides a useful step toward observationally constrained land-surface emulation, and the manuscript is generally clear and comprehensive. I recommend minor revision, primarily to clarify several methodological choices and further discuss potential future developments.
Comments
-
The current study adopts an MLP architecture, which provides a good balance between computational efficiency and predictive performance. However, the current architecture does not explicitly represent spatial or longer-term temporal dependencies. It would be useful to expand the Discussion on potential future ML developments, such as architectures that can better capture spatiotemporal context (e.g., Transformer-based approaches) and the incorporation of explicit physical constraints, including energy and water conservation. Such discussion would help clarify how aiLand could evolve toward a more physically consistent and scalable land-surface emulator.
-
Section 2.3.2: The rationale for introducing a new data-driven “biome” classification should be clarified. Since the clusters are defined using climate and land-surface variables, including soil moisture and Bowen ratio, they may be better described as environmental or climate–land-surface regimes rather than ecological biomes. I also encourage the authors to discuss why established classifications such as PFTs or Köppen–Geiger classes are insufficient and, if possible, examine whether the main transferability conclusions are robust to the choice of classification.
-
The cross-biome transferability analysis may not provide a fully equivalent comparison because the biome-specialist models are evaluated on unseen biomes, whereas the global model has already been trained on all biomes, including the target biome. Thus, its better performance may partly reflect direct exposure to the target regime rather than stronger generalization. If computationally feasible, a leave-one-biome-out experiment, in which the global model is trained on all but the target biome, would provide a more rigorous test of transferability to unseen environmental regimes.
4. Line 253: Please specify which validity metrics were used to determine the optimal number of clusters, (k=7).
Citation: https://doi.org/10.5194/egusphere-2026-3620-RC2 -
Viewed
| HTML | XML | Total | BibTeX | EndNote | |
|---|---|---|---|---|---|
| 148 | 73 | 11 | 232 | 9 | 10 |
- HTML: 148
- PDF: 73
- XML: 11
- Total: 232
- BibTeX: 9
- EndNote: 10
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
Dear authors,
Unfortunately, after checking your manuscript, it has come to our attention that it does not comply with our "Code and Data Policy".
https://www.geoscientific-model-development.net/policies/code_and_data_policy.html
First, you provide the training data for your model by storing them in ECMWF servers. However, the ECMWF service does not fulfil GMD’s requirements for a persistent data archive because:
- It does not appear to have a published policy for data preservation over many years or decades (some flexibility exists over the precise length of preservation, but the policy must exist).
- It does not appear to have a published mechanism for preventing authors from unilaterally removing material. Archives must have a policy which makes removal of materials only possible in exceptional circumstances and subject to an independent curatorial decision.
If we have missed a published policy which does in fact address this matter satisfactorily, please post a response linking to it. If you have any questions about this issue, please post them in a reply.
The GMD review and publication process depends on reviewers and community commentators being able to access, during the discussion phase, the code and data on which a manuscript depends, and on ensuring the provenance of replicability of the published papers for years after their publication. Please, therefore, publish the mentioned data in one of the appropriate repositories and reply to this comment with the relevant information (link and a permanent identifier for it (e.g. DOI)) as soon as possible. We cannot have manuscripts under discussion that do not comply with our policy. If for some reason copying the mentioned data to a new server is not feasible (for example, due to a large size (e.g. more than 100 GB)), then, reply to this comment with the necessary information, and we could grant an exception to the application of the policy.
Also, I would like to note here that for the FluxNet data, the reference that you provide does not contain a permanent handler to access the data. The permanent handler in this case is a DOI: 10.5281/zenodo.14808331, and the corresponding link to the repository is https://doi.org/10.5281/zenodo.14808331. Please, fix the mentioned issue in any potentially reviewed version of your manuscript in the future, and if your manuscript is accepted for publication.
I must note that if you do not fix the mentioned issues, we cannot accept your manuscript for publication in GMD.
Juan A. Añel
Geosci. Model Dev. Executive Editor