the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
Enhancing Parameter Calibration in Land Surface Models Using a Multi-Task Surrogate Model within a Differentiable Parameter Learning Framework
Abstract. Land surface models (LSMs) are essential for simulating terrestrial processes and their interactions with the atmosphere, but parameter calibration remains challenging because of complex process coupling, parameter uncertainty, and limited observational constraints. In this study, we introduce multi-task differentiable parameter learning (MdPL), a deep learning framework that combines a multi-task neural surrogate with a differentiable parameter generator for efficient and diagnostically transparent LSM parameter calibration. The framework was evaluated at 20 PLUMBER2 sites spanning four plant functional types (PFTs). Because MdPL optimizes multiple parameters simultaneously through a neural surrogate, we first assessed surrogate–ILS consistency under joint parameter perturbations. Eighteen of the 20 sites satisfied the consistency criterion, defined as median Pearson’s R > 0.9 for both sensible and latent heat fluxes, and were retained for the main calibration analysis. Across these consistency-screened sites, the MdPL-calibrated Integrated Land Simulator (ILS) reduced RMSE by 14.1 % and 13.9 % at the PFT-mean level for sensible and latent heat flux simulations, respectively, relative to the default parameter set. Benchmarking with the PLUMBER2 dataset showed that the MdPL-calibrated ILS achieved competitive performance relative to standard LSMs, including CLM5, JULES, Noah, and GFDL, and performance comparable to an LSTM benchmark, although the relative advantage varied across variables, sites, and temporal resolutions. A comparison with a traditional calibration baseline based on Sobol sequence sampling at four representative sites further showed that MdPL identified competitive parameter solutions without exhaustive full-model evaluation, but its advantage remained site- and metric-dependent. An out-of-target evaluation using gross primary productivity (GPP) suggested that calibration based on sensible and latent heat fluxes did not lead to systematic degradation of a non-target carbon-related flux. Parameter plausibility analysis showed that some calibrated solutions approached predefined parameter bounds, indicating limited parameter identifiability under flux-only observational constraints. Therefore, the calibrated parameters should be interpreted as effective parameters under the given model structure and observational constraints, rather than as uniquely identifiable physical quantities. Overall, MdPL provides a scalable and diagnostically transparent approach for improving parameter calibration in complex LSMs.
- Preprint
(1907 KB) - Metadata XML
-
Supplement
(585 KB) - BibTeX
- EndNote
Status: open (until 23 Oct 2026)
- RC1: 'Comment on egusphere-2026-4009', Anonymous Referee #1, 19 Sep 2026 reply
-
RC2: 'Comment on egusphere-2026-4009', Anonymous Referee #2, 02 Oct 2026
reply
This study presents MdPL, a calibration framework pairing a multi-task neural surrogate of the ILS land surface model with a differentiable parameter generator. Twelve PFT-level parameters are calibrated against sensible and latent heat fluxes at 20 PLUMBER2 sites covering four plant functional types, with a surrogate reliability screening, a benchmark against four LSMs and an LSTM, a comparison with Sobol sampling, and an evaluation against GPP as a non-target flux.
It is an interesting and timely study, relevant to HESS. The manuscript is clearly written, and I was glad to see a surrogate reliability check and a parameter plausibility diagnostic included at all, since both are often missing from this kind of study. Nevertheless, I have a few comments that must be addressed before the manuscript can be published.
Most importantly, I could not find any out-of-sample evaluation. The generator is trained on the flux observations and the parameters are averaged over the full training period (L320–321), so the 14% RMSE reductions and the comparisons against the other LSMs appear to be calculated over the same period used for the calibration. As it stands, these numbers show how well the model can be fitted rather than how well it predicts. Could the authors state which periods were used for training and evaluation, and add a hold-out test? Related to this, the only test of generalisation is the leave-one-site-out analysis in the supplementary material, which shows that parameters transfer poorly, worse than the default parameters for woodland latent heat and for both fluxes at cultivation sites. Since scalability is claimed in the abstract and conclusions, I think this analysis belongs in the main text.
I am also not convinced that the surrogate screening tests the property the calibration relies on, and no results are reported with uncertainty. Finally, I feel there’s a missed opportunity to discuss what the calibrated parameters themselves mean for the model. I believe one of the main strengths of parameter calibration is what it teaches us about the base model and how we can use that information to guide future model development and understanding. For example, several parameters that shift in the same direction at nearly every site can tell us something about the model rather than just about the fit.
I hope my comments will help clarify and strengthen the manuscript. Below are my specific comments:
Specific comments
Abstract
L37: "did not lead to systematic degradation of a non-target carbon-related flux" rests on one site per PFT (L550), and GPP degrades at IE-Dri (L557). The conclusions (L586–590) are more careful, and the abstract should be brought in line with them.
L41: "scalable", see my comment on the transferability analysis below.
Introduction
L108, L133, L195 and L213–217: the claims that the gates reduce overfitting and improve stability and generalisability, that the day-of-year encoding gives a robust representation, and that the framework is robust under sparse and noisy observations, are not tested anywhere I could see. These should be phrased as motivation, unless the authors can add an ablation with and without the gates
Data and ILS (Sect. 2.3)
L229 and L240–245: the site selection needs more justification. The records are approximately 1–3 years, and four sites (AR-SLu, DK-Fou, PT-Mi1, US-Bar) have only a single year, which is short for training the surrogate and for calibration, and leaves no room for a hold-out period. Twenty sites is also a small subset of what is now available, both within PLUMBER2 (42 sites) and in the larger flux collections that have appeared since (FluxnetShuttle). The authors state that the subset balances ecosystem diversity, data completeness and computational feasibility, but it would help to say which of these is the binding constraint, presumably the roughly 150 ILS simulations needed per site for pretraining and screening. How many sites in the archive would have met the data requirements? This matters for the transferability analysis in Sect. S2, where only 3–5 sites per PFT are available, and where more sites per PFT would be the obvious way to make the test meaningful.
L247–249: ILS is described in two sentences. Which version and configuration were used, what are the forcing variables and units, and what is the time step?
I could not find any mention of spin-up. Parameters such as psicr, vgcov and vmax0 change how quickly soil water is drawn down, and therefore the equilibrium soil moisture. Was each perturbed simulation spun up separately, for the 48 pretraining runs, the 100 screening runs, the 256 Sobol runs and the final calibrated run? If not, differences between runs will partly be drift from inconsistent initial states, which over 1–3 years could be a large part of the signal, and it would also add to the cost of the Sobol baseline.
Pretraining and surrogate diagnostics (Sect. 2.4–2.5)
Fig. S1: the sensitivity analysis appears to have been run at a single site, ES-LgS, which is not one of the 20 study sites. Could the authors explain this, and justify using one site to select parameters for all four PFTs? The size of the candidate pool would also be useful.
L269–275: the surrogate is pretrained on one-at-a-time perturbations only, then used to optimise 12 parameters jointly. Can the authors comment on the consequences, and ideally show whether adding some joint samples to the pretraining changes the calibrated parameters?
L283–291: I am not convinced this screening tests what the calibration depends on. Pearson's R between two half-hourly flux series is dominated by the diurnal cycle, so a surrogate responding only weakly to the parameters could still pass a threshold of 0.9, and R ignores bias, which is what the MSE-based calibration acts on. Could the anomalies relative to the default run be compared instead, (ILS(p) − ILS(p₀)) against the same for the surrogate, with RMSE or bias reported alongside R? It would also help to justify the 0.9 threshold, to say how the 100 joint samples were drawn, and to report the surrogate error at the calibrated parameters themselves, since that is where it matters most.
On a related point, the surrogate sees only forcing, parameters and day of year over 48-step sequences, with no soil moisture or other ILS state. How can it then represent parameters that act through soil water storage over weeks to months, as psicr does (L264–266)? Is the LSTM hidden state carried over between sequences or reset?
Experimental design (Sect. 2.6)
L320–321: the generator produces time-varying parameters but ILS is run with their average, so the vector actually used was never itself optimised. It would help to see p(t) at a few sites, and to compare surrogate performance under p(t) with that under the average.
Eqs. 1–2: how are the two losses weighted, and are both computed on standardised fluxes?
L345–348: excluding the two sites that fail the screening will bias the reported means upwards so could the results for all 20 sites also be given? ID-Pag failed the screening but is still used in the multi-task comparison (Table S5), where it shows the largest gain of the four sites, and in the leave-one-site-out analysis. Why?
L364, L366 and L534–540: are the Sobol ranges the same as the Table 2 ranges available to MdPL? The Sobol sets are also ranked on RMSE(Qh) + RMSE(Qle) in physical units, while MdPL minimises a standardised loss. The efficiency argument would be much more convincing with wall-clock or CPU/GPU costs for both approaches, and I am not sure pretraining "can be amortized… across multiple sites" when it uses 48 runs per site and one model is trained per site.
Results (Sect. 3)
L419: "significantly improved" reads as a statistical claim, but no test is reported.
No results are given with any uncertainty: no repeated seeds, generator ensembles or confidence intervals. Given the equifinality the authors rightly discuss, the spread of p* across random initialisations would be interesting in itself and would show whether the gains are robust. Site-level metrics would also be easier to read as a table than from the Taylor diagrams.
L420–428: the PFT-mean weights each PFT equally despite 3–6 sites per PFT; could the site-mean also be given? The decrease in latent heat correlation at cultivation sites (−4.8%) is interesting and could be explored a little more.
Fig. 4: the diagnostic is confounded for vgcov, whose default of 1 is also its upper bound, so the default itself counts as boundary-proximal. Counts per parameter would help.
My main comment on this section is that the calibrated parameters are called "effective parameters" and then left there. In Tables S1–S4 several move the same way almost everywhere: vgcov is reduced at 17 of 18 sites, cdl increased at 16, vmax0 reduced at 15. That consistency suggests a systematic bias rather than site-specific noise. Lower cover and lower photosynthetic capacity would be consistent with the model over-estimating transpiration, or mis-partitioning energy between canopy and soil. Alternatively, since tower Qh + Qle usually falls short of the available energy, calibrating to uncorrected fluxes would push the parameters to suppress turbulent fluxes: which flux product was used here, corrected for energy balance closure or not? Could these shifts be reported and related to the biases of the default run in Qh, Qle and their partitioning? That would make the calibration genuinely useful for model development. And if vmax0 falls nearly everywhere, would we not expect GPP to fall too?
Supplementary material
Sect. S1: the multi-task and single-task results are close and go in both directions (mean Qh RMSE 42.82 vs 41.43 W m⁻², Qle 37.8 vs 39.5). With four sites and no uncertainty, I am not sure this supports "strong collaborative effects between key variables" (S1, L51–52).
- Sect. S2: as above, this analysis belongs in the main text. A few questions on it:
- Why is grassland excluded, and why are only 3 of the 5 cultivation sites used?
- How are the held-out site's parameters produced: as one set per PFT, or by applying g_z to that site's attributes?
- How well does the multi-site generator fit the sites it was trained on? Only the held-out performance is given.
- Is the surrogate trained per site, or pooled across sites?
This is also the only place where the site attributes can matter. In the single-site experiments A is constant during training (L160–163), so it cannot inform the parameters.
Discussion and conclusions
The framework is described as scalable (L584–585, L597), but as implemented it is trained per site, needs flux observations and depends on passing the surrogate screening. How would it be applied to grid cells where no flux observations exist? I would also welcome a paragraph on the limitations of the study; those given in Supplement Sect. S2 would sit more naturally here.
Minor comments and typos
L45: I do not think Exbrayat et al. (2014) supports this statement, that paper is about microbial decomposition and spin-up in CMIP5 models.
L69–71: something has gone wrong with this sentence: "dPL, e.g., the variable infiltration capacity (VIC) model, leverages…".
L76: HBV is attributed to Tibangayuka et al. (2022); the original HBV reference should be cited.
L122–123: CLM5 is cited to "NCAR, 2020" and JULES to McNeall et al. (2024); the model description papers would be more appropriate.
L226: 20 sites concentrated in Europe and East Asia do not really "ensure geographic and ecosystem diversity", I would soften this.
L17: ILS is used before it is defined (L19).
L105 and L108: the g_z notation has not rendered properly – "(g_z(\cdot))".
L128: "leave-one-out" should be "leave-one-site-out".
L188–195: L, the number of encoding frequencies, is never given.
L177 and Fig. 1: the generator's outputs are called "parameter priors" but are used as point estimates, which suggests a Bayesian treatment that is not carried out.
L230: the forcing variables and their units are not listed.
L326: 48 half-hourly time steps is 1 day, not 2 days.
L350: SFE is used without being described – a sentence would help.
Eqs. 16–17: ORI is not defined at first use.
L552 and Fig. 7 caption: observed GPP is attributed to PLUMBER2 in the text and to FLUXNET2015 in the caption.
Table 1: AU-Emr has no climate class; PT-Esp, PT-Mi1 and AR-SLu are listed as "Subtropical", which is not one of the classes defined in the note; and PL-wet, a wetland site, is grouped with grassland.
Table 2: the units of cdl and chl are given as "s m-2", which does not look right for exchange coefficients. Fortran notation ("9.82d-2") is also mixed with plain decimals.
References: please check the entries for Gupta (2008), Kuppel (2014), Vrugt (2005) and Rosolem (2012) and Raoult (2024), which is no longer a preprint
Citation: https://doi.org/10.5194/egusphere-2026-4009-RC2 -
RC3: 'Comment on egusphere-2026-4009', Anonymous Referee #3, 05 Oct 2026
reply
This study presents a multi-task differentiable parameter learning (MdPL) framework for calibrating parameters in the Integrated Land Simulator (ILS). The framework combines a multi-task neural surrogate with a parameter-generation network, and the resulting parameter sets are evaluated using the original ILS. The study includes surrogate-ILS consistency screening under joint parameter perturbations, benchmarking against other land surface models and an LSTM, comparison with a Sobol-based calibration baseline, and an out-of-target evaluation using GPP.
Overall, I found the manuscript well written and well organised, and the methodological workflow is generally easy to follow. The consideration of surrogate reliability and non-target flux behaviour is valuable. However, several questions concerning independent evaluation, the construction of the final parameter vector, and the interpretation of the methodological and generalisation claims need to be addressed before the conclusions can be fully supported. I therefore recommend major revision.
Some comments below request clarification of the existing experiments, whereas others suggest additional diagnostics. These diagnostics are intended to strengthen particular claims rather than require an extensive new experimental programme. Where additional analyses are not feasible, clearer limitations and more cautious interpretation would be appropriate.
(1) Lines 320-335, Eq. (2), and Fig. 1. Please clarify the temporal partitioning used for surrogate training, parameter-generator training, validation or model selection, and final evaluation. The current description does not establish whether the main performance estimates are independent of calibration. For the fixed-parameter application evaluated here, an independent temporal test should use parameters estimated from the calibration period and held fixed during evaluation, with standardisation statistics calculated using training data only. Fig. 1 indicates that historical flux observations are supplied to the generator; please explain their temporal relationship to the evaluation targets and reconcile this description with the definition of the auxiliary inputs in the text.
(2) Lines 280-292 and Fig. 2. The joint-perturbation screening is useful, but correlation between complete flux time series does not adequately characterise bias or errors in parameter-dependent responses. Shared meteorological and diurnal variation could produce high correlations even when parameter effects are imperfectly represented. Please supplement R with RMSE or bias, preferably including comparisons of changes relative to the default simulation and surrogate errors at the final calibrated parameters. It would also be helpful to describe how the 100 joint samples were generated and justify the median R > 0.9 threshold.
(3) Eqs. (1)-(2), around lines 145-173. Please explicitly define how the Qh and Qle losses are combined during surrogate training and parameter calibration, including whether standardised values are used, whether relative weights are applied, and how missing observations are handled. The Sobol baseline selects parameter sets using RMSE(Qh) + RMSE(Qle), which is not generally equivalent to minimising a sum of MSE losses. Although common performance metrics are subsequently reported, please explain whether differences in the optimisation or selection objectives affect the interpretation of the comparison.
(4) Lines 320-321, Table 2, and Tables S1-S4. The supplementary captions describe clipping unconstrained parameter estimates, but Table S2 reports rlfn values of 0.23 and 0.22 for DK-Fou and IE-Ca1, below the lower bound of 0.30 in Table 2. Please resolve this discrepancy and explain the order of optimisation, clipping, and temporal averaging, identifying the final parameter vectors used in ILS. Because the surrogate is nonlinear, optimisation using time-dependent parameters is not equivalent to optimisation using their mean. For a few representative sites, it would therefore be useful to compare surrogate performance using the generated parameter sequence and the final constant vector, together with a brief description of parameter variability before averaging.
(5) Table 2, Fig. 4, and lines 591-595. The methods appropriately describe boundary proximity as a potential indication of limited identifiability, but the conclusions use stronger wording. Please retain the same caution throughout: boundary-proximal solutions may reflect parameter compensation, unsuitable bounds, or a constrained optimum and do not alone establish non-identifiability. The interpretation of vgcov should also account for its default value already lying at the upper bound. In addition, please clarify whether relevant joint physical constraints are enforced. For example, if rlfn and tlfn represent reflectance and transmittance in the same spectral band, their sum should not exceed one, whereas the listed independent ranges permit combinations above this value.
(6) Lines 320-338 and 458-461. The comparison between multi-task and single-task configurations is useful, but differences in parameter count and learning-rate strategy limit attribution of the results specifically to task sharing or gating. Please distinguish evidence supporting the complete framework from evidence supporting its individual components. If improved generalisation or reduced overfitting is attributed specifically to the gates, a focused ablation would be informative. Otherwise, these statements should be presented as design motivations rather than demonstrated findings. I do not consider a broader architecture search necessary.
(7) Lines 240-245 and 334-335, Supplementary Section S2, and the conclusions. The leave-one-site-out analysis provides useful evidence of transferability, including its limitations, and its principal findings should be summarised in the main text. Please clarify how parameters are generated for the held-out site and how the framework would operate without local flux observations. Since the principal implementation trains a separate model at each site, computational scalability should be distinguished from demonstrated spatial generalisation. The abstract and conclusions should reflect the scope of the available spatial validation.
(8) Lines 345-355, Fig. 5, and Section 3.2. Please describe the training and calibration information available to the LSTM and LSM benchmarks, because common meteorological forcing alone does not establish equivalent evaluation conditions. SFE and the other methods also use different site sets; comparisons over common sites and valid timestamps would be preferable for ranking their performance. Please briefly describe SFE and its implementation, together with temporal aggregation and missing-data rules. When interpreting performance across resolutions, distinguish error cancellation caused by aggregation from a sustained advantage relative to the benchmarks, and qualify the temporal-robustness claim accordingly.
(9) Lines 545-562 and Fig. 7. The manuscript already acknowledges site-dependent GPP responses and the need for broader evaluation. Nevertheless, the magnitude of the IE-Dri deterioration deserves explicit discussion: RMSE increases from 1.95 to 3.93 micromoles of CO2 per square metre per second, and NSE decreases from 0.64 to -0.46. Please discuss the associated seasonal response or parameter changes where possible and distinguish relative improvement from absolute performance. The abstract should reflect the mixed outcomes and limited site coverage as carefully as the conclusions. Please also identify the GPP reference product, explain its relationship to PLUMBER2 and FLUXNET2015, and specify the temporal resolution used for evaluation.
(10) Lines 226-245 and Table 1. The manuscript already explains the broad considerations behind site selection and appropriately identifies the study as a proof of concept. Please supplement this with more specific inclusion criteria and soften the statement that the subset "ensures" geographic and ecosystem diversity. It would also be useful to distinguish available observation records from the periods actually analysed. In particular, the description of approximately 1-3 years does not clearly correspond to all listed periods, including JP-SMF and IT-BCi. Please report the effective data coverage where it differs from the nominal record span.
(11) Sections 2.3 and 2.6.1. Please list the meteorological forcing variables and their units, distinguish the inputs supplied to ILS and the two networks, and specify the ILS version and configuration. At line 326, 48 half-hourly time steps correspond to one day rather than two days, unless resampling was applied. Please clarify the actual input frequency and sequence duration. It would also be helpful to explain sequence construction, recurrent-state handling across separate simulations, the final-model or checkpoint-selection rule, and whether the rolling training strategy involves anything beyond the reported learning-rate adjustment.
(12) Lines 188-195 and Eqs. (3)-(4). The motivation for representing seasonality is clear, but please report the number of frequencies, L, and explain the choice f_l = 2^l. Was this representation compared with a simpler annual sine-cosine encoding? If no comparison was conducted, a design rationale and appropriately cautious wording would be sufficient; its superiority or robustness should not be implied as an established result.
(13) Fig. 1(b) and Eqs. (7)-(8). The diagram appears to concatenate the shared LSTM representation with a pathway carrying parameter and DOY information before gating, whereas the equations generate gates from X_combined and apply them element-wise to H_shared. Please make these descriptions consistent with the implementation and provide the relevant tensor dimensions. Since gating operates on learned features, statements about selecting individual physical parameters should also be qualified.
(14) Fig. 3, Section 2.6.5, and line 419. The manuscript defines ordinary RMSE in Eq. (11) and identifies the Taylor-diagram contours as centred RMSE. For clarity, please state explicitly in the Fig. 3 caption that the relative RMSE improvement annotations use the ordinary RMSE defined in Eqs. (11) and (16). Reporting mean bias would provide a useful complementary diagnostic. For key architecture or performance comparisons, variability across repeated training runs would help assess stability, since cross-site variability does not establish robustness to network initialisation. Please replace "significantly improved" with descriptive wording unless statistical significance is demonstrated, and qualify robustness claims that were not directly tested.
(15) Please introduce Integrated Land Simulator (ILS) at its first occurrence in the abstract, correct the rendering of the parameter-generator notation around lines 105-108, and use "leave-one-site-out" where sites are the validation units. Although ILS_ORI is defined earlier, briefly restating the meaning of ORI alongside Eqs. (16)-(17) would make the equations easier to read. In Fig. 1, please replace "parameter priors" with "parameter estimates" if the outputs are point estimates rather than probabilistic priors.
Citation: https://doi.org/10.5194/egusphere-2026-4009-RC3
Data sets
PLUMBER2 benchmarking dataset Anna Ukkola https://researchdata.edu.au/plumber2-forcing-evaluation-surface-models/1656048.
Dataset and results for dPL and MdPL Experiments Wenpeng Xie https://zenodo.org/records/15753067
Model code and software
All code, data subsets, and environment specifications required to reproduce the ILS-MdPL experiments and figures Authors/Creators Xie, Wenpeng (Data manager)1 ORCID icon Wenpeng Xie https://zenodo.org/records/15748737
Viewed
| HTML | XML | Total | Supplement | BibTeX | EndNote | |
|---|---|---|---|---|---|---|
| 211 | 89 | 51 | 351 | 45 | 62 | 67 |
- HTML: 211
- PDF: 89
- XML: 51
- Total: 351
- Supplement: 45
- BibTeX: 62
- EndNote: 67
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
This study presents a multi-task differentiable parameter learning (MdPL) framework for calibrating parameters in the Integrated Land Simulator (ILS). The framework first trains a differentiable multi-task surrogate to emulate ILS sensible and latent heat fluxes under different parameter settings, and then uses a parameter-generation network to infer parameter sets constrained by flux observations. The ILS calibrated through MdPL is then evaluated against the original/default ILS and several reference approaches. The study includes multiple evaluation experiments, including surrogate-ILS consistency under joint parameter perturbations, benchmarking against other land surface models and an LSTM, comparison with a Sobol-based calibration baseline, and an out-of-target GPP evaluation.
Overall, I found the manuscript well written, clearly explained, and well organised. The methodological workflow is generally easy to follow. I have a few recommendations and requests for clarification that I believe would strengthen the manuscript; none of these should be considered major.
A few comments suggest additional parameter diagnostics. These are intended as recommendations, as they could provide useful insight into parameter behaviour and identifiability where the relevant outputs or analyses are readily available.