Attribution of aDGVM mismatches in biomass and tree cover across Africa
Abstract. Dynamic global vegetation models (DGVMs) are commonly evaluated against satellite-derived products such as biomass or tree-cover. Yet, most studies provide error metrics and maps that reveal where models fail rather than why. Here, we develop a residual-analysis framework to diagnose systematic mismatches between aDGVM results and remotely sensed aboveground biomass and tree cover and attribute them to climate, fire, and human influence. We computed grid-cell residuals between aDGVM simulations and remote-sensing estimates and modeled these residuals using generalized additive models (GAMs) with bioclimatic variables, fire activity, human footprint and population density, as well as a spatial smooth. Models were checked for concurvity and basis-dimension adequacy, and evaluated using spatial block cross-validation. For aboveground biomass, a GAM with twelve smooth terms explained 80.3 % of residual deviance (adjusted R2 = 0.798), while for tree cover, a GAM with similar structure explained 78.5 % of residual deviance (adjusted R2 = 0.78). Strikingly, bioclimatic variables dominated residuals in 82.9 % of AGB grid cells, not the processes known to be absent from aDGVM, such as nitrogen cycling and land use. This suggests that re-calibrating existing climate responses would reduce model bias more than adding missing process modules. Fire and human pressure played a proportionally larger role for tree cover (42.3 % of grid cells) than for biomass (17.3 %), reflecting distinct ecological controls on net primary production versus canopy dynamics. A persistent spatial signal after accounting for all predictors (unique R2 = 11.0 % AGB, 9.6 % TC) points to unmeasured processes including soil properties, wild megafauna, and land-use history. The framework transforms descriptive DGVM error maps into quantitative, spatially explicit process attribution directly translating model-data mismatches into prioritized development targets for next-generation vegetation models.
This manuscript presents a novel and carefully designed framework for diagnosing DGVM deficiencies by modeling data-model residuals with generalized additive models (GAMs). Rather than simply quantifying model performance or mapping model-data differences, the study attempts to identify the environmental drivers underlying these residuals and to translate them into priorities for future model development. I consider this shift from descriptive model evaluation to process-oriented diagnosis to be an important methodological contribution. The statistical framework is generally rigorous, particularly the use of spatial block cross-validation and variance partitioning, and the manuscript is overall well written and logically organized. I therefore enjoyed reading this study. My comments below are intended primarily to improve the clarity of the methodology and to moderate several interpretations that, in my opinion, go somewhat beyond what the analyses directly support.
___ Major Issue ___
(1) I appreciate that the authors evaluated the models using spatial block cross-validation rather than conventional random cross-validation. This is an important strength of the manuscript, as it provides a more realistic assessment of model generalization in the presence of spatial autocorrelation. However, the implementation of the spatial block cross-validation is currently insufficiently described, making it difficult to evaluate or reproduce the analysis.
In particular:
[a] The manuscript states that blocks of 100 km2 were used. This appears inconsistent with the 0.5° resolution of the analysis, and I wonder whether the intended block size was 100 km rather than 100 km2. Please clarify.
[b] More generally, additional details of the blockCV implementation would be helpful. For example, how were the blocks constructed (regular grid or otherwise)? How was the block size selected? Was it chosen based on the spatial autocorrelation range of the residuals or by another criterion? How were blocks assigned to the five folds?
Because spatial block cross-validation is a key methodological component of this study, a clearer description would substantially improve the manuscript's transparency and reproducibility.
(2) Sign consistency between residual definition and partial-effect interpretation (Section 3.1, Figs. 4-5)
Residuals are defined as r = aDGVM - RS (Eq. 1), so a positive residual means overestimation and a negative residual means underestimation. Checking Fig. 4 against this definition, several interpretations in Section 3.1 appear to have the sign reversed, and this is not limited to a single instance:
bio12: partial effect is positive at high values (implying overestimation), but the text states this indicates "underestimation of biomass in wetter regions." Notably, bio18 is discussed in the same sentence with the opposite sign (negative at high values), yet both are used to support the same conclusion.
Burned area: partial effect decreases (becomes negative) with increasing burned area, implying underestimation, but the text states this indicates "overestimation of woody biomass and cover."
HFI: partial effect is negative at high HFI, implying underestimation, but the text concludes the model "overestimates vegetation in heavily human-modified landscapes."
Population density and N deposition: partial effects are positive at high values (implying overestimation), but the text states aDGVM "underestimates biomass and tree cover" in these regions.
I may be missing an intended sign convention. Still, since this pattern recurs across nearly all disturbance and anthropogenic predictors in this paragraph, I would ask the authors to systematically re-check every sign attribution in Section 3.1 against Eq. (1), rather than correcting individual sentences.
This matters beyond wording: these sign interpretations directly support the development priorities in Section 4.3 (e.g., claims about the fire module or under-representation of grazing/land degradation). If signs need to be reversed, the corresponding conclusions in the Discussion (and possibly the Abstract) may need revision as well. A supplementary table listing, for each predictor, the sign of the effect at high/low values, alongside the stated over- or underestimation conclusion, would help verify consistency.
___ Minor Technical Issues___
(1) The manuscript contains an unresolved cross-reference ("Table ??") in Section 2.4. Please correct the table reference.
(2) Please clarify the interpretation of the effective degrees of freedom (edf) of the spatial smooth in Section 3.1. A large edf reflects the complexity of the fitted spatial surface, but does not by itself indicate that a large fraction of the residual variance is explained by spatial structure. This conclusion seems to be more directly supported by the variance partitioning analysis presented later.
(3) The first paragraph of Section 3.1 could be clarified. The statement that "spatial patterns of residuals predicted with GAMs agreed well with residuals derived from aDGVM and remote sensing" initially gave the impression that the manuscript was referring to the residuals remaining after fitting the GAM. However, this sentence actually compares the original data-model residuals with the residuals predicted by the GAM. Since the following sentences then switch to discussing the residuals of the fitted GAM, I suggest making this distinction more explicit to improve readability.
(4) In Section 3.2, stating that variance partitioning "confirms" the contribution of unmeasured spatially structured processes seems too strong. This analysis quantifies the spatial smooth's unique contribution (11.0% AGB, 9.6% TC) but cannot identify its cause. Notably, this is in tension with Section 4.4, where the authors acknowledge that spatially structured uncertainty in the remote-sensing products "may generate patterns in the GAM residual surface that reflect measurement error rather than aDGVM process deficiencies." The same spatial signal therefore appears to be interpreted somewhat differently in the two sections. I suggest softening "confirming" to "suggesting" in Section 3.2, cross-referencing Section 4.4, and applying similarly cautious wording to "The third hypothesis is also confirmed" in Section 4.1, which relies on the same evidence.