Potential natural vegetation as overlooked determinant of land-use change flux estimates
Abstract. The net carbon dioxide (CO2) flux from land use and land-use change (FLUC) is a major driver of anthropogenic climate change and central to mitigation strategies for achieving global emission reduction targets. Despite its importance, estimates of FLUC are characterized by large uncertainties. In models quantifying FLUC, the spatial distribution of potential natural vegetation (PNV) is a key component, but its influence on FLUC estimates has not been systematically quantified. Here, we address this gap by combining pollen-based biome reconstructions and observation-based datasets of environmental conditions with machine learning to derive global PNV maps. Compared to existing PNV maps, our approach improves the representation of biome-specific spatial heterogeneity and provides sensitivity maps for quantifying how assumptions about potential forest and grassland distribution propagate into FLUC estimates. Implementing the PNV maps as Plant Functional Types (PFTs) in the bookkeeping model BLUE, we find global cumulative FLUC for the period 1850–2023 to be 16 % (6–27 %) higher than the default estimate. Differences at the regional scale are often even larger. Our results demonstrate that uncertainties in PNV distribution represent a substantial and previously overlooked source of uncertainty in FLUC estimates, comparable in magnitude to other key sources. Accurate PNV mapping is therefore essential for robust FLUC estimates, particularly at regional scales, which are required for understanding the global carbon cycle, improving FLUC modeling, and informing effective climate mitigation policies.
The manuscript "Potential natural vegetation as overlooked determinant of land-use change flux estimates" is a very interesting and cool study. The authors introduce pollen-based biome reconstructions into a machine-learning framework to derive global potential natural vegetation maps and then translate these maps into PFTs for implementation in the BLUE bookkeeping model. This is a critical contribution because the spatial representation of potential natural vegetation is often treated as a fixed background assumption in land-use carbon flux modelling, even though this manuscript shows that it can substantially affect FLUC estimates. Overall, I think the study conceptually strong and potentially important for the carbon-cycle and land-use modelling communities.
The abstract doesn’t clearly summarize what the main improvement is compared with existing PNV or PFT maps. I suggest adding one or two concrete results, for example the difference in forest, shrub, grass, or abiotic PFT areas. The abstract already reports that cumulative FLUC during 1850–2023 is 16% higher than the default estimate, with a sensitivity range of 6–27%. However, this result would be clearer if the absolute values were also included, for example 292 Pg C versus 252 Pg C. This would help readers immediately understand the magnitude of the effect.
In the PNV/PFT results section, the manuscript would be strengthened by adding a figure that quantitatively summarizes the area of each biome or PFT class under the new best-guess map and the default BLUE map. Currently, the spatial maps are informative, but readers would benefit from a bar plot or stacked-area comparison showing the global area differences among the 11 PFTs and/or the 16 biome classes.
The authors use nested cross-validation with stratified sampling, which is useful for class imbalance, but spatial autocorrelation among pollen sites may inflate model performance. The authors acknowledge this limitation, but the issue is important enough that I suggest adding a spatial block cross-validation or leave-one-region-out validation as a sensitivity test.
Line comments:
Line 5: "Observation-based datasets of environmental conditions" is slightly vague. Consider specifying such as "climate, soil and geographic predictors."
Lines 70–75: The statement that model-based approaches simplify the relationship between vegetation and environment is reasonable, but machine learning used in this study is also model-based. Please do a bit more clarification.
Lines 1430-149: The aggregation from 32 biome classes to 16 classes is reasonable, but the ecological implications of this aggregation should be discussed more explicitly. Some ecotonal or mixed vegetation types may be lost or simplified.
Lines 150-157: The increase in training data relative to previous studies is a major strength. This result should be highlighted earlier, possibly in the abstract or introduction.
Lines 155-157: The sparse sampling in Africa and Australia is important. Please consider adding a short paragraph explaining how this affects confidence in predicted savanna, dryland, and shrubland systems.
Lines 215-234: The reclassification from biome classes to BLUE PFTs is central to the manuscript. I suggest adding more details here, especially for mixed classes.
Figure 5: This figure contains many important results but is visually dense. Consider splitting the annual fluxes and cumulative flux components into separate figures or improving readability.