Quantifying multiple sources of uncertainty in national-scale soil organic matter prediction
Abstract. Large-scale assessment of soil organic matter (SOM) relies on spatial prediction to extend sparse observations to unsampled locations, but extensive soil databases do not necessarily ensure reliable predictions. Predictive uncertainty can arise from limitations in sampling, response-variable perturbation, and model specification, yet their relative contributions are rarely quantified. Here, we developed nationwide models of SOM in paddy soils across South Korea and used an ensemble random forest framework to decompose predictive uncertainty into sampling, response-variable perturbation, and model-specification components. We compiled 283,618 observations with soil, terrain, climate, and vegetation predictors. The model explained 50 % of the variation in independent SOM observations but underpredicted localized areas with high SOM. Model specification was the largest contributor to total predictive variance (41.9 %) and was the dominant uncertainty source at 55.8 % of locations, followed by sampling (29.3 %) and response-variable perturbation uncertainty (28.8 %). Exchangeable calcium and annual precipitation were the most influential predictors, although their importance rankings varied across model configurations. These results show that predictive performance alone does not reveal the sources of uncertainty in large-scale SOM assessment. Large datasets do not necessarily provide equal predictive information across the response distribution. Explicit decomposition of predictive uncertainty can identify whether improvements should prioritize additional sampling, improved response-variable measurements, or alternative model specifications, providing a basis for more targeted strategies to improve the reliability of large-scale SOM predictions.