the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
A Weakly Supervised Deep Learning Framework for Estimating Above-Ground Biomass for Non-Forest Landscapes From Optical Images
Abstract. Above-ground biomass (AGB) maps are essential for carbon accounting and sustainable land management, yet AGB for non-forest landscapes remains poorly accounted for in global datasets. Here, we make use of deep learning and high-resolution PlanetScope imagery to introduce the concept of AGB contribution maps, which are high-resolution AGB predictions that capture local patterns. These maps can be predicted at any resolution from 1 to 100 m, providing insights into the spatial features included in the coarser resolution AGB maps, being essential for mapping trees outside forests. Our method employs a weakly supervised hybrid framework that transfers information from an existing 100 m global AGB map to high‑resolution optical satellite imagery, enabling the interpretation of detailed spatial patterns. We demonstrate that our map achieves detailed and spatially consistent patterns of woody vegetation in African savanna landscapes comparable to UAV-based LiDAR. Aggregated AGB values are well aligned with independent in-situ measurements (r2 = 0.71, bias 1 %), which is contrary to the original coarse AGB map used for training (r2 = 0.17, bias 48 %), indicating the capability of our approach to refine the existing map towards a higher accuracy for estimating tree biomass outside forests. This suggests that our model has learned tree-level information that is not present in the original AGB training data, providing a framework to refine existing coarse-resolution AGB maps. The granular and multi-resolution results provide no contribution to global efforts in sustainable land management of non-forest landscapes at any preferred scale and resolution.
- Preprint
(6153 KB) - Metadata XML
-
Supplement
(1185 KB) - BibTeX
- EndNote
Status: open (until 01 Oct 2026)
- RC1: 'Comment on egusphere-2026-2924', Anonymous Referee #1, 11 Aug 2026 reply
-
RC2: 'Comment on egusphere-2026-2924', Anonymous Referee #2, 13 Sep 2026
reply
This manuscript addresses an important and timely problem in ecological remote sensing, namely, how to obtain spatially detailed above-ground biomass (AGB) information for non-forest landscapes where high-resolution field observations or LiDAR data are generally limited. The proposed framework combines 3 m PlanetScope imagery with an existing 100m AGB product under a weakly supervised learning framework to obtain an AGB contribution map at fine-scale spatial resolution. The use of independent UAV LiDAR-based datasets provides encouraging evidence for the potential of the approach. Overall, this study has potentially important value for biomass mapping in heterogeneous non-forest ecosystems. However, several aspects of the manuscript would benefit from further clarification and revision. The specific suggestions are as follows:
1. The current Introduction (Lines 15–84) is relatively long and contains very few paragraph breaks, which makes it difficult to follow and to identify a clear logical structure. Please restructure the Introduction into several clearly separated paragraphs to improve the logical flow.
In addition, the current content in Section 2.1 (Lines 106–115), particularly the description of the ecological characteristics and mapping challenges of non-forest landscapes, could be considered to move to the Introduction.
2. The concept of the “AGB contribution map” needs a much clearer definition. The AGB contribution map appears to be a central concept of the manuscript, but its meaning remains somewhat ambiguous. I tried and maybe I missed, but I failed to find a clear definition.
For example, the manuscript describes it as an “intermediate high-resolution prediction output (Line 73)” that captures fine-scale spatial variability, while elsewhere it is described as a “product that can be aggregated to generate AGB maps at different spatial resolutions (Line80 and the caption of Figure 1)”.
Please provide a precise mathematical and physical definition of the AGB contribution map. In particular, the authors should clarify:
- whether the output has physical units such as Mg/ha;
- whether the values represent biomass density, relative contribution (%), or normalized spatial weights;
- where the contribution map is produced in the network architecture;
- how it differs from the final high-resolution AGB map;
- and how the contribution map is transformed into the 30, 50, or 100 m AGB products.
Lines 147-148 seems central as it states that the AGB contribution map represents its relative contribution to the total AGB within a 100m X 100m tile. This can be ambiguous because total AGB is the product of AGB density and the area (1ha in this case) but the AGB map at a high resolution (1m or 3m) derived finally in the study is the AGB per area (i.e., density). The AGB contribution map can be in unit of percentage (%)? These confusions need to be clarified. Other places that are relevant include line 166, where ‘the mean AGB’ should have a unit of Mg/ha if the ‘mean’ means being averaged over the area of the tile. This point is also relevant for Fig. 4: the panel c cannot be understood in a naïve manner. If as explained in the caption, does it mean that the values over 100 (10x10) 10-m pixels should sum to 100%?
Hence the authors should in particular check AGB and be clear when they mean a density (Mg/ha) or the total value over a given area (Mg).
3. The manuscript should clarify what the model actually learns from PlanetScope imagery. An important question is: Although PlanetScope images can help the model determine where the AGB should be located, does it really know what the AGB value should be at each location? The model may learn the “spatial distribution” associated with vegetation structure and biomass distribution, but this does not necessarily mean that it can accurately estimate “the true biomass value” at each fine-scale location.” These are two entirely different things.
Since biomass estimation depends not only on spatial patterns but also on structural,biophysical properties of vegetation and so on, it is important to distinguish between the ability of PlanetScope imagery to identify spatial patterns and the ability of the model to quantitatively estimate true fine-scale biomass. Please clarify this distinction and moderate the relevant statements where necessary.
Line 201-202 seems to say that the model can only predict the relative contribution of AGB at fine spatial scale, but it needs the AGB scalar values (the original 100m AGB) to derive the AGB density at fine scale. If this is the case, the authors may need to think revise their title as it is in essence a ‘downscaling’ approach, different from what is implied in the title that AGB can be directly predicted from PlanetScope imagery.
4. The manuscript uses the term “weak supervision”, but the concept and its specific implementation in this study are not introduced in sufficient detail for readers who are less familiar with weakly supervised learning. Please add a short explanation of what is meant by weak supervision in this study.
5. In Figure 3, the original Xu et al. (2021) AGB product shows a relatively limited value range of approximately 0~60 Mg/ha (Fig. 3e), while the proposed high-resolution prediction (Fig. 3b) and its aggregated 100 m product by spatial averaging (Fig. 3c) appear to cover a broader range of approximately 0-100 Mg/ha. At the same time, the manuscript states that the model is trained using the Xu et al. (2021) 100 m AGB product and that the spatial mean of the high-resolution prediction is constrained to match the coarse-scale target.
This raises an important methodological question: to what extent does the 100 m aggregation of the model output actually constrained to reproduce the original Xu et al. (2021) 100 m values?
Please report, preferably in a quantitative way:
- the agreement between the original 100 m supervisory labels and the model predictions aggregated to 100 m;
- the distribution of their differences;
- the training/validation loss associated with the coarse-scale constraint;
- whether the original 100 m labels and the final aggregated predictions are compared on exactly the same spatial grid and using the same spatial extent and masking procedure.
These details are particularly important because the manuscript reports that the original 100 m product has an R2 =0.17 against the UAV data, whereas the proposed 100 m product reaches an R2 = 0.71 (Lines 259-261). If the two products are expected to be similar after strict mean-pooling supervision, the substantial difference in their agreement with the UAV reference requires further clarification.
6. The independent validation results from the UAV LiDAR are encouraging. Could you provide the spatial distribution of these validation plots? Are these validation plots spatially independent?
Additionally, it would also be helpful to explain more clearly how the 20 × 20 m UAV-derived plots are spatially matched to the 10m/100 m AGB products. In particular, please clarify how multiple 20 × 20 m plots within a 100 × 100 m pixel are handled and whether the comparison uses averaging, weighting, or another aggregation methods.
7. The figures are currently relatively small and some labels and legends are difficult to read. Please enlarge the figures and associated text in the final version where possible.
8. The references to the Xu et al. (2021) / Saatchi AGB product are not consistent across the manuscript. For example, Figure 3(e), Table 1, and Table 2 use slightly different descriptions, such as “Saatchi et al. Mg/ha” and “Saatchi et al. (Xu et al., 2021)”. Please standardize the naming of this product throughout the manuscript.
9. Please consider adding the original 100 m AGB product (Xu et al., 2021) to Figure 4. Including the original 100 m AGB distribution would facilitate a direct comparison with the proposed high-resolution prediction and the aggregated 100 m product, and would make the spatial refinement introduced by the proposed framework more evident.
10. The ordering of panels in Figure 5 could be improved. You may consider reorganizing the panels as a–c–d–b–e to improve the logical flow of the visual comparison.
11. The manuscript alternates among “AGB contribution map,” “high-resolution AGB map,” and “AGB density map.” These terms can bring in confusions. Please clearly define the physical meaning and units of each product and use the terminology consistently throughout the manuscript.
12. The manuscript refers to training, validation, hold-out test, and independent datasets. Please report the number of samples in each subset and specify whether the splitting was random, spatially blocked, or geographically separated.
13. The native spatial resolution of PlanetScope imagery is approximately 3 m, while the input imagery is resampled to 1 m using bilinear interpolation. Nevertheless, the statement in the Abstract that “These maps can be predicted at any resolution from 1 to 100 m (Line 4)" may give the impression that the method directly provides genuine 1 m AGB observations. In line 164, what do you mean by ‘an unsampled resolution of 1m’?
Citation: https://doi.org/10.5194/egusphere-2026-2924-RC2
Viewed
| HTML | XML | Total | Supplement | BibTeX | EndNote | |
|---|---|---|---|---|---|---|
| 110 | 54 | 17 | 181 | 30 | 19 | 22 |
- HTML: 110
- PDF: 54
- XML: 17
- Total: 181
- Supplement: 30
- BibTeX: 19
- EndNote: 22
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
General comments:
This paper presents a weakly supervised deep learning model for Above Ground Biomass (AGB) estimation for non-forest areas. The model uses Planet images at 3m resolution as the input and AGB reference maps at 100m to predict 10, 30, 50, or 100m resolution AGB products. The authors highlight that most researchers focus on forest areas, ignoring small non-forest open-canopy areas. Studying this areas is important in the current context of climate change. The authors insist that their main contribution do not consist in the production of high resolution maps, but rather in the focus on mapping those non-forest areas. Moreover, the authors state that the used spatial pooling method improves the 100m resolution AGB predictions thanks to the extraction of the fine-grained spatial features from Planet imagery.
Overall, the paper addresses an interesting and relevant research problem, particularly because AGB mapping is known to be challenging due to the lack of reliable ground-truth data. Hence, it is understandable that training and evaluating an AGB prediction model can be challenging. The authors provide a quantitative evaluation of their method against one independent UAV-based reference dataset, with generally encouraging results. However, this evaluation alone is not sufficient to assess the robustness of the method. The article lacks some standard evaluation protocols and metrics commonly used for model training and evaluation. In addition, some of the conclusions and discussions are not fully supported by the presented results, and some aspects of the results are not sufficiently discussed.
Specific comments:
General:
1.
Line 310. You say: “However, the key output is not the high-resolution of the AGB contribution map. [...] The major achievement is the fact that by aggregating the high-resolution contribution map to the scale of a desired operational AGB map (typically > 30 × 30 m resolution), we show a highly improved relationship with UAV data as compared to previous maps, including the map that has been used for training/supervision (Xu et al., 2021).”
Moreover, the abstract states : “The granular and multi-resolution results provide no contribution to global efforts in sustainable land management of non-forest landscapes at any preferred scale and resolution.”
At the same time, on the line 83 of the introduction you state: “Therefore, it offers a scalable solution for mapping high-resolution AGB in landscapes where such information has historically been unavailable or incomplete.”
which is contradictory to the information in the abstract and discussion.
The paper is presented as using a weakly supervised learning approach. If the main contribution is producing AGB maps at 100 m resolution, it is unclear what the role of the weakly supervised approach is in this contribution, since the model ultimately produces the same type of output as the target data. The authors should clarify the main contribution of the paper and explain more clearly what is achieved by the proposed weakly supervised approach.
Dataset:
2.
Did you divide your dataset into train/test/validation? How many patches were used? The more detailed information should be given in the dataset section.
Overall, there is no clear information about all the products used for evaluation.
Suggestion: It could be easier for readers, if you make an overview table that describes all the data that was used with months/year, resolution and surface.
Moreover, sometimes it is not clear if during the comparison you used the products from the same year.
Finally, you do not indicate which month the Planet image used for model training was takes. How do you deal with seasonality while making predictions?
Model
3.
The reproducibility of your method is not possible with the description given in the article. Hence, the model should be better explained: 1/Figure 1 needs to be detailed (the encoder-decoder part, bilinear upscaling, ….); 2/The experimental setup should be presented: the number of layers and other hyperparameters, as well as training time.
4.
Convolution margin effect. The paper does not appear to account for the convolution margin effect when applying the model to generate the AGB maps. Since predictions near patch boundaries can be affected by the limited spatial context and padding, this could introduce spatial artefacts in the resulting maps. Please clarify how this effect is handled and, if it is not explicitly addressed, discuss its potential impact on the results.
Validation of the method
5.
Quantitative evaluation. You present a quantitative evaluation on an independent dataset; however, you do not report any evaluation of the model on the test set of the original dataset. This should be reported as part of the standard evaluation protocol. Without this information, it is difficult to assess the model performance and interpret the subsequent evaluations.
For example, at line 260 you state: “The original map used for training has a weak relationship (R² = 0.17) with the UAV data and a high underestimation bias of 48%. When aggregated to 100 m, our map has a considerably higher (R² = 0.71) and a lower bias −1% at the 100 m plot level.” This raises the question of how well the model performs with respect to the original reference data. Please report the results on the test set of the original dataset, including the relevant evaluation metrics, so that the model's performance can be properly assessed.
Moreover, RMSE and MAPE should be reported in all experiments to assess the magnitude of the errors, as R² alone does not provide information about the actual prediction error, especially on the extreme values.
Line 254. You state: “The strong relationship is evident over the entire value range, almost following the 1:1 regression line, see Fig. 3.” However, most of the values appear to be concentrated in the lower range, and the points show substantial scatter around the 1:1 line, particularly relative to the magnitude of the values. Therefore, the strength of the relationship over the entire value range is difficult to assess from the figure. Reporting RMSE and MAPE would help quantify the magnitude of these errors and provide a more complete assessment of the model performance. Moreover, a density scatterplot (e.g., using point density or hexagonal binning) would be more appropriate for visualizing the distribution of points in this case.
It is understandable that lower AGB values can be difficult to predict, and your algorithm shows the overall improvement over the existing maps, but the accuracy of the prediction of your method should be assessed and discussed for different value ranges.
Line 307: “Considering that the error is below 10 Mg/ha, multitemporal maps may provide indications on the carbon accumulation of plantations already after a few years of growth.” – giving the previous comment, I am not sure this is fully correct.
6.
Line 255 “Interestingly, when aggregating our map to 100 m and comparing both the original 100 m AGB map used for training and our map aggregated to 100 m with the UAV derived AGB, our map shows an improved relationship with higher correlation values.”
How did you compare the original AGB map with the UAV-derived AGB if they do not have the same resolution? You state that the UAV-derived biomass is computed at the 20 m plot level by delineating individual tree crowns. What about small vegetation that does not have a clearly identifiable crown? Is it considered as zero biomass?
How did you perform the resampling of the UAV-based map from 20m to 100m?
7.
High resolution maps. It is known that the presented method (Weak supervision via spatial pooling) tend to produce homogeneous (or only weakly varying) high-resolution predictions within each low-resolution pixel. From the presented results, it is not clear whether your method overcomes this limitation.
Figure 4 presents the standard deviation of the pixels within a hectare. However, there are several issues with this figure. First, the figure does not include a colorbar, making it impossible to assess the magnitude of the reported spatial variability. If the colorbar from subfigure (d) is intended to apply to the subfigure (e), this should be made explicit and the color bar should be centered appropriately.
Second, it is difficult to visually compare your high resolution results to the satellite imagery without a proper zoom level on the high resolution results: I suggest selecting a couple of 2x2 or 3x3 pixels low resolution patches for of heterogenous areas and present the corresponding high resolution results with sufficient zoom level. For biomass maps, the 0 value should be set to white to better see textures.
8.
The paper does not provide any quantitative evaluation of the spatial consistency of the super-resolved biomass maps, although, you claim that “The high spatial resolution of our map gives confidence and clarity in what features are included in the AGB calculations.”
In Section 3.2, the evaluation is performed only on 20 × 20 pixel plots containing vegetation. However, this does not show whether the model correctly predicts zero biomass in non-vegetated areas.
This could be evaluated quite easily. For example, the predicted biomass maps could be converted into vegetation/non-vegetation masks and compared with masks derived from the LiDAR reference data. The spatial consistency could then be reported using metrics such as the F1 score or mIoU.
9.
National-scale mapping. I do not understand which year the national-scale predictions correspond to. Were the predictions generated for the same year and spatial extent as the data used for training? If so, the higher correlation with the Xu et al. map compared with the other maps is not surprising, since the Xu et al. map was used as the training target. This comparison therefore does not provide an independent assessment of the model.
Moreover, the reported MAPE of 366% at 100 m resolution appears extremely high and raises questions about the quality of your model predictions. Please, discuss this result.
Finally, the maps used for comparison appear to correspond to different years. This makes a direct comparison difficult, as differences between the maps may partly reflect temporal variability rather than differences in mapping performance. Ideally, the maps should be compared for the same year, or the potential impact of the temporal mismatch should be explicitly discussed.
Figure 6. Why your model prediction values are all concentrated around 18000 Mg per 1km2? How it is then possible that your model has higher total AGB in table 2 if other models predict higher values? Please, such homogenious prediction values of your model.
10.
Overall, the authors should provide at least one additional quantitative comparison with independent reference data. Otherwise, the generalizability of the method cannot be properly assessed.
Figures
11.
Figure S2. Optical images (especially 1st row) can not be analysed properly due to the chosen color normalization. Please, normalize properly or use false color scheme for the images, so the vegetation is better seen.
On the second row, it can be seen that two images do not correspond to the same year (the absence of the vegetation on the right). On the first row, the predicted vegetation pattern clearly do not match with Google Earth image: the forest gap in the middle of the image is not represented in the prediction. Please, explain. Does the Google Earth image match the Planet image?
Technical corrections:
Figure 5 - “(a) Land-cover map showing the non-forest classes preserved in the analysis.” This is not correct, you show land cover map for the first row and satellite images for other rows. Please, change.
The text in sections 1 and 2 is not divided into paragraphs. Moreover, most Figures do not have a colormap legend and/or a scale; the chosen colormaps make it complicated to visually analyse the results. Some results need to be presented in tables for clarity.
For colormaps – use white for “no biomass” zero values and colors for biomass. It is difficult to see vegetation patterns in high resolution predictions with magenta colorscheme.
Please, add scale to all maps.
The results from Section 3.2 should be put in a table for the clarity.
You use different author’s names for the method used for the reference AGB map - Saatchi and XU – please, use the same name, it is can be confusing for readers.