Pushing the Boundaries of NGGM and MAGIC Derived Terrestrial Water Storage Anomalies with Unsupervised Deep-Learning Downscaling
Abstract. This work presents an unsupervised deep-learning approach for the spatial downscaling of Terrestrial Water Storage Anomalies (TWSA) from their native coarse resolution to a 1° grid, using simulated Next-Generation Gravity Mission (NGGM) and ESA–NASA Mass-change And Geosciences International Constellation (MAGIC) products obtained from realistic end-to-end (E2E) closed-loop experiments carried out in the context of ESA studies, together with real Gravity Recovery and Climate Experiment (GRACE) and Follow-On (GRACE-FO) observations. A U-Net Convolutional Neural Network integrates hydro-climatic variables from ERA5-Land and topographic information to retrieve fine-scale TWSA patterns while keeping consistency with the original coarse gravimetric observations. Results over five regions of interest – Africa, the Amazon basin, Europe, the Ganges basin, and the Mississippi basin – have been evaluated through a composite score combining pixel-wise and structure-wise metrics. The analysis shows that the quality and native resolution of the input gravimetric products strongly influence the downscaled fields. In the simulated experiments, NGGM and MAGIC generally achieve higher high-resolution composite scores than the GRACE-C-like configuration, while monthly products outperform the corresponding 5-daily solutions. At the native satellite resolution, the aggregated downscaled fields remain highly consistent with the input products, with composite scores generally above 0.90 in the simulated and real scenarios. Water balance equation (WBE) analyses over the Congo and Danube basins show that the downscaled products generally preserve the basin-scale closure behavior of their corresponding satellite inputs, particularly for the 5-daily simulations, where the errors of the aggregated and 1° downscaled products remain close to the satellite baseline. The improved quality of the simulated observing scenarios is also reflected in lower WBE errors: from the GRACE-C-like to the MAGIC configuration, the 1° downscaled RMSE decreases from 4.82 to 3.32 mm over Congo and from 14.47 to 7.10 mm over Danube. For real GRACE(-FO) observations, the aggregated downscaled products reproduce the satellite-scale WBE residuals with NRMSE values of 0.014 over Congo and 0.007 over Danube. Although the fine-scale redistribution remains underconstrained in the absence of independent high-resolution observations, these findings demonstrate that unsupervised deep-learning downscaling can improve the spatial detail of satellite-derived TWSA while preserving the large-scale gravimetric and hydrological consistency of both simulated future-mission products and real GRACE(-FO) observations.
In this study, Goracci et al. (2026) present a deep learning framework for downscaling gravimetry-derived terrestrial water storage anomalies (TWSA) from their native resolution to a 1° grid. A U-Net is driven by ERA5-Land hydro-climatic variables (with temporal lags), topography, a surface-water fraction, and a seasonal encoding, and is trained with a loss that compares the area-weighted aggregate of the 1° prediction with the coarse satellite field. The framework is applied to (i) realistic end-to-end (E2E) closed-loop simulations of GRACE-C-like, NGGM, and MAGIC products, in 5-daily (3°) and monthly (2°) configurations, which allow evaluation against the ESA ESM 3.0 HISBED reference at 1°; and (ii) real COST-G GRACE(-FO) monthly TWSA at 3°. Performance is assessed over five regions with a weighted composite score of pixel-wise and structure-wise metrics, a water balance equation (WBE) analysis over the Congo and Danube basins, and a comparison between the simulated GRACE-C-like products and GRACE(-FO) over 2007–2020.
The manuscript provides valuable insights into downscaling MAGIC-like TWSA products. Although the spatio-temporal resolution is expected to be substantially improved by combining the polar constellation with the Bender constellation, downscaling approaches remain important for regional analyses. However, the manuscript in its current form requires major revisions to address several concerns.
An overall major issue: the manuscript attempts to report too many aspects, and as a result, its main focus is lost. The most valuable contribution of the work is the assessment of how the quality of future gravimetry mission products propagates into the downscaled fields, so the authors may consider highlighting this part and downweighing the others.
Major comments
1. The overall writing flow needs to be improved. Some general suggestions that the authors may consider are as follows:
[1] All abbreviations need to be defined at their first occurrence.
[2] I highly recommend moving the methodology section after the data section. The reading flow will be much smoother if readers are already familiar with the data you use.
[3] Please provide the necessary references for the techniques you use, such as CNNs, U-Net, the Adam optimizer, deep ensembles, etc..
2. Overall, the quality of the figures and tables needs to be improved. For example, the caption of Fig. 1 does not provide much information. Moreover, Fig. 1 includes many subplot titles that are redundant and not easily readable. Similar issues occur repeatedly, including but not limited to Fig. 2 (inappropriate font sizes), Fig. 3 (caption and unreadable high-resolution inputs) and Fig. 9 (spatial details are not readable; consider providing enlarged regional maps). In addition, several tables do not convey information effectively and need to be redesigned. Table 1 does not provide much information and could be combined with Fig. 1, Table 3 does not have an appropriate caption, and Tables 5 and 6 include too much information and should be streamlined (e.g., by highlighting key values).
3. I have several concerns regarding the methodology.
First, "self-supervised learning" is the correct term rather than "unsupervised learning", since you do compute the MSE between predictions and labels, which then serves as the supervision signal. "Unsupervised learning" refers to approaches without any supervision signal.
Second, there is an inconsistency between the "real coarse-resolution satellite data" and your "coarse-resolution aggregated results": all the satellite data (simulations) you used are based on spherical harmonics, which yield spatially continuous fields. However, you computed TWSA_LR using an area-weighted average, which will very likely introduce discontinuities in the spatial field. Using a filter rather than simple averaging would be better.
Third, I do not fully understand how the model can generate high-resolution information. The only supervision signal comes from the "averaged MSE", so it can only force the model's predicted values to agree with the satellite data at their native resolution. How do you ensure that the model generates plausible high-resolution information?
Finally, you include a large number of inputs (Table 3). Have you performed any ablation tests or feature importance analyses to understand which features are important and which may be unnecessary?
4. The study applies the method to only five regions, although it should be easily extendable to the whole globe. I highly recommend that the authors extend their analysis to the global TWSA field so that everyone in the world can benefit from it. Moreover, the water balance equation closure analysis should also be extended to include more basins.
5. The study includes too many evaluation metrics (Table 4), which are unnecessary and actually introduce more confusion than insight. Since the authors have a simulated ground truth, a more meaningful analysis could be achieved through a more careful interpretation of a few "basic" metrics, such as RMSE or NSE, rather than reporting more than ten metrics.