the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
Deep Learning-Enhanced Background Error Covariance Estimation for Massive-Ensemble Kalman Filter in Sea Surface Temperature Forecasting
Abstract. In recent years, deep learning-based ocean forecasting has become a prominent research focus. However, recent studies often rely on operational ocean forecast systems to provide initial conditions. Operational ocean forecast systems typically use Ensemble Data Assimilation (EDA) to generate these initial conditions. Nonetheless, the high computational cost of numerical models limits the ensemble sizes in EDA, resulting in rank deficiencies in the background error covariance matrix and introducing spurious correlations. To address these challenges and advance the operationalization of deep learning-based ocean forecasting, we propose Deepcov-EnKF, a deep learning-enhanced background error covariance method for massive Ensemble Kalman Filter (EnKF) applications in Sea Surface Temperature (SST) forecasting. The proposed method incorporates a deep learning-based SST forecasting model to generate approximately 5,000 ensemble members. It then directly maps this high-dimensional perturbation set into a background error covariance matrix using a deep neural network. This approach reduces the computational cost of covariance estimation by more than 200-fold compared to conventional techniques. Experimental results demonstrate that the forecasting model can initialize reliable 60-day SST predictions with a Root Mean Square Error (RMSE) of approximately 0.6 °C. Furthermore, assimilation diagnostics reveal that Deepcov-EnKF effectively resolves the spurious correlations in the covariance matrix. The method also exhibits robust stability and outperforms advanced numerical assimilation methods during a 360-day cycling forecasting experiment. This study confirms that Deepcov-EnKF overcomes key limitations of traditional EDA frameworks, significantly enhances the accuracy of SST assimilation, and lays the foundation for high-precision marine forecasting systems.
- Preprint
(4387 KB) - Metadata XML
- BibTeX
- EndNote
Status: open (until 02 Nov 2026)
-
CEC1: 'Comment on egusphere-2026-3488 - No compliance with the policy of the journal', Juan Antonio Añel, 23 Sep 2026
reply
-
AC1: 'Reply on CEC1', Baoxu Li, 25 Sep 2026
reply
Dear Executive Editor Añel,
Thank you very much for your careful assessment of our manuscript and for explaining GMD’s Code and Data Policy. We sincerely apologize that our original availability statement cited the sources of the raw data without clearly identifying a persistent archive for the precise processed data used in our study. We are grateful for the opportunity to correct this omission.
We have now archived the materials relevant to the study in Zenodo. The model code is available at https://doi.org/10.5281/zenodo.20993446. The processed global 1.5° atmospheric and ocean reanalysis data and ocean observation data used for training are available at https://doi.org/10.5281/zenodo.22908836. The saved forecast and B-matrix models, together with the experimental output arrays, are available at https://doi.org/10.5281/zenodo.22930636. The output record also provides a manifest and instructions for reconstructing the model files that were divided into parts for upload. The ERA5, Copernicus Marine, and NOAA references identify the original data providers, while the Zenodo records identify the specific processed inputs and outputs used in our work. Zenodo’s preservation and removal policies are published at https://about.zenodo.org/policies/.
We fully appreciate why access to these materials during public discussion is essential. We will revise the manuscript’s “Code and data availability” section and bibliography to cite all three records clearly when a revised version is requested. We are very keen to have our work evaluated through GMD’s public discussion and peer-review process, and we respectfully ask that the review be allowed to continue now that the code, processed inputs, models, and outputs have been archived. We would be grateful for the opportunity to address any further concerns you may have.
Sincerely,
The authorsCitation: https://doi.org/10.5194/egusphere-2026-3488-AC1 -
CEC2: 'Reply on AC1', Juan Antonio Añel, 25 Sep 2026
reply
Dear authors,
Thanks for addressing this issue so quickly. I have checked the repositories and we can consider now the current version of your manuscript in compliance with the code policy of the journal.
Juan A. Añel
Geosci. Model Dev. Executive Editor
Citation: https://doi.org/10.5194/egusphere-2026-3488-CEC2
-
CEC2: 'Reply on AC1', Juan Antonio Añel, 25 Sep 2026
reply
-
AC1: 'Reply on CEC1', Baoxu Li, 25 Sep 2026
reply
Viewed
| HTML | XML | Total | BibTeX | EndNote | |
|---|---|---|---|---|---|
| 221 | 77 | 32 | 330 | 31 | 28 |
- HTML: 221
- PDF: 77
- XML: 32
- Total: 330
- BibTeX: 31
- EndNote: 28
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
Dear authors,
Unfortunately, after checking your manuscript, it has come to our attention that it does not comply with our "Code and Data Policy".
https://www.geoscientific-model-development.net/policies/code_and_data_policy.html
The sites that you cite to access the data used in your study do not fulfil GMD’s requirements for a persistent data archive because:
-They do not appear to have a published policy for data preservation over many years or decades (some flexibility exists over the precise length of preservation, but the policy must exist).
-They do not appear to have a published mechanism for preventing authors from unilaterally removing material. Archives must have a policy which makes removal of materials only possible in exceptional circumstances and subject to an independent curatorial decision,
-They do not appear to issue a persistent identifier such as a DOI or Handle for each precise dataset.
If we have missed a published policy which does in fact address this matter satisfactorily, please post a response linking to it. If you have any questions about this issue, please post them in a reply. Also, if the size of the datasets used is too big (of the order of hundreds of GB or TBs) to share it in an appropriate repository, please, provide justification for it. Also, you must publish the output datasets produced in your work, and currently it is not clear if the Zenodo repository that you have cited include them. Please, clarify it.
The GMD review and publication process depends on reviewers and community commentators being able to access, during the discussion phase, the code and data on which a manuscript depends, and on ensuring the provenance of replicability of the published papers for years after their publication. Please, therefore, publish your code and data in one of the appropriate repositories and reply to this comment with the relevant information (link and a permanent identifier for it (e.g. DOI)) as soon as possible. We cannot have manuscripts under discussion that do not comply with our policy.
Later, if the Topical Editor decides to continue with the review or publication process of your manuscript and you are requested to upload a new version of it, then The 'Code and Data Availability’ section of your manuscript must also be modified to cite the new repository locations, and corresponding references added to the bibliography.
I must note that if you do not fix this problem, we cannot continue with the peer-review process or accept your manuscript for publication in GMD.
Juan A. Añel
Geosci. Model Dev. Executive Editor