the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
Deep Learning-Based Prediction of Marine Heatwaves in the East China Sea
Abstract. Accurate sea surface temperature (SST) prediction in the East China Sea remains challenging because of its highly dynamic oceanic and atmospheric conditions, yet it is essential for regional fisheries management and marine hazard early warning. Here, we propose SwinTrans-ConvLSTM, a spatiotemporal deep-learning framework tailored for SST forecasting in the East China Sea. The model couples the global representation capability of the Swin Transformer with the local temporal-evolution modeling strength of ConvLSTM, incorporates air–sea temperature contrast and vector wind-field features, and is optimized using a curriculum-learning strategy constrained by a physics-informed gradient loss. Experiments show that SwinTrans-ConvLSTM achieves a mean absolute error of only 0.071 °C and a root mean square error of 0.156 °C on the test set, reducing prediction errors by approximately 30 % relative to state-of-the-art baselines. Crucially, hindcasts of the extreme 2022 marine heatwave event further demonstrate that the model can reproduce heatwave occurrence frequency and duration while substantially mitigating the systematic underestimation of extreme SST peaks inherent in purely data-driven models. These results highlight the critical role of thermodynamic-variable reconstruction and training-strategy optimization in improving the robustness of marine extreme-event prediction, and provide a promising technical pathway for high-resolution operational forecasting under complex ocean conditions.
- Preprint
(3947 KB) - Metadata XML
- BibTeX
- EndNote
Status: open (until 08 Oct 2026)
-
CC1: 'Comment on egusphere-2026-4434', Le Gao, 12 Aug 2026
reply
The comment was uploaded in the form of a supplement: https://egusphere.copernicus.org/preprints/2026/egusphere-2026-4434/egusphere-2026-4434-CC1-supplement.pdfReplyCitation: https://doi.org/
10.5194/egusphere-2026-4434-CC1 -
AC1: 'Reply on CC1', Zefang Ma, 19 Aug 2026
reply
Dear Dr. Le Gao,
Thank you very much for your thorough, rigorous, and constructive comment on our preprint. We deeply appreciate your critical insights, which have significantly highlighted areas requiring refinement to ensure the scientific rigor and clarity of our work.
We have prepared a detailed, point-by-point response addressing all of your major concerns—particularly clarifying the boundaries of our methodological contribution, the precise terminology, and confirming the strict absence of future atmospheric forcing in our experimental setup.
Please find our complete and formatted response in the attached PDF document.
Best regards,
Zefang Ma (on behalf of all authors)
-
AC1: 'Reply on CC1', Zefang Ma, 19 Aug 2026
reply
-
RC1: 'Comment on egusphere-2026-4434', Anonymous Referee #1, 13 Sep 2026
reply
The comment was uploaded in the form of a supplement: https://egusphere.copernicus.org/preprints/2026/egusphere-2026-4434/egusphere-2026-4434-RC1-supplement.pdf
-
RC2: 'Comment on egusphere-2026-4434', Anonymous Referee #2, 30 Sep 2026
reply
Summary and General Comments:
This manuscript assesses the prediction of SST extremes/MHWs using SwinTrans-ConvLSTM. In terms of datasets the authors are using OSTIA SST and ERA5 to represent the associated physical parameters and the study area is East China Sea. U and V winds and ΔT (SST – 2m air temperature) are used as dynamical and thermodynamical physical constrains. Definition and metrics from Hobday et al., 2016, Oliver et al., 2018 and Xu et al., 2025 are used to characterize MHWs. Four metrics were introduced to characterize MHWs: mean duration, mean intensity, occurrence frequency, and total MHW days. Metrics like MAE/RMSE/R^2/PCC account for SST field prediction whereas SEDI/POD/FAR measure the discrete extreme even detection skill. The authors deploy a set of deep learning (DL) models to build a comparison between them and assess the best detection skill of all: LSTM, ConvLSTM, Transformer, Swin Transformer and SwinTrans-ConvLSTM. They conclude that the hybrid DL framework, SwinTrans-ConvLSTM is best to predict SST extremes with lead-times up to 7 days with high accuracy.
Therefore, the manuscript presents an extensive set of experiments, figures and metrics, but lacks a consistent conceptual line. Specifically, the relationship between SST/SST anomalies and MHW event prediction is never clearly established and shifts across sections and metrics. I’m finding it difficult to exactly establish the manuscript’s contribution lane. Combined with the technical issues described below (written as suggested revisions, I recommend rejection, rather than revision, since it’ll need a full restructuring instead of incremental fixes.
Major Comments:
- Inflated physics-informed terminology throughout the manuscript. I agree on the MHW rarity and the need of an extra processing step to balance the classes, and L_grad used as spatial/sharpness smoothness here is a good choice. However, toning down the language and naming of the method regarding the physics-informed concept is necessary.
- SEDI equation seems to have a miscalculation: In the original paper that defined SEDI (Ferro and Stephenson, 2011), F is defined as FP/FP+TN, representing the rate out of actual non-occurrences, how many time did we guess wrong, named ‘false alarm rate’. In the cited paper (Shu et al 2025), F is defined as FP/FP+TP and named ‘false alarm rate’. In the reviewed manuscript, the authors call FAR (F) as ‘false alarm ratio’ (and later on rate) and define it as FP/FP+TP. Is the substitution of TN with TP intentional in the manuscript – simply following Shu et al.? In Shu et al there is no documented reasoning for this substitution either. Do the authors consider something different than what the original Ferro and Stephenson calculation considers? Could the authors please double check the sources, their definitions and goals with the metric? Also, the explanation of ‘ "Hits" denote correctly predicted marine heatwave days, "Misses" represent missed observations, and "False Alarms" correspond to false detections ’ might be misleading (to me as well), could the authors possibly add the traditional TP, TN, FP and FN naming to make this calculation clearer? Where: TP = True Positive, TN = True Negative, FP = False Positive, and FN = False Negative. Full reference of original paper: Ferro, C. A. T., and D. B. Stephenson, 2011: Extremal Dependence Indices: Improved Verification Measures for Deterministic Forecasts of Rare Binary Events. Wea. Forecasting, 26, 699–713, https://doi.org/10.1175/WAF-D-10-05030.1.
- I understand that the paper compares four deep learning benchmark models, however a simple SST persistence model to use as a baseline would be very valuable to make the deep learning choice more solid. At minimum a naive persistence Y(t+n) = X(t) without parameters and training – to represent a baseline for the highly autocorrelated SST. Another suggestion would be the anomaly persistence: Y(t+n)= Climatology(t+n) + (SST(t) -Climatology(t)) – should also be straight forward given that daily climatology has already been calculated in the manuscript. An AR(1) baseline fit per gridpoint on the SST anomaly could be another option. Any one of the above would be sufficient as a baseline model and fulfill this review point. These results would not need to be presented in the main manuscript (could fit well in a supplementary information), but the skill comparison between the simple baseline and DL models should be seen in the text (somewhere in the results section around table 3).
- Eq. 2 is missing the multiple lead times component. Given that 𝑌𝑡+n = 𝑋𝑡 + F(𝑋𝑡−𝑘:𝑡), I’d consider that the DL model predicts all lead times at once – still independently but using the same starting point t. However, looking at Fig 3, a second option may have been considered which every predicted day is calculated based on the one before it so the Eq. 2 should incorporate the lead times such as that: with the base case . In the second option the moving 14-day window would switch one day ahead for every predicted day. The interpretation of Fig 3 and Table 4 would also be affected depending on which of the above two options is valid. Please clarify which one of the two the model implements and revise Eq 2, Fig 3 and table 4 explanations if necessary.
- MHW metrics evaluation missing. Section 2.4 and Table 1 describe 4 metrics on which the model would be evaluated upon predicting MHWs following Hobday et al., 2016 and Xu et al., 2025 definitions. These are: mean duration, mean intensity, frequency and total MHW days. However, the Results section 3.5 presents only the intensity (Figs 10-13) as a predicted vs observed comparison. The rest of the metrics are not shown at all. This creates some confusion throughout the manuscript, as to whether we are assessing the skill of predicting MHW events or SST extremes as a continuous field and adding the MHWs characterization on top. I’d recommend either (a) extending the results to show the rest of the MHWs metrics comparisons between predicted and observed occurrences, or (b) if that’s out of scope, clarify in the text that the manuscript’s primary contribution is detection of SST/SST anomaly fields with an additional binary MHW-detection skill (SEDI), and reframe results’ section 3.5 accordingly.
Minor Comments:
- Please add a few sentences on the potential drawbacks of the selected ML models combination. E.g. the combination is more expensive, or it is hard to generalize on a global scale vs a regional study.
- Add stand-alone figure captions. Many figure captions are just a short line of naming the panels. I’d argue that figure captions need to be stand-alone descriptions of the figures thinking that readers that would just want to skim through the paper could understand/interpret in isolation. Please expand the captions to be more self-contained.
- Grammar/language polishing throughout (e.g. ‘this study adopt’ in line 132; ‘As seasons progresses’ in line 385; etc).
Citation: https://doi.org/10.5194/egusphere-2026-4434-RC2
Viewed
| HTML | XML | Total | BibTeX | EndNote | |
|---|---|---|---|---|---|
| 325 | 116 | 33 | 474 | 25 | 26 |
- HTML: 325
- PDF: 116
- XML: 33
- Total: 474
- BibTeX: 25
- EndNote: 26
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1