the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
Flood nowcasting based on deep learning and radar rainfall estimates: A reliable and efficient framework for diverse flood regimes
Abstract. Floods pose a severe risk to lives, infrastructure, and ecosystems, necessitating accurate and timely forecasts to support early warning and emergency response. Flash floods in particular are among the most destructive flood hazards due to their rapid onset, short response times, and thus limited time for warning. However, physics-based or conceptual hydrological models often struggle to deliver reliable short lead-time predictions, particularly in small or fast-responding catchments where complex and rapidly evolving hydrological processes are at play. Additionally, in the case of flashy catchments, there is often insufficient time to run numerical models and issue timely warnings. This study explores the use of encode-decode Long Short-Term Memory (LSTM) networks for short-term flood forecasting using high-frequency radar rainfall and river level data across three locations representing diverse flood regimes in the Hunter Valley in Australia. Validation across these locations demonstrates strong agreement between predicted and observed flood levels, with an average RMSE of 0.09 m and MAE of 0.06 m at a 90-minute lead time, and 0.45 m (RMSE) and 0.24 m (MAE) at a 12-hour lead time. By coupling LSTM-predicted water levels with airborne LiDAR-derived digital elevation models (DEMs), inundation maps were generated to translate point-based flood level predictions into spatially distributed flood extent information. These maps showed high agreement with Sentinel-2 and Sentinel-1derived flood products, achieving over 94 % overall accuracy and up to 84.5 % critical success index at a 12-hour lead time. We also demonstrated the workflow’s operational capability, achieving accurate flood level forecasts supported by radar-based rainfall nowcasts and numerical weather predictions, albeit lower quality longer 12-hour forecasts when input precipitation forecasts have high uncertainty and error. Overall, this study presents a scalable, data-driven approach for real-time flood nowcasting, providing a practical tool to support early warning systems and inform emergency planning in vulnerable regions.
- Preprint
(2507 KB) - Metadata XML
-
Supplement
(655 KB) - BibTeX
- EndNote
Status: final response (author comments only)
- RC1: 'Comment on egusphere-2026-3401', Anonymous Referee #1, 21 Sep 2026
-
RC2: 'Comment on egusphere-2026-3401', Anonymous Referee #2, 27 Sep 2026
This manuscript proposes an encoder–decoder LSTM-based framework for flood nowcasting, using radar rainfall estimates and forecast products, combined with water level observations and LiDAR DEMs, to perform water level prediction and inundation mapping. The manuscript has a sufficient workload and some potential for operational application. However, it has certain shortcomings in terms of core novelty; the encoder–decoder LSTM architecture, radar rainfall forcing, and other components have already appeared in prior literature. At the methodological level, several issues require further clarification, such as the use of STEPS-NWP ensemble information, error decomposition of the HAND method, and the selection of validation events. In addition, there is partial overlap with the first author’s MODSIM2025 conference paper, and the authors must explain this issue. Therefore, my recommendation is Major Revision, with specific comments as follows.
Major Comments
- The novelty is unclear and needs to be explicitly defined relative to existing work. The core methodological combination in this manuscript—encoder–decoder LSTM, radar rainfall forcing, and HAND-based inundation mapping—has been extensively covered in the existing literature. For example, Kao et al. (2020) systematically validated the effectiveness of encoder–decoder LSTM for multi-step flood forecasting, and this work is also cited in the manuscript. A study in the Tamanduateí catchment in Brazil implemented LSTM coupled with HAND-based inundation mapping, also targeting a rapidly responding catchment and using a short antecedent window. Kratzert et al. (2018, 2019) and Nearing et al. (2024), among others, have demonstrated the strong capability of LSTM in rainfall–runoff modeling, even achieving global-scale flood forecasting with lead times of up to five days. The question raised in the Introduction—“whether models trained solely on real-time radar rainfall data can effectively nowcast floods”—is more a choice of experimental design than a fundamental theoretical or methodological breakthrough. Applying an existing architecture to a new region and validating its performance under three flood regimes is closer to regional application validation than to methodological innovation. It is recommended that the authors clearly state the novelty of this study in the Introduction.
- The overlap with the MODSIM2025 conference paper must be clearly clarified. The first author, Jiawei Hou, has published a conference paper entitled “Flash flood nowcasting based on deep learning and radar rainfall estimates” at MODSIM2025, and its abstract substantially overlaps with this manuscript. The authors need to explain this issue, clarifying which parts are a continuation of the conference paper and which parts are new.
- The use of STEPS-NWP ensemble information requires further justification. In Sections 2.3.2 and 3.3, the manuscript mentions that STEPS-NWP provides eight ensemble members, but the study used the ensemble mean as the decoder input “to maintain consistency with the modeling framework.” This choice discards probabilistic forecast information, whereas ensemble information is crucial for uncertainty quantification in operational flood forecasting. The authors need to explain this issue and are encouraged to supplement experiments using ensemble members (or at least ensemble spread) to drive the LSTM, comparing forecast performance between ensemble-mean and ensemble-member inputs. In addition, it is recommended to add a qualitative or quantitative analysis in Section 3.3 or 3.4 of how STEPS-NWP forecast errors propagate to water level forecast errors, so as to enhance the credibility of the framework for operational applications.
- The sources of error in the HAND method have not been decomposed. Section 3.2 presents comparisons of inundation extents with Sentinel-1/2 (OA 94–96%, CSI 63–85%), and Section 3.4 discusses several limitations of the HAND method, such as the uniform hydraulic gradient assumption, neglect of temporary storage, and the inability of LiDAR to penetrate water surfaces. However, the manuscript does not evaluate the specific contribution of these simplifications to inundation mapping errors. In the current results, errors may arise simultaneously from LSTM water level prediction errors, HAND terrain simplification errors, and errors in the satellite products themselves (e.g., omission within river channels in Sentinel-1 and vegetation occlusion in Sentinel-2). The authors are advised to add an error decomposition analysis, for example, by driving HAND mapping with observed water levels rather than LSTM-predicted water levels and comparing the results with those driven by LSTM predictions, so as to separate the contributions of water level errors and terrain simplification errors. If full decomposition is not possible, at least clearly state the limitations of the current validation in the discussion and provide a qualitative judgment of the error sources. In addition, consideration could be given to using a high-accuracy hydrodynamic model for comparison over selected events to evaluate the impact of HAND simplifications.
- The selection of validation events and the temporal split of training data require further justification. Section 2.3.2 states that the historical prediction experiments used data from 2020 and 2022–2025 for training, skipping the entire year 2021, and then used the March 2021 flood event as an independent validation case. Although this approach avoids data leakage, it also means that the validation event is not within the temporal neighborhood of the training distribution. For data-driven models such as LSTM, the temporal coverage of training data directly affects their ability to generalize to extreme events. In addition, the manuscript notes that radar rainfall data currently cover only approximately 2020–2026, so the amount of training data is relatively limited. The authors are advised to explain why 2021 was skipped rather than using other splitting strategies (e.g., proportional random splitting or event-based splitting); if 2021 was excluded because of data quality or event-specific characteristics, this should be clearly stated.
- The interpretation of operational evaluation metrics needs to be more cautious. Table 3 presents operational metrics for the 90-minute and 12-hour lead times. For the 90-minute configuration, the threshold exceedance lead times are “+2 h 40 min” (minor) and “+20 min” (moderate), both earlier than the observed exceedance times; for the 12-hour configuration, the minor threshold lead time is as long as “+1 d 1 h 50 min.” The manuscript considers that “these detections remain true positives” but that “excessively early warnings may lead to premature mobilization of emergency resources.” This interpretation is reasonable, but a more in-depth discussion is needed: the cost of excessively early warnings in operational flood warning can be very significant, such as reduced public trust and wasted emergency resources, and the manuscript needs to discuss how to balance “not missing alarms” and “not issuing false alarms.” For the 12-hour configuration, the flood peak timing error is “+4 h 50 min” (early bias) and the flood peak magnitude error is “+0.14 m”; for operational applications, is this timing error acceptable? This should be discussed in the context of the response time of the specific catchment. In addition, the comparison with the zero-skill benchmark (Figure 7) shows that the LSTM outperforms persistence forecasts at all lead times, but the absolute skill gain at 12 hours is about 0.29 m; the significance of this gain for actual warning decisions needs to be discussed.
Citation: https://doi.org/10.5194/egusphere-2026-3401-RC2
Viewed
| HTML | XML | Total | Supplement | BibTeX | EndNote | |
|---|---|---|---|---|---|---|
| 209 | 115 | 24 | 348 | 36 | 26 | 20 |
- HTML: 209
- PDF: 115
- XML: 24
- Total: 348
- Supplement: 36
- BibTeX: 26
- EndNote: 20
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
The manuscript presents an encoder–decoder LSTM framework using radar rainfall and gauge observations to forecast water levels across three contrasting flood regimes in the Hunter Valley, Australia, and further converts predicted water levels to inundation extents using high resolution LiDAR-derived HAND. The study evaluates historical predictions and additionally tests operational nowcasting using STEPS-NWP rainfall forecasts. This combination of high-frequency operational radar products, water level forecasting, operationally relevant metrics, and inundation mapping gives the study practical relevance.
The manuscript addresses an important problem, and it contains a useful framework and encouraging results. I particularly appreciate the authors’ effort to move beyond conventional error statistics by considering operationally relevant measures such as flood threshold warning lead time, peak magnitude and timing, as well as comparison against a benchmark from a zero-skill model.
Overall, I find the study promising and potentially suitable for publication in the special issue “Early Warning Systems from Research to Operations: Status, Innovations and Multi-Hazard Applications”. However, several aspects of the methodology require further clarification and justification, particularly regarding model reproducibility, the validation strategy, and the distinction between historical prediction and operational forecasting. I’ve provided my comments below to help strengthen the manuscript.
1. More detail is needed on the model architecture, training procedure and reproducibility. The methods provide a useful overview of the encoder–decoder LSTM framework, but the description of the actual model implementation and training procedure is currently rather limited. More information is needed on the model architecture and configuration, training and optimisation strategy, and how key hyperparameters were determined. Providing these details would help readers better understand how the framework was implemented and assessed and would improve the reproducibility of the study. It should also be clarified whether separate models were developed for each site, lead time, and input sequence configuration, or whether a common model setup was applied across the different experiments.
2. The validation strategy requires further justification. For the historical experiment, the model is trained using data from 2020 and 2022–2025 and evaluated on the March 2021 flood event. Although holding out an entire flood event is preferable to randomly dividing a highly autocorrelated time series, validation against only one event makes it difficult to assess robustness across different event magnitudes and hydrometeorological conditions. The manuscript also acknowledges the limited radar record and number of available flood events. I recommend either conducting event-based cross-validation/leave-one-event-out testing where possible, or providing a stronger justification for the current split and moderating claims of generalisability accordingly.
3. The distinction between historical prediction and operational forecasting needs to be made more explicit. In the historical evaluation, observed radar rainfall is used as a proxy for future forecast rainfall. Therefore, the very good 6 h and 12 h results in this experiment do not include rainfall forecast uncertainty and are not directly comparable with operational forecasts at the same lead times. This is particularly important because the nowcasting experiment shows larger and more variable errors at 12 h as the rainfall forcing transitions towards NWP. I suggest making this distinction clear throughout the Abstract, Results, Discussion and Conclusions, so that potential predictability using observed rainfall is not confused with forecast skill using forecast rainfall.
Specific comments:
Line 85: This sentence needs a bit more context, e.g., under what kind of hydrological/catchment conditions
Line 130: Perhaps categorise those three catchments based on their hydrological complexities, for example, from relatively simple to complex settings as described in Table 1
Line 160 Fig.1: replace 15m, 90m with 15min and 90min, and briefly describe OA, CSA, RMSE and MAE in the caption
Line 169: Typo - Fig.2
Line 170 Section 2.1 study area: this section would benefit from a general description of rainfall conditions for your study region, e.g., mean annual rainfall and/or typical rainfall intensity
Line 210 Data: A general summary of datasets used would be helpful, such as radar derived rainfall, water level, DEM, etc. Also it would be helpful to briefly describe the resolution of the datasets.
Line 220: Please clarify the abbreviation ACCESS-CE
Line 225: It is unclear to me whether your water level datasets were directly sourced from the Bureau online portal or from other partners.
Line 255-260: The description states that the “future segment was used by the decoder” and that the decoder takes “radar-based rainfall nowcasts” as input. It is unclear whether this description applies to both model training and prediction. During historical training, are the decoder inputs observed future rainfall, while during operational forecasting they are STEPS-NWP rainfall forecasts? Please clarify this distinction.
Line 281: This should move to the data section
Lines 289–290: Please clarify “15-minute total precipitation … by aggregating observations over upstream catchment areas.” Does the total refer to temporal accumulation of the original 5-min rainfall data, spatial aggregation over the catchment, or both? If spatially aggregated, please specify whether the catchment mean, sum, or another statistic was used and provide the corresponding units.
Line 315-320: The description of the STEPS NWP product should move under the data section
Line 340: A table showing the different thresholds for minor, moderate, and major flooding would be helpful here.
Lines 428–445: The statement that the sensitivity analysis “eliminates the possibility” that model performance arises solely from water level autocorrelation seems too strong. The results show that increasing the length of the historical water level sequence substantially improves performance at longer lead times. I suggest revising this to say that the analysis provides evidence that the model does not rely solely on water-level autocorrelation. Similarly, the claim that the model “effectively learns hydrological response patterns” should be phrased more cautiously unless this is demonstrated independently.
Line 474: Was the “significant difference” tested statistically? If not, please consider rewording.