the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
Global Sub-national Impact-based Forecasting for Tropical Cyclones Using Open Data: Combining Machine Learning and Exposure-based Approaches
Abstract. Tropical cyclones (TCs) cause substantial and uneven impacts across regions, driven by differences in exposure and vulnerability. While anticipatory action (AA) systems aim to mitigate these impacts, they are typically based on hazard thresholds rather than predicted consequences, limiting their effectiveness and consistency. Impact-based forecasting offers a promising alternative, but existing approaches are often region-specific or rely on non-transferable data. In this study, we develop a global, sub-national impact-based forecasting framework that predicts affected-population fractions using only openly available data. The model integrates hazard, exposure, and contextual features within a two-stage XGBoost architecture and is evaluated across 780 historical TC events using decision-relevant metrics aligned with operational thresholds. Our results show that machine learning improves the detection and spatial localization of impacts, but does not outperform simpler exposure-based approaches in identifying severe events. This reveals a fundamental trade-off between coverage and conservative severity detection, suggesting that hybrid strategies combining both approaches are better suited for operational use. We position this system as a first-generation global benchmark for impact-based forecasting: it demonstrates the feasibility of transferable, sub-national predictions using open data, while clarifying the limitations that must be addressed for reliable deployment in anticipatory action systems.
- Preprint
(3910 KB) - Metadata XML
- BibTeX
- EndNote
Status: final response (author comments only)
- RC1: 'Comment on egusphere-2026-1996', Bernard Alan Racoma, 24 Jun 2026
-
RC2: 'Comment on egusphere-2026-1996', Anonymous Referee #2, 03 Jul 2026
The comment was uploaded in the form of a supplement: https://egusphere.copernicus.org/preprints/2026/egusphere-2026-1996/egusphere-2026-1996-RC2-supplement.pdf
-
RC3: 'Comment on egusphere-2026-1996', Anonymous Referee #3, 27 Jul 2026
The manuscript provides a useful benchmark and is refreshingly open about its limited skill in predicting severe impacts. However, its main claim of globally transferable, grid-scale forecasting is not fully supported because the grid-level targets are derived by redistributing ADM1 totals rather than from actual local observations.
Recommendation: Major revision
Major comments
- The 0.1° targets are not based on observed local impacts. They are created by distributing ADM1 totals according to population. This means grid-scale skill cannot really be validated, and the study may be better framed as an ADM1-level model with disaggregated outputs. Please provide more discussion on this and maybe reframe some of the statements used throughout the manuscript
- The forecast experiment uses observed IMERG rainfall within ±48 hours of landfall, including rainfall that would not have been available at earlier lead times. This is therefore not a true anticipatory forecast. The authors can explore the use forecast/modeled rainfall or present the experiment only as a track/wind sensitivity test.
- Using 2020 population and 2025 urbanization data for events from 2000–2022 creates a temporal mismatch and possible leakage. Historical or nearest-year data would be more appropriate. Or maybe provide a strong justificatio on the use of 2020 as reference year
- Unreported ADM1 impacts should not automatically be treated as zero, given the uneven reporting in EM-DAT. Add more discussion on the limitations of EM-DAT and differences in cross-country reporting and how that might affect your analysis
- The prev_events_5years variable may partly reflect reporting capacity and could disadvantage data-poor regions. Results without this variable should be given more emphasis.
- The 15% high-impact threshold is not clearly justified for global use. The authors should test other percentage thresholds and consider absolute affected-population thresholds.
- Grid cells with the same ADM1 label are not independent. The analysis should account for clustering and unequal weighting across administrative units.
- Overall performance remains modest, with low precision and many severe impacts missed. Claims about operational use and early warning should therefore be toned down.
- Stronger statistical and hazard-based baselines are needed to show whether XGBoost provides meaningful added value, particularly for operational Impact-based Forecasting.
Minor comments
- Title: Make it clearer that the study is a retrospective benchmark or proof of concept rather than a fully operational forecasting system.
- Throughout the manuscript, especially L15–45 and 95–145: Use “risk,” “hazard,” “exposure,” “impact,” and “vulnerability” consistently, as these terms are not interchangeable.
- L19–21: The statement that climate change is shifting TC occurrence toward previously less-exposed regions is too broad. Qualify it by basin, metric, and level of confidence. “Increasing their intensity” should distinguish among observed changes, projected changes, and changes in the proportion of intense TCs.
- L38–40: “Sub-national (grid) resolution” mixes two different spatial concepts. A 0.1° grid is not an administrative level.
- L45–50: The main finding seems to pre-empt the results and discussion. Consider moving this material to Section 4 or the conclusions.
- L54 and elsewhere: Use the acronym “TC,” since it has already been defined at first instance
- L77–78: Explain why the Philippines is singled out here.
- L93–94: Briefly discuss the limitations of using open datasets, including differences in data collection and reporting systems across countries.
- L96: Spell out IBTrACS, NASA GPM/IMERG, SRTM, and the other datasets at first mention. Also specify the dataset versions used.
- L98–102: Bilinear interpolation is not an aggregation method. Clarify whether the continuous rasters were resampled, averaged, or interpolated.
- L105–114 and 280–284: Define the TC selection criteria more clearly, including the intensity threshold, basin treatment, and EM-DAT–IBTrACS matching procedure.
- L110–114 and 300–307: Use “landfall or closest approach” consistently where both are included.
- L140–144: Revise the statement that socioeconomic vulnerability data are not globally available. Several global proxies exist, so the authors should explain why these were not used.
- L143–144: The manuscript states that 19 of the 72 countries were missing from the SHDI dataset. Explain how this gap was handled.
- L154–157: Using 2020 WorldPop and 2025 urbanization data for events from 2000–2022 creates a temporal mismatch and possible future-information leakage. Historical or nearest-year data should be used.
- L180–200 and 280–284: Clarify how a TC affecting several countries is counted i.e. as one event, several country-events, or multiple ADM1–event observations.
- L238–243 and 263–269: Report the numbers and proportions of no-, low-, and high-impact samples.
- Table 1, around L345–350: “15% (some vs. high impact)” is unclear. Consider “≤15% versus >15%” or “non-high versus high.”
- L285–295 and throughout Section 4: Provide uncertainty intervals for all main performance metrics, not just the rolling results.
- L427–434: Consider bootstrap resampling done at the cyclone-event level to preserve dependence among grid cells and ADM1 observations.
- L507–550: The case studies should not be treated as performance evidence unless their selection criteria were defined beforehand.
- L507–550: Include at least one Philippine failure case, given the country’s large contribution to the dataset.
- L470–478 and Figure 7: Use “Kammuri (Tisoy)” at first mention.
- L470–478: Specify the landfall time and location used for Kammuri/Tisoy, which crossed the Philippine archipelago several times.
- L245–255: Note that the Saffir–Simpson scale is not operationally used by PAGASA or most western North Pacific agencies.
- L245–255 and 509–515: The 33 m s⁻¹ threshold represents hurricane-force winds and may miss damaging tropical-storm-force winds and rainfall-dominated of relatively less intense events.
Citation: https://doi.org/10.5194/egusphere-2026-1996-RC3
Viewed
| HTML | XML | Total | BibTeX | EndNote | |
|---|---|---|---|---|---|
| 354 | 156 | 23 | 533 | 27 | 24 |
- HTML: 354
- PDF: 156
- XML: 23
- Total: 533
- BibTeX: 27
- EndNote: 24
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
This manuscript presents a timely and well-executed study on global, sub-national impact-based forecasting for tropical cyclones using a two-stage XGBoost framework and openly available data. The work is particularly valuable in its integration of hazard, exposure, and contextual predictors at global scale, as well as its comparison against operationally relevant exposure-based baselines. The inclusion of decision-oriented evaluation (e.g., 0% vs. 15% thresholds), sensitivity analyses (temporal and geographic), and exploratory forecast-based experiments strengthens the contribution and demonstrates careful methodological consideration. Overall, the paper addresses an important problem and provides meaningful insights into the trade-offs between machine learning and simpler rule-based approaches.
At the same time, several aspects of the manuscript would benefit from clarification and further refinement. In particular, some methodological elements would be clearer with additional explanation, including the definition and use of key concepts such as the ‘impact fraction’, the selection of decision thresholds (0% and 15%), and the role of hyperparameters in the XGBoost models. Similarly, certain components—such as the description of the TIGGE forecasts, the operational interpretation of the two-stage framework, and the aggregation to ADM1 units—would benefit from more explicit documentation or justification, especially given the global and heterogeneous nature of the dataset.
Additionally, there are opportunities to strengthen the interpretation and presentation of results. For example, reporting class distributions more explicitly would improve understanding of model performance under strong class imbalance, and the conclusions and future work sections could be expanded to more clearly articulate the methodological contribution, novelty, and implications of the findings (e.g., the complementarity of machine learning and threshold-based approaches, and the role of rainfall-driven impacts).
I have also attached an annotated version of the manuscript with specific comments and suggestions provided directly in the text.
Overall, these are relatively minor revisions focused on clarity, transparency, and positioning, and I recommend the manuscript for publication pending minor revisions.