the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
Meteorological and Land-Cover Controls on Grassland Fire Behaviour by WRF v4.4-SFIRE v0.1 with the computational optimization
Abstract. Under global warming, grassland wildfire risk has increased substantially. The WRF-SFIRE model was applied to the extreme grassland wildfire that occurred in Inner Mongolia in 2023, with satellite-derived fire perimeters used to evaluate the effects of meteorological forcing, land cover, and key driving factors on fire behavior, as well as to optimize model performance. For meteorological forcing, the FNL 6 h two-way coupling configuration achieved the highest spatial agreement with observations, with a recall of 39.1 %, whereas ERA5 was more sensitive to short-term meteorological variability but tended to overestimate fire spread. For land cover, FROM_GLC30 (2017) showed relatively high overall agreement but tended to overexpand the fire, GlobeLand30 (2020) produced more conservative simulations, and GLC_FCS30D (2022) achieved the best overall performance. Fire-atmosphere feedback significantly enhanced fire spread rate and burned area. Wind speed exerted the strongest influence on burned area and spread distance, relative humidity mainly controlled spread rate and flame length, and air temperature played a secondary role. Among the fuel-related parameters, fuel load mainly affected flame intensity, fuel depth dominated burned area variation, and fuel moisture effects mainly regulated spread rate. In addition, the optimized computational configuration, particularly the combined use of PNetCDF and asynchronous I/O, reduced runtime by 29.3 % relative to the baseline case. These results clarify the differential effects of input data and key driving factors on grassland fire behavior simulations and provide more reliable support for fire risk warning and management.
- Preprint
(1573 KB) - Metadata XML
-
Supplement
(755 KB) - BibTeX
- EndNote
Status: open (extended)
- RC1: 'Comment on egusphere-2026-3195', Anonymous Referee #1, 23 Sep 2026 reply
Viewed
| HTML | XML | Total | Supplement | BibTeX | EndNote | |
|---|---|---|---|---|---|---|
| 210 | 62 | 24 | 296 | 36 | 22 | 23 |
- HTML: 210
- PDF: 62
- XML: 24
- Total: 296
- Supplement: 36
- BibTeX: 22
- EndNote: 23
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
General comments
Specific comments
The first sentence of the abstract states that “under global warming, grassland wildfire risk has increased substantially.” I do not think this claim is necessary to motivate the study and the cited material does not appear to establish that general statement for global grassland wildfire risk. The paper is about evaluation of WRF SFIRE for one Inner Mongolia event; the motivation should be tied directly to that problem.
The statement that WRF SFIRE was applied using satellite-derived perimeters “to evaluate the effects of meteorological forcing, land cover, and key driving factors … as well as to optimize model performance” does not follow logically. Satellite observations can be used for model evaluation but they have no direct connection to compiler selection, domain decomposition, or I/O optimization. The abstract combines two substantially different objectives.
The manuscript describes an image from 26 April and a post-fire image from 10 May from which the burned area was extracted using BAI and then visually refined. That is a final burn scar not an observed time sequence of fire perimeters. Yet Fig. 5 compares simulations with an “observed burned area perimeter at 13:00 UTC.” What evidence establishes that this final burn scar represents the fire extent at 13:00 UTC? If there is an independent perimeter observation at that time, it needs to be described. Otherwise this comparison is not a valid evaluation of fire progression.
The abstract gives a Recall value of 39.1% without defining what that is. More importantly, Recall is later defined as overlap area divided by observed burned area. This doesn’t penalize simulated burned area outside the observed scar. Recall alone is not a sufficient metric of spatial agreement. Calling the 39.1% case the “best agreement in both extent and location” seems stronger than this metric supports. Some measure of false positive area / precision or an intersection-over-union type measure would be more informative.
Fig. 5 is, to me, a major problem. The observed burned area has a substantial lateral/easterly extension, whereas the simulations predominantly spread northward/northwestward. Even the best scoring simulation does not reproduce the dominant observed geometry. This should be treated as a model failure to understand, not mainly as a basis for ranking the simulations. In a coupled fire-atmosphere model, getting the fire-driving flow wrong is central to the evaluation. The statement that the FD simulations “more closely reproduce the observed fire geometry and spread pathway” needs to be reconsidered in light of Fig. 5.
The introduction states three aims: identify suitable input datasets, identify controlling factors, and improve computational efficiency. These are tasks rather than a clearly stated scientific hypothesis or evaluation question. What, specifically, is being tested? For example, is the hypothesis that one meteorological forcing represents the observed fire-driving flow better? That two-way coupling improves simulation of the observed spread? That a particular land cover representation improves the simulation for a physically identifiable reason? The experiment design should follow from a stated question.
The manuscript correctly states that both FNL and ERA5 are used at 0.25° spatial resolution; their principal stated resolution difference is temporal, 6 h versus 1 h. However, differences between FNL and ERA5 cannot simply be attributed to temporal resolution. These are different analysis/reanalysis systems with different atmospheric states and data assimilation. The FNL–ERA5 comparison a not a clean temporal resolution experiment. The manuscript states that ERA5 produces larger burned areas and higher spread rates and concludes that higher temporal resolution is “more favorable for representing fire spread.” That conclusion does not follow from the comparison. A larger or faster simulated fire is not necessarily a better simulation.
Wind speed is perturbed by ±2 and ±4 m s⁻¹, RH by ±5 and ±10 percentage points, and temperature by ±2 and ±4°C throughout the entire domain and forcing period. The manuscript then compares the resulting percentage response of the fire metrics and labels the largest response the “dominant” control. But these perturbations are arbitrary and are not normalized to comparable climatological variability, uncertainty, variance, or nondimensional sensitivity. Thus, conclusions such as “wind speed was the dominant meteorological factor” are conditional on the perturbations chosen. They should not be presented as general physical rankings.
I do not understand the interpretation that air temperature “played a secondary role.” What physical pathway is being tested here? SFIRE uses the Rothermel spread formulation, and the manuscript separately enables/disables a fuel moisture model. The authors need to explain how perturbing atmospheric temperature affects fire spread in this configuration and whether that response occurs through fuel moisture, atmospheric dynamics, stability, surface fluxes, or some other pathway. Without that explanation, comparing the magnitude of a temperature perturbation response with wind or RH does not establish the relative physical importance of temperature.
Several reported findings appear largely to reflect the model formulation itself: increasing wind increases spread; modifying fuel load changes intensity; modifying fuel depth changes behavior according to the Rothermel fuel representation. The manuscript needs to distinguish between a genuinely emergent coupled-model result and a response already prescribed by the fire spread equations. Otherwise, it is difficult to determine what was learned from the experiments.
The three land cover products are not merely alternative spatial maps. They are converted to different numbers and distributions of Anderson fuel classes: two produce four mapped fuel classes, whereas GLC_FCS30D produces nine. Consequently, differences between the simulations are not solely “effects of the land cover dataset.” They also reflect the chosen land cover to fuel mapping. This distinction is especially important because the paper draws conclusions about which land cover dataset performs best.
The manuscript says FROM_GLC30 gives better overall spatial agreement with the observed burn scar, while GLC_FCS30D better identifies fine-scale nonburnable patches. The abstract nevertheless states that GLC_FCS30D achieved the “best overall performance.” What metric defines “overall performance”? That needs to be stated consistently.
The manuscript concludes that fire-atmosphere feedback significantly enhances spread and burned area. But in the time series the FD and noFD cases cross, and later the noFD burned area can exceed the FD area. The physical interpretation therefore seems more complicated than simply “feedback enhances spread.” The paper should distinguish time-dependent and process-dependent effects rather than collapsing them into a single statement.
I understand what is being done: compiler comparisons, MPI versus MPI+OpenMP configurations, grid decomposition, and NetCDF/PNetCDF/asynchronous I/O tests. The question is what scientific or generally transferable result comes from these experiments. The compiler differences are only a few percent and are specific to the Intel versions and hardware tested. The conclusion that compiler versions should be benchmarked on the target architecture is unsurprising.
The PNetCDF/asynchronous-I/O and decomposition results may be more useful. If these are intended as a contribution to GMD, the manuscript should make clearer what is specific to WRF SFIRE, why the behavior occurs, and how transferable the recommended configurations are to other domains, processor counts, and architectures.
The manuscript repeatedly connects these results to early warning, emergency response, and operational simulation. That claim seems premature when the simulated fire does not reproduce a major feature of the observed spread direction and the observational evaluation itself is uncertain. Computational speed is only useful operationally if the forecast has demonstrated physical and predictive skill.
Technical corrections / presentation
These are mostly presentation issues rather than scientific ones: