the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
A Statistical-Physical Prior-Guided Framework for Multi-Class Landslide Susceptibility Mapping
Abstract. Regional landslide susceptibility mapping generally treats landslide/non-landslide discrimination as a binary learning task and subsequently divides predicted probabilities into several qualitative classes. This two-stage strategy does not directly represent the susceptibility gradient and may therefore reduce the spatial concentration of historical landslides within the very high susceptibility zone. This study proposes a statistical-physical prior-guided framework for direct five-class landslide susceptibility mapping along the Ancient Qin-Shu Roads in the Qinling-Daba Mountains. A landslide inventory containing 1,407 records and 14 conditioning factors was compiled. First, certainty factor (CF) analysis was used to generate a initial susceptibility map. The Shallow Landsliding Stability model (SHALSTAB) was then introduced to provide a physically meaningful stability mask with which the zonation was corrected and samples of very high, high, moderate, low, and very low susceptibility were constructed. Logistic regression (LR), support vector machine (SVM), extreme gradient boosting (XGBoost), random forest (RF), Feature Tokenizer Transformer (FT-Transformer), and Tabular Prior-Data Fitted Network (TabPFN) were compared under binary and multi-class settings. The results show that RF performed well in both tasks, attaining an area under the receiver operating characteristic curve (AUC) of 91.08 % for binary classification and 92.53 % for multi-class classification. In the binary result, the landslide-prone zone occupied 19.3 % of the study area and contained 87.3 % of the historical landslides, yielding a Hit Optimization Index (HOI) of 1.68. In the corresponding multi-class result, the very high susceptibility zone occupied only 13.5 % of the study area while containing 91.8 % of the historical landslides, with an HOI of 1.78. These findings indicate that multi-class learning provides greater historical-landslide coverage and spatial concentration. Compared with labels derived from CF alone, the SHALSTAB-constrained scheme increased accuracy from 65.05 % to 68.60 %, AUC from 91.30 % to 92.53 %, and HOI from 1.72 to 1.78, demonstrating that physical constraints can improve susceptibility zonation. Shapley additive explanations (SHAP) showed that elevation, land use and land cover (LULC), and distance to roads were the three principal controls on identification of very high susceptibility zones.
- Preprint
(3709 KB) - Metadata XML
- BibTeX
- EndNote
Status: open (until 01 Oct 2026)
- RC1: 'Comment on egusphere-2026-4550', Anonymous Referee #1, 07 Aug 2026 reply
-
RC2: 'Comment on egusphere-2026-4550', Anonymous Referee #2, 09 Aug 2026
reply
-
The comparison among LR, SVM, RF, XGBoost, FT-Transformer, and TabPFN lacks a systematic hyperparameter-tuning procedure. Using one fixed configuration for each model may bias the conclusion that RF performs best. The authors should perform hyperparameter optimization using a validation set separate from the final test data and report the search space, objective, budget, stopping criterion, random seeds, and selected parameters. The study “A hyperparameter tuning strategy for an LSTM model to simulate reservoir outflows: large-scale evaluation across 441 dams in the CONUS” (DOI: 10.1016/j.jhydrol.2026.136204) provides a useful Bayesian optimization/TPE framework that could be adapted here.
-
The random 70/30 split is not sufficient for a spatial landslide dataset because nearby samples may appear in both training and testing sets, leading to overly optimistic performance. Spatial-block cross-validation, group-wise spatial validation, or leave-one-corridor/region-out testing should be included. Random-split results can be retained, but spatially independent validation is necessary.
-
The construction of the five-class target may introduce circularity. The CF map and SHALSTAB results are derived from the landslide inventory and conditioning factors, while similar factors are later used as ML predictors. Therefore, high accuracy may partly reflect reproduction of the CF/SHALSTAB labeling scheme rather than independent landslide prediction. This concern is stronger if the full landslide inventory was used before the train-test split.
-
The target-generation and validation procedure should be redesigned or clearly justified. CF generation, SHALSTAB calibration, and sample labeling should rely only on training data, while final evaluation should use landslide observations not involved in any previous modeling or labeling step. Spatially, temporally, or event-based independent validation would provide stronger evidence.
-
Differences among the best-performing models are small. For example, RF reports AUC = 92.53% and F1 = 68.65%, while TabPFN reports AUC = 92.30% and F1 = 68.25%. These differences may not be statistically meaningful.
-
The authors should report uncertainty intervals and paired statistical comparisons using identical validation folds and seeds. Confidence intervals or paired bootstrap comparisons would help determine whether the differences between models are robust. Claims that RF “outperformed” the other models should be moderated unless statistical significance is demonstrated.
Citation: https://doi.org/10.5194/egusphere-2026-4550-RC2 -
-
RC3: 'Comment on egusphere-2026-4550', Anonymous Referee #3, 12 Aug 2026
reply
The manuscript presents an interesting statistical-physical prior-guided framework for direct multi-class landslide susceptibility mapping. The integration of a statistical prior derived from the Certainty Factor method with a SHALSTAB-based physical constraint is potentially valuable, and the comparison among conventional machine-learning and deep tabular-learning models is relevant. However, several methodological and conceptual issues should be addressed before the manuscript can be considered for publication.
Comments for the authors:
- The manuscript treats the landslide inventory largely as a homogeneous population, although landslide susceptibility may depend strongly on landslide typology. Different landslide mechanisms can respond to different combinations of topographic, geological, hydrological and anthropogenic conditioning factors. Silva et al. (2018)*, in a susceptibility assessment, demonstrated that models developed separately for different landslide typologies performed better than models based on the total inventory, because the statistical relationships between landslides and predisposing factors may become dominated by the most spatially abundant landslide typology. The authors should consider citing this article. The authors should therefore report the landslide typologies represented in the 1,407-event inventory and discuss whether aggregating all landslides into a single class is appropriate. They should also clarify whether all inventoried landslides are compatible with the shallow-landslide assumptions underlying SHALSTAB.
*Silva, R. F., Marques, R., & Gaspar, J. L. (2018). Implications of Landslide Typology and Predisposing Factor Combinations for Probabilistic Landslide Susceptibility Models: A Case Study in Lajedo Parish (Flores Island, Azores -Portugal).
Geosciences, 8(5), 153. https://doi.org/10.3390/geosciences8050153- The manuscript reports that the historical landslides are used to construct the very-high-susceptibility class and subsequently evaluates the resulting susceptibility maps using historical-landslide coverage and HOI. The authors should clarify whether the same landslide records used to construct the labels are also included in the evaluation of the final maps. If so, the reported HOI and historical-landslide hit rates should not be interpreted as fully independent predictive validation.
- The use of 14 conditioning factors should be more strongly justified. Several factors may represent related environmental processes or may be spatially correlated. The authors should provide a correlation/multicollinearity assessment. This would help determine whether the proposed framework benefits from the full set of factors or whether some factors provide limited additional information.
- The manuscript should moderate the general statement that multi-class classification provides greater historical-landslide coverage and spatial concentration. The results demonstrate this advantage for the present study area, inventory, class-construction procedure and evaluation framework, but they do not necessarily establish that multi-class susceptibility modelling is generally superior to binary modelling. This distinction should be reflected in the conclusions.
- The manuscript harmonizes all conditioning factors to a 30-m grid, although the original datasets have substantially different spatial resolutions and mapping scales. The authors should discuss the implications of this harmonization, particularly for rainfall, NDVI and geological data, and clarify the effective spatial resolution of the susceptibility model. The representation of individual landslides relative to the 30-m mapping unit should also be discussed.
Citation: https://doi.org/10.5194/egusphere-2026-4550-RC3
Viewed
| HTML | XML | Total | BibTeX | EndNote | |
|---|---|---|---|---|---|
| 50 | 16 | 4 | 70 | 7 | 7 |
- HTML: 50
- PDF: 16
- XML: 4
- Total: 70
- BibTeX: 7
- EndNote: 7
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
The study presents a promising direct multi-class susceptibility framework that combines certainty factor analysis, SHALSTAB constraints, six learning algorithms, and SHAP interpretation. The following technically focused comments are intended to strengthen its methodological validity and reproducibility.