Preprints
https://doi.org/10.5194/egusphere-2026-3866
https://doi.org/10.5194/egusphere-2026-3866
28 Sep 2026
 | 28 Sep 2026
Status: this preprint is open for discussion and under review for Natural Hazards and Earth System Sciences (NHESS).

A catalogue-verified multiverse audit of machine-learning seismic susceptibility modelling in western Yunnan, China

Guowei Wang, Ze Fan, Guodong Wang, Xin Wei, Ye Fan, and Lu Xu

Abstract. Machine-learning (ML) classifiers are increasingly applied to regional seismic susceptibility map- ping, frequently reporting high discrimination metrics that imply a reliable link between surficial environmental proxies and earthquake locations. We argue that such claims are often conditioned on undocumented analytical choices rather than on a stable physical signal. Using the active Western Yunnan fault belt (97.0°–103.5° E, 22.0°–28.5° N) as a test bed, we execute a transparent, reproducible methodological audit of a complete ML susceptibility workflow. The audit proceeds through four linked stages: catalogue-provenance verification, leakage-aware predictor screening, predictor-coverage quality assurance (QA), and a multiverse evaluation of model behaviour across model families, spatial cross-validation (CV) designs, background-sampling strategies and target definitions. We find that a legacy working inventory contained 173 of 214 records (80.8 %) that could not be verified against the formal China Earthquake Networks Center (CENC) bulletin, and we replace it with a formally verified catalogue spanning 2009–2023. Naïve spatial joining silently discarded the majority of mainshocks; coverage reconstruction recovered the matched sample from 123 to 330 of 331 events (99.7 %). Across twelve defensible analytical branches, the mean spatial area under the receiver-operating-characteristic curve (AUC) ranged from 0.55 to 0.79, with the highest values attached to the least stable configurations. Independent spatial point-process intensity models confirmed that the full surficial predictor stack provided little measurable incremental gain (∆D2 = +0.002) over a simple distance-to-fault baseline. We conclude that, under the tested non-circular surficial predictors, the apparent skill of regional ML susceptibility models is highly conditional on analytical specification. We provide an auditable reporting checklist and a predictor roadmap that prioritises deep geodetic and tectonic covariates for future work.

Publisher's note: Copernicus Publications remains neutral with regard to jurisdictional claims made in the text, published maps, institutional affiliations, or any other geographical representation in this paper. While Copernicus Publications makes every effort to include appropriate place names, the final responsibility lies with the authors. Views expressed in the text are those of the authors and do not necessarily reflect the views of the publisher.
Share
Guowei Wang, Ze Fan, Guodong Wang, Xin Wei, Ye Fan, and Lu Xu

Status: open (until 09 Nov 2026)

Comment types: AC – author | RC – referee | CC – community | EC – editor | CEC – chief editor | : Report abuse
Guowei Wang, Ze Fan, Guodong Wang, Xin Wei, Ye Fan, and Lu Xu

Data sets

Processed data and result archive for the western Yunnan seismic susceptibility audit G. Wang et al. https://doi.org/10.5281/zenodo.21067519

Guowei Wang, Ze Fan, Guodong Wang, Xin Wei, Ye Fan, and Lu Xu
Metrics will be available soon.
Latest update: 28 Sep 2026
Download
Short summary
Earthquake susceptibility maps made with machine learning can look convincing, but their reliability depends on data choices. We audited a western Yunnan case using verified earthquake catalogues, non-circular predictors, coverage checks, spatial validation, and alternative models. The results show weak or conditional evidence rather than a robust map, highlighting the need for transparent checks before such methods inform hazard studies.
Share