the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
Considering mineralization local-global geological features: An interpretable DCN-Transformer hybrid model with attribution for mineral prospectivity mapping
Abstract. Mineral prospectivity mapping is a critical task in mineral exploration, effectively integrating which demands models capable of capturing both local and global geological features. While deep learning models excel in this domain, their "black-box" nature often limits the trust and insights geologists can derive from their predictions. This paper introduces a novel and interpretable hybrid model to explicitly address this challenge. Our architecture synergistically combines a deformable convolutional network for adapting spatially varying local mineralization features, such as geochemical anomalies, a Transformer module for modelling long-range global-scale spatial features governing mineral deposition. A pivotal innovation is the incorporation of an attribution branching network that generates significance scores for each input predictive factor to the final prospectivity probability. These scores not only provide a direct interpretation of factor relevance but are also fed back to dynamically modulate the key values in the Transformer's attention mechanism, effectively injecting prior geological knowledge into the local and global feature learning process. This design fosters a more geologically informed integration of local and global representations. The model's performance is evaluated against benchmark models including standalone deformable convolutional network, Transformer, and a hybrid model with deformable convolutional network and Transformer. Results demonstrate a superior predictive accuracy and more geologically plausible prospectivity maps. Furthermore, we provide a multi-faceted interpretation framework: the attribution branching network quantifies the contribution of each evidence layer, while gradient-weighted class activation mapping visualizes the discriminative local regions highlighted by the convolutional components and attention maps reveal the long-range feature relationships prioritized by the Transformer to elucidate how the model hierarchically integrates local and global features to arrive at its decisions. This dual-path interpretation strategy demystifies the model's decision-making process, offering geologists tangible insights into both "where" and "why" the model identifies high-potential zones, thereby bridging the gap between high-performing deep learning and actionable c intelligence.
- Preprint
(3106 KB) - Metadata XML
- BibTeX
- EndNote
Status: final response (author comments only)
-
RC1: 'Comment on egusphere-2026-2616', Anonymous Referee #1, 26 Jul 2026
- AC1: 'Reply on RC1', Yihui Xiong, 20 Sep 2026
-
RC2: 'Comment on egusphere-2026-2616', Anonymous Referee #2, 16 Aug 2026
The manuscript presents an attribution-guided DCN-Transformer framework for mineral prospectivity mapping, a relevant modelling problem within the scope of GMD. Combining deformable convolution, attention, and model explanation may be useful. However, substantial concerns currently limit confidence in the central claims of local-global modelling, predictive generalization, interpretability, and methodological novelty. The manuscript needs a more rigorous evaluation and a substantially more complete, precise description before its conclusions can be assessed reliably.Major comments1. The abstract is substantially too long and reads more like an extended introduction and methods summary than a conventional abstract. It should be shortened to a concise, self-contained version accounting for the research question, the specific methodological contribution, the study setting/data, the evaluation design, the principal quantitative result with appropriate qualification, and the main supported conclusion. All the background details, claims, and unsupported statements should be removed.2. It is difficult to read the manuscript because the geological setting chapter (2) is not connected with the methodology. The authors should clearly explain how the geological setting and data characteristics motivate the modelling choices. Partly this is addressed in the introduction, but due to the length of the data chapter, the connection is lost.3. The manuscript presents a potentially useful integration of deformable convolution, Transformer-based spatial modeling, and attribution-guided attention for mineral prospectivity mapping. However, to my opinion its scientific novelty is currently moderate because the principal architectural and interpretability components are largely established. The authors should clarify how the attribution module is genuinely distinct from DeepLIFT, define the geological prior knowledge incorporated into the model, and moderate the novelty claims unless these issues are resolved.More specific comments1. The explanation method needs a clearer and more accurate description. Equations (1)-(4) appear to define a learned branch that identifies features differing from the average training data. This does not look like the established DeepLIFT method, although the manuscript calls it DeepLIFT-based. The manuscript later also uses Integrated Gradients, but it is unclear what each approach contributes. The authors should clearly explain what each method does, what result it produces, and how it influences the Transformer. They should also state whether the claimed geological prior knowledge comes from independent geological expertise or is learned from the data. Learned scores should not automatically be described as geological prior knowledge. Finally, the authors should show that the explanations are reliable, for example by checking whether they remain similar across repeated training runs and after reasonable changes to the input data.2. The manuscript should include a systematic ablation study separating the contributions of the deformable convolution, Transformer module, attribution-guided attention, and attribution regularization. The authors should add component-removal experiments and compare the complete model with appropriate reduced variants to demonstrate the effectiveness and necessity of each proposed component. Without such analysis, the claimed methodological improvements and scientific novelty remain insufficiently supported.3. The input is a 9 x 9 x 42 neighbourhood (chapter 4) centered on each target pixel. With the reported 1 km grid, direct attention is consequently confined to approximately a 9 km local window; it cannot directly learn relationships among locations hundreds of kilometres apart, as repeatedly claimed in the manuscript. Producing a full-map prediction or visualizing map-wide scores does not itself establish that the model attended jointly to distant locations. The authors should document the tokenization and full-map inference procedure, including how window-level attention is mapped to regional outputs. I think one sohuld either define and consistently use a window-scale meaning of "global", or demonstrate that the model incorporates the regional-scale context.4. In the results authors describes 23 positive and 23 negative locations, augments these to 196 samples per class, and only then partitions the data 8:2 into training and validation sets. This sequence risks placing augmented versions, or strongly overlapping neighbourhoods, of the same original location in both subsets and can materially inflate reported performance. I would recommend to split original deposits/locations before any augmentation and apply augmentation only inside each training fold. And a completely untouched outer test set, with the split identifiers disclosed, is needed for final performance claims.5. The reported AUC improvement over the DCN-Transformer baseline is 0.960 versus 0.951, which is a relatively small difference.At the same time, no repeated-run variability, confidence interval or statistical test analysis is reported. The manuscript should provide per-run results over multiple seeds, uncertainty intervals, raw numerical metric tables, and ROC and precision-recall curves. Ideally any paired comparison must be performed on a genuinely independent test set. Until such evidence is supplied, claims of "statistically significant" improvement, "state-of-the-art" performance, and a "new benchmark" should be substantially moderated.6. Overall, the authors should position the work explicitly against close prior work on CNN-Transformer mineral prospectivity mapping, deformable convolution for geochemical anomaly recognition, and interpretable or attention-branch MPM models. A comparison table should identify the exact new component and its intended benefit. Controlled ablations should isolate the effects of DCNv2, Transformer fusion, and attribution guidance, rather than only comparing whole architectures. This is necessary to demonstrate novelty, give proper credit to related work, and establish that the contribution is more than a new combination of existing modules.ÂMinor comments1. Figure 4 should use a common color scale bar, and the symbols repeat for all 4 panels, so it is better to place one legend for all panels2. In Eq. (11), the term "FTN" is undefined and appears to be a typographical error for "TN" (True Negative). In addition, the expression for the expected agreement, P_e, uses n/2 for both classes. This is valid only if the evaluation set contains exactly equal numbers of positive and negative samples; otherwise, the standard expression based on the observed class and prediction totals should be used.3. P19L408 - The statement that the threshold-dependent metrics were calculated at an "optimal probability cutoff" is also insufficient. Accuracy, precision, recall, F1, Kappa, and MCC all depend on this cutoff, whereas AUC does not. The authors should state the numerical cutoff, the criterion used to select it, and the data used for selection. The authors should also report the resulting confusion matrix for the test set.4. Would be good to supplement the radar chart in Fig. 5 with a numerical table of all metrics and their uncertainty. The sample count underlying every reported metric should be stated.5. The manuscript should be edited throughout for grammar as well as for imprecise or promotional language. In particular, "global", "geological prior knowledge", "DeepLIFT-based", "interpretability", "statistically significant", and "state-of-the-art" should be used only where the method and evidence support those terms.Citation: https://doi.org/
10.5194/egusphere-2026-2616-RC2 - AC2: 'Reply on RC2', Yihui Xiong, 20 Sep 2026
Viewed
| HTML | XML | Total | BibTeX | EndNote | |
|---|---|---|---|---|---|
| 158 | 94 | 24 | 276 | 14 | 16 |
- HTML: 158
- PDF: 94
- XML: 24
- Total: 276
- BibTeX: 14
- EndNote: 16
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
This paper proposes an interpretable DCN–Transformer hybrid model with attribution for mineral prospectivity mapping. The proposed framework is novel and potentially valuable. I recommend the authors address the following minor issues to further improve the clarity and quality of the manuscript.
1. Â Â It is suggested to adjust the subheadings of the methods section in Section 3. The proposed framework consists of four major components, whereas the current subsection structure only reflects three of them. I suggest revising the subsection organization so that it better corresponds to the overall framework and more clearly presents the model architecture and workflow.Â
2. Â Â Please check whether the title of Figure 4 is accurate. The current title refers to geochemical anomalies associated with mineralization, whereas the figure appears to present mineral prospectivity maps according to the description in the manuscript. The terminology should be made consistent.Â
3. Â Â Figure 7 presents the attribution results for ten mineral deposits and states that these deposits correspond to those shown in Figure 2. However, the deposits are not numbered in Figure 2, making it difficult for readers to establish the correspondence between the two figures. I suggest adding deposit numbers to Figure 2 or providing another clear way to link the deposits shown in the two figures.Â
4. Â Â In line 320: "Prior to model training, data augmentation, which was adopted in Yang and Zuo (2024), ..." Since data augmentation is an important part of the training process, I recommend briefly describing the augmentation strategy (e.g., augmentation operations and sample expansion) rather than relying solely on the citation. This would improve the completeness and reproducibility of the methodology.Â
5. Â Â The manuscript would benefit from careful proofreading. There are several grammatical and typographical errors:
1) the first sentence of the Abstract ("Mineral prospectivity mapping is a critical task in mineral exploration, effectively integrating which demands models capable of capturing both local and global geological features") contains a grammatical error.Â
2) the phrase "actionable c intelligence" in the Abstract appears to contain a typographical error.Â
3) There are grammatical errors in Line 73 "However, the above deep learning architecture which proficiency with local features also imposes a fundamental limitation..."
4) There are grammatical errors in Line 77 "The above deep learning architecture (CNN, GNN, and DCN) analyzing..."
A thorough language check is recommended before publication.