the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
The Stippled Gridpoints are Statistically Significant: (Mis)uses of False Discovery Rate Correction for Geospatial Data
Abstract. Peer-reviewed articles in the geosciences routinely assess statistical significance in spatially distributed data. Statistical significance is often assessed independently at each grid point, while formal adjustment for multiple testing is applied less consistently. Although several approaches to account for multiple testing exist, their application to geosciences data is not always straightforward, as these data often exhibit spatially coherent signals.
In this work, we revisit multiple-testing correction in the context of spatially structured datasets. We first highlight how neglecting multiple testing correction can substantially inflate the number of false positives. We further show that the global false discovery rate (FDR) approach, proposed in literature for application in geosciences, can yield counterintuitive and potentially misleading results when applied to spatially coherent signals. To illustrate the latter point, we provide an example based on near-surface air temperature composites following sudden stratospheric warmings. We show that when anomalies are spatially coherent, restricting the spatial domain can increase the FDR-adjusted significance threshold. Consequently, the same underlying field can appear more statistically significant solely due to domain selection, despite unchanged data. We explain this behavior from the rank-based structure of the FDR procedure and discuss its implications for spatial inference and uncertainty quantification in the geosciences.
Building on these insights, we outline practical recommendations for transparent and robust significance assessment in geoscientific applications. These include clearly documenting multiple-testing corrections when adjusted pointwise significance is shown, cautious interpretation of adjusted thresholds, and considering spatially aware alternatives such as regional or cluster-based inference when appropriate.
Overall, our results highlight both the need to account for multiple-testing and potential issues with a naïve application and interpretation of the FDR correction. We hope that our work may contribute to more robust statistical testing in the geosciences.
- Preprint
(3165 KB) - Metadata XML
- BibTeX
- EndNote
Status: open (until 30 Sep 2026)
-
RC1: 'Comment on egusphere-2026-2203', Anonymous Referee #1, 15 Jun 2026
reply
-
AC1: 'Reply on RC1', Michael Schutte, 13 Aug 2026
reply
We thank the reviewer for the constructive and very helpful comments. Our answers can be found in the file attached.
-
AC1: 'Reply on RC1', Michael Schutte, 13 Aug 2026
reply
-
RC2: 'Comment on egusphere-2026-2203', Anonymous Referee #2, 15 Sep 2026
reply
General Evaluation
This manuscript addresses a very timely, practical, and important topic for the geoscientific and geographical community. The misuse and over-interpretation of statistical significance maps (and multiple-testing corrections) is a widespread issue, and the paper clearly highlights key pitfalls that researchers encounter when applying False Discovery Rate (FDR) corrections to spatially coherent fields. Overall, the study is well-motivated and provides a valuable discussion.
However, I have several major conceptual points and specific recommendations that I believe should be addressed to strengthen the manuscript and ensure statistical accuracy before publication.
Major Comments
- Inherent Nature of FDR Control vs. FWER Control
The authors should more explicitly clarify the fundamental mechanism of False Discovery Rate (FDR) control. By definition, controlling FDR means that the threshold adjusts dynamically depending on the proportion of rejections: it permits identifying more false rejections when the analysis contains many rejections (percentage-wise), and fewer false rejections when only a small fraction of tests are rejected. This property directly explains why restricting the analysis domain to a region with widespread anomalies results in a high critical threshold (p* = 9.8%), whereas analyzing the global domain (where significant tests constitute a small percentage) yields a strict threshold (p* = 0.99%).
If a researcher finds this adaptive behavior undesirable or problematic for their application, FDR control is simply not the appropriate tool. In such cases, Family-Wise Error Rate (FWER) control should be recommended, as it strictly guards against any false rejection and eliminates these domain-dependent threshold fluctuations.
- Clarification on FDR Threshold vs. Pointwise Significance (Alpha Selection)
Throughout the text, it is mentioned that FDR control can lead to a higher adjusted threshold than nominal pointwise testing (and thus reject more hypotheses). It is crucial to clarify that under identical nominal significance levels (alpha_FDR = alpha_pointwise), the FDR threshold p* can never exceed alpha_pointwise (since p_(i) <= (i/N) * alpha <= alpha).
The observed phenomenon where p* > 0.05 occurs solely because two different nominal levels were chosen (alpha_pointwise = 0.05 vs. alpha_FDR = 0.1, following Wilks, 2016). Therefore, this is not a statistical failure or anomaly of the FDR procedure itself, but rather a direct artifact of choosing a higher initial alpha for FDR. This distinction must be made explicit throughout the manuscript.
- Domain Selection, Post-Hoc Analysis, and the Role of Total Test Count (N)
While the authors attribute the threshold shift across domains primarily to p-value ranking, the main driver is the reduction in total test count N relative to the proportion of small p-values. Selecting a sub-domain after visually inspecting the data represents post-hoc selection ("data fishing").
In other fields facing similar spatial/high-dimensional multiplicity problems (such as genetics), when testing a post-hoc selected region, standard practice dictates dividing by the total number of candidate tests (N) from the initial dataset, rather than the subset size n. Because the researcher selects the domain based on prior knowledge of the data, the rest of the domain is implicitly assumed non-significant. If the authors keep N equal to the global domain size when evaluating the subset, the FDR threshold for the selected domain will remain consistent with the full domain.
I strongly encourage the authors to emphasize to the geoscientific community that post-hoc spatial domain selection without retaining the full N is statistically invalid—analogous to selecting only small values in group A and large values in group B in a two-sample test and declaring a significant difference.
- Spatial Dependence and N_eff vs. Correlated Multiple Testing Procedures
Regarding the effective sample size estimation (N_eff in Eq. 2), substituting N_eff fundamentally alters the nature of the test (turning it into a Walker-type field significance test rather than FDR control). To properly handle spatial autocorrelation within the FDR framework, I suggest referring to multiple testing corrections specifically designed to account for test dependence (e.g., Benjamini & Yekutieli, 2001; Mrkvicka & Myllymaki, 2023).
- Recommendation to Report Effect Sizes
I recommend adding a recommendation that geoscientists present effect sizes (e.g., physical anomaly magnitude or standardized effect size) alongside p-value maps. Showing effect sizes would mitigate unnecessary debates over whether a statistically significant grid point carries practical or scientific relevance.
Minor / Specific Comments
- Line 48:
Quote: "In some applications, the resulting FDR threshold may even be higher than the uncorrected threshold."
Comment: As noted in Major Comment 2, this statement is misleading. The threshold is only higher because different nominal significance levels were chosen (alpha_FDR = 0.1 vs. alpha_pointwise = 0.05). Please clarify this sentence to reflect that this is a consequence of parameter choice rather than an intrinsic property of FDR.
- Lines 84–86:
Quote: "In practice, selecting alpha_FDR approximately twice as large as the target field significance level can lead to corrected thresholds that exceed the nominal local significance level. This may produce the counterintuitive outcome that more grid points appear statistically significant following FDR correction than when using uncorrected local tests."
Comment: Please clarify here as well that this counterintuitive outcome arises strictly from applying different nominal alpha values for local and FDR tests, not from FDR itself.
- Lines 128–130:
Quote: "The number of stippled gridpoints is even larger than the uncorrected significance testing case (cf. Fig. 1d to Fig. A1)."
Comment: Again, please remind the reader that Fig. 1d uses alpha = 0.1 while Fig. A1 uses alpha = 0.05.
- Lines 156–157:
Quote: "Yet, this behavior only applies to domains dominated by points with relatively large amplitudes, i.e., the FDR correction reduces the significance threshold p* as expected for domains containing only few grid points with strong anomalies."
Comment: It is not the "amplitude" per se that dictates this behavior, but rather the relative proportion of small p-values (i.e., the overall distribution of test statistics and N). I suggest rephrasing this to focus on p-value proportions rather than physical signal amplitudes.
- Lines 191–198:
Quote: "For example, one could reduce an initial alpha_in = 0.1 to alpha_eff = alpha_in * exp(-N_sig,0.05 / N)..."
Comment: Adjusting nominal alpha post-hoc based on data-derived properties (alpha_eff) is statistically problematic. The significance level alpha must be set a priori; modifying it after observing the fraction of significant tests invalidates standard probabilistic interpretations of the error rate.
- Line 218:
Quote: "When a large fraction of grid points exhibit strong anomalies, the resulting FDR threshold can become relatively permissive..."
Comment: Please see Major Comment 1. This permissiveness is a feature of FDR control when signal density is high, rather than a flaw.
Citation: https://doi.org/10.5194/egusphere-2026-2203-RC2 -
RC3: 'Comment on egusphere-2026-2203', Anonymous Referee #3, 20 Sep 2026
reply
The manuscript discusses the use of false discovery rate (FDR) multiple testing control for assessing statistical significance in spatially distributed data in geosciences. The authors highlight important issues to consider, such as considering multiple testing, avoiding post-hoc domain changes, and reporting the assumptions of statistical tests done. The main issue they discuss is that the use of the FDR control for multiple testing can be confusing for spatially coherent/dependent data. They suggest cautious interpretation of adjusted thresholds. The topic is highly relevant for analyses in geosciencies, but some issues need clarification and enhanced description. Below I give some specific comments as well as a few technical details.
The FDR is defined as the proportion of false discoveries among the discoveries (rejections of the null hypothesis). To best of my understanding, the authors do *not* claim that this control would not hold in some cases, but this is not very clearly stated. Instead the authors want to highlight that if one applies/restricts the analysis to a region where many true discoveries are present, the number of false discoveries can be surprisingly high - even though the proportion of false discoveries remains controlled. It would be good to describe the FDR control idea more precisely, to emphasize that the FDR control works as supposed, and then discuss that there are situations where one might like to introduce a stricter control instead of FDR. Hereby, it would be good to also define/explain more precisely what is meant by "relatively permissive".
The performance of the methods is illustrated with a data example. This example shows that the test significances are dependent on the domain choice. It is, however, not possible to show how the number of false/true positives is changing, as the truth is not known. To illustrate differences in false/true positives along domain change and highlight the points that the authors wish to do, simulated experiments where the true null and alternative hypotheses are known could be prepared.
The nominal local significance levels exceedances that are discussed apparently occur only if alpha_FDR is set two times as large as the target and local field significance level. It would be good to state this clearly whenever discussed.
Regarding the discussion on the alternatives to the FDR control. 1) If FDR is considered "too permissive", an alternative would also be to use the stricter family-wise error rate (FWER) control. 2) It is generally not a good strategy to adjust the significance levels after seeing the data and its preliminary significances, so I would be careful on suggesting such a strategy.
The data example uses a resampling approach to test the statistical significance of the t2m anomalies. Here, the 60-day mean t2m anomaly following the Sudden Stratospheric Warming (SSW) dates is computed. Also, 10000 randomly selected 60-day periods from extended winter (November–March) are selected and the mean t2m computed from these. The task is then to compare the t2m anomaly after SSW to the random t2m means. As resampling was done, I would expect here a resampling test being computed for the mean t2m anomaly after SSW to be colder than the random mean t2m. However, the authors used a Welch’s t-test. The basic form of this test is for comparing two samples based on the sample means and their standard errors (not available here). Consequently, it should be clarified what test exactly is used.
Some details:
page 3, lines 78-79: This sentence needs rephrasing on what exactly is meant. To what is the FDR method compared to here?
page 3, lines 84-85: Please explicitly state the different significance levels that you either set or wish to have.
page 4, line 113: What does 'composite' mean in this connection?
page 5, line 127-128: 'anomalies...appear'
Citation: https://doi.org/10.5194/egusphere-2026-2203-RC3
Viewed
| HTML | XML | Total | BibTeX | EndNote | |
|---|---|---|---|---|---|
| 281 | 121 | 30 | 432 | 23 | 21 |
- HTML: 281
- PDF: 121
- XML: 30
- Total: 432
- BibTeX: 23
- EndNote: 21
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
This paper points out potential problems associated with use of the FDR procedure in the context of spatially correlated hypothesis tests. The authors show that the rule of thumb recommended in the 2016 Wilks paper, which was derived from a particular synthetic data setting with relatively few locally significant gridpoints, apparently behaves badly in the extreme restricted-domain example highlighted in the present paper, and so is evidently not optimal in general.
The paper is fairly short, and would be stronger if it were to include a spectrum of synthetic-data simulations aimed at quantifying optimized parameterization of the ratio of the FDR to the local test level, perhaps as a function of the proportion of local tests that are nominally significant (for example as suggested on line 196), or possibly in terms of the relationship between the domain size and the spatial autocorrelation length scale. Figure 2 seems to indicate that setting the FDR level closer to 0.05 might yield more consistent results in Figure 1d. At minimum, it would be interesting to see the counterpart of Figure 1d with equality of the FDR level and the local test level (i.e., 0.05).
A few more specific comments:
para beginning line 65. Not quite accurate: FDR is actual, not expected number of incorrectly rejected nulls. The Benjamini-Hochberg procedure limits (“controls”) the proportion of such rejections, in expectation. So alpha-FDR = 0.1 implies that, on average, no more than 10% of rejected nulls are false positives, and indeed there may be fewer than this. (The characterization on line 91 is correct).
line 110. The test setup here appears to assume implicitly that SSW events are uniformly distributed in the November-March data window, which is substantially longer than 60 days. Is this the case in the observations? Also, is there a physical justification for use of the 60-day period, external to the test data? What is the effect of temperature nonstationarity during November-March?
line 125. The choice of the small Northern European gridbox appears to have been made a posteriori, after calculation of the initial hemispheric analysis. The several papers cited, apparently to justify this choice, presumably were based on the same or substantially overlapping historical data. The possible impact on the second, spatially restricted, analysis should be discussed more fully.