the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
Too good to be true: Underdispersion in geochronology
Abstract. Statistical hypothesis testing is widely used in geochronology to assess whether multiple analyses of a sample are consistent with a single age. Failure of such tests is evidence for excess scatter ('overdispersion'), suggesting geological complexity or faulty data. In contrast, this paper highlights the opposite and largely overlooked problem of 'underdispersion', in which datasets agree unrealistically well with the null hypothesis and appear "too good to be true". Underdispersion can arise from incorrect propagation of analytical uncertainties or from over-zealous outlier rejection that inflates p-values and suppresses genuine geological variability. This paper introduces a simple graphical diagnostic for identifying systematic underdispersion across collections of geochronological studies, based on the empirical cumulative distribution of p-values from chi-squared tests of data homogeneity. While overdispersion shifts p-value distributions above and to the left of the 1:1 line in cumulative probability space, there are no natural mechanisms that should produce an excess of very high p-values. The area below the 1:1 line in a cumulative probability plot defines a 'forbidden zone' and provides a meta-analytical signature of underdispersion.
The proposed method is demonstrated using synthetic examples and applied to extensive compilations of published geochronological data, including fission-track analyses, mass-spectrometer-based chronometers reported in high-profile journals, datasets underlying the geologic time scale, and detrital zircon U–Pb age spectra. These case studies reveal widespread and sometimes extreme underdispersion, particularly in fission-track data and 40Ar/39Ar age plateaux, at levels that cannot be explained by chance alone. Underdispersion matters because it creates unwarranted confidence in apparently precise ages at the expense of accuracy and obscures meaningful geological information carried by excess scatter. The diagnostic plot introduced here provides a practical step towards recognising when data have been over-processed, and towards restoring a more balanced treatment of dispersion in geochronology.
Competing interests: The author is a member of the editorial board of Geochronology.
Publisher's note: Copernicus Publications remains neutral with regard to jurisdictional claims made in the text, published maps, institutional affiliations, or any other geographical representation in this paper. While Copernicus Publications makes every effort to include appropriate place names, the final responsibility lies with the authors. Views expressed in the text are those of the authors and do not necessarily reflect the views of the publisher.- Preprint
(851 KB) - Metadata XML
- BibTeX
- EndNote
Status: final response (author comments only)
- RC1: 'Comment on egusphere-2026-2783', Marco Colombo, 15 Jul 2026
-
CC1: 'Review of Vermeesch 2026, Too good to be true: underdispersion in geochronology', B. Schoene, 18 Aug 2026
Publisher’s note: this comment is a copy of RC2 and its content was therefore removed on 20 August 2026.
Citation: https://doi.org/10.5194/egusphere-2026-2783-CC1 -
RC2: 'Review of Vermeesch 2026, Too good to be true: underdispersion in geochronology', B. Schoene, 19 Aug 2026
The manuscript by Pieter Vermeesch is an interesting read that tackles the widespread practice of culling geochronologic data under the premise that the community sees analytically disperse data as a failure on the analysts part rather than a lens into geologic processes. In other words, Vermeesch’s assertion is that in too many publications, if the data don’t fit the assumed model (be it instantaneous crystallization or cooling of a population of minerals or colinearity on an isochron), that this is generally seen as a bad thing and that data should be culled until they do fit that model. He argues that this practice is common and bad science. I tend to agree with him generally and there is a paper trail of others stating so as well over the last decade or two, though less blatantly than this one. So in one sense, Vermeesch is trodding on a well-worn path here, but he is not wrong to point out that this problem persists. Additionally, he brings the issue to us a way that I haven’t seen published before, which is to quantify just how pervasive data culling is in the literature, in particular for a few different geochronologic methods and/or publications. So perhaps someone such as Vermeesch, with statistical gravitas, can remake this point to the community in a way that will have a positive effect. So I support publishing this paper, but have a few comments/questions that I’ll go through below that may serve as a way to improve some of the points or result in a more nuanced paper.
1) The paper comes across a bit like it was produced in a near vacuum, where data culling decisions are made with the singular goal of fitting the desired statistical model using only the ratios or dates themselves. To be fair, the paper hints that there are decisions being made to cull data that are based on geological inference or complimentary data that cannot be entirely captured by statistical analysis. However, there are many examples where geochronologists use data such as mineral textures, geochemical information, or geologic relationships to aid in data analysis, and the result is that data culling can be guided by things other than (or in addition to) simply attaining an acceptable MSWD. One example would be using geochemistry of zircons to pick which ones to include in a weighted mean. You could argue that this process should still not result in data that is “too good to be true”, which would be a fair point, but I do think a little more discussion of this nuance would enrich the paper a bit.
2) On the other and, there are certainly examples of authors taking over-disperse data sets and deleting data until the MSWD is deemed acceptable. An example from my own community is the practice of taking overdisperse zircon data from an ashbed and subjectively deleting older data points until there is a young population which yields a weighted mean with a reasonable MSWD. The temptation to do so is that you get a higher precision date by taking a weighted mean. The assumption in doing so is that there is some population of zircons that crystallized instantly before eruption and that this ideal population is captured in our data but obscured by pre-eruptive zircon growth that is also sampled. In my reading of Vermeesch’s paper, he is not arguing that culling data to yield an appropriate MSWD is bad, but that people are doing it badly. He suggests a method by which people do so to not consistently yield data that are unrealistically underdispersed (section 6.2, Fig. 3).
This is all good, but I would caution against setting a tone with readers that if you choose to cull data to fit your model assumption, as long as you do so following these guidelines, you’re good to go. I think that Vermeesch agrees with this based on his cited 2025 paper. We should all know that the population of zircons that crystallized instantly before eruption does not actually exist, and I would think that if we had infinitely high-precision dates, this would become obvious. Over and over again we see that as analytical precision improves in geochronology, populations of crystals that once appeared to meet a single population (i.e., pass the chi-squared test) no longer do. There are some published examples that could be cited that show how high-N weighted means can result in statistically significant but inaccurate dates. So while there are better ways to cull/interpret overdisperse data (I am fond of published Bayesian approaches, which of course don’t yield MSWDs), I would still recommend noting in the paper more strongly that just because you do get a good MSWD, it doesn’t mean you get an accurate age.
3) A perhaps minor point that might be worth noting. Due to the issues with weighted means from overdisperse data outlined in Vermeesch’s paper and others, there are now [an increasing proportion of?] papers that simply do not calculate weighted-means and therefore do not report MSWDs. If I understand the search criteria laid out in this paper, those would not show up in the database and therefore the analysis is biased towards bad players. My guess is this bias is tiny. I can only think of one paper with U-Pb data in Nature journals that opt out of using weighted-means, but there are likely more, and I could list many such papers in non-Nature journals.
On that note, it might be interesting to plot a figure like Fig. 8 or 9 from say 2000-2010 compared to 2015-2025 and see if there’s any difference. My own experience is that 20 years ago (when I was a graduate student), there was certainly a strong notion that if you produced an overdispersed U-Pb dataset then you failed somehow. But that changed once precision increased to where data dispersion was the undeniable and the norm and became a tool to learn more about magmatic and metamorphic systems. But again, the papers that take that approach may not report MSWDs so would not be in the database.
4) The tack this paper takes of pitching itself as introducing a tool for identifying underdispersion falls a little flat for me. The paper more aptly reads to me as one that uses that tool to highlight the problem of overculling and arguing that it is quite pervasive and not good. I don’t imagine a bunch of folks will use this new tool; though I hope some do for their own communities or lab groups. Regardless it won’t need to be used very often so maybe isn’t the best way to pitch the paper.
some minor points:
50: scales with 1/sqrt(N) if one takes a weighted mean, yes. Which not every study does.
145: typo
231: typo
section 6: Are all the equations necessary? They are standard stuff…and also built into most software that people would use to do any of this. I suppose the author could make an argument for why it’s worth repeating in this paper?
342: typo, Fig. 8
486 and throughout: It is stated that overdispersed datasets are more likely to be rejected from a journal than underdispersed ones. Does the author have data to back this up or is this based on personal experience? If the latter, I’d recommend against making statements like this based on one or more personal experiences. My own contrary personal experience includes many published papers with overdispersed datasets. How one interprets those datasets is likely the sticking point.
section 7.3: It is perhaps beyond the scope of this paper to ask for concrete examples where overculling of data resulted in an overly precise or inaccurate date that heavily influences the interpretations and/or conclusions of a paper. But…this would be the section to do this in because studies focussed on the GTS are really pushing data interpretation to result in the highest precision dates - simply because the questions being asked around these boundaries require the highest precision dates.
Citation: https://doi.org/10.5194/egusphere-2026-2783-RC2
Data sets
FAIR code and data Pieter Vermeesch https://doi.org/10.5281/zenodo.20186900
Viewed
| HTML | XML | Total | BibTeX | EndNote | |
|---|---|---|---|---|---|
| 147 | 76 | 17 | 240 | 8 | 7 |
- HTML: 147
- PDF: 76
- XML: 17
- Total: 240
- BibTeX: 8
- EndNote: 7
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
See the attached PDF document.