the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
Statistical Assessment of the Representative Elementary Area for Areal Fracture Intensity (P21) in Digital Outcrop Models
Abstract. The definition of a representative elementary volume (REV) or area (REA) for a target parameter is a fundamental step toward the upscaling of fracture network properties generated by discrete fracture network models (DFN) to an equivalent continuous medium for in situ applications in engineering geology, hydrogeology, and structural geology. The target parameter of this work is the areal fracture intensity (P21), a key metric often used as a stopping criterion in stochastic DFN simulations, that is derived directly from surface data collected at natural outcrops. We propose a novel approach to define the REA as a range bounded by a lower and an upper limit. The upper limit, often overlooked but nonetheless theorized, identifies the largest representative domain, which is crucial for optimizing computational efficiency. We evaluate the REA range based on three statistical parameters, namely: the shape, mean, and variance of the P21 distributions obtained with progressively increasing scan area sizes. Each statistical parameter is assessed by combining formal statistical tests and diagnostic plots. Within a multi-parametric framework, the method enables a detailed analysis of the statistical behaviour of the dataset, supporting informed decisions in defining the REA range. The methodology is tested on two fractured limestone outcrops with markedly different characteristics: (i) an abandoned quarry in the Murge Plateau (Puglia, Italy) and (ii) the Lilstock Benches in the southern Bristol Channel basin.
- Preprint
(3724 KB) - Metadata XML
- BibTeX
- EndNote
Status: final response (author comments only)
- CC1: 'Comment on egusphere-2026-2409', Jie Liu, 10 Jun 2026
-
RC1: 'Comment on egusphere-2026-2409', Jie Liu, 12 Jun 2026
The manuscript "Statistical Assessment of the Representative Elementary Area for Areal Fracture Intensity (P21) in Digital Outcrop Models" focuses on the range (lower and upper bounds) of REA by investigating the shape, mean, and variance of P21 distributions obtained with progressively increasing scan area sizes. The authors applied the approach at two field sites with distinct fracture characteristics and conducted subsequent analyses and discussions. The methods, results, and analyses appear reasonable; however, there are three major problems with the manuscript.
I. The practical motivation for defining an upper REA bound is not well established.
The paper presents the determination of an upper bound of the REA range as a key novelty. However, the practical motivation for defining such an upper bound is not well established. The REV/REA literature has historically focused on the lower bound, largely because the engineering requirement is to identify the smallest scale at which a property is still representative. The need for a statistically defined upper bound is far less clear: a scale larger than the lower bound remains representative in principle, and grid resolution in reservoir models is typically constrained by geological boundaries (e.g., lithological contacts, fault zones) rather than by the upper limit of statistical homogeneity.
If the upper limit of REA is intended to define the maximum cell size in numerical modeling (Lines 67–72), the result for the Lilstock outcrop (0.3 m) is counterintuitive when viewed against Fig. 2b, even though it may be justifiable on purely statistical grounds. Using the maximum cell size of 0.3 m to perform numerical modeling of this area (Fig. 2b) is not reasonable from the perspective of a numerical modeller. The paper would benefit from a more critical discussion of why an upper REA bound, derived solely from P21 statistics, should be preferred over, or even considered alongside, geologically based cell-size decisions.
II. The manuscript's organizational logic needs improvement.
1) Placing a subsection entitled "Data collection" within the "Method" section is confusing.
2) In Line 161, Eq. [4] is referred to, yet Eq. [4] is not presented until Line 216, several subsections later.
3) Figures 4 to 7 are mentioned in Section 3 (Method), whereas these figures are explained in Section 4 as results.
4) The Method section should introduce the concepts and methods used in the study using purely mathematical language and expressions. The subsequent section can then describe data acquisition (including the studied sites) and implementation details, accompanied by necessary schematic figures, but without forward references to figures that belong to the Results section. By this standard, the implementation of the analyses has not been explained clearly.
5) In the paragraph from Line 324 to Line 332, the discussion order is confusing: Fig. 9C and Fig. 9D are discussed first, followed by Fig. 10, and then Fig. 9A and Fig. 9B.
III. The manuscript requires much more concise writing.
Some statements are repeated several times across different paragraphs or sections, for example, regarding the calculation of P21. The sentence in Line 175 ("P21 is not calculated using …") is unnecessary.
More broadly, the manuscript contains many unnecessarily detailed interpretations and self-explanatory expressions, which greatly reduce readability. The two paragraphs from Line 145 to Line 158 serve as a representative example:
In the first paragraph (Line 145), the rationale that random sampling ensures independence is unnecessarily elaborated for a specialist audience. The boundary-handling procedure is also described in excessive detail, including a bullet-like enumeration within a sentence and an explicit justification for allowing tangency. Furthermore, the explanation that larger scan areas lead to more frequent boundary rejections and that areas larger than the interpretation boundary cannot be placed is largely self-explanatory. This paragraph could be substantially shortened without any loss of methodological rigor.
In the second paragraph (Line 154), the comparison with grid sampling is introduced via a cumbersome, literature-review-style sentence, when a simple "Compared to grid sampling" would suffice. The subsequent justification for equal sample sizes, while correct, expands into a general discussion of statistical power and the risk of misleading comparisons, complete with a citation. A single sentence stating that equal sample sizes were maintained to ensure comparability of statistical tests would convey the same message more efficiently.
I recommend a thorough revision of these and similarly verbose passages throughout the manuscript. The focus should be on reporting what was done and why in the briefest possible terms that a domain expert would require, eliminating over-justification and redundant explanation.
IV. Miscellaneous issues
1) It is not necessary to include P21 in the title, nor in the abstract.
2) Fig. 11 could be plotted in a single panel using narrow bars.
Citation: https://doi.org/10.5194/egusphere-2026-2409-RC1 -
RC2: 'Comment on egusphere-2026-2409', Roberto Emanuele Rizzo, 29 Jun 2026
Overall assessment
The manuscript addresses an important and practically relevant problem: how to define a Representative Elementary Area (REA) for areal fracture intensity, P21, from digital outcrop fracture maps. The attempt to define the REA as a range, bounded by lower and upper limits, rather than as a single threshold, is definitely a worthwhile contribution. The manuscript is also valuable because it combines exploratory plots, Shapiro-Wilk tests, Levene test, ANOVA, and diagnostic plots rather than relying on a single numerical criterion.However, in my opinion, the current version of the manuscript still needs some methodological refinement before it can be accepted for publication. The main issue relates directly with the statistical inference described in the manuscript. Currently, the manuscript treats repeated non-rejection of statistical tests as evidence for representativity, while several assumptions behind those tests are either not clearly justified or partly contracted in the same text. More specifically, the treatment of (datum) independence, spatial autocorrelation, overlapping scan areas, multiple testing, and finite-boundary effects requires (substantial) revision.
In my opinion, the manuscript would benefit from narrowing the central claim, as in its current form the described method does not fully demonstrate it can robustly identify a true geological lower and upper REA limit in all cases. I think the manuscript provides a diagnostic framework for identifying scale intervals over which P21 distributions appear relatively stable under the described sampling design. To me this is a more honest and more accurate claim for the work.
The manuscript’s own results show why this distinction matters. In the case of Cava Pontrelli, the authors conclude that the REA range is between 7 and 10 m radius, but there they also acknowledge that the apparent upper bound is influenced by the finite outcrop geometry, “no-data” zones, and strong overlap of large windows, rather than by a “true” geological REA limit. This is a central issue and not a minor limitation.
General Comments
(I) The strongest novelty of the proposed method is the concept of a multi-parameter REA range. However, this novelty is currently overclaimed in two ways.First, the manuscript presents the workflow as a method for defining lower and upper REA limits, but in at least one case study (i.e., Cava Pontrelli) the upper bound is not a geological upper REA limit. Second, the method is presented as more informative than previous or simpler criteria, but it falls short in fully demonstrating the advantages. The manuscript often states that existing approaches are less informative and do not define upper limits, whereas the proposed method is able to detect upper limits. This statement should be qualified because in the case of Cava Pontrelli – as stated by the authors – the upper bound is not a true upper REA limit.
(II) The manuscript currently states, or implies, that random sampling with replacement ensures independence among P21 measurements; however, this is hard to defend. Randomly selecting scan-window centres does not remove spatial autocorrelation in the fracture network, and overlapping scan windows will necessarily share fracture traces. Therefore, the resulting P21 values are likely to be non-independent, particularly at larger scan radii. This directly affects the validity of the Shapiro–Wilk, Levene, and ANOVA framework, because ordinary forms of these tests assume independent observations or residuals.I would recommend the authors to frame this issue using standard geostatical methods. Spatial geostatical data are not generally independent, and spatial continuity is commonly quantified using variograms, correlograms, and related geostatistical tools. A useful reference point would be Pyrcz and Deutsch (2014) which provides a detailed background on spatial continuity, scale, uncertainty, and model checking. For specific resampling issue, I think that spatial bootstrap or block-resampling – see for example Lahiri et al. (2006) and Pardo-Igúzquiza et al. (2012) – can help addressing the issue of non-independence.
A potential approach could be to estimate the autocorrelation range of P21 for each scan radius, for example using an empirical variogram or correlogram of P21 values as a function of distance between scan-window centres (possibly repeating the REA analysis using one or more spatially constrained approaches: non-overlapping windows, spatial blocks larger that the estimated autocorrelation range, or block permutation procedures. This would reduce pseudo-replication and provide stronger uncertainty estimates for the REA bounds. At minimum, the authors should report:
- The proportion of overlap among sampled scan windows as a function or radius;
- The effective number of independent windows at each radius;
- The spatial autocorrelation of P21 values, for example using correlograms, variograms, or distance-based covariance analysis;
- A sensitivity test comparing the current random-with-replacement design with a non-overlapping or spatially thinned sampling design.
(III) The manuscript repeatedly uses the term “acceptance” of Shapiro–Wilk, Levene, and ANOVA tests. This is rather confusing. A non-significant p-value only indicates that the test did not reject the null hypothesis at the chosen significance level and sample size; it does not prove that the null hypothesis is true. This is a known limitation of null-hypothesis significance testing: large p-values should not be interpreted as evidence in favour of the null hypothesis, and "absence of evidence is not evidence of absence" (Altman and Bland, 1995; Wasserstein and Lazar, 2016; Greenland et al., 2016).This issue is particularly important because the method uses repeated tests across many radii, many moving windows, and 100 realizations. The final REA range is selected after inspecting patterns of test non-rejection. This is acceptable as a diagnostic or exploratory procedure, but it should not be presented as formal proof of representativity.
For example, the REA range for Cava Pontrelli is selected because Shapiro–Wilk, Levene, ANOVA, diagnostic plots, and exploratory plots converge on a 7–10 m interval. This is plausible, but the inference should be described as “the most supported stable interval under the adopted sampling/test framework,” not as a statistically proven REA. Similarly, for Lilstock, the manuscript reports ANOVA acceptance for two groups: 0.211–0.27 m and 0.241–0.3 m. It is not clear whether the final REA is the union of these intervals, their overlap, or an interpreted continuous range from 0.211 to 0.3 m.
I would suggest replacing “accepted” with “not rejected” throughout the manuscript, unless the authors can implement a formal equivalence testing.
(IV) The manuscript uses an acceptance ratio around 80% as a criterion for plausibility of normality and other test outcomes. This threshold is central to the method, yet its justification is not sufficiently developed. The authors correctly note that sample size affects the power of normality tests, and that large samples may reject approximate normality even when the distribution is visually reasonable. However, this does not justify the specific 80% threshold. The final REA interval could be sensitive to whether the threshold is 70%, 80%, or 90%.I would suggest adding a sensitivity analysis showing how the inferred REA bounds change under different thresholds, e.g., 70%, 80%, and 90%. Alternatively, avoid a hard threshold and report probability bands or uncertainty intervals for the REA bounds. The current method appears overly dependent on a user-defined cut-off.
(V) The workflow presented in the manuscript involves many Shapiro–Wilk tests, many Levene tests, several group sizes, multiple moving windows, and repeated ANOVA tests. The manuscript then selects intervals based on the resulting patterns. This creates a multiple-testing and post hoc selection problem.
I am not arguing that repeated testing should be avoided, in my opinion, the issue is that the manuscript presents the selected REA interval as if it were a formal outcome, but the decision process is partly exploratory. Without reporting uncertainty in the selected bounds, the reader cannot evaluate the robustness of the REA range.
I would suggest providing a clear “decision tree" before presenting the results. The decision rule should specify:
- how candidate radius intervals are selected;
- how Shapiro–Wilk and Levene outcomes are combined;
- how diagnostic plots are used;
- whether ANOVA is applied only after assumptions are met;
- how borderline radii are treated;
- how uncertainty in lower and upper bounds is reported.
(VI) The manuscript assumes that within the REA range, P21 values should be approximately normally distributed. This can be a very interesting finding, but, as currently presented, it is not sufficiently justified. P21 is non-negative, may be zero-inflated at small radii, may be skewed by sparse fractures, and may become artificially compressed at large radii due to boundary constraints.
The manuscript sometimes discusses normality of ANOVA residuals and sometimes normality of P21 values within each radius. These are not identical. ANOVA assumes normally distributed residuals, not necessarily raw P21 values in every group. The methodological description should be clearer about exactly what is being tested.
The Lilstock case is especially revealing. In the initial radius discretisation, Shapiro–Wilk normality is never fully achieved, with the highest acceptance ratio around 70%, yet probability plots are interpreted as showing approximate symmetry. This suggests that the method is not purely test-based but partly relies on visual judgment. That is acceptable, but the manuscript should state this explicitly.
In my opinion, the manuscript would benefit from clarifying if the Shapiro–Wilk test is applied to raw P21 values, group residuals, or model residuals. Then justify why approximate normality should indicate REA behaviour.
(VII) Possibly, the manuscript’s most important conceptual weakness is the insufficient distinction between: (1) the lower REA limit; (2) a true upper REA limit caused by geological heterogeneity at larger scales; (3) a finite-outcrop or sampling-geometry limit caused by boundary effects, no-data zones, and window overlap.
I think that this distinction is essential. In the case of Cava Pontrelli, the authors explicitly state that beyond approximately 14 m radius, the P21 distribution converges toward a single value because the largest scan areas are constrained by boundary geometry and overlap strongly. However, in the Discussion it is stated that the “upper limit” for Cava Pontrelli is instead a representativity limit and possibly a lower bound to the true upper limit. This might suggest to the readers that this is not a true geological upper REA limit.
I believe that for the sake of clarity, this limitation should be stated more clearly in the Abstract, Methods, Results, and Conclusions. The conclusion currently says that the proposed REA range is relevant for reservoir modelling because it defines a range of cell-size resolutions from smallest representative cell to largest low-cost cell. That claim is too strong unless the upper bound is demonstrably geological rather than imposed by the dataset.
I would suggest revising the terminology throughout the manuscript and using “dataset-imposed maximum representative scan radius” or “finite-outcrop representativity limit” for the upper bound in Cava Pontrelli. The authors should reserve “upper REA limit” for cases where the data show renewed heterogeneity at larger scales that is not an artefact of outcrop geometry or overlap.
(VIII) The two natural case studies are useful and provide a valuable demonstration that the proposed workflow can be applied to real fracture networks with contrasting characteristics. However, because the true REA range is not independently known in either case, these examples are better described as demonstrations of applicability rather than full validation of the method. The inferred REA ranges appear plausible, but their accuracy cannot be independently assessed from the natural datasets alone.
The manuscript would therefore be strengthened by including a synthetic or semi-synthetic benchmark in which the underlying spatial structure and expected scale behaviour are controlled. For example, the authors could test the workflow on:
- homogeneous Poisson fracture networks;
- clustered fracture networks;
- anisotropic single-set networks;
- multi-set networks with known spacing and orientation distributions.
Such tests would help evaluate whether the workflow consistently identifies stable intervals, detects imposed upper-scale heterogeneity where present, and distinguishes geological upper limits from finite-domain or sampling-geometry artefacts.
Consider adding a synthetic benchmark section, or alternatively moderate the validation language throughout the manuscript. The current natural case studies convincingly demonstrate practical applicability, but they do not by themselves provide a controlled validation of the inferred REA bounds.
(VIII) In a recent paper, Zwarts and Lesueur (2024) propose a “convergence cone” framework in which the scatter of an effective property decreases systematically with increasing sample size. Although their application concerns permeability REV in digital rock physics, the underlying approach might be relevant to the present study because P21 distributions in 2D fracture maps also show scale-dependent changes in mean, variance, and distribution shape.
This type of convergence-envelope framework could be useful for 2D fracture networks because it provides a way to model the scale-dependent reduction in variability explicitly, rather than relying only on a sequence of hypothesis tests across scan radii. For example, the authors could explore whether the spread of P21 values across circular scan windows follows a predictable convergence trend as scan area increases. If so, fitting an envelope to the P21 convergence behaviour could provide complementary estimates of uncertainty around the inferred REA bounds.
I am not suggesting to use the “convergence cone” framework for replacing the proposed Shapiro–Wilk, Levene, and ANOVA workflow. Instead, it could be used as an additional diagnostic layer. The current workflow identifies candidate intervals where P21 distributions appear stable in shape, variance, and mean. A convergence-envelope analysis could then test whether this stability is part of a broader, predictable scale-dependent convergence pattern. This would be particularly helpful for distinguishing genuine REA behaviour from finite-domain effects, especially where large scan windows become constrained by outcrop boundaries or no-data zones.
The framework of Zwarts and Lesueur may also be useful because it highlights two issues that are directly relevant to 2D fracture-network sampling: uncertainty in the estimated representative scale and the importance of independent or minimally overlapping subsamples.
I would suggest considering adding a short discussion of whether convergence-envelope modelling could be adapted to P21-based REA estimation in 2D fracture networks. This discussion could address whether the decrease in P21 variance with increasing scan-window size can be modelled explicitly, whether such a model could provide uncertainty bounds on the inferred REA limits, and how independent or minimally overlapping sampling windows would be required for such an approach.
Specific Comments
Figures – The figures are very clear, but I feel that several could better guide the reader through the REA selection process.
- For Cava Pontrelli, please consider adding shaded bands or vertical markers for the inferred 7–10 m REA interval in Figs. 4–7, 9, and 10. It would also help to mark the larger-radius range where boundary effects and scan-window overlap become important, especially in Fig. 3 and, where relevant, Figs. 4–6 and 9.
- For Lilstock, the link between the initial coarse sampling and the refined small-radius analysis should be clearer. In Figs. 11–13 and Figs. 17–19, consider adding annotations or shaded intervals showing why additional radii between 0.005 m and 0.31 m were introduced and how the final candidate REA interval was selected.
- Tables 1 and 2 would be more informative if they included p-value summaries, effect sizes, or uncertainty ranges for the inferred REA bounds, rather than only acceptance/rejection counts.
Overall, the figures contain much of the necessary evidence, but annotations showing the selected REA intervals, candidate bounds, and boundary-controlled regions would make the decision logic easier to follow.
Language, clarity, and readability – The manuscript is generally well written, but it might benefit from some language editing. There are a few repeated grammatical inaccuracies, inconsistent terminology, and statistical wording that is too strong.
- Use “radii” rather than “radiuses” throughout.
- Use either British or American spelling consistently: “behaviour” and “behavior” both appear.
- Avoid repeated “However” at the beginning of consecutive sentences.
- Several figure captions are too long and they might benefit from some shortening.
- Line 23: “The Geometrical, mechanical and hydraulic properties” should be “The geometrical, mechanical, and hydraulic properties.”
- Lines 27–29: “REV/REA … represent” and “REV/REA mark” should be revised depending on whether REV/REA is treated as singular or plural. For example, use either “The REV/REA concept represents/marks…” or “REVs/REAs represent/mark…”
- Line 132: “needed collect P21 values” should be “needed to collect P21 values.”
- Line 154: “With regard to grid sampling strategy” should be revised to “Compared with a grid sampling strategy.
- Line 183 / Fig. 3 caption: “One hundred scan area are randomly distributed” should be “One hundred scan areas are randomly distributed.”
- Line 222: “the scan-area radiuses” should be “the scan-area radii.” More generally, replace “radiuses” with “radii” throughout. Similar occurrences appear around lines 414 and 546.
- Line 313: “the best performing scan area radius are 8m and 9m” should be “the best-performing scan-area radii are 8 m and 9 m.”
- Lines 277–290: Replace “acceptance,” “accepted group,” and “ANOVA test is accepted” with statistically more precise language such as “non-rejection,” “group for which the null hypothesis was not rejected,” or “the ANOVA test did not reject the null hypothesis.”
- Line 339 and Tables 1–2: Similarly, “accepted” should be replaced with “not rejected” or “null hypothesis not rejected.” This is especially important in the ANOVA result tables.
- Lines 565–567: “a significant clue to a situation” should be revised to “an indication of a situation” or “evidence for a situation.” In the same sentence, “spatial autocorrelation becomes too much important” should be “spatial autocorrelation becomes increasingly important.”
- Line 610: “a more informed and approach” should be “a more informed approach.”
- Line 623: “since scan areas are circular windows to avoid orientation bias. larger radii…” should be corrected by replacing the period with a comma or restructuring the sentence.
- Line 674: “We also warmly thanks prof. Sophie Viseur” should be “We also warmly thank Prof. Sophie Viseur.”
References:- Pyrcz, M. J., and Deutsch, C. V. (2014). Geostatistical Reservoir Modeling (2nd ed.). Oxford University Press.
- Lahiri, S. N., & Zhu, J. (2006). Resampling methods for spatial regression models under a class of stochastic designs. Annals of Statistics, 34, 1774–1813.
- Pardo-Igúzquiza, E., & Olea, R. A. (2012). VARBOOT: a spatial bootstrap program for semivariogram uncertainty assessment. Computers & Geosciences, 41, 188–198.
- Altman, D. G., & Bland, J. M. (1995). Absence of evidence is not evidence of absence. BMJ, 311, 485;
- Wasserstein, R. L., & Lazar, N. A. (2016). The ASA statement on p-values: context, process, and purpose. The American Statistician, 70(2), 129–133;
- Greenland, S., Senn, S. J., Rothman, K. J., Carlin, J. B., Poole, C., Goodman, S. N., & Altman, D. G. (2016). Statistical tests, P values, confidence intervals, and power: a guide to misinterpretations. European Journal of Epidemiology, 31, 337–350.
- Zwarts, S., & Lesueur, M. (2024). Predicting the Representative Elementary Volume by determining the evolution law of the convergence cone. Geomechanics for Energy and the Environment, 40, 100594.
Citation: https://doi.org/10.5194/egusphere-2026-2409-RC2 -
EC1: 'Comment on egusphere-2026-2409', Christoph Schrank, 15 Jul 2026
Dear Dr Casiraghi, dear authors,
We received two comprehensive reviews of your manuscript “Statistical Assessment of the Representative Elementary Area for Areal Fracture Intensity (P21) in Digital Outcrop Models”. I would like to thank both reviewers cordially for their thorough work.
While the reviewers agree that your work is suitable for Solid Earth in principle, both recommend significant major revisions before publication can be considered. I agree with their careful and very helpful assessments. I encourage you to implement the reviewers’ technical and editorial suggestions and to submit a significantly revised version of the work.
Many thanks for sharing your work with Solid Earth.
With the best wishes from Brisbane,
Chris Schrank
Citation: https://doi.org/10.5194/egusphere-2026-2409-EC1
Data sets
Pontrelli quarry dataset Stefano Casiraghi and Andrea Bistacchi https://doi.org/10.5281/zenodo.17483588
Model code and software
FracArea Andrea Bistacchi and Stefano Casiraghi https://doi.org/10.5281/zenodo.19735298
Viewed
| HTML | XML | Total | BibTeX | EndNote | |
|---|---|---|---|---|---|
| 245 | 65 | 21 | 331 | 15 | 16 |
- HTML: 245
- PDF: 65
- XML: 21
- Total: 331
- BibTeX: 15
- EndNote: 16
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
Publisher’s note: the content of this comment was removed on 12 June 2026 since the comment was posted by mistake.