the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
A protocol for the biodiversity model intercomparison project: BMIP 1.0
Abstract. Model intercomparisons are emerging as a powerful approach to improve our understanding of biodiversity and ecosystem responses to global changes, thereby supporting policy design and decision-making. However, efforts to date have focused on climate change impacts on species richness at a global scale, while regional impacts remain comparatively understudied. To address this gap, we propose a protocol for a biodiversity model intercomparison project focused on regional-to-continental scale patterns of species abundance and occurrence across space and time. This protocol covers a continuum of modeling approaches, ranging from correlative to process-explicit models. We detail the standardized input data used for model calibration –climate, land-use, and biodiversity data– and outline a common framework for model validation and performance assessment. Our protocol is structured in two phases. The first phase involves model specification, calibration, and validation using historical data, together with counterfactual experiments designed to attribute observed changes in population abundance and occurrence to different environmental drivers. In the second phase, model projections are compared under future climate and land-use scenarios. This two-phase approach enables direct comparison of model predictive performance on common benchmark data, providing an empirical basis for assessing confidence in future projections. By combining regional-scale biodiversity time series, common calibration and out-of-sample validation, and explicit representation of ecological processes, this protocol advances regional biodiversity model intercomparisons and sets the stage to reveal fundamental insights into which models, processes, and drivers determine historical and future biodiversity dynamics.
- Preprint
(780 KB) - Metadata XML
-
Supplement
(615 KB) - BibTeX
- EndNote
Status: open (until 28 Oct 2026)
- RC1: 'Comment on egusphere-2026-5033', Anonymous Referee #1, 18 Sep 2026 reply
Viewed
| HTML | XML | Total | Supplement | BibTeX | EndNote | |
|---|---|---|---|---|---|---|
| 220 | 118 | 25 | 363 | 34 | 31 | 29 |
- HTML: 220
- PDF: 118
- XML: 25
- Total: 363
- Supplement: 34
- BibTeX: 31
- EndNote: 29
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
General comments
This manuscript presents a protocol for comparing biodiversity models using common regional datasets, historical validation, counterfactual experiments, and future scenarios. Its focus on species-level occurrence and abundance, together with the inclusion of correlative and process-explicit approaches, addresses an important limitation of existing intercomparisons. The temporal holdout design, inclusion of simple reference models, and emphasis on performance across ecological contexts are valuable features. The manuscript is generally well organized and provides a useful overview of the participating approaches.
I recommend reconsideration after major revisions. Several rules governing whether comparisons will be fair and reproducible remain ambiguous, resulting in consequential inconsistencies between the main text and the supplement. These concern the independence of validation data, the definition and direction of evaluation scores, comparability of predicted and observed quantities, and interpretation of counterfactual differences. The benchmark resources and future experiment configurations also need a more precise specification. I have assessed the paper as a model-experiment description and an intercomparison protocol. A completed intercomparison is not needed to address these concerns, but a versioned implementation specification and a compact worked example would make the proposed procedures assessable and usable by independent teams.
Specific comments
1 Protect the independence of the historical evaluation
Main text lines 273-279 and 314-319; Supplement S1, lines 178-183 and 297-324.
The commitment to withhold the final 30% of years is a strength, but the supplement does not consistently preserve this separation. The MSM is explicitly defined using the past 10 years of the validation dataset (S1, lines 298-300). If implemented literally, this gives the reference model access to held-out observations. Please define this baseline using training data only, or specify a separate rolling-origin experiment in which every model receives the same information at each forecast origin. The DOM description also permits selection on validation data followed by refitting to the whole dataset. Please clarify that any such selection occurs within the first 70% of years and distinguish it from any later refitting for Phase B. The same restriction should explicitly cover hyperparameter tuning, suitability-to-abundance conversion, binary thresholds, initialization, and supplementary occurrence records. State the dates used to filter GBIF records. These are inconsistencies or ambiguities in the written protocol; they do not establish that leakage has already occurred in an implementation.
2 Define common prediction targets and their relationship to observations
Main text lines 198-213 and 309-319; Table 1; Supplement S1, lines 69-99, 118-123, and 162-184.
The required outputs combine quantities that are not automatically interchangeable: habitat suitability, occurrence probability, survey detection, abundance indices, and population abundance. S1 already addresses imperfect detection for DOMs and observation-dependent scaling for metaRange, but a common evaluation rule across model families is still needed. For each benchmark, define the response variable, units, sampling effort, spatial support, and handling of repeated visits, non-detections, and missing surveys. Explain how each model's ecological state is mapped to that observed response. Otherwise, differences in observation assumptions could be mistaken for differences in ecological predictive ability.
The mandatory conversion from suitability or binary occurrence to abundance also needs an explicit procedure and justification. Such a conversion adds a fitted component and uncertainty. Specify its training-only calibration and how its contribution is reported. The exposure model requires particular clarification: S1 describes exposure as a warning indicator without assumed demographic consequences, whereas Table 1 and the common output requirements imply suitability or abundance predictions. Define a defensible conversion or a separate evaluation task for outputs that do not estimate the same quantity.
3 Reconsider the universal spatial resampling rule
Main text lines 291-308 and 337-343.
Bilinear interpolation from 10 km to 1 km changes the grid spacing but does not establish ecological information at 1 km. Its suitability depends on the output: interpolating population totals generally does not conserve abundance, and occupancy probabilities are defined over a particular spatial support. Please distinguish abundance per cell, density, and occurrence probability, and prescribe an appropriate transformation for each. Evaluation at a common coarser support, or a documented observation operator, would provide a useful primary comparison; finer outputs could be assessed separately where justified. Also specify how prediction grids are aligned with surveys and how the 100 km evaluation buffer is constructed, including whether held-out occurrence locations influence it. Explain the ecological and computational rationale for this buffer and distinguish the evaluation domain from any larger simulation domain required for dispersal.
4 Correct the score definitions before using them for comparisons
Main text Table 2 and lines 389-394; Supplement S2, lines 451-503, 505-524, and 533-544.
The CRPS equations in S2, lines 484 and 487, use the negative of the usual non-negative loss, while Table 2 states that lower values are better. The negative of this loss can be used as a reward when higher values are better, but combining it with the stated direction reverses the comparison. Please adopt a single convention consistently across the equations, tables, software, and interpretation. CRPS and the logarithmic score assess probabilistic forecast quality, whereas the C-index assesses ordering; the latter is not interchangeable with a proper probabilistic score for defining confidence tiers.
Other discrepancies also affect implementation. The moving-window calculation compares a model's error in successive windows, rather than its error against a reference forecast evaluated in the same window. Rename this as a change in error or define a genuine reference-based skill score. Resolve the reversed predictor/response wording for the calibration regression (S2, lines 452-454), add the factor of 100 to Percent Bias or rename it relative bias, and specify the pseudo-R-squared measure promised by the section heading. A short set of worked scoring examples would help verify these definitions.
5 Make uncertainty outputs comparable and define confidence tiers
Main text lines 320-324, Table 2, lines 389-394 and 439-443; Supplement S2, lines 464-474.
Posterior parameter uncertainty, uncertainty in a mean prediction, stochastic population variability, and predictive uncertainty for a future observation are different quantities. Please specify which components each submission must include and how they will be evaluated. The statement that missing uncertainty is treated as no uncertainty should be revised. If a point forecast is evaluated as a degenerate predictive distribution, identify this explicitly; it does not demonstrate the absence of uncertainty in the ecological system.
There is also a mismatch in coverage: Table 2 refers to observations falling within prediction intervals, whereas S2 defines confidence-interval coverage of an unknown parameter. Provide the empirical predictive-coverage calculation appropriate to the chosen observation target, its nominal levels, and a measure of interval width or an appropriate interval score. Finally, define the rules for assigning confidence tiers before examining the held-out results. Historical predictive skill can inform confidence, but the limits of using it to judge performance under novel future conditions should be stated.
6 Specify the attribution estimand and factorial experiments
Main text lines 353-398, especially 363-370.
Please express the attribution calculation mathematically and list the required factual and counterfactual runs. If prediction errors are signed residuals against the same observations, subtracting the counterfactual residual from the factual residual simply yields the counterfactual prediction minus the factual prediction. This is a model-based scenario contrast, but the subtraction alone does not control or quantify predictive uncertainty. If absolute or squared errors are intended, the resulting comparison measures relative predictive fit and require a different interpretation. Explicitly define the estimand, sign convention, temporal aggregation, and uncertainty propagation.
To distinguish climate and land-use contributions, specify simulations with both drivers varying, each varying separately, and both held at their counterfactual baselines, or justify an alternative design. Explain how interactions are retained or allocated. State the initial conditions and calibration used across paired runs, the exact detrending/reference periods, and treatment of additional predictors. Clarify how the 1901-based detrending and dataset-specific starting dates are related. The winner/loser categories also need a stated reference for the temporal trend and a rule for near-zero or uncertain effects. These additions would support a transparent attribution analysis while making clear its dependence on model assumptions and omitted drivers.
7 Provide a reproducible benchmark and implementation specification
Main text Sections 3.1-3.4, lines 288-297 and 519-523; Supplement S1.
The four benchmark datasets and 72 selected species are central to the protocol, but essential information is deferred to a manuscript in preparation. Please provide a benchmark table containing species identities, source versions, observation periods, sampling coverage, response units, and exact training/evaluation years. Describe how trend and range-size categories are calculated and which period they use. Selection based on full-period trends would condition the benchmark on those trajectories and should be disclosed when assessing generality.
The availability statement should identify the released BMIP resources and distinguish them from planned resources. Provide versioned inputs, accessible metadata, and preprocessing scripts or an equivalently precise implementation specification for grid alignment, aggregation, derived climate variables, and output formats. Explain the access and release arrangements for centrally held evaluation data. A small worked example using a benchmark subset or clearly labeled synthetic data, with representative outputs, would demonstrate that the harmonization and scoring procedures work together. This is consistent with GMD's expectations for model experiment descriptions (https://www.geoscientific-model-development.net/about/manuscript_types.html) and need not await the full project.
8 Resolve forcing provenance and temporal alignment
Main text lines 228-251, 363-367, and 519-523; reference list lines 1046-1048.
The HILDA+ specification contains two verifiable inconsistencies. The official v2.0 record (https://doi.pangaea.de/10.1594/PANGAEA.974335) identifies a 2020 ESA WorldCover base map, whereas line 247 gives 2000. The availability statement links to the earlier release, 10.1594/PANGAEA.921846, while the reference list correctly gives v2.0 as 10.1594/PANGAEA.974335. Please confirm the intended product and harmonize the description and links. For CHELSA-daily, identify the exact release or download snapshot and the complete period available for each variable. Explain how differences in the climate, land-use, and biological observation periods are handled, including any pre-1960 or post-2020 land-use assumptions needed by the selected experiments. These decisions matter for both hindcasts and counterfactual baselines.
9 Distinguish model comparisons from tests of individual processes
Main text lines 161-166, 215-225, 288-297, and 451-480; Supplement S1.
The protocol aims to assess the influence of ecological processes, but models can also differ in trait information, imputation, calibration, resolution, and the suitability of the models they support. I appreciate the requirement to share additional environmental predictors among teams (lines 256-262). Please extend the reporting framework to identify these other differences and shared components, and explain which contrasts support process-specific conclusions. Where practical, matched configurations or a shared suitability input could be used to test selected mechanisms. Where this is infeasible, conclusions should concern the performance of complete modeling approaches under their stated information requirements. Since computational efficiency is part of the motivation, reporting comparable runtime and resource requirements would also help interpret any gains in complexity.
10 Freeze the core scenario and comparison rules
Main text lines 325-343 and 402-443.
The requirement to use a broad range of scenarios and to move to CMIP7 as data become available leaves the core experiment open to change during participation. Please define a minimum common scenario matrix for this version, including forcing products, climate-model members, paired land-use pathways, baseline and projection periods, and stochastic replication. Treat later forcing generations as documented extensions or protocol versions. The manuscript correctly recognizes the need for consistent climate and land-use scenarios; their concrete pairing still needs to be specified.
Similarly, define a common set of primary metrics, reference models, and aggregation rules, with optional diagnostics clearly identified. Explain how unequal sampling, missing years, differences in abundance scale, and dependence among sites and years affect comparisons. Distinguish averaging predictions before scoring from averaging scores. Temporal holdout evaluates prediction in later years; spatially summarized scores alone do not establish transfer to unsampled regions. Clarify that scope or add a separate spatial-transfer experiment if such claims are intended.
Technical corrections
1. Main text line 283: replace the reference to Section 4.3 with Section 4.1.3.
2. Main text line 310: add the missing space before the citation in "intercomparison(Zurell et al., 2026)".
3. Main text line 521: remove the stray "l" after "CHELSA V2.1".
4. Supplement S1, line 4: use "BMIP 1.0" consistently. At line 67, replace "Apendix B" with the intended reference to Section S2.
5. Supplement S2, lines 482 and 488: use "cumulative distribution function" for F and Phi; identify phi as the probability density function. Use the indexed quantities consistently in the CRPS equations.
6. Supplement S2, line 572: correct "pair si comparable" to "pair is comparable".
7. Remove duplicated DOI prefixes in the main references (lines 1035 and 1075) and the supplement (line 744).