Preprints
https://doi.org/10.5194/egusphere-2026-4074
https://doi.org/10.5194/egusphere-2026-4074
20 Aug 2026
 | 20 Aug 2026
Status: this preprint is open for discussion and under review for Earth Observation (EO).

Bayesian evaluation of deep learning architectures and sensor modalities for remote sensing-based driftwood segmentation in the Mackenzie Delta, Arctic Canada

Carl Stadie, Ingmar Nitze, Begüm Demir, Martin Stefan Brandt, and Guido Grosse

Abstract. Automated mapping of driftwood deposits along Arctic coastlines is a challenging task due to spectral ambiguity, strong depositional heterogeneity, and limited training data. Systematic comparisons of sensor modality and deep learning architecture choices remain absent for this application, leaving practitioners without evidence-based guidance on which combination to deploy. Here we present a statistical evaluation of three architectures, U-Net, Swin-U-Net, and the TerraMind foundation model, across aerial (15 cm), PlanetScope (3 m), and Sentinel-2 (10 m) imagery acquired over ten target areas in the Mackenzie Delta, Arctic Canada. Each combination was trained ten times and evaluated within a Bayesian hierarchical framework to account for run-to-run variability inherent to stochastic training. Choice of the sensor is the dominant performance driver, having an effect approximately four times larger than the choice of architecture in Intersection over Union and a total sensor spread of 0.40 IoU between aerial and Sentinel-2 imagery. Architecture choice is of limited practical consequence at sub-metre and intermediate resolution, but becomes a first-order concern when constrained to coarse imagery: U-Net performs poorly on Sentinel-2 with a posterior mean IoU of 0.127, while transformer-based architectures show a more gradual performance decline. Swin-U-Net paired with PlanetScope imagery is the most competitive accessible alternative to aerial acquisition, with a 94 % posterior probability of practical equivalence to the top-ranked configuration. Conventional single-run evaluation missed this combination as a practical alternative, typically placing it at ranks 4–5. Probabilistic multi-run evaluation is therefore a necessary condition for reliable model selection in spectrally ambiguous remote sensing benchmarks, and the framework presented here could be directly transferable to similar Arctic mapping targets.

Publisher's note: Copernicus Publications remains neutral with regard to jurisdictional claims made in the text, published maps, institutional affiliations, or any other geographical representation in this paper. While Copernicus Publications makes every effort to include appropriate place names, the final responsibility lies with the authors. Views expressed in the text are those of the authors and do not necessarily reflect the views of the publisher.
Share
Carl Stadie, Ingmar Nitze, Begüm Demir, Martin Stefan Brandt, and Guido Grosse

Status: open (until 01 Oct 2026)

Comment types: AC – author | RC – referee | CC – community | EC – editor | CEC – chief editor | : Report abuse
Carl Stadie, Ingmar Nitze, Begüm Demir, Martin Stefan Brandt, and Guido Grosse
Carl Stadie, Ingmar Nitze, Begüm Demir, Martin Stefan Brandt, and Guido Grosse
Metrics will be available soon.
Latest update: 21 Aug 2026
Download
Short summary
Driftwood piles up along Arctic coasts, storing carbon and shaping habitats, but mapping it automatically is difficult. We compared three artificial intelligence models on aerial photos and satellite images, repeating each test many times to capture natural training variation. Image resolution mattered far more than which model was used, and testing only once, as is common, gave misleading results. Our approach offers a more reliable way to choose mapping tools for similar Arctic features.
Share