Preprints
https://doi.org/10.5194/egusphere-2026-4514
https://doi.org/10.5194/egusphere-2026-4514
08 Oct 2026
 | 08 Oct 2026
Status: this preprint is open for discussion and under review for Geoscientific Model Development (GMD).

Lossy compression by dimension reduction – Evaluation of methods applied to 2-d meteorological fields for data-based modeling

Uwe Ehret, Jieyu Chen, Fedor Scholz, Sam Allen, and Sebastian Lerch

Abstract. Data sets in the geosciences continue to grow in size, and data compression methods can facilitate efficient data storage, transfer, and utilization. Related workflows in geoscience modelling typically involve data generation (e.g. via weather forecasting), followed by compression, transfer, decompression, extraction of local subsets, and integrations into downstream tasks such as hydrological modelling. Recently, data-driven models have been introduced into geoscientific workflows with great success, allowing input data sources and types to be used in much more flexible ways than in classical numerical models, e.g. by directly using the compressed data. In this study, we therefore address the questions on 1) how various lossy compression algorithms compare in terms of general properties such as computational effort and their ability to extract spatial subsets directly from the compressed representation, and 2) how they compare in terms of compression efficiency and reproduction error. We focus on dimension reduction methods, as the resulting low-dimensional representations require fewer input channels in downstream data-driven models, thereby reducing computational costs and the risk of overfitting. We compare several dimension reduction methods in an application to five years of hourly air temperature and rainfall data, in the form of 2-d gridded spatial fields spanning a 100,000 km² domain in central Germany. The methods we compare are: Block Averaging, Principal Component Analysis, an Autoencoder, and the Ramer-Douglas-Peucker algorithm. We measure compression by the number of distinguishable objects in the compressed representation, rather than by file size, keeping in mind the potential use of the compressed data as input to data-driven models. This approach directly relates dimension reduction/compression to the number of input channels of a data-driven model, which is typically limited to avoid overfitting. Our results indicate that the methods differ substantially in terms of training and computational overhead, preservation of field-scale statistics, and the possibility to extract spatial subsets directly from the compressed representation. All methods demonstrated very good compression efficiency, with low reconstruction errors even for 99 % compression. For the spatially smooth temperature fields, Principal Component Analysis performed best, followed by the Autoencoder. For the spatially heterogeneous rainfall fields, the Autoencoder and the Ramer-Douglas-Peucker algorithm were most effective. The choice of the best method therefore depends on the specific application. This study provides a step towards a more efficient use of geoscientific data in downstream applications, either by directly using compressed data as input to data-driven models, or by efficiently extracting spatial subsets from compressed data, and using these subsets as input to local numerical models.

Publisher's note: Copernicus Publications remains neutral with regard to jurisdictional claims made in the text, published maps, institutional affiliations, or any other geographical representation in this paper. While Copernicus Publications makes every effort to include appropriate place names, the final responsibility lies with the authors. Views expressed in the text are those of the authors and do not necessarily reflect the views of the publisher.
Share
Uwe Ehret, Jieyu Chen, Fedor Scholz, Sam Allen, and Sebastian Lerch

Status: open (until 03 Dec 2026)

Comment types: AC – author | RC – referee | CC – community | EC – editor | CEC – chief editor | : Report abuse
Uwe Ehret, Jieyu Chen, Fedor Scholz, Sam Allen, and Sebastian Lerch
Uwe Ehret, Jieyu Chen, Fedor Scholz, Sam Allen, and Sebastian Lerch
Metrics will be available soon.
Latest update: 08 Oct 2026
Download
Short summary
Data sets in the geosciences are often very large, therefore data compression methods are required. In this study, we test four such methods: Block Averaging, Principal Component Analysis, Autoencoder, and Ramer-Douglas-Peucker algorithm. "Lossy" means that the original data cannot fully be restored from the compressed data. We tested the methods on rainfall and air temperature data. All methods worked very well, and for the air temperature and rainfall, different methods were best.
Share