Preprints
https://doi.org/10.5194/egusphere-2026-3036
https://doi.org/10.5194/egusphere-2026-3036
28 Jul 2026
 | 28 Jul 2026
Status: this preprint is open for discussion and under review for Atmospheric Measurement Techniques (AMT).

A Self-Supervised and Distillation-Guided Framework for Robust Cloud-Layer Identification in Multiwavelength Polarization Raman Lidar Networks

Weijie Zou, Zhenping Yin, Tingxin Sun, Zhichao Bu, Yaru Dai, Haofei Wang, Longlong Wang, Yun He, Shuangliang Li, Jiarui Cheng, Xuan Wang, Siwei Li, and Detlef Müller

Abstract. Cloud-layer identification is a key prerequisite for automated ground-based lidar processing, but remains challenging in Raman lidar networks because of incomplete near-range overlap, weak high-level cloud echoes, day–night signal-to-noise ratio (SNR) contrasts, energy fluctuations, and aerosol–cloud ambiguity. Leveraging the China Aerosol Raman Lidar NETwork (CARLNET), we propose a multi-channel cloud-layer identification framework that operates on 355/532/1064 nm elastic signals and 355/532 nm volume depolarization ratio, reducing dependence on absolute calibration and retrieved optical products. The framework combines an anisotropic encoder–decoder architecture with a three-stage training strategy, including self-supervised pretraining, traditional-algorithm-guided probabilistic distillation, and fine-tuning with limited expert refinement. Strict cross-site and cross-time evaluation on held-out sites shows that, without using any test-site labels, the framework achieves an F1 score of 0.9371 and reduces the mean absolute errors of cloud-base and cloud-top heights to 113 m and 213 m, respectively, outperforming training without pretraining by approximately 50 m and 70 m and the conventional baseline for cloud-top height by about 400 m. A labeled-data-size ablation further shows improved label efficiency, with near-plateau performance reached at about 40 labeled days. Consistency checks against co-located radiosonde moist-layer indications support the realism of identified high-level weak-echo clouds and the robustness of cross-site deployment. These results demonstrate a transferable and calibration-decoupled cloud-layer identification technique for network-scale Raman lidar processing.

Publisher's note: Copernicus Publications remains neutral with regard to jurisdictional claims made in the text, published maps, institutional affiliations, or any other geographical representation in this paper. While Copernicus Publications makes every effort to include appropriate place names, the final responsibility lies with the authors. Views expressed in the text are those of the authors and do not necessarily reflect the views of the publisher.
Share
Weijie Zou, Zhenping Yin, Tingxin Sun, Zhichao Bu, Yaru Dai, Haofei Wang, Longlong Wang, Yun He, Shuangliang Li, Jiarui Cheng, Xuan Wang, Siwei Li, and Detlef Müller

Status: open (until 02 Sep 2026)

Comment types: AC – author | RC – referee | CC – community | EC – editor | CEC – chief editor | : Report abuse
Weijie Zou, Zhenping Yin, Tingxin Sun, Zhichao Bu, Yaru Dai, Haofei Wang, Longlong Wang, Yun He, Shuangliang Li, Jiarui Cheng, Xuan Wang, Siwei Li, and Detlef Müller
Weijie Zou, Zhenping Yin, Tingxin Sun, Zhichao Bu, Yaru Dai, Haofei Wang, Longlong Wang, Yun He, Shuangliang Li, Jiarui Cheng, Xuan Wang, Siwei Li, and Detlef Müller
Metrics will be available soon.
Latest update: 28 Jul 2026
Download
Short summary
Cloud layers can be identified from ground-based laser observations, but stable and consistent results across different instruments and sites remain challenging. This study develops a transferable artificial intelligence method using multi-wavelength lidar signals. It improves cloud boundary detection, works with limited expert labels, and is tested at independent sites with weather balloon checks. It supports reliable automated cloud and aerosol processing.
Share