the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
A Self-Supervised and Distillation-Guided Framework for Robust Cloud-Layer Identification in Multiwavelength Polarization Raman Lidar Networks
Abstract. Cloud-layer identification is a key prerequisite for automated ground-based lidar processing, but remains challenging in Raman lidar networks because of incomplete near-range overlap, weak high-level cloud echoes, day–night signal-to-noise ratio (SNR) contrasts, energy fluctuations, and aerosol–cloud ambiguity. Leveraging the China Aerosol Raman Lidar NETwork (CARLNET), we propose a multi-channel cloud-layer identification framework that operates on 355/532/1064 nm elastic signals and 355/532 nm volume depolarization ratio, reducing dependence on absolute calibration and retrieved optical products. The framework combines an anisotropic encoder–decoder architecture with a three-stage training strategy, including self-supervised pretraining, traditional-algorithm-guided probabilistic distillation, and fine-tuning with limited expert refinement. Strict cross-site and cross-time evaluation on held-out sites shows that, without using any test-site labels, the framework achieves an F1 score of 0.9371 and reduces the mean absolute errors of cloud-base and cloud-top heights to 113 m and 213 m, respectively, outperforming training without pretraining by approximately 50 m and 70 m and the conventional baseline for cloud-top height by about 400 m. A labeled-data-size ablation further shows improved label efficiency, with near-plateau performance reached at about 40 labeled days. Consistency checks against co-located radiosonde moist-layer indications support the realism of identified high-level weak-echo clouds and the robustness of cross-site deployment. These results demonstrate a transferable and calibration-decoupled cloud-layer identification technique for network-scale Raman lidar processing.
- Preprint
(2733 KB) - Metadata XML
- BibTeX
- EndNote
Status: final response (author comments only)
-
RC1: 'Review on egusphere-2026-3036', Anonymous Referee #1, 14 Sep 2026
- AC1: 'Reply on RC1', Weijie Zou, 20 Sep 2026
-
RC2: 'Comment on egusphere-2026-3036', Anonymous Referee #2, 14 Sep 2026
General Comments:
This manuscript presents a novel approach for cloud layer identification using a Raman lidar network system named CARLNET. The method utilizes multi-channel wavelengths, as well as an anisotropic encoder–decoder architecture, self-supervised pretraining, and traditional algorithm guidance. This approach aligns well with the scope of the AMT journal.
Overall, this paper introduces a new automated method for cloud layer identification, which is acceptable for publication in AMT. However, some clarifications in the representation, as well as minor technical corrections, would enhance the quality of the manuscript, as mentioned in the specific comments and technical corrections section.
Specific comments:1. In section 2.5, as “Radiosonde Observations and Cloud Identification,” the authors utilized a relative humidity thresholding method to define the boundaries of clouds by radiosondes. However, the authors also mentioned that mapping a single-column profile at a specific time is not feasible due to wind drift. Would it be possible to include a Skew-T diagram for one or a few days to validate the results? Specifically, this would help in identifying cloud bases, corresponding not only to relative humidity but also to temperatures and pressure levels.
2. In Section 4.3, as “Independent Validation with Radiosonde Observations,” it is recommended that the results be presented more clearly and systematically. In particular, it is suggested to include a table summarizing the results for improved clarity and ease of comparison.
Technical correction:
1. The terms (e.g., Figure/Fig., Section/Sect.) should follow the Copernicus manuscript guidelines and be used consistently throughout the manuscript.
2. Lines 156-157 and 217, 220: Please use standard distances for units and mathematical signs, and ensure consistency throughout the manuscript.
3. In Table 2, it is recommended to reform the table using abbreviations for better structure.
4. In Figure 5, it is recommended to include every annotation used in the figure within the caption as well.
5. In Figure 6b, in terms of "noise misclassified as ..." is something missing after "as"?
Please also define the green dashed contour, as it is not visible on the plot.6. In A5, line 665: replace “hile” to “while”
Citation: https://doi.org/10.5194/egusphere-2026-3036-RC2 - AC2: 'Reply on RC2', Weijie Zou, 20 Sep 2026
-
RC3: 'Comment on egusphere-2026-3036', Anonymous Referee #3, 16 Sep 2026
This manuscript presents a three-stage framework combining self-supervised pre-training, VDE-based prior distillation, and limited expert-annotated fine-tuning for cloud detection in multi-wavelength polarized Raman lidar networks, with a systematic multi-site and cross-temporal evaluation on CARLNET, but it requires major revisions regarding novelty positioning, transparency of label generation and potential VDE-induced biases, justification of methodological choices, statistical robustness, and conclusions that currently exceed the supporting evidence. 本稿件提出了一个三阶段框架,结合了自监督预训练、基于VDE的先驱蒸馏以及有限的专家注释微调,用于多波长极化拉曼激光雷达网络中的云检测,并对CARLNET进行了系统性多站点和跨时段评估,但需要在新颖性定位、标签生成透明度及潜在VDE诱导偏差、方法学选择的合理性以及统计鲁棒性方面进行重大修订。 以及目前支持证据的得出。
本稿件提出了一个三阶段框架,结合了自监督预训练、基于VDE的先驱蒸馏以及有限的专家注释微调,用于多波长极化拉曼激光雷达网络中的云检测,并对CARLNET进行了系统性多站点和跨时段评估,但需要在新颖性定位、标签生成透明度及潜在VDE诱导偏差、方法学选择的合理性以及统计鲁棒性方面进行重大修订。 以及目前支持证据的得出。本稿件提出了一个三阶段框架,结合了自监督预训练、基于VDE的先驱蒸馏以及有限的专家注释微调,用于多波长极化拉曼激光雷达网络中的云检测,并对CARLNET进行了系统性多站点和跨时段评估,但需要在新颖性定位、标签生成透明度及潜在VDE诱导偏差、方法学选择的合理性以及统计鲁棒性方面进行重大修订。以及目前支持证据的得出。- AC3: 'Reply on RC3', Weijie Zou, 20 Sep 2026
Viewed
| HTML | XML | Total | BibTeX | EndNote | |
|---|---|---|---|---|---|
| 302 | 96 | 40 | 438 | 25 | 25 |
- HTML: 302
- PDF: 96
- XML: 40
- Total: 438
- BibTeX: 25
- EndNote: 25
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
The manuscript proposes a self‑supervised and AI framework for cloud‑layer identification in multi‑wavelength Raman‑polarization lidar networks. The work addresses an important gap in operational lidar processing. Therefore, the topic is generally very well suited for AMT.
By operating directly on uncalibrated elastic backscatter and calibrated volume depolarization ratio channels from the Chinese lidar network, the approach aims to reduce dependence on absolute calibration and improve robustness across heterogeneous network sites with the potential to be applied to other lidar networks as well.
The three‑stage training pipeline—self‑supervised reconstruction pretraining, VDE‑guided probabilistic distillation, and fine‑tuning with limited expert labels—is to my opinion a complete novel and very sophisticated approach. It seems solid, and the reported gains are promising. Therefore, well suited for publication in this journal for the atmospheric research community.
However, for a good understanding, the manuscript needs additional methodological description (more details) and clearer justification of certain design choices (especially the descriptions given in the Annex). Thus, while generally well written, the methodological section needs to be strengthened and therefore recommend major revisions before the paper can be accepted for publication.
General comments:
I am not sure if distilling is the correct word for what you want to say even though I understand the meaning. I guess only a native speaker could help there.
1. The self‑supervised pretraining tasks (channel‑wise masking, temporal frame prediction) are described only at a high level. Precise formulation of loss functions, masking ratios, and the balance between reconstruction and prediction objectives is needed.
2. Reported metrics include IoU, F1 score, detection rates, …. Partly the reader should be more guided for the interpretation of these metrics – description is a bit missing and should be extended.
3. Numbers used for the approach like penalties are given moistly without any justification. This should be improved. If the choice was based on empirical studies, it is ok, but should be stated.
4. Cloud top: It remains unclear for me, how cloud tops are defined. As there might be a lot of cases, where signal attenuation is strong and a cloud top height cannot be detected with lidar, I wonder how these circumstances are treated. The other way around: Under which conditions are you able to identify cloud top heights properly
5. Radiosonde comparison: It is not clear for me if radiosondes are available at all sites and have been used for evaluation or only on the selected ones. Please clarify.
6. Temporal context window: The fixed 10‑profile window: It remains unclear if it is 5 min or 50 min in total. Please clarify.
7. In the evaluation of the proposed methodology, models trained with and without pretraining are compared. It is unclear to me which parts of the 3-stage training process is included in the pretraining. Only stage one—reconstruction, or stage one and two—reconstruction and distillation. Please clarify this.
Labels of Figures partly need to be extended. As a general rule, figures including their labels should be as self-explaining as possible. E.g. in Fig 3., it is not described what the numbers in brackets mean and just by reading all text one can guess what is meant. See further specific comments below. Please change this.
The CLARNET network looks impressive: In general, it would be great if CARLNET data could be shared in future and the initiative joins activities within GAW (GALION) and/or contribute to other global activities like satellite validation.
The approach presented seems to have great potential to be checked against other lidar networks. Therefore, I would insist to share the code and training data as promised during the publication process. For a peer‑review setting, providing a DOI of the code would be key.
Specific Comments
● Abstract – The reported F1 improvement (0.9281 → 0.9371) is presented without context and statement if it is significant. For my understanding. the improvement is negligible. Please clarify.
● Figure 1: The small embedded map in the lower right is not needed and rather misleading. I would recommend to remove it… In general, the dashed lines are not explained, if needed, please do so - otherwise remove.
● Figure 2: The measurements in the panels you provide look impressive. Congratulations. Nevertheless, It would be important to know how you performed calibration (e.g. reference height) to retrieve the high resolution backscatter coefficient. (panels f- h). Please provide this information, e.g. in the respective Section line 195 ff or the Annex.
● Sec.2.2.: It is stated that “each site performs routine maintenance and periodic calibration following operational protocols”. Can you provide more information on that or give references? Furthermore, the lidar description you give seems to be valid for all 47 sites, is this correct? Please state more explicitly.
● Line 158: “and 407 Raman channels” missing “nm” after “407”, please add this.
● Test data (line 236): Why was as test data set only one month chosen, and not data distributed over a whole year?
● Sec 2.5. Radiosondes available at all sites? Please clarify.
● Figure 3 (pipeline) – The show graph has great potential but still needs some more clarification: Also, the caption must be extended. The captions describing the pooling operation (all the way to the left) needs to be rechecked as I believe there are several typos here. Firstly, both the uppermost and the middle pooling approach are stated to be applied for “Range gate > 1600”. However, the middle approach should be applied from “800 < Range gate < 1600” from the description in the appendix. Please fix this. I would also recommend to include the actual height as an additional information instead of only the range gates, i.e., 0-12km, 12-24km, and 24-30km respectively for the different pooling approaches (from the lowermost to the uppermost). Secondly, it is stated in the uppermost approach that a pooling of 16 bins is used. However, in the Appendix, a pooling of 8 bins is stated for the same region. Please fix this inconsistency.
● Lines 397-398: “Any experiment that introduces a small amount of labeled data from the test sites is explicitly treated as an adaptation ablation rather than as part of the main validation protocol.” What is meant by “adaptation ablation”? I do not get this. Please consider rephrasing this point to make it clearer for the reader.
● Lines 402 – 403: The x and y-axis of Figure 4 is introduced before the actual figure is introduced. Please move this to after you introduce the figure.
● No Pretrain: The evaluation category “no pretrain” sounds a bit strange. It could be worth renaming it to “no pretraining” to be more consistent with the second category “pretraining”.
● Table 2 – Many values are only marginally different compared to others, yet the text emphasizes the benefit of e.g., including test‑site labels. Please clarify and probably rephrase to avoid overstating this effect.
● F1: As you use the F1 score as the main metric, the interpretation of this value should be explained in the beginning. It would be also of interest to which extent (digital precision) differences are really significant. I.e., 0.928 vs. 0.937 seem to be the same for me and no significant difference, but maybe I am wrong. Please clarify in the manuscript.
● Lines 421-422: “Overall, the results suggest that pretraining not only reduces reliance on large-scale manual annotations for cloud-layer identification, but also provides a more controllable training cost for cross-site deployment.” I cannot follow this conclusion. Can you guide more with the interpretation of your results?
● Lines 426-428: You state that an early-stop strategy is used to mitigate overfitting by terminating the training when the validation performance does not improve for a number of consecutive epochs. Please state how many epochs were used as the threshold to stop training. This could be added in the Appendix.
● Lines 429-430: The categories “Test sites used in pretraining” and “Test-site labels used in training” are not category names used in table 2. Please change this to be consistent with the names used in the table.
● Table 3 – Base‑height bias is reported as –10 m for the pretrained model, which is negligible considering the vertical resolution of the lidar.
● Sec 4.3: Several results from the radiosonde evaluation are stated in this section like DL and VDE identification rates for multilayer and first layer clouds without any statistics backing them up. I would recommend explaining the setup of this experiment in more detail and adding some statistics in the form of a table to support these claims.
● Appendix:
E.g. what does it mean: gaussian noise with probability 0.3 and SD 0.5?
But also the other values are too briefly explained for full understanding and in the appendix you have enough space to do so – so no need to leave out information.
Technical comments
● Some figure sub panels (e.g. 7 & 8 d-e) are not referred to in the text.
● There are two He’s et al. 2021, references which significantly differ. Please separate by 2021a and 2021b
● The numbers given for the 3 reference stations are not explained. What do these numbers stand for, e.g. 58847 (Fuzhu)
● Line 165: I think there is a mix up. It should be the 355-parallell and the 1064 total (elastic) channel, correct?
● Line 254: “Therefore, radiosonde-based … “. Consider changing from "Therefore" to “Thus” to avoid using "Therefore” to start two consecutive sentences.
● Line 282: “... guides the Deep model” typo of “deep learning model”?
● Line 346. What is a “segmentation head”? Please try to explain in a way that it is more understandable for a non-AI expert.
● Fig. 6: Please check the label. I do not see a green dashed contours.
● 665: hile --> while?