the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
A Hardware-Aware Deep Learning Framework for Wind Retrieval from Raw Doppler Spectra of Radar Wind Profilers
Abstract. Radar wind profilers (RWPs) traditionally rely on heuristic statistical procedures for signal processing and quality control, often suffering from inherent trade-offs between data availability and retrieval reliability. To address these limitations, this study presents an end-to-end deep learning framework that retrieves wind vectors directly from raw Doppler spectra. Validated against a three-year (2022–2024) dataset of raw spectra from 437.5-MHz and 1290-MHz profiler systems and collocated radiosonde observations, the proposed framework demonstrates robust performance, particularly when trained on the full spectral continuum without filtering ambiguous samples. Notably, we found that hybridizing deep learning with conventional post-processing is highly contingent upon hardware-specific error characteristics; it effectively mitigated transient outliers in the 1290-MHz systems but provided negligible benefits for the range-smeared, vertically coherent artifacts dominating the 437.5-MHz systems. Furthermore, the framework reasonably reproduced seasonal, diurnal, and vertical atmospheric structures without the excessive smoothing typical of conventional statistical approaches. Although constrained by a slight underestimation of extreme wind speeds due to the regression-to-the-mean effect inherent in mean-squared-error optimization, the model notably enhanced operational data availability and integrity, demonstrating the feasibility of hardware-aware deep learning as a viable alternative or complement to conventional RWP processing pipelines.
- Preprint
(2067 KB) - Metadata XML
- BibTeX
- EndNote
Status: final response (author comments only)
-
RC1: 'Comment on egusphere-2026-3168', Anonymous Referee #1, 11 Sep 2026
-
AC1: 'Reply on RC1', Kyung Hun Lee, 17 Sep 2026
We sincerely thank the reviewer for the careful evaluation of our manuscript and for the constructive comments. Our responses to each point are provided below in the order presented.
- In Module 1, KWPP identifies up to two peaks based on spectral peak magnitude, with the dominant peak assigned as the primary peak. The primary peak is then evaluated through the subsequent spectral-level QC procedures. Under clear-air conditions, if the primary peak fails the checks in Modules 3–7, the second-ranked peak is considered instead. When precipitation is identified from the vertical-beam observation in Module 2, the dominant peak is retained and Modules 3–7 are bypassed.
- KWPP does not apply a separate procedure to decompose partially overlapping signal components. If partially overlapping signals appear as a single spectral feature under the peak-detection procedure, the feature is treated as a single peak for subsequent moment estimation.
- We agree that such an overlay would be useful for visualizing the spectral regions selected by the conventional processing scheme. However, the proposed deep learning framework does not identify a spectral peak or retrieve beam-wise radial velocities. Instead, it directly maps the five-beam Doppler spectra to the horizontal wind components (u and v). Therefore, a radial velocity corresponding to the deep learning output is not uniquely defined without additional assumptions, particularly regarding the vertical velocity component. Reconstructing beam-wise radial velocities would consequently introduce information or assumptions that are not part of the proposed framework.
- “Spectral points” refers to the number of Doppler spectral bins actually stored in the raw data, whereas “FFT points” refers to the full FFT size used during signal processing.
- The RWP alternates between low- and high-mode observations, with each mode requiring approximately 5 minutes. Thus, the observation sequence is Low–High–Low–High…, and the recurrence interval for a given mode is approximately 10 minutes. Therefore, the 10-minute temporal resolution does not represent an average of two observation cycles.
- Only one low-mode RWP observation cycle corresponding to the radiosonde launch time is used for each comparison. Each low-mode cycle consists of approximately 5 minutes of measurements. As described above, the 10-minute interval refers to the recurrence interval of the low-mode observations resulting from the alternating low- and high-mode sequence, rather than to the duration of data used for each comparison.
- The application of the framework over a broader altitude range was also considered during the initial development of the study. However, as the reviewer noted, the present analysis was restricted to 0.5–2.0 km to maintain a consistent analysis range across the 437.5- and 1290-MHz systems and to account for reduced data availability at higher altitudes, particularly during winter, as well as increasing RWP–radiosonde representativeness differences associated with radiosonde drift.
- As the reviewer noted, in this study the term “consensus check” refers to a temporal continuity check applied to radial velocity. We agree that the term “consensus” alone does not explicitly indicate the time domain, and we will clarify this point in the manuscript.
- We agree that the wording “filtered out” is clearer and will adopt this revision. The “Uncertain” category is defined based on processing history rather than a specific physical cause. These regions are initially filtered out by either the low-SNR check in Module 8 or the vertical-discontinuity check in Module 9, but are subsequently restored through the temporal consensus procedure in Module 10. Because the reliability and physical origin of the restored signals cannot be determined unambiguously, they are classified as “Uncertain.” The original description referring to the coexistence of meteorological and non-meteorological echoes did not accurately convey this intended meaning. We will revise this description to more clearly reflect the uncertainty associated with signals restored through the QC procedure.
- For DL-v1, the model is trained using samples for which at least one of the five beam spectra at a given range gate is classified as “Uncertain.” DL-v2 is trained using only samples for which all five beams are classified as “Certain.” DL-v3 is trained using the combined set of “Certain” and “Uncertain” samples, while “Rejected” samples are excluded. The KWPP-QC10 flag values are not used as input features of the deep learning models; they are used only for dataset classification and sample selection. The models themselves use only the raw Doppler spectra as input.
- Modules 11–13 correspond to a vertical continuity inspection of wind speed and direction, vertical linear interpolation of u and v, and temporal consensus averaging of u and v, respectively. These procedures are described in Sect. 3.2, but we agree that their roles should also be clarified when they are first introduced in the Methods section.
- As suggested, we will spell out RMSE as “root mean square error” and MBE as “mean bias error” when the abbreviations are first introduced in the manuscript.
- KWPP-QC10 denotes the output obtained after processing through Module 10 of the full KWPP-QC pipeline; Modules 11–13 are not applied at this stage. In Sect. 3.1, KWPP-QC10 was intentionally selected as the conventional baseline because the three standalone deep learning models (DL-v1, DL-v2, and DL-v3) are trained using raw spectral samples selected according to the flags generated through Modules 8–10. This allows the conventional and deep learning retrievals to be compared at the same processing stage. Applying Modules 11–13 only to the conventional retrieval at this stage would combine differences in the retrieval method with the effects of subsequent post-processing. The effects of Modules 11–13 are evaluated separately in Sect. 3.2. In this section, KWPP-QC10 and KWPP-QC13 are compared to evaluate the effect of conventional post-processing, while DL-v3 and DL-QCv3 are compared to evaluate the effect of applying the same post-processing procedures to the deep learning retrieval. Here, DL-QCv3 represents the DL-v3 retrieval after Modules 11–13 have been applied. This separation allows the retrieval-stage performance and the additional effects of post-processing to be evaluated independently.
- We understand the reviewer’s concern that several points are closely clustered in Figure 4. This clustering reflects the relatively similar normalized standard deviations and correlation coefficients among the evaluated models. The Taylor diagrams are intended to provide an overview of these statistical relationships relative to the RS reference, while the related quantitative performance metrics are provided separately in Table 3 and discussed in Sect. 3.1. We will revise Figure 4 to include enlarged views of the clustered regions while retaining the original Taylor diagrams for an overview of the overall statistical relationships.
Citation: https://doi.org/10.5194/egusphere-2026-3168-AC1
-
AC1: 'Reply on RC1', Kyung Hun Lee, 17 Sep 2026
-
RC2: 'Comment on egusphere-2026-3168', Anonymous Referee #2, 15 Sep 2026
This paper compares a deep learning framework to more traditional methods for retrieval and quality control of radar wind profiler data. The paper both utilizes and contrasts Lee et al (2026), which describes a manufacturer agnostic quality controlled output for the Korean radar wind profiler network. While the Lee paper uses traditional signal processing, the current paper uses a framework trained to produce wind vectors directly from the Doppler spectra. This framework is applied to data from two instruments at 437.5 MHz and two instruments at 1290 MHz, and goes on to combine the framework with quality control components from the Lee paper. Key differences between the two frequencies are found, which are attributed to oversampling and subsequent range smearing.
The paper is interesting, and highly topical as machine learning techniques are increasingly used. I believe the manuscript can be improved with some clarifications and minor edits as outlined below.
Firstly, section 2 could be improved by making the procedure clearer to the reader. On page 6, line 114 you state sonde observations were used to train the proposed deep learning models, implying the models were trained on sonde data only. Line 121 on page 7 and line 152 on page 8 both state the same thing, that the primary input is raw Doppler spectra, which implies just profiler data is used. On line 161 you state model training is performed by minimizing the mean squared error between the retrieved wind components and the spatiotemporally collocated RS observations. In section 2.3.3.3 we learn profiler data is classified into three categories based on the outputs of the traditional processing method in the Lee paper. Clearly you need both spectra and a wind profile to train the model, and you have described all components, but I think could be made much clearer to the reader. Perhaps you could add a flow chart, and then describe each component in turn?
Secondly, as this work both utilises components of Lee et al (2026) it would be helpful to the reader if further details are provided in the current paper. This is particularly true for section 2.3.3.3 where you are using outputs of modules 8, 9 and 10. A description of these, perhaps including a discussion on these being the moment level outputs, would greatly assist the reader.
As a note only, it would be interesting to know if the hardware difference comes about due to oversampling and the resulting range smearing alone, or are there other factors? The 437.5 MHz profilers would be less sensitive to scatter from precipitation than the 1290 MHz, so there may be some rainfall events where the lower frequency sees two peaks where the higher does not. Could this be contributing? It is beyond the scope of the current paper, but have you considered either running a short campaign sampling at the pulse length, or looking for other datasets without oversampling to compare?
Minor comments
Page 2 line 28 – suggest adding ‘such as’ or similar to the list of signal processing techniques.
Page 4 line 84 onwards, this section describing both frequencies makes no mention of precipitation as a target? While it is likely wind profiles are not retrieved during rainfall periods, this should be made explicit.
Page 5 Figure 2 could be improved by zooming in on the peaks, say plot the x axis from -100 to 100, and labelling features. You could also potentially overlay the peaks selected by conventional processing to demonstrate your point more clearly?
Page 6 table 1 – suggest you list the beam directions, and which one is repeated in the table description. Similarly does 10 minute temporal resolution refer to the production of a single wind profile, or an averaged product? The table suggests less than one minute per beam/dwell so it is likely two complete cycles averaged, but this should be made clear to the reader. Further it is implied but not made clear both frequencies sample at 50 m. This should be included in the table.
Page 7 beginning line 129 – you state you line up the profiler observation with the sonde launch time, but do not mention how much profiler data you use? Is it 10 minutes worth of data? Assuming a sonde ascent rate of 5 m/s the sonde will reach 2 km in a little under 7 minutes, is this why you are averaging two profiles?
Page 7 line 136 – as mentioned above, since you extensively compare to KWPP-QC presented in the Lee et al (2026) paper it would be appropriate to include a brief description here, enough to allow the reader to understand the premise, and then refer to the paper for further details.
Page 16 figure 4 – there is a lot of information on a very small portion of the plot. Like figure 2, this could be improved by zooming in on the values with reduced axes.
Page 18 figure 5 – I assume the left and right panels are u and v? These should be explicitly labelled.
Citation: https://doi.org/10.5194/egusphere-2026-3168-RC2 -
AC2: 'Reply on RC2', Kyung Hun Lee, 17 Sep 2026
We sincerely thank the reviewer for the careful evaluation of our manuscript and for the constructive comments. Our responses to each point are provided below in the order presented.
Major comments
- We agree that the roles of the different datasets and processing components were not sufficiently clear in the original manuscript. We will revise Sect. 2 to more clearly distinguish the three main components of the framework: the raw five-beam RWP Doppler spectra are used as model inputs, the collocated RS u and v wind components are used as supervised target variables and reference observations for performance evaluation, and the KWPP-QC quality flags are used to classify and select spectral samples for constructing the DL-v1, DL-v2, and DL-v3 training datasets. We will also clarify the temporal and vertical collocation procedures between the RWP and RS observations and reorganize the relevant descriptions in Sects. 2.2 and 2.3 to present the training workflow more consistently.
- We will expand the description of the KWPP-QC procedure in Sects. 2.2 and 2.3.3.3. In particular, Modules 8–10 will be explicitly identified as the moment-level QC stage of KWPP-QC. Module 8 applies low-SNR filtering, Module 9 evaluates vertical discontinuities in radial velocity, and Module 10 performs temporal consensus averaging of radial velocity. We will also clarify how the processing history through these modules defines the Certain, Uncertain, and Rejected categories and how these flags are subsequently used to construct the model-specific training datasets.
- We appreciate the reviewer’s suggestion that precipitation-related spectral structures may also contribute to the occurrence of multiple peaks at 437.5 MHz. Such double-peak structures are certainly possible when clear-air and precipitation echoes coexist. However, the retrieval issue discussed in this study is primarily associated with inconsistent peak selection among multiple spectral components, which can produce discontinuous wind estimates. During precipitation, precipitation echoes are often stronger than clear-air echoes and are therefore more likely to remain the dominant spectral component selected by conventional processing. In such cases, the same spectral component would tend to be selected consistently rather than switching irregularly between multiple peaks. Therefore, although precipitation may contribute to the occurrence of double-peak spectra, we consider it less likely to be the primary explanation for the discontinuous retrieval behavior examined in this study. Our interpretation was based on the vertically coherent multi-peak structures observed in the 437.5-MHz data and their consistency with the expected effects of oversampling and range smearing. We did not conduct a dedicated short-pulse experiment or compare the results with an independent non-oversampled dataset, as the interpretation was derived from the characteristics of the available operational observations.
Minor comments
- The sentence will be revised as suggested.
- Precipitation periods were not excluded from the present analysis. We will revise Sect. 2.1 to clarify the influence of precipitation echoes on the two profiler frequency bands. The 437.5-MHz profilers are relatively less sensitive to precipitation scattering than the 1290-MHz profilers, although precipitation echoes can coexist with clear-air Bragg echoes as distinct spectral components. In contrast, the higher operating frequency of the 1290-MHz profilers results in greater sensitivity to precipitation scattering, allowing precipitation echoes to contribute more strongly to the observed Doppler spectra.
- Figure 2 will be revised to provide an enlarged view of the relevant spectral region so that the multiple spectral peaks can be more clearly distinguished.
- We will add the actual beam sequence (E, N, V, W, S, V) to Table 1. We will also clarify in Sect. 2.2 that the RWPs alternate between low- and high-mode observations, with each mode requiring approximately 5 min, such that low-mode observations recur at approximately 10-min intervals. The 50-m gate spacing for all four profiler systems was already listed in Table 1 and remains unchanged.
- We will clarify the temporal matching procedure in Sect. 2.2. For each RS launch, a single low-mode RWP observation sequence corresponding to the launch time is selected. Because the RWPs alternate between low- and high-mode observations, successive low-mode profiles occur at approximately 10-min intervals. Thus, the comparison does not involve averaging two consecutive low-mode profiles. The RS profiles are subsequently resampled vertically to the 50-m RWP range gates.
- A brief overview of the KWPP-QC framework will be added to Sect. 2.2, including its role as a 13-module processing and QC system spanning spectral processing, moment estimation, and wind retrieval. Additional details of the moment-level QC Modules 8–10 and their use in the present deep learning experiments will be provided in Sect. 2.3.3.3.
- Figure 4 will be revised to include enlarged views of the clustered data region, allowing the differences among the retrieval methods to be more clearly identified.
- The figure will be revised to explicitly label the left and right columns as the u and v wind components, respectively.
Citation: https://doi.org/10.5194/egusphere-2026-3168-AC2
-
AC2: 'Reply on RC2', Kyung Hun Lee, 17 Sep 2026
Viewed
| HTML | XML | Total | BibTeX | EndNote | |
|---|---|---|---|---|---|
| 176 | 71 | 50 | 297 | 39 | 39 |
- HTML: 176
- PDF: 71
- XML: 50
- Total: 297
- BibTeX: 39
- EndNote: 39
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
This manuscript evaluates a deep learning framework for retrieving wind profiles directly from the Doppler spectra of Radar Wind Profilers (RWPs). The framework is trained with the objective of minimising the mean square differences from the wind componets for colocated radiosonde measurements. The same framework has been applied to data from RWPs operating at both 437.5 MHz and 1290 MHz. The best of 3 deep learning models results in an improvement in data quality compared to that of a conventional processing scheme, which has previously been described by Lee et al. (2026). Interestingly, a further improvement in quality is achieved by applying the final quality control modules from the conventional processing scheme to the output of the deep learning model.
This is an interesting approach and the analysis presented is thorough. I think the manuscript will be acceptable for publication after relatively minor changes are made.
It should be noted that the improvements in data quality seen with the deep learning framework reflect, at least in part, limitations of the Lee et al. (2026) scheme. We know from Lee et al. (2026) that the scheme is at least as good as the schemes provided by the manufacturers of the RWPs. However, we do not know how it compares with other processing schemes described in the literature. It would be useful to include a few more details about the Lee et al. (2026) scheme in order to give a feel of how robust it is likely to be. Specifically:
1) How many signal components are identified within each spectrum? Figure 2 of Lee et al. (2026) implies two, although this is not explicitly stated in the paper. Note that Griesser and Richner (1998) show the usefulness of idenitfying more than 2 components for observations made at 1290 MHz.
2) Does the Lee et al. (2026) scheme attempt to separate partially overlapping signal components? Note that Griesser and Richner (1998) show the usefulness of this for observations made at 1290 MHz.
3) It would be useful to overlay on Figure 2 of the manuscript the regions of the spectra that have been selected by the conventional processing scheme and (for the oblique beam spectra) the radial velocities implied by output of the deep learning scheme. See Figure 2 of Griesser and Richner (1998) for an example of what I mean. This will give some indication of how effective the Lee at al. (2026) scheme is at discriminating between clear-air and unwanted signal components. Ideally, an equivalent figure should be added for a cycle of observation made by a 1290 MHz RWP.
And here are my comments about specific parts of the manuscript.
4) page 6, Table 1, row 5: What does "spectral points" mean?
5) page 6, Table 1: What is the observation strategy for producing data at 10 minute intervals? The numbers in Table 1 suggest dwell lengths of approximately 46 s and line 122 (section 2.2) indicates that "the standard observation cycle includes six sequential measurements across five beam directions", which should take approximately 5 minutes. Do 10 minute interval data represent the averages of 2 cycles of observation?
6) page 7, section 2.2: Are only 10 minutes' worth of RWP data used for comparisons with radiosonde data?
7) page 7, section 2.2: Have the authors tested the deep learning scheme on the 437.5 MHz RWP observations for altitudes above 2 km? Clearly it makes sense to restrict primary attention to the altitude region 0.5 - 2.0 km to allow 1290 MHz observations to be included. However, since 437.5 MHz observations extend up to 8 km, it would be interesting to know something about the potential utility of the deep learning framework for the broader altitude range.
8) page 11, line 209 onwards. I think the authors are using the term "consensus check" to imply a time continuity test? This should be clarified since, although the term "consensus averaging" is typically used to imply a process carried out in the time domain, the word "consensus" itself does not specifically imply the time domain.
9) page 11, line 212. I think the word "out" needs to be added as shown in the following sentence since the term filtering can apply to either what is removed or what is retained. "Regions initially filtered OUT by Modules 8 or 9 but subsequently restored through the consensus procedure in Module 10." Note it is not necessary for "meteorological and non-meteorological echoes [to] coexist" in order for a region to be classified as "uncertain". A low signal-to-noise ratio would have the same effect.
10) page 12, lines 226 - 235. Some clarification is needed for how each model is trained. For example, it is stated that the DL-v1 model (Uncertain-only) is "trained exclusively using spectral samples containing at least one beam classified as 'Uncertain.'" Does this really mean one beam or does it mean at least one range gate (for one beam pointing direction within a cycle of observation) has been deemed to contain an "uncertain" spectrum? Moreover, are KWPP-QC10 flag values used as input to the model or are the flag values used only to decide which spectra are used as input to the model?
11) page 12, line 242: Brief details should be given as to what modules 11–13 of KWPP-QC do. This is directly relevant to subsequent analysis.
12) page 13, line 255: The full meaning of RMSE should be given here, where the abbreviation is first used. Similarly the full meaning of MBE should be given on line 259.
13) page 13, line 252: The manuscript does not fully describe the scope of KWPP-QC10. Presumably it does not make use of QC modules 11-13, which are subsequently applied to DL-v3? If so, why have they not been used in the comparisons against DL-v1, DL-v2, and DL-v3, which would provide a more meaningful comparison of conventional and deep learning retrieval schemes?
14) page 16, Figure 4. It's not so easy to distinguish between the different plot points since they are quite small and are located very close together in a small part of the diagram. It would be easier to distinguish them if the diagrams were expanded for a sub-portion of the values shown.