the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
A Hardware-Aware Deep Learning Framework for Wind Retrieval from Raw Doppler Spectra of Radar Wind Profilers
Abstract. Radar wind profilers (RWPs) traditionally rely on heuristic statistical procedures for signal processing and quality control, often suffering from inherent trade-offs between data availability and retrieval reliability. To address these limitations, this study presents an end-to-end deep learning framework that retrieves wind vectors directly from raw Doppler spectra. Validated against a three-year (2022–2024) dataset of raw spectra from 437.5-MHz and 1290-MHz profiler systems and collocated radiosonde observations, the proposed framework demonstrates robust performance, particularly when trained on the full spectral continuum without filtering ambiguous samples. Notably, we found that hybridizing deep learning with conventional post-processing is highly contingent upon hardware-specific error characteristics; it effectively mitigated transient outliers in the 1290-MHz systems but provided negligible benefits for the range-smeared, vertically coherent artifacts dominating the 437.5-MHz systems. Furthermore, the framework reasonably reproduced seasonal, diurnal, and vertical atmospheric structures without the excessive smoothing typical of conventional statistical approaches. Although constrained by a slight underestimation of extreme wind speeds due to the regression-to-the-mean effect inherent in mean-squared-error optimization, the model notably enhanced operational data availability and integrity, demonstrating the feasibility of hardware-aware deep learning as a viable alternative or complement to conventional RWP processing pipelines.
- Preprint
(2067 KB) - Metadata XML
- BibTeX
- EndNote
Status: open (until 29 Sep 2026)
- RC1: 'Comment on egusphere-2026-3168', Anonymous Referee #1, 11 Sep 2026 reply
-
RC2: 'Comment on egusphere-2026-3168', Anonymous Referee #2, 15 Sep 2026
reply
This paper compares a deep learning framework to more traditional methods for retrieval and quality control of radar wind profiler data. The paper both utilizes and contrasts Lee et al (2026), which describes a manufacturer agnostic quality controlled output for the Korean radar wind profiler network. While the Lee paper uses traditional signal processing, the current paper uses a framework trained to produce wind vectors directly from the Doppler spectra. This framework is applied to data from two instruments at 437.5 MHz and two instruments at 1290 MHz, and goes on to combine the framework with quality control components from the Lee paper. Key differences between the two frequencies are found, which are attributed to oversampling and subsequent range smearing.
The paper is interesting, and highly topical as machine learning techniques are increasingly used. I believe the manuscript can be improved with some clarifications and minor edits as outlined below.
Firstly, section 2 could be improved by making the procedure clearer to the reader. On page 6, line 114 you state sonde observations were used to train the proposed deep learning models, implying the models were trained on sonde data only. Line 121 on page 7 and line 152 on page 8 both state the same thing, that the primary input is raw Doppler spectra, which implies just profiler data is used. On line 161 you state model training is performed by minimizing the mean squared error between the retrieved wind components and the spatiotemporally collocated RS observations. In section 2.3.3.3 we learn profiler data is classified into three categories based on the outputs of the traditional processing method in the Lee paper. Clearly you need both spectra and a wind profile to train the model, and you have described all components, but I think could be made much clearer to the reader. Perhaps you could add a flow chart, and then describe each component in turn?
Secondly, as this work both utilises components of Lee et al (2026) it would be helpful to the reader if further details are provided in the current paper. This is particularly true for section 2.3.3.3 where you are using outputs of modules 8, 9 and 10. A description of these, perhaps including a discussion on these being the moment level outputs, would greatly assist the reader.
As a note only, it would be interesting to know if the hardware difference comes about due to oversampling and the resulting range smearing alone, or are there other factors? The 437.5 MHz profilers would be less sensitive to scatter from precipitation than the 1290 MHz, so there may be some rainfall events where the lower frequency sees two peaks where the higher does not. Could this be contributing? It is beyond the scope of the current paper, but have you considered either running a short campaign sampling at the pulse length, or looking for other datasets without oversampling to compare?
Minor comments
Page 2 line 28 – suggest adding ‘such as’ or similar to the list of signal processing techniques.
Page 4 line 84 onwards, this section describing both frequencies makes no mention of precipitation as a target? While it is likely wind profiles are not retrieved during rainfall periods, this should be made explicit.
Page 5 Figure 2 could be improved by zooming in on the peaks, say plot the x axis from -100 to 100, and labelling features. You could also potentially overlay the peaks selected by conventional processing to demonstrate your point more clearly?
Page 6 table 1 – suggest you list the beam directions, and which one is repeated in the table description. Similarly does 10 minute temporal resolution refer to the production of a single wind profile, or an averaged product? The table suggests less than one minute per beam/dwell so it is likely two complete cycles averaged, but this should be made clear to the reader. Further it is implied but not made clear both frequencies sample at 50 m. This should be included in the table.
Page 7 beginning line 129 – you state you line up the profiler observation with the sonde launch time, but do not mention how much profiler data you use? Is it 10 minutes worth of data? Assuming a sonde ascent rate of 5 m/s the sonde will reach 2 km in a little under 7 minutes, is this why you are averaging two profiles?
Page 7 line 136 – as mentioned above, since you extensively compare to KWPP-QC presented in the Lee et al (2026) paper it would be appropriate to include a brief description here, enough to allow the reader to understand the premise, and then refer to the paper for further details.
Page 16 figure 4 – there is a lot of information on a very small portion of the plot. Like figure 2, this could be improved by zooming in on the values with reduced axes.
Page 18 figure 5 – I assume the left and right panels are u and v? These should be explicitly labelled.
Citation: https://doi.org/10.5194/egusphere-2026-3168-RC2
Viewed
| HTML | XML | Total | BibTeX | EndNote | |
|---|---|---|---|---|---|
| 92 | 38 | 24 | 154 | 25 | 23 |
- HTML: 92
- PDF: 38
- XML: 24
- Total: 154
- BibTeX: 25
- EndNote: 23
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
This manuscript evaluates a deep learning framework for retrieving wind profiles directly from the Doppler spectra of Radar Wind Profilers (RWPs). The framework is trained with the objective of minimising the mean square differences from the wind componets for colocated radiosonde measurements. The same framework has been applied to data from RWPs operating at both 437.5 MHz and 1290 MHz. The best of 3 deep learning models results in an improvement in data quality compared to that of a conventional processing scheme, which has previously been described by Lee et al. (2026). Interestingly, a further improvement in quality is achieved by applying the final quality control modules from the conventional processing scheme to the output of the deep learning model.
This is an interesting approach and the analysis presented is thorough. I think the manuscript will be acceptable for publication after relatively minor changes are made.
It should be noted that the improvements in data quality seen with the deep learning framework reflect, at least in part, limitations of the Lee et al. (2026) scheme. We know from Lee et al. (2026) that the scheme is at least as good as the schemes provided by the manufacturers of the RWPs. However, we do not know how it compares with other processing schemes described in the literature. It would be useful to include a few more details about the Lee et al. (2026) scheme in order to give a feel of how robust it is likely to be. Specifically:
1) How many signal components are identified within each spectrum? Figure 2 of Lee et al. (2026) implies two, although this is not explicitly stated in the paper. Note that Griesser and Richner (1998) show the usefulness of idenitfying more than 2 components for observations made at 1290 MHz.
2) Does the Lee et al. (2026) scheme attempt to separate partially overlapping signal components? Note that Griesser and Richner (1998) show the usefulness of this for observations made at 1290 MHz.
3) It would be useful to overlay on Figure 2 of the manuscript the regions of the spectra that have been selected by the conventional processing scheme and (for the oblique beam spectra) the radial velocities implied by output of the deep learning scheme. See Figure 2 of Griesser and Richner (1998) for an example of what I mean. This will give some indication of how effective the Lee at al. (2026) scheme is at discriminating between clear-air and unwanted signal components. Ideally, an equivalent figure should be added for a cycle of observation made by a 1290 MHz RWP.
And here are my comments about specific parts of the manuscript.
4) page 6, Table 1, row 5: What does "spectral points" mean?
5) page 6, Table 1: What is the observation strategy for producing data at 10 minute intervals? The numbers in Table 1 suggest dwell lengths of approximately 46 s and line 122 (section 2.2) indicates that "the standard observation cycle includes six sequential measurements across five beam directions", which should take approximately 5 minutes. Do 10 minute interval data represent the averages of 2 cycles of observation?
6) page 7, section 2.2: Are only 10 minutes' worth of RWP data used for comparisons with radiosonde data?
7) page 7, section 2.2: Have the authors tested the deep learning scheme on the 437.5 MHz RWP observations for altitudes above 2 km? Clearly it makes sense to restrict primary attention to the altitude region 0.5 - 2.0 km to allow 1290 MHz observations to be included. However, since 437.5 MHz observations extend up to 8 km, it would be interesting to know something about the potential utility of the deep learning framework for the broader altitude range.
8) page 11, line 209 onwards. I think the authors are using the term "consensus check" to imply a time continuity test? This should be clarified since, although the term "consensus averaging" is typically used to imply a process carried out in the time domain, the word "consensus" itself does not specifically imply the time domain.
9) page 11, line 212. I think the word "out" needs to be added as shown in the following sentence since the term filtering can apply to either what is removed or what is retained. "Regions initially filtered OUT by Modules 8 or 9 but subsequently restored through the consensus procedure in Module 10." Note it is not necessary for "meteorological and non-meteorological echoes [to] coexist" in order for a region to be classified as "uncertain". A low signal-to-noise ratio would have the same effect.
10) page 12, lines 226 - 235. Some clarification is needed for how each model is trained. For example, it is stated that the DL-v1 model (Uncertain-only) is "trained exclusively using spectral samples containing at least one beam classified as 'Uncertain.'" Does this really mean one beam or does it mean at least one range gate (for one beam pointing direction within a cycle of observation) has been deemed to contain an "uncertain" spectrum? Moreover, are KWPP-QC10 flag values used as input to the model or are the flag values used only to decide which spectra are used as input to the model?
11) page 12, line 242: Brief details should be given as to what modules 11–13 of KWPP-QC do. This is directly relevant to subsequent analysis.
12) page 13, line 255: The full meaning of RMSE should be given here, where the abbreviation is first used. Similarly the full meaning of MBE should be given on line 259.
13) page 13, line 252: The manuscript does not fully describe the scope of KWPP-QC10. Presumably it does not make use of QC modules 11-13, which are subsequently applied to DL-v3? If so, why have they not been used in the comparisons against DL-v1, DL-v2, and DL-v3, which would provide a more meaningful comparison of conventional and deep learning retrieval schemes?
14) page 16, Figure 4. It's not so easy to distinguish between the different plot points since they are quite small and are located very close together in a small part of the diagram. It would be easier to distinguish them if the diagrams were expanded for a sub-portion of the values shown.