Wind Profile Retrieval Using Ensemble Learning and Mobile Brightness Temperature over the Tibetan Plateau
Abstract. Sparse and infrequent upper‑air observations over the Three‑River Source Region (TRSR) of the Tibetan Plateau (TP) severely limit the capability for extensive and continuous wind profile (WP) monitoring. A WP retrieval model was developed via ensemble learning (XGBoost+CatBoost), trained on TB data from 14 fixed-site ground-based microwave radiometers (GMWR) and WP radiosonde observations (Raob), and then applied to mobile GMWR observations over the TRSR. After strict QC of TB data, three distribution‑alignment methods (Quantile Transformer, PCA Whitening, RankGauss) were used to mitigate the TB discrepancy between fixed and mobile GMWR. A nonlinear weighted regression loss function was introduced to reduce the central tendency effect in long‑tailed wind‑speed distributions. Retrieved WP from 14 fixed‑site GMWR test data agreed well with Raob (R = 0.96, RMSE = 2.24 m s−1, MAE = 1.57 m s−1). Retrieval errors in the 600–700 hPa layer were lower at 08:00 BT than at 20:00, and smaller in lower‑altitude regions. Mobile GMWR retrievals reproduced WP variability with mean biases below 2 m s−1 relative to Raob, though errors increased for upper‑level strong winds and at high altitudes; against Raob‑corrected ERA5, the RMSE and MAE were 2.66 and 2.25 m s−1, respectively. This study offers a novel pathway to optimize ground‑based vertical remote sensing, enhance meteorological support for the low‑altitude economy, and improve wind energy assessment and forecasting.
This manuscript presents a novel approach for retrieving wind profiles (WP) from ground-based microwave radiometer (GMWR) brightness temperature (TB) observations over the Tibetan Plateau using an ensemble learning framework (XGBoost + CatBoost). The study addresses a critical gap—the sparse and infrequent upper-air observations over the TRSR—by leveraging both fixed-site and mobile GMWR data. The methodology incorporates rigorous quality control, distribution alignment techniques, and a weighted regression loss function to mitigate the central tendency effect in long-tailed wind speed distributions. Results show promising retrieval accuracy (R=0.96, RMSE=2.24 m/s for fixed sites) and demonstrate the feasibility of mobile GMWR-based WP retrieval over complex terrain.
Overall, this is a well-structured and technically sound study with significant implications for atmospheric monitoring over the Tibetan Plateau. However, several issues need to be addressed before publication.
Major Concerns
1. The manuscript claims novelty in three aspects: (1) using radiosonde observations rather than ERA5 for training, (2) applying distribution alignment to address domain shifts between fixed and mobile GMWR, and (3) introducing a weighted loss function for long-tailed wind distributions. While these are valuable contributions, the novelty is somewhat incremental given previous work by Chen et al. (2024) and Jiang et al. (2026). The authors should more clearly articulate what new scientific understanding or methodological advancement their work provides beyond these existing studies, particularly regarding the physical interpretability of why the ensemble approach outperforms single models.
2. The mobile GMWR validation against radiosonde data relies on only 22 matched profiles (Section 5.2.1). This is a very small sample size for drawing robust conclusions about model generalization. The authors acknowledge this limitation but then proceed to use Raob-corrected ERA5 as an alternative reference (204 profiles). While this is a reasonable workaround, the correction of ERA5 using only 186 Raob profiles introduces additional uncertainties. The authors should:
Provide confidence intervals for the mobile validation statistics;
Discuss how the small sample size affects the statistical significance of their conclusions;
Consider bootstrapping or other resampling techniques to estimate uncertainty;
And more explicitly state the limitations imposed by this small sample size in the conclusions.
3. The authors evaluate three distribution alignment methods and select the "optimal method independently for each pressure level based on retrieval errors". However:
Which method was chosen for which level? This information is not provided.
Why does the optimal method vary by pressure level? A physical explanation would strengthen the manuscript.
Was this selection performed on the validation set or test set? If performed on the test set, this introduces data leakage. The authors state the transformers were "fitted on the 70% training set and then applied to the validation/test sets", but the selection of the "optimal" method based on retrieval errors is ambiguous regarding which dataset was used for selection.
4. The three-level weighting scheme (w=1 for <10 m/s, 1.5 for 10-20 m/s, 2 for ≥20 m/s) is somewhat arbitrary. The authors should:
Provide a sensitivity analysis showing how different weighting schemes affect retrieval performance
Justify the specific thresholds (10 and 20 m/s) based on the data distribution or physical considerations
Consider whether a continuous weighting function (e.g., inversely proportional to sample frequency) might be more appropriate
5. The ERA5 correction method is described only briefly: "A linear correction model, based on 186 Raob WPs... was applied to correct the ERA5 WPs." This is insufficient for reproducibility:
What type of linear correction? (e.g., linear regression, bias correction, quantile mapping?)
Were the correction coefficients pressure-level specific?
What are the statistics of the correction (R, RMSE before/after correction)?
How were the 186 Raob profiles distributed across Xining, Dari, and Yushu?
Minor Issues
1. The selection of 12 common channels is mentioned but not justified. Why these specific channels? Do they represent the optimal subset for WP retrieval, or were they simply the intersection of all available channels?
2. Line 120: The QC procedures for fixed-site and mobile GMWR are described as "applied sequentially," but the specific order of operations for fixed-site QC is not clearly stated (unlike mobile QC which is more detailed).
3. Line 150: The 3-sigma method (k=3) is used for outlier detection. Is k=3 appropriate for all channels and all stations? Given the different variability of K-band vs. V-band channels, should different k values be considered?
4. Equations 4-5: The thresholds for the temporal consistency check (0.8 and 0.5 for low- and high-frequency channels) are stated to be "based on the statistical distribution under normal conditions" (line 154). Provide more detail on how these thresholds were determined.
5. Section 3.3: SG filtering is applied to both input features and training labels. Applying smoothing to the labels (radiosonde WS) may remove physical variability. The authors should justify why label smoothing is appropriate and discuss potential impacts on the retrieval of extreme wind events.
6. Figure 5a: The retrieved WS is consistently lower than Raob at most levels. The authors attribute this to "uncertainties in TB measurements, limited training samples, and radiosonde drift". Could there be a systematic bias in the retrieval model itself? This should be explored further.
7. Table 2: The relationship between TCC and retrieval errors is not monotonic—RMSE decreases from 2.18 to 2.12 between the 0.2-0.4 and 0.4-0.6 ranges. Is this physically meaningful or statistical noise?
8. Section 5.1.2(3): The RMSE increases with elevation, but R remains high (>0.96) across all elevation ranges. This suggests the correlation is maintained but the absolute error grows. The authors should discuss whether this is due to increased variability at higher altitudes or systematic bias in the retrievals.
9. Figure 9c: The systematic bias (overestimation at low WS, underestimation at high WS) is attributed to "domain shift" between fixed-site training data and mobile observations. However, this is also a classic symptom of regression to the mean. The authors should discuss whether the weighted loss function adequately addresses this issue and what additional strategies might be needed.
10. Line 25: The abstract uses "BT" for Beijing Time without defining it. Define at first use.
11. Equation 2 (Outlier Detection): The threshold for K-band channels (8.0 K) seems quite large given that typical TB variations in these channels are on the order of a few Kelvin. The authors should justify this threshold and show that it does not remove genuine meteorological variability.
12. The 3-sigma method assumes normally distributed data. Are the TB observations normally distributed? If not, the 3-sigma rule may not be appropriate. The authors should either justify the normality assumption or use a more robust outlier detection method.
13. The manuscript states that vehicle speed affects TB accuracy and that observations were removed based on INS speed measurements. However, the specific speed threshold used for removal is not provided. What speed was considered "abnormal"?
14. The antenna pointing correction is described as following Corbella et al. (2001), but the actual correction methodology is not explained. Given that mobile observations are a key contribution of this study, the antenna pointing correction deserves more detailed description.