Improving representativeness of microwave radiometer brightness temperatures for data assimilation by complementing cloud detection with cloud clearing
Abstract. This study introduces two new retrievals for ground-based microwave radiometers (MWR) and demonstrates that, together, they solve the misrepresentation issue that makes it challenging to assimilate cloudy-sky brightness temperatures. Multiple studies have shown that assimilating MWR brightness temperatures is beneficial to numerical weather prediction, despite rejecting substantial amounts of valid observations from the assimilation due to the presence of liquid water clouds. Cloud detection complemented with a cloud clearing retrieval makes this rejected data available for direct assimilation.
We scrutinize the introduced cloud detection retrieval and cloud clearing retrieval using multiple approaches. For one, we contextualize the new retrievals with reference retrievals. Secondly, we determine the retrieval sensitivity to instrument errors using a Monte Carlo method. And most importantly, we assess the representativeness of the retrieval products, which is crucial for data assimilation, through observation minus background statistics.
Our analysis reveals that the new cloud detection retrieval predicts less false positive cloudy-sky cases (-13 % over two years) compared to the established reference. Additionally, the cloud-clearing retrieval improves the representativeness between observation and background to the extent that artificial clear-sky and native clear-sky statistics almost match. Considering the minimal instrumentation required, both retrievals perform surprisingly well for elevation angles from zenith down to 4.8 °.
Overall, our findings demonstrate that combining cloud detection and cloud clearing retrievals improves the representativeness between observations and model. This retrieval combination enables the direct assimilation of cloudy-sky brightness temperatures.
General comments
This manuscript presents two neural-network-based retrievals for ground-based microwave radiometers: Cloud-RERA5 for cloud detection and Clearing-RERA5 for removing the spectral contribution of cloud liquid water from observed brightness temperatures. The topic is timely and relevant to the broader use of microwave-radiometer observations in data assimilation, particularly at off-zenith viewing angles. The manuscript is generally well organized, and the results indicate that the proposed methods have potential to increase the number of microwave-radiometer observations that can be used under cloudy conditions.
My main concern is the interpretation of the evaluation results. Throughout the manuscript, the terms observation error, representativeness, representativeness error, and O–B statistics are sometimes used interchangeably. However, the analyses presented in this study directly evaluate O–B statistics, which contain contributions from observation error, background error, representativeness error, and observation-operator error. Therefore, a reduction in O–B standard deviation does not necessarily demonstrate a corresponding reduction in observation error alone. I recommend revising the terminology throughout the manuscript and consistently distinguishing between the quantities directly evaluated and their implications for data assimilation.
A second concern is that several conclusions are stronger than supported by the available validation data. Comparisons among Cloud-RERA5, STD30, STD6, and LWP-RRS quantify agreement among retrieval algorithms rather than absolute performance, because none of these methods provides independent ground truth. More cautious wording would improve the scientific precision of the manuscript without changing its main findings.
Major comments
L38, L41, L65–68, L423, L439: The term “representativeness” is used with several different meanings throughout the manuscript. The discussion appears to combine subgrid-scale variability, cloud-location errors in the model, systematic cloud-water biases, observation-operator errors, and cloud contamination. These are distinct sources of uncertainty. Please define “representativeness error” more precisely and distinguish it consistently from model error, forecast displacement error, observation-operator error, and instrumental observation error.
L114-115: The description of the quality-control procedure is insufficiently quantitative. The sentence “Carefully flagged data are removed from our analysis” does not specify which flags were applied or how much data were excluded. Please report the number or percentage of observations removed because of precipitation, radome degradation, instrument failure, and calibration issues. Given the extended exclusion period from August 2021 to January 2022, please also discuss whether the remaining dataset is sufficiently representative of the intended evaluation period.
L133-134: The procedure used to extend the ICON-D2 atmospheric profiles to 30 km is not described in sufficient detail. Please explain how the standard-atmosphere profiles were merged with the ICON-D2 profiles, at which altitude or pressure level the transition was applied, and how continuity in temperature, pressure, and humidity was ensured across the transition.
L137, L144–145: The radiative-transfer configuration requires further discussion. First, only pressure, temperature, humidity, and liquid water content are mentioned as atmospheric inputs. Please clarify whether frozen hydrometeors were neglected and justify this assumption, particularly for the higher-frequency HATPRO channels. Second, off-zenith brightness temperatures are computed from a single vertical model column under the assumption of horizontal homogeneity. Please discuss the implications of this approximation for cloudy scenes, where horizontal gradients and cloud displacement may be substantial, and comment on whether it could affect the evaluation of the proposed retrievals.
L165-166: The neural-network development and model-selection procedure is not described in sufficient detail for reproducibility. Please specify the hyperparameters that were explored, their ranges, the validation metric used for selection, and the architecture of the final models. In particular, the phrase “selected the best neural network manually” should be replaced by an objective description of the selection procedure.
L168–174: The bias-correction procedure requires further explanation. Please provide more information on the architecture, inputs, outputs, and training procedure of the internal bias-correction network. The rationale for using a 24 h correction window should also be justified, including whether the retrieval performance is sensitive to this choice. Furthermore, the manuscript states that the correction is forward propagated when no clear-sky observations are available in the next time bin. Please explain how long this propagation is allowed to continue and how the correction is reinitialized following prolonged cloudy periods.
Because the bias correction relies on Cloud-RERA5 to identify clear-sky periods, please also discuss how cloud-detection errors propagate into the bias correction and subsequently into Clearing-RERA5.
L190–194: The definitions and evaluation of the retrieval uncertainties should be clarified. The terms “retrieval error” and “product error” are currently ambiguous. Please specify the reference against which retrieval error is defined and explain how retrieval error and sensitivity to instrument errors are combined.
In addition, Cloud-RERA5 and Clearing-RERA5 are evaluated using different approaches: Monte Carlo perturbations for Cloud-RERA5 and O–B statistics for Clearing-RERA5. Please explain why different methodologies are required and whether the resulting uncertainty estimates are intended to be quantitatively comparable.
L199–203 and throughout Sections 2.6, 3.1.3, and 3.2.2: The interpretation of the O–B standard deviation should be revised throughout the manuscript. O–B statistics contain contributions from observation error, background error, representativeness error, and observation-operator error. Therefore, they do not provide a direct estimate of observation error alone.
The use of O–B standard deviation as a practical indicator of retrieval performance is reasonable, but it should be described explicitly as a proxy for the combined uncertainty reflected in the O–B differences. The statement that forecast errors are “included into the observation error” is potentially misleading. A more accurate interpretation is that the individual error contributions cannot be separated using O–B statistics alone.
L267 and Section 3.1: The selection of the LWP threshold of 5 g m⁻² requires stronger quantitative justification. The manuscript explains the qualitative trade-off between probability of detection and false-alarm rate, but it remains unclear why 5 g m⁻² is the preferred threshold. Please show the sensitivity of the relevant performance metrics, and preferably the O–B statistics, to the selected threshold or provide independent evidence supporting this choice.
Section 3.1.1: The Monte Carlo experiment assumes a static Gaussian brightness-temperature bias with σ = 0.5 K. Please justify why this perturbation is representative of realistic instrumental uncertainties. In particular, discuss whether random noise, calibration drift, temporally varying biases, and interchannel error correlations could affect the conclusions.
L302: The statement that “cloud detection is perfect” is too strong. The results show a POD of 100% above the specified LWP range under the assumed perturbation model. The conclusion should be restricted to those experimental conditions.
L315–345: Cloud-RERA5 is used as the shared reference when comparing the cloud-detection retrievals. Since Cloud-RERA5 is itself a retrieval rather than independent ground truth, it may be helpful to clarify that disagreements with Cloud-RERA5 do not necessarily correspond to false detections or missed detections by STD30, STD6, or LWP-RRS. Instead, the reported statistics could be described as reflecting the level of agreement among the retrievals.
L345: For the same reason, the statement that the differences “mostly favor Cloud-RERA5” appears stronger than supported by the comparison. A more neutral interpretation should be adopted unless independent observational validation is available.
L349–351, L365–367: The analysis shows that Cloud-RERA5 and STD30 produce similar clear-sky O–B standard deviations despite selecting different numbers of clear-sky scenes. This is a useful result. However, it does not directly demonstrate that the two methods yield the same (approximated) observation error. It would be more precise to phrase the conclusion in terms of the similarity of the resulting clear-sky O–B statistics.
L353–355: Cases classified as clear sky by the retrievals but cloudy by ICON-D2 were excluded from the analysis. Please report the fraction of observations removed by this criterion and discuss whether this filtering could reduce the apparent differences between Cloud-RERA5 and STD30.
L376–377, L401–408: The evaluation of Clearing-RERA5 relies on linearly interpolated brightness temperatures between the last clear-sky observation before a cloud passage and the first clear-sky observation after it. This is a useful practical reference, but it is not independent ground truth. Convective cloud passages may be associated with systematic evolution in temperature and humidity, violating the linear-interpolation assumption. Please discuss more explicitly how this limitation may affect the inferred retrieval performance.
L401–408: The statement that the remaining residuals can be largely explained by atmospheric variability appears somewhat stronger than the presented evidence. The observed correlation with IWV fluctuations is consistent with this interpretation, but it does not exclude residual retrieval errors or errors in the interpolated reference. It may therefore be helpful to phrase this conclusion more cautiously.
L421–424, L434–437: The reduction in O–B standard deviation after cloud clearing is an important result, but it should not be described directly as a reduction in random observation error. The analysis demonstrates a reduction in cloudy-sky O–B variability after both the observed and simulated brightness temperatures have been transformed into an artificial clear-sky framework. It would therefore be more precise to describe the result in terms of the resulting O–B statistics rather than the observation error itself.
L422–424: Similarly, the statement that the cleared cloudy observations have a representativeness comparable to native clear-sky observations appears stronger than the presented evidence. The results show comparable O–B standard deviations, but they do not isolate or directly estimate representativeness error.
L419–420: Please also explain why the cleared O–B standard deviation is occasionally smaller than the native clear-sky value. This may be relevant for interpreting whether the cloud-clearing transformation changes variability unrelated to cloud liquid water.
L446–450: Several statements in the Conclusions could be moderated. In particular, the claim that Cloud-RERA5 “identifies all clouds relevant for data assimilation” is not supported by an independent cloud-truth dataset. It would be more accurate to describe the retrieval performance under the evaluation conditions used in this study.
L458–460: The statement that cloud clearing “could double the effect of assimilating zenith MWR TB observations” appears somewhat speculative because no data-assimilation experiments are presented. The reduction in O–B variability suggests potential benefits for data assimilation, but the actual impact on the analysis and forecast should be evaluated in future assimilation experiments.
Minor comments
L44–46: Please clarify that observation-error inflation and VarQC reduce observational influence through different statistical mechanisms.
L65: Consider replacing “representativeness of MWR TBs and ICON-D2” with “consistency between MWR observations and ICON-D2 simulations.”
L99: Please briefly explain the role of the IR-radiometer when it is first introduced.
L109: Please clarify what is meant by uncertainty. If instrumental errors are intended, measurement uncertainty or instrument error may be more appropriate.
L111: Please briefly describe the method of Löffler (2024).
L116–117: Please define more specifically what is meant by consistent, stable, and climatologically representative.
L123: Please verify the stated ERA5 horizontal resolution 0.02 degree.
L123–126: Please clarify whether the ERA5 profiles were interpolated onto regular 25-hPa pressure levels.
L124: Please explain which numerical instability is avoided by truncating the profiles at 30 km.
L125: Please revise the wording, as ERA5 provides surface variables separately and the current phrasing may be misleading.
L135: Consider replacing “independent model reference” with “independent model dataset” or “independent evaluation dataset.”
L146: Please briefly justify neglecting antenna beam width and receiver bandwidth.
L161: Please specify which dimensionality-reduction methods were tested.
L171: The statement that the differences "should ideally be zero" appears too strong.
L180: Consider replacing "response to instrument errors" with "sensitivity to instrument errors."
L181: Please clarify the relevance of Szegedy et al. (2013), which primarily concerns adversarial perturbations, to realistic microwave-radiometer instrument errors.
L210–216: Please distinguish representativeness errors from geometric effects (e.g., beam width, beam tracing, and viewing geometry), as these appear to be discussed together.
L223–240: Please clarify whether the benchmark retrievals were used in their original form or retuned for the present dataset.
232–240: Please discuss how the zenith-only benchmark methods limit the evaluation of off-zenith retrievals.
L238–240: Please justify the selected LWP threshold (5 g m⁻²) or provide an appropriate reference.
L252: Consider replacing "impact on the approximated observation error" with "impact on the O–B statistics."
L252: Please clarify what is meant by "conceptually works."
L259–262: Please clarify whether all available MWR channels or only the channels used by Cloud-RERA5 were included in the analysis.
L265–267: Consider replacing "ideal threshold" with "selected threshold" or "practical threshold", as the threshold reflects a trade-off rather than a unique optimum.
L276–281: Please justify the use of 100 Monte Carlo realizations and explain why this number was considered sufficient.
L273–295: Please define POD_local when it is first introduced, rather than only when its mathematical definition is given.
L315: Please explain why the detailed comparison is restricted to summer 2021.
L326–332: Please provide quantitative support for the explanation of increased false detections (e.g., the effects of temporal averaging and humidity-dependent thresholds).
L313–345: Please confirm that identical quality-control criteria were applied to all retrievals before comparison.
L353–358: Please clarify whether the reported cloud-cover fractions were calculated after all exclusion criteria had been applied.
L362–363: Please define native clear-sky STD when it is first introduced.
L353–365: Please report the absolute number of retained clear-sky samples.
L362–365: Please support the statement that the differences are negligible with a quantitative criterion.
L376–380: Please explain how cloud passages were identified.
L377–380: Please justify the selected duration and LWP thresholds.
L380–384: Please report the number of cases in different LWP ranges.
L396–404: Please clarify whether the IWV retrieval is independent of the brightness temperatures used by Clearing-RERA5.
L405–408: Please replace "yields plausible cleared TBs" with a more objective statement.
L389–394: Please briefly discuss the increased variance above approximately 54 GHz.
L413–414: Please define native clear-sky, native cloudy-sky, and artificial clear-sky more explicitly when these terms are first introduced.
L410–413: Please clarify which parts of the ICON-D2 state were modified during cloud clearing.
L424–433: Please discuss the origin of the increased O–B standard deviation in the upper V-band.
L463–465: Please clarify what is meant by “low requirements for additional instrumentation.”
L452–453: Consider replacing “pioneers the potential” with a more objective expression.
L458–465: Please distinguish the demonstrated results from the expected implications for future data assimilation.
L466–470: Please mention that part of the evaluation relies on retrieval-to-retrieval comparisons and interpolated brightness temperatures rather than independent observational truth.