the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
A Deep Convolutional Neural Network for Retrieving Tropospheric Temperature and Moisture Profiles from Refractivity over Tropical Oceans: Framework Development and Characterization
Abstract. Retrieving profiles of temperature and water vapor from atmospheric refractivity over tropical oceans constitutes an inherently underdetermined problem. Conventional one-dimensional variational methods resolve this through numerical weather prediction (NWP) priors, potentially propagating model biases into retrieved profiles and limiting the independence desirable for climate monitoring applications. We present a deep learning retrieval that substitutes learned statistical constraints for model-dependent priors. A convolutional neural network, trained on approximately 20,800 high resolution radiosonde profiles —combining reference-grade GCOS Reference Upper-Air Network (GRUAN) measurements with quality-controlled operational GCOS Upper Air Network (GUAN) and field campaign data— predicts dry refractivity and partial pressure of dry air as intermediate targets; temperature, water vapor pressure, and relative humidity are derived analytically. The model achieves water vapor pressure root-mean-squared errors of ~0.5 hPa near the surface (decreasing with height) and relative humidity errors below 6 % from 100 m to 10 km, performance comparable in magnitude to the mean measurement uncertainty of state-of-the-art radiosondes. Temperature errors are ~1.5 K near the surface, improving to ~0.5 K in the 10–15 km range where dry refractivity dominates. Evaluation against more than 27,400 geographically independent radiosonde ascents across the tropical Pacific, Indian Ocean, and Atlantic—including observations from the 1992–93 TOGA-COARE field campaign, two decades predating the training period—demonstrates robust generalization within the tropical marine atmosphere. The framework accepts only refractivity vertical structure as input, with no dependence on geographic coordinates or NWP background states. This paper establishes the retrieval framework and characterizes performance under ideal input conditions; Part 2 addresses application to satellite observations.
- Preprint
(2425 KB) - Metadata XML
-
Supplement
(2804 KB) - BibTeX
- EndNote
Status: final response (author comments only)
- RC1: 'Comment on egusphere-2026-2027', Anonymous Referee #1, 24 Jun 2026
-
RC2: 'Comment on egusphere-2026-2027', Anonymous Referee #2, 02 Jul 2026
Recommended decision: Major revision
General comments
- This is an interesting paper. The manuscript addresses an important and timely problem in atmospheric measurement science: how to retrieve temperature and water-vapor/moisture profiles below 20 km from refractivity N over tropical oceans. The authors develop a deep convolutional neural network framework that uses radiosonde-derived refractivity as input and radiosonde thermodynamic variables as targets. The paper uses high-resolution radiosonde records, GUAN/GRUAN-related observations, and field-campaign soundings from DYNAMO and TOGA-COARE to evaluate whether the model can retrieve T and W without using an NWP background during inference. This topic is suitable for AMT because it concerns retrieval methodology, characterization, and potential observational applications.
2) The key and unique contribution is the attempt to train a model that takes N as input and retrieves T and W below 20 km. In my view, the novelty lies not only in the use of a CNN. The more important contribution is whether the authors have found a defensible way to learn the information contained in refractivity and to separate the temperature, pressure, and moisture contributions in a physically meaningful way. Therefore, two questions (as mentioned in the paper) are central for this paper. i) what is the retrieval accuracy? ii) how much does the AI approach learn beyond climatological covariance, radiosonde sampling priors, and simple regression relationships? I expect authors to provide solid evidence to demonstrate i) and ii), so that this approach will be useful for the future inversion for deriving T and W using real RO observations.
My evaluations for i) and ii) presented in the current paper are:
i) Evidence for accuracy. The manuscript provides useful evidence: RMSE, bias, standard deviation, profile examples, confidence intervals, station/campaign evaluations, and comparisons among training, GUAN, DYNAMO, and TOGA-COARE datasets. These results show that the framework can reproduce radiosonde thermodynamic profiles from radiosonde-derived N under ideal input conditions. However, this is still a closed-loop evaluation. Because the input N and target T/W are derived from the same radiosonde ascents, the evidence for accuracy does not yet demonstrate retrieval performance on actual GNSS-RO refractivity profiles.
Please explain how the current setting can be adapted to actual GNSS-RO refractivity profiles?
I understand that for real RO data, vertical smoothing, horizontal averaging, super-refraction, critical refraction, lower-tropospheric N biases, noise, and missing low-level information may substantially change retrieval errors.
However, at least, you may add vertical N uncertainty to your test set then do the 1D var exercise.
ii) Evidence for how much the AI model learns. The paper provides evidence that the model learns nontrivial empirical structure, particularly through cross-station/campaign tests, profile examples, and the perturbation experiment. These tests are useful, but they do not quantify the incremental value of the AI approach. To support the claim that the AI model has learned useful retrieval information, the paper needs baseline comparisons and ablation tests. At a minimum, the authors should compare against altitude-only climatology, basin/season climatology, multilinear regression, simpler MLP/random-forest retrievals, direct prediction of T/W/RH, and the same CNN without the intermediate targets or wavelet-transform input. Without these tests, it is difficult to know how much skill comes from the proposed AI architecture and how much comes from the limited tropical state space and strong T-W-N covariability.
Here I summarize my claim-evidence assessment
Claim or issue
Evidence provided
Assessment
Needed revision
Retrieval accuracy
RMSE, bias, standard deviation, confidence intervals, profile examples, and independent station/campaign evaluations.
Useful for ideal-input characterization, but not sufficient for real RO retrieval accuracy.
Add RO-like smoothing/noise/bias tests and real GNSS-RO validation.
AI learning
Cross-station evaluation, WCTN/intermediate-target design, perturbation experiment, and profile examples.
Suggestive, but not quantitative. The paper does not show the gain over simple baselines.
Add climatology, regression, MLP/random forest, direct-output CNN, and ablation tests.
NWP independence
The retrieval does not use an NWP background during inference.
Partly supported. The model still depends on the empirical radiosonde prior used in training.
Use “NWP-independent during inference” and avoid “prior-free” language.
Satellite RO applicability
The paper motivates GNSS-RO use and states that real RO validation will be Part 2.
Not yet demonstrated. Radiosonde-derived N is not real RO N.
Keep claims restricted to framework characterization under ideal input conditions.
Uncertainty
RMSE, bias, standard deviation, and confidence intervals.
Incomplete. RMSE is not per-profile retrieval uncertainty and CIs likely assume independent profiles.
Add uncertainty propagation, block bootstrap, sensitivity tests, and per-profile uncertainty estimates.
Key problems that should be addressed
1) The writing sometimes presents claims more strongly than the evidence supports. The manuscript should distinguish what is demonstrated from what remains a future application. The current evidence supports the ideal-input radiosonde-to-radiosonde framework characterization. It does not yet support broad claims of general-purpose retrieval, independent observational products, input-source independence, or satellite RO applicability. The authors should revise the abstract, introduction, discussion, and conclusions so that each claim is tied directly to the evidence presented. (Some of the related lines can be seen on 16-40, 23-31, 33-40, 181-190, 605-618, 741-744, 888-895, 950-972, 1019-1025, and 1067-1073.
2) Ideal inputs and retrievals do not prove that the approach works for real RO data. The paper uses radiosonde-derived refractivity as the input. This is internally consistent and useful for testing the framework, but it is not equivalent to GNSS-RO refractivity. Real RO profiles have vertical smoothing, horizontal averaging, tracking errors, super-refraction, critical refraction, lower-tropospheric negative N bias, and missing/low-level data. These effects directly influence the N-to-T/W inversion. Therefore, the present results should not be described as validation for real satellite RO retrievals. (see 16-40, 33-40, 83-113, 315-323, 605-618, 745-781, and 1067-1073).
3) The paper does not explain why temperature retrievals below 20 km are not better. Below 20 km, temperature retrieval from refractivity is underdetermined because N combines dry-air and water-vapor contributions. If the model retrieves temperature with only limited improvement, the paper should explain whether the limitation comes from T-W ambiguity, the dominance of moisture variability in the lower troposphere, radiosonde representativeness, the architecture, the intermediate-target design, or the limited information content of N. The manuscript needs a physical interpretation of where the retrieval contains information and where it is mostly constrained by the learned empirical prior.
4) Several results are overstated. Terms such as “independent observational products,” “prior-free,” “input-source independent,” “platform-agnostic,” and “general-purpose learned inversion operator” should be revised. More defensible language would be “NWP-independent during inference,” “trained with an empirical radiosonde prior,” and “evaluated under ideal radiosonde-derived input conditions.” (some of the lines includes 16-40, 145-152, 169-173, 181-190, 203-210, 605-618, 820-860, 888-895, 950-972, 1019-1025, and 1067-1073)
Detailed comments
- Lines 16-40: Abstract and scope. The abstract should clearly state that the evaluation uses radiosonde-derived refractivity as ideal input. The phrase that the framework is independent of NWP background states is acceptable only for inference. The radiosonde training distribution still constrains the model. The abstract should not imply satellite RO validation.
- Lines 23-31: Accuracy versus measurement uncertainty. Please do not compare RMSE directly with radiosonde measurement uncertainty as if they are the same quantity. RMSE includes retrieval error, representativeness error, sampling effects, and possible residual QC effects. The coverage factor should be defined consistently. The paper should avoid implying that the retrieval reaches the measurement limit unless this is demonstrated by uncertainty propagation.
- Lines 33-40 and 249-263: Ideal-input framework characterization. The scope paragraph is important and scientifically honest. The paper should move this idea into the abstract and conclusions. The present paper establishes a framework for behavior under ideal input conditions; it does not yet demonstrate performance for real RO profiles.
- Lines 83-113: Lower-tropospheric GNSS-RO challenges. The GNSS-RO motivation should introduce super-refraction, multipath, horizontal gradients, critical refraction, vertical smoothing, and lower-tropospheric N biases earlier. These are not secondary details; they directly determine whether the proposed N-to-T/W retrieval can work for real RO data.
- Lines 114-152 and 203-210: Prior information. The manuscript should not frame the approach as prior-free. The relevant comparison is not between an NWP prior and no prior. It is an NWP model prior versus an empirical radiosonde sampling prior. Both are priors, each with error structures and representativeness limitations.
- Lines 169-173: Radiosonde reference statement. The statement that colocated radiosondes are superior to any NWP analysis or reanalysis at the same footprint is too strong. Radiosondes are traceable in situ measurements, but they are point observations with drift, sensor errors, and representativeness limitations. Please revise to a more defensible statement.
- Lines 181-190: Claim of largest independent assessment. The claim that this is the largest independent in situ assessment should be substantiated with a short comparison to prior studies or softened. “To our knowledge” is not sufficient unless the relevant literature is adequately covered.
- Lines 266-279: WCTN input and 150 m dilation. The wavelet-transform input is central to the method. The rationale for the 150 m dilation should be given in the main text, and the authors should show performance without WCTN and with alternative dilation scales. Otherwise, it is hard to know whether this design is physically necessary or simply a tuned feature representation.
- Lines 284-298 and 508-510: Vertical grid definition. The vertical grid appears inconsistent. The text refers to 100 m to 20 km, 10 m spacing, and 2,000 vertical elements. Please define the exact lower boundary, upper boundary, spacing, endpoint convention, and number of levels.
- Lines 293-298 and 861-887: Intermediate targets. The choice of dry refractivity and dry-air partial pressure as intermediate targets may be useful, but this has not been proven. Because these variables are not independently observed components of total N, their retrieval still depends on learned prior information. Please add ablation experiments against direct T/W/RH prediction and against a single multi-output network.
- Lines 315-323: Validation wording. Because input N and target T/W are derived from the same radiosonde ascent, this is not validation against independent observations of the same atmospheric state. Please use “evaluation under internally consistent ideal-input conditions” or similar wording unless independent real-RO validation is shown.
- Lines 324-413 and 493-514: QC reproducibility. The data description is extensive, but the QC implementation is not fully reproducible. Please provide profile counts before and after each QC step, RH clipping counts, altitude-threshold exclusion counts, and objective criteria for visually identified anomalies.
- Lines 414-419: Island radiosonde representativeness. A small island should not be described as a large ship permanently anchored in the ocean. Island boundary layers, terrain, surface heating, and local circulations may affect low-level profiles. Please quantify whether errors change when the lowest 0.5-1 km is excluded.
- Lines 474-477: Geographic independence. Manus and Nauru are not geographically independent if they appear in both the training and TOGA-COARE evaluation sets. Please report metrics excluding these locations or describe the test as temporally/instrumentally independent rather than geographically independent.
- Lines 523-565 and 1080-1120: Reproducibility of architecture and training. The main manuscript should include sufficient information for scientific interpretation: the number of parameters, normalization, loss weights, optimizer, learning-rate schedule, train/validation/test splits, early stopping, random seeds, and hyperparameter ranges. Code, trained weights, preprocessing scripts, and evaluation scripts should be archived in a persistent repository with a DOI.
- Lines 563-565 and 672-687: Confidence intervals and dependence. The confidence intervals appear to assume independent profiles. This is likely unrealistic for repeated soundings from the same stations, campaigns, seasons, and synoptic regimes. Please recompute uncertainty using station/campaign/day block bootstrap or a comparable dependence-aware method.
- Lines 577-585: Need for baseline models. A similar RMSE between the training and evaluation data is encouraging, but it does not prove that the model learned a general inversion. Please compare against altitude-only climatology, season/basin climatology, multilinear regression, MLP/random forest, and direct-output CNN baselines.
- Lines 605-618: Independent observational products. The sentence that all model retrievals constitute independent observational products is too strong. They are independent of NWP input during inference, but not independent of the radiosonde training prior. They are also not yet validated against real satellite-observed refractivity.
- Lines 624-658 and 688-710: Temperature and moisture interpretation. The interpretation of retrieval skill should be strengthened. Please explain why temperature retrievals below 20 km are not better and identify altitude ranges where N contains useful independent information versus ranges where the output is dominated by the learned empirical prior.
- Lines 728-740: Regime-dependent statistics. The statement that differences are less than 5% under deep-convective and dry conditions should be quantified rather than inferred primarily from visual inspection. Please provide statistics for moist, dry, convective, suppressed, day/night, boundary-layer, and free-tropospheric regimes.
- Lines 741-744 and 995-1004: DYNAMO temporal overlap. The temporal overlap issue should be handled with a sensitivity test. Please exclude all DYNAMO evaluation profiles from headline statistics or report results with and without them.
- Lines 745-781: Perturbation experiment. The perturbation experiment is useful, but it should not be interpreted as evidence of 10 m resolving power. A two-grid-level RH spike is an idealized diagnostic. Please repeat the test after RO-like smoothing and realistic N perturbations.
- Lines 782-819: What the model learns. This is one of the stronger parts of the manuscript because it recognizes that the model learns training-distribution statistics rather than explicit physical laws. Please connect this discussion more directly to limitations in retrieving sharp inversions, boundary-layer caps, and low-level humidity structures.
- Lines 820-860: Robustness claims. Cross-radiosonde and cross-campaign generalization are useful, but they are not equivalent to transferability to satellite RO observations. Please moderate claims of spatial, temporal, and instrument robustness.
- Lines 888-895 and 1019-1025: NWP independence. The paper should use “NWP-independent during inference” rather than “genuine NWP independence.” The model inherits radiosonde sampling, processing, QC, station distribution, and climate-regime priors.
- Lines 950-972 and 1067-1073: Platform-agnostic statement. The architecture may accept input profiles from different sources, but performance is source-dependent. GNSS-RO profiles have different vertical resolution, horizontal averaging, bias, noise, and missing-data characteristics. Please revise the platform-agnostic wording.
Specific comments
- Lines 16-40: Add one sentence in the abstract stating that real GNSS-RO validation is not included in this Part 1 study.
- Lines 18-19: Use consistent hyphenation: “one-dimensional” and “model-dependent.”
- Lines 23-24: Use “high-resolution” as a compound adjective.
- Lines 23-31: Define whether the reported radiosonde uncertainty is standard uncertainty or expanded uncertainty, and do not imply that RMSE equals measurement uncertainty.
- Lines 33-40: Add a sentence such as: “Part 2 will determine how this performance changes for actual GNSS-RO inputs.”
- Line 83: Capitalize as “Global Navigation Satellite System.”
- Lines 88-90: Consider changing “of the order of a few hundred meters” to “on the order of a few hundred meters.”
- Line 106: Check the spelling and capitalization of “PlanetIQ” and the cited author name.
- Lines 114-144: Revise the 1D-Var discussion so that the contrast is not NWP prior versus no prior, but NWP-model prior versus empirical radiosonde prior.
- Lines 145-152 and 203-210: Replace “prior-free” language with “NWP-independent during inference” or “constrained by an empirical radiosonde prior.”
- Lines 169-173: Revise the radiosonde-superiority statement and acknowledge representativeness limitations.
- Lines 181-190: Substantiate or soften the “largest independent in situ assessment” claim.
- Line 199: Insert a space in “domain.The.”
- Lines 207-210: Revise the sentence beginning “The approach seeks...” for grammar and clarity; it may be better as two sentences.
- Line 247: Change “The Part 2” to “Part 2.”
- Lines 249-263: Move the main limitation of ideal-input testing into the abstract and conclusions.
- Line 255: Avoid “etc.” in a formal methods section; list the factors or write “among others.”
- Lines 266-279: Give the rationale for 150 m WCTN dilation and show sensitivity to the dilation scale.
- Lines 284-298: Define the exact vertical grid and resolve the apparent inconsistency in the number of levels.
- Lines 286-287: Use consistent section-reference formatting: “Sect. 2.3, Sect. 2.4, and Sect. S1.”
- Lines 293-298: Do not imply that Nd and Pd are straightforward targets simply because they are linearly related to N; their decomposition still depends on a learned prior.
- Lines 307-313: Please check Eq. (8). The denominator inside the Haar wavelet argument appears to use h rather than the dilation a.
- Line 311: Correct “and and Zb” to “and Zb.”
- Lines 315-323: Avoid using “validation” without qualification because the input and output come from the same radiosonde ascent.
- Line 316: Insert a space in “RHprovided.”
- Lines 324-413: Provide station-by-station counts before and after QC and document objective criteria for visual anomaly removal.
- Lines 361-365: The text mixes standard and expanded uncertainty. Please define k = 1 and k = 2 consistently.
- Lines 414-419: Remove or revise the “large ship” analogy; quantify island effects in the lowest 0.5-1 km.
- Lines 474-477: Do not describe Manus and Nauru as geographically independent if they are included in both training and evaluation locations.
- Lines 493-514: Quantify sensitivity to RH clipping, gap interpolation, truncation below 100 m, and exclusion of profiles not reaching 20 km.
- Lines 511-514: Change “Unrealistically high positive (negative)” to “Unrealistically high positive (negative) RH values”; change “was small” to “were small.”
- Lines 523-565: Add number of trainable parameters, normalization strategy, train/validation split, early stopping criteria, optimizer, loss weights, random seeds, and hyperparameter ranges.
- Lines 532-543: Resolve the inconsistency between 0-20 km and 100 m-20 km output ranges.
- Lines 563-565 and 672-687: Recompute confidence intervals using a station/campaign/day block bootstrap or another method that accounts for dependence.
- Lines 577-585: Add baseline models to quantify the actual AI contribution.
- Lines 605-618: Revise “independent observational products”; the retrievals are not independent of the prior radiosonde training.
- Lines 624-658: Add a compact main-text table showing RMSE and bias for 0.1-2, 2-5, 5-10, 10-15, and 15-20 km for each evaluation dataset.
- Lines 688-710: Bring a concise station-by-station summary into the main text, including the best and worst sites and possible dependence on basin, season, sonde type, and local time.
- Line 704: Clarify the meaning of “~15% RMSE 10-12 km.”
- Lines 728-740: Quantify the less-than-5% statement with regime-conditioned statistics instead of relying on visual inspection.
- Lines 741-742: Change to “It should be noted, however, that a temporal overlap exists...”
- Lines 741-744 and 995-1004: Support the conclusion about temporal overlap with a sensitivity test excluding DYNAMO evaluation profiles.
- Lines 745-781: Repeat the perturbation experiment after RO-like vertical smoothing and realistic refractivity noise.
- Lines 782-819: State more clearly that sharp real thermal and moisture structures may be smoothed by the learned prior.
- Lines 820-860: Tone down claims of robustness and distinguish radiosonde generalization from satellite RO transferability.
- Lines 861-887: Add an ablation test against direct prediction of T/W/RH and against a single multi-output network.
- Lines 888-895 and 1019-1025: Use “NWP-independent during inference” consistently and avoid “genuine NWP independence.”
- Line 929: The phrase “behavioral property of the model” is acceptable but should be connected to retrieval information content.
- Lines 950-972: The limitations section should include missing baselines, no per-profile uncertainty, no real-RO validation, correlated evaluation data, subjective QC, and incomplete code/model availability.
- Lines 1000-1004: The statement about temporal overlap not conferring a measurable advantage appears repeated; remove one occurrence.
- Lines 1021-1022: Use “NWP-derived” consistently.
- Lines 1067-1073: Revise the platform-agnostic statement because performance is source-dependent.
- Line 1069: Use “general-purpose” as a compound adjective.
- Lines 1080-1120: Strengthen the code and data availability statement. Trained weights, preprocessing scripts, architecture, normalization constants, and evaluation scripts should be placed in a DOI repository.
- Line 1131: Change “reference radio soundings” to “reference radiosonde soundings.”
Final statement
Overall, I find this manuscript interesting, relevant to AMT, and potentially publishable after major revision. The key value of the paper lies in its attempt to learn an N-to-T/W retrieval relationship at distances below 20 km without using an NWP background during inference. However, the present evidence supports only an ideal-input framework characterization using radiosonde-derived refractivity. It does not yet demonstrate that the method works for real GNSS-RO refractivity profiles. The authors should revise the claims, add baseline and ablation experiments, improve the treatment of uncertainty, explain the information content and limitations of temperature retrieval below 20 km, and strengthen reproducibility. If these issues are addressed, the paper could make a meaningful contribution to atmospheric retrieval methodology and to the development of future GNSS-RO profile retrieval.
Citation: https://doi.org/10.5194/egusphere-2026-2027-RC2
Viewed
| HTML | XML | Total | Supplement | BibTeX | EndNote | |
|---|---|---|---|---|---|---|
| 110 | 55 | 8 | 173 | 32 | 7 | 9 |
- HTML: 110
- PDF: 55
- XML: 8
- Total: 173
- Supplement: 32
- BibTeX: 7
- EndNote: 9
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
The manuscript by Muralidharan et al. presents a neural-network-based retrieval framework for deriving thermodynamic profiles from GNSS-RO refractivity observations over tropical oceanic environments. The topic is relevant for data assimilation and for demonstrating the potential impact of GNSS-RO measurements within the global climate observing system, as well as for assessing the value of upper-air reference observations as a collateral outcome of their investigation. The proposed deep-learning methodology itself is not fundamentally novel, as similar machine-learning approaches have already been explored for a variety of atmospheric retrieval and inversion problems. Nevertheless, its application to GNSS-RO retrievals and upper-air observations remains potentially valuable and the study hasmay have the potential to provide insights that extend beyond the machine-learning methodology itself. However, the current experimental design does not yet fully exploit this opportunity.
A key limitation is the absence of a rigorous comparison with established 1D-Var GNSS-RO retrieval frameworks, which remain the standard physically constrained approach for combining observations with prior information. Without such a benchmark, it is not possible to assess whether the proposed method provides a genuine improvement over existing methodologies in terms of retrieval accuracy, vertical structure representation, or uncertainty characteristics.
Another key limitation is the restricted geographical representativeness of the validation framework. The analysis relies primarily on island-based tropical stations, which may not adequately sample the variability of open-ocean atmospheric conditions. In particular, boundary-layer processes and moist convection regimes in the tropical Atlantic and broader oceanic regions are not explicitly tested. This raises concerns about the claimed generalization to “tropical ocean environments”.
The study does not also convincingly demonstrate true spatial or instrumental generalization beyond a relatively constrained set of island-based and campaign-oriented radiosonde observations. Although the use of multiple datasets spanning different periods is presented as a strength, in practice the evaluation framework mixes heterogeneous observing systems without a systematic stratification of measurement quality, correction level, or representativeness errors.
In particular, the role of GRUAN observations as a reference standard is not fully exploited or critically assessed. While GRUAN is appropriately treated as a high-quality dataset, the analysis does not isolate the specific benefit of using GRUAN versus non-GRUAN radiosondes, nor does it quantify how mixed-quality observations may influence both training and evaluation outcomes.
Moreover, the inclusion of historical datasets (e.g., TOGA-COARE) together with modern operational radiosondes introduces potential inhomogeneities due to known long-term changes in radiosonde instrumentation and correction practices (e.g., solar radiation bias, humidity sensor dry bias). These issues are not explicitly addressed through homogenization or sensitivity analysis.
Therefore, the manuscript contains a promising methodological contribution and is potentially suitable for publication in AMT. However, substantial additional analysis is required before the conclusions can be considered robust.
Specific comments are provided below.
Specific Comments
- retrieval accuracy,
- sensitivity to prior information,
- vertical structure representation,
- and robustness across regimes.
In this first manuscript (where the authors note that a second paper will include real GNSS-RO measurements), it is essential to provide statistical comparisons with GNSS-RO data produced by the authors for the exact same region analyzed, using a suitable collocation criteria. Without this, the claimed advantages of the neural network approach cannot be verified.