the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
Aircraft Sampling Representativeness for LASSO Domain Averages over the ARM Southern Great Plains (SGP) Site
Abstract. Aircraft observations are widely used to evaluate cloud simulations; however, their inherently localized sampling may not fully represent the domain-scale cloud properties resolved by large-eddy simulations (LES). This study quantifies how aircraft-like sampling strategies influence the representativeness of cloud water content (CWC) derived from LES output. Using LASSO simulations for the ARM Southern Great Plains site (SGP), we emulate aircraft trajectories by prescribing multiple horizontal flight patterns (circle, diagonal, straight, combo, and zigzag) and vertical profiles (flat, sine wave, staircase, and zigzag) under different speed categories. Results show that sampling geometry substantially influences the degree to which aircraft-like measurements capture the magnitude and variability of domain-scale CWC. Sampling representativeness depends jointly on flight speed and trajectory geometry. For the flat, sine wave, and staircase vertical profiles, increasing flight speed generally improves agreement with the layer-restricted domain mean by increasing sampling density and reducing sensitivity to localized heterogeneity. In contrast, the zigzag pattern does not show systematic improvement with speed because its phase-locked vertical–horizontal coupling constrains effective spatial coverage. Cloud-boundary analysis further indicates the tendency to underestimate cloud-top height, while cloud-base height is more consistently captured, with higher speeds reducing both bias and variability. These findings establish a systematic framework for interpreting aircraft-based cloud measurements in the context of LES output and provide practical guidance for designing flight strategies that improve sampling representativeness in model–observation comparisons.
- Preprint
(3610 KB) - Metadata XML
-
Supplement
(2449 KB) - BibTeX
- EndNote
Status: final response (author comments only)
- RC1: 'Comment on egusphere-2026-2786', Anonymous Referee #1, 25 Jul 2026
-
RC2: 'Comment on egusphere-2026-2786', Anonymous Referee #2, 03 Aug 2026
Review of 'Aircraft Sampling Representativeness for LASSO Domain Averages over the ARM Southern Great Plains (SGP) Site'
This manuscript uses LASSO large-eddy simulations over the ARM SGP site to ask how well aircraft-like sampling reproduces domain-average cloud water content (CWC). Synthetic trajectories built from five horizontal patterns (four derived from HI-SCALE flight tracks, plus a zigzag), four vertical profiles, and three speed classes are flown through the model fields, and the sampled CWC is compared with layer-restricted domain means through regression slopes, correlations, and normalized error metrics. The authors conclude that sampling representativeness is governed mainly by vertical trajectory design and flight speed, with faster flight generally improving the agreement with the layer-restricted domain mean and with horizontal geometry playing a secondary role.
My questions center on interpretation and reporting rather than on the design itself. Most of them trace to one difficulty: as written, it is hard to tell how much of the headline findings reflects the physics of aircraft sampling and how much follows from choices made in constructing the emulation. I recommend minor revision. Reviewer 1 has already covered the realism of the flight profiles, the organization of the text, and the case for deeper analysis; I have tried not to repeat those points and note below where I second them.
Main Comments
Sect. 2.1.3 says 'a subset of the simulation library has been used in this study to make the processing manageable' (L118). Table S1 in the Supplement does list the analyzed case days, but the main text never states this number, and I could not find the criteria used to select the subset or any note on whether the selection could influence the results; the only rationale given concerns the choice among forcing options rather than the choice of cases. Could the authors state the case and ensemble-member counts in Sect. 2.1.3, explain how the subset was chosen, and comment on whether the results are expected to be sensitive to that choice?
Two related details would also help: are the 2015–2016 cases taken from Alpha2 or from Version 1, and could the Version 1 domain size (25.5 km, currently mentioned only in passing at L181–182) be stated in Sect. 2.1.2? On the sampling parameters, Sect. 3.3 gives characteristic speed ranges for the three aircraft classes (L276–283), but unless I missed it, the specific airspeed values used to construct the slow, moderate, and fast trajectories are never stated; please provide them.
Please also define the threshold on qc (or CWC) that identifies cloud base and cloud top along the track, state how the domain-mean cloud top and base in Fig. 7 are computed (over cloudy columns only?), and clarify what happens to the ±100 m offsets when the local cloud depth is under 200 m, in which case base +100 m would lie above top −100 m. Extending Table S1 with the ensemble members and the release for each case would cover much of this.
The description at L240–243, where each 10-min output is repeated at 1-s resolution, together with the fixed 6-h window sampled at 1 Hz (L290–292), left me unsure what actually differs between the speed categories. If every trajectory collects one sample per second for six hours, the sample count is the same at every speed; what increases with speed is the distance covered, i.e. the fraction of the domain seen. So when the results attribute the improvement to 'sampling density' (L341, L389–390), is it not more accurate to say that a faster aircraft covers more of a snapshot that is frozen for 10 minutes at a time, which makes convergence toward that snapshot's domain mean almost guaranteed? I second Reviewer 1's request for a clearer description of this expansion step. Beyond the description, I would suggest tuning down the stated benefit of the higher speed categories. What limits fast platforms in reality is that the cloud field evolves while the aircraft transits, with shallow cumulus turning over on timescales of tens of minutes, and that effect cannot operate in this emulation.
One more question here: L361–362 gives zigzag sample counts rising from 19410 (slow) to 27617 (fast). At 1 Hz over a fixed 6 h that cannot be the count for a single trajectory. Could the authors state what is being counted (in-cloud samples only? Or pooled over cases?)?
In Table 1, every nBias entry is positive: 0.49–0.67 for the flat, sine wave, and staircase patterns, up to 1.27 for zigzag, and 1.38–1.51 for the flat pattern in the profile comparison of Sect. 4.3. The text calls this 'limited systematic deviation' (L354), but a sampled average running 50–150% above the reference would be a substantial systematic offset. Furthermore, the vertical profiles are selected to stay between cloud base +100 m and cloud top −100 m, the example track is 'constructed to deliberately cross regions of enhanced CF' (L301–302), and the reference is a layer mean over the full domain including clear air, so the trajectories conditionally sample cloudy air while the reference does not. Therefore, I echo Reviewer 1's question about why the all-sky domain mean is the appropriate comparison target, and would take it one step further. Could the authors compute the same metrics against an in-cloud conditional domain mean over the same layers? All required fields appear to be in hand, and the comparison would separate the cloud-conditional offset from genuine trajectory effects. It would also show whether the bias mainly comes from comparing in-cloud samples with all-sky averages. If it does, no flight strategy can remove it, and the remedy lies in how the comparison statistic is constructed rather than in how the aircraft flies.
How is the quantity labeled nRMSE in Table 1 calculated? Most entries are negative (e.g. −0.11), which no standard RMSE definition allows, yet the same metric is positive in Sect. 4.3 (1.32–2.05). Could you please give the formulas for nBias and nRMSE, including the normalization, and apply them consistently between the two sections? Is the identical zigzag moderate and fast column (0.37, 0.20, 1.27 in all three metrics, to two decimals) a copy error, or do the two speeds genuinely produce the same statistics despite different sample counts?
Could the correlation values quoted in Sect. 4.1 also be checked against Table 1? The text says r ranges 'from about 0.69 to 0.88' (L348–349), while Table 1 contains several values of 0.60–0.66, and 'the lowest correlation occurs for the diagonal pattern at slow speed (r ≈ 0.72–0.73)' (L353) does not match the table either (diagonal-slow is 0.74 for flat and 0.60 for sine wave and staircase). Last, L408–409 attaches a single p-value (p = 0.032) to three different correlation coefficients. Can the authors confirm what this p-value refers to, report the three values separately, and note how the effective sample size was handled given that the profile points are adjacent 30-m model levels with strong vertical autocorrelation?
Minor Comments
L73–74 states the study aims to quantify sampling of 'key cloud and aerosol properties', but only CWC is analyzed; other cloud and aerosol properties are deferred to future work (L481–483). I echo Reviewer 1's comment on the title here: the aims statement should match the delivered scope.
Could an explicit formula for CWC be given in Sect. 2.3 (CWC = ρqc?). And does the CWC sampled from the model include only cloud water, or also mass that the Thompson scheme classifies as rain? Since in-situ probes on a real aircraft would count drizzle-sized drops as part of the measured liquid water, whether qr is included affects how the model-sampled CWC maps onto what an aircraft would report (e.g., King Probe), and one clarifying sentence would help.
Sect. 2.2 gives the horizontal coverage of the aircraft sampling as around 100 km (L143–144), yet the LES domains are only 14.4 and 25.5 km across. Since the sampling paths are idealized, how were they derived from the observed latitude–longitude tracks, if at all: truncated to the domain, rescaled to fit it, or redrawn from scratch? This affects how literally the patterns in Fig. 1b–f can be read as HI-SCALE geometries. Please clarify.
L181–182 notes that horizontal sampling is performed separately for the two domain sizes. Are the statistics in Table 1 and Fig. 6 then pooled across both domains (and both LASSO releases)? Representing a 25.5 km domain from a localized track is presumably harder than a 14.4 km one, so I wonder whether domain size could confound the pattern and speed comparisons; a sentence on this, or numbers stratified by domain, would help.
Could the regression intercept be reported alongside the slope for Fig. 5 and the configurations in Fig. 6? Unless I missed it, neither the text nor the figures give the intercept, and with slopes near unity and uniformly positive nBias the offset must reside there, so reporting it would connect Figs. 5–6 with Table 1. Please also state whether the scatter points come from all cases and ensemble members.
In the paragraph on Fig. 6 (L342–346), it is not always clear which vertical profile the horizontal-pattern statements refer to; for example, 'for the circle horizontal pattern, the slope value is greater than 0.94' presumably applies to the staircase row only. Could each quoted number be anchored to its row in the figure?
The case grouping by cloud fraction (L457–464) appears for the first time in the conclusion, as a finding ('no systematic dependence of sampling performance on cloud amount'), with the supporting material only in Fig. S2. Since Fig. S2 shows scatter plots without per-category metrics, could this analysis be presented in Sect. 4, with slope or r for each CF bin? The statement that differences among CF categories 'primarily reflect variations in the number of available cases' is also hard to assess without those case counts.
Technical Comments
At L279–281 the C-130 appears twice in the moderate-speed category ('C-130/LC-130' and 'Lockheed C-130'). And is the BAe-146 correctly placed in the 200–250 m s−1 class? Please correct me if I am wrong: from what I can find online, its indicated airspeed during sampling is about 100 m s−1, corresponding to a true airspeed of roughly 100–120 m s−1 (Sanchez-Marroquin et al., 2019, https://doi.org/10.5194/amt-12-5741-2019), which would place it with the moderate-speed platforms.
The text at L403–404 says Fig. 8 shows the 'flat, sine wave, and zigzag' vertical patterns; the caption says 'flat, sine wave, and staircase'. The following paragraphs then discuss all four. Please reconcile which patterns are plotted.
Citation: https://doi.org/10.5194/egusphere-2026-2786-RC2
Viewed
| HTML | XML | Total | Supplement | BibTeX | EndNote | |
|---|---|---|---|---|---|---|
| 39 | 34 | 8 | 81 | 10 | 3 | 4 |
- HTML: 39
- PDF: 34
- XML: 8
- Total: 81
- Supplement: 10
- BibTeX: 3
- EndNote: 4
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
Review for “Aircraft Sampling Representativeness for LASSO Domain Averages over the ARM Southern Great Plains (SGP) site” (egusphere-2026-2786) submitted by Liang et al.
Summary:
The manuscript takes up an important challenge – how well do models represent in-situ aircraft sampling of cloud water content (chosen as a variable to represent cloud properties) if the aircraft sampling strategies are represented in selecting model data as opposed to using domain averages.
This is an important topic, and the manuscript tackles this problem with a thoughtful methodology. Nevertheless, I felt the manuscript was missing many pieces and the text could benefit from better organization. Often I found key information was missing, or the authors made statements/arguments before introducing the details that could potentially support those statements/arguments (some of my comments below get answered partially by the text but much after they arise). Specifically, improvements can be made to the description of the flight tracks, a better link can be established with real-world aircraft observations, and issues associated with the “flat profile” type should be addressed, etc. Finally, the authors have a wealth of information given the work emulate "flight tracks" within model output - many details lost within aircraft observations can be inferred here to really do a more comprehensive evaluation of not just the model output relative to domain averages, but also the aircraft sampling itself.
I recommend major revisions before the manuscript is reconsidered with my detailed comments below.
Detailed comments:
Title: The manuscript title suggests the work is broader than it is, especially given the length of the results section which has essentially three figures, two of which examine CWC. The abstract represents this directly in Line 10, and I suggest the title be reworded to mention “CWC” specifically, or similar.
Line 29: I believe the authors can do a better job motivating the need for observations and simulations to be combined to answer cloud-related questions in ESMs. This sentence comes a bit abruptly.
Line 35: This should be re-phrased. Active remote sensing does penetrate clouds and can often get comparable vertical resolution.
Line 43: I recommend citing this recent article that summarizes the latest AAF campaigns and a dataset from ARM AAF - https://essd.copernicus.org/articles/16/5429/2024/.
Line 60: “These aircraft measurements” is confusing, the AAF deployment during HI-SCALE should be introduced first?
Line 84: The definition of WRF should be moved to line 78.
Line 120: This citation refers to ERA-Interim – is that the version of the reanalysis that was used here? This should be mentioned here in the text explicitly and the variable(s) used should be provided or the methodology, if novel, should be addressed here directly.
Line 128: typo for “hydrometeor”.
Line 129: add references?
Line 140: I’m a bit confused here. How can there be consistency between the datasets when the LASSO cases used are not overlapping with HI-SCALE aircraft observations according to lines 67-69: “The LASSO simulations analyzed in this study correspond to different cases aside from those sampled during HI-SCALE and are not directly constrained by the HI-SCALE aircraft observations.”
Section 2.2: Some of this information belongs in the introduction section when the campaign is discussed (or at least this section should be referred to over there).
Line 150: Are these the only parameters used from the AAF dataset? If so, this should be specified much sooner. If not, this sentence should be edited.
Line 157: I’m not sure how CWC directly provides an estimate of cloud thickness, perhaps the initial part of the sentence can be updated to “CWC and its vertical profile/integral…”?
Line 162: I like this figure and the corresponding discussion. However, it is missing any reference to the height of the aircraft during these legs. Are these constant altitude legs or is the aircraft altitude changing? If so, how much? If only the lat-lon are used from these patterns and CWC is examined at multiple model levels, that should be mentioned.
Fig 1 and 2: The font sizes for the labels and legends need to be increased for accessibility.
Line 200: I’m not sure if “flat” refers to the altitude being constant or something else? The term “flat profile” on Line 214 seems confusing.
Line 227: Should it be “horizontal tracks”?
Line 238: This detail needs to be described much sooner in the manuscript.
Line 240: “As the LASSO simulations archive output at 10 min intervals, model fields were temporally expanded to 1 s resolution by repeating the corresponding 10 min output (mean fields) across each interval.” I apologize but I’m struggling to understand how this works. A better description or schematic may be needed here as this is critical to the study. The subsequent discussion of sampling density as a function of aircraft speed then follows this confusion.
Line 245: This ties back to the need for using only the 6-12h model outputs as this sentence suggests that temporal constraint wouldn’t be necessary.
Section 3.1 and 3.2: A lot of the text describes the sampling strategies – however one key element is missing – how much time does the “aircraft” spend in-cloud during these legs? Or what does the model cloud field need to look like along the “aircraft track” for the case to be used in this study? Is that considered (or should it be considered) during case selection or during flight strategy determination? I’m assuming this is accounted for when the authors remove instances of cloud edge, but it should be properly discussed here in the context of the flight tracks.
Line 248: The first type of “profile” has major issues. Semantically speaking, this is not “flat” and the terminology is misleading as “flat profile” is an oxymoron. Additionally, this is an impossible flight track in the real world, judging by Fig. 3a, this implies the “aircraft” instantaneously rises or drops up to 500 m? While ascending/descending, an aircraft will still propagate horizontally, and any “flat” profile will look closer to profile types shown in panels (b) and (c) depending on the pitch angle. I understand the authors’ intent here, but if they want “aircraft sampling representativeness”, the profile in panel (a) is extremely unrealistic.
Line 258: I would assume this is the case in most clouds even with extreme entrainment. I’m not sure if the authors meant to state this as a result derived from their figure?
Line 262: “As the aircraft ascends or descends through a cloud, it samples a coherent cloud column over a short horizontal distance”, I don’t believe think this statement is accurate. The horizontal air speed of the aircraft would typically be greater than the ascent/descent rate. I would suggest the authors follow such arguments (throughout the manuscript) with an example from actual or representative horizontal and vertical speeds of the G1 aircraft from HI-SCALE.
Line 264: While the authors don’t intend to compare exact profiles, this argument deserves the use of actual aircraft observations since those data are readily available.
Figure 3: The figure is discussed already but even by Line 275, there is no description of the June 9, 2015 case that is being plotted. I’m not sure why or how that day is chosen or what it represents. I’m not entirely sure if this is an actual flight or model simulation and the tracks being based on a flight?
Section 3.3: This section represents the type of considerations to be made, and it is a nice addition. However, it feels incomplete because it ties partly to the temporal resolution and somewhat to the model domain size but does not relate the potential (mis)match with the model grid spacing. If the model grid spacing is 100 m and the aircraft speed is 100 m s-1, a temporal resolution of 1 s directly matches the aircraft measurements. But otherwise, there is a mismatch, and a proper, direct comparison becomes unlikely.
Line 301: While it is fine to construct a flight track that crosses regions of enhanced CF, this would be improbably to replicate in real life during an aircraft campaign. How would flight planning work to ensure enhanced CF along the track unless the forecasts were highly accurate or CF was high across the domain? “Aircraft sampling representativeness” needs to be thought of as a real-time scenario. Otherwise, the authors are creating an artificial selection bias by selecting tracks based on known model output. Would it not be more strictly representative if the flight tracks were randomly assigned to model case days? I’m not sure how one answers this question but I’m curious to see the authors comments.
Line 315: Is the comparison with domain means conducted because those are the values typically used for model-observational comparisons? The discussion would benefit from some references or examples here to illustrate why this specific comparison is conducted.
Figure 5: This is an interesting figure.
Figure 6: This is a very nice way of succinctly providing essentially a sensitivity test for different sampling profiles and aircraft speeds!
Table 1: The average and std. dev. Values for the model and aircraft sampled CWC should be provided for each scenario.
Figure 8 – with the amount of noise in the CWC values as a function of altitude, I’m wondering if the authors tried looking at CWC as a function of height above cloud base or normalized height above cloud base? Where the latter is defined as Z – CTH / CTH – CBH. Surely the advantage of using model data is that they can retrieve cloud base heights for all “aircraft” transects even if the aircraft did not sample the cloud base. This is one major limiting factor in examining CWC profiles when looking at real-world aircraft observations. I suspect their correlations will drastically improve.
Line 436: I don’t mean to sound harsh or trivialize the authors’ work, but it seems a lot more could have been done here with the model and “flight track” setup. Did the authors look to examine other cloud properties? One gets the feeling there’s just not enough analysis here yet to draw meaningful conclusions. I’m hesitant to recommend publication of a revised manuscript without more detailed assessment of the dataset. I recommend investigating liquid water path and adding process rates that aircraft observations can lack. The authors have an opportunity to answer questions such as “what are aircraft observations missing” and “how can models inform better aircraft sampling strategies”, etc. The lack of any real aircraft comparisons or such insights essentially boils the work down to comparisons between model domain average CWC and CWC simulated along a few different categories of “line tracks” within the model.
Line 463: “Additional factors such as trajectory scale and orientation relative to the mean wind may also modulate sampling performance through their interaction with cloud morphology.” Similar to my previous comment – the authors should have these data from the model simulation and this can be examined instead of having a speculative statement.