the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
Observing formation and early evolution of contrails formed by IAGOS aircraft using high-resolution LEO satellite imagery
Abstract. Persistent contrails and contrail cirrus are estimated to be a major contributor to the climate impact of aviation. The mitigation of these impacts by means of technological or operational changes benefits from the ability to skillfully model the formation, evolution, and impacts of contrails. Although these models can be evaluated and improved by use of observations of contrails obtained from remote sensing instruments, these comparisons are hindered by uncertainty in the required meteorological data (such as relative humidity) and limitations in the method of observation (such as younger contrails not being observable in geostationary satellite imagery). To address these challenges, we collocate aircraft equipped with in-situ humidity sensors from the IAGOS fleet in high-resolution (10–30 m) satellite imagery obtained by instruments aboard the low Earth orbit Sentinel-2 and Landsat missions. The resulting dataset consists of 543 IAGOS aircraft found in satellite imagery (51 % of which form contrails), which we use to evaluate predictions of contrail formation by the Schmidt-Appleman criterion (SAC) as well as predictions of contrail growth by the CoCiP model. When accounting for uncertainty in the IAGOS measurements of humidity and temperature, we find that the SAC correctly explains 98.3 % of the observations. Disagreement between predictions and observation increases when using meteorological data from the ERA5 reanalysis, with only 92.1 % of the observations being explained correctly. Out of the 195 annotated contrails, 48.2 % of these contrails were found to persist for longer than 10 s (approximately the jet phase) and 8.7 % longer than 120 s (approximately the vortex phase). The relative humidity with respect to ice is found to correlate most strongly with observed contrail lifetime, exhibiting an R2 value of 0.49 with the logarithm of contrail age. The observed horizontal growth during the jet and vortex phases is consistent with previous observations and contrail model results. Although the limited lifetimes of the annotated contrails prevent robust statistical conclusions for the dispersion phase, three example cases show horizontal growth rates consistent with simulations by CoCiP and that of observations in literature. Overall, this study demonstrates the potential of high-resolution LEO satellites to create observational datasets for evaluating and improving models of contrail formation and early evolution.
- Preprint
(24868 KB) - Metadata XML
- BibTeX
- EndNote
Status: final response (author comments only)
-
RC1: 'Comment on egusphere-2026-1171', Anonymous Referee #1, 14 Jul 2026
-
RC3: 'Reply on RC1', Anonymous Referee #2, 17 Jul 2026
Dear RC1,
I read your comprehensive review with great interest. While I agree with several of your technical points, I would like to offer a different perspective regarding your general comment that checking the Schmidt-Appleman Criterion (SAC) is no longer a very important topic since it was established years ago. I also want to share some observations regarding your suggestion to include Vazquez-Navarro et al. (2010) and the associated detection methods.
- The Relevance of Evaluating the Contrail Formation Pipeline
While the fundamental thermodynamics of the SAC are indeed well-established, I must respectfully disagree that checking its operational application is unnecessary. This manuscript is doing much more than just testing the classic SAC equation; it is evaluating our entire contrail formation prediction pipeline.
Crucially, the authors are testing the accuracy of the underlying ERA5 reanalysis model against a real-world counterfactual provided by the in-situ IAGOS measurements. As we move toward operational contrail avoidance, the community relies heavily on numerical weather prediction models like ERA5 to forecast supersaturated regions. Demonstrating how often this meteorological input leads to false positives or false negatives when fed into the SAC, and proving that in-situ corrections significantly improve these predictions, is highly relevant.
Furthermore, as engine architectures evolve toward modern lean-burn and ultra-high bypass ratio designs, their altered microphysical emissions fundamentally change the applicability of the SAC. As demonstrated by recent microphysical modeling (Ponsonby et al., 2025) and flight campaign measurements (Voigt et al., 2021), soot-poor emissions from newer engines fail to effectively scavenge gaseous precursors. Consequently, droplet activation relies on highly volatile particles, which pushes the activation threshold to much colder temperatures (changes T_SAC) and alters the active ice crystal emission index (AEI_ice). High-fidelity simulations paired with ground-based observations (Quibén Figueroa et al., 2025) further corroborate that the classical SAC struggles to accurately predict formation in these soot-depleted regimes. Therefore, continuously validating our predictive pipelines across these evolving aircraft parameters remains a mandatory and ongoing challenge for the community.
- Applicability of Mannstein's CDA and Vazquez-Navarro et al.
Regarding the suggestion to mention Vazquez-Navarro et al. and the tracking of LEO-detected contrails in geostationary imagery, I took some time to explore the authors' labeled dataset provided with this preprint. Having reviewed the high-resolution (10 to 30 m) imagery, I applied Mannstein's Contrail Detection Algorithm (CDA) used in Vazquez-Navarro et al., and this showed that the method is not suited to detect the contrails at the stage of their lifecycle that is under scrutiny in this study.
The CDA was explicitly designed to address the detection of large, mature, and persistent contrails. Conversely, validating formation and early lifecycle models requires observing much smaller, thinner, and younger contrails right at the aircraft's wake. This represents an entirely different physical and algorithmic problem. While citing Vazquez-Navarro et al. as broad historical context certainly wouldn't hurt, it is not directly transferable to the methodology here. If the authors were to expand their literature review in that direction, they would also need to cite Mansteinn’s CDA that ACTA uses, as well as all the more recent LEO-specific contrail detection literature, such as the Google Landsat dataset, Modis based contrail studies, VIIRS contrail studies and so on, as well as other contemporary works focusing on high-resolution satellite annotations.
Thank you again for your review; this interactive discussion highlights exactly why I believe this preprint is such a thought-provoking contribution to our field.
Citation: https://doi.org/10.5194/egusphere-2026-1171-RC3
-
RC3: 'Reply on RC1', Anonymous Referee #2, 17 Jul 2026
-
RC2: 'Comment on egusphere-2026-1171', Anonymous Referee #2, 17 Jul 2026
REFEREE REPORT
Title: Observing formation and early evolution of contrails formed by IAGOS aircraft using high-resolution LEO satellite imagery
Manuscript No.: egusphere-2026-1171
- GENERAL COMMENTS
First, I would like to sincerely congratulate the authors on a truly excellent and innovative piece of work. The authors develop and apply a highly commendable methodology to observe contrail formation and early evolution by collocating IAGOS aircraft in-situ data with high-resolution LEO satellite imagery from Sentinel-2 and Landsat. This research addresses a critical observational gap in the field, as geostationary satellites often lack the spatial resolution necessary to detect the very early, thin stages of contrail lifecycles. By painstakingly building a dataset of flight collocations the authors have provided a fantastic new empirical foundation for the atmospheric science community.
I particularly want to highlight the rigorous attention to detail demonstrated in the data methodology. Accurately geolocating aircraft in push-broom satellite imagery is a surprisingly difficult task. The authors' corrections for parallax displacement and their intricate accounting for detector-specific scanning time offsets represent a significant and impressive technical achievement. Furthermore, leveraging this unique dataset to evaluate the Schmidt-Appleman Criterion (SAC) and the CoCiP model against both in-situ (IAGOS) and reanalysis (ERA5) meteorological data yields valuable insights into how input data quality directly impacts contrail prediction performance.
Overall, this manuscript is fundamentally strong, well-written, and makes a substantial contribution toward bridging the gap between in-situ measurements and large-scale remote sensing. However, to ensure the conclusions are perfectly framed for the community and to prevent potential over-extrapolation of the results, there are a few structural assumptions and methodological constraints that would benefit from further contextualization. Specifically, I recommend expanding the discussion regarding fleet representativeness, particularly concerning the SAC's applicability to modern engine architectures, alongside addressing the statistical limitations in the dispersion phase analysis and acknowledging some hardware latencies. Addressing these points will easily elevate this great manuscript and solidify its findings for future Air Traffic Management and climate mitigation applications.
- SPECIFIC COMMENTS (MAJOR SCIENTIFIC ISSUES)
[Fleet Representativeness and SAC Assumptions]
I note that the dataset currently focuses on legacy wide-body aircraft (specifically the Airbus A330 and A340 families) operating primarily on specific route structures. Furthermore, because Sentinel-2 and Landsat are sun-synchronous, the observations are naturally skewed toward standard departure schedules that cross the satellite path at fixed local times.
Your results beautifully demonstrate that the Schmidt-Appleman Criterion (SAC) performs exceptionally well for predicting contrail genesis in these older aircraft types used in the IAGOS fleet. However, it is crucial to avoid making definitive statements about the SAC's performance across the global fleet.
Recent studies using ground camera images (including findings presented in a recent CICONIA Outreach session) show that SAC prediction performance can vary significantly, dropping from an F1 score of around 0.9 for older engines (e.g., CFM56 on the A320ceo) to an F1 score of around 0.75 for modern lean-burn engines (e.g., Leap 1-A on the A320neo).
Recent modeling and empirical research supports this discrepancy. For instance, Ponsonby et al. (2025) demonstrated via microphysical models that while the SAC critical temperature error (Delta T_SAC) is rather small (less than 2 K) for rich-burn (soot-rich) engines where the diameter is greater than 20 nm, the classical SAC struggles with lean-burn (soot-poor) engines. When soot emissions drop below 10^13 to 10^14 per kg, droplet activation relies on highly volatile particles (diameter less than 5 nm), which pushes the activation threshold to much colder temperatures and severely miscalculates the active ice crystal emission index (AEI_ice).
Additionally, flight campaigns like ECLIF II/ND-MAX (Voigt et al., 2021) highlight that in modern low-soot/lean-burn engine plumes, low soot concentrations fail to effectively scavenge gaseous precursors. Volatile aerosols thus grow at temperatures significantly below the standard T_SAC. High-fidelity 3D Large-Eddy Simulations of early-stage jet-vortex interactions paired with ground-based camera imagery validation further corroborate this.
Recommendation: It is certainly not the responsibility of this paper to improve or fix the SAC. However, I kindly request that you add a statement in the discussion or limitations section explicitly acknowledging that while these results validate the SAC's high performance for older engine types, literature shows its predictive capability is significantly reduced for modern lean-burn, ultra-high bypass ratio, and SAF-burning engines. Consequently, operational generalizations regarding the global fleet should be carefully tempered.
[Statistical Insignificance in the Dispersion Phase]
I appreciate the authors' transparency regarding the limited sample size, correctly noting that three contrail cases survived longer than 600 seconds. While the observed horizontal growth rates align nicely with CoCiP simulations, drawing operational or definitive modeling conclusions from three isolated cases is difficult. I highly recommend not drawing any conclusions out of these results, just keep them as an example case study.
[Manual Annotation Subjectivity]
The manual annotation of contrails using polygons and linestrings is a solid, practical approach. However, defining the edges of a dispersing contrail by the human eye in 10 m or 30 m resolution imagery is inherently subjective. A small note acknowledging the potential for non-reproducible biases in the width evolution metrics would be a great addition.
[Wind Advection and Spatial/Temporal Mismatch]
The explicit acknowledgment that ignoring horizontal wind advection induces an error in contrail age of approximately 20% (for a 50 m/s wind speed) is much appreciated. Briefly mentioning that this spatial/temporal mismatch could introduce noise into the correlation between the modeled early evolution and the true physical state would round out this section nicely.
[In-Situ Sensor Latency]
The choice not to apply a correction to the IAGOS capacitive hygrometer (ICH) to avoid artificially amplifying noise makes complete practical sense. However, because this leaves an uncorrected temporal lag (up to 120 seconds at 210 K), it would be beneficial to briefly discuss how this lag might affect the relative humidity values used to validate the SAC threshold within the context of your error margins.
[Aircraft performance modeling Assumption Uncertainty]
Applying a generic 15% uncertainty buffer to bridge the gap between the BADA 4.2 and Poll-Schumann overall engine efficiency models is a reasonable compromise, though it masks the localized thrust variations that physically govern condensation. A short sentence recognizing this constraint would be helpful.
- TECHNICAL CORRECTIONS AND MINOR ISSUES
[Incomplete Annotations]
There are a few examples in the dataset where only part of the contrail was annotated. Examples include:
- Case 1: 2019032209415402_LC82070182019081LGN00
- Case 2: 2020013109252202_LC80190202020031LGN00
- Case 3: 2014040110181403_LC81990232014091LGN01
This is completely understandable given the complexity of the imagery, but it might slightly skew the metric stating that only 3 examples lasted much longer than the small section that was labeled. Please consider adding a short disclaimer regarding partial annotations.
[Unclear Imagery]
A significant amount of images in the dataset didn’t allow to distinguish contrails, even though there is a label.
For example : the case LC82050312013210LGN01, it is quite difficult to see anything on the cirrus band png file provided in the dataset, and only a very diffuse shape is visible on the thermal band. Could you add a brief explanatory note detailing why this specific case presents this way? Note that these are much more visible when using the uncompressed data directly from the gcp repo.
[Proposed Table Addition (Formation vs. Persistence)]
If space & time permits, it would be highly informative and interesting to see a table outlining the scores for the contrails regarding which ones were predicted to be persistent (SAC + ISSR) versus contrails that effectively lasted until the start of the diffusion phase (or those lasting more than 5 minutes).
REFERENCES
Poll, D. I. A., and Schumann, U. (2021). An estimation method for the fuel burn and other performance characteristics of civil transport aircraft during cruise: part 2, determining the aircraft's characteristic parameters. Aeronautical Journal, 125(1284), 296-341. https://doi.org/10.1017/aer.2020.124
Ponsonby, J., Teoh, R., Kärcher, B., and Stettler, M. (2025). An updated microphysical model for particle activation in contrails: the role of volatile plume particles. EGUsphere [preprint]. https://doi.org/10.5194/egusphere-2025-1717
Quibén Figueroa, R., Ferreira, T., Gorlé, C., and Soler Arnedo, M. (2025). Comparing high-fidelity LES of early contrail formation with ground-based images. EGU General Assembly 2025, Vienna, Austria, 27 Apr–2 May 2025, EGU25-17191. https://doi.org/10.5194/egusphere-egu25-17191
Voigt, C., Kleine, J., Sauer, D., et al. (2021). Cleaner burning aviation fuels can reduce contrail cloudiness. Nature Communications, 12(1), 5461. https://doi.org/10.1038/s41467-021-25574-2
Technical Corrections
L 74: "...either due the difficulty of performing measurements..." Missing preposition. Please correct to "...due to the difficulty...".
L 87: Typo in "(in the case one was forrmed)". Please correct to "formed".
L 322: "...with help of other true colour and false colour images..." Missing article. Change to "...with the help of...".
L 325: "Such annotations are flagged accordingly, to avoid potential misinterpretation of the annotation." The phrasing is slightly repetitive. Consider simplifying to: "Such annotations are flagged accordingly to avoid potential misinterpretation".
L 353: "we make us" As also noted in another comment, please correct to "we make use of"[cite: 3, 4].
L 380: "...collocated in both climb (54 times), cruise (455 times) and descent (34 times) phases." The word "both" is incorrectly applied to three items. Please delete "both".
L 384: "...which has since been fixed, but results in validity flags of 2 for all repaired data and are therefore excluded from the analysis." Is the subject for the latter half of the sentence is missing/mismatched? I don’t really understand the sentence…
L 407: "Over the whole dataset, we found of a mean difference of -0.583 K (TERA5 − TIAGOS, with a standard deviation..." Remove the extra "of" ("we found a mean difference") and add the missing closing parenthesis after TIAGOS.
L 424: "...the probability of finding the observed differences in precision and recall are 0.005 and 0.06 respectively." Subject-verb agreement error. Please change "probability" to "probabilities".
L 498: "This is also visible in a lower R2 value 0.23 between..." Missing preposition. Change to "lower R2 value of 0.23".
L 510: "...47.3 m for both aircraft types ." There is an unnecessary extra space before the period.
L 589: "...does not allow for robust statistical conclusion on the rate of contrail growth..." Missing pluralization. Please change to "robust statistical conclusions".
L 633: "...that the observed differences are explainable by random sampling along." Typo. Please correct "along" to "alone".
L 655-656: "...with on PS computing engine efficiencies 0.01 to 0.02 higher than BADA 4.2." Typo. Please remove "on" to read "...with PS computing...".
L 707: "...since the azimuth angle changes discontinously across the detector border..." Typo. Please correct "discontinously" to "discontinuously".
L 733: "TW, VM, and ZE developed the metholodogy and creation of the dataset." Typo in "methodology". Additionally, "developed the... creation" is awkward phrasing; suggest changing to "...developed the methodology and created the dataset".
Figure 6 caption: "...even when accounting for uncertainty are shown in purple. with lower opacity." Typo with the punctuation. Please change the period before "with" to a comma, or remove it entirely.
Citation: https://doi.org/10.5194/egusphere-2026-1171-RC2
Viewed
| HTML | XML | Total | BibTeX | EndNote | |
|---|---|---|---|---|---|
| 124 | 79 | 7 | 210 | 4 | 6 |
- HTML: 124
- PDF: 79
- XML: 7
- Total: 210
- BibTeX: 4
- EndNote: 6
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
General comments
The authors develop and apply a new method to observe contrail formation and their early evolution with LEO satellites. This is a welcome step forward, since contrail formation and early stages cannot be observed with geostationary satellties due to their low spatial resolution. Almost two decades ago, Vazquez-Navarro et al. (2009) used LEOS to detect contrails that were then tracked in image series from a geostationary satellite. This might be mentioned in the paper. Here the authors use the observations to check the prediction whether a contrail forms or not (via the Schmidt-Appleman criterion SAC). IAGOS data (in situ) are used as the ground thruth of RHi, while ERA5 reanalyses are used for the remaining met. data. CoCiP is used to simulate the lateral expansion of three of the persistent contrails, for up to 20 min.
Overall, I like the description of method how the contrails are detected and geolocated which seems to be surprisingly difficult. However, the conclusion of this paper and sketch of further work (text after L 593) is not something that I would rate as a very important topic. I think, it is not necessary to check the SAC again, since this has been done in extenso many years ago (there is a figure in the 1999 special IPCC report). There are other interesting problems that I would see as more important, in particular the strange fact that sometimes contrails survive for a while in subsaturated or even quite subsaturated cases. How is this possible? There are some examples in this study and a thorough investigation of these cases would be a better follow-on project.
Another possibility could be to make this method available for further tracking with geostationary images. You should think about it.
Apart from this general comment, I have one major issue (major, because it appears often in the paper), and a couple of minor points.
Major issue:
1) Engine efficiency: This misleading expression is used throughout the paper. It is misleading since it indicates that it is a property of the engines. This is not correct. It is rather a property of the whole system Aircraft. In Schumann's (1996) original paper it is termed "overall propulsion efficiency" and defined as FV/Qm_f, that is it is the driving power of the aircraft divided by the energy per second necessary to maintain this power. Of course, the engine is an important part of the system, but the flying body itself with its drag is important as well. Later in the paper the problem appears that BADA and PS models give differing results for descending aircraft. Perhaps, η is no longer well defined in such a situation, since both F (thrust) and m_f (fuel flow rate) approach zero (I don't know this, however).
Specific comments:
L 26: don't mix up "climate effect" (Delta T, see level rise, etc.) with radiative effects (ERF, RF). What you describe is the latter.
L 39: I suggest to delete "For example" and start the sentence right with "Geostationary".
LL 60-64: This discussion is not entirely clear. There are three possibilities for SAC fulfilled-no contrail observed: either the forecast was wrong, or the forecast is correct but no aircraft flying, or the observation is not good enough.
LL 66-69: The expression "lower performance of the SAC" is misleading. I think the SAC itself rests on firm thermodynamic grounds. So if forecast and observations don't agree, the error must be somewhere either in the input data for the SAC or in the measurements, but not in the SAC.
Sect. 2.1: when I read that first, I thought that some info on the different channels (i.e. for what they are good) would be nice. This info is then given later, which is ok. Perhaps you can indicate here, that this info will follow.
L 147: What is a "first-difference standard deviation"?
L 210: Please correct. According to Schumann (1996) T_LM is in °C.
L 220: The word "can" surprised me. Perhaps this can be made more concrete (say for the 250 hPa level).
Eq. 6: Please check the units. They don't combine to a length.
L 235: the effective time scale is not defined.
L 244: What effect does the angle between the flight direction and the swath direction have.
LL 247 ff: It remains unclear how the IAGOS data are used within CoCiP. The problem is that IAGOS has a 4 s time step, while the other met data that are used as input have a 1 hr time step.
L 353: "we make us" ???
Table 5: I was surprised by the similar magnitude of the TP and TN cases. The probable reason is that this table is on contrail FORMATION, not on contrail persistence. It would be helpful to mention this in the table caption.
LL 437 ff: Is the interpretation correct that there are then 81% cloud free cases and that in the cloudy cases the increase of RHi was not large? This could be mentioned (if true).
LL 445-457: The explanation invoking the difference in spatial resolution of the sensors sounds plausible at first reading, but on further thinking, questions arise. If the resolution is 10 or 30 m, but the typical distance of the vortex centres is, say, 60-70 m, then the resolution should in both cases suffice to detect a contrail. So perhaps there is a different reason, for instance the initial optical thickness of the contrail in relation to the threshold contrast necessary for the sensors to detect anything. As you have the relative humidty, you could perhaps try to estimate the optical thickness.
L 471: It is unclear whether these three points are a subset of the few purple ones or of the orange ones.
Fig. 6b: There is one of the orange bars that does not cross the zero line. Please check. In the last line of the caption there is an incomplete sentence.
L 501: Just a comment. This points to a tremendous small-scale variability of the UT humidity field, isn't it?
LL 525-526: Why this, that is why vortex separation (a horizontal distance)? Contrail spreading in the dispersion regime is mainly due to vertical wind shear. Thus the vertical extension of the contrail at the end of the vortex phase is important. As far as I know this is coded in CoCiP.
Sect. 3.3.3: I find this section with its three examples a bit thin. What is the message of this section? What happens in case 1? What are these regions of sublimation? Or are these short paths where SAC is not fulfilled?
LL 575-581: Please see my comment to LL 445-457.
L 634: alone (not along).
References: some entries have formatting problems: namely Knight et al., Neis et al., Tompkins et al.
Reference:
Vazquez-Navarro, M., Mannstein, H., Kox, S. (2015): Contrail life cycle and properties from 1 year of MSG/SEVIRI rapid-scan images. Atmospheric Chemistry and Physics (15), 8739-8749. doi: 10.5194/acp-15-8739-2015.