Development of an automated contrail-to-flight attribution using Meteosat Second Generation satellite data
Abstract. Persistent contrails are a major contributor to the effective radiative forcing from aviation. Optimizing flight trajectories to avoid regions prone to the formation of warming contrails has therefore been proposed as a mitigation strategy to achieve the international climate targets for the aviation sector. However, the precise attribution of observed contrails to individual flights remains a challenge for both the evaluation of avoidance measures and the validation of contrail models. This study presents a fully automated evaluation framework designed for matching contrails to flights using observations from the Spinning Enhanced Visible and Infra-Red Imager (SEVIRI) instrument aboard the geostationary Meteosat Second Generation (MSG) satellite and flight trajectories. The method utilizes an automated contrail detection algorithm. Although such automated detections designed for MSG/SEVIRI often face a trade-off between high detection efficiency and low false alarm rate, the proposed matching process substantially improves detection reliability through a multi-step verification: contrails are preselected based on their spatial and temporal occurrence as well as their spatial orientation with respect to the flight trajectories, followed by temporal tracking. A new aspect is the development of a life cycle-based confidence score to derive a quantitative matching score. In contrast to previous approaches, the proposed method is entirely observation-driven, requires no model input, and enables rapid, large-scale application. The framework’s performance is demonstrated through two specific case studies. Future applications include the assessment of contrail avoidance trials across the Europe-African airspace and adaptation to other satellite configurations.
This paper describes an automated contrail-to-flight attribution algorithm. A few such algorithms exist in the literature (Chevalier 2023, Geraedts 2024, Sarna 2025). The main innovation of this paper is to apply its algorithm to the MSG satellite, which introduces new challenges due to its relatively low resolution. This is valuable since Europe is one of the most contrail-dense regions in the world and prior to the availability of MTG in late 2024 MSG was the only way to detect contrails in this region by geostationary satellite. Automated contrail-to-flight attribution systems are challenging to develop, so developing one and reporting on its performance is a valuable contribution. The paper therefore should be published after some revisions.
I have one major criticism of this paper which is that they do comparably little analysis of how their algorithm actually performs. They make the statement “The method demonstrates high potential for automated, observation-based verification of contrail mitigation strategies in practice” but I don’t think they’ve done enough validation to justify this statement. Producing validation as thorough as that in Sarna 2025 might not be feasible, but the authors should do better than just the 4 case studies presented in this work. Even increasing the number of case studies to 10 would be an improvement. Another approach, similar to Chevalier/Geraedts, would be to run their system on a few hundred flights and check if aggregate statistics of things like contrail age, diurnal/seasonal patterns, fraction of flights making contrails, are plausible. Ideally upon resubmission the authors should include some more efforts towards validating their methods. Failing that they should not stress much more strongly that they do not have a good estimate of the model's performance.
I have a number of minor comments listed below:
From what little validation the authors do, their method doesn’t appear to do very well. When compared to manual labels, their method only agrees on one out of four contrails, implying a precision and recall of 0.25. This is far less than what is achieved by Chevalier 2023, Geraedts 2024 or Sarna 2025 (as evaluated in Sarna 2025) [it's a small sample size and not an apples-to-apples comparison, but still I think this is a bad sign]. This does not mean that the work does not merit publication. Even if the performance does not match the best algorithms already in the literature, sharing how different design decisions affect performance is useful to inform future algorithm development. Furthermore the poor performance could in part be because of the poor resolution of SEVIRI. But I think the authors should at least call out this disparity, the statement ““The method demonstrates high potential for automated, observation-based verification of contrail mitigation strategies in practice” does not seem justified. It would be good if people reading this paper who might be interested in using the algorithm should at least be cautioned that it might not give the same results as some other methods in the literature. (For example, the authors plan to use this algorithm to validate a trial. If such a study gives a negative result, I would not necessarily say this implies the trial did not work, it might just be that the algorithm is not good enough to see the impact)
The authors claim that their method is scalable to large numbers of flights. This claim is somewhat undercut by the fact that the authors ran the algorithm on a total of 4 flights. (c.f. Geraedts 2024, which also claims a scalable algorithm and runs on >250 000 flights). The authors should prove that the algorithm is scalable by running on at least 100 flights. If this is too big a burden, then the algorithm is not actually scalable.
The authors made a design decision to limit the maximum advection speed to 50 km/h. The authors claim this is a fine decision, but I don’t agree. They point out that 50 km/h is higher than the mean, which may be true, but that’s not the right metric. The authors should instead compute what fraction of the time the wind is higher than 50km/h. The authors later try a larger wind error. They find in total 26 matching contrails instead of 17. They claim this is a small change, I don’t agree. I think ideally the authors would recompute all their main results with a larger wind threshold. Failing that they should at least change their claim that 50 km/h is sufficient. I would like to see increasing this threshold as a possible direction for future work.
Another design decision the authors make is to not advect the flights. They claim this is a strength of the paper, I think instead it is a weakness that likely contributes to the relatively poor performance of their method compared to existing literature. The authors correctly point out that wind data is uncertain. But having no data at all about the wind is more uncertain! For example the authors consider all contrails within 50 km/h of the flight (which I claim is not big enough). The wind uncertainty is only 20km/h, so if they used the advected flight instead they would need to consider a lot less flights. They have a similar issue with angle. In the introduction they authors claim that sedimentation + wind shear is a reason not to use advection, but then in their algorithm they assume a constant advection speed, which would not be a reasonable assumption of sedimentation + wind shear was really a problem. I think the authors' claims that not advecting is a virtue need to be toned down, and I think they should say that adding advection to their algorithm is a possible direction for future work.
The authors method does not consider multiple flights that might match the same contrail, and the authors correctly highlight this as a drawback of their work as well as discuss why it is hard. A reader might be left with the impression that such consideration is therefore infeasible, but in fact Chevalier, Geraedts and Sarna all do it. The authors should inform the reader of this.
In some of the case study figures I find the label ‘no decision’ confusing. I think the authors mean that these are flights that were within the max velocity of the flight but got a score of 0. Please be more explicit about these.