the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
Transferable Hourly Ozone Forecasting with Transformers
Abstract. We investigate the suitability of a transformer-based approach for air-quality forecasting, focusing on 4-day ahead hourly predictions of surface ozone (O3). The study employs Google’s Temporal Fusion Transformer (TFT) to integrate meteorological predictors, historical pollutant observations, and static station metadata, using an open source implementation with minimal domain-specific preprocessing. The analysis addresses two questions: (1) how efficiently a transformer model can be deployed for regional air quality forecasting, and (2) how well the learned representations transfer across geophysically distinct regions.
Model performance is evaluated against state-of-the-art regional chemical transport model Copernicus Atmosphere Monitoring Service (CAMS) ensemble forecast using observations from Germany. The TFT consistently achieves lower bias and higher forecast skill across all lead times. Suburban monitoring sites exhibit the highest skill relative to CAMS based on RMSE and SMAPE-based metrics. Urban stations show moderate skill against CAMS baseline, while rural stations have reduced skill in comparison but remain positive across the full 96 h forecast, with the strongest improvements observed at shorter lead times. Post–day-1 results indicate a clear separation of performance by station type; suggesting increasing performance stratification by station type beyond day 1, with larger relative gains at urban and suburban sites and smaller but consistently positive skill at rural locations.
Geographic transferability is assessed by adapting a model trained over Germany to South Korea by retraining region-specific metadata embeddings while preserving learned temporal representations. Forecast errors increase by only 5–10 %, indicating that the model captures meteorological drivers of O3 variability that generalize across contrasting anthropogenic and climatic regimes. Ablation experiments further demonstrate the robustness of the chosen experimental configuration for both forecasting performance and cross region transferability.
- Preprint
(4315 KB) - Metadata XML
- BibTeX
- EndNote
Status: open (until 07 Aug 2026)
- RC1: 'Comment on egusphere-2026-1562', Anonymous Referee #1, 31 May 2026 reply
-
RC2: 'Comment on egusphere-2026-1562', Anonymous Referee #2, 25 Jul 2026
reply
This study presents a transformer based model, the Temporal Fusion Transformer (TFT), for surface O3 prediction up to 4 days. The TFT model outperforms the regional chemical transport model (CTM), in this case CAMS, with better agreements with observations. This study also examines model’s spatial generalization by applying the trained model, which is based on datasets in Germany, to another region South Korea. Overall, the scientific objective of this paper is clear while the structure needs extra refinement. I would suggest declination for now as additional work and in-depth analysis are necessary, but resubmission is recommended.
Major comments:
- The transformer architecture, including TFT, has been used for various forecasting applications and the findings in this work are not new (e.g. median forecast). The key innovation should be the geographic transferability from Germany to South Korea. However it is not fully discussed in the paper. For instance, Figure 6 shows the comparison of O3 predictions in South Korea based on different fine-tuning strategies. The reasons why changing specific strategies can improve the results are not presented as well as the impacts of input variables on O3 prediction. Some discussions on the differences between Germany and South Korea (e.g. geographic/meteorological characteristics and O3 sources) should be warranted. In addition, RMSE of 11 ppb seems quite large. Say the average daily O3 is 40 ppb based on Figure 6, then a bias of 11 ppb contributes about 25% of uncertainty, which may highly affect the reliability of model results. Also it does not physically make sense that O3 concentration 10 days ago shows high impacts on current-day O3. Please elaborate.
- The choice of TFT model and data selection is not identified. Google’s TFT is not the only transformer model although it provides reliable performance in timeseries prediction. Did the authors test other models? If so, please elaborate. As for meteorological data, should TOAR provide site-specific weather variables (e.g. T, RH, WS)? The spatial resolution of ERA-5 is quarter degree, which is much coarse compared to point data. Since this study focuses on site dependent predictions, using site observation as input could valid the uncertainties due to data gridding and processing. CAMS has a resolution of 40 km, even coarser than ERA-5. Spatial interpolation could highly affect the site-specific values and thus a direct comparison is unfair. Also the selection of sites is unclear. The authors mentioned “The stations were chosen for the study to cover the full spatial region and also such that they are equally balanced across different types of locations”. However, there are around 100 sites being removed (386 out of 493) which is about 20% of the total. Are these sites randomly chosen? Will site selection affect model performance? Please clarify.
- The quantile analysis demonstrates that trained model follows the behavior of median forecast. The associated discussions are somehow misleading. The “median forecast” referring from ECMWF is for ensemble product. The idea is that the median values from various predictions generated by a collection of models (with different configs) usually provide the best performance. However, the median shown in this work is not the median value from various models but the average stage presented by the deep learning model. Moreover, unfortunately, an air quality prediction model (such as O3 prediction) may not be operational applicable if it fails to capture the extreme cases especially in urban and suburban regions. It would be interesting to see an overview of results from all selected sites, e.g. average timeseries. It may better represent the predictive capacity of the trained model.
- I felt a little confused while reading Section 3. Many important information such as variable lists, data sources and model structure/configurations are in appendices. These should be in the main context since they are the soul of the proposed model. Please consider moving some materials to the main context. For example, an extended Table A1 summarizing all input variables with abbreviations, long names and data sources. Full descriptions like Table A2 can stay in appendix. Same for model architecture. A general description, even though TFT is widely used, is warranted.
- All figures need refinement. I assume “horizon step” represents forecast time. Please consider unifying corresponding labels to the same term, “time step” could be a good choice. The fontsize is too small especially figure resolution is quite bad. Figures should be fully described and presented in order. Figure 3 is shown but not mentioned in the main context. Figure 8 is first mentioned in Section 4.2, prior to Figure 6. Other figure-specific comments can be found below. The writing needs to be polished as well. I am not a native English writer but current writing is somehow distracting and should be carefully reviewed. There are many random spaces and commas that should not present. Also please use italics and capitals cautiously.
Minor comments:
- Line 43: “most do not provide hourly resolution of forecasts” -> This statement is questionable. Probably not for Germany specific but there should be many DL models working on hourly O3 predictions.
- Section 2: Please consider revising this section as many things are not fully described. For instance, the authors mention foundation models but do not describe the definition and application in this type of work. Since previous works like MLAir established the scope of this work, a brief description of these works should be warranted, not just a list of key findings. Should sections 2.0.1 and 2.0.2 be just 2.1 and 2.2?
- Section 3: Descriptions of data sources should be included. A site map showing the locations of selected sites would also be helpful.
- Line 128-135: Objectives should be mentioned in the intro.
- Figure 1: Does it show all selected 386 sites? If not, why showing a German map with all sites? Same as the South Korea sites.
- Figure 2: external prediction?
- Figure 3: There seems to be a diurnal pattern in RMSE which can also be seen in Figure 4 and 8. It indicates that TFT model may have issues predicting daytime O3 peaks especially near noon time. O3 photochemical formation is strong during the day, as well as NOx emissions (O3 precursor). This is not necessarily a bad thing because the diurnal pattern probably can be a evidence that TFT model capture the overall O3 variability.
- Line 310: where does the 35% come from?
- Line 312: “In rural regions, where ozone variability is more strongly driven by large-scale meteorology” -> It might be true, but please elaborate and/or provide references.
- Figure 6: consistent yaxis scale
- Line 345-351 and Figure 7: Population density shows the highest importance probably associated with urban environments. Cities usually show higher O3 due to more NOx emissions. Is it possible to verify the importance of urban, suburban and rural sites separately? The contributions of meteorology might be more significant. In addition, what does decoder/encoder importance mean? Do you use different input variables in decoder and encoder? Also it is weird to see station code (a string I assume) has such high importance. In fact, I am surprised to see station code is used as an input. With this information, the trained model may only be able to predict O3 at these given sites and thus show poor performance beyond the region. It probably can explain the high RMSE (11 ppb) in South Korea since the station code is beyond the model range. I would suggest removing station code from inputs as it is not informative regarding O3 concentration.
- Line 357-359: The model is trained based on surface data. Comparing with total-column ozone is not a fair comparison as the environment, source and physical/chemical processes are totally different. “The goal is to verify that the model captures large-scale synoptic/seasonal ozone variability,” -> current work does not present the seasonality of surface O3 given the 4 day timeseries. Synoptic variability usually refers to spatial variability which is not shown in current work either. Please verify.
- Section 4.4: merge to Section 5?
- Line 401-403: move to discussion?
Citation: https://doi.org/10.5194/egusphere-2026-1562-RC2
Data sets
Model Checkpoints, Inference Outputs and Plots for Transferable Hourly Ozone Forecasting Sindhu Vasireddy https://zenodo.org/records/19151740?token=eyJhbGciOiJIUzUxMiIsImlhdCI6MTc3NDE5MjkzOSwiZXhwIjoxODA1NTg3MTk5fQ.eyJpZCI6ImYzYmVjZmM1LWViNTEtNDA1Yi05Yzk2LTQyNmI2NzQ4YTMxNiIsImRhdGEiOnt9LCJyYW5kb20iOiI5MDFhZjM4ODRjOTZiMTJlN2RlMWE4MWZiYTkwN2FiZSJ9.4nlYEJ7KbA_KYUoo7xtMHpev4erpA3A2PE5UEOVK8zW1yuj6c_TO-k-pUp38cEAn1ELnRI1m54Xz7BTu44rZXg
Model code and software
TOAR Ozone Data Processing Pipeline for Transformer-Based Forecasting Sindhu Vasireddy https://zenodo.org/records/19151435?token=eyJhbGciOiJIUzUxMiIsImlhdCI6MTc3NDEzMzI4MCwiZXhwIjoxODA1OTMyNzk5fQ.eyJpZCI6IjgyMTBiYTE2LWFjYmMtNDcwMi05NWExLTlmOWI1MmVlOGE3MyIsImRhdGEiOnt9LCJyYW5kb20iOiJiMmQ4MjU1OTRkN2YxYjQ3Mzg3NDczZjJkMDM3OWI2MyJ9.J_1oreVJATEgxDxlKzon0elpPqbTnF_0qskg8Xuy3ezCAUt1Hoe2QPykJCG63VMDcFYGwyJp6vl-2PdDcI4Nuw
Transformer-Based Framework for Transferable Hourly Ozone Forecasting Sindhu Vasireddy https://zenodo.org/records/19151703?token=eyJhbGciOiJIUzUxMiIsImlhdCI6MTc3NDEzMzIwNSwiZXhwIjoxODA1OTMyNzk5fQ.eyJpZCI6IjE1OGZmOTgwLTgyYTAtNGJlMS04ZDJjLTVmMDE2MTMwNzMwYiIsImRhdGEiOnt9LCJyYW5kb20iOiIyMzY2YjJhNWM4YzE2NjBjYzYxZGVmYTQ5YjI5ZjFlMiJ9.eUvbvfHopML2Y
Viewed
| HTML | XML | Total | BibTeX | EndNote | |
|---|---|---|---|---|---|
| 295 | 152 | 24 | 471 | 22 | 21 |
- HTML: 295
- PDF: 152
- XML: 24
- Total: 471
- BibTeX: 22
- EndNote: 21
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
This paper presents the application of a popular deep learning architecture – Temporal Fusion Transformer (TFT) – to ozone forecasting. The study initially focuses on Germany where long ozone records are available for training, and where TFT predictions can be compared to state-of-the-art physics-based ozone forecasts from the CAMS regional ensemble. The study then explores the geographic transferability of the trained TFT model to another region with fewer observations, in this case South Korea. Overall, the authors report improved skills compared to CAMS and reasonably good geographical transferability.
Over the last years, the TFT architecture has been used for a variety of time series forecasting applications, apparently with a reasonably good success rate. This justifies the interest of exploring its skills on air pollution forecasting, and although the application of this specific type of model is not new (e.g. Hickman et al., 2023), the authors are still proposing here some refinements (e.g. station-level anthropogenic metadata). Therefore, the innovation of the paper is probably more on the side of the geographical transferability, although this should come with a more extended discussion of the results.
Overall, the paper is clear and well written (although some specific parts could be improved, see minor comments), and falls in the scope of GMD. I suggest accepting the publication but after addressing the major issues described below, which I think could strengthen the study.
Major comments:
The first major comment is related to the set-up chosen for the AI-versus-CAMS comparison, which I think currently represents a significant limitation given that the AI forecasting model is using as known future the ERA5 reanalysis, which would evidently not be available in an operational context, and should thus have been replaced by a meteorological forecast. At least this is what I understood, but it is still partially confusing because the authors are mentioning several times the importance of “meteorological forecast” as known future inputs (L114 and L722), but the data description only mentions the use of ERA5 meteorological reanalysis. If the TFT model does not rely on meteorological forecast but on meteorological reanalysis, then the comparison against the CAMS operational air quality forecast – that only relies on meteorological forecasts – is unfair. Consequently, we can expect the AI forecast model to be less (not) affected by error accumulation on the meteorology, that represents a key driver of the O3 variability, as mentioned by the authors. Given that this comparison against CAMS is quite central in the paper, it would be important to ensure a fairer comparison, using meteorological operational forecast as known future covariates (e.g. IFS or equivalent) (and eventually meteorological operational analysis for past known covariates). At the very least, the authors should make a very clear statement about this strong limitation, but to me the paper would be much stronger replacing ERA5 reanalysis by some meteorological forecast.
As a side comment, given that the authors are evaluating the uncertainties obtained with the TFT model, for a more comprehensive AI-versus-CAMS comparison it would have been useful and very informative to compare them to the uncertainties of the CAMS ensemble, as derived from the spread of the different individual members. Finally, I don’t think it is completely fair to compare an observation-based forecast relying on local station-based information to a pure CTM-based forecast at 10-km (thus quite coarse) resolution. A more appropriate comparison would have required using some CAMS forecast bias-corrected with local observations. I think CAMS is already providing CAMS-MOS forecast, but maybe only at a limited number of stations and probably not in 2023. Here again, this limitation should be highlighted more clearly.
The second major comment is related to the results of the transfer learning. If I understand correctly, please correct me if I am wrong and adjust the text accordingly to avoid confusion, the RMSE on urban stations in South Korea is around 11 ppb (Fig. 6) while it is around 2 ppb in Germany (Fig. 3). This is a strong difference, roughly a factor 4, and therefore I don’t understand why is the abstract is talking about “only 5-10% increase of the forecast errors” when passing from Germany to South Korea? It is crucial to clarify that point.
Minor comments:
All figures: Please revise all figures and include systematically the units of the variable or metric shown. The font size and resolution of the figures should be increased so that to be readable without having to zoom. Besides that, the quality of some figures could be generally improved, especially the multi-panels plots that are not aligned.
L151: About the use of forward-filling, why not doing a simple linear interpolation? This seems to me already much better than repeating the last value.
L170: The authors mention that they split their dataset into train-validation-test sets along the temporal dimension, but they mention only train-test split along the station dimension. Does it mean that the model tuning is performed only the “temporal” validation set (April 2015 to December 2022) but still considering the same 386 stations used for training? Please clarify, and if so, please explain why no independent stations were kept also not only for testing but also for validation.
L179: Only the test set allows providing an unbiased estimate of the skills of the predictive model. Therefore, although summer is indeed the most relevant season for O3 episodes, it would still be useful to have an idea of the performance of the model all along the year, considering that it has been trained with samples distributed all along the year and not only in summer. Footnote 5 suggests that the AI model performs similarly to CAMS in spring and winter. Could you provide more quantitative results during the spring/winter/fall seasons (and ideally some plots in Appendix or Supplement)? Even if ozone episodes occur mostly in summer, I think this is still a relevant aspect given that some ozone episodes can occur outside summer season along early or late heat waves which are more frequent under climate change. (In an operational context, if the AI model is better than CAMS only in summer, this raises the question of when exactly the AI model starts and stops to be more skilful.)
L205: Sample grouping: I am not sure to understand the new feature introduced here, please clarify this paragraph. Do you mean that in practice this new feature takes values from 1 to S with S the total number of stations? If so, it would correspond to a unique identifier for each station, so what would be the difference with the station code that already encodes such unique information for each station?
Fig. 1: I don’t think the red rectangle brings much here.
Fig. 2: Please correct some missing punctuation in the legend. Also, I don’t understand why the authors are describing stations shown in panels d-e-f as “seen a-priori during training or validation” given that they previously said summer 2023 was used only for testing. And why “a-priori one or the other”? Please clarify.
Fig. 3: I don’t understand why for a given station type, for instance urban stations, the RMSE of TFT shown in panel c (around 2.5 ppbv) is much lower than the one shown below for the percentile 50 (around 5-7 ppbv). RMSE also differ for CAMS. Please explain. Also, it would be better to keep the same colour for TFT and CAMS across all figures and panels, here they are inversed.
Also, results of ozone forecast across all stations (here and in elsewhere) are shown in terms on RMSE averaged over all stations. This tells little about the stations where forecast may be the least skilful, could the authors provide some information regarding the distribution (and not only the mean) of the RMSEs across the different stations?
Sect. 3.3.4: In this section it is not clear which components remain frozen. To illustrate more easily which components of the initial TFT model trained over Germany are frozen and which ones are fine-tuned with South Korean data, it would be useful to replicate fig. B1 indicating clearly for instance with a specific colour the components that are retrained.
It would be useful to explain in more detail how transfer learning and gating layers work, so that non-experts on AI can still understand.
L280: The authors should also provide results on the CRPS metric it is the most used in atmospheric forecasting applications. This would facilitate comparisons against other studies.
L298-304: Revise these sentences, the formulation is unclear and the English quite poor. In particular, “performance peak” does not mean much to me in the context of this paragraph.
L310: Please provide the percentage improvement at rural stations.
L312-313: TFT still shows substantially better performance than CAMS here (I would say maybe roughly 15% improvement), this is not so much reflected by the tone of this sentence.
Footnote 7: Change for “anthropogenic forcing”.
L314: Missing reference to Fig. 4.
Fig. 5: Increase the resolution and font size of the figure (and resolution could probably be increased in several other figures).
L335: Fig. 8 should probably come before to be consistent with the order of appearance in the text.
Sect. 4.3: The authors are highlighting the geographical transferability as one of the main contributions of the paper, but the analysis/discussion of the transfer learning results remains very short in the main text (only a few lines, from L337 to L344). I would suggest extending it a bit, maybe including part of the results discussed in Appendix in the main document. One specific issue of this section is that they are no benchmark forecast to compare the TFT model fine-tuned with local data. The authors could eventually consider training their same TFT model directly on South Korea, in the same way as they trained in over Germany and compare the performance of both approaches, or use a simpler approach.
L342-343: I don’t understand the part on “an RMSE of 11 ppb […] when compared against CAMS deterministic global forecasts”, these global CAMS forecasts have not been introduced before and are not shown on the figures. Please reformulate in a clearer way.
L345: Why is this discussed here in the section of transferability of the model to South Korea? This does not seem related and should probably be placed in a dedicated section on variable importance. It is not clear if this variable importance concerns the model trained over Germany or the one fine-tuned over South- Korea. Please clarify.
L357: I really don’t see the interest of this comparison against CANMS stratospheric ozone. At least the authors should have considered the CAMS ozone tropospheric column, not the total column, or better the surface concentration, this comparison against the total column does not make a lot of sense to me.
L363: Where are AI and CAMS comparisons made on elevated ozone episodes? As far as I understand, results are mostly evaluated using RMSE which does not provide insights on these episodes. Some categorical metrics on ozone exceedances above regulatory threshold would be required here to support this statement.
L365: Which minor degradation are the authors referring to here? (see my previous comment on the RMSE increased by x4 in South Korea compared to Germany).
Table A.2: Why using a climatological mean emission, which is likely not the most accurate information about emissions around a given station on a given year?
D1: The authors should refer to this figure at the beginning of Appendix D, include the AI model performance before the ablation to facilitate the comparison. Also, it would be interesting to know how this ablation affects the probabilistic forecast through the WIS for instance. Finally, I don’t understand why D2 is not merged with D1.1, both treating the same ablation aspects, this should be reorganised as it is a bit confusion right now, with transfer learning (D1.2) in the middle.
L711: Do the authors tried longer windows? Does it improve the skills?
L717: This aligns with the limitation mentioned before, that the TFT model probably benefits significantly from relying on reanalysis meteorology instead of forecast meteorology. I don’t understand why now the authors are saying “These results highlight the importance of conditioning on forecast meteorology” if they are not using such forecast but the ERA5 reanalysis, please reformulate to avoid confusion.
L775: “…product, a proxy…”.