the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
Transferable Hourly Ozone Forecasting with Transformers
Abstract. We investigate the suitability of a transformer-based approach for air-quality forecasting, focusing on 4-day ahead hourly predictions of surface ozone (O3). The study employs Google’s Temporal Fusion Transformer (TFT) to integrate meteorological predictors, historical pollutant observations, and static station metadata, using an open source implementation with minimal domain-specific preprocessing. The analysis addresses two questions: (1) how efficiently a transformer model can be deployed for regional air quality forecasting, and (2) how well the learned representations transfer across geophysically distinct regions.
Model performance is evaluated against state-of-the-art regional chemical transport model Copernicus Atmosphere Monitoring Service (CAMS) ensemble forecast using observations from Germany. The TFT consistently achieves lower bias and higher forecast skill across all lead times. Suburban monitoring sites exhibit the highest skill relative to CAMS based on RMSE and SMAPE-based metrics. Urban stations show moderate skill against CAMS baseline, while rural stations have reduced skill in comparison but remain positive across the full 96 h forecast, with the strongest improvements observed at shorter lead times. Post–day-1 results indicate a clear separation of performance by station type; suggesting increasing performance stratification by station type beyond day 1, with larger relative gains at urban and suburban sites and smaller but consistently positive skill at rural locations.
Geographic transferability is assessed by adapting a model trained over Germany to South Korea by retraining region-specific metadata embeddings while preserving learned temporal representations. Forecast errors increase by only 5–10 %, indicating that the model captures meteorological drivers of O3 variability that generalize across contrasting anthropogenic and climatic regimes. Ablation experiments further demonstrate the robustness of the chosen experimental configuration for both forecasting performance and cross region transferability.
- Preprint
(4315 KB) - Metadata XML
- BibTeX
- EndNote
Status: final response (author comments only)
-
RC1: 'Comment on egusphere-2026-1562', Anonymous Referee #1, 31 May 2026
-
AC3: 'Author Reply on RC1 comments', Sindhu Vasireddy, 14 Sep 2026
We thank RC1 for the constructive assessment and for identifying places where the current presentation can lead to ambiguity. Several comments reveal that distinctions central to the study, particularly the role of ERA5, the meaning of the probabilistic quantiles, and the comparison of German and South Korean errors, need to be stated more explicitly.
A point-by-point response is provided below. Because several comments from RC1 concern overlapping methodological and interpretative issues, we additionally provide an attached addendum that consolidates the authors' response across 12 major aspects of the manuscript. The corresponding clarifications and revisions will be incorporated into the revised manuscript.
Reviewer comment: Major comment 1, ERA5 reanalysis as known-future meteorology and fairness of the TFT–CAMS comparison.
Response: We agree that the manuscript must unambiguously distinguish historical meteorological inputs from meteorological information available over the forecast horizon. ERA5 reanalysis values were supplied to TFT as known-future covariates, the experiment should not be described as a strictly operational head-to-head comparison with CAMS, because CAMS necessarily propagates meteorological forecast error whereas a reanalysis-driven TFT would not. In that case, our intention is to frame CAMS as an external physics-based forecast benchmark under different information constraints, and to state this limitation prominently. A fully operational comparison of future ERA5 with an operational NWP forecast is not in the scope of this study due to limited data availability. Conversely, if the implemented TFT uses ERA5 only for historical covariates and not future realized meteorology, the present wording has caused a misunderstanding and will be corrected. We will clarify the implemented treatment of ERA5 in the revised manuscript based on the existing model configuration.
Revision/action: Rewrite the data/model-input description so that every variable is explicitly classified as static, historical observed, or known-future. Remove ambiguous references to 'meteorological forecast' unless forecast meteorology was used. Add an explicit operational-comparability limitation where appropriate.
Reviewer comment: Side comment, compare TFT predictive uncertainty with the CAMS ensemble spread.
Response: The suggestion is valuable, but the two uncertainty objects are not identical. TFT quantiles estimate the conditional predictive distribution of O3 at a given forecast horizon through quantile regression. CAMS ensemble spread reflects dispersion among constituent CTM forecasts and therefore primarily represents inter-model ensemble variability. Treating the ensemble spread as numerically equivalent to the TFT predictive interval would conflate different uncertainty definitions. In the present study, the CAMS ensemble product is therefore used as the closest available operational forecast benchmark for the central TFT prediction. We will clarify this distinction and identify a calibrated probabilistic TFT–CAMS comparison as future work rather than implying that the two uncertainty estimates are directly interchangeable.
Revision/action: Clarify the uncertainty definitions and temper claims about probabilistic comparability with CAMS.
Reviewer comment: Side comment, station-based TFT versus coarse gridded CAMS; CAMS-MOS would be a fairer station-level benchmark.
Response: We agree that the spatial support differs. TFT uses local station observations and station metadata, whereas CAMS is a gridded CTM product whose interpolation to a monitoring location does not make it a true station-scale forecast. The comparison was intended to benchmark the data-driven forecast against an independent operational regional air-quality system, not to claim identical observational support. We will state this representativeness mismatch explicitly. A bias-corrected/MOS CAMS product would indeed be a more directly comparable station-level benchmark where a consistent product is available for the study period and stations.
Revision/action: Reframe CAMS as an independent operational benchmark and explicitly state the spatial-representativeness limitation.
Reviewer comment: Major comment 2, approximately 11 ppb RMSE in South Korea versus approximately 2 ppb in Germany, and the abstract's 5–10% increase.
Response: We agree that the absolute RMSE alone does not provide a scale-independent comparison of predictive performance between Germany and South Korea, particularly if the concentration ranges and variability differ between the two domains. We will therefore clarify the absolute RMSE values and, where supported by the existing evaluation outputs, report a corresponding relative/normalized RMSE to facilitate cross-domain interpretation. The previously stated degradation will be rephrased against reference baseline and supported by the existing results or similar appropriate error quantities to describe the output clearly.
Revision/action: We will clarify the evaluation basis of the reported 5–10% statement and ensure that the corresponding baseline, aggregation and units are explicitly stated in the manuscript. State the baseline, station subset, horizon aggregation and units explicitly; revise or remove the percentage if it is not supported under a like-for-like comparison.
RC1, Minor comments
Reviewer comment: All figures, units, font size, resolution, panel alignment.
Response: Agreed. All figures will be regenerated/reformatted with explicit units, larger fonts, improved resolution, aligned panels and consistent visual conventions.
Reviewer comment: L151, forward filling versus linear interpolation.
Response: We will clarify why forward filling was used. Forward filling avoids using information from a future observation to reconstruct a missing historical value, whereas ordinary linear interpolation can introduce look-ahead information if applied across the forecasting origin. Where interpolation is used strictly within already observed history, we will state this explicitly and ensure no information leakage occurs.
Reviewer comment: L170, temporal validation versus independent validation stations.
Response: We will clarify the two split dimensions and exactly which stations are present in training, validation and test sets. Hyperparameter tuning must be described with respect to the validation set actually used. If station generalisation was reserved only for the final test, we will state this design choice and its implication explicitly.
Reviewer comment: L179, performance outside summer.
Response: We agree that year-round skill is useful for assessing operational applicability. We will clarify the scope of the present evaluation and expand the discussion of the seasonal results already reported, including the limitation that the principal evaluation focuses on the ozone-relevant summer period. We will clarify the scope of the present evaluation and expand the discussion of the existing seasonal results, including the limitation that the principal evaluation focuses on the ozone-relevant summer period.
Reviewer comment: L205, sample grouping versus station code.
Response: We will rewrite this paragraph to distinguish the grouping/indexing mechanism from station metadata and explain exactly how each is represented. A station identifier is not intended to act as a physically meaningful predictor; its role, if retained, is to allow the model to represent station-specific effects within the trained network.
Reviewer comment: Fig. 1, red rectangle.
Response: Agreed; it can be removed unless it is required to communicate the spatial domain.
Reviewer comment: Fig. 2, punctuation and 'seen a-priori' wording.
Response: Agreed. The caption/text will be rewritten so that 'seen' refers unambiguously to whether a station identity/location was represented during model fitting, rather than whether the summer-2023 target period itself was seen. The test period remains unseen temporally.
Reviewer comment: Fig. 3, inconsistent RMSE magnitudes between panels and colour conventions.
Response: We will explicitly state the aggregation and subset represented by each panel. Values from panels with different aggregation should not be compared as though they were the same statistic. TFT/CAMS colours will be made consistent throughout.
Reviewer comment: Distribution of station-wise RMSE.
Response: We agree that an aggregated RMSE alone can obscure substantial station-to-station heterogeneity. Since station-specific RMSE values are available from the existing evaluation, we will add a frequency distribution of station-specific RMSE values to show the spread of predictive performance across the evaluated stations. This will complement the aggregated RMSE and make it possible to assess whether the reported overall performance is representative across the station network or influenced by stations with particularly high or low errors.
Reviewer comment: Sect. 3.3.4, frozen versus fine-tuned components.
Response: Agreed. The transfer-learning description will identify which components are frozen and which are updated. The architecture schematic will be revised to indicate clearly which components are frozen and which are fine-tuned.
Reviewer comment: Explain transfer learning and gating for non-experts.
Response: Agreed. We will add a concise conceptual explanation of transfer learning, gating and variable selection in the main Methods.
Reviewer comment: L280, CRPS.
Response: CRPS is an important probabilistic verification metric. The present study evaluates probabilistic performance using quantile-based metrics consistent with the model outputs. We will clarify the rationale for the selected metrics and acknowledge CRPS as an alternative metric for future work.
Reviewer comment: L298–304, unclear wording / 'performance peak'.
Response: Agreed; these sentences will be rewritten in precise statistical language.
Reviewer comment: L310, percentage improvement at rural stations.
Response: Agreed; the corresponding percentage improvement will be stated explicitly based on the existing reported results, provided the compared values use the same evaluation definition.
Reviewer comment: L312–313, tone understates TFT improvement.
Response: Agreed; the wording will be aligned with the quantitative result rather than using qualitative language that understates the difference.
Reviewer comment: Footnote 7, 'anthropogenic forcing'.
Response: Agreed, will be corrected.
Reviewer comment: L314, missing Fig. 4 reference.
Response: Agreed, will be corrected.
Reviewer comment: Fig. 5, resolution/font.
Response: Agreed; corrected together with the global figure revisions.
Reviewer comment: L335, Fig. 8 ordering.
Response: Agreed; figures and citations will be reordered consistently.
Reviewer comment: Sect. 4.3, transfer-learning discussion too short and no South-Korea-only benchmark.
Response: We agree that the transfer-learning discussion should be expanded because geographic transfer is a central contribution. We will move relevant interpretation from the Appendix into the main text. A South-Korea-only baseline would constitute a separate new modelling experiment and is therefore outside the scope of the present manuscript revision. We will instead clarify the scope and interpretation of the existing transfer-learning experiments.
Reviewer comment: L342–343, unexplained CAMS deterministic global forecast.
Response: Agreed. The sentence will be rewritten or removed unless that benchmark is clearly introduced and supported by the presented evaluation.
Reviewer comment: L345, variable importance placement and model identity.
Response: Agreed. Variable-importance results will be placed in a dedicated subsection and explicitly tied to the Germany-trained or South-Korea-fine-tuned model.
Reviewer comment: L357, comparison with total-column/stratospheric ozone.
Response: We agree that total-column ozone is not a like-for-like benchmark for surface O3. The text will be revised so that no direct validation claim is made from this comparison; if it does not provide a defensible diagnostic purpose, it will be removed.
Reviewer comment: L363, elevated ozone episodes.
Response: Agreed. RMSE alone does not establish skill for threshold exceedances. Because the present evaluation does not include dedicated categorical exceedance metrics, we will remove claims of demonstrated performance specifically during elevated-ozone episodes and state this limitation explicitly.
Reviewer comment: L365, 'minor degradation'.
Response: Agreed. This wording will be replaced by a quantitatively defined comparison consistent with the corrected transfer-learning metrics.
Reviewer comment: Table A.2, climatological mean emissions.
Response: We will clarify that the climatological emission information is used as station-level contextual metadata rather than as a year-specific emissions forecast and acknowledge the resulting limitation.
Reviewer comment: Appendix D1/D2 organization and probabilistic ablation.
Response: Agreed that the appendix can be reorganized and the unablated baseline should be shown alongside ablations. The Appendix will be reorganized so that the existing unablated baseline and existing ablation results are presented together more clearly. Additional probabilistic ablation experiments are outside the scope of the present revision.
Reviewer comment: L711, longer windows.
Response: We will clarify whether longer encoder windows were tested. We will not imply an optimization over window length unless such experiments were actually performed.
Reviewer comment: L717, 'forecast meteorology' despite ERA5 reanalysis.
Response: Agreed. This wording will be corrected to match the actual meteorological input. This is linked directly to Major Comment 1.
Reviewer comment: L775, punctuation.
Response: Agreed; will be corrected.
-
AC3: 'Author Reply on RC1 comments', Sindhu Vasireddy, 14 Sep 2026
-
RC2: 'Comment on egusphere-2026-1562', Anonymous Referee #2, 25 Jul 2026
This study presents a transformer based model, the Temporal Fusion Transformer (TFT), for surface O3 prediction up to 4 days. The TFT model outperforms the regional chemical transport model (CTM), in this case CAMS, with better agreements with observations. This study also examines model’s spatial generalization by applying the trained model, which is based on datasets in Germany, to another region South Korea. Overall, the scientific objective of this paper is clear while the structure needs extra refinement. I would suggest declination for now as additional work and in-depth analysis are necessary, but resubmission is recommended.
Major comments:
- The transformer architecture, including TFT, has been used for various forecasting applications and the findings in this work are not new (e.g. median forecast). The key innovation should be the geographic transferability from Germany to South Korea. However it is not fully discussed in the paper. For instance, Figure 6 shows the comparison of O3 predictions in South Korea based on different fine-tuning strategies. The reasons why changing specific strategies can improve the results are not presented as well as the impacts of input variables on O3 prediction. Some discussions on the differences between Germany and South Korea (e.g. geographic/meteorological characteristics and O3 sources) should be warranted. In addition, RMSE of 11 ppb seems quite large. Say the average daily O3 is 40 ppb based on Figure 6, then a bias of 11 ppb contributes about 25% of uncertainty, which may highly affect the reliability of model results. Also it does not physically make sense that O3 concentration 10 days ago shows high impacts on current-day O3. Please elaborate.
- The choice of TFT model and data selection is not identified. Google’s TFT is not the only transformer model although it provides reliable performance in timeseries prediction. Did the authors test other models? If so, please elaborate. As for meteorological data, should TOAR provide site-specific weather variables (e.g. T, RH, WS)? The spatial resolution of ERA-5 is quarter degree, which is much coarse compared to point data. Since this study focuses on site dependent predictions, using site observation as input could valid the uncertainties due to data gridding and processing. CAMS has a resolution of 40 km, even coarser than ERA-5. Spatial interpolation could highly affect the site-specific values and thus a direct comparison is unfair. Also the selection of sites is unclear. The authors mentioned “The stations were chosen for the study to cover the full spatial region and also such that they are equally balanced across different types of locations”. However, there are around 100 sites being removed (386 out of 493) which is about 20% of the total. Are these sites randomly chosen? Will site selection affect model performance? Please clarify.
- The quantile analysis demonstrates that trained model follows the behavior of median forecast. The associated discussions are somehow misleading. The “median forecast” referring from ECMWF is for ensemble product. The idea is that the median values from various predictions generated by a collection of models (with different configs) usually provide the best performance. However, the median shown in this work is not the median value from various models but the average stage presented by the deep learning model. Moreover, unfortunately, an air quality prediction model (such as O3 prediction) may not be operational applicable if it fails to capture the extreme cases especially in urban and suburban regions. It would be interesting to see an overview of results from all selected sites, e.g. average timeseries. It may better represent the predictive capacity of the trained model.
- I felt a little confused while reading Section 3. Many important information such as variable lists, data sources and model structure/configurations are in appendices. These should be in the main context since they are the soul of the proposed model. Please consider moving some materials to the main context. For example, an extended Table A1 summarizing all input variables with abbreviations, long names and data sources. Full descriptions like Table A2 can stay in appendix. Same for model architecture. A general description, even though TFT is widely used, is warranted.
- All figures need refinement. I assume “horizon step” represents forecast time. Please consider unifying corresponding labels to the same term, “time step” could be a good choice. The fontsize is too small especially figure resolution is quite bad. Figures should be fully described and presented in order. Figure 3 is shown but not mentioned in the main context. Figure 8 is first mentioned in Section 4.2, prior to Figure 6. Other figure-specific comments can be found below. The writing needs to be polished as well. I am not a native English writer but current writing is somehow distracting and should be carefully reviewed. There are many random spaces and commas that should not present. Also please use italics and capitals cautiously.
Minor comments:
- Line 43: “most do not provide hourly resolution of forecasts” -> This statement is questionable. Probably not for Germany specific but there should be many DL models working on hourly O3 predictions.
- Section 2: Please consider revising this section as many things are not fully described. For instance, the authors mention foundation models but do not describe the definition and application in this type of work. Since previous works like MLAir established the scope of this work, a brief description of these works should be warranted, not just a list of key findings. Should sections 2.0.1 and 2.0.2 be just 2.1 and 2.2?
- Section 3: Descriptions of data sources should be included. A site map showing the locations of selected sites would also be helpful.
- Line 128-135: Objectives should be mentioned in the intro.
- Figure 1: Does it show all selected 386 sites? If not, why showing a German map with all sites? Same as the South Korea sites.
- Figure 2: external prediction?
- Figure 3: There seems to be a diurnal pattern in RMSE which can also be seen in Figure 4 and 8. It indicates that TFT model may have issues predicting daytime O3 peaks especially near noon time. O3 photochemical formation is strong during the day, as well as NOx emissions (O3 precursor). This is not necessarily a bad thing because the diurnal pattern probably can be a evidence that TFT model capture the overall O3 variability.
- Line 310: where does the 35% come from?
- Line 312: “In rural regions, where ozone variability is more strongly driven by large-scale meteorology” -> It might be true, but please elaborate and/or provide references.
- Figure 6: consistent yaxis scale
- Line 345-351 and Figure 7: Population density shows the highest importance probably associated with urban environments. Cities usually show higher O3 due to more NOx emissions. Is it possible to verify the importance of urban, suburban and rural sites separately? The contributions of meteorology might be more significant. In addition, what does decoder/encoder importance mean? Do you use different input variables in decoder and encoder? Also it is weird to see station code (a string I assume) has such high importance. In fact, I am surprised to see station code is used as an input. With this information, the trained model may only be able to predict O3 at these given sites and thus show poor performance beyond the region. It probably can explain the high RMSE (11 ppb) in South Korea since the station code is beyond the model range. I would suggest removing station code from inputs as it is not informative regarding O3 concentration.
- Line 357-359: The model is trained based on surface data. Comparing with total-column ozone is not a fair comparison as the environment, source and physical/chemical processes are totally different. “The goal is to verify that the model captures large-scale synoptic/seasonal ozone variability,” -> current work does not present the seasonality of surface O3 given the 4 day timeseries. Synoptic variability usually refers to spatial variability which is not shown in current work either. Please verify.
- Section 4.4: merge to Section 5?
- Line 401-403: move to discussion?
Citation: https://doi.org/10.5194/egusphere-2026-1562-RC2 -
AC2: 'Author Reply on RC2 comments', Sindhu Vasireddy, 14 Sep 2026
We thank RC2 for the detailed comments. We agree that the manuscript should highlight the geographic-transfer contribution and the probabilistic nature of the TFT outputs substantially clearer to distinguish from point forecasts which are prevalent in the field. Several points also identify terminology that can lead to a point-forecast interpretation of a work that is primarily intended as a quantile forecast.
A point-by-point response is provided below. Because several comments from RC2 concern overlapping methodological and interpretative issues, we additionally provide an attached addendum that consolidates the authors' response across 12 major aspects of the manuscript. The corresponding clarifications and revisions will be incorporated into the revised manuscript.
Reviewer comment: Major comment 1, novelty, transferability discussion, 11 ppb RMSE, and the apparent importance of O3 from 10 days earlier.
Response: We agree that applying TFT itself is not the principal novelty. The study does not claim novelty from simply producing a median point forecast. TFT is trained to estimate multiple conditional quantiles of future O3; the 0.5 quantile is only the central member of this predictive distribution and is highlighted because it provides a natural central estimate for comparison with deterministic/reference forecasts. We will revise the manuscript so that this distinction is explicit and place greater emphasis on geographic transfer from Germany to South Korea.
The transfer-learning section will be expanded to explain what is frozen/fine-tuned and why different strategies can change target-domain performance. We will add physical context on differences between Germany and South Korea, including meteorological regime, seasonality, station environments and O3 formation/precursor conditions, while avoiding causal claims that are not directly tested.
We also agree that approximately 11 ppb RMSE is not small and should be discussed transparently. However, RMSE is not bias: an RMSE of 11 ppb relative to a characteristic concentration of 40 ppb cannot be described as an 11 ppb systematic bias or directly as 25% bias. We will report RMSE consistently as a prediction-error metric and revise the manuscript to avoid confusing RMSE with systematic bias.
Finally, a high importance assigned to a lag near ten days should not be interpreted as a direct causal effect of O3 ten days earlier on present O3. A temporal model can use lagged O3 as a proxy for persistent or recurring atmospheric states, synoptic regimes, seasonality, precursor conditions and autocorrelation. TFT variable/attention importance indicates information used by the model, not physical causality. We will make this distinction explicit. We will therefore revise the manuscript to distinguish explicitly between model-derived feature importance and physical causality and will avoid interpreting the identified lag as evidence of a direct causal influence.Revision/action: Rewrite novelty claims; expand geographic-transfer discussion; correct RMSE-versus-bias terminology; explain feature importance as model reliance rather than causality.
Reviewer comment: Major comment 2, TFT choice, ERA5 versus station meteorology, CAMS resolution, and station selection.
Response: TFT was not selected because it is the only Transformer capable of forecasting or the only architecture that can be trained with quantile loss. Alternative forecasting architectures are discussed in the literature review. TFT was retained because it provides, in one architecture, multi-horizon quantile forecasting, static station information, historical and known-future covariates, variable selection and interpretable gating/attention mechanisms. An additional motivation was to evaluate whether an established, general-purpose time-series architecture could be applied to multi-station O₃ forecasting without developing a bespoke deep-learning architecture specifically for this application. TFT was therefore used as an established forecasting framework whose existing capabilities matched the requirements of the study, rather than as a newly developed architecture. We will make this methodological rationale explicit and avoid overstating architectural uniqueness.
Regarding meteorology, ERA5 provides a spatially and temporally consistent predictor set over the complete network. Station meteorological observations can offer better local representativeness, but availability, completeness and measurement consistency can vary by station, since only very few air quality stations report meteorological measurements at all. We agree that ERA5 cannot resolve all station-scale meteorological variability and will state this as a limitation. The use of spatially consistent ERA5 rather than heterogeneous station-level meteorological observations is a methodological choice of the present study. We will explain this rationale and explicitly acknowledge the resulting limitation in representing station-scale meteorological variability as part of main manuscript away from appendix for ease of context.
We also agree that CAMS and station observations have different spatial support. Interpolation of a coarse CTM field to a station introduces representativeness uncertainty. The CAMS comparison will therefore be described as an independent regional operational benchmark rather than a perfectly matched station-scale comparison.
The station subset was not chosen according to model performance. Stations were selected randomly as subject to maintaining spatial coverage and an approximately balanced representation of urban, suburban and rural/location categories, with the reduced sample motivated by computational/data-size constraints. Sensitivity to the amount of station data is already addressed by ablation experiments in the Appendix. We will move the selection procedure and the relevance of these ablations into the main Methods so that this is visible without relying on the Appendix.Revision/action: Expand TFT rationale, station-selection algorithm and ablation cross-reference; state ERA5/CAMS spatial-resolution limitations.
Reviewer comment: Major comment 3, 'median forecast' terminology, ensemble median, extremes, and overview across sites.
Response: We agree that the current use of 'median forecast' can be misleading. The 0.5 output in this study is the median of the model's estimated conditional predictive distribution, not the median across an ensemble of independently configured models. We will consistently call it the 0.5 predictive quantile and explicitly distinguish it from an ensemble median.
The model is also not restricted to the central quantile: upper and lower quantiles are intended to represent different parts of the conditional O3 distribution. Nevertheless, the reviewer is correct that operational usefulness during extreme/high-O3 events cannot be established merely from median RMSE. Claims about extremes will therefore be tied to all the quantile results shown. We will improve the presentation of the existing multi-station results along with inclusion of discussion statements on other quantile metrics presented (currently only plot images of these quantiles are provided but will include qualitative discussion as well) to make the overall predictive behavior across the selected stations clearer.Revision/action: Replace ambiguous 'median forecast' terminology and strengthen/qualify evaluation of high-O3 conditions.
Reviewer comment: Major comment 4, important data/model information is relegated to appendices.
Response: Agreed. Detailed variable descriptions can remain in the Appendix for space, but the main Methods should contain the information required to understand the experiment. We will add a compact table listing variable abbreviation, full name, data source and TFT role (static, historical observed, known-future where applicable), and a concise architecture description covering variable selection, encoder/decoder inputs, gating, temporal processing and quantile outputs.
Revision/action: Move a compact input-variable/source table and a general TFT architecture description into the main text.
Reviewer comment: Major comment 5, figure refinement, terminology/order, and language.
Response: Agreed. We will standardize forecast-horizon terminology, increase font sizes and resolution, ensure every figure is introduced before it appears, correct figure ordering, expand captions, and perform a full language/formatting pass covering punctuation, spacing, italics and capitalization.
RC2, Minor comments
Reviewer comment: L43, 'most do not provide hourly resolution of forecasts'.
Response: Agreed that this is too broad. The statement will be narrowed to the specific literature/domain surveyed or removed.
Reviewer comment: Section 2, foundation models and prior works insufficiently described; numbering.
Response: Agreed. We will define the term as used in this context, explain the relevance of representative prior studies rather than listing findings only, and correct subsection numbering to 2.1/2.2 as appropriate.
Reviewer comment: Section 3, data-source descriptions and site map.
Response: Agreed. Data sources will be summarized in the main Methods. The existing map will be checked/revised so that the displayed stations correspond clearly to the selected datasets.
Reviewer comment: L128–135, objectives belong in Introduction.
Response: Agreed; the objectives will be moved or restated in the Introduction.
Reviewer comment: Figure 1, does it show all 386 selected sites?
Response: The caption and map will be revised to state exactly which stations are displayed. If it currently shows a broader network than the analyzed subset, the figure will be corrected or the distinction made explicit.
Reviewer comment: Figure 2, 'external prediction'.
Response: The terminology will be defined or replaced with a more precise term describing prediction at stations/regions outside the relevant training exposure.
Reviewer comment: Figure 3, diurnal RMSE pattern and daytime O3 peaks.
Response: We agree that the diurnal error structure is scientifically informative. We will discuss it as evidence of horizon/time-of-day dependent difficulty, particularly during daytime photochemical O3 evolution, without overinterpreting it as proof that the model has learned the underlying chemistry.
Reviewer comment: L310, origin of 35%.
Response: The percentage will be explicitly derived from the reported metrics or removed if it cannot be reproduced.
Reviewer comment: L312, rural O3 and large-scale meteorology.
Response: The statement will be qualified and supported with appropriate references or rewritten as an interpretation rather than an established fact.
Reviewer comment: Figure 6, consistent y-axis scale.
Response: Agreed; will be corrected.
Reviewer comment: L345–351/Fig. 7, population density, urban/rural importance, encoder/decoder importance, station code.
Response: We will clarify that feature importance measures model reliance, not causal contribution. Encoder importance refers to variables used over the historical context, while decoder importance refers to variables supplied over the prediction horizon; the exact input sets will be listed. Station code/grouping will be explained as an identifier/representation of station-specific effects, not a physical O3 driver. The concern that such an identifier may hinder out-of-domain transfer is valid and will be discussed. The manuscript will clarify the interpretation and limitations of the existing variable-importance analysis, including that it should not be interpreted as a causal attribution or as a station-type-specific analysis.
Reviewer comment: L357–359, total-column ozone comparison.
Response: Agreed that total-column ozone is not directly comparable to surface O3 because the governing processes and vertical domain differ. Any language presenting this as validation of surface O3 skill will be removed; the comparison will be removed entirely if it lacks a defensible diagnostic purpose.
Reviewer comment: Section 4.4, merge with Discussion.
Response: Will re-phrase and check if this improves flow; the section structure will be revised.
Reviewer comment: L401–403, move to Discussion.
Response: Agreed; will be moved where appropriate.
-
AC1: 'Author Reply on RC1 comments', Sindhu Vasireddy, 14 Sep 2026
We thank RC1 for the constructive assessment and for identifying places where the current presentation can lead to ambiguity. Several comments reveal that distinctions central to the study, particularly the role of ERA5, the meaning of the probabilistic quantiles, and the comparison of German and South Korean errors, need to be stated more explicitly.
A point-by-point response is provided below. Because several comments from RC1 concern overlapping methodological and interpretative issues, we additionally provide an attached addendum that consolidates the authors' response across 12 major aspects of the manuscript. The corresponding clarifications and revisions will be incorporated into the revised manuscript.
Reviewer comment: Major comment 1, ERA5 reanalysis as known-future meteorology and fairness of the TFT–CAMS comparison.
Response: We agree that the manuscript must unambiguously distinguish historical meteorological inputs from meteorological information available over the forecast horizon. ERA5 reanalysis values were supplied to TFT as known-future covariates, the experiment should not be described as a strictly operational head-to-head comparison with CAMS, because CAMS necessarily propagates meteorological forecast error whereas a reanalysis-driven TFT would not. In that case, our intention is to frame CAMS as an external physics-based forecast benchmark under different information constraints, and to state this limitation prominently. A fully operational comparison of future ERA5 with an operational NWP forecast is not in the scope of this study due to limited data availability. Conversely, if the implemented TFT uses ERA5 only for historical covariates and not future realized meteorology, the present wording has caused a misunderstanding and will be corrected. We will clarify the implemented treatment of ERA5 in the revised manuscript based on the existing model configuration.
Revision/action: Rewrite the data/model-input description so that every variable is explicitly classified as static, historical observed, or known-future. Remove ambiguous references to 'meteorological forecast' unless forecast meteorology was used. Add an explicit operational-comparability limitation where appropriate.
Reviewer comment: Side comment, compare TFT predictive uncertainty with the CAMS ensemble spread.
Response: The suggestion is valuable, but the two uncertainty objects are not identical. TFT quantiles estimate the conditional predictive distribution of O3 at a given forecast horizon through quantile regression. CAMS ensemble spread reflects dispersion among constituent CTM forecasts and therefore primarily represents inter-model ensemble variability. Treating the ensemble spread as numerically equivalent to the TFT predictive interval would conflate different uncertainty definitions. In the present study, the CAMS ensemble product is therefore used as the closest available operational forecast benchmark for the central TFT prediction. We will clarify this distinction and identify a calibrated probabilistic TFT–CAMS comparison as future work rather than implying that the two uncertainty estimates are directly interchangeable.
Revision/action: Clarify the uncertainty definitions and temper claims about probabilistic comparability with CAMS.
Reviewer comment: Side comment, station-based TFT versus coarse gridded CAMS; CAMS-MOS would be a fairer station-level benchmark.
Response: We agree that the spatial support differs. TFT uses local station observations and station metadata, whereas CAMS is a gridded CTM product whose interpolation to a monitoring location does not make it a true station-scale forecast. The comparison was intended to benchmark the data-driven forecast against an independent operational regional air-quality system, not to claim identical observational support. We will state this representativeness mismatch explicitly. A bias-corrected/MOS CAMS product would indeed be a more directly comparable station-level benchmark where a consistent product is available for the study period and stations.
Revision/action: Reframe CAMS as an independent operational benchmark and explicitly state the spatial-representativeness limitation.
Reviewer comment: Major comment 2, approximately 11 ppb RMSE in South Korea versus approximately 2 ppb in Germany, and the abstract's 5–10% increase.
Response: We agree that the absolute RMSE alone does not provide a scale-independent comparison of predictive performance between Germany and South Korea, particularly if the concentration ranges and variability differ between the two domains. We will therefore clarify the absolute RMSE values and, where supported by the existing evaluation outputs, report a corresponding relative/normalized RMSE to facilitate cross-domain interpretation. The previously stated degradation will be rephrased against reference baseline and supported by the existing results or similar appropriate error quantities to describe the output clearly.
Revision/action: We will clarify the evaluation basis of the reported 5–10% statement and ensure that the corresponding baseline, aggregation and units are explicitly stated in the manuscript. State the baseline, station subset, horizon aggregation and units explicitly; revise or remove the percentage if it is not supported under a like-for-like comparison.
RC1, Minor comments
Reviewer comment: All figures, units, font size, resolution, panel alignment.
Response: Agreed. All figures will be regenerated/reformatted with explicit units, larger fonts, improved resolution, aligned panels and consistent visual conventions.
Reviewer comment: L151, forward filling versus linear interpolation.
Response: We will clarify why forward filling was used. Forward filling avoids using information from a future observation to reconstruct a missing historical value, whereas ordinary linear interpolation can introduce look-ahead information if applied across the forecasting origin. Where interpolation is used strictly within already observed history, we will state this explicitly and ensure no information leakage occurs.
Reviewer comment: L170, temporal validation versus independent validation stations.
Response: We will clarify the two split dimensions and exactly which stations are present in training, validation and test sets. Hyperparameter tuning must be described with respect to the validation set actually used. If station generalisation was reserved only for the final test, we will state this design choice and its implication explicitly.
Reviewer comment: L179, performance outside summer.
Response: We agree that year-round skill is useful for assessing operational applicability. We will clarify the scope of the present evaluation and expand the discussion of the seasonal results already reported, including the limitation that the principal evaluation focuses on the ozone-relevant summer period. We will clarify the scope of the present evaluation and expand the discussion of the existing seasonal results, including the limitation that the principal evaluation focuses on the ozone-relevant summer period.
Reviewer comment: L205, sample grouping versus station code.
Response: We will rewrite this paragraph to distinguish the grouping/indexing mechanism from station metadata and explain exactly how each is represented. A station identifier is not intended to act as a physically meaningful predictor; its role, if retained, is to allow the model to represent station-specific effects within the trained network.
Reviewer comment: Fig. 1, red rectangle.
Response: Agreed; it can be removed unless it is required to communicate the spatial domain.
Reviewer comment: Fig. 2, punctuation and 'seen a-priori' wording.
Response: Agreed. The caption/text will be rewritten so that 'seen' refers unambiguously to whether a station identity/location was represented during model fitting, rather than whether the summer-2023 target period itself was seen. The test period remains unseen temporally.
Reviewer comment: Fig. 3, inconsistent RMSE magnitudes between panels and colour conventions.
Response: We will explicitly state the aggregation and subset represented by each panel. Values from panels with different aggregation should not be compared as though they were the same statistic. TFT/CAMS colours will be made consistent throughout.
Reviewer comment: Distribution of station-wise RMSE.
Response: We agree that an aggregated RMSE alone can obscure substantial station-to-station heterogeneity. Since station-specific RMSE values are available from the existing evaluation, we will add a frequency distribution of station-specific RMSE values to show the spread of predictive performance across the evaluated stations. This will complement the aggregated RMSE and make it possible to assess whether the reported overall performance is representative across the station network or influenced by stations with particularly high or low errors.
Reviewer comment: Sect. 3.3.4, frozen versus fine-tuned components.
Response: Agreed. The transfer-learning description will identify which components are frozen and which are updated. The architecture schematic will be revised to indicate clearly which components are frozen and which are fine-tuned.
Reviewer comment: Explain transfer learning and gating for non-experts.
Response: Agreed. We will add a concise conceptual explanation of transfer learning, gating and variable selection in the main Methods.
Reviewer comment: L280, CRPS.
Response: CRPS is an important probabilistic verification metric. The present study evaluates probabilistic performance using quantile-based metrics consistent with the model outputs. We will clarify the rationale for the selected metrics and acknowledge CRPS as an alternative metric for future work.
Reviewer comment: L298–304, unclear wording / 'performance peak'.
Response: Agreed; these sentences will be rewritten in precise statistical language.
Reviewer comment: L310, percentage improvement at rural stations.
Response: Agreed; the corresponding percentage improvement will be stated explicitly based on the existing reported results, provided the compared values use the same evaluation definition.
Reviewer comment: L312–313, tone understates TFT improvement.
Response: Agreed; the wording will be aligned with the quantitative result rather than using qualitative language that understates the difference.
Reviewer comment: Footnote 7, 'anthropogenic forcing'.
Response: Agreed, will be corrected.
Reviewer comment: L314, missing Fig. 4 reference.
Response: Agreed, will be corrected.
Reviewer comment: Fig. 5, resolution/font.
Response: Agreed; corrected together with the global figure revisions.
Reviewer comment: L335, Fig. 8 ordering.
Response: Agreed; figures and citations will be reordered consistently.
Reviewer comment: Sect. 4.3, transfer-learning discussion too short and no South-Korea-only benchmark.
Response: We agree that the transfer-learning discussion should be expanded because geographic transfer is a central contribution. We will move relevant interpretation from the Appendix into the main text. A South-Korea-only baseline would constitute a separate new modelling experiment and is therefore outside the scope of the present manuscript revision. We will instead clarify the scope and interpretation of the existing transfer-learning experiments.
Reviewer comment: L342–343, unexplained CAMS deterministic global forecast.
Response: Agreed. The sentence will be rewritten or removed unless that benchmark is clearly introduced and supported by the presented evaluation.
Reviewer comment: L345, variable importance placement and model identity.
Response: Agreed. Variable-importance results will be placed in a dedicated subsection and explicitly tied to the Germany-trained or South-Korea-fine-tuned model.
Reviewer comment: L357, comparison with total-column/stratospheric ozone.
Response: We agree that total-column ozone is not a like-for-like benchmark for surface O3. The text will be revised so that no direct validation claim is made from this comparison; if it does not provide a defensible diagnostic purpose, it will be removed.
Reviewer comment: L363, elevated ozone episodes.
Response: Agreed. RMSE alone does not establish skill for threshold exceedances. Because the present evaluation does not include dedicated categorical exceedance metrics, we will remove claims of demonstrated performance specifically during elevated-ozone episodes and state this limitation explicitly.
Reviewer comment: L365, 'minor degradation'.
Response: Agreed. This wording will be replaced by a quantitatively defined comparison consistent with the corrected transfer-learning metrics.
Reviewer comment: Table A.2, climatological mean emissions.
Response: We will clarify that the climatological emission information is used as station-level contextual metadata rather than as a year-specific emissions forecast and acknowledge the resulting limitation.
Reviewer comment: Appendix D1/D2 organization and probabilistic ablation.
Response: Agreed that the appendix can be reorganized and the unablated baseline should be shown alongside ablations. The Appendix will be reorganized so that the existing unablated baseline and existing ablation results are presented together more clearly. Additional probabilistic ablation experiments are outside the scope of the present revision.
Reviewer comment: L711, longer windows.
Response: We will clarify whether longer encoder windows were tested. We will not imply an optimization over window length unless such experiments were actually performed.
Reviewer comment: L717, 'forecast meteorology' despite ERA5 reanalysis.
Response: Agreed. This wording will be corrected to match the actual meteorological input. This is linked directly to Major Comment 1.
Reviewer comment: L775, punctuation.
Response: Agreed; will be corrected.
Citation: https://doi.org/10.5194/egusphere-2026-1562-AC1
Data sets
Model Checkpoints, Inference Outputs and Plots for Transferable Hourly Ozone Forecasting Sindhu Vasireddy https://zenodo.org/records/19151740?token=eyJhbGciOiJIUzUxMiIsImlhdCI6MTc3NDE5MjkzOSwiZXhwIjoxODA1NTg3MTk5fQ.eyJpZCI6ImYzYmVjZmM1LWViNTEtNDA1Yi05Yzk2LTQyNmI2NzQ4YTMxNiIsImRhdGEiOnt9LCJyYW5kb20iOiI5MDFhZjM4ODRjOTZiMTJlN2RlMWE4MWZiYTkwN2FiZSJ9.4nlYEJ7KbA_KYUoo7xtMHpev4erpA3A2PE5UEOVK8zW1yuj6c_TO-k-pUp38cEAn1ELnRI1m54Xz7BTu44rZXg
Model code and software
TOAR Ozone Data Processing Pipeline for Transformer-Based Forecasting Sindhu Vasireddy https://zenodo.org/records/19151435?token=eyJhbGciOiJIUzUxMiIsImlhdCI6MTc3NDEzMzI4MCwiZXhwIjoxODA1OTMyNzk5fQ.eyJpZCI6IjgyMTBiYTE2LWFjYmMtNDcwMi05NWExLTlmOWI1MmVlOGE3MyIsImRhdGEiOnt9LCJyYW5kb20iOiJiMmQ4MjU1OTRkN2YxYjQ3Mzg3NDczZjJkMDM3OWI2MyJ9.J_1oreVJATEgxDxlKzon0elpPqbTnF_0qskg8Xuy3ezCAUt1Hoe2QPykJCG63VMDcFYGwyJp6vl-2PdDcI4Nuw
Transformer-Based Framework for Transferable Hourly Ozone Forecasting Sindhu Vasireddy https://zenodo.org/records/19151703?token=eyJhbGciOiJIUzUxMiIsImlhdCI6MTc3NDEzMzIwNSwiZXhwIjoxODA1OTMyNzk5fQ.eyJpZCI6IjE1OGZmOTgwLTgyYTAtNGJlMS04ZDJjLTVmMDE2MTMwNzMwYiIsImRhdGEiOnt9LCJyYW5kb20iOiIyMzY2YjJhNWM4YzE2NjBjYzYxZGVmYTQ5YjI5ZjFlMiJ9.eUvbvfHopML2Y
Viewed
| HTML | XML | Total | BibTeX | EndNote | |
|---|---|---|---|---|---|
| 342 | 201 | 35 | 578 | 34 | 26 |
- HTML: 342
- PDF: 201
- XML: 35
- Total: 578
- BibTeX: 34
- EndNote: 26
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
This paper presents the application of a popular deep learning architecture – Temporal Fusion Transformer (TFT) – to ozone forecasting. The study initially focuses on Germany where long ozone records are available for training, and where TFT predictions can be compared to state-of-the-art physics-based ozone forecasts from the CAMS regional ensemble. The study then explores the geographic transferability of the trained TFT model to another region with fewer observations, in this case South Korea. Overall, the authors report improved skills compared to CAMS and reasonably good geographical transferability.
Over the last years, the TFT architecture has been used for a variety of time series forecasting applications, apparently with a reasonably good success rate. This justifies the interest of exploring its skills on air pollution forecasting, and although the application of this specific type of model is not new (e.g. Hickman et al., 2023), the authors are still proposing here some refinements (e.g. station-level anthropogenic metadata). Therefore, the innovation of the paper is probably more on the side of the geographical transferability, although this should come with a more extended discussion of the results.
Overall, the paper is clear and well written (although some specific parts could be improved, see minor comments), and falls in the scope of GMD. I suggest accepting the publication but after addressing the major issues described below, which I think could strengthen the study.
Major comments:
The first major comment is related to the set-up chosen for the AI-versus-CAMS comparison, which I think currently represents a significant limitation given that the AI forecasting model is using as known future the ERA5 reanalysis, which would evidently not be available in an operational context, and should thus have been replaced by a meteorological forecast. At least this is what I understood, but it is still partially confusing because the authors are mentioning several times the importance of “meteorological forecast” as known future inputs (L114 and L722), but the data description only mentions the use of ERA5 meteorological reanalysis. If the TFT model does not rely on meteorological forecast but on meteorological reanalysis, then the comparison against the CAMS operational air quality forecast – that only relies on meteorological forecasts – is unfair. Consequently, we can expect the AI forecast model to be less (not) affected by error accumulation on the meteorology, that represents a key driver of the O3 variability, as mentioned by the authors. Given that this comparison against CAMS is quite central in the paper, it would be important to ensure a fairer comparison, using meteorological operational forecast as known future covariates (e.g. IFS or equivalent) (and eventually meteorological operational analysis for past known covariates). At the very least, the authors should make a very clear statement about this strong limitation, but to me the paper would be much stronger replacing ERA5 reanalysis by some meteorological forecast.
As a side comment, given that the authors are evaluating the uncertainties obtained with the TFT model, for a more comprehensive AI-versus-CAMS comparison it would have been useful and very informative to compare them to the uncertainties of the CAMS ensemble, as derived from the spread of the different individual members. Finally, I don’t think it is completely fair to compare an observation-based forecast relying on local station-based information to a pure CTM-based forecast at 10-km (thus quite coarse) resolution. A more appropriate comparison would have required using some CAMS forecast bias-corrected with local observations. I think CAMS is already providing CAMS-MOS forecast, but maybe only at a limited number of stations and probably not in 2023. Here again, this limitation should be highlighted more clearly.
The second major comment is related to the results of the transfer learning. If I understand correctly, please correct me if I am wrong and adjust the text accordingly to avoid confusion, the RMSE on urban stations in South Korea is around 11 ppb (Fig. 6) while it is around 2 ppb in Germany (Fig. 3). This is a strong difference, roughly a factor 4, and therefore I don’t understand why is the abstract is talking about “only 5-10% increase of the forecast errors” when passing from Germany to South Korea? It is crucial to clarify that point.
Minor comments:
All figures: Please revise all figures and include systematically the units of the variable or metric shown. The font size and resolution of the figures should be increased so that to be readable without having to zoom. Besides that, the quality of some figures could be generally improved, especially the multi-panels plots that are not aligned.
L151: About the use of forward-filling, why not doing a simple linear interpolation? This seems to me already much better than repeating the last value.
L170: The authors mention that they split their dataset into train-validation-test sets along the temporal dimension, but they mention only train-test split along the station dimension. Does it mean that the model tuning is performed only the “temporal” validation set (April 2015 to December 2022) but still considering the same 386 stations used for training? Please clarify, and if so, please explain why no independent stations were kept also not only for testing but also for validation.
L179: Only the test set allows providing an unbiased estimate of the skills of the predictive model. Therefore, although summer is indeed the most relevant season for O3 episodes, it would still be useful to have an idea of the performance of the model all along the year, considering that it has been trained with samples distributed all along the year and not only in summer. Footnote 5 suggests that the AI model performs similarly to CAMS in spring and winter. Could you provide more quantitative results during the spring/winter/fall seasons (and ideally some plots in Appendix or Supplement)? Even if ozone episodes occur mostly in summer, I think this is still a relevant aspect given that some ozone episodes can occur outside summer season along early or late heat waves which are more frequent under climate change. (In an operational context, if the AI model is better than CAMS only in summer, this raises the question of when exactly the AI model starts and stops to be more skilful.)
L205: Sample grouping: I am not sure to understand the new feature introduced here, please clarify this paragraph. Do you mean that in practice this new feature takes values from 1 to S with S the total number of stations? If so, it would correspond to a unique identifier for each station, so what would be the difference with the station code that already encodes such unique information for each station?
Fig. 1: I don’t think the red rectangle brings much here.
Fig. 2: Please correct some missing punctuation in the legend. Also, I don’t understand why the authors are describing stations shown in panels d-e-f as “seen a-priori during training or validation” given that they previously said summer 2023 was used only for testing. And why “a-priori one or the other”? Please clarify.
Fig. 3: I don’t understand why for a given station type, for instance urban stations, the RMSE of TFT shown in panel c (around 2.5 ppbv) is much lower than the one shown below for the percentile 50 (around 5-7 ppbv). RMSE also differ for CAMS. Please explain. Also, it would be better to keep the same colour for TFT and CAMS across all figures and panels, here they are inversed.
Also, results of ozone forecast across all stations (here and in elsewhere) are shown in terms on RMSE averaged over all stations. This tells little about the stations where forecast may be the least skilful, could the authors provide some information regarding the distribution (and not only the mean) of the RMSEs across the different stations?
Sect. 3.3.4: In this section it is not clear which components remain frozen. To illustrate more easily which components of the initial TFT model trained over Germany are frozen and which ones are fine-tuned with South Korean data, it would be useful to replicate fig. B1 indicating clearly for instance with a specific colour the components that are retrained.
It would be useful to explain in more detail how transfer learning and gating layers work, so that non-experts on AI can still understand.
L280: The authors should also provide results on the CRPS metric it is the most used in atmospheric forecasting applications. This would facilitate comparisons against other studies.
L298-304: Revise these sentences, the formulation is unclear and the English quite poor. In particular, “performance peak” does not mean much to me in the context of this paragraph.
L310: Please provide the percentage improvement at rural stations.
L312-313: TFT still shows substantially better performance than CAMS here (I would say maybe roughly 15% improvement), this is not so much reflected by the tone of this sentence.
Footnote 7: Change for “anthropogenic forcing”.
L314: Missing reference to Fig. 4.
Fig. 5: Increase the resolution and font size of the figure (and resolution could probably be increased in several other figures).
L335: Fig. 8 should probably come before to be consistent with the order of appearance in the text.
Sect. 4.3: The authors are highlighting the geographical transferability as one of the main contributions of the paper, but the analysis/discussion of the transfer learning results remains very short in the main text (only a few lines, from L337 to L344). I would suggest extending it a bit, maybe including part of the results discussed in Appendix in the main document. One specific issue of this section is that they are no benchmark forecast to compare the TFT model fine-tuned with local data. The authors could eventually consider training their same TFT model directly on South Korea, in the same way as they trained in over Germany and compare the performance of both approaches, or use a simpler approach.
L342-343: I don’t understand the part on “an RMSE of 11 ppb […] when compared against CAMS deterministic global forecasts”, these global CAMS forecasts have not been introduced before and are not shown on the figures. Please reformulate in a clearer way.
L345: Why is this discussed here in the section of transferability of the model to South Korea? This does not seem related and should probably be placed in a dedicated section on variable importance. It is not clear if this variable importance concerns the model trained over Germany or the one fine-tuned over South- Korea. Please clarify.
L357: I really don’t see the interest of this comparison against CANMS stratospheric ozone. At least the authors should have considered the CAMS ozone tropospheric column, not the total column, or better the surface concentration, this comparison against the total column does not make a lot of sense to me.
L363: Where are AI and CAMS comparisons made on elevated ozone episodes? As far as I understand, results are mostly evaluated using RMSE which does not provide insights on these episodes. Some categorical metrics on ozone exceedances above regulatory threshold would be required here to support this statement.
L365: Which minor degradation are the authors referring to here? (see my previous comment on the RMSE increased by x4 in South Korea compared to Germany).
Table A.2: Why using a climatological mean emission, which is likely not the most accurate information about emissions around a given station on a given year?
D1: The authors should refer to this figure at the beginning of Appendix D, include the AI model performance before the ablation to facilitate the comparison. Also, it would be interesting to know how this ablation affects the probabilistic forecast through the WIS for instance. Finally, I don’t understand why D2 is not merged with D1.1, both treating the same ablation aspects, this should be reorganised as it is a bit confusion right now, with transfer learning (D1.2) in the middle.
L711: Do the authors tried longer windows? Does it improve the skills?
L717: This aligns with the limitation mentioned before, that the TFT model probably benefits significantly from relying on reanalysis meteorology instead of forecast meteorology. I don’t understand why now the authors are saying “These results highlight the importance of conditioning on forecast meteorology” if they are not using such forecast but the ERA5 reanalysis, please reformulate to avoid confusion.
L775: “…product, a proxy…”.