the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
Reactivity Over Abundance: Unveiling the True Kinetic Drivers of Urban Ozone Using Process-Informed Machine Learning
Abstract. Ground-level ozone remains a persistent challenge in East Asian urban centers, where concentrations continue rising despite significant reductions in precursor emissions. Designing effective mitigation is complicated by the non-linear relationship between precursor abundance and reactivity. Standard data-driven approaches often suffer from "survivor bias," systematically undervaluing highly reactive precursors that are rapidly depleted. We introduce a process-informed machine learning framework that uses net chemical consumption (ΔVOC) rather than ambient concentrations to resolve this attribution failure. Applied to measurements from a NOₓ-saturated roadside site in Seoul, South Korea, the approach reveals a robust quantitative relationship between intrinsic hydroxyl radical reactivity (kOH) and ozone formation sensitivity.
Across a comprehensive suite of precursors—including oxygenated VOCs (e.g., formaldehyde, acetone) and aromatics—the framework identifies reactive aromatics (trimethylbenzenes, xylenes) and biogenics (isoprene, monoterpenes) as the dominant kinetic drivers. Whereas static metrics such as OFP rank precursors by stoichiometric capacity under idealized accumulated conditions, the process-informed framework shows that realized ozone production in fresh urban plumes is governed by kinetic turnover rather than abundance, with this kinetic selectivity further amplified during high-ozone episodes. These results indicate that mass-based VOC inventories and OFP-style rankings, when applied without kinetic context, can systematically misallocate control priorities in NOₓ-saturated urban regimes. Because the framework requires only sub-hourly co-located VOC and ozone observations and no prescribed mechanism, it offers a complementary empirical pathway in settings where explicit mechanism-based modeling is constrained by incomplete VOC speciation or unmeasured radical precursors.
- Preprint
(894 KB) - Metadata XML
-
Supplement
(894 KB) - BibTeX
- EndNote
Status: open (until 11 Sep 2026)
- RC1: 'Comment on egusphere-2026-2647', I. Pérez, 24 Jun 2026 reply
-
RC2: 'Comment on egusphere-2026-2647', Anonymous Referee #2, 18 Aug 2026
reply
The comment was uploaded in the form of a supplement: https://egusphere.copernicus.org/preprints/2026/egusphere-2026-2647/egusphere-2026-2647-RC2-supplement.pdf
-
RC3: 'Comment on egusphere-2026-2647', Anonymous Referee #3, 03 Sep 2026
reply
In their manuscript 'Reactivity Over Abundance: Unveiling the True Kinetic Drivers of Urban Ozone Using Process-Informed Machine Learning', Hu et al. present a process-informed machine learning framework to predict urban ozone formation and demonstrate its use with data from a roadside air quality monitoring station in Seoul, South Korea. They demonstrate that this approach is more reliable in NOx-saturated environments, where other methods relying on ozone precursors' concentrations suffer from "survivor bias".
I find that the results are presented clearly and that the arguments in the conclusions are satisfactory, demonstrating the usefulness of the proposed framework. However, while the substance of the manuscript is new and interesting, the form could be improved in my opinion. If the authors can address these minor comments, the manuscript should be accepted for publication in ACP.Here are a few comments and suggestions:
- As it seems a pretty important caveat that this framework is useful in NOx-saturated (VOC-limited) environments, I wondered if it would be worth including this in the title somehow.
- The last part of the introduction (lines 60 to 75) seems to me to fit more as an abstract, while some sentences of the abstract are quite general and would fit better in the introduction to leave space in the abstract to describe the presented study and results.
- Section 2.2.2: The configuration of the model and the introduction of shapley additive explanations (SHAP) is condensed in one sentence/two lines. SHAP being such an essential part of the manuscript (every single plot includes SHAP values), it might be worth describing the concept more in the main text. I would suggest bringing the text S7 to the manuscript, including Equation (2) (1?). (Note that it is the only Supplementary Text that is not referred to in the main text. More technical observations below.) Especially for a journal like ACP whose readership might not be entirely familiar with machine learning concepts, it would be useful.
- Section 3.1: The authors describe how they use a fixed ratio from another campaign for C8H10 as PTR-ToF-MS cannot chemically resolve ethylbenze and xylenes. Table S2 indicates ranges of kOH values for xylenes, trimethylbenzenes, and monoterpenes and in the text there is a mention of kOH of "36.0 × 10^(-12)" as "average of three isomers", but it is not clear if it is a weighted average or how the average was calculated as it does not seem consistent with Table S2. Also, it is not clear for monoterpenes which kOH is used and why and how it affects the results. Clarifying these numbers would help readers who would like to apply the framework in a consistent way with their own data.
- The 'Data availability' statement does seem to be sufficient according to ACP guidelines and the authors should justify why the data and/or model cannot be made available through a FAIR-aligned reliable public data repository.In addition, I found some small inconsistencies in the way abbreviations are used and introduced. Here a few examples of these technical corrections:
- In the abstract, ΔVOC is introduced as "net chemical consumption", but VOC is not defined. Nor is NOx or OFP. In the introduction, ΔVOC is introduced as "net VOC consumption" and later as "net change in VOC concentration".
- "Δ" itself is used on its own once after 'processed-informed' and once as 'rate of change', which seems unnecessary to be mentioned like that. Also, "delta model" (not even Δ as defined earlier) is used as shorthand for "processed-informed model", which I would discourage. It would be better to keep a unified nomenclature throughout the manuscript.
- "OHR" is introduced on line 124, but only defined later on line 134.
- Of the compounds presented in the figures using chemical formulas, only monoterpenes seem to be defined in the main text and while Table S2 summarizes the compounds and abbreviations, it might be good to at least define them once in the main text as well. For instance, on line 93, the authors define "formaldehyde (HCHO)", but not "acetone".
- Some concepts are capitalized before the definition of their abbreviation (e.g. SHAP), but this is not according to ACP guidelines.Citation: https://doi.org/10.5194/egusphere-2026-2647-RC3
Viewed
| HTML | XML | Total | Supplement | BibTeX | EndNote | |
|---|---|---|---|---|---|---|
| 76 | 56 | 20 | 152 | 15 | 17 | 15 |
- HTML: 76
- PDF: 56
- XML: 20
- Total: 152
- Supplement: 15
- BibTeX: 17
- EndNote: 15
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
This paper investigates the ozone photochemical production by a machine learning model where the key variable is formed by the volatile organic compounds. The main role is devoted to the chemical processes whereas physical processes are ignored. The measurement campaign extended during one month, from 1 to 30 April 2022 and the study site is inside the city, although the sampling site is surrounded by vegetation and close to a hill (the Namsan Mountain). In consequence, some minor changes should be introduced to determine the paper restrictions.
Since the measurement extension is one month, some comments about the data representativeness should be included. For instance, the authors should consider the prevailing atmospheric circulation with the suitable meteorological charts. Moreover, the authors should comment if the measurements and the model response are similar under varied meteorological conditions, such as front passages or precipitation events.
The sampling site is close to a hill. The authors should comment if the airflow is modified by this hill. In addition, the wind rose presented in Fig. S1c shows that the prevailing wind directions come from this hill, whereas the city contribution looks like secondary.
Moreover, the study site is surrounded by vegetation. Perhaps the precursor composition does not respond to that from the city, i.e. measurements may be quite local.
The authors present a comparison between measured and predicted O3 deltas. Perhaps, potential readers would wonder about the reason for such comparison instead of measured and predicted concentrations. The authors highlight Dt=18 minutes in the text. Perhaps, they should indicate the reason, since Dt=30 minutes determines a better correlation in Fig. S7.
Some of the graphs in Fig. 5 may be misunderstood since most of the space in them is white, without dots. To see a clear result Y axes should be modified, between -5 and 5 in Fig. (b), perhaps between -7.5 and 5 in Fig. (c) and perhaps -2 and 3 in Fig. (c). These changes would allow the linear relationship to be observed and, perhaps, questioned.
Finally, the authors should determine the potential readers of this research, which is quite focused, and they should establish the way to increase them.