the Creative Commons Attribution-NonCommercial 4.0 International License.
the Creative Commons Attribution-NonCommercial 4.0 International License.
Two decades of kilometer-scale daily PM2.5 from satellite observations and machine learning reveal geographically diverging exposure in Ghana
Abstract. Exposure to fine particulate matter (PM2.5) is a major contributor to global burden of disease, yet air quality data remain sparse in many low- and middle-income countries, limiting nationwide monitoring and effective policy development. We address this gap by developing a high-resolution gridded (1 km × 1 km) dataset for daily surface PM2.5 concentrations in Ghana from 2005 to 2025 by training multiple machine learning (ML) models built on ground-based monitoring, satellite observations, and reanalysis products for atmospheric composition and meteorological parameters. Estimates from these models were evaluated with measurements from reference-grade monitors and a large network of calibrated low-cost sensors deployed across Ghana. XGBoost showed the strongest performance among all ML algorithms and best captured spatial and temporal variability in PM2.5 levels. SHapley Additive exPlanations (SHAP) analysis for model predictors indicates that both meteorological variables and aerosol optical properties are key contributors to model performance. The long-term gridded PM2.5 dataset reveals a unique north-south exposure disparity in Ghana, with northern regions of the country experiencing substantially higher PM2.5 concentrations compared to the South, that may be widening over the 21-year period by over 0.2 µg m-3 yr-1. This study provides the first long-term high-resolution PM2.5 exposure levels for Ghana and presents a scalable framework for generating air quality information in data-sparse regions to support air pollution relevant health impact assessment and evidence-based mitigation policies.
Status: final response (author comments only)
-
RC1: 'Comment on egusphere-2026-3370', Anonymous Referee #1, 09 Aug 2026
The comment was uploaded in the form of a supplement: https://egusphere.copernicus.org/preprints/2026/egusphere-2026-3370/egusphere-2026-3370-RC1-supplement.pdfCitation: https://doi.org/
10.5194/egusphere-2026-3370-RC1 -
RC2: 'Comment on egusphere-2026-3370', Anonymous Referee #2, 17 Aug 2026
Kudos to the author team for putting together this impressive paper estimating daily surface-level PM2.5 in Ghana combining a number of satellite- and model-derived datasets and an impressively large in the African context and multi-sensor/diverse network of low-cost instruments through machine learning.
The writing is excellent; figures are clear (I especially like figure 1 and think it does a good job at distilling the complicated procedure for generating the authors’ dataset); and I appreciate the US-Ghanaian partnerships demonstrated through this paper’s slate of authors and the deployment of the in-situ monitoring network leveraged within. To me, this study is well on track for publication in AMT. I have included several comments below that, when addressed, should make the paper in excellent shape for publication.
Introduction:
- Lines 35ff: The author might consider, in addition to mentioning the overall burden of disease from air pollution in Africa, stating the estimated burden in Ghana alone. I believe these values should be readily accessible via the Global burden of Disease’s/IHME’s Global Health Data Exchange (GHDx). It might help contextualize the importance of this particular word better.
- Lines 85ff: The sentence beginning with “Furthermore, …” feels a little misplaced. Prior to this, the authors are discussing recent advances in high resolution PM2.5 modeling, so it feels weird to pivot back to the deficiencies of MERRA-like products. This sentence might be better suited earlier in the Introduction when discussing challenges of limited monitoring.
Methods:
- Section 2.2, Figure S2: The low-cost sensors have uneven spatiotemporal coverage and are available only from 2018 onwards. Are there ways that the co-variations in AOD and other measures of pollution, meteorology, and other inputs to their machine learning model are unique to the 2018-2025 period (b/c climate change, b/c teleconnections, etc.) such that applying these machine-learned insights to previous periods spanning 2005-2018 could be inappropriate if the coupling among these drivers have changed, especially in nonlinear ways? It would be good for the authors to discuss this potential and even better if there are quantitative ways to demonstrate the impact that this could have on results (assuming I’m thinking through it correctly).
- In Sections 2.3.1, 2.3.2, and 2.3.3, I appreciated the detailed information about how the authors applied the QA/QC flags recommended for the different products to their study. It would be beneficial for readers to understand the fractions of data filtered out by these flags to better understand how often the ML methods rely on the MERRA model. Perhaps a schematic akin to Figure S2 could illustrate this nicely (although I understand it’s a little more difficult than replicating Fig. S2 exactly since there’s a spatial dimension to the satellite measurements). Relatedly, the authors allude to missing satellite data in Lines 369ff but I don’t see any quantitative information on how common this is.
- Can the authors comment on the appropriateness of generating their PM2.5 product at ~1km2 when all the native inputs besides, I believe, AOD appear to be in the 5km2 or greater range? Unlike with some LUR products I’ve seen that have several fine scale (< 1km2) inputs like road locations and NDVI, this product seems to inherent all of the fine-scale features from AOD alone, which might be decoupled from surface-level PM.
- Line 387: I believe the authors meant to reference Figure S6, not S4, here.
- Section 2.6: The use of 1-9km buffers to define the extent of cities feels overly simplistic. I am not familiar with most of the cities studied here, but I assume some might be bigger than 18km from north to south or east to west? I also just looked up Kintampo on Google Earth for the fun of it, and its built-up area appears to be quite elongated and not a circle. The authors might consider using products like GHS-UCDB or GHS-SMOD to more precisely define areas over which they will average PM2.5. Alternatively, if any Ghanaian agencies have city shapefiles/boundaries, those would be great too.
- Sections 2.7.1, 2.7.2, and 2.7.3: I appreciate the detailed information about the statistical procedures but, in my opinion, feel like it could be shortened significantly or even removed. These statistical tests are fairly standard, and unless the authors modified them in some major way, I don’t know readers need almost 2 pages devoted to them.
Results:
- For Figure 2 and the associated discussion, while the MB improves nearly 20-folder for the TROPOMI model compared with the MERRA model, the MAE, RMSE, and correlation are nearly identical and not that much different from the OMI model. I think the authors attributed this similar performance to the complete daytime estimates from MERRA compared with the early afternoon measurements from TROPOMI and OMI, but could it also be that that PM2.5 is mostly explained by regional phenomena that are captured by the coarse MERRA product rather than needing the finer scale data provided by TROPOMI? Or does it potentially relate back to my earlier comment about the lack of fine-scale inputs?
- Line 692: how are the northern and southern regions of Ghana shown in Figure 3B, or is this a typo?
- Figure 5: Allowing each subplot to have its own y-axis limit based on its data range or using a consistent one that doesn’t extend all the way down to 0 µgm-3 might allow readers to gauge variations a little better, even if it means omitting the WHO guidelines.
- Lines 740ff: why are p-values only quoted for some cities? Are p-values even necessary if it was already stated that trends are significant?
- Paragraph beginning at line 767: When discussing contrasting trends in northern versus southern Ghana, the authors highlight how improved technology e.g., cleaner vehicles, cleaner cooking, etc.) may have reduced PM in southern cities but then point out how northern cities such as Wa and Tamale have positive non-Harmattan trends and suggest that local emissions may be at play in driving the increases. Do the authors have any reason to believe that there is uneven rollout of these technologies in the north versus south or any regional initiatives that would create this dipole by which they would conclude that improved technology only reduces PM in the south?
- Furthermore, in Lines 892-896 the authors stated that declines in southern population centers are consistent with improvements in vehicle fuel standards, emissions controls, and incremental adoption in cleaner household energy. Similar to my question above, is there uneven rollout of these initiatives in the northern cities? Or why wouldn’t this argument be used to understand Ghana-wide urban trends?
- Line 833ff: Isn’t it the *major* wet season that shows a partial recovery of the north-south gradient? The minor wet season PM plot in Fig. 7E looks somewhat anthropogenic/correlated with population density and biomass burning in Figs. S1, S10.
Supplement:
- Table S2: Just curious/checking - is Kintampo supposed to be “rural”?
- Text S1: For Airnote, where are the corrections coefficients obtained from? There doesn’t appear to be a BAM in Kumasi and Somanya from which scalings were obtained? I didn’t look up the Owusu-Tawiah et al. (2025) reference to see if it’s in there, but - even if it is - I think more explanation should be made available here so if anyone wants to reproduce/extend this works, it’s clear to understand the assumptions of important equations like these.
- Fig. S3: should the x- and y-axis tick label currently written as “NO” be “NO_2”?
Citation: https://doi.org/10.5194/egusphere-2026-3370-RC2
Data sets
Gridded Daily PM2.5 at 1 km × 1 km Resolution from 2005–2025: Ghana Abhishek Anand, Joe Adabouk Amooli, and Daniel M. Westervelt https://zenodo.org/records/19636051
Viewed
Since the preprint corresponding to this journal article was posted outside of Copernicus Publications, the preprint-related metrics are limited to HTML views.
| HTML | XML | Total | BibTeX | EndNote | |
|---|---|---|---|---|---|
| 136 | 0 | 1 | 137 | 0 | 0 |
- HTML: 136
- PDF: 0
- XML: 1
- Total: 137
- BibTeX: 0
- EndNote: 0
Viewed (geographical distribution)
Since the preprint corresponding to this journal article was posted outside of Copernicus Publications, the preprint-related metrics are limited to HTML views.
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1