the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
Intercomparison of low-cost sensors via simultaneous atmospheric measurements: a case study
Abstract. The adoption of low-cost sensors (LCS) is growing steadily due to their affordability, ease of use, and broad applicability. However, concerns remain regarding their reliability, prompting continued investigations into their performance and proper handling of measurements.
This study uses a three week field campaign in a urban area in central Italy carried out during the winter holiday season. Atmospheric physical and chemical parameters, temperature, relative humidity, pressure, concentration of carbon monoxide (CO), nitric oxide (NO), nitrogen dioxide (NO2), ozone in the form of O3 and OX and particulate matter PM2.5 and PM10, have been measured by three different commercial LCS platform (Vaisala AQT, AirSensEUR and Libelium Smart Environment PRO) in their factory primary calibration, to assess their initial performance. The LCS have been placed in a site close to two meteorological stations hosting standard certified reference instruments, which have been used for the intercomparison process. Additionally a 2B Ozone Monitor, EPA-certified Federal Equivalent Method, has been mounted next to the LCS, to add ozone to the evaluated variables. Due to the absence of a CO reference dataset, only a comparison between LCS has been performed to asses consistency for this measurement.
Meteorological measurements showed high correlation (R ∼ 0.9) across all LCS with the reference data, except for a discrepancy in temperature and relative humidity for AirSensEUR. The concentrations of NO and NO2 exhibited a good correlation (R ≥ 0.75) with reference instrument, although some discrepancies and deviations from the ideal linear relationship were observed. Differently ozone comparison had a good similarity only for Vaisala AQT (R ∼ 0.8), while for the remaining two the differences are noticeable (R ∼ 0.5). CO time series across the three low-cost sensors are almost the same. Finally, both PM values, available from the reference only as daily averages, showed a reasonable level of agreement with the reference instrument for AirSensEUR, albeit with greater variability.
The LCS data acquired in the atmosphere was also analysed in relation to nearby pollution sources. Workday versus holiday daily comparison and wind pollutant correlation have been executed with the aim to evaluate the ability of these LCS to recognize daily patterns and attribute pollutant sources.
Results show the potential information-driven applications of these commercial low-cost sensors, detecting emission patterns during rush hours and holidays.
Competing interests: At least one of the (co-)authors is a member of the editorial board of Atmospheric Measurement Techniques.
Publisher's note: Copernicus Publications remains neutral with regard to jurisdictional claims made in the text, published maps, institutional affiliations, or any other geographical representation in this paper. While Copernicus Publications makes every effort to include appropriate place names, the final responsibility lies with the authors. Views expressed in the text are those of the authors and do not necessarily reflect the views of the publisher.- Preprint
(4154 KB) - Metadata XML
-
Supplement
(651 KB) - BibTeX
- EndNote
Status: final response (author comments only)
-
RC1: 'Comment on egusphere-2026-1262', Anonymous Referee #1, 16 May 2026
-
AC1: 'Reply on RC1', Lorenzo Gentile, 18 Sep 2026
This work by Gentile at al. presents an interesting intercomparison exercise between low-cost sensors for meteorological and air pollution measurements. The topic is relevant for AMT, particularly in light of the continuously growing interest in low-cost technologies for air quality monitoring and research applications.
In addition to provide insights into the quality and reliability of data from three specific commercial low-cost sensors, the authors also discuss their fitness for purpose by exploring selected use cases, such as the investigation of the typical diel variability of pollutants and source attribution through combined analysis with near-surface wind variability.
The manuscript is generally well organized and easy to follow. I recommend publication after addressing the following technical and minor comments.
Response: We thank the reviewer for the positive assessment of the manuscript and for the constructive technical and editorial comments. We have addressed each point below. References to line numbers currently refer to the submitted preprint; the final line numbers will be updated after implementation of all revisions in the revised manuscript.
Specific Comments
- Introduction
The following WMO report on the use of low-cost sensors in air quality monitoring networks should be included to strengthen the reference list: https://library.wmo.int/records/item/68924-integrating-low-cost-sensor-systems-and-networks-to-enhance-air-quality-applications
It would be valuable if the authors discuss how their work relates to the findings and recommendations presented in this report.
Response: We agree. We have cited the WMO report and added a short paragraph in the Introduction explaining how the present study relates to its recommendations, particularly the need for fit-for-purpose evaluation, transparent reporting of sensor characteristics, field validation under representative environmental conditions, and caution when LCS data are used without site-specific calibration. This discussion has been added around the current lines 55–69.
- Materials and methods
The Vaisala AQT sensor is not described with the same level of detail as the other LCSs in this section. I recommend providing a proper introduction of the Vaisala AQT sensor.
Response: We agree. We have expanded the Materials and methods section to introduce the Vaisala AQT530 at the same level of detail as ASE and SEP, including its integrated meteorological, electrochemical gas, and optical particulate-matter measurements, its proprietary processing/calibration, and the variables used in this study.
No specifications are provided regarding measurement uncertainty, precision, or stability for the three sensors. These characteristics are usually reported in the instrument manuals. Please include them in Table 1 for each measured variable.
Response: We agree. We have expanded Table 1 with the manufacturer-declared uncertainty/accuracy, precision or repeatability, and stability/drift information wherever these quantities are available in the relevant manuals or datasheets. Where a manufacturer does not report a quantity for a specific variable, the table explicitly states "not specified by the manufacturer" rather than inferring a value. This revision is also discussed critically in the Conclusions. Where exact firmware and hardware revisions could not be retrospectively confirmed, the manuscript now states explicitly that specifications refer to manufacturer datasheets current at the time of purchase (2023) and may not capture subsequent firmware updates.
Please specify what Ox stands for.
Response: We agree. Ox is defined as the sum of O3 and NO2; we have added this definition at the first occurrence of the term (in the Abstract). The Alphasense OX-A431 electrochemical sensor is sensitive primarily to O3 and NO2, so its signal is the sum of those two species. We have also clarified that AirSensEUR, which uses this sensor, reports its raw channel response as Ox, whereas the Smart Environment PRO manufactured by Libelium, even though it uses the same sensor, treats the channel as O3 under its proprietary calibration assumptions. This distinction is reasonable because O3 concentrations are typically much higher than NO2, so a bias would only arise when NO2 levels are comparatively high. This clarification has been added at the first appearance of Ox and in Table 1.
Line 85: Please clarify what is meant by “open source” with reference to the ASE. Does this represent an added value compared to the other sensors?
Response: We agree that the wording was ambiguous. "Open source" refers to the AirSensEUR platform architecture, for which hardware/software documentation and data-handling tools are openly available, rather than to the individual sensing elements. We have clarified this and stated that the main added value is transparency and user control over acquisition and processing; it does not by itself guarantee superior metrological performance. The sentence at the current line 85 has been revised accordingly.
Line 106: More details on the site setup (including ITAF and ARTA) are needed. Please specify sampling heights above ground level and (for your site) presence of nearby obstacles.
Response: We agree and have expanded the description of the experimental setup. Based on our recollection of the field deployment, the LCS platforms and the OM205 ozone monitor were installed at approximately 4–5 m above ground level, with a free field of view of over 180° toward the airport, while the wall of the residential building was located behind the measurement setup. These figures are approximate, as no measured record of installation height was retained during deployment. We have also clarified that the comparison sites were approximately 0.87 km from the ARTA station and 1.17 km from the ITAF station and therefore did not constitute a strict collocation.
Table 1: Include declared measurement uncertainty, precision, and stability. Add the acronyms used throughout the manuscript (AQT, SEP, ASE). Ensure consistent use of sensor naming (acronyms vs full names). Include the measurement principle/sensor type for each parameter
Response: We agree. We have reorganized Table 1 to include the acronyms AQT, ASE and SEP, used consistent instrument names throughout, identified the measurement principle/sensor type for each variable, and reported manufacturer-declared uncertainty/accuracy, precision/repeatability and stability/drift where available. Missing specifications are explicitly marked as not reported.
Line 108: Please add definitions and formulas for the statistical indicators used.
Response: We agree. We have added definitions and equations for Pearson’s correlation coefficient (R), normalized mean squared error (NMSE), fractional bias (FB), factor of two (FA2), mean bias, and the slope of the least-squares regression. We have briefly explained the aspect of performance represented by each metric and identified the ideal value or range. This material has been inserted immediately before Section 2.1.
Line 114: Specify that daily averages are only calculated for PM10 and PM2.5.
Response: We agree. We have revised the sentence to state explicitly that hourly averaging was used for the gaseous-pollutant/reference comparisons, whereas daily averaging was applied only to PM2.5 and PM10 because the ARTA particulate-matter data were available as daily means.
Line 115: Clarify what is meant by “the two subsequent analyses”.
Response: We agree. We have clarified that "the two subsequent analyses" refers to the workday-versus-holiday diurnal-cycle analysis and the wind-direction/wind-speed analysis. We have clarified that the diurnal-cycle analysis uses a composite average: observations are grouped into 5-minute time-of-day bins (e.g., 08:00–08:05) and averaged across all workdays, and separately across all holidays, rather than averaged sequentially along the time series. The wind analysis instead uses a standard sequential 30-minute averaging of the continuous time series, to match the temporal resolution of the wind data. Both procedures have been stated explicitly at the current lines 115–117.
Section 2.1.1
Equation (1): Which temperature (T) and relative humidity (RH) data are used? If they are taken from ASE, discuss how the detected inaccuracies in T and RH retrieval could affect gas concentration estimates.
Response: T and RH in Eq. (1) are temperature and relative humidity recorded by ASE. We have stated this explicitly and have identified this as a limitation: the model can exploit covariance with the local environmental signal, but the physical interpretation and portability of the coefficients are reduced when the covariates themselves are biased, since the offsets and nonlinear behavior observed in the ASE meteorological channels may propagate into the stepwise calibration coefficients. This clarification has been added in Section 2.1.1 and revisited in the Discussion.
Line 132 (and throughout the manuscript): Please reconsider the use of excessive significant digits when reporting statistical indicators.
Response: We agree. We have rounded statistical indicators consistently to a meaningful precision, generally two decimal places (with additional digits retained only where needed to distinguish very small values), throughout the tables, figures and running text.
Section 3.1.1
Line 142: The phrase “(from 0.8767… respectively)” is unclear. Please specify how the “average” is calculated.
Response: We agree. We have rewritten the sentence without the confusing parenthetical comparison, reporting the values instrument by instrument and, where an average across AQT and the two SEP units is useful, explicitly stating which instruments are included and how the arithmetic mean was calculated. This revision concerns the current lines 140–145.
Figure 3: The comparison between ITAF and ASE/SEP shows evident non-linearity, which implies a concentration-dependent bias. This should be clearly emphasized. Moreover, the use of a linear model to assess the performance of the low cost sensors for RH should be critically discussed, as it may not be consistent with the observed behavior.
Response: We agree. We have explicitly described the curvature in the ITAF–ASE and ITAF–SEP comparisons, particularly for RH, as evidence of a concentration-/humidity-dependent bias, and clarified that the linear fit is retained only as a compact descriptive metric that does not represent the full response function. The limitations of R, slope and linear regression for this nonlinear behavior are discussed in Section 3.1.1, and the Conclusions avoid characterizing high correlation as equivalent to accuracy.
Section 3.1.2
Line 160: In addition to SEP NO₂, NO also appears to be poorly reproduced by SEP.
Response: We agree. We have revised the discussion to state that SEP performs poorly not only for NO2 but also for NO in terms of absolute agreement, despite a moderate correlation coefficient, emphasizing the very large bias/NMSE and the mismatch in scale.
Line 170: The statement “The bias for AQT and ASE are generally low” is not supported by Fig. 5. The figure shows large biases (e.g., >25 ppb for NO2 at higher concentrations) for both AQT and ASE. Similar issues are observed for NO.
Response: We agree. Rather than removing the statement, we have qualified it: the manuscript now explicitly notes that the tabulated mean bias for AQT and ASE is numerically low for most compounds, but that this summary metric masks a concentration-dependent residual clearly visible in Fig. 5, with both instruments increasingly underestimating NO and NO2 at higher reference concentrations. The manuscript states that moderate-to-high R values therefore coexist with substantial, concentration-dependent bias rather than uniform agreement. This revision has been made in Section 3.1.2 and in the Conclusions.
Figure 4: Some fixed values appear in the OM205 time series. Please clarify their origin and whether these values were excluded from the comparison analysis.
Response: We agree. We decided to retain these values in the time series plots after inspection of the quality-control flags and raw records confirmed them to be acquisition artifacts. Where OM205 data showed fixed values for short periods, we left them visible in the time series to document the issue transparently but excluded that data from the comparison analysis.
Figure 5: The NO correlation appears strongly influenced by a few high-concentration data points. I suggest repeating the analysis limiting NO values to 0–30 ppb and discuss differences (if any). In the O3 AQT plot there are some fix values for OM205 at around 20 ppb. I think they should be removed.
Response: We agree and have repeated the NO comparison after restricting the reference range to 0–30 ppb (Supplement, Table S3 and Figure S8). This sensitivity test showed no visual improvement in the scatter plots; the recalculated statistical indices were, for the most part, close to the original ones, although some deterioration was observed: R decreased from 0.76 to 0.45 for AQT and from 0.78 to 0.58 for ASE, while the AQT slope decreased from 0.56 to 0.38. This suggests that these LCS more accurately reproduce NO concentration peaks than lower, more ordinary values. The fixed OM205 values near 20 ppb were investigated and confirmed as acquisition artifacts; they were removed from the quantitative comparison, as described in our response above.
Section 3.1.3
Line 201: Please specify which differences are being referred to.
Response: We agree. We have rewritten the sentence to specify that the OPC-N3/SEP showed closer agreement with ARTA data for PM10 than for PM2.5, while AQT substantially underestimated both size fractions and the PMS5003/ASE generally overestimated them.
I recommend including a summary table for PM2.5 and PM10 comparisons, reporting mean differences (with min–max range) and standard deviation of differences
Response: We agree and have added a summary table (Table 5) reporting the PM2.5 and PM10 comparison with ARTA for each instrument, including the mean sensor-minus-reference difference, the standard deviation of the differences, and the minimum–maximum range, based on daily means over the 19 days of paired data.
Section 3.2
Line 206: Better introduce Figures 8 and 9, clearly explaining their content.
Response: We agree. Section 3.2 now begins with a clearer explanation that Figs. 8 and 9 show composite diel cycles calculated separately for workdays and holidays, using 5-min time-of-day bins across the available days; the lines/shading and the number of contributing days are defined in the text and captions. We have also stated that the analysis is exploratory because only five holidays are available.
Figure 8: If the evening NO peak is attributed to traffic emissions, why do NO2 and CO not decrease during holidays compared to weekdays? The NO2/NO ratio changes between weekdays and holidays. Could this indicate changes in emission sources?
Response: We agree that the original traffic-only interpretation was too strong. In the revised manuscript, we have pointed out that, while traffic remains the main contributor to the NO peak (confirmed by its evident decrease during holidays), NO2 and CO do not show a comparable reduction during holidays. This may reflect the role of other sources, such as the airport, and changes in aging, source mix, oxidation chemistry, background concentrations, and meteorological dispersion, rather than a uniform reduction of traffic emissions. We have discussed the changed NO2/NO ratio as suggestive, but not conclusive, evidence of a different balance between fresh combustion emissions and more aged/secondary NO2.
The diurnal ozone peak is likely influenced by vertical mixing and entrainment from higher atmospheric layers under conditions of strong atmospheric mixing. This interpretation is supported by Fig. 10, where ozone behaves differently compared to primary pollutants. Including wind speed data from ITAF in the plot would help disentangle the role of boundary layer dynamics.
Response: We agree and have added vertical mixing and entrainment from residual/background layers as plausible contributors to the daytime ozone maximum, in addition to photochemical production. Wind-direction/speed composites stratified by workday and holiday have been generated and are provided in the Supplement (Figures S9–S12); the discussion explicitly separates meteorological and emission-related interpretations, and notes that disentangling boundary-layer dynamics from photochemistry would require dedicated vertical-profile measurements not available in this campaign.
Section 3.3
Line 248: The attribution to traffic emissions from the E80 corridor appears still consistent with south-west winds. Have the authors considered differences in traffic flow directions between morning and evening rush hours?
Response: We agree. We have acknowledged in the revised text that south-westerly winds can also be consistent with transport from the E80 corridor and that traffic direction/intensity may differ between morning and evening rush periods. Because direction-resolved traffic counts are not available for the campaign, we have presented this as an alternative interpretation rather than a resolved attribution.
Line 250: Please clarify what is meant by “This is expected for the NOx/Ozone daily cycle.”
Response: We agree. We have replaced the sentence with a specific explanation: primary NOx tends to accumulate under weak dispersion and fresh-emission conditions, whereas O3 is depleted by NO titration and often increases with stronger daytime mixing and photochemistry. We have also stated that this qualitative pattern does not uniquely identify a source.
Lines 252–255: It appears that AQT results are more consistent with ASE calibrated data, while SEP aligns better with ASE raw data. Please add some discussions about this observation.
Response: We agree. We have added a discussion explaining that the apparent similarity partly reflects data processing: the calibrated ASE gas series was trained against AQT, whereas the raw ASE response can preserve a scale/shape that happens to resemble SEP for some pollutants. Consequently, agreement between AQT and calibrated ASE cannot be treated as an independent validation, and similarity between SEP and raw ASE does not imply accuracy relative to the reference. This is now stated in Sections 2.1.1, 3.1.2 and 3.3.
Conclusions
In general, it should be interesting that you critically discuss the performance of the three sensors with the characteristics provided by the manufacturer in the manual/data sheet in terms of declared measurement uncertainties.
Response: We agree. The Conclusions now compare observed field performance with manufacturer-declared specifications, while clearly noting that datasheet values are often obtained under controlled conditions and may not be directly transferable to this non-collocated field deployment. We have distinguished correlation, bias, precision and agreement, and avoided using the specifications as pass/fail criteria where test conditions differ.
Line 265: The concentration-dependent bias observed for NO and NO2 (particularly for AQT and ASE) should be explicitly mentioned.
Response: We agree. The Conclusions now explicitly state that NO and NO2 show concentration-dependent bias, particularly for AQT and ASE, with increasing departure from the 1:1 line at higher concentrations despite moderate-to-high correlation coefficients.
Line 270: Please provide possible explanations for the different behavior observed in PM10 (overestimation) versus PM2.5 (underestimation). Additional information on PM composition at the site would be valuable if available: the authors should discuss whether their findings can be generalized to environments with different PM composition.
Response: We agree. The discussion now explains that different PM10 and PM2.5 behavior can arise from size-dependent optical response, particle composition/refractive index, density, hygroscopic growth, and the manufacturer's conversion algorithm. We have stated that no speciated PM composition was measured during this campaign, so a composition-based explanation cannot be tested and the results should not be generalized to sites or seasons with different aerosol mixtures without additional validation.
Technical Comments
The reference list formatting is inconsistent and should be standardized.
Response: We agree. We have regenerated the bibliography from a standardized BibTeX database, with consistent author, year, title, journal, volume, pages and DOI formatting according to the Copernicus style.
Line 3: “Harwey” (year missing) should be verified; it does not appear to be a peer-reviewed reference. Please also check formatting issues in other references (e.g., “Organization”, “of Science et al.”).
Response: We agree and have corrected the malformed opening references and replaced them with complete, verifiable sources. The bibliography has been regenerated from a standardized BibTeX database, with organization names (e.g., World Health Organization, Academy of Science of South Africa) enclosed in double braces to ensure correct rendering, and all in-text citations have been checked against the reference list for consistency.
In several cases, references are written as “XXXX at al., yyyy”. The correct format should be “XXXX et al. (yyyy)”.
Response: We agree. We have standardized all in-text citations through the appropriate LaTeX citation commands so that narrative citations appear as "Author et al. (year)" and parenthetical citations follow the journal style. Typographical instances of "at al." have been corrected.
Figure 5: Some axis labels are partially obscured. Please shift the SEP O3 plot to align more clearly with the ASE Ox plot.
Response: We agree, and have regenerated Figure 5 with adequate margins and consistently aligned panels so that no axis labels are clipped or obscured; the SEP O3 and ASE Ox panels are now aligned.
Figures 8–9: Explain what the shaded areas represent in the captions. Ensure that color schemes are accessible to color-blind readers
Response: We agree. The captions of Figures 8 and 9 now define the shaded areas explicitly as the standard deviation around the composite mean, and state the number of contributing days (n=14 workdays, n=5 holidays) and the time resolution (5-minute time-of-day bins).
Citation: https://doi.org/10.5194/egusphere-2026-1262-AC1
-
AC1: 'Reply on RC1', Lorenzo Gentile, 18 Sep 2026
-
RC2: 'Comment on egusphere-2026-1262', Anonymous Referee #2, 24 Jul 2026
Overall comments
In “Intercomparison of low-cost sensors via simultaneous atmospheric measurements: a case study,” the authors evaluate the performance of three low-cost sensor (LCS) models during a field campaign in central Italy. They also consider the ability of the factory-calibrated LCS to capture pollutant patterns in an area with several pollution sources.
The manuscript contributes useful information about the field performance of three LCS devices in one specific environment. It also provides a valuable perspective about the data quality of off-the-shelf (un-calibrated) sensors and what uses they are appropriate for.
However, the manuscript does not adequately address some areas of uncertainty including (1) the distance between the LCS and reference site and (2) the limited sample size of measurements for the analysis of holiday/weekday patterns and wind direction.
Additionally, while the study aims to assess factory-calibrated sensors, one of the three sensor models (ASE) was calibrated directly against a second sensor (AQT) in the field, undermining independent evaluation of the ASE. There also appear to be some discrepancies in the data presented that should be reconciled (see my specific comments about Table 4 below).
Therefore, I would recommend the following revisions:
Specific comments
Title
To distinguish this paper, consider including in the title:
- the location of measurements
- specify that the paper evaluates factory-calibrated sensors
Abstract
The abstract could use more nuance about the performance results of the sensors. While R statistics show relatively good agreement, the un-corrected sensors have large bias (especially for NO2).
Introduction
Line 55: DeSouza et al. (2022) and Peters et al. (2022) might also be good citations that focus on the role of calibration in achieving air sensor objectives
Line 61: Zimmerman et al. (2018) is also a relevant citation here
Line 65: Best practices typically include calibrating or correcting sensors. It would strengthen the framing to add more justification about why you decided to use un-corrected data from sensors and how it contributes to the aims of the research.
Line 65-69: It would be valuable to add citations that put this study into context – are there other works that test these sensor models?
Materials and methods
Line 88: Please clarify what you mean by black-box. Does this mean a proprietary algorithm?
Line 105: The distance between the sites, especially with known nearby emissions sources, introduces uncertainty into the performance assessment as it is not a true collocation between the LCS and reference. For example, one site could be downwind of airport or highway emissions at a time when the other is not. This could introduce noise (if levels between sites vary differently) or bias (if one site tends to face higher pollution). The authors should include a discussion of this uncertainty and whether sensitivity analyses were performed to understand the impacts of distance between sites.
Line 109: The statistics named very briefly here are an important part of the study. It would be beneficial to introduce each statistic and the type of information it provides so the reader can more readily interpret the results tables.
Line 110: Why was a 3-sigma filter applied? Were these suspected measurement artifacts? If valid measurements were excluded it could affect the performance results.
Section 2.1.1: The calibration of the ASE sensor using a different sensor (AQT) complicates the performance evaluation results and the framing of the paper. While the paper is stated to examine un-calibrated sensors, in reality the ASE sensor has been calibrated to match the results of the AQT sensor as closely as possible. This would not be possible when buying ASE off the shelf. As a result it is not surprising that the gas performance results for ASE and AQT are very similar. To truly evaluate off-the-shelf performance as the study aims, it would make more sense to evaluate the digital signal directly (does it correlate and show trends properly) or use a correction from the literature. This way the ASE could truly be evaluated independently of the AQT.
Results
Line 181: Is there any evidence in the literature of similar effects on this sensor?
Figure 6: The 1:1 line is in black, not red. Also make sure the x axis labels are not cut off.
Line 185: The agreement does not look so good between SEP and AQT – there is a nonlinear shape in the scatter plot.
Table 3: Similar to my comment above, it is not surprising that AQT and ASE have very similar results because ASE has been calibrated to match AQT. It is difficult to see this as an independent evaluation of ASE.
Table 4: Why are statistics for AQT vs ASE CO comparisons different than Table S2? The AQT vs SEP R value is highest in the table although these two variables appear to have worse agreement in Figure 6 scatter. Please double check that these are the correct statistics for each pair of sensors.
Figure 7: The daily column chart is useful, but it is difficult to assess performance in aggregate. I’d suggest including a supplemental table with summary statistics. This would support the statements on line 195.
Figure 8: Given that the number of holidays is small (5 days), it is difficult to conclude that the difference between days is caused by emissions patterns and not changes in weather that happened to occur on certain days. Did the authors perform any sensitivity analyses to understand the potential influence of weather (e.g. increased dispersion that reduces pollutant concentrations)? Including this and discussing the potential role of weather would strengthen the argument that emission differences were detected.
Figure 8: Specify what the shaded region represents.
Section 3.3 Wind Analysis: Have the authors filtered the wind analysis to ensure a minimum number of data points per wind speed/direction bin? In some cases, the points on the figure might represent only a single moment in time, making generalization difficult. Consider filtering to a minimum number of hours to improve the robustness.
Line 241: The higher concentrations at low wind speeds are also consistent with the buildup of pollution during periods of stagnation overnight
Conclusions
The discussion of sensor performance results could use more clarity and specificity to inform the reader. Currently it is qualitative and discusses some positives (good correlation) and some negatives (poor agreement with ARTA). The authors could consider comparing performance results to benchmarks (such as those from US EPA or defined by the authors) to provide a more definitive judgement of how each sensor performed in terms bias and accuracy.
Line 274: The discussion of how the sensors can be used despite imperfect data is useful. It would be interesting to also consider analyses that would not be appropriate based on the data quality (like comparisons to regulatory or health thresholds).
Minor comments
Decimals should be rounded throughout to something meaningful like 2 places
I would suggest specifying in tables and figure captions the time resolution of the data that were used to generate the figure or statistics.
Check figure captions to ensure the correct colors are used to describe features (like the 1:1 line).
Citation: https://doi.org/10.5194/egusphere-2026-1262-RC2 -
AC2: 'Reply on RC2', Lorenzo Gentile, 18 Sep 2026
Overall comments
In “Intercomparison of low-cost sensors via simultaneous atmospheric measurements: a case study,” the authors evaluate the performance of three low-cost sensor (LCS) models during a field campaign in central Italy. They also consider the ability of the factory-calibrated LCS to capture pollutant patterns in an area with several pollution sources.
The manuscript contributes useful information about the field performance of three LCS devices in one specific environment. It also provides a valuable perspective about the data quality of off-the-shelf (un-calibrated) sensors and what uses they are appropriate for.
However, the manuscript does not adequately address some areas of uncertainty including (1) the distance between the LCS and reference site and (2) the limited sample size of measurements for the analysis of holiday/weekday patterns and wind direction.
Additionally, while the study aims to assess factory-calibrated sensors, one of the three sensor models (ASE) was calibrated directly against a second sensor (AQT) in the field, undermining independent evaluation of the ASE. There also appear to be some discrepancies in the data presented that should be reconciled (see my specific comments about Table 4 below).
Response: We thank the reviewer for identifying these central limitations. The revised manuscript now more clearly distinguishes (i) evaluation of factory-processed AQT and SEP outputs from (ii) conversion/calibration of ASE raw digital signals, and no longer presents the ASE gas results as an independent off-the-shelf validation. The non-collocated design and the limited holiday/wind-bin sample have been treated as explicit sources of uncertainty. We have also verified and reconciled the statistics in Table 4 and the Supplement.
Therefore, I would recommend the following revisions:
Specific comments
Title
To distinguish this paper, consider including in the title:
- the location of measurements
- specify that the paper evaluates factory-calibrated sensors
Response: We agree. We propose revising the title to: “Field intercomparison of factory-processed low-cost air-quality sensors in central Italy: a case study.” This title specifies the field setting and location while avoiding the inaccurate implication that all three devices provide factory-calibrated concentration outputs, since ASE reports raw digital signals for the gas channels. The final wording has been checked for consistency with the revised framing.
Abstract
The abstract could use more nuance about the performance results of the sensors. While R statistics show relatively good agreement, the un-corrected sensors have large bias (especially for NO2).
Response: We agree. We have rewritten the Abstract to distinguish correlation from agreement and to report the substantial biases and departures from the 1:1 relationship, especially for NO2 and SEP gas measurements. It now also clarifies that ASE gas signals required an empirical conversion using AQT and therefore do not constitute an independent off-the-shelf performance test.
Introduction
Line 55: DeSouza et al. (2022) and Peters et al. (2022) might also be good citations that focus on the role of calibration in achieving air sensor objectives
Response: We agree that the calibration literature should be strengthened. We have evaluated and added DeSouza et al. (2022) and Peters et al. (2022) where they directly support the role of calibration/correction in achieving intended sensor objectives. The exact references and statements have been verified before inclusion in the BibTeX database.
Line 61: Zimmerman et al. (2018) is also a relevant citation here
Response: We agree. We have cited Zimmerman et al. (2018) in the discussion of data-driven/multivariate calibration approaches, placing this study within the broader development of field calibration methods.
Line 65: Best practices typically include calibrating or correcting sensors. It would strengthen the framing to add more justification about why you decided to use un-corrected data from sensors and how it contributes to the aims of the research.
Response: We agree. The Introduction now states that best practice commonly includes collocation and correction, but that the original objective was to assess the information obtainable from devices as supplied/processed by their manufacturers before site-specific correction. We have revised this framing because ASE is a special case: its gas channels do not provide calibrated concentrations and required conversion. The revised manuscript therefore distinguishes "factory-processed outputs" from "raw signals converted empirically" and describes the study as an exploratory field intercomparison rather than a uniform test of uncorrected sensors.
Line 65-69: It would be valuable to add citations that put this study into context – are there other works that test these sensor models?
Response: We agree. We have added relevant studies evaluating the same or closely related platforms/sensor modules (AQT530, AirSensEUR, Libelium/Alphasense configurations, PMS5003 and OPC-N3), while carefully distinguishing platform-level evaluations from studies of individual sensing elements. Existing citations already used in the manuscript have been reorganized to make this context clearer, and additional references have been verified before inclusion.
Materials and methods
Line 88: Please clarify what you mean by black-box. Does this mean a proprietary algorithm?
Response: We agree. We have replaced "black-box" with "manufacturer-proprietary internal conversion algorithm," meaning that the PMS5003 provides a factory-processed output but the coefficients and full algorithm are not disclosed to the user.
Line 105: The distance between the sites, especially with known nearby emissions sources, introduces uncertainty into the performance assessment as it is not a true collocation between the LCS and reference. For example, one site could be downwind of airport or highway emissions at a time when the other is not. This could introduce noise (if levels between sites vary differently) or bias (if one site tends to face higher pollution). The authors should include a discussion of this uncertainty and whether sensitivity analyses were performed to understand the impacts of distance between sites.
Response: We agree and have explicitly stated that this is a near-site intercomparison rather than a regulatory collocation, noting that spatial separation can introduce both random mismatch and systematic bias when plumes from roads, the airport, or residential sources affect the two sites differently. As a partial check, periods were qualitatively stratified by wind direction relative to the relative positions of our site and the ARTA/ITAF stations, to assess whether the largest LCS-reference discrepancies co-occurred with wind directions likely to place the sites on opposite sides of a local plume. No systematic, direction-dependent amplification of bias was identified; however, given the relatively short distance between sites (<1.2 km) and the absence of co-located 3D wind measurements, this check is presented as indicative rather than a formal sensitivity analysis.
Line 109: The statistics named very briefly here are an important part of the study. It would be beneficial to introduce each statistic and the type of information it provides so the reader can more readily interpret the results tables.
Response: We agree. We have added definitions, equations, interpretation and ideal values for all performance statistics before the Data Preparation subsection, as described in our response to RC1. This allows the reader to distinguish association (R), normalized error (NMSE), mean directional bias (FB/bias), factor-of-two agreement (FA2) and scale response (slope).
Line 110: Why was a 3-sigma filter applied? Were these suspected measurement artifacts? If valid measurements were excluded it could affect the performance results.
Response: The 3-sigma filter was intended to remove isolated values suspected to be instrument/acquisition artifacts after flag checking, not to truncate valid pollution episodes. We have quantified its impact: the filter removed at most 0.007% of the remaining data points for any given variable, and in most cases removed none at all. Given this negligible impact, we consider the filter efficient at removing artifacts without materially affecting the dataset, and no unfiltered-data sensitivity comparison was deemed necessary.
Section 2.1.1: The calibration of the ASE sensor using a different sensor (AQT) complicates the performance evaluation results and the framing of the paper. While the paper is stated to examine un-calibrated sensors, in reality the ASE sensor has been calibrated to match the results of the AQT sensor as closely as possible. This would not be possible when buying ASE off the shelf. As a result it is not surprising that the gas performance results for ASE and AQT are very similar. To truly evaluate off-the-shelf performance as the study aims, it would make more sense to evaluate the digital signal directly (does it correlate and show trends properly) or use a correction from the literature. This way the ASE could truly be evaluated independently of the AQT.
Response: We agree with the reviewer's central concern. ASE gas channels are not factory-calibrated concentration outputs; our conversion against AQT makes the resulting ASE–AQT similarity non-independent. We have revised the aims, methods, results and conclusions accordingly, and the main analysis now clearly labels ASE concentrations as AQT-referenced empirical conversions rather than independent evidence of accuracy. We have additionally evaluated the raw, uncalibrated ASE digital signals against ARTA, OM205, and AQT (Supplement, Section S2, Figures S2–S3): this comparison showed strong correlation with AQT (|R| ≥ 0.86) and correlation coefficients close to the calibrated counterparts when compared against the independent reference instruments, offering a more thorough evaluation of the ASE.
Results
Line 181: Is there any evidence in the literature of similar effects on this sensor?
Response: We searched the literature for evidence of humidity interference and failure modes specific to the Alphasense NO2-A43F/OX-family electrochemical sensors, but were unable to find direct confirmation of this specific failure mode. We have therefore softened our water-ingress hypothesis in the manuscript, explicitly labelling it as a plausible but unconfirmed interpretation rather than a documented effect.
Figure 6: The 1:1 line is in black, not red. Also make sure the x axis labels are not cut off.
Response: We agree. We have corrected the Figure 6 caption to state that the 1:1 line is black, and adjusted the figure margins so that all x-axis labels are fully visible.
Line 185: The agreement does not look so good between SEP and AQT – there is a nonlinear shape in the scatter plot.
Response: We agree. The revised text no longer describes the SEP–AQT agreement as uniformly good. We have stated that the time series are strongly correlated but the scatter is nonlinear, indicating scale-dependent disagreement, and have reported this distinction together with the relevant error and slope metrics.
Table 3: Similar to my comment above, it is not surprising that AQT and ASE have very similar results because ASE has been calibrated to match AQT. It is difficult to see this as an independent evaluation of ASE.
Response: We agree. Table 3 and its discussion now explicitly state that the ASE gas concentrations were calibrated against AQT and thus cannot provide an independent comparison with AQT; comparisons of ASE with ARTA/OM205 are also interpreted cautiously because the AQT-based transformation can transfer AQT behavior into the ASE series. The independent raw-signal evaluation described above has been added.
Table 4: Why are statistics for AQT vs ASE CO comparisons different than Table S2? The AQT vs SEP R value is highest in the table although these two variables appear to have worse agreement in Figure 6 scatter. Please double check that these are the correct statistics for each pair of sensors.
Response: We agree, and have fully audited Table 4 and Table S2 against a single version-controlled dataset. This audit revealed that the AQT–SEP and AQT–ASE pairs had been mislabeled in Table 4 of the submitted version: the statistics originally reported under "AQT vs ASE" corresponded to the AQT–SEP comparison, and vice versa. This has been corrected in the revised Table 4. We also confirmed that the remaining numerical differences between Table 4 and Table S2 are explained by the different temporal resolution used (hourly averages in Table 4, consistent with the ARTA and OM205 comparisons, versus original per-minute acquisition frequency in Table S2 for the AirSensEUR calibration); this distinction is now stated explicitly in the revised manuscript (Sect. 3.1.2). The corrected Table 4 values are now consistent with the scatter plots shown in Fig. 6.
Figure 7: The daily column chart is useful, but it is difficult to assess performance in aggregate. I’d suggest including a supplemental table with summary statistics. This would support the statements on line 195.
Response: We agree. We have added a supplementary PM summary table, as described in our response to RC1 (Table 5), with paired daily sample size, mean difference, standard deviation and min–max range for PM2.5 and PM10, providing aggregate support for the qualitative discussion of Fig. 7.
Figure 8: Given that the number of holidays is small (5 days), it is difficult to conclude that the difference between days is caused by emissions patterns and not changes in weather that happened to occur on certain days. Did the authors perform any sensitivity analyses to understand the potential influence of weather (e.g. increased dispersion that reduces pollutant concentrations)? Including this and discussing the potential role of weather would strengthen the argument that emission differences were detected.
Response: We agree. With only five holidays, the composite comparison is exploratory and cannot fully isolate emission changes from coincident meteorology. We have quantified and compared temperature, RH, and pressure between workdays and holidays (Supplement, Figure S14). The overall diurnal shape was similar between the two groups for all three variables, but holidays were characterized by systematically higher night-time temperature and RH (by approximately 1–2°C and a few percentage points RH) and by lower atmospheric pressure (by approximately 2–3 hPa) throughout the day. This indicates that the two periods did not share an identical meteorological regime, and that reduced nocturnal dispersion under these slightly milder, more humid conditions may have contributed, alongside emissions, to the higher night-time pollutant concentrations observed during holidays. We have revised the text to describe these differences as observed associations rather than confirmed causal effects, given the limited holiday sample size.
Figure 8: Specify what the shaded region represents.
Response: We agree. The Fig. 8 caption now explicitly defines the shaded region as the standard deviation, states how it was calculated, and reports the number of workdays/holidays contributing to each composite.
Section 3.3 Wind Analysis:
Response: We agree. The wind-direction analysis reports the number of observations in each speed/direction bin and applies a minimum-count threshold of four observations before displaying or interpreting a bin; bins below this threshold are masked, and the threshold and temporal resolution are stated in the Methods and figure captions. Given time constraints, sensitivity to alternative thresholds was not formally tested; we have noted this explicitly as a caveat, particularly relevant for sparsely populated bins.
Line 241: The higher concentrations at low wind speeds are also consistent with the buildup of pollution during periods of stagnation overnight
Response: We agree. The discussion now adds that high concentrations at low wind speed are consistent with reduced dispersion and overnight/stagnation buildup, not necessarily proximity to a directional source; this interpretation is presented separately from the source-direction arguments.
Conclusions
The discussion of sensor performance results could use more clarity and specificity to inform the reader. Currently it is qualitative and discusses some positives (good correlation) and some negatives (poor agreement with ARTA). The authors could consider comparing performance results to benchmarks (such as those from US EPA or defined by the authors) to provide a more definitive judgement of how each sensor performed in terms bias and accuracy.
Response: We agree that the conclusion needed clearer performance criteria. We have provided a concise instrument-by-variable synthesis separating correlation, bias, error and concentration-dependent behavior, and compared observed performance against manufacturer-declared datasheet specifications where applicable. We did not perform a formal comparison against regulatory or fit-for-purpose benchmarks (e.g. US EPA sensor performance metrics), because the averaging times, concentration ranges and testing protocols underlying those benchmarks differ substantially from this non-collocated field campaign; we have explained explicitly why such a pass/fail comparison would be inappropriate here.
Line 274: The discussion of how the sensors can be used despite imperfect data is useful. It would be interesting to also consider analyses that would not be appropriate based on the data quality (like comparisons to regulatory or health thresholds).
Response: We agree. The Conclusions now state explicitly that these data are not suitable for regulatory compliance assessment, direct comparison with health-based limit values, or precise exposure estimation without appropriate collocation, calibration and uncertainty characterization. More defensible uses include detecting broad temporal patterns, identifying episodes, and exploratory source-pattern analysis, subject to the limitations described.
Minor comments
Decimals should be rounded throughout to something meaningful like 2 places
Response: We agree. We have standardized decimal precision throughout, generally to two decimal places for statistical results, with scientifically justified exceptions.
I would suggest specifying in tables and figure captions the time resolution of the data that were used to generate the figure or statistics.
Response: We agree. Tables and figure captions now state the temporal resolution and aggregation used (raw/5-min, 30-min, hourly or daily) and the number of paired observations or days where relevant.
Check figure captions to ensure the correct colors are used to describe features (like the 1:1 line).
Response: We agree. We have checked all captions against the regenerated figures so that line colors/styles, the 1:1 line, regression line and uncertainty intervals are described correctly and consistently.
Citation: https://doi.org/10.5194/egusphere-2026-1262-AC2
Viewed
| HTML | XML | Total | Supplement | BibTeX | EndNote | |
|---|---|---|---|---|---|---|
| 444 | 250 | 59 | 753 | 82 | 38 | 49 |
- HTML: 444
- PDF: 250
- XML: 59
- Total: 753
- Supplement: 82
- BibTeX: 38
- EndNote: 49
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
This work by Gentile at al. pesents an interesting intercomparison exercise between low-cost sensors for meteorological and air pollution measurements. The topic is relevant for AMT, particularly in light of the continuously growing interest in low-cost technologies for air quality monitoring and research applications.
In addition to provide insights into the quality and reliability of data from three specific commercial low-cost sensors, the authors also discuss their fitness for purpose by exploring selected use cases, such as the investigation of the typical diel variability of pollutants and source attribution through combined analysis with near-surface wind variability.
The manuscript is generally well organized and easy to follow. I recommend publication after addressing the following technical and minor comments.
Specific Comments
1. Introduction
The following WMO report on the use of low-cost sensors in air quality monitoring networks should be included to strengthen the reference list: https://library.wmo.int/records/item/68924-integrating-low-cost-sensor-systems-and-networks-to-enhance-air-quality-applications
It would be valuable if the authors discuss how their work relates to the findings and recommendations presented in this report.
2. Materials and methods
The Vaisala AQT sensor is not described with the same level of detail as the other LCSs in this section. I recommend providing a proper introduction of the Vaisala AQT sensor.
No specifications are provided regarding measurement uncertainty, precision, or stability for the three sensors. These characteristics are usually reported in the instrument manuals. Please include them in Table 1 for each measured variable.
Please specify what Ox stands for.
Line 85: Please clarify what is meant by “open source” with reference to the ASE. Does this represent an added value compared to the other sensors?
Line 106: More details on the site setup (including ITAF and ARTA) are needed. Please specify sampling heights above ground level and (for your site) presence of nearby obstacles.
Table 1: Include declared measurement uncertainty, precision, and stability. Add the acronyms used throughout the manuscript (AQT, SEP, ASE). Ensure consistent use of sensor naming (acronyms vs full names). Include the measurement principle/sensor type for each parameter
Line 108: Please add definitions and formulas for the statistical indicators used.
Line 114: Specify that daily averages are only calculated for PM10 and PM2.5.
Line 115: Clarify what is meant by “the two subsequent analyses”.
Section 2.1.1
Equation (1): Which temperature (T) and relative humidity (RH) data are used? If they are taken from ASE, discuss how the detected inaccuracies in T and RH retrieval could affect gas concentration estimates.
Line 132 (and throughout the manuscript): Please reconsider the use of excessive significant digits when reporting statistical indicators.
Section 3.1.1
Line 142: The phrase “(from 0.8767… respectively)” is unclear. Please specify how the “average” is calculated.
Figure 3: The comparison between ITAF and ASE/SEP shows evident non-linearity, which implies a concentration-dependent bias. This should be clearly emphasized. Moreover, the use of a linear model to assess the performance of the low cost sensors for RH should be critically discussed, as it may not be consistent with the observed behavior.
Section 3.1.2
Line 160: In addition to SEP NO₂, NO also appears to be poorly reproduced by SEP.
Line 170: The statement “The bias for AQT and ASE are generally low” is not supported by Fig. 5. The figure shows large biases (e.g., >25 ppb for NO2 at higher concentrations) for both AQT and ASE. Similar issues are observed for NO.
Figure 4: Some fixed values appear in the OM205 time series. Please clarify their origin and whether these values were excluded from the comparison analysis.
Figure 5: The NO correlation appears strongly influenced by a few high-concentration data points. I suggest repeating the analysis limiting NO values to 0–30 ppb and discuss differences (if any). In the O3 AQT plot there are some fix values for OM205 at around 20 ppb. I think they should be removed.
Section 3.1.3
Line 201: Please specify which differences are being referred to.
I recommend including a summary table for PM2.5 and PM10 comparisons, reporting mean differences (with min–max range) and standard deviation of differences
Section 3.2
Line 206: Better introduce Figures 8 and 9, clearly explaining their content.
Figure 8: If the evening NO peak is attributed to traffic emissions, why do NO2 and CO not decrease during holidays compared to weekdays? The NO2/NO ratio changes between weekdays and holidays. Could this indicate changes in emission sources?
The diurnal ozone peak is likely influenced by vertical mixing and entrainment from higher atmospheric layers under conditions of strong atmospheric mixing. This interpretation is supported by Fig. 10, where ozone behaves differently compared to primary pollutants. Including wind speed data from ITAF in the plot would help disentangle the role of boundary layer dynamics.
Section 3.3
Line 248: The attribution to traffic emissions from the E80 corridor appears still consistent with south-west winds. Have the authors considered differences in traffic flow directions between morning and evening rush hours?
Line 250: Please clarify what is meant by “This is expected for the NOx/Ozone daily cycle.”
Lines 252–255: It appears that AQT results are more consistent with ASE calibrated data, while SEP aligns better with ASE raw data. Please add some discussions about this observation.
Conclusions
In general, it should be interesting that you critically discuss the performance of the three sensors with the characteristics provided by the manufacturer in the manual/data sheet in terms of declared measurement uncertainties.
Line 265: The concentration-dependent bias observed for NO and NO2 (particularly for AQT and ASE) should be explicitly mentioned.
Line 270: Please provide possible explanations for the different behavior observed in PM10 (overestimation) versus PM2.5 (underestimation). Additional information on PM composition at the site would be valuable if available: the authors should discuss whether their findings can be generalized to environments with different PM composition.
Technical Comments
The reference list formatting is inconsistent and should be standardized.
Line 3: “Harwey” (year missing) should be verified; it does not appear to be a peer-reviewed reference. Please also check formatting issues in other references (e.g., “Organization”, “of Science et al.”).
In several cases, references are written as “XXXX at al., yyyy”. The correct format should be “XXXX et al. (yyyy)”.
Figure 5: Some axis labels are partially obscured. Please shift the SEP O3 plot to align more clearly with the ASE Ox plot.
Figures 8–9: Explain what the shaded areas represent in the captions. Ensure that color schemes are accessible to color-blind readers