the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
Rainfall monitoring based on spectral analysis of sound recordings in different geographical locations
Abstract. Recent interest in geophony-based approaches for rainfall monitoring has highlighted its potential as a low-cost complement to traditional methods. Yet, most existing techniques rely on model training and lack cross-site validation. In this study, we introduce and evaluate a simple acoustic metric for rainfall characterization based on deviations in power spectral density from a local baseline under dry-weather conditions. The method was applied to eight datasets containing audio recordings from tropical and temperate forests, as well as urban and semi-urban environments, and validated using rain gauge measurements. Results show consistently high correlations (≈ 0.7) between the proposed metric and rainfall rate for the tropical (Amazonian and West African) datasets when low-frequency bands are chosen (0.2–0.5 kHz), as opposed to the
open-air dataset, which did not present significant correlation values. Additionally, for the temperate forest dataset (France), higher correlations are obtained on the higher part of the spectra (3.2–3.5 kHz), likely due to wind and technophonic-related interferences (jet engine background noise) at lower frequencies. Although common frequency bands are identified across multiple sites, the amplitude of the proposed metric varies substantially between regions, indicating that a universal physical relationship for direct rainfall rate estimation may be difficult to achieve. These results highlight both the potential and the limitations of a simple, training-free acoustic metric for rainfall monitoring across diverse environments.
- Preprint
(9568 KB) - Metadata XML
- BibTeX
- EndNote
Status: open (until 07 Oct 2026)
- RC1: 'Comment on egusphere-2026-3457', David Dunkerley, 14 Sep 2026 reply
-
RC2: 'Comment on egusphere-2026-3457', Anonymous Referee #2, 02 Oct 2026
reply
General
The manuscript introduces a method to obtain quantitative estimates of rainfall from acoustic sensors. Strong aspects of the study are the focus on quantitative rainfall estimates (i.e. not just classification of events), the use of a simple and transparent algorithm and the focus on cross-site validation.
In my view the study itself is worth publishing, but the manuscript as a whole requires substantial technical improvements before it can be considered for publication.
Main comments
The authors make the point that one of the strengths of the here proposed method is that it is training free (Sect. 1, l. 53). At the same time, they underline the importance of site-specific considerations when selecting the optimal frequency windows (Sect. 4, l. 506). It would be useful for readers to understand more clearly under which circumstances the method presented here has advantages over alternative published methods.
The authors should make more clear why acoustic techniques could make an important contribution to reducing observational gaps in (world-wide) monitoring of surface precipitation. Is the technique more cost effective than conventional methods? Does it require less maintenance? Is there a potential for using existing networks?
In Sect. 3.5 a comparison with satellite observations (GPM) is presented, but only a small subset of the entire dataset of acoustic observations is used for this. I would advise either omitting this comparison or expanding the analysis substantially by including all acoustic sites and all satellite overpasses. In the latter case, the rain-gauge observations should also be included in the comparison, considering both event classification and quantitative rainfall estimates.
The method presented here appears to assume that differences between the dry acoustic baseline and observations during rainfall are primarily caused by the sound generated by the rainfall. This assumption may not always be valid. For instance, wet surfaces surrounding the measurement site may have different acoustic properties (e.g. absorption and reflection), while sound propagation from more distant sources may also be affecting during heavy rainfall. The potential implications for the proposed metric, including possible nonlinear effects should be discussed.
The sensors CH, LV and HV are in close proximity to one another. This provides a unique opportunity for inter-comparison of acoustic sensors through analysis of events, and statistical analysis of relative consistency and accuracy.
Technical comments
The method presented here deserves a name and/or acronym. The manuscript currently uses many variations, including "the proposed (acoustic) metric", "empirical metric", "metric", "rainfall metric", "the proposed acoustic baseline method", which is rather informal and potentially confusing. The manuscript also lacks a clear definitions of measured or derived quantities, their symbols and units. The manuscript could improve substantially through consistent application of these names, symbols and units in text, tables and figures.It is unclear how σ is defined and what quantity it represents. Please provide a clear mathematical definition and specify its name and units (if applicable). In particular, please distinguish clearly between the percentile level used to derive the threshold and the resulting threshold value.
The second sentence of the Introduction ("However, due... accurate prediction") should be rewritten as it attempts to connect a wide range of subjects all related to rainfall: spatio-temporal variability, measurement reliability (accuracy? representativeness?), prediction. Please reformulate.
Introduction l. 35. "Local approaches" please replace by: "in-situ techniques"
l.37: "weather radars are expensive to operate and maintain" I am not convinced that this is the most relevant limitation of weather radar in this context, particularly considering the large spatial coverage, high temporal resolution and range of variables provided by a single radar. More relevant limitations for precipitation monitoring may include the indirect nature of radar observations, the assumptions required for quantitative precipitation estimation (QPE) and the dependence of QPE accuracy on factors such as the distance from the radar.
Section 2.2.1 describes two different sensor types. Both sensors are described by different parameters (e.g. signal-to-noise ratio, sampling frequency) and can therefore not really be compared. Please provide a characterisation for both sensors that includes all relevant parameters.
Figure 1: both the figure and the caption should be improved. It is suggested to replace the overview maps (e.g. satellite view of Africa, Brazil, France) with plotted maps of country borders, rivers, coasts and oceans for the same spatial regions. This enhances also the contrast with the satellite imagery of the local surroundings. Also the connection between the maps and the photos of the direct sensor surroundings should be explained more clearly to the reader in the figure caption.
Sect. 2.2.2, l. 125. "Data collection... rainfall events." The sentence is unclear (use of "data" is confusing), please reformulate.
Table 1. Please consider to reformulate the caption, e.g. "Specification of measurement sites and sensor characteristics and statistical summary of datasets"
Table 1. For column "Total rainfall" it should be made clear (also in the caption) that this is derived from rain gauges.
Sect. 2.3 Eq. (2) contains a typo: the horizontal bar denoting the hourly average should be above the S in the denominator.Sect. 2.3 l. 160-162. "For the Amazonian....effect negligible." Please reformulate to more than one sentence.
Figure 2: Please remove typo from title of panels b) and c)
Figure 2: Please replace label under panel a to "Hour of day". Also the last hour of the day is missing in the plot.Please explain in more detail how the reader can identify from the plotted lines that Fig. 2. panel b) shows light rain and panel c) shows heavy rain.
Figure 3: please specify a) and b) for top and bottom plot respectively and consider to rename the label "metric" of the left vertical axis for panel b) by an appropriate name or symbol (see first technical comment above).
Figure 3: please consider to change the labels on the horizontal (time) axis (e.g. 18:00 etc.) and add a title to the axis (hour of day). Also the year should be mentioned in the figure or caption.Sect. 2.5 First, introduce symbols in the text before using them in equations. Second, Sect. 2.5 uses the word "mean" in equations to denote the mean, whereas previously a horizontal bar was used to denote a mean (e.g. Eq. (1)).
Sect 2.5, l. 211-212: "where the authors defined the error rate mean as a weighted mean between the false positives and false negatives" please reformulate, e.g. "where the authors defined the error rate mean as a weighted mean of the error rates for false positives and false negatives."
Sect. 3 and Sect. 3.1: These two sections have identical first paragraphs. Please correct.
Figure 4: Please change the vertical-axis label “data points” to “number of data points” or “frequency”, and define clearly in the text or caption what constitutes a “data point”. Please also consider changing the caption from “Rainfall data distribution...” to “Rainfall intensity distributions for each measurement site”.
Figure 5: Please include "Event" in the title for panel b): "Event total rainfall". Caption: consider to replace "dataset" by "measurement site" or "location".
Sect. 3.3, Fig. 8: The layout of the figure should be improved: some of the panels (e.g. upper left) have a missing label along the vertical axis, others seem to have inconsistencies: label for vertical axis for rows 2 and 5. Another consideration could be to use a common color scale and color bar for each row. This would also allow better visual comparison between the different panels on each row, which now cannot be compared visually because of different ranges for the same colours.
Sect 3.3., l. 340-341: Please explain in more detail to the reader how the figure shows that the most consistent results are obtained for a threshold between 0.97 and 0.98.
Sect. 3.4, l. 427: "manual audio-based validation" please rephrase (same for Sect 4., l. 500). An alternative could be "human-assisted classification".
Sect. 4. The final sentences of the Conclusions appear somewhat redundant.
Fig. A1. Please add a unit for time to the horizontal axis.
Citation: https://doi.org/10.5194/egusphere-2026-3457-RC2
Data sets
Replication Data for: Rainfall monitoring based on spectral analysis of sound recordings in different geographical locations de Souza Xavier et al. https://doi.org/10.23708/UMPMKI
Viewed
| HTML | XML | Total | BibTeX | EndNote | |
|---|---|---|---|---|---|
| 128 | 97 | 72 | 297 | 66 | 55 |
- HTML: 128
- PDF: 97
- XML: 72
- Total: 297
- BibTeX: 66
- EndNote: 55
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
A fundamental issue with this manuscript is the unexplained attempt to record rainfall under vegetation canopies, including tropical forest (and indeed, under what appears to be a metal roof). This is of course essentially impossible, and what was being recorded (except beneath the roof) would largely have been the sound of throughfall produced by the interaction of rainfall with the vegetation canopy. Processes involved included drip and splash, rather than direct rain droplet impact. In forest hydrology, rainfall would normally be recorded either in a clearing, or on a mast of sufficient height that it extended above the vegetation canopy. There should be no expectation that the open-field rainfall rate should match the sub-canopy throughfall rate, especially in rain of low to moderate intensity, such as the authors report. After all, interception loss consumes part of the open-field rainfall, and throughfall has an altered drop size distribution such that the sounds generated by the drops are not the same as the sounds that would be generated by rain droplets. The median drop size is normally considerably larger in throughfall than in open-field rainfall, and of course in appropriate weather conditions there can be fog drip even when no rain is falling.
It is surprising that the authors say nothing at all about these issues. They certainly warrant some attention. The title of the paper is of course also misleading, since the authors were clearly not monitoring “rainfall” but throughfall.
The authors neglect the issue of the area surrounding the sound recorders that was being sampled – which might be calle the “sound catchment area”. This is important because it bears on the degree to which the audio data can be taken to represent the highly spatially variable throughfall field typical of forests and other plant canopies. This area might have had a radius of 2 m or 20 m, or something else, depending on how well sound carried through the vegetation, and on the sensitivity of the microphones and recording systems. Can the authors provide an estimate? How many plant stems might have been included in the area from which the audio signal was gathered? The area being sampled by audio recordings is important because the rain gauge data are essentially point-based, whereas the audio signal will be from a much larger surrounding area. The two kinds of data are therefore not strictly comparable and it should not tacitly be assumed that they are.
A related question is where the rain gauges were placed – the authors do not bother to mention this important aspect. Were the gauges located under a canopy gap? Were they located under a branch or a dense area of foliage that might have dripped into the gauge? Again since the rain gauges were at most sites recording throughfall, not rainfall, this is important: throughfall is notoriously variable – sometimes less than the rainfall above (if the gauge is sheltered by a branch) and sometimes more (if the gauge was under a drip point). Were the rain gauges moved during different periods, to sample the throughfall field more systematically, or were they left at a fixed location? (Throughfall gauges are often roving, moved to different locations at regular intervals).
A major issue that arises because the authors recorded the sound of throughfall is that they were unable to detect the start of rainfall (since throughfall would not begin until the vegetation canopy was wetted) or the end of rainfall (since the wet canopy can continue to drip for a significant time after rain ceases). These are further reasons not to attempt to measure rainfall beneath a forest canopy and I remain puzzled as to why the authors would do this.
There are systematic problems with the recording of rainfall in 5-minute steps. Naturally, rain and throughfall continue to fall during the transition from one 5-minute period to the next. However, the number of raingauge tips is discrete. This is a well-known limitation of tipping-bucket rain gauges, and is one of the reasons for preferring acoustic recording of rainfall, that can be essentially continuous.
“The metric” is referred-to throughout the paper, but is never referred to by a name or a symbol. The axes of many graphs are also simply labelled “metric” with no clue as to what this is, nor are the units of measurement specified. This needs to be rectified, and the relevant equation number given when “the metric” is referred to. The relevant units in which “the metric” is expressed must also be clearly explained. In Figure 3 for example the vertical axis is simply labelled “metric”. No equation cited, no units given. This needs to be fixed throughout the manuscript.
Correlations of “the metric” with cumulative rainfall in events are poor – averaging about 0.7 (line 497), This means that the relationship accounts for less than half the variance of the data (i.e., r2 = 0.49) which is poor. Can the authors comment on this, and why they used simple correlation rather than r2?
More specific queries:
Why was sound recording interrupted such that only 1 minute was recorded every 5 minutes (lines 106-107) etc? This requires explanation. Such intermittent recording would inevitably fail to detect intensity peaks and short periods of intermittency.
Why was one audio device installed under a metal roof (as shown in Figure 1)? Why was this not explained in lines 84-88 where the site ABJ was introduced?
Is MIT 1 hour (line 127) or 3 hours (lines 274-275)? The authors need to be consistent.
Line 291: there is no blue colour, as claimed in the Figure 6 caption. It is black and grey.
Judging from Figure 1, the audio recorders were installed at different heights above the ground (different distances below the plant canopy). Is this correct, and if so, why was a constant height not adopted?
Line 398: what is the definition of “violent rainfall”?
The authors need to adopt a more consistent terminology for sound sources. They variously refer to “anthropophonic” and “geophonic” (line 508) as well as to “technophonic” (line 442) and to “bioacoustic activity” (line 168). I was left unsure of the difference in meaning of ‘anthropophonic’ and ‘technophonic’, for instance. Oddly pehaps, given this array of nomenclature, the authors do not adopt a term such as ‘hydrophonic’ to refer to water-related sounds.
What is ‘violent rainfall’ (line 398)? Most intensities mentioned in the manuscript are low. Indeed, what are the definitions of ‘light’, ‘moderate’, and ‘heavy’? The classification used needs to be explained.
Line 423 refers to “spatial separation between the recorder and the gauge”. What was this separation, for each field site? It should be documented. How do the authors know that this separation contributed to low correlations between event rainfall and the audio analysis?
In line 156 the authors note that they adopted 100 samples of dry sound levels for each hour of the day. They do not present any verification that this was a suitable number. Why not 50 samples or 500 samples? Some basis for this choice is needed.
Figure 2 has curves labelled “light rain” and “heavy rain”. How are these terms defined?
In lines 199-200 the authors claim that the sound data “accurately captures the beginning and end of the rainfall event”. How accurately could this be achieved, given that the authors were recording throughfall, not rainfall, with delays expected from canopy wetting -up time and post-rain dripping?
In line 241 the authors report that “the metric” was averaged for each event and correlated with the mean rainfall rate of the event. Why is the arithmetic mean appropriate here? Are sound levels and rainfall rates known to be normally distributed in the data? Can the authors present some evidence? Might they not be logarithmically distributed, suggesting that the geometric mean might be more appropriate? Can the authors defend this procedure as it was seemingly applied to entire events?
Lines 254-257 confirm that most rain was of low intensity – less than 4 mm/h. Such rain is not well recorded by tipping bucket gauges, and presumably the “sound catchment area” would be smaller for low intensity rain than for high intensity rain. The authors need to comment.