MuscaDet – A multiscale and uncertainty-aware detector of cloud-free periods in time series of 1 min measurements of solar irradiance at ground level
Abstract. Detecting cloud-free periods in time series of measurements of solar irradiance at ground level is essential in many domains, from meteorology to management of photovoltaic systems, including verification of so-called clear-sky models such as the McClear service of the Copernicus Atmosphere Monitoring Service. However, current cloud-free detectors call themselves on clear-sky models to discard measurements that are not close enough to the model outputs, introducing potential biases that pose a problem for the quarterly verification of the McClear service carried out in our institute on behalf of Copernicus. Accordingly, a new multiscale detector called MuscaDet was devised that is based on the correlations between measurements and a locally fitted Lambert-Beer-type function computed over a cascade of temporal windows of different lengths. This approach ensures that the detection remains independent of any predefined clear-sky model while still leveraging the dependence of the irradiance on the solar zenithal angle in cloud-free conditions. Measurement uncertainty and natural atmospheric variability are accommodated through a dynamic tolerance depending on the solar zenithal angle. MuscaDet is designed to operate on 1 min measurements of irradiance for the specific requirements of evaluation of clear-sky models. It needs measurements of the irradiance and its direct or diffuse component to discard overcast conditions. For its verification, time series spanning several years were collected at seven stations of the Baseline Surface Radiation Network covering various climates. Visual inspections of the measurements and outputs of MuscaDet show that the detector is accurate even in cases of short-lived isolated clouds or ramping cloudy events. Most cloud-free instants are detected and no visually-detected clouds are falsely noted as cloud-free by MuscaDet. A comparison with the currently used cloud-free detector indicates that MuscaDet provides a more balanced statistical distribution of solar zenithal angles, thus improving the verification of the McClear service. MuscaDet is easy to implement and needs no external input except the solar zenithal angle which is easily computed using existing libraries.
The article introduces a new, multi-scale detection framework named MuscaDet for the detection of cloud-free periods in 1-min resolution time series of solar surface irradiance. Novel aspects include the use of multiple time windows increasing in length from 15min to 2h, and to not rely on a clear-sky irradiance model as a priori information. In particular, the latter aspect avoids biases in detection efficiency arising from the choice of clear-sky model depending on the specific conditions at a chosen site.
Having applied detection algorithms to irradiance time series myself before, I strongly support these goals. In particular, the article is well-written, it contains a solid verification of the proposed method, and is well within the scope of AMT. There are, however, a number of aspects of the MuscaDet approach which should be discussed / potentially improved, to clarify the advances with respect to the current state of science, given that the field of clear-sky detection methods has been widely covered before and is rather mature, see e.g. Bright et al., (2020) and references therein.
I thus recommend publication of the article after minor revisions, in particular after addressing my general and specific comments listed below.
General comments:
1) The article mainly considers global horizontal irradiance for the classification. Given that it focuses on BSRN data, this choice should at least be better motivated, or better yet, the effect of using both direct and global/direct irradiance quantified (I do expect that adding direct irradiance would make the algorithm much more sensitive to small clouds in front of the sun / short shadow events). I believe the current approach could easily be applied to the direct/diffuse irradiance components.
2) Related to the previous comment, the distinction between „cloudless skies“ and „clear sun“ situations as introduced by Bright et al., 2020, should be mentioned and discussed. Which of these scenarios are the main focus of the present work?
3) Also applicable to many other prior studies, I have always been puzzled by the fact that clear-sky detection algorithms / their thresholds are mostly formulated in terms of absolute values of solar irradiance (instead of e.g. atmospheric transmittance). In particular, this leaves variability due to changes in sun-earth distance over the year unaccounted for. For closer sun / higher levels of extraterrestrial solar irradiance (e.g. in northern winter), variability thresholds will be stricter than for larger sun-earth distances (e.g. in Northern hemispheric summer). Technically, it seems easy enough to at least scale irradiances to mean earth-sun distance prior to classification.
4) Use of the correlation / Eq.4: at first sight, this indeed seems like a new and promising innovation, and I do agree with the statement that (L592): „The normalized nature of rho and 𝑅^2 makes them particularly attractive …“. Note however that a fundamental challenge for using correlation is mentioned and implicitely addressed in the current formulation based on Eq.4 / Appendix: the slow change of solar zenith angle around solar noon / for some polar observation. This implies that the MLB-function has little value for predicting changes in irradiance. For me, this raises an important question: is it really the linear correlation coefficient as a goodness-of-fit criterium which serves as main source of information for the classification? Focusing on this quote: „It can be observed that Eq. (4) reduces to the standard Pearson correlation coefficient when 𝑉𝑎𝑟(𝐸_𝑚𝑒𝑎𝑠) > > sigma_tol**2“. While this sentence is indeed true, it remains unclear whether this limit is actually relevant in the present application. Some thoughts on this:
* A correlation of rho=0.999 suggests an extremely good fit (RMSE of fit residual is 4.4%, or sqrt(1-rho**2)). How does the PDF of the right factor / square-root term in Eq.A-9b look like? To me, this suggests that Eq.4 no longer should be interpreted as correlation.
* Due to the division by zero in Eq.4 if Var(E_meas) <=rho_min**2, Eq.4 essentially becomes a threshold test for low correlations / around noon. For which fraction of clear-sky cases is this threshold criterium met, thus Eq.4 acting fas threshold test? My guess would be that this is the majority of cases.
* One alternative suggestion: instead of the correlation / Var(E_meas), wouldn’t it make sense to consider Var(E_meas-F_MLB), i.e. the residual of the MLB fit?
5) The article argues that the use of the correlation „allows the definition of a decision threshold that is independent of levels in irradiance (L593)“. However, in Eq.4, the threshold used is calculated from the cosine of the solar zenith angle. Doesn’t this invalidate the stated advantage?
6) The article emphasizes the point that it does not rely on a clear-sky model. It does, however, make use of the MLB function to approximate observations (instead of the power law used by Long and Ackermann, 2000). What is the difference between this choice, and reliance on a clear-sky model? How does the algorithm performance depend on the choice of this function / its functional form? I do think this is an aspect which deserves a bit of discussion.
Specific comments:
* L11: „for the quarterly verification“: Please explain what this actually means.
* L278: „The Python implementation of MuscaDet is openly available on GitHub“. Please do mention its license here (LGPL), openly is ambiguous.
* I suggest to add a DOI to the code / github repo, e.g. via Zenodo
* Finally, I encourage the authors to make available also the underlying data including the MuscaDet classification, to make it easier for subsequent studies to compare their results against the present work. At this stage, only the Python code and some example data are included in the github repository.