the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
High-latitude auroral and cloudiness occurrence from automatic image classification
Abstract. We have investigated auroral and cloudiness occurrence over Kjell Henriksen Observatory (KHO) in Svalbard using full-colour all-sky images from 2016-2025. Our approach focused on constructing a high-quality manually labelled training set. Images were classified as ClearAurora, ClearNoAurora, CloudyAurora, or CloudyNoAurora based on their content. As there is natural overlap between these classes, we carried out several iterative validation rounds to increase the number of high-quality sample images while removing images with unclear contents. We then evaluated different Convolutional Neural Network topologies and selected the best performing network to classify all images between January 2016 and December 2025 (over 8 million images in total). In addition to the validation accuracy with the ground truth, we also estimated the classification accuracy based on a random selection of classified images. Our final classifier, called KHOnet2026, results in accuracies from 94% to 98% depending on the score and
image class.
We investigated auroral occurrence over Kjell Henriksen Observatory in Svalbard in 2016-2025 with data based on automatic classification of full-colour all-sky images (8.2 million images in total). We used a simple-to-use classification algorithm with several rounds of manual labelling of randomly selected individual images in 4 classes: ClearAurora, ClearNoAurora, CloudyAurora, CloudyNoAurora. In each iteration, images which were not obviously belonging to any of the four classes were removed to minimise the confusion. We therefore acknowledge that our classes naturally overlap, and that the overlap determines the highest achievable accuracy of our method, in this case 96%.
We found that most of our image data is cloudy (60-70%). A validation of the cloud occurrence results was performed with an independent dataset from a co-located cloud sensor. We found a good agreement between the two datasets at a monthly average level with a correlation coefficient of 0.86. Auroral occurrence over Svalbard is of the order of 25% of the imaging time, and it shows no solar cycle correlation but is rather modulated by the cloudiness. The portion of clear skies without aurora is only about 10%. The statistically clearest month at KHO is January, and the cloudiest is November. This automatic classification routine is set to run in real-time and further expand the database of classified images to aid researchers in finding images with aurora. This knowledge allows for a far more efficient use of computer time in analysis of the structural evolution of the aurora, when cloudy data can be excluded. Furthermore, the automatically classified images provide a very useful proxy for all other optical instruments hosted by KHO.
- Preprint
(4133 KB) - Metadata XML
- BibTeX
- EndNote
Status: final response (author comments only)
-
RC1: 'Comment on egusphere-2026-2388', Masatoshi Yamauchi, 03 Jun 2026
-
AC1: 'Reply on RC1', Noora Partamies, 13 Jul 2026
We thank the reviewer for careful reading of the manuscript and much appreciate the detailed feedback. Please find out point-to-point answers below in bold.
Currently the readability of particularly section 3 is not sufficient.
Very true. This section is what we have struggled most with as the process itself has taken many years with many iterations making it challenging to document in a sensible way, but we will do our best to improve it for the revised version and appreciate the external opinions.
Location of KHO: please add latitudes (and longitude)At the beginning of Section 2, where KHO is introduced, we give the geographic coordinates and the geomagnetic latitude for the station.
Wording: "carefully construct" => Although the author who has long been worked on aurora image is one of the most reliable scientists (in terms of carefulness") to construct the "ground true" dataset, the word "careful" is unfortunately a subjective word. So, it is better to use "objective" numbers of iteration and numbers of year that the author has classified aurora. Similarly, best => best possibleGood point. These wordings will be changed to “construct through a number of iterations”, and “the best-scoring classifier”
Section 1
Solar cycle dependence: At Svalbard which is located far north of nightside aurora, no SSN dependence is reasonable. For nightside aurora, average size of the aurora oval (this depends on solar cycle) changes the probability of diffuse/discrete aurora at one fixed station. This note should be already mentioned in the introduction and discussion.Svalbard is located at the poleward boundary of the auroral oval, so as the size of the oval correlates with the solar activity, it would be reasonable to expect an anti-correlation of auroral occurrence with the solar activity, provided that other time sectors would not be strongly affected by the solar activity. However, this is not seen in our results. True enough, we are not separating diffuse or discrete aurora in this study, but particularly at the poleward boundary these two types separate less than at the equatorward boundary. There are little evidence for any of these things in the literature, which motivates our study. We will add a comment on the lack of solar cycle evolution of high-latitude aurora in the revised introduction.
Section 3.2
"aurora expert": Did only Partamies checked manually, or any other too. Since Partamies is knows as expert (readers can guess from reference list), I recommend to explicitly write as "The author, as the aurora expert"Yes, that is a fair point and will be changed accordingly.
In example in Fig 2, only last column shows the moon case, and in Fig A1-A4, one two moon cases are found. When I looked at many aurora images for judging, the moon appears as often cloud because moons at second and third quarters are high during winter (dark month) at high latitude. How often is moon included in the training set and ML labelled result? I do not require exact number (not asking authors to classify all samples in Table 2, but just rough filling)
The moon does appear in the training set of all classes. To estimate how often that happens we define twilight as the time when the Sun is less than 12 degrees below the horizon. Then roughly 15% of the ground truth images represent twilight conditions. The Moon is above the horizon in roughly 31% of images and a combination of twilight and Moon in 42% of images. Here we consider only the times when the Moon is more than 50% illuminated. In the revised version of the manuscript, we state the presence of the Moon and twilight in the training set in connection to Figure 5.
Line 151: "chose random images": from which years? Or is it from 2019-2020 unclear cases that was classified in the first manual labelling? Since this is not clear, I could not understand how to read Table 2. Is explanation in Section 4.1 (224-234) for Table 2?
For the two first rounds of manual labelling, the images were from 2019 and 2020. These correspond to the first two rows of Table 2. For all other iterations, random images were chosen from the entire dataset. We will add the years for all the rows in the revised Table 2 for clarity. The text on lines 224-234 describes the classification process and sets the scene for what we think is achievable. The labelled dataset mentioned there refers to the training data for the final KHOnet2026, which is an outcome of the iterations listed at the bottom row in Table 2. We will make this connection clear in the revised text.
Table 2 and 3 "title" (technical issue): Please change paragraph after the title (Unlike Figures, Tables do not have caption sentences). If you need caption, this should be moved to the footnote of table (with *1, *2, *3,,, marking)
According to the submission guidelines at https://www.annales-geophysicae.net/submission.html#manuscriptcomposition , tables need concise but descriptive captions. We will streamline the table captions for the revised version.
Table 2
Also, where is "parentheses"? Is images for shaded part are all originally from 5000 images in the third row? How did you define ground truth? (classifier's result of which samples, and manually checked one last time?)Apologies for the “parentheses”, that is a remnant from an earlier version and will be removed.
The images in the shaded cells of Table 2 are not from the previous set of 4x5000 random images. There is a new classification run between every iteration, meaning that the 5000 images per class where visually re-labelled, only the unambiguous images were added into the training set of the next classification round. Once the classifier was trained on the unambiguous dataset, the whole image dataset was classified with it, and another 2000 random images were chosen from each class for visual check and re-labelling. Therefore also the last visual check included 2000 random images from each of the four classes.
This part of the text will be reworked in the revised version.
Line 179 and Table 3: "2000 random images per class": If each category has 2000 samples, why Table 3 summation is not 2000? Could you explain more.
This is where we got ourselves tangled multiple times as well. The full training set for the final version of the classifier includes all the images counted in the shaded cells of Table 2. We have gathered all unambiguous examples for the four classes throughout the iterations to get a large enough of training dataset that represents the different sky conditions as well as possible. These are the ground truth numbers on the bottom row per class. From these images we used 70% for training and 30% for validation, and the numbers in Table 3 correspond to the validation part of the dataset.
This is mentioned in the connection to Table 3 (line 175), but we will try to rephrase it in a clearer way for the revised version.
Line 183-187: Unclear and incomplete sentences. Please fix. Errors are also found from section 4.
Incomplete sentences will be fixed and much of these text passages will be rephrased for clarity in the revised version.
It would also be better to switch the order of explanation (cloudy<->clear confusion is not surprising and better to explain first).If this comment relates to the text on lines 182-187, it roughly follows the order of the classes presented in the confusion matrix and mostly talks about cloudy versus clear conditions. We would therefore keep the order of the items, but will keep the comment in mind when re-working the text for the revision.
Figure 4: Although this is a good example of "reason for wrong classification", I was bit surprised that you run your classification even for twilight (noon) case in the figure. The classification scheme drastically improves if you remove the twilight (SZA = 90-102 degree or 90-96 degree) cases. Could you explain why you include twilight cases?
The aim of the classification is to be robustly used on all our data, not just certain times of the day or certain seasons but on everything that we have images for. We recognise that there are conditions where the classifier fails, but these are the cases where human experts cannot tell the ground truth either, and that is what we need to accept as a classification error. This is admittedly a long-term study focussed view, but it seems artificial to us to start excluding data, which do not behave well.
Figure 4 actually shows that the classifier works surprisingly well in twilight conditions too, as the pre-noon part of the day is more cloudy and the post-noon part clear. Probabilities for the cloudy and clear classes ramp down and up respectively. If one wanted to do a temporal smoothing of the class probabilities as a post-processing step, it would give a very reasonable timeline. Here we merely just show the “raw data” as an example.
Line 194 "exposure time": The aurora judging scheme depends on exposure time. Isn't the exposure time just before the abrupt appearance of aurora a good value to stop operation?Often after the change of the exposure time, the image content becomes much easier to see as the contrast between aurora and the background increases. This is why the exposure times were changed in the first place. In reality, our ability (as humans or as ML classifiers) to detect aurora depends very much on the cloudiness level and the type of clouds, colour of the background sky and the level of auroral activity. There is therefore no one simple and obvious threshold or indicator on when to stop imaging with certain set of imaging parameters. The current imaging software does not contain exposure time changes at all. As explained in the previous answer, we want the entire dataset classified no matter what the imaging mode was, rather than implementing (and testing and validating) additional ways to pre-prune the data.
In the aurora quantification method (more explanation below), exposure time is used as one of the parameter to judge twilight time and moon (to adjust detected intensity). This case clearly show that "aurora could have existed but not detectable". From users viewpoint, such "twilight" should not marked as "no aurora”.We think it is fair to gather all detectable aurora into the class of aurora, while the rest is no aurora. The images that we as human experts or ML classifiers cannot classify correctly are the classification error we cannot improve. While the pixel values change drastically during twilight conditions, a CNN classifier will detect aurora when the contrast to the background sky is strong enough for the structures to be seen. This is comparable to the “visual aurora” concept introduced in Yamauchi et al. 2023.
Figure 4: The image is strongly affected by the moon after around 20 UT, and I suspect the some percentage of "ClearAurora" (particularly after 2330) could be no aurora. Could you show original images from around 2240 UT and 2340 UT?
There is actually some aurora at low elevation from about 22:40 to about 23:50 UT, as shown by the attached sample images. The class probability for ClearAurora only reaches 0.4 in that time frame, so this would not be counted as aurora in our statistics, where we assign the images to the class with the highest probability.
It would be useful to mention the other types of automated aurora classification in addition to machine leaning type. Oldest example is photometer count (filtered 5577) to classify intensity, and recently, pixel-level classification to quantify the aurora activity was introduced (Yamauchi+ https://doi.org/10.5194/gi-12-71-2023). While the ML type provides good classification of morphology, manual type provides quantified values, and combining both methods will improve in evaluating the aurora activity in future. This work is a good reference for such work.
This work will be included in the revised version of the introduction as an example of other types of automatic detection tools.
Line 210: 0.1 second per image is a good performance fitting to real-time operation of burst-mode (high time resolution) observation. Please mention such advantage.
We will mention the realtime capability in the revised version at the end of section 3. In fact, the classifier was internally test-run in realtime during the winter season 2025-2026.
Section 4.1
Line 235-247 "Sony and Nikon": The RGB sensitivity and its balance are quite different between Sony and Nikon, and is also a big problem for automated aurora quantification method too (above reference). Could you make similar table (appendix is ok) as Table 2 and Figure 3 just for this purpose (1000 samples)? After reading your result, I am quite confident that statistic (like Figure 6) should be separated obtained for Nikon camera and Sony camera. In this respect, I could not catch what authors want to say this paragraph: Excuse why the statistics is limited to 2016-2025 (I agree), or it is feasible to correct the judging of Nikon camera result in terms of Sony camera (I personally do not agree)?The sensitivity and the colour balance are indeed very different between Sony and Nikon, or any two different camera models. This was a test run for classifying Nikon data with the method we used for Sony. The results are not very bad, but the method clearly requires to be trained with some Nikon data in order to perform at a level that would be more comparable to the level of Sony classification. As the aim is just to prune all our colour image data for more detailed analysis and at the same time provide some statistical insights. Extending the AO/CO timeline back to 2008 will have its value for facilitating further auroral studies on the older data. The training of the classifier will allow it to learn the colour balance of the old data, while the sensitivity issue will remain and must be kept in mind, as we can only teach a classifier to detect aurora where we can see and detect the aurora ourselves.
We will include the image numbers for the Nikon camera test in the same format as we have for Sony images in the revised version of the manuscript. The human accepted image numbers from the 1000 random samples per class are: 785 for ClearAurora, 450 for ClearNoAurora, 632 for CloudyAurora, and 884 for CloudyNoAurora.
Section 4.2
In the statistics (Fig 6), was twilight case (cf Fig 4 center part) included or excluded?Twilight is included. All the images the camera took in 2016-2025 are included.
Figure 6: The "low" probability of AO at 12 MLT (which is close to local noon) and remaining double peak might be due to twilight that prevents recognising the aurora.
This was tested by excluding the twilight (solar zenith angles larger than -12) and moonlit (Moon illumination larger than 50%) conditions. The biggest change is that the data numbers drop to half and one daytime bin goes to zero. AO increases by 6-7%, while cloudiness and clear skies decrease by 7% and 3%. The double peak in AO still remains, as shown by the figure in the supplement file.
Figure 6: It is useful figure but please add a comment that this is not for "solid AO probability", because part of the cloudy cases (thick cloud through which the aurora is not possible to see) should be removed from total number for such probability. Even ClearAurora/(ClearAurora+ClearNoAurora) will not give the probability because ClearAurora allows 3/8 of the sky to be covered by cloud)
Yes, the revised version will emphasise that these occurrences include the cloudiness, which we cannot remove. This is rather the detectable ground-based AO.
Figure 7 and 8: It is easier to place 11-12-1-2 from top to bottom.
This would work well for a plot with one winter season, as in Figures 7 and 8, but we wanted to keep the format extendable to multiple winter seasons in Figure 9.
Figure 8: I worry that high cloud rate under daylight/twilight means that cloud detection depends on the sky brightness, so that you should remove the twilight case from the GroundTrue. Could you comment? Please also compare ClearAurora/ClearNoAuora ratio and CloudyAurora/CloudyNoAurora
We will not remove twilight from the ground truth, because the images where it is possible to distinguish between clear and cloudy with and without aurora will help the classifier to learn the reality. We acknowledge that some bright images are too bright for this distinction, but that is by no means all twilight and moonlit images. These are the images where also the manual labelling becomes impossible, and the images were classified as Unclear.
It is not clear to us what the ratios of ClearAurora/ClearNoAurora and CloudyAurora/CloudyNoAurora would tell more than their occurrence variability.
Line 311 "sound agreement": The slope < 1 (ML method overestimate the cloud) is better explicitly written rather than just saying a "discrepancy".
True. The revised version will say that the agreement is good. Because the cloudiness estimates are so very different, we should not compare the absolute numbers between the axes. Cloud sensor value 60 does not mean the same as 60% CO. This is why we think that a correlation of 0.86 is good.
Line 330: For the next step, how about combining intensity of aurora (quantification of aurora using pixel size and integrated intensity of auroral pixels)? A general scheme is already presented in Yamauchi+ 2023.
This is also an option we can mention in the revised text. A possible combination of these methods could be to first prune out the clouds with our ML routine and then use the pixel-level method to detect the local brightenings.
Line 348 "Computer science point of view": What do you exactly mean? CPU time?
This means “from the methodological point of view”, and will be changed to say so, meaning that currently there are more options that just CNN. Vision Transformers seem particularly promising and should be tested.
Lin 358 "allow the Moon and twilight in": As described above, they causes change of exposure, and drastically decrease the chance of recognising aurora, particularly the diffuse one. If they are included, they should also be classified as "MoonNoAurora" "MoonAurora". I am sure that ratio of MoonAurora/MoonNoAurora is smaller than ClearAurora/ClearNoAurora.
The Moon in the images is not causing a change in the exposure time in our data. The twilight may do that depending on the imaging mode and the threshold value used for the solar zenith angle. The sky illumination conditions determine how dim aurora can be detected by human eye and by a classifier, and these two approaches can do about equally well. An ML approach is not sensitive to changes in the pixel-level colour balance, but rather the contrast between the aurora and the background. When a human cannot detect aurora, we call the conditions NoAurora acknowledging that there may still be faint diffuse aurora. In a long-term statistical approach, we still get good results, and the detection of the aurora is still helpful for further studies.
We have included the zenith angles for the Sun and the Moon, as well as the Moon illumination in our classification results, so that anyone using the classification results can decide on thresholding of those parameters as it best suits the purpose of their study.
Line 375 "This suggest...": Not necessary true according to Table 2. This illusion comes from different mother group "clearNoAurora" vs "cloudyNoAurora.
What we want to convey is that the largest class of images is the CloudyNoAurora class. Even with the twilight and moonlit conditions excluded, there is still more cloudy data than anything else. Therefore the biggest limitation for observing aurora is the cloudiness. While AO includes both ClearAurora and CloudyAurora, CO is CloudyNoAurora only, and clear skies is ClearNoAurora only. All these super-classes or mother-groups are then normalised the same way so that they are directly comparable. This is why we think that the statement is correct.
Line 389 "more than 20%": The difference also comes from twilight data (see comment on Figure 6). Also the purpose of the aurora classification is different from Nanjo+ 2022. This paper aims to search aurora with blind information for the other observation teams, whereas the previous works are rather for aurora scientists (including citizen scientists) to search more classified aurora.
The purpose of the classifier will define much of the structure of the implementation, which is something that we will recognise in the revised text, along with the inclusion/exclusion of the twilight data. In our data, the largest class is cloudy skies, with or without the twilight and moonlit conditions. We therefore want to prune out the cloudiness and detect the aurora and clear skies as the first analysis step to make more detailed classifications and studies more efficient. These differences do not allow direct comparison with our results and those from Nanjo et al. although their method is an impressive piece of work.
Line 404 "only one other": Actually the quantification of aurora activity from ASC (Yamauchi+ 2023) is already in operation, including ESA spaceweather site. Also, Nanjo's TromsoAI also includes "cloudy" probability
This statement was referring to ML methods, but can be extended to include other approaches and thus include the local brightening detection based on pixel values. A little downside with it is that ESA space weather site requires a log in. The IRF site we were not aware of, but it seems like the classification can be followed through that interface as well. We will also emphasise the different motivations guiding the different operational approaches. Our motivation is to detect all aurora, so that this cleaner set of images containing aurora can be more efficiently analysed in more detail.
-
AC1: 'Reply on RC1', Noora Partamies, 13 Jul 2026
-
RC2: 'Comment on egusphere-2026-2388', Anonymous Referee #2, 02 Jul 2026
This paper describes an algorithm for automatic identification of aurora and cloudiness in full-color all-sky images dating 2016-2025 from the Kjell Henriksen Observatory. The algorithm is based on the use of a convolutional neural network for obtaining the image classifications. The paper is very well written and the results are convincing. Convolutional neural networks have been used in this area in the past, and the significance of this work lies in the extensive, manually annotated and rigorously validated dataset of auroral images that the authors constructed in order to train the model. The validation process includes comparison with cloud occurrence data obtained independently from a co-located cloud sensor and shows high correlation between the two datasets. As one of the largest auroral image datasets to have been manually annotated with ground-truth labels, this dataset is likely to be of great value to researchers in the field.
One area needs correction: the method of Johnson et al. described in lines 52 - 64 is not correctly explained: the teacher-student method described here and also in the Xie et al. 2020 reference is a fundamentally different approach than the method used by Johnson and described in the Johnson et al. 2024 paper and in the Rani et al. 2023 self-supervision overview paper. Johnson's method uses self-supervision with a contrastive learning objective to first learn latent representations of all-sky images, then finetunes the resulting network for classification using a small set of manually annotated images.
Citation: https://doi.org/10.5194/egusphere-2026-2388-RC2 -
AC2: 'Reply on RC2', Noora Partamies, 13 Jul 2026
We thank the reviewer for picking up an important item in our manuscript. We will correct and clarify the method by Johnson et al. in the revised version.
Citation: https://doi.org/10.5194/egusphere-2026-2388-AC2
-
AC2: 'Reply on RC2', Noora Partamies, 13 Jul 2026
Viewed
| HTML | XML | Total | BibTeX | EndNote | |
|---|---|---|---|---|---|
| 68 | 25 | 12 | 105 | 5 | 7 |
- HTML: 68
- PDF: 25
- XML: 12
- Total: 105
- BibTeX: 5
- EndNote: 7
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
Review of "High-latitude auroral and cloudiness occurrence from automatic image classification" (egusphere-2026-2388) by Partamies and Syrjäsuo
The paper reports construction of dataset of sky condition at KHO in Svalbard. Although the method itself is not new, the content is important and should be published in a scientific journal like AnnGeo, because (1) the paper correctly pointed out that key for the supervised machine leaning (ML) is how to select correct training images of highly variable aurora/cloud activity, and showed how to make it (explanation of Fig 3 is excellent), and (2) Svalbard is a special location for auroral observation (this os different from other location) and the resultant database will be useful for all related study on aurora such as solar cycle dependence. To achieve (2), the authors note the importance of constructing "ground truth" dataset (by "interactive manual labelling"), which is needed these days after many supervised machine leaning (ML) methods are presented.
I have several point (major and minor) to be addressed before this paper is published. Currently the readability of particularly section 3 is not sufficient.
-----
Location of KHO: please add latitudes (and longitude)
Wording: "carefully construct" => Although the author who has long been worked on aurora image is one of the most reliable scientists (in terms of carefulness") to construct the "ground true" dataset, the word "careful" is unfortunately a subjective word. So, it is better to use "objective" numbers of iteration and numbers of year that the author has classified aurora. Similarly, best => best possible
Section 1
Solar cycle dependence: At Svalbard which is located far north of nightside aurora, no SSN dependence is reasonable. For nightside aurora, average size of the aurora oval (this depends on solar cycle) changes the probability of diffuse/discrete aurora at one fixed station. This note should be already mentioned in the introduction and discussion.
Section 3.2
"aurora expert": Did only Partamies checked manually, or any other too. Since Partamies is knows as expert (readers can guess from reference list), I recommend to explicitly write as "The author, as the aurora expert"
In example in Fig 2, only last column shows the moon case, and in Fig A1-A4, one two moon cases are found. When I looked at many aurora images for judging, the moon appears as often cloud because moons at second and third quarters are high during winter (dark month) at high latitude. How often is moon included in the training set and ML labelled result? I do not require exact number (not asking authors to classify all samples in Table 2, but just rough filling)
Line 151: "chose random images": from which years? Or is it from 2019-2020 unclear cases that was classified in the first manual labelling? Since this is not clear, I could not understand how to read Table 2. Is explanation in Section 4.1 (224-234) for Table 2?
Table 2 and 3 "title" (technical issue): Please change paragraph after the title (Unlike Figures, Tables do not have caption sentences). If you need caption, this should be moved to the footnote of table (with *1, *2, *3,,, marking)
Table 2
Also, where is "parentheses"? Is images for shaded part are all originally from 5000 images in the third row? How did you define ground truth? (classifier's result of which samples, and manually checked one last time?)
Line 179 and Table 3: "2000 random images per class": If each category has 2000 samples, why Table 3 summation is not 2000? Could you explain more.
Line 183-187: Unclear and incomplete sentences. Please fix. Errors are also found from section 4.
It would also be better to switch the order of explanation (cloudy<->clear confusion is not surprising and better to explain first).
Figure 4: Although this is a good example of "reason for wrong classification", I was bit surprised that you run your classification even for twilight (noon) case in the figure. The classification scheme drastically improves if you remove the twilight (SZA = 90-102 degree or 90-96 degree) cases. Could you explain why you include twilight cases?
Line 194 "exposure time": The aurora judging scheme depends on exposure time. Isn't the exposure time just before the abrupt appearance of aurora a good value to stop operation?
In the aurora quantification method (more explanation below), exposure time is used as one of the parameter to judge twilight time and moon (to adjust detected intensity). This case clearly show that "aurora could have existed but not detectable". From users viewpoint, such "twilight" should not marked as "no aurora"
Figure 4: The image is strongly affected by the moon after around 20 UT, and I suspect the some percentage of "ClearAurora" (particularly after 2330) could be no aurora. Could you show original images from around 2240 UT and 2340 UT?
It would be useful to mention the other types of automated aurora classification in addition to machine leaning type. Oldest example is photometer count (filtered 5577) to classify intensity, and recently, pixel-level classification to quantify the aurora activity was introduced (Yamauchi+ https://doi.org/10.5194/gi-12-71-2023). While the ML type provides good classification of morphology, manual type provides quantified values, and combining both methods will improve in evaluating the aurora activity in future. This work is a good reference for such work.
Line 210: 0.1 second per image is a good performance fitting to real-time operation of burst-mode (high time resolution) observation. Please mention such advantage.
Section 4.1
Line 235-247 "Sony and Nikon": The RGB sensitivity and its balance are quite different between Sony and Nikon, and is also a big problem for automated aurora quantification method too (above reference). Could you make similar table (appendix is ok) as Table 2 and Figure 3 just for this purpose (1000 samples)? After reading your result, I am quite confident that statistic (like Figure 6) should be separated obtained for Nikon camera and Sony camera. In this respect, I could not catch what authors want to say this paragraph: Excuse why the statistics is limited to 2016-2025 (I agree), or it is feasible to correct the judging of Nikon camera result in terms of Sony camera (I personally do not agree)?
Section 4.2
In the statistics (Fig 6), was twilight case (cf Fig 4 center part) included or excluded?
Figure 6: The "low" probability of AO at 12 MLT (which is close to local noon) and remaining double peak might be due to twilight that prevents recognising the aurora.
Figure 6: It is useful figure but please add a comment that this is not for "solid AO probability", because part of the cloudy cases (thick cloud through which the aurora is not possible to see) should be removed from total number for such probability. Even ClearAurora/(ClearAurora+ClearNoAurora) will not give the probability because ClearAurora allows 3/8 of the sky to be covered by cloud)
Figure 7 and 8: It is easier to place 11-12-1-2 from top to bottom.
Figure 8: I worry that high cloud rate under daylight/twilight means that cloud detection depends on the sky brightness, so that you should remove the twilight case from the GroundTrue. Could you comment? Please also compare ClearAurora/ClearNoAuora ratio and CloudyAurora/CloudyNoAurora
Line 311 "sound agreement": The slope < 1 (ML method overestimate the cloud) is better explicitly written rather than just saying a "discrepancy".
Line 330: For the next step, how about combining intensity of aurora (quantification of aurora using pixel size and integrated intensity of auroral pixels)? A general scheme is already presented in Yamauchi+ 2023.
Line 348 "Computer science point of view": What do you exactly mean? CPU time?
Lin 358 "allow the Moon and twilight in": As described above, they causes change of exposure, and drastically decrease the chance of recognising aurora, particularly the diffuse one. If they are included, they should also be classified as "MoonNoAurora" "MoonAurora". I am sure that ratio of MoonAurora/MoonNoAurora is smaller than ClearAurora/ClearNoAurora.
Line 375 "This suggest...": Not necessary true according to Table 2. This illusion comes from different mother group "clearNoAurora" vs "cloudyNoAurora.
Line 389 "more than 20%": The difference also comes from twilight data (see comment on Figure 6). Also the purpose of the aurora classification is different from Nanjo+ 2022. This paper aims to search aurora with blind information for the other observation teams, whereas the previous works are rather for aurora scientists (including citizen scientists) to search more classified aurora.
Line 404 "only one other": Actually the quantification of aurora activity from ASC (Yamauchi+ 2023) is already in operation, including ESA spaceweather site. Also, Nanjo's TromsoAI also includes "cloudy" probability