the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
Convective Environments Over the Arabian Peninsula in Current and Future Climates: A Machine Learning Approach
Abstract. Climate change is intensifying extreme rainfall and flash floods across the arid Arabian Peninsula (AP). Adaptation requires a large ensemble of high-resolution precipitation projections obtained from dynamical downscaling, which remains computationally prohibitive. Here, we present the first component of a statistical-dynamical downscaling framework designed to reduce this computational burden. Machine learning models were trained to identify convective environments (CEs) using predictors derived from ERA5 reanalysis and a binary predictand from IMERG precipitation data. The best-performing model was applied to the CMIP6 ensemble to assess CE occurrence and components under current and future climates. At the current regional warming level, CEs are less frequent in the CMIP6 ensemble relative to ERA5, accompanied by drier and more stable environments, suppressed updraft range, and stronger wind shear. At an additional +1 °C regional warming, CE occurrence generally increases, excluding spring, accompanied by a diurnal shift towards nocturnal and morning periods. A general decrease is projected at an additional +3 °C, excluding winter. Nevertheless, CE-related moisture and instability exhibit monotonic increases with warming, suggesting more extremes and a shift toward a higher contribution of extremes to total rainfall. The resulting CE probability dataset enables targeted event selection for convection-permitting dynamical downscaling across the AP.
Competing interests: At least one of the (co-)authors is a member of the editorial board of Natural Hazards and Earth System Sciences.
Publisher's note: Copernicus Publications remains neutral with regard to jurisdictional claims made in the text, published maps, institutional affiliations, or any other geographical representation in this paper. While Copernicus Publications makes every effort to include appropriate place names, the final responsibility lies with the authors. Views expressed in the text are those of the authors and do not necessarily reflect the views of the publisher.- Preprint
(8805 KB) - Metadata XML
-
Supplement
(10693 KB) - BibTeX
- EndNote
Status: open (until 14 Sep 2026)
- RC1: 'Comment on egusphere-2026-3288', Anonymous Referee #1, 20 Aug 2026 reply
-
RC2: 'Comment on egusphere-2026-3288', Anonymous Referee #2, 03 Sep 2026
reply
Overall assessment:
The manuscript entitled “Convective Environments Over the Arabian Peninsula in Current and Future Climates: A Machine Learning Approach” presents itself as the first part of a study aiming at inferring precipitation extremes over the Arabian Peninsula in a future warmer climate through a machine learning and statistical downscaling framework. In this manuscript, the authors presents their methods to detect “convective environments” using three different Machine Learning algorithms trained on more than a hundred of usual convective parameters including CAPE, vertical wind shear, or vertical velocity from ERA5 reanalysis which is validated against the satellite precipitation product IMERG. They then apply their algorithm to CMIP6 model ensembles to deduce spatio-temporal properties of convective environments in CMIP6 historical and future climate projections. They discuss biases between CMIP6 and ERA5 convective environment and projected changes in CMIP6 convective environments compared to current climate.
This is a well thought and promising project. The manuscript appears to contain a substantial amount of relevant analyses. However, I find the overall presentation difficult to follow. This may be partly a matter of individual reading experience, but I think it is sufficiently pronounced to warrant substantial restructuring. Otherwise the manuscript presents a substantial amount of unclear formulations and arguments.
Please find below my major and specific comments.
Major Comment:
I found that the large number of analyses and the way they are sequentially presented makes it difficult to identify the main scientific storyline. In several places, the manuscript moves from one analysis or figure to another without making sufficiently clear the interpretation of the results, what scientific question motivates the transition, and how the new analysis relates to the preceding result. This creates the impression of a succession of analyses rather than a progressively developing argument. I think the manuscript would benefit considerably from a clearer hierarchy between the main results and supporting analyses.
I also found the separation between the Results and Discussion sections problematic in this respect. The Results section contains a large number of observations, but the motivation and broader significance of these observations only become clearer later in the Discussion. The reader therefore has to retain and connect a large amount of information before understanding how the individual results contribute to the overall conclusions. I would therefore strongly encourage the authors to reconsider the structure of the manuscript.
In my view, the manuscript would be much more effective if it would include some interpretation/discussion of the results inside the main Results section.
Specific Comments:
L. 1-3: This sentence feels a bit too absolute, as if dynamical downscaling were the only solution for adaptation.
L. 9-10: “ excluding winter” , “excluding spring” → what happens for these seasons? Do CE occurrence increases/decreases or remains the same?
L. 10-12: I find this assertion (“shift towards a higher contribution of extremes to total rainfall”) a bit too speculative, when just an increase in moisture and convective instability are provided as argument.
L. 16-17: Not clear what is “combined with regional aridity, high temperatures, and limited freshwater resources”, please reformulate.
L. 36-41: The difference between explicit and implicit approaches is not so clear, when convective indices could also be used in CPMs to derive convective environments. Please precise.
L. 42-44: Here, the reader would expect the authors to specify which approach is adopted in this manuscript.
L. 80: The term “agglomeration” is somewhat unusual in this context and may be unclear to readers + the second part of this sentence is unclear with “resolved” referring to “questions” instead of “precipitation”.
L. 106-107: Not clear to me, are pressure levels less numerous?
Section 2.1: It would be good to reconnect each data set to its purpose in this manuscript.
L. 115-117: Please provide the number of hybrid levels (optionally compared to the number of pressure levels) to make this sentence clearer.
L. 121-124: With the current phrasing, it is not so clear whether you performed the calculation to retrieve the average regional warming or if you used a value from the literature.
L. 124-125: Not clear in this sentence that you defined the periods so that they reach a predefined averaged warming value. In addition, the reason why this approach was adopted is missing here.
L. 125-126: When truncating at 2014, are the years post 2014 dropped from the calculation of the regional warming or not, implying that the final truncated period would not exactly reach the set warming of 1.22oC compared to preindustrial levels?
L. 133: Table S2 is mentioned before table S1. Please correct here and in other instances where the order is not correct.
L. 142-146: It is not clear why the data was averaged over a 1o x 1o grid (when convection is usually more local), why the local median (both local and median) was chosen for detecting convective precipitation (why e.g. not another percentile defined over the whole AP). The argument about virga is also not clear as it is not precised that IMERG could potentially identify virga as precipitation reaching the surface, and probably refers more to the 1 mm/6h threshold than to the median.
L 149: Are the instantaneous convective indices computed at e.g. 0600 to predict convective rainfall occurrence centered at 0600 (that is, for IMERG, precipitation falling from 0300 to 0900) or there is a time offset between the predictor and the predictand? It should be considered that convective rainfall may feedback negatively on many convective predictors (e.g. CAPE).
L 149-151 : Here as well I am not totally convinced with the method using relative thresholds instead of absolute thresholds and of the justification to enhance the signal of convection.
Take for example the lifted index (LI), if the monthly-diurnal average is e.g. +10oC (so overally stable), but one day it is +5oC then it is a strong negative anomaly of -5oC but it is still stable. Then this day will be given the same LI predictor than a day with -10oC in a very unstable month/hour of average -5oC. Yet, in the former configuration, I see hardly any possibilities for deep convection compared to the later configuration.
Besides, if I am not mistaking, the convective precipitation threshold (defined as the median) is defined from the whole climatology, and is not made relative to the hour of the day or the month in the year. Therefore predictor and predictand do not really match in my opinion.
L. 153: Please specify the interpolation method.
L 157: The text does not say the used reference data set for the MBCn (I assume ERA5).
L. 182-83: The threshold used to convert predicted probabilities to a binary classifier could be precised here (I see it is only defined later in the text) for more clarity.
L. 188-190: The letters “h”, “q”, “f”, and “m” need to be defined.
L. 198: As mentioned, this could have been precised earlier in the text. Please also precise the data set used to calculate the HSS (whole climatology, training, test, or validation).
L. 208: It is not clear which imbalance do you refer to, and the reason why it is mentioned here. Please specify.
L. 210: Please clarify if the frequency of CEs in each parts (a, b, c) is the same as the whole dataset (i.e. equal to 0.66%).
L. 212-214: Please cite relevant literature for this approach.
L. 214: I wonder why the complete train dataset was still used for DL despite its large size and imbalance.
L. 216: As mentioned above, this could have been precised earlier in the text. Looking at Fig. S1, since the performance of the ML models were not tested for percentiles below the 50th percentile, I am not convinced that the optimal percentile is 50. Could it be expanded to lower percentiles or could it be justified why the lower percentiles were not considered?
L. 218-219: The same remark holds here: why sampling ratios < 15% were not included in the plot to illustrate the optimality at 15%?
Eq. 2, 3: Both the adjusted probability estimates and the calibration appear to use the CE frequency of the climatological dataset. Since the data are split using stratified sampling that preserves the climatological CE distribution, the prevalence of CEs in the test set is therefore known and consistent with the frequency used in the probability adjustment/calibration. This may result in performance that is more favorable than would be expected when the model is applied to an independent dataset with a different or unknown CE frequency, such as a future-climate dataset. Please discuss this issue and clarify how the proposed probability adjustment and calibration are expected to perform when the underlying CE frequency changes.
L. 237: Please add the missing parentheses.
L. 241-242: Please specify how the probabilities from CMIP6 were adjusted and calibrated.
L. 262: Please include the reference to Fig. S5 at the end of this sentence (instead of the next one) as this sentence appears to already be commenting on this figure.
Fig. S4: Please clarify why the AUCPR baseline is missing on this plot.
Fig. S5: Please consider adding the frequency map of CEs for comparison/discussion.
L 267-268: If I understand well, the “climatological” data mentioned in L. 229 is only the training and the validation data set? It is not so clear in my opinion until this passage (L. 267).
L 271 and Fig. 3: The term “Multivariate Model” is not sufficiently clear, as it does not indicate that the model is based on CMIP6 data. Given that MBCn is applied to the CMIP6 convection indices, I suggest explicitly including “CMIP6” in the label to make the distinction clearer throughout the results. L. 273: Please define FutI and FutIII at their first occurrence in the manuscript.
L 273-274: I do not see an obvious increase in occurrence for probabilities below ~10-2. Please reformulate this sentence.
Please note that the projected changes in FutI are not consistent across all ensemble members, with some members showing increases and others showing decreases.
L 280-281: Please clarify what “it” refers to, as it appears not to refer to the diurnal cycle. Also, is the underestimation exactly 0.001, or is this an approximate value? If it is approximate, please indicate this accordingly.
L 283 : The term “earlier” (or “afterwards” in L. 275) is unclear, particularly in the context of diurnal-cycle plots, where it could be interpreted as referring to an earlier time of day. Please specify what is meant by “earlier” (e.g., at a lower value of the relevant variable or threshold).
L. 302-304: Could the authors briefly explain why the fact that the ERA5 test data represent only 20% of the total period (2001–2024) is expected to result in an underestimation of CE frequency?
Sec 3.3: The section on biases in the CMIP6 convective indices (Section 3.3) feels somewhat disconnected from the subsequent results. I suggest considering an exchange of Sections 3.2 and 3.3, so that the discussion of projected changes in CE frequency (new Section 3.3) is followed directly by the projected changes in convective indices conditioned on CE (Section 3.4). This ordering would provide a more natural flow between the occurrence of CEs and their associated convective environments.
L. 327-328: Please clarify how the CMIP6 models with grid spacing coarser than 1o were regridded to the 1o grid.
L331 332: “Dry bias” is unclear here. Please specify what the PW and 2-m dew-point temperature biases are relative to (e.g., ERA5).
L. 350: Fig. 8a appears to show the mean bias in vertical profile not the profiles themselves. Please correct.
L 381-382: The reported increase in w500 is not readily apparent from Fig 10c.
L. 388-389: If the diurnal-cycle plots are included, it would be useful to discuss the relevant features of the diurnal cycle in the text. Otherwise, the purpose of presenting these plots is unclear.
L. 417: “the the” → Please correct.
L 417 – 420: Could the authors briefly explain why would the ML model performs the best when “these convective situations are influenced by large-scale environments”?
L 432-435: I do not fully understand the relevance of these references in this context. The authors identify CEs based on environmental variables rather than convective precipitation from the models. Therefore, the finding that CMIP6 models require more moisture to transition from drizzle to intense convection does not necessarily imply that these models underestimate moisture in convective environments. Please clarify how the cited studies support the interpretation made here.
L 435-438: A similar comment applies here. The fact that CMIP6 models struggle to represent elevated convection does not necessarily imply that this is due to biases in the environmental variables. These are two distinct issues, and evidence would be needed to establish that the misrepresentation of elevated convection results specifically from biases in the environmental conditions. Please clarify the basis for this interpretation.
L. 457: The statement “The consistent increase in trends during DJF” is unclear. Please specify what is increasing and whether this increase is consistent across the models. Moreover, since this study does not analyse temporal trends, the use of the term “trends” seems inappropriate here. Please clarify or revise this statement.
L. 461: “Both futures” is somewhat vague. Please reformulate this statement to explicitly specify the two future scenarios being referred to. L. 465: Hypothesis d. Please see my previous comment (L. 435-438)
L. 475-476: Please check the stated relationship between low-level moisture and LCL. At constant RH (as stated in the next sentence), an increase in specific humidity requires an increase in temperature, and it is not clear that this would result in a lower LCL. Please clarify the physical basis for this relationship.
L 471-481: The link between CE intensification and the subsequent statements about a decrease in the number of wet days and an increased contribution of extreme precipitation to the total precipitation is not clear. CE intensification does not directly imply changes in either wet-day frequency or the relative contribution of extreme precipitation. Please provide evidence supporting these relationships or clarify the basis for these conclusions.
L 481-483: I find the argument difficult to follow, as this statement and the one in L. 475 appear to be contradictory. The authors first state that tropospheric RH generally does not change with warming, but subsequently argue that slower ocean warming limits moisture supply needed to maintain continental RH, resulting in a drying influence. If these statements refer to different atmospheric layers or spatial domains (e.g., tropospheric versus near-surface/continental RH), this distinction should be clearly explained. Otherwise, please clarify how these two arguments are consistent.
L. 500: Please add the missing parentheses.
I would suggest adding as a limitation the method used to identify convective events. In particular, precipitation exceeding the median over a 6-hour period at a 1° × 1° grid cell does not necessarily imply that the precipitation is convective. Future versions of the approach could consider incorporating a more direct indicator of convection, such as lightning observations.
L 505: The ending “is recommended” is awkward and appears grammatically inconsistent with the beginning of the sentence (“Future research should focus on…”). Please revise the sentence for clarity.
L. 515: The statement that the study “addresses a critical gap in the literature by focusing on convectively favourable conditions rather than the precipitation itself” seems to overstate the novelty of the study. As written, it suggests that previous studies have primarily focused on precipitation and have not investigated environments favourable for convection, which is not the case (e.g., Lepore et al., 2021). Please clarify what specifically is novel about the present approach and revise this statement accordingly.
L 518: The statement “it is still possible that a CE does not rain” is unclear, please clarify.
Figures:
Fig. 2c: “Basline” should be corrected to “Baseline.” In addition, the caption does not describe what each curve/color represents. Please provide a clear description of the different curves (here and in the other figures where necessary).
Fig. 9a: Both the x- and y-axes are labelled MUCAPE, whereas MUCIN is referred to in the text.
Supplementary Document
L. 45: Please remove the extra space.
Figure captions: Please provide sufficient information in the captions to make the figures understandable independently. In particular, describe the meaning of the different lines/colors and add y-axis labels where necessary.
Citation: https://doi.org/10.5194/egusphere-2026-3288-RC2
Data sets
Dataset for "Convective Environments Over the Arabian Peninsula in Current and Future Climates: A Machine Learning Approach" Ahmed Homoudi https://doi.org/10.5281/zenodo.18514414
Viewed
| HTML | XML | Total | Supplement | BibTeX | EndNote | |
|---|---|---|---|---|---|---|
| 117 | 41 | 14 | 172 | 36 | 13 | 13 |
- HTML: 117
- PDF: 41
- XML: 14
- Total: 172
- Supplement: 36
- BibTeX: 13
- EndNote: 13
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
The manuscript argues that statistical downscaling framework using machine learning can help identify convective environment for the current and future climate condition over the Arabian Peninsula. This approach is computationally less expensive than that of physics-base dynamical downscaling. The investigation is well thought. First the authors tested evaluated the performance of Random Forest, XGB, and Deep Learning. Based on the evaluation, XGB outperforms the other ML models. The investigation of convective environment is utilized the XGB model. The results exhibit a landscape of potential forecast skills of the XGB model in identifying the convective environment for the current and future climate over the Arabian Peninsula. This ML approach is promising in terms of reducing computation expenses and time. The manuscript is also well written though some figures are not clear and not easy to read and understand. Some arguments and discussion still need clarification.
Here are my major comments:
Here are my minor comments: