the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
What surface radiative fluxes reveal about Arctic cloud modelling accuracy
Abstract. Low-level clouds exert a strong control on the Arctic surface energy budget, yet their representation in regional atmospheric models remains a major source of uncertainty. We evaluate the Weather Research and Forecasting (WRF) model against observations from the Norwegian Young Sea Ice Experiment (N-ICE2015), conducted north of Svalbard from polar night to polar day. The analysis focuses on downward surface shortwave (SW↓) and longwave (LW↓) radiation under synchronous cloudy conditions to diagnose cloud-related radiative biases. While near-surface meteorology is generally well reproduced, pronounced seasonal radiative errors emerge. A dominance analysis based on a simplifed two-layer emission framework shows cloud emissivity, primarily controlled by liquid water path (LWP), is the leading contributor to LW↓ errors. During spring transition, the model underestimates cloud occurrence and simulates optically too thin clouds, leading to excessive SW transmission and insuffcient LW trapping. During polar day, a marked negative SW↓ bias develops. Radiative errors are largest for LWP below 30–40 g.m-2, where cloud optical properties are highly sensitive to variations in liquid water content. Sensitivity experiments demonstrate that improved representations of sea ice cover and surface albedo reduce polar day SW↓ biases, while modifying prescribed cloud droplet number concentration alters optical thickness but introduces compensating errors. Clouds diagnosed as surface-decoupled exhibit lower LWP and larger radiative biases, and this regime is overrepresented in the model. These results highlight the need for consistent representation of surface properties, boundary-layer structure and mixed-phase microphysics to improve simulations of Arctic surface radiation.
- Preprint
(1560 KB) - Metadata XML
-
Supplement
(524 KB) - BibTeX
- EndNote
Status: final response (author comments only)
-
RC1: 'Comment on egusphere-2026-2516', Anonymous Referee #1, 27 Jun 2026
-
AC1: 'Reply on RC1', Yaël Le Gars, 28 Sep 2026
We thank the reviewer for a careful and constructive reading of our manuscript. The comments identified important weaknesses in the presentation and structure of the study, as well as in the justification of several methodological choices. We take these comments very seriously and outline below, point by point, how we have addressed each one in the revised manuscript. In particular, we have updated our cloud diagnostic methodology, extended our evaluation framework, and incorporated new sensitivity experiments (WRF-neXtSIM, WRF-Li, WRF-50CDNC, WRF-10CDNC). The two-layer emission model / dominance-analysis section has been moved to the Supplement, reducing its role in the main text to a short, clearly caveated paragraph presenting it as supporting evidence only. We also separated more clearly the presentation of the methodology, model evaluation, and model sensitivity experiments and new developments. These changes aim to clarify the objectives of our study to the reader: first, to evaluate the WRF model, a regional model widely used to simulate the Arctic climate, second, to identify the likely physical origin of these model biases, and third, to explore the sensitivity of, and correct, the representation of these physical processes, or identify additional development perspectives for future work. For this reason, we also suggest a change of title to "Identifying and correcting causes of Arctic cloud radiative biases over sea ice in the WRF model". Reviewer comments are reproduced in italics; our responses follow in plain text.
Response to Reviewer 1
General commentsReviewer comment: I am struggling to understand why the authors chose to limit their analysis to comparisons with data from the N-ICE2015 campaign when there have now been sever other observational campaigns collecting radiation data in the Arctic sea ice that span multiple years, seasons, and some of which have more comprehensive datasets than N-ICE (see MOSAiC, SHEBA, AO2018, ARTofMELT). Some of these campaigns have collected measurements of liquid water path – which would go along way towards helping to answer some of the outstanding questions in this study, as well as co-located measurements of cloud base height and more frequent radiosonde launches. The data are all published and easy to access, the analysis would be same, but the statistical results would be much more compelling. Perhaps there is a resource limitation with model runs, but this ought to be less of an issue now than it used to be - I don’t think there is a good excuse to not include more data in this analysis.
Author response: We agree that additional campaigns, in particular MOSAiC and ARTofMELT, offer more comprehensive microphysical retrievals (e.g. independent LWP, cloud-base height) that would strengthen the statistical robustness of our conclusions. Our choice of N-ICE2015 was motivated by
(i) its unique continuous coverage from polar night to polar day within a single drift, which is essential for our season-resolved bias-attribution framework,
(ii) by the availability of co-located IAOOS lidar and radiosonde data along the same track,
(iii) and the regional specificity. Unlike SHEBA or the central Arctic path of MOSAiC, N-ICE2015 took place north of Svalbard, in the European/Atlantic sector of the Arctic. This region is a crucial gateway characterized by massive winter advancements of warm and moist air masses (as seen during the P1 storm periods) and a rapid transition toward the Marginal Ice Zone (MIZ). Evaluating WRF in this specific regime, characterized by thin, young sea ice, is complementary to studies focused on the perennial central Arctic pack ice.
We have clarified this rationale in the revised Introduction, added a paragraph to the Discussion (Sect. 5) acknowledging the limitation of a single-campaign evaluation, and identify a multi-campaign extension (particularly incorporating MOSAiC/ARTofMELT LWP retrievals) as a priority for future work. A full multi-campaign reproduction is, however, beyond the scope of the resources available for the present revision.
Reviewer comment: Although I found this study very interesting, I am finding it difficult to see how this drives the community forwards towards improving the representation of the modelled cloud radiative effect at the surface in the Arctic. Many of the results support things which we already know (the high sensitivity of cloud radiative properties to small changes in liquid water path when liquid water path is low, the importance of albedo and thermodynamic profiles, the seasonal variations in physical drivers). It’s not completely clear to me (a) why the authors have targeted a WRF-NICE2015 evaluation specifically in order to glean more information that can be used to improve Arctic cloud and energy budget simulations, (b) how these results can now be used to help address the hard problem of improving the simulations, and (c) how these results generalise more broadly to some of the models discussed in the introduction (GCMs, regional models). I think these deficiencies can be addressed simply by adding a more thorough literature review of what previous model evaluations studies have found, adding some specific motivation for why they have selected the model and analysis they have and how it fits into the bigger picture, and adding some further discussion about what the next steps forward should be now. I think this would greatly improve the study.
Author response: We have expanded the Introduction with a more thorough review of previous regional and global model evaluations of Arctic cloud radiative biases, we expanded the discussion on our choice of the NICE2015 campaign, and we justified our use of the WRF model. We now explicitly motivate the choice of the WRF model (given its widespread use in Arctic reanalyses like ASRv2 and regional ensembles like Arctic-CORDEX and AMAP) and explain how our evaluation framework (combining LWnet-based cloud identification, coupling-state diagnostics, and targeted sensitivity experiments (WRF-neXtSIM, WRF-Li, CDNC) provides actionable physical levers to improve model performance. We expanded the conclusions to summarize our findings as actionable recommendations for other modeling groups, including what the next steps forward should be.
Reviewer comment: The organisation of the text makes it a bit hard to follow at times. Specifically, the results section (section 3) is very long, and many of the subsections here contain background, methodological description, and discussion in addition to results. I think it needs reorganising so that all of the methods are in the methods section, and discussion in the discussion section ect – this would help make it more readable and also more useful. I don’ think that the introduction and the data and methods section accurately prepare the reader for what to expect currently. This should be easy to fix with some re-organisation.
Author response: We have followed this advice and restructured the manuscript so that all background and scientific rationale are moved to the Introduction, methodological justifications (including all sensitivity experiments) and diagnostic criteria are unified in Sect. 2, the evaluation and bias diagnosis of the control simulation are in Sect. 3, and the incremental impact of the physical levers tested in the sensitivity experiments are presented in Sect. 4.
Reviewer comment: I think the data and methods section is deficient. There is very little information on why the specific methods have been selected, there is no information or discussion of measurement uncertainties and how they might impact the analysis. There are a lot of methodological choices made in the results section that are not introduced in the data / methods sections – these should be outlined in advanced to help the reader make sense of the direction and purpose of the paper.
Author response: We have expanded Sect. 2 : (i) explicit justification for each methodological choice (LWnet cloud-identification threshold, CLT threshold synchronous cloudy sets, RH-diagnosis cloud layer & virtual potential temperature gradients for coupling state, and the MB/NMB statistics), (ii) a discussion of N-ICE2015 instrumental uncertainties following Walden et al. (2017) (radiometer calibration and radiosonde temperature/humidity uncertainties), and (iii) all methodological descriptions currently embedded within Sect. 3 subsections, consolidated here.
Reviewer comment: There are several occasions where sentences are ambiguous or lacking in specific detail that leaves the results open to different possible interpretations, these should be addressed and
hopefully I’ve mentioned them all in specific comments below. Please check that every sentence, especially when stating results, is clear and unambiguous – in particular by clearly stating what you are referring to in each sentence (models / observations / wrf runs / variables) rather than relying on the reader interpreting context from previous sentences. Adding specific numbers rather than words such as ‘strong’ or ‘less accurately’ would also help.
Author response: We have systematically revised the text to ensure that every statement explicitly names the variable, period (P1, P2, P3), simulation run (WRF-CTRL, WRF-neXtSIM, WRF-Li, etc.), or observational dataset being referenced, replacing subjective descriptors with exact quantitative metrics (NMB, MB, RMSE, r, percentage occurences…). Specific instances raised by the reviewer are addressed individually below.
Specific commentsReviewer comment: Since the cloud fraction comparison is discussed in detail, and included in the abstract and conclusion, I think some data for this needs to be included – I would prefer to see a figure in
the main text. But there should at least be one in the supplement.
Author response: We have moved Table B1 into the main text, where it is now referred as Table 3, showing the initial cloud occurrence statistics in the control simulation. We have also added an extra figure (Fig. 7) (cloud occurrence statistics) into the main text (Sect. 4 ; see also Reviewer 2's related comment), which compares observed vs simulated cloud occurrence frequencies across all the three periods for both LW_net and CLT criteria across WRF-CTRL, WRF-neXtSIM, and WRF-Li.
Reviewer comment: Line 35: A formal definition of the cloud radiative effect at the surface is warranted here.
Author response: We have added a formal definition of surface CRE : “the difference between the all-sky and clear-sky net (downward minus upward) surface radiative fluxes” in the Introduction,
Reviewer comment: Line 51: What do you mean by “derive model cloud evaluation”? Can you be more specific about what you’re trying to achieve here?
Author response: We have rephrased this in Sect. 1 to clearly state our two explicit objectives : First, to evaluate clouds and associated radiative fluxes at the sea ice surface in the WRF model and identify the processes driving model biases and in SW↓ and LW↓ radiation, and, second, to improve model results by modifying the representation of these processes.
Reviewer comment: Line 55: Why WRF? Why NICE?
Author response: We have added explicit justification in Sect. 1 and 2: WRF is a benchmark regional model used extensively across the Arctic (e.g., ASRv2, CORDEX, AMAP), and N-ICE2015 is selected for its unique continuous observational coverage spanning polar night to polar day over thin, young sea ice in the European Arctic region, with co-located radiative, meteorological, radiosonde and lidar (IAOOS) measurements, enabling a season-resolved process evaluation within a single consistent simulation and observational framework.
Reviewer comment: Figure 1: Please describe what the sea ice cover concentration map is showing in the caption and include the reference. I think it would also be good to have mean sea ice cover during N-
ICE plotted on the background of panel B?
Author response: Figure 1 caption has been expanded to specify that it displays sea ice cover concentrations on 20 January 2015, from the NSIDC-0079 dataset (panel A), and shows the Lance drift trajectory with color-coded dates and floe breakup breaks (panel B), along with the mean sea ice cover during the N-ICE2015 campaign.
Reviewer comment: Section 2.1: A short discussion of the uncertainties in the radiation measurements are warranted here, to contextualise the magnitude of the bias results.
Author response: We have added the instrumental uncertainty values reported by Walden et al. (2017), Cohen et al. (2017) and Kayser et al. (2017) for the N-ICE2015 broadband radiometers, surface meteorological instrumentation and radiosondes, noting that these instrumental uncertainties are small relative to the simulated model biases.
Reviewer comment: Section 2.2: Can you justify why you have selected these particular parameterisation schemes? And what influence you expect these decisions to have on the results of this study and how they might be interpreted?
Author response: We have added justification for the RRTMG radiation / MYNN2 PBL / Noah LSM / FNL configuration, selected based on optimal performance for Arctic boundary layer meteorology and vertical profiles, demonstrated in Marelle et al. (2021, 2025) and Ahmed et al. (2023). We also acknowledge the limitation of using Thompson’s 1-moment liquid treatment of CDNC sensitivity testing (Sect. 2.3 and 4.3). See also Reviewer 2's comment on Line 81.
Reviewer comment: Figure 2:
o Why are air temperature and RH shown with an hourly resolution but the radiation shown with daily resolution. All quantities show daily cycles, so it seems odd to not have them at the same resolution. I think the hourly resolution would be more informative for the radiation data.
o Are all the data gaps during the ship repositioning? Or are there other reasons for the missing data?
o Why does the simulated RH drop to zero briefly sometimes? It would be nice to start the y-axis at 40% RH so you could see more of the main data variability.
o I think this figure could be improved by highlighting the three periods described in the text, and indicating where the storms are that are discussed for P1. Is missing data completely due to ship movement? Or are there other periods of missing data in there?
Author response: We address each point below :
- The SW↓ and LW↓ fluxes are shown as 24-hours averaged in Fig. 2c,d because their strong diurnal cycle makes the hourly time series difficult to read at the scale of the full campaign. The daily averages remain informative regarding the seasonal trends discussed in the text. All statistics reported in Table 2 and 5 are nonetheless computed from hourly values, consistent with T2m and RH2m.
- We have clarified this point in Sect. 2.1: the two major gaps visible in Fig. 2a,b,d (17 Feb-1 Mar and 15 Mar-20 Apr) correspond to floe breakup and repositioning of the research vessel. Shorter, minor gaps result from the exclusion of radiative flux measurements flagged as low-quality (described in Sect. 2.1)
- The brief episodes of near-zero simulated RH are unexplained. Following the reviewer’s suggestion, we have restricted the y-axis to start at 50% RH, to improve the readability.
- The boundaries of periods P1, P2 and P3 are now indicated in panel (a) of Fig.2.
Reviewer comment: Are the three periods you discuss the same as other studies have used?
Author response: Our definition of P1 (polar night) is consistent with the winter regime described for N-ICE2015 by Cohen et al. (2017). However, the P2/P3 boundary we had originally chosen (25 May), based on a SW↓ criterion, differed from the seasonal transition identified by Cohen et al. (2017), who place the start of summer, associated with snowmelt onset, at 1 June. We have revised our P2/P3 boundary to 1 June to align with this definition. This boundary is now used throughout the manuscript, and the figures and tables that use the seasonal delimitation have been updated accordingly.
Reviewer comment: Clear sky thresholding (line 162) – where these numbers determined by statistically
comparing the distributions? Or just eyeballed? Also I don’t think you need to use ‘centered around’ – just give the actual specific value..
Author response: Values are now identified taking local maxima of kernel density estimated on observed LW_net PDFs (Sect. 2.4). The exact modal peak values are now reported (instead of “centered-around”): clear-sky modes at -44, -72 and -72 W/m2 for P1, P2 and P3; opaque-cloud modes at -3, -14 and -10 W/m2.
Reviewer comment: Line 166: I would like to see some more about this in the discussion section – optically thin clouds can have an outsized and non-linear impact on the surface radiation components as
you demonstrate. Can you speculate what impact underestimating the occurrence of optically thin clouds might have on the results of this analysis and their interpretation?
Author response: We have added discussion in Sect. 2.4 : while LW_net thresholding reliably isolates optically thick, liquid clouds, thinner ice clouds (especially in winter/P1) fall into the inter-modal region. Consequently, the retained cloudy subset is intentionally biased toward radiatively dominant, optically thick clouds. Thus, the radiative errors reported here (under optically thick cloud conditions) are not affected by the optically thin clouds, excluded from the analysis. The model’s ability to represent this thinner cloud regime falls outside the scope of the present study, and remains an open question for future assessment.
Reviewer comment: Line170: “The latter results from an underestimation of the upwelling component.” Please expand on this, how do you know this? Why is this the case?
Author response: It has been clarified in Sect. 3.2 : in P1 clear-sky conditions, the simulated upwelling LW mode is shifted by -16 W/m2 relative to observations, partly explained by a near-surface cold bias, which shifts the clear-sky LW_net mode to less negative values (-34 simulated vs -44 W/m2 observed).
Reviewer comment: Line 174: Please describe how you are determining whether or not the simulated and observed distributions “agree well” and why? The uncertainties of the measurements and model are important for determining whether they actually agree well or not.
Author response: We have replaced this qualitative statement in Sect. 3.2 with explicit quantitative modal peak differences (5 W/m2 shift in P2 modal locations) and NMB values.
Reviewer comment: Line 180: It would be nice to see how sensitive the overall study results are to this
thresholding on both datasets
Author response: As already noted in Sect. 3.2, varying the CLT threshold from 0.95 to 0.75 alters simulated cloud frequencies moderately (e.g. from 39% to 53% in P1, and 50% to 55% in P2) without altering downstream radiative error trends.
Reviewer comment: Line 185: “Primarily” – replace with likely? I don’t think you’ve actually shown this is true or have the data to do so
Author response: We have amended the wording accordingly in Sect. 3.2.
Reviewer comment: Line 189: A plot of how the cloud occurrence varies with the threshold would be helpful here, even in the supplement.
Author response: We have not added a full sensitivity curve. However, Table 3 now reports cloud occurrence statistics using the LW_net threshold (the same criterion applied to the observations) which allow a direct, like-for-like comparison of cloud occurrence between observations and the model under a single and common criterion; something the previous comparison could not offer, since it relied on two different criteria (LW_net for the observations and CLT for the model), and therefore sensitive in a certain extent to the choice of CLT threshold. The occurrence fractions obtained with the CLT criterion, which is the criterion used throughout the rest of the study, are also reported.
Also, a new figure (Fig. 7) addresses how these occurrences vary with the sensitivity experiments.
Reviewer comment: Line 189: “and shows weak temporal overlap (50 %) “ – what does this mean? How do you calculate a percentage temporal overlap? Please show this if you’re going to draw conclusions from it.
Author response: We now explain this more clearly in the text. This percentage refers to the proportion of observed cloudy cases that are also classified as cloudy in the model at the exact same time. When we later refer to the set of cloudy cases with temporal overlap, we refer to these synchronous observed/modeled cloudy cases. This is distinguished explicitly in the manuscript from analysis where exact synchronicity is not needed, and where we might want for example to verify if the model reproduces the right statistics of cloud occurrence during a given period even if the timing is not synchronous.
Reviewer comment: Line 189: Change to: indicating that “the observed” opaque clouds are rarely simulated
Author response: We adopted this suggested phrasing in Sect. 3.2.
Reviewer comment: Line 193: What does slight to moderate mean and how have you determined this? Please show these results in the supplementary.
Author response: It has been quantified in Sect. 3.2 with exact numerical ranges.
Reviewer comment: Section 3.2.1: The discussion of cloud occurrence in models versus obs needs supporting with data / plots if discussed in this much detail, but I understand it’s perhaps not the main focus, so could move the discussion to the supplement. I prefer the former though because I think a lot of readers might be interested in this.
Author response: We agree, and cloud occurrence is now treated as a substantive element of the analysis rather than a secondary result. The cloud occurrence table, originally placed in the appendix (Table B1), has been moved into the main text as Table 3, and is now supported by the addition of Fig. 7, which also shows how the sensitivity experiments affect the occurrence fractions of opaque clouds.
Reviewer comment: Throughout, there is not much discussion of the fact that cloud height and vertical temperature profile errors could be another reason for errors in downwelling longwave variation in the model compared to observations. I understand that this is difficult to validate with the data you are currently using, but I think you need to caveat that this might still be the case in the abstract, discussion, and conclusion. You are effectively adding any errors in cloud height and temperature profiles into errors in cloud emissivity by assuming that the T cloud base = T surface, and therefore possibly overestimating the sensitivity to cloud emissivity. This is a really important point to discuss, since it is one of your main conclusions.
Author response: We agree that this is an important caveat, and have addressed it in coordination with Reviewer 2’s related comment on the Tcb = T2m assumption. We have reduced the scope of the two-layer emissivity model dominance analysis, which relied on this temperature assumption on the observational side: it is no longer presented as a main-text result, but has been moved to the Supplement (Table~S1) , to avoid overstating the relative contribution of cloud emissivity given this simplifying assumption.
Also, we now explicitly note that the cloud emission temperature (itself potentially biased by errors in cloud-base height or in the vertical temperature profile) also contributes to LW↓ variability (not quantified), while making clear that the SW↓/LW↓ error pattern still points to cloud optical thickness and cloud emissivity as the dominant driver to first order.
Reviewer comment: Line 199: “The distribution shape is retrieved” – what does this mean? Please be specific. Please can you give the specific value instead of “about” ?
Author response: We have clarified this sentence. “The distribution shape is retrieved” now explicitly states that the model reproduces both modes seen in the observed LW↓ distribution during P1. We have replaced the vague "about 10 W/m2" with the exact value (“underestimated by 7 W/m2”).
Reviewer comment: Line 206: P3 radiative analysis. Low values of downwelling longwave are overestimated and high values are underestimated. But for the shortwave, low values are overestimated while high values are underestimated. This suggests that in the model, you have more clouds with a high optical depth (less shortwave transmission) but less longwave emission (i.e. colder clouds) – could this be errors in cloud height? Too many high clouds? Or not enough liquid water? Maybe add some of this to the discussion if it’s not already covered.
Author response: We have added a discussion in Sect. 3.3 : during P3, SW↓ underestimation at high LWP occurs alongside saturated LW↓, indicating that while clouds behave as blackbodies in LW↓ above ~40 g/m2, excessive cloud optical depth (or CDNC/multiple scattering errors) continues to reduced SW↓.
Reviewer comment: Line 209: “Likely reflect biases in cloud optical depth” –why likely? what else could they be? What about multiple scattering and cloud edge enhancement? Does surface albedo impact downwelling shortwave by a significant amount?
Author response: We broaden the discussion of possible SW↓ bias sources in Sect. 3.3 and specifically in Sect. 3.5. Surface albedo, including multiple reflections between the high-albedo snow/ice surfaces and the cloud base, and its impact on SW↓, is discussed in detail in Sect. 3.5. We also mentioned 3D cloud-edge enhancements in Sect. 3.3 as a plausible contributing factor.
Reviewer comment: Please make sure you clearly define what you mean by CRE somewhere – there can be several different definitions of this.
Author response: It has been explicitly defined in Sect. 1. See the response above.
Reviewer comment: Line 222: The IAOOS dataset should be introduced and described in the datasets / methods section.
Author response: We have moved the IAOOS dataset description to Sect. 2.1.
Reviewer comment: Line 225: I think the assumption that the cloud base temperature = surface temperature in the observations is awkward and problematic in this analysis. Some specific comments to
address:
- a) How does the Maillard identification of cloud presence line up with the observed cloud presence from the radiometers (are they seeing the same clouds?). I ask this because the 808 nm lidar will be biased towards seeing low clouds anyway as it will suffer from attenuation higher up.
- b) “We do not have any direct measurements of cloud base temperature” – is sounds like you actually do, at least twice a day from radiosondes. At the very least you should validate your assumption using these data.
- c) There are several other campaigns that have more frequent atmospheric temperature profiles as well as better measurements of cloud base height, like MOSAIC and ARTOFMELT for example.
- d) Another way of testing if you’re assumption is reasonable: If you’re assuming all the clouds have the same temperature as the surface, you’re assuming that all the net longwave variation in the cloudy mode of figure 3 is due to changes differences in cloud and surface emissivity – can you demonstrate whether or not that assumption is reasonable? How much do you expect these to change?
- e) By making the assumption above, perhaps it’s not surprising that you find changes in cloud emissivity to be of such high importance? Can you comment on this in the discussion and what it might mean for your results?
- f) Even if all the cloud bases were located below 90 m, that is not a good argument for assuming that the cloud base temperature and 2 m temperature are the same. Low level temperature inversions are common in the Arctic, you should be able to have a look how relevant these are for NICE by examining the radiosonde profiles, and show some data for that in the supplement.
Author response: We agree that the assumption Tcb = T2m is.
(a) IAOOS lidar cloud occurrences (Maillard et al., 2021) are consistent with radiometer-based opaque cloud states during N-ICE2015; (b-f) (b-f) Maillard et al. (2021) demonstrated that cloud bases were overwhelmingly below 90 m with weak temperature gradients (0.6-2°C in the lowest 100 m), making T at 2 m a reasonable bulk proxy. Radiosondes launched twice daily confirmed that sub-cloud temperature inversions are weak during cloudy episodes.
In order to better reflect the uncertainties in the two-layer model decomposition, which were also highlighted by reviewer 2, we now explicitly state that the two-layer model is an illustrative variance-decomposition tool, and note that errors in cloud-base height are subsumed into the emissivity term. We also de-emphasized the discussion of these results, since our main findings in Section 3, that motivate later changes, are that errors in LWP, boundary layer structure, and surface processes seem to be driving model errors - the 2-layer model discussion being additional but not load-bearing evidence for the LWP issue, even though it could also point to issues in the simulation of cloud effective radius contributing to motivating our exploration of the sensitivity to CDNC in Section 4. For this reason this analysis is now mainly presented in the supplement and only the main results are mentioned in the text as supporting evidence.
Reviewer comment: Line 233: Please can you explain why you have chosen these thresholds? Are they based on literature from elsewhere? Are they a conservative threshold for cloud presence compared to
the LW emission threshold from the observations or not? Perhaps most importantly, are the results sensitive to this choice?
Author response: We thank the reviewer for pointing out the potential ambiguity of using distinct hydrometeor mass mixing ratio thresholds (10-5 kg/kg for LWC and 10-6 kg/kg for ISWC). To ensure full methodological consistency across the manuscript and eliminate this issue, we have updated the cloud-base layer identification in the supplementary two-layer dominance analysis (Supplement Sect. S1) by adopting a unified relative humidity threshold (RH > 95% with respect to liquid water), following the reviewer's recommendation.
Reviewer comment: Line 237: Is this from both the observations and the model derivations?
Author response: We have clarified this in Supplement Sect.S1. Instances with diagnosed epsilon_c > 1, resulting from strong surface-based inversions in the simplified two-layer framework) were excluded from both observational and model emissivity calculations.
Reviewer comment: Table 2: Please make the caption more descriptive / specific. You’ve defined this in the text but make it easier for your reader to interpret your results by being explicit in the captions. Something like “Relative contribution of the four variables from the two-layer atmospheric emission model to the variance in the difference between the modelled and observed downwelling longwave radiation during cloudy scenes (∆LW↓ ) “
Author response: We have adopted this suggested caption wording in Supplement Table S1.
Reviewer comment: Line 240: Do you have a reference for this exact analysis? If not please describe it in more detail with equations (ideally in the methods section).
Author response: The analysis and their comments were moved to the supplement, as detailed above. There, we also detail our methodological choices.
Reviewer comment: Line 244: The properties you list are specifically cloud microphysics as opposed to cloud properties more generally which might include cloud height, depth, temperature, so change
“cloud properties” à “cloud microphysics”
Author response: We make this terminology correction throughout the new manuscript (see Sect. S1).
Reviewer comment: Line 245: The values in table 2 are reflecting your assumption that T2m=Tcb in the
observations as well as biases or problems in the model, how can you tell how much of this just comes from the former or the latter? I think some further analysis is warranted here.
Author response: We removed the former Table 2 from the main manuscript. We agree that Table 2 (now referred as Table S1) cannot cleanly separate genuine model biases from the effect of the Tcb = T2m assumption. in the observational side. We did not attempt a full quantitative attribution of this kind, and instead treat the dominance analysis conservatively, as stated in Sect. 3.4.
Reviewer comment: Line 246: You go from discussing the contribution of different variables to the variance in the model-obs bias (I think), and then you talk about “errors in Tcb” – are you now talking about the contribution of Tcb to that variance – or are you talking about something else? The difference between Tcb in the model and the assumed Tcb in the observations perhaps? Please be very specific and consistent with your language here. If you’re talking about the latter, you probably need a new paragraph, and it would also help to see some of this data so the reader can interpret what you mean by substantial (or at least some numbers).
Author response: We have restructured this paragraph in Sect. S1 to clearly distinguish the statistical variance contribution of T_cb from direct physical biases in model temperature profiles and cloud-base heights.
Reviewer comment: Line 247: The “temperature profile misrepresentation” could presumably either be a misrepresentation in the model, or an incorrect assumption that T2m == Tcb in the observations – please add that detail here.
Author response: We make this distinction explicit in the revised text (Sect. S1).
Reviewer comment: Line 248: Change “well captured” to “well captured by the model” – it’s important to be specific. Also change “with 95% within the first model level” to “with 95% of observed cloud base heights falling within the first model level” (if that is really what you mean).
Author response: We have adopted the suggested wording in Sect. S1.
Reviewer comment: Line 249: “Thus, Tcb ≈ T2m during P1” - do you mean in the model here? Please be specific! Do you mean less accurately captured compared to the IAOOS observations? If so, please show some data, and also introduce the IAOOS as a dataset that you’re using for analysis in the data and methods section. In what way are they less accurately captured? Is the IAOOS data showing more low cloud bases and the model showing higher? Or the other way around?
Author response: We clarified in Sect. S1 that during P2 and P3, simulated cloud bases are higher (mean ~200-300 m) than IAOOS observations (<90 m), likely introducing a T_cb bias.
Reviewer comment: Line 251: Versus what value for the observations? Please show some of this data / comparison if you’re using it to derive results.
Author response: Observational comparison values from IAOOS lidar are explicitly stated in Sect. S1.
Reviewer comment: Line 256: Why do you choose modelled cloud water specifically here? When you have also stated that radiative errors could be driven by particle size? I think that this is the right thing to focus on because the impact of cloud water on the LW emissions is so large – but guide a reader who might be less familiar with the relative radiative effects of changes in cloud microphysics.
Author response: We have added an explanatory sentence in Sect. 3.4 noting that surface LW↓ and SW↓ flues exhibit first-order sensitivity to LWP in the low-LWP regime (<30-40 g/m2), while particle size plays a secondary modulating role.
Reviewer comment: Line 259: this could easily be solved by using one of the many Arctic field campaigns with microwave radiometer data…
Author response: We note this explicitly in Sect. 5 (Conclusion) as a key motivation for extending this methodology to campaigns like MOSAiC or ARTofMELT.
Reviewer comment: Table 3 caption: add ‘relative to observations’ and define which way you’ve calculated the bias (model minus obs?)
Author response: Table (now) 4 caption has been updated to explicitly specify “model minus observations”.
Reviewer comment: Table 3: Please explain in the methods section why you have chosen to focus the evaluation on the statistics MB, NMB and NMAE.
Author response: We have added this justification to Sect. 2.6 : MB and NMB quantify systematic directional errors in absolute and relative terms across seasonal periods with contrasting solar irradiances, while RMSE (in replacement of NMAE) captures absolute error magnitude.
Reviewer comment: Line 267-268: Please be specific about what you mean by low. Instead of using the word strong, can you put some numbers to this? The word ‘strong’ is quite subjective without relative context.
Author response: We have linked this phrase (“[...] low LWP”) to the following one, which contains the LWP range values we are referring to. The qualitative terms have been substantiated with the corresponding numerical bias values (e.g. SW↓ overestimation up to +75% and LW↓ underestimation down to -30% at low LWP) in Sect. 3.4.
Reviewer comment: Line 269: What do you mean by mirroring?
Author response: We have rephrased in Sect. 3.4 for clarity to “ a comparable reduction of SW↓ and LW↓ errors with increasing LWP”.
Reviewer comment: Figure 5: Can you use the same colour bars and y-axis for the all the lw plots and the same colour bars for all the shortwave plots so that the reader can compare the magnitude of the errors between the periods?
Author response: The colour scales have not been modified, as an homogenisation would result in loss of visible contrast for periods relative to others with a much larger range of bias values. Y-axis ranges within each radiative band across periods have been standardised.
Reviewer comment: Line 286: I believe this could also come from too much ice? I.e. clouds that are too deep generally?
Author response: We have added in Sect. 3.4 this alternative interpretation : excessive cloud optical depth in summer can also result from deep ice/mixed-phase clouds.
Reviewer comment: Line 290: Is this still the case when you only consider cases when ISWP is > 0 ? Certainly, radiative biases in the liquid only cloud are not related to ISWP misrepresentation, but I could
imagine that downwelling SW biases for deep mixed phase clouds could be related to misrepresentations of cloud depth and total water content (liquid +ice), can you address that?
Author response: We have added analysis in Sect. 3.4 and Supplementary Fig. S2 : restricting cases to ISWP>0 confirms that ISWP shows no systematic correlation with surface radiative errors, demonstrating that LWP remains the primary driver.
Reviewer comment: Section 3.3, I am a bit confused about the ordering of this. WRF-SEAICE clearly does a much better job of representing the surface albedo – surely it would make sense to show this first,
and then present and discuss the radiative bias results from the WRF-SEAICE run in the main paper rather than the WRF-CTRL run? With the WRF-CTRL results in the supplement?
Author response: In the revised manuscript, Sect. 3 evaluated the baseline performance of WRF-CTRL and identifies the origins of these biases. These findings motivate in Sect. 4 the sequential sensitivity experiments led by realistic sea ice representation. This type of structure (model evaluation and analysis, followed by model developments to address shortcomings, then updated model evaluation) is typical of model evaluation and development studies, and reflects the actual path of our investigations behind this paper. We have updated the manuscript’s structure and clarified the motivations to better explain why the study is structured this way. We also note here that we have replaced the sea ice fields origin from AMSR2 to neXtSIM, a high resolution reanalysis product assimilating the AMSR2 product, that has the advantage of also predicting sea ice albedo, snow depth and ice depth for sea ice.
Reviewer comment: Line 319: Please reference figure 6 here. It took me a while to make sure I was understanding this result correctly due to the lack of specificity in the sentence, ‘reduced SW underestimation’ is a double negative and the result is unclear (is it now an over estimation? – I see now from fig 6 that it is not), and a ‘systematic negative bias persists for LWP’ – the immediate interpretation of this is that there as a bias in LWP – but I think you mean ‘systematic negative bias in downwelling shortwave persists for LWP values greater than..’ . Please be very specific about what you are saying both here and elsewhere.
Author response: Fig. 6 along with the section 3.3.3 have been removed in the new manuscript. Sect 4.1 now presents the effect of improving the surface conditions.
Reviewer comment: Line 320: Later in the text you also say this could be due to phase partitioning?
Author response: We cross-referenced this discussion point in Sect. 4.1 and 4.3.
Reviewer comment: Figure 7 – the caption says WRF-CTRL but the figure says WRF-SEAICE?
Author response: Figure 7 is no longer presented in the manuscript.
Reviewer comment: Line 335: You haven’t shown or mentioned the LW changes in the CDNC runs – maybe add to the supplement? It would be interesting to know how the magnitude of the CDNC forced
changes in downwelling LW compared to the biases.
Author response: We added a discussion in Sect. 4.3, and added a figure (Fig. S9) in the Supplement to show the effect on the LW↓ errors.
Reviewer comment: Line 337: Maybe for the discussion section – what do we know about the variability of this on an event basis? The reference value of 100 cm-3 might be good for a seasonal mean, but the radiative biases are event specific, and I suspect you can get very low aerosol events and very high aerosol events in all seasons. Can you comment on whether getting the event scale CDNC closer to reality might help improve these biases?
Author response: We added this discussion point in Sect. 4.3, noting that prescribed uniform CDNC values do not capture synoptic aerosol intrusions or event-scale CDNC variability, which represents an additional source of uncertainty. We note, however, that within the results presented here, CDNC already emerges as a secondary lever, unable to resolve the radiative biases observed at either low or high LWP (Sect. 4.3)
Reviewer comment: Line 340: change “too efficiently” to “too efficiently in the model”
Author response: We changed our initial wording in Sect. 4.3.
Reviewer comment: Sect. 3.3.5: this section starts by some background / literature references, and then some methodology. This paper would be more readable if you could move the background / literature from each subsection into the introduction, and the methods from each subsection of section 3 into the data and methods in section 2.
Author response: Agreed; addressed under the general reorganisation described above : background moved to Sect. 1, coupling diagnostic methods moved to Sect. 2.5, and results presented in Sect. 3.6.
Reviewer comment: Figure 8: I find it very difficult to distinguish between the purple circles and blue squares or to see any real difference in these distributions from these plots, so I was left interpreting differences from the NMB calculations. I think you can at least make the two sets of points more contrasting colours and shapes. Perhaps you could also add a centroid like you have on the previous plots?
Author response: We revised the figure with distinct markers. We chose not to add the centroids: a key feature of this figure is the uneven density of points within each bin, in particular the marked over-representation of decoupled cases in the low-LWP tail of the distribution, which the centroids would not convey but in fact obscure.
Reviewer comment: Line 370, can you do a statistical test between the two different distributions to confirm whether these differences are statistically significant? I suspect that they are, but that would help support the point.
Author response: Mann-Whitney U and Kolmogorov-Smirnov test results have been added in Sect. 3.6, along with the Table S2 (Supplement), confirming statistically significant differences (p<0.01) between coupled and decoupled regimes for SW↓ errors, LW↓ errors and LWP.
Reviewer comment: Line 387: Presumably also CDNC would be different between coupled and decoupled clouds – it would be interesting to so how varying the CDNC impacts the radiative bias separately for coupled versus decoupled clouds.
Author response: We did not perform this specific analysis of the CDNC effect by coupling regime. However, we did analyse the effects of surface-condition changes (WRF-neXtSIM), the experiment that produced the largest changes in cloud occurrence and radiative biases, and that is most likely to have altered the near-surface temperature profile, on the coupled/decoupled distribution of clouds, discussed in Sect. 4.1
Reviewer comment: Line 399: Be clear on what is a novel result of this study and what is not – I think the dominance analysis identified that cloud emissivity was the leading driver of the LW errors – but the fact that this is primarily driven by LWP is not a direct result of this study, but just based on physics / past studies? If I’m correct, then please rephrase this sentence.
Author response: This sentence has been removed from Sect. 5, consistent with our responses to previous comments on these limitations, and with the removal of the dominance analysis from the main body of the manuscript.
Reviewer comment: Line 405: add “when averaged over the season” (i.e. there is still an overestimation of SW at low LWP in all seasons)
Author response: We revised the sentence to make explicit that, if the SW↓ overestimation exists at low LWP across all seasons, it is by far most pronounced during P2.
Reviewer comment: Line 410: I think you need to look at total cloud water content before you can state this as a conclusion – there may be no systematic relationship when you just look at ice and snow water paths because you’re considering both ice and mixed-phase clouds, but there could still be a relationship with total cloud water content (ice + liquid), i.e. it might be the ice in the mixed-phase that is important for the cloud radiation errors in the high LWP cases. These will be the deepest clouds – if the model is not capturing the ice phase the clouds might be too optically thin even if there are capturing the liquid phase correctly.
Author response: Total condensed water path (CWP = LWP+IWP) was evaluated in Sect. 3.4, confirming that liquid water dominated CWP and drives radiative errors.
Reviewer comment: Line 415: I wasn’t convinced that this was a ‘strong’ relationship – some statistical testing on the differences would help. As would more data from a wider range of campaigns.
Author response: Statistical significance has been reported in Sect. 3.6 and 4.1, and we adapted the wording in Sect. 5.
Reviewer comment: Code data availability: Where can we find the IOASS data?
Author response: We have added the appropriate IAOOS data DOI/repository reference to the Code and Data Availability section.
Citation: https://doi.org/10.5194/egusphere-2026-2516-AC1
-
AC1: 'Reply on RC1', Yaël Le Gars, 28 Sep 2026
-
RC2: 'Comment on egusphere-2026-2516', Anonymous Referee #2, 08 Jul 2026
This paper addressed surface radiative fluxes over the Arctic and uses observations from the N-ICE2015 campaign to evaluate WRF model simulations and diagnose the causes of model radiation biases. The topic is an important one as cloud-radiation-system interactions are one of the factors that influence Arctic climate, they continue to pose challenges for models, and there is a need to extract whatever information is possible from the limited existing observations. In these regards, the paper is an important contribution and is thematically within the scope of ACP. However, this manuscript has several major shortcomings that lead me to question its fundamental findings. Many employed methods have unquantified and undiscussed uncertainties and could be readily replaced with more consistent methods or more reliable data. Thus, while the manuscript’s goal of explaining the specific contributions to model bias is a good one, and some of the results could provide useful insight, I do not believe those results should even be considered until the major methodological deficiencies are resolved and the analysis can be re-done. This will require significant additional work in most parts of the study. At a minimum, this would be a case of “Major Revisions.” Specific comments are included below, distinguished into some of the major methodological issues and then more general comments.
Major Methodological Issues
Line 162: Why does the threshold change so much between the different periods of observations? During MOSAiC, Shupe et al. (2026) applied this same LWN threshold approach but found that -25 W/m2 worked for all seasons. Why the difference here? Is there a physical explanation or something markedly different about the atmospheric scenes in this part of the Arctic during winter? Also, these different definitions will make the different periods incomparable with each other and hard to interpret. I believe this difference manifests itself into some of the results presented later in the analysis, but its role is not discussed.
Line 176-181: As described in this paragraph, it is apparent that the cloud conditions are identified using LWN in the observations but CLT in the model. The reasoning presented for why this difference is implemented is the role of spatial variability in cloudiness that could impact the radiation measurements. That is a valid concern, however, the observational dataset does not have information on spatial distribution of cloudiness. That is an unavoidable limitation of these observations. But given the observations, it makes more sense to at least apply consistent definitions. Why not just apply the same LWN threshold approach to both observations and models so the different definition does not add further, unquantified, uncertainty to the comparison? The comparisons that are shown later clearly reveal some influence from this difference in definition, but that influence is not discussed nor quantified. Moreover, it really doesn’t make sense. “Clouds” have many definitions. Thus, to have any hope of legitimately assessing a model, it is important to at least define them in the same way between observations and models.
Line 201-205: It is a bit strange that the model shows a lot of low LWD values that are not observed. I believe this figure has only the cloudy sky subset of data described in the previous section (although that is not entirely clear). If so, these low simulated LWD values would have to come from clouds that are vastly thinner and/or colder than those that were observed. This signature would be consistent with a discrepancy in the definition of clouds. Since the observations are simply based on LWN threshold, they will all have a high LWD. Since the model results are based on this CLT, there is no guarantee they will have high LWD. In fact, based on later parts of this manuscript, it appears that “clouds” are also defined where there is ice/snow via the modeled ISWC. Thus, if I had to speculate, I would say that these low LWD values are from times when there is only ice/snow in the model with little to no liquid water. This is a demonstration of the apples-and-oranges comparison that results from the different cloud definitions.
Line 214-238. There are a lot of problems with this general approach to assessing the clouds and radiation, as outlined in the following six points:
1) Eq 1 includes some of the terms that contribute to LWD and is mostly correct when Ec is approaches 1. When Ec is significantly less than 1 this means there will be contributions from the atmosphere above the cloud (possibly from multi-layer clouds that are at a different temperature), which are not included here. This point is kind of mentioned at the end of the paragraph but is not included in the equation. Overall, this equation is approximately true only when Ec is high.
2) Having "almost all" cloud base heights less than 90m would be a remarkable result for observations over nearly 5 months extending from winter to early summer. It is divergent from other observations, including those from MOSAiC. I do not believe this statement is true. I have tried to assess the results of the Maillard et al. (2021) paper, but there is no actual lidar data shown with which to independently verify that the lidar really identified almost every cloud base to be below 90m. Thus, without further information, my interpretation of this result is that the “cloud base” referred to here is actually the base of “hydrometeors”, including falling ice. In that case, yes, the base of hydrometeors would most commonly be near or at the surface because there is quite often falling precipitation. However, this base is not always (or not often) the base that is important for radiation. I believe N-ICE had a ceilometer; why not use actual cloud base height measurements from a widely trusted system like that? Ceilometer reported cloud base would be more consistent with the radiatively-important cloud base.
3) Building further on cloud base: In the description of the model analysis, it is stated that cloud base is defined when LWC exceeds 10^-5 kg/kg or when ice/snow exceeds 10^-6 kg/kg. As outlined in the Introduction, ice and snow have vastly different interactions with atmospheric radiation (for a given amount of mass) due to microphysical properties. By defining things in this way, you will get artificially low cloud base heights under all-ice and mixed-phase conditions that will be identified based on weak ice/snow precipitation, which has a low total emissivity. Thus, it is not a good approximation of the location where most emission occurs. i.e., the assumed cloud emission temperature will be wrong. This incorrect identification of cloud base could be part of the reason you identify most modeled cloud bases to occur in the first model level, but I do not have enough information to know for sure.
4) On a related topic: The assumption that cloud base temperature and surface temperature are equal is highly uncertain, condition dependent, and effectively forces the implied LWN to be 0 W/m2 under thick clouds. In reality, there is usually a negative temperature lapse rate over some layer below cloud base (i.e., usually adiabatic ascent helps to form the cloud), which could overlie a variety of temperature structures at lower levels. In the best-case scenario (for the assumption made here) with a low, surface-coupled cloud the temperature difference might only be a couple tenths of a degree. However, for real cloud bases that will often be significantly >100 m above the surface, the difference becomes important. We already know that Eo is small and in the limit where Eo approaches 0, Eq 2 becomes simply LWD=Ec*sigma*T2m^4! This is a poor approximation.
5) The analysis uses Eq 1 for model data and Eq 2 for observational data. How then can the results be compared to each other? The method alone will have a large uncertainty that could mask other issues in the model. To truly evaluate the model, you must apply the same approaches to both observations and model results.
6) Simply throwing out cases with Ec > 1 is not a sufficient remedy for the shortcomings of this approach. In general, I would expect the smallest uncertainties to occur when Ec is highest (i.e, that is when Eq 1 is most accurate). Ec values > 1 do indeed indicate a bias, but it could be a smaller bias than cases with lower Ec values (which are retained in the analysis). The main point here is that the whole analysis is highly uncertain because the equations and assumptions used are suspect. The impact of these uncertainties has not been evaluated or discussed thoroughly. As a result, I have little confidence in all the downstream results that are presented.
Overall, because of all these points, I do not see how this approach can legitimately be used as presented here.
Line 239-242: I do not understand what information is being shown by this analysis. First, while the analysis examines the errors (LWD model-obs), it does a dominance analysis using Eo, T2m, Tcb, and Ec. I assume those parameters are from the model? Second, Ec was derived from T2m, Tcb, and LWD for the model results (Eq 1). Thus, the respective Ec values are directly dependent on the other parameters. Their contributions to LWD variability are pre-determined by the relationship through which they were derived, not based on independent physical variations. Thus, it is not clear to me how this analysis reveals information about the physics of Ec. Lastly, the assumption that T2m=Tcb (in the observational analysis) is undercut by the different results for T2m and Tcb. These temperatures apparently do not vary the same nor impart the same influence on LWD. If true, the assumption that T2m and Tcb equal each other in Eq 2 neglects some important physics.
Line 272: This point might indeed explain the results in Figure 5 for this dataset. However, based on LWP observations from MOSAiC it was found that the opaque cloudy state (i.e., LWN > -25W/m2) contained a lot of LWP values that were less than 30 g/m2 (see Dahlke et al. 2025, Shupe et al. 2026). Moreover, the fact that a LWN threshold of -10 W/m2 is used in P1 instead of -25 W/m2 means that the “clouds” identified in that period would naturally have higher LWP than the other periods (in a season that routinely has lower LWP and has lower background moisture such that the radiative effects of thin clouds are accentuated relative to the moist summer). First, I do not think there should be different LWN thresholds applied to the different periods. But if there are, the impact of this difference should certainly be discussed thoroughly.
Line 380-383: I have concerns about the accuracy of the “observed” coupling state results as they are based on a fixed cloud base height of 100m that may or may not be correct for any given case. The effect of this assumption on the results has not been discussed. Moreover, with a radiosonding dataset (and only 55 soundings) it should be readily possible to perform a robust determination of coupling state for the available cases. In soundings it is easy to identify a liquid saturated layer (typically a threshold of ~96% RH is used). Then, you can determine the equivalent potential temperature at the base of that cloud layer and the surface. Their difference gives you a clear delineation of the coupling state. This approach is far preferable to making an uncertain assumption about cloud base height.
General Comments
Full text: There are many issues with the written text in terms of grammar, sentence structure, and other issues. While I started noting these in my review, there are so many that I will not include them here. Before proceeding, this manuscript needs a full editorial review.
Line 21: “positive effect” on what? They have a net warming effect on the surface or a positive effect on the surface radiative balance. That should be clarified.
Line 22-23: This statement about “short period of surface cooling in summer” is true over highly reflective surfaces like sea ice. Over less reflective surfaces like land (after snow melt) or open ocean, the cooling period extends for multiple months. This cannot be considered “short.”
Line 24: Cloud radiative parameters are important of course. But this statement is more accurate as “Cloud radiative properties along with relevant properties of the surface and atmosphere……” Ultimately the surface albedo plays a substantial role in the net effect of clouds seasonally that can overwhelm cloud radiative properties.
Line 81: Having only 1-moment liquid leaves substantial deficiencies in representing the liquid phase, as has been shown in prior studies. This is particularly true when you are attempting to examine the effect of, for example, aerosol indirect effects where the second piece of information about the liquid droplet size distribution can be quite important.
Line 116: The model had a rather large warm bias during 14-16 March. What is the cause of this and does it provide any insight into the model issues?
Line 118-120: Importantly, this good agreement during P3 should not be surprising since the surface is melting at the time.
Line 121-126: Since RH is influenced by both the amount of moisture and the temperature, it is better to examine specific humidity to allow for the moisture and temperature to be evaluated separately. It is not clear here how much the temperature biases impact the RH compared to specific humidity biases.
Line 125-126: True, the near-surface RH biases might impact the representation of low clouds under circumstances when the cloud layer is coupled to the surface. However, under decoupled conditions this is likely not the case. Additionally, even for coupled cases, a bias at the surface may not manifest in the same way at the cloud level because of temperature effects. I would be more comfortable with this statement if it were based on specific humidity and with the qualifier about coupling state.
Line 128-129: The point of the paper is to understand the model representation of clouds and their impacts on surface radiation. Thus, it seems important to examine higher resolution data since clouds vary on sub-daily timescales. Daily averages could mask competing / compensating biases. Why not use the higher resolution data in the analysis?
Line 146-148: I do not understand this feedback. If the near surface air temperature is too cold, with all else unchanged (i.e., the overlying atmospheric temperature and the amount of turbulent mixing), the near-surface temperature structure would become more stable, which would diminish upward sensible heat flux (or increase downward sensible heat flux). This then effectively works against the cold bias rather than reinforcing it. Perhaps I’m missing something to this argument, so that should be spelled out more clearly.
Line 172-173: Something is wrong with this sentence. It is only true if “exceeding” means a greater negative value than -25 W/m2, which is not really the correct use of exceeding. Additionally, the clear sky mode (values that are more negative than -25) is not discussed until the next sentence. Thus, maybe here you mean “underestimated”?
Line 174-175: “better representation of the cloudy mode frequency”. Strictly speaking the representation of that mode is better than for P2. However, it is still not good and I do not think the distributions “agree well.”
Line 181: I think this appendix B and Table B1 should be in the main text. The cloud occurrence information is important for the main paper and the reader should not have to search for it at the back of the paper.
Line 193-194: I do not agree with this statement. There are quite different clouds identified using the different approaches and it is not clear how these two datasets do overlap or should overlap with each other. For example, the fact that the observations and model show “quite similar cloud occurrence frequencies” during P3 could be for vastly different reasons due to the different methods for identifying clouds. Even the “synchronous cloudy set” discussed in the following sentence is not very satisfying because clouds that exist in that set could be there for different reasons. Simply use the same method for identifying clouds to eliminate that inconsistency and uncertainty.
Line 198: is this section presenting only results for cloudy periods identified in the previous section? That is not clearly stated in the text, but the caption of the figure seems to indicate so.
Line 199: “retrieved” is the wrong word here. Maybe “simulated”
Line 200: This last part of the sentence is unclear. I believe you mean that the emission of these warm and opaque clouds is underestimated. Correct?
Line 201: Instead of “magnitude” it is better to say “fractional occurrence”
Line 230: Ec is used in the equations while Ecloud is used in the text. Make consistent.
Table 2: The concept of the information presented in this table is very compelling. However, the vast uncertainties in the method, and lack of discussion of how those uncertainties impact the results, leaves me with a notion that these values cannot be trusted. I would need discussion of the role played by interdependence of Ec on other parameters, on the removal of all cases when derived Ec >1, on the identification of cloud base height, on the identification of cloudiness in general, etc. etc.
Table 3 and associated discussion: These are the statistics that go with the distributions discussed in Section 3.2.2 and Figure 4. Why not include this information together in the description of results? Certainly the distributions provide useful information for interpreting the biases and statistics given here.
Line 275: What is liquid water fraction? This has not been defined as far as I recall. Also, just because the LWF ~ 1 does not mean that model phase partitioning is not an issue. Phase partitioning impacts the properties of the liquid and its radiative effects in non-linear ways.
Line 299-300: One important bias that is not stated directly here but is important to highlight is the fact that the observations are at a specific location on solid ice (representing ~10 m^2) while the model grid size is 15x15 km^2. The observed surface will be on the high end of the albedo distribution within a model grid cell, which also includes non-level ice, leads, etc. There has been some work to evaluate such spatial differences when, for example, comparing surface point sources of albedo to albedo derived from satellites with a large spatial footprint.
Section 3.3.4: In this section we have learned that the model can represent the basic first and second indirect effects of aerosols. However, the argument is made that the baseline CDNC is correct. Thus…. Is the conclusion that errors in CDNC are NOT the cause of radiation errors?
Line 359-360: Does this mean that you do not also use the ISWC threshold that was described in Section 3.3.1?
Line 363: How is ABL height defined here?
Line 384-386: This has not been shown. However, it could be shown by making direct comparisons of the model results to the measured radiosonde profiles. Rather than speculate using an indirect approach with uncertain assumptions, it is better to just use the direct measurements that are available.
Line 402: By “clouds are too often underrepresented” do you mean that the model does not produce clouds as often as they are observed? Please clarify.
Line 409-410: Yes, but phase partitioning can impact the properties of the liquid component, which dominates the radiative signal.
Section 4: There is a marked increase in typographical and grammar errors in this final section. I re-iterate the need for a full editorial review of the manuscript.
Citation: https://doi.org/10.5194/egusphere-2026-2516-RC2 -
AC2: 'Reply on RC2', Yaël Le Gars, 28 Sep 2026
We thank the reviewer for a very thorough and technically rigorous review. The reviewer's comments have identified important weaknesses in the presentation, structure, and the underlying methodology of the cloud-emissivity/dominance-analysis approach. The reviewer's major methodological concerns, in particular regarding the consistency of the cloud-identification criteria between observations and model, and the assumptions underlying the two-layer emission model and dominance analysis, are well taken. We take these comments very seriously and outline below, point by point, how we have addressed each one in the revised manuscript, describing a substantially revised methodology that we believe directly resolves these concerns. In particular, we have updated our cloud diagnostic methodology, we extended our evaluation framework, we incorporated new sensitivity experiments (WRF-neXtSIM, WRF-Li, WRF-50CDNC, WRF-10CDNC). The dominance-analysis section has been moved to the Supplement, reducing its role in the main text to a short, clearly caveated paragraph presenting it as supporting evidence only. We also separated more clearly in the manuscript the presentation of the methodology, model evaluation, and model sensitivity experiments and new developments. These changes aim to clarify the objectives of our study to the reader: first, to evaluate the WRF model, a regional model widely used to simulate the Arctic climate, second to identify the likely physical origin of these model biases, and third, to explore the sensitivity and correct the representation of these physical processes, or identify additional development perspectives for future work. For this reason, we also suggest a change of title to "Identifying and correcting causes of Arctic cloud radiative biases over sea ice in the WRF model". Reviewer comments are reproduced in italics; our responses follow in plain text.
Response to Reviewer 2
We thank the reviewer for a very thorough and technically rigorous review. The reviewer's major methodological concerns, in particular regarding the consistency of the cloud-identification criteria between observations and model, and the assumptions underlying the two-layer emission model and dominance analysis, are well taken. We describe below a substantially revised methodology that we believe directly resolves these concerns, together with our response to each general comment.
Major Methodological IssuesReviewer comment: Line 162: Why does the threshold change so much between the different periods of observations? During MOSAiC, Shupe et al. (2026) applied this same LWN threshold approach but found that -25 W/m2 worked for all seasons. Why the difference here? Is there a physical explanation or something markedly different about the atmospheric scenes in this part of the Arctic during winter? Also, these different definitions will make the different periods incomparable with each other and hard to interpret. I believe this difference manifests itself into some of the results presented later in the analysis, but its role is not discussed.
Author response: As demonstrated by the observed LW_net PDFs during N-ICE2015 (Fig. 3), atmospheric conditions in the Atlantic sector north of Svalbard present a distinct winter regime. During P1 (polar night), clear-sky surface cooling is weaker (clear-sky mode peaking at -44 W/m2 compared to P2 and P3 (clear-sky mode peaking at -72 W/m2. Applying a fixed -25 W/m2 threshold during P1 would place the cutoff in the middle of the inter-modal transition zone rather than between the distinct clear and opaque-cloud modes. We explicitly clarify this physical rationale in Sect. 2.4.
Reviewer comment: Lines 176-181: As described in this paragraph, it is apparent that the cloud conditions are identified using LWN in the observations but CLT in the model. The reasoning presented for why this difference is implemented is the role of spatial variability in cloudiness that could impact the radiation measurements. That is a valid concern, however, the observational dataset does not have information on spatial distribution of cloudiness. That is an unavoidable limitation of these observations. But given the observations, it makes more sense to at least apply consistent definitions. Why not just apply the same LWN threshold approach to both observations and models so the different definition does not add further, unquantified, uncertainty to the comparison? The comparisons that are shown later clearly reveal some influence from this difference in definition, but that influence is not discussed nor quantified. Moreover, it really doesn’t make sense. “Clouds” have many definitions. Thus, to have any hope of legitimately assessing a model, it is important to at least define them in the same way between observations and models.
Author response: We agree with the reviewer that using different cloud definitions creates ambiguity. In the revised manuscript (Sect. 2.4) we restructured our diagnostic framework into two rigorous distinct comparisons :
- To assess whether the model produces radiatively opaque clouds at the correct frequency and timing, we apply the exact same LW_net threshold to both observed and simulated net longwave fluxes. Table 3 now reports cloud occurrence statistics using the LW_net threshold (the same criterion applied to the observations) which allow a direct, like-for-like comparison of cloud occurrence between observations and the model under a single and common criterion; something the previous comparison could not offer, since it relied on two different criteria. Also, a new figure (Fig. 7) addresses how these occurrences vary with the sensitivity experiments.
- To evaluate radiative biases when a cloud is present, filtering simulated output with LW_net > -25 W/m2 would artificially discard model clouds that are physically present but optically too thin (a core failure mode). Thus, we define a synchronous cloudy set comprising hours with an observed
opaque cloud (LW_net obs criterion) and simulated cloud cover in the column (CLT > 0.95), an operational definition independent of simulated surface fluxes (Sect. 2.4).
Reviewer comment: Line 201-205: It is a bit strange that the model shows a lot of low LWD values that are not observed. I believe this figure has only the cloudy sky subset of data described in the previous section (although that is not entirely clear). If so, these low simulated LWD values would have to come from clouds that are vastly thinner and/or colder than those that were observed. This signature would be consistent with a discrepancy in the definition of clouds. Since the observations are simply based on LWN threshold, they will all have a high LWD. Since the model results are based on this CLT, there is no guarantee they will have high LWD. In fact, based on later parts of this manuscript, it appears that “clouds” are also defined where there is ice/snow via the modeled ISWC. Thus, if I had to speculate, I would say that these low LWD values are from times when there is only ice/snow in the model with little to no liquid water. This is a demonstration of the apples-and-oranges comparison that results from the different cloud definitions.
Author response: The reviewer’s insight is correct. Under the synchronous CLT > 0.95 criterion, model cases containing minimal liquid water (with either predominantly ice/snow condensate or with low total condensed water) yield low LWP and low LW. This explicitly confirms our diagnosis in Sect. 3.3 and 3.4: spring (P2) radiative errors in WRF-CTRL stem from simulated clouds that are present in the column but optically too thin due to insufficient liquid water content, rather than complete cloud absence.
Reviewer comment: Line 214-238. There are a lot of problems with this general approach to assessing the clouds and radiation, as outlined in the following six points:
1) Eq 1 includes some of the terms that contribute to LWD and is mostly correct when Ec approaches 1. When Ec is significantly less than 1 this means there will be contributions from the atmosphere above the cloud (possibly from multi-layer clouds that are at a different temperature), which are not included here. This point is kind of mentioned at the end of the paragraph but is not included in the equation. Overall, this equation is approximately true only when Ec is high.
2) Having "almost all" cloud base heights less than 90m would be a remarkable result for observations over nearly 5 months extending from winter to early summer. It is divergent from other observations, including those from MOSAiC. I do not believe this statement is true. I have tried to assess the results of the Maillard et al. (2021) paper, but there is no actual lidar data shown with which to independently verify that the lidar really identified almost every cloud base to be below 90m. Thus, without further information, my interpretation of this result is that the “cloud base” referred to here is actually the base of “hydrometeors”, including falling ice. In that case, yes, the base of hydrometeors would most commonly be near or at the surface because there is quite often falling precipitation. However, this base is not always (or not often) the base that is important for radiation. I believe N-ICE had a ceilometer; why not use actual cloud base height measurements from a widely trusted system like that? Ceilometer reported cloud base would be more consistent with the radiatively-important cloud base.
3) Building further on cloud base: In the description of the model analysis, it is stated that cloud base is defined when LWC exceeds 10^-5 kg/kg or when ice/snow exceeds 10^-6 kg/kg. As outlined in the Introduction, ice and snow have vastly different interactions with atmospheric radiation (for a given amount of mass) due to microphysical properties. By defining things in this way, you will get artificially low cloud base heights under all-ice and mixed-phase conditions that will be identified based on weak ice/snow precipitation, which has a low total emissivity. Thus, it is not a good approximation of the location where most emission occurs. i.e., the assumed cloud emission temperature will be wrong. This incorrect identification of cloud base could be part of the reason you identify most modeled cloud bases to occur in the first model level, but I do not have enough information to know for sure.
4) On a related topic: The assumption that cloud base temperature and surface temperature are equal is highly uncertain, condition dependent, and effectively forces the implied LWN to be 0 W/m2 under thick clouds. In reality, there is usually a negative temperature lapse rate over some layer below cloud base (i.e., usually adiabatic ascent helps to form the cloud), which could overlie a variety of temperature structures at lower levels. In the best-case scenario (for the assumption made here) with a low, surface-coupled cloud the temperature difference might only be a couple tenths of a degree. However, for real cloud bases that will often be significantly >100 m above the surface, the difference becomes important. We already know that Eo is small and in the limit where Eo approaches 0, Eq 2 becomes simply LWD=Ec*sigma*T2m^4! This is a poor approximation.
5) The analysis uses Eq 1 for model data and Eq 2 for observational data. How then can the results be compared to each other? The method alone will have a large uncertainty that could mask other issues in the model. To truly evaluate the model, you must apply the same approaches to both observations and model results.
6) Simply throwing out cases with Ec > 1 is not a sufficient remedy for the shortcomings of this approach. In general, I would expect the smallest uncertainties to occur when Ec is highest (i.e, that is when Eq 1 is most accurate). Ec values > 1 do indeed indicate a bias, but it could be a smaller bias than cases with lower Ec values (which are retained in the analysis). The main point here is that the whole analysis is highly uncertain because the equations and assumptions used are suspect. The impact of these uncertainties has not been evaluated or discussed thoroughly. As a result, I have little confidence in all the downstream results that are presented.
Author response: We thank the reviewer for this detailed and constructive critique We have revised and contextualized this analysis in Sect. 3.4 and Supplement Sect. S1. We understand that this simplified 2-layer approach has inherent limitations. As a result, we also de-emphasized the discussion of these results in the manuscript, since our main findings in Section 3, that motivate later changes, are that errors in LWP, boundary layer structure, and surface processes seem to be driving model errors - the 2-layer model discussion being additional but not load-bearing evidence for the LWP issue, but that could interestingly also point to issues in the simulation of cloud effective radius contributing to motivating our exploration of the sensitivity to CDNC in Section 4.
(1) We explicitly state that Eq. 1 is a simplified two-layer model that becomes less accurate as cloud emissivity drops below unity or under complex multi-layered structures. We present this two-layer model as a qualitative variance-attribution too rather than an exact physical inversion. This is now explicitly said in Sect. 3.4.
(2) We clarified in Sect. 2.1 and 3.4 that the IAOOS lidar measurements during N-ICE2015 (Maillard et al., 2021) demonstrated that cloud bases were overwhelmingly below 90 m with a weak temperature gradient (0.6-2°C in the lowest 100 m), supporting T2m as a reasonable proxy for low-level Arctic stratus. To address potential errors in cloud-base height and sub-cloud temperature inversion structure, we explicitly note throughout Sect. S1 (and flag it in the main text; Sect 3.4) that any discrepancy in cloud altitude or temperature profile in the model is subsumed into the diagnosed emissivity term, and we caveat our conclusions accordingly.
Regarding the reviewer's suggestion to use ceilometer/lidar-derived cloud-base height: N-ICE2015 indeed deployed an ARM MicroPulse Lidar (MPL) aboard the R/V Lance. We contacted the data owner in the first steps of this study, requesting access to this dataset, but never received a response. We further note that the only publicly orderable product (level B1) contains raw, uncalibrated backscatter signal counts. Deriving products such as cloud fraction, cloud-base height, hydrometeor phase would require additional retrieval processing beyond the scope of this study. We were therefore unable to independently verify our cloud-base height assumption against this ceilometer/lidar dataset.
(3) We thank the reviewer for pointing out the potential ambiguity of using distinct hydrometeor mass mixing ratio thresholds (10-5 kg/kg for LWC and 10-6 kg/kg for ISWC). To ensure full methodological consistency across the manuscript and eliminate this issue, we have updated the cloud-base layer identification in the supplementary two-layer dominance analysis (Supplement Sect. S1) by adopting a unified relative humidity threshold (RH > 95% with respect to liquid water), following the Reviewer 2's recommendation.
(4) We explicitly acknowledge in Sect. 3.4 that assuming T_cb = T2m introduces uncertainty when cloud bases lie above the shallow boundary layer or within inversions. However, during N-ICE2015, radiosonde profiles confirm that the temperature gradient in the lowest 100 m was weak (0.6-2°C in 90% of cases). We state clearly that this two-layer formulation serves as a qualitative diagnostic tool and rely primarily on direct LWP vs. flux error relationships (Fig. 5).
(5) We agree that using different forms of the two-layer equation for model and observations (because Tcb is directly diagnosed in the model but not observed) introduces an asymmetry that limits direct comparability of the two emissivities. We do not believe this can be fully resolved with the available observations, since atmospheric profiles are only available twice daily, too scarce to allow a cloud-base diagnosis at the same temporal resolution as the hourly model evaluation.
(6) We have thoroughly expanded the discussion of these methodological uncertainties in Sect. 3.4 and Supplement Sect. S1. Instances of diagnosed epsilon_c > 1 reflect cases where surface T2m underestimates cloud radiating temperature during strong surface inversions. We clarify that the two-layer dominance analysis is included only as supporting evidence for variance attribution, while all primary conclusions in the paper rest independently on direct flux error distributions, LWP sensitivity curves (Fig. 5, Fig. 8), and physical sensitivity experiments (WRF-neXtSIM, WRF-Li, CDNC).
We believe that, taken together, these revisions directly address the reviewer's concerns and substantially increase confidence in this part of the analysis; we, in addition, temper the strength of the conclusions drawn from Table 2 throughout the manuscript to reflect the remaining uncertainty.
Reviewer comment: Line 239-242: I do not understand what information is being shown by this analysis. First, while the analysis examines the errors (LWD model-obs), it does a dominance analysis using Eo, T2m, Tcb, and Ec. I assume those parameters are from the model? Second, Ec was derived from T2m, Tcb, and LWD for the model results (Eq 1). Thus, the respective Ec values are directly dependent on the other parameters. Their contributions to LWD variability are pre-determined by the relationship through which they were derived, not based on independent physical variations. Thus, it is not clear to me how this analysis reveals information about the physics of Ec. Lastly, the assumption that T2m=Tcb (in the observational analysis) is undercut by the different results for T2m and Tcb. These temperatures apparently do not vary the same nor impart the same influence on LWD. If true, the assumption that T2m and Tcb equal each other in Eq 2 neglects some important physics.
Author response: We clarify, in Sect. 3.4 and Supplement Sect. S1, that c is an inverted diagnostic parameter rather than an independent prognostic variable. The dominance analysis applies multiple linear regression variance decomposition to attribute error variance among terms without assuming independence. We explicitly state that c absorbs residual processes (including cloud height and microphysical errors), which is why we connect this diagnostic directly to independently simulated LWP distributions in Sect. 3.4.
Reviewer comment: Line 272: This point might indeed explain the results in Figure 5 for this dataset. However, based on LWP observations from MOSAiC it was found that the opaque cloudy state (i.e., LWN > -25W/m2) contained a lot of LWP values that were less than 30 g/m2 (see Dahlke et al. 2025, Shupe et al. 2026). Moreover, the fact that a LWN threshold of -10 W/m2 is used in P1 instead of -25 W/m2 means that the “clouds” identified in that period would naturally have higher LWP than the other periods (in a season that routinely has lower LWP and has lower background moisture such that the radiative effects of thin clouds are accentuated relative to the moist summer). First, I do not think there should be different LWN thresholds applied to the different periods. But if there are, the impact of this difference should certainly be discussed thoroughly.
Author response: We explicitly discuss this in Sect. 2.4 and 3.4. The -10 W/m2 threshold in P1 selects optically thick clouds during winter storm events, which naturally exhibit higher LWP (see the response above on the sensitivity to Winter results to the threshold value used).
Based on Shupe et al. (2026), we acknowledge that a portion of the MOSAiC opaque state (LW_net > −25 W m⁻²) does correspond to LWP values in the 5–30 g m⁻² range (this range, however, falls largely within the ~20 g m⁻² uncertainty of derived LWP), compensated by a substantial ice water content that keeps the cloud radiatively opaque. We have revised the text in Sect 3.4 to make this connection explicit: we now directly follow this MOSAiC comparison with our own test of whether ice content (ISWP) explains the radiative errors in our WRF-N-ICE2015 dataset, and find no such correlation (Fig. S2), indicating that this ice-compensation mechanism does not appear to be what drives the errors we diagnose here. Deficiencies in liquid water content remain the better explanation, further supported by the liquid water fraction (LWF) close to 1 found for the largest-error, lowest-LWP cases (Fig. S3)
Reviewer comment: Line 380-383: I have concerns about the accuracy of the “observed” coupling state results as they are based on a fixed cloud base height of 100m that may or may not be correct for any given case. The effect of this assumption on the results has not been discussed. Moreover, with a radiosounding dataset (and only 55 soundings) it should be readily possible to perform a robust determination of coupling state for the available cases. In soundings it is easy to identify a liquid saturated layer (typically a threshold of ~96% RH is used). Then, you can determine the equivalent potential temperature at the base of that cloud layer and the surface. Their difference gives you a clear delineation of the coupling state. This approach is far preferable to making an uncertain assumption about cloud base height.
Author response: We updated our boundary-layer coupling analysis in Sect. 2.5 and Sect. 3.6 following the reviewer's recommendation and Gierens et al. (2020). For all radiosondes, we identify cloud bases using a 95% relative humidity threshold with respect to liquid water and evaluate virtual potential temperature profiles. Clouds are then classified as surface-decoupled or surface-coupled using the method from Gierens et al. (2020). Comparisons between the soundings to co-located model profiles are added to Sect. 3.6 (WRF-CTRL) and Sect 4.1 (WRF-neXtSIM). This same threshold is now used also to characterise the surface-coupling state in all points from the modelled RH.
General CommentsReviewer comment: Full text: There are many issues with the written text in terms of grammar, sentence structure, and other issues. While I started noting these in my review, there are so many that I will not include them here. Before proceeding, this manuscript needs a full editorial review.
Author response: We have arranged a full English-language and editorial review of the entire manuscript prior to resubmission.
Reviewer comment: Line 21: “positive effect” on what? They have a net warming effect on the surface or a positive effect on the surface radiative balance. That should be clarified.
Author response: We clarified this explicitly as “a net warming effect on the surface”.
Reviewer comment: Line 22-23: This statement about “short period of surface cooling in summer” is true over highly reflective surfaces like sea ice. Over less reflective surfaces like land (after snow melt) or open ocean, the cooling period extends for multiple months. This cannot be considered “short.”
Author response: We qualified this statement in the Introduction to specify it refers to high-albedo, snow/ice-covered Arctic Ocean surfaces, consistent with the scope of this sea-ice-focused study.
Reviewer comment: Line 24: Cloud radiative parameters are important of course. But this statement is more accurate as “Cloud radiative properties along with relevant properties of the surface and atmosphere……” Ultimately the surface albedo plays a substantial role in the net effect of clouds seasonally that can overwhelm cloud radiative properties.
Author response: We revised the sentence to read “The magnitude of the CRE is driven by cloud optical properties, together with the properties of the surface and of the overlying atmosphere” as suggested.
Reviewer comment: Line 81: Having only 1-moment liquid leaves substantial deficiencies in representing the liquid phase, as has been shown in prior studies. This is particularly true when you are attempting to examine the effect of, for example, aerosol indirect effects where the second piece of information about the liquid droplet size distribution can be quite important.
Author response: We agree and have added an explicit caveat in Sect. 2.3, noting that single-moment liquid microphysics with prescribed CDNC limits the physical representation of droplet size distribution adjustments.
Reviewer comment: Line 116: The model had a rather large warm bias during 14-16 March. What is the cause of this and does it provide any insight into the model issues?
Author response: The reviewer is right. However, we have not investigated the cause of this warm bias. We do not want to speculate on a mechanism (warm-air intrusion, FNL artefact, or otherwise).
Reviewer comment: Line 118-120: Importantly, this good agreement during P3 should not be surprising since the surface is melting at the time.
Author response: We agree and have added this caveat in Sect. 3.1 : agreement in P3 is largely constrained by the melting-point limit of snow and ice, which caps both observed and simulated surface temperatures near 0°C.
Reviewer comment: Line 121-126: Since RH is influenced by both the amount of moisture and the temperature, it is better to examine specific humidity to allow for the moisture and temperature to be evaluated separately. It is not clear here how much the temperature biases impact the RH compared to specific humidity biases. Line 125-126: True, the near-surface RH biases might impact the representation of low clouds under circumstances when the cloud layer is coupled to the surface. However, under decoupled conditions this is likely not the case. Additionally, even for coupled cases, a bias at the surface may not manifest in the same way at the cloud level because of temperature effects. I would be more comfortable with this statement if it were based on specific humidity and with the qualifier about coupling state.
Author response: RH does conflate temperature and moisture effects, in principle. However, under the near-saturation conditions that dominate this dataset, specific humidity is itself tightly constrained by temperature, so q2m and T2m are largely redundant rather than independent in this case. We provide, attached to our responses, the temporal series of observed and simulated q2m for the reviewer’s own judgement (see Qvapor.pdf, attached).
Reviewer comment: Line 128-129: The point of the paper is to understand the model representation of clouds and their impacts on surface radiation. Thus, it seems important to examine higher resolution data since clouds vary on sub-daily timescales. Daily averages could mask competing / compensating biases. Why not use the higher resolution data in the analysis?
Author response: The SW↓ and LW↓ fluxes are shown as 24-hours averaged in Fig. 2c,d because their strong diurnal cycle makes the hourly time series difficult to read at the scale of the full campaign. The daily averages remain informative regarding the seasonal trends discussed in the text. All statistics reported in Table 2 and 5 are nonetheless computed from hourly values, consistent with T2m and RH2m.
Reviewer comment: Line 146-148: I do not understand this feedback. If the near surface air temperature is too cold, with all else unchanged (i.e., the overlying atmospheric temperature and the amount of turbulent mixing), the near-surface temperature structure would become more stable, which would diminish upward sensible heat flux (or increase downward sensible heat flux). This then effectively works against the cold bias rather than reinforcing it. Perhaps I’m missing something to this argument, so that should be spelled out more clearly.
Author response: It has been clarified in Sect. 3.1: near-surface cold biases reduced turbulent heat transport from the surface, reinforcing decoupling and surface-base inversions (Hong and Jiang, 2024).
Reviewer comment: Line 172-173: Something is wrong with this sentence. It is only true if “exceeding” means a greater negative value than -25 W/m2, which is not really the correct use of exceeding. Additionally, the clear sky mode (values that are more negative than -25) is not discussed until the next sentence. Thus, maybe here you mean “underestimated”?
Author response: “Exceeding” was referring to greater positive value than -25 W/m2. We understand that this was not clear and this has been rephrased in Sect. 3.2.
Reviewer comment: Line 174-175: “better representation of the cloudy mode frequency”. Strictly speaking the representation of that mode is better than for P2. However, it is still not good and I do not think the distributions “agree well.”
Author response: Following Reviewer 1’s comment on the choice of the periods. We changed the ending date of P2 from 25 May to 1 June. This alters the LW_net PDF for P2 and P3, as well as all the following figures and tables that use the seasonal delimitation, which have been updated accordingly. We can now state that the model reproduces the observed distribution well during P3.
Reviewer comment: Line 181: I think this appendix B and Table B1 should be in the main text. The cloud occurrence information is important for the main paper and the reader should not have to search for it at the back of the paper.
Author response: This is also raised by Reviewer 1 and we agree. We moved the former Table B1 to the main text as Table 3 (Sect. 3.2), and now discuss it thoroughly .
Reviewer comment: Line 193-194: I do not agree with this statement. There are quite different clouds identified using the different approaches and it is not clear how these two datasets do overlap or should overlap with each other. For example, the fact that the observations and model show “quite similar cloud occurrence frequencies” during P3 could be for vastly different reasons due to the different methods for identifying clouds. Even the “synchronous cloudy set” discussed in the following sentence is not very satisfying because clouds that exist in that set could be there for different reasons. Simply use the same method for identifying clouds to eliminate that inconsistency and uncertainty.
Author response: We agree with the reviewer that using different cloud definitions creates ambiguity when assessing model performance. In the revised manuscript (Sect. 2.4), we have restructured our cloud evaluation into two distinct and complementary steps to eliminate this inconsistency:
- We apply the exact same LW_net threshold to both observed and simulated net longwave fluxes (Table 3 and Fig. 7), allowing a direct, like-for-like comparison of cloud occurrence between observations and the model under a single and common criterion. This evaluates whether the model simulates radiatively opaque clouds at the right time and frequency. Under this unified threshold, WRF-CTRL captures 12 % of opaque clouds in P1 (vs. 19 % observed), 36 % in P2 (vs. 69 % observed), and 90 % in P3 (vs. 87 % observed). The new WRF-neXtSIM experiment (Sect. 4.1, in replacement of the former Sect. 3.3.3 and the WRF-SEAICE experiment) increases P2 opaque cloud occurrence to 66%, along with better temporal overlap, demonstrating that surface boundary conditions drive a large share of this spring cloud deficit.
- To evaluate radiative biases when clouds are present without discarding model clouds that are physically simulated but optically too thin (a core failure mode), we retain the synchronous cloudy set (observed LW_net opaque cloud criterion and CLT > 0.95 in the model). We now explicitly state the operational rationale, overlap fractions, and limitations of this subset in Sect. 2.4 and Sect. 3.2.
Reviewer comment: Line 198: is this section presenting only results for cloudy periods identified in the previous section? That is not clearly stated in the text, but the caption of the figure seems to indicate so.
Author response: We have updated the opening text of Sect. 3.3 to state explicitly in the main text (and not only in figure captions) that the radiative bias distributions (SW↓ and LW↓), LWP relationships, and statistical metrics (NMB, MB, RMSE) are evaluated exclusively for the synchronous cloudy-sky subset defined in Sect. 2.4.
Reviewer comment: Line 199: “retrieved” is the wrong word here. Maybe “simulated”
Author response: We have corrected the wording in Sect. 3.3, replacing "retrieved" with "reproduced" (e.g., " the distribution shape of LW↓ is reproduced: two modes appear [...]").
Reviewer comment: Line 200: This last part of the sentence is unclear. I believe you mean that the emission of these warm and opaque clouds is underestimated. Correct?
Author response: Confirmed; we have rewritten this sentence in Sect. 3.3 to state clearly that “the main mode, corresponding to the emission of warm and opaque clouds occurring during winter storms, is underestimated by 7 W/m2”
Reviewer comment: Line 201: Instead of “magnitude” it is better to say “fractional occurrence”
Author response: We have adopted this correction in Sect. 3.3 and replaced "magnitude" with "fractional occurrence" when referring to cloud frequency distributions across seasonal periods.
Reviewer comment: Line 230: Ec is used in the equations while Ecloud is used in the text. Make consistent.
Author response: We have standardized the notation throughout this section now in the Supplement (S1) : c is now used consistently in both the equations and the text to designate cloud emissivity.
Reviewer comment: Table 2: The concept of the information presented in this table is very compelling. However, the vast uncertainties in the method, and lack of discussion of how those uncertainties impact the results, leaves me with a notion that these values cannot be trusted. I would need discussion of the role played by interdependence of Ec on other parameters, on the removal of all cases when derived Ec >1, on the identification of cloud base height, on the identification of cloudiness in general, etc. etc.
Author response: As stated above, this section has been moved to the Supplement (S1), and its role in the main text reduced to a short paragraph at the end of Sect. 3.4, given the limitations of this method and the non-essential role of this result, neither in the rest of this study nor our conclusions. We have added uncertainty discussions in Sect. 3.4 and Supplement Sect. S1. We emphasize that Table S1 provides qualitative support within a simplified framework, while primary conclusions rest on direct flux comparisons and physical sensitivity runs.
Reviewer comment: Table 3 and associated discussion: These are the statistics that go with the distributions discussed in Section 3.2.2 and Figure 4. Why not include this information together in the description of results? Certainly the distributions provide useful information for interpreting the biases and statistics given here.
Author response: Former Table 3 (now referred as Table 4) was discussed in Sect. 3.3.2 (which now corresponds to Sect. 3.4). This is moved to the section referred by the reviewer (now Sect 3.3).
Reviewer comment: Line 275: What is liquid water fraction? This has not been defined as far as I recall. Also, just because the LWF ~ 1 does not mean that model phase partitioning is not an issue. Phase partitioning impacts the properties of the liquid and its radiative effects in non-linear ways.
Author response: We added a formal definition of LWF in Sect. 3.4 as the ratio of cloud liquid water path to total condensed water path.
Regarding the reviewer's second point: we agree that a near-unity LWF does not, by itself, guarantee that phase partitioning is not an issue. However, when LWF is close to 1, there is very little ice mass to be mispartitioned in the first place, so a LWP bias under these conditions cannot primarily reflect an error in the liquid/ice mass split, but instead points to a deficit in the total condensed water mass produced by the model. We acknowledge that this does not rule out biases in the microphysical properties of the liquid itself (e.g. droplet size distribution), which can affect its radiative properties independently of total mass - this is a secondary effect that we do not attempt to quantify here. We have modified our statement accordingly.
Reviewer comment: Line 299-300: One important bias that is not stated directly here but is important to highlight is the fact that the observations are at a specific location on solid ice (representing ~10 m^2) while the model grid size is 15x15 km^2. The observed surface will be on the high end of the albedo distribution within a model grid cell, which also includes non-level ice, leads, etc. There has been some work to evaluate such spatial differences when, for example, comparing surface point sources of albedo to albedo derived from satellites with a large spatial footprint.
Author response: It has been explicitly discussed in Sect. 3.5 : point-based radiometer measurements on floes reflect local high albedo, whereas model grid cells aggregate heterogeneous surfaces (leads, melt ponds, marginal ice), contributing to albedo discrepancies, especially when coarse sea ice products (FNL) are used. This is supported by the WRF-neXtSIM experiment (Sect. 4.1), where using a higher-quality sea ice concentration product reduces this discrepancy.
Reviewer comment: Section 3.3.4: In this section we have learned that the model can represent the basic first and second indirect effects of aerosols. However, the argument is made that the baseline CDNC is correct. Thus…. Is the conclusion that errors in CDNC are NOT the cause of radiation errors?
Author response: It has been clarified in Sect. 4.3 : while reducing prescribed CDNC alters cloud optical thickness via Twomey/Albrecht effects, realistic CDNC reductions (from 100 to 50 cm-3) do not remove high-LWP SW↓ deficits but rather redistribute errors, and increasing CDNC to 200 cm-3 does not correct the low-LWP radiative errors either, while degrading the high-LWP SW↓ bias. Only an unrealistically low value (10 cm-3), well below the observationally supported range, markedly reduces the frequency of high-LWP cases, partly through enhanced precipitation efficiency and shorter cloud lifetime. This confirmed that CDNC misrepresentation is not the primary driver of the model radiative biases in our study.
Reviewer comment: Line 359-360: Does this mean that you do not also use the ISWC threshold that was described in Section 3.3.1?
Author response: We have modified our criterion to use the level at which RH with respect to liquid water exceeds 95%, which is used to characterise the surface coupling state in Sect. 3.6, as suggested by the Reviewer.
Reviewer comment: Line 363: how is ABL height defined here?
Author response: It is defined in Sect. 2.5 as the boundary layer height diagnosed prognostically by the MYNN2 PBL scheme.
Reviewer comment: Line 384-386: This has not been shown. However, it could be shown by making direct comparisons of the model results to the measured radiosonde profiles. Rather than speculate using an indirect approach with uncertain assumptions, it is better to just use the direct measurements that are available.
Author response: It has been replaced in Sect. 3.6 and Sect. 4.1 by direct comparisons of model profiles to radiosondes. We adapted our conclusion regarding this new analysis
Reviewer comment: Line 402: By “clouds are too often underrepresented” do you mean that the model does not produce clouds as often as they are observed? Please clarify.
Author response: We clarified this explicitly in the Sect. 5: ”During P2, opaque clouds are markedly underrepresented".
Reviewer comment: Line 409-410: Yes, but phase partitioning can impact the properties of the liquid component, which dominates the radiative signal.
Author response: It has been added in Sect. 5 : we noted that phase partitioning indirectly affects liquid cloud optical depth through ice-growth water consumption.
Reviewer comment: Section 4: There is a marked increase in typographical and grammar errors in this final section. I re-iterate the need for a full editorial review of the manuscript.
Author response: Acknowledged; The entire text underwent a complete language polish.
-
AC2: 'Reply on RC2', Yaël Le Gars, 28 Sep 2026
Viewed
| HTML | XML | Total | Supplement | BibTeX | EndNote | |
|---|---|---|---|---|---|---|
| 190 | 57 | 33 | 280 | 32 | 29 | 25 |
- HTML: 190
- PDF: 57
- XML: 33
- Total: 280
- Supplement: 32
- BibTeX: 29
- EndNote: 25
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
Please see my full referee comment in the file attached.