the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
TOPAZ5: A high-resolution ocean and sea-ice model for the Arctic and North Atlantic
Abstract. Accurate simulation of coupled ocean-sea ice dynamics is important for understanding Arctic climate variability. This paper evaluates TOPAZ5, a novel high-resolution (6–10 km) hindcast configuration of the HYCOM-CICE model optimized for the Arctic and North Atlantic, over the period 2009–2019. TOPAZ5 is validated against satellite altimetry, in situ observations, and reanalysis products, assessing sea surface temperature (SST), sea surface salinity (SSS), sea level anomaly (SLA), eddy kinetic energy (EKE), mixed layer depth (MLD), volume transports, and sea ice properties. TOPAZ5 captures the large-scale spatial patterns of SST and SSS with strong agreement relative to the OSTIA dataset and WOA climatology; however, persistent warm and saline surface biases remain, particularly during summer and over Arctic shelf regions. SLA trends are generally consistent with satellite altimetry and reanalyses, though positive anomalies in the Norwegian and Barents Seas are underestimated. Compared to observations, the model reproduces similar spatial patterns of EKE but underestimates their amplitude, especially in the Lofoten Basin and Labrador Sea. MLD variability shows reasonable agreement with observed seasonal stratification, though the model tends to produce overly deep winter mixing in the Labrador Sea. Sea ice seasonality is well represented, but the model overestimates the central Arctic thickness and underrepresents spatial and temporal variability in marginal ice zones. While volume transports through the Bering Strait are realistically simulated, the flows through the Fram Strait and Barents Sea Opening are underestimated. Overall, TOPAZ5 demonstrates sufficient skill in simulating large-scale hydrography and sea ice dynamics to support operational forecasting applications in the North Atlantic and Arctic.
- Preprint
(34225 KB) - Metadata XML
- BibTeX
- EndNote
Status: final response (author comments only)
- RC1: 'Comment on egusphere-2026-1520', Anonymous Referee #1, 18 Jun 2026
-
RC2: 'Comment on egusphere-2026-1520', Anonymous Referee #2, 10 Aug 2026
The authors present TOPAZ5, a new high-resolution (6-10 km) configuration of the HYCOM-CICE model covering the Arctic Ocean, Nordic Seas, and North Atlantic. The forced ocean-sea ice simulation is validated against surface, hydrographic, transport, and sea-ice observations and reanalysis products for the period 2009-2019.
The paper could make an absolutely valuable contribution to the community by establishing a state-of-the-art hybrid-coordinate configuration for Arctic ocean-ice modelling and forecasting.
However, in its current form, the manuscript cannot be recommended for publication without major revision.My main comments are:
- The model domain is focused on the Arctic and North Atlantic, however there is no validation whatsoever for the Arctic Ocean hydrography and the T/S sections used for validation are all in the Nordic Seas outside of the Arctic Ocean. While in-situ data is indeed sparse in the Arctic Ocean, other studies have provided extensive model validation for the different basins in this region using for example PHC3 or WOA18 climatology (Khosravi et al., 2022; Shu et al., 2023; Wang et al., 2024). A similar analysis is the minimum requirement for presenting a new model for this specific region. There are also existing CTD, ITP and mooring observations for the Arctic Ocean, which can be used for validation.
- The model validation in this paper is done in several cases against observational data that is also assimilated into the reanalysis products used at the lateral boundaries (GLORYS12V1) or the atmospheric forcing (ERA5). While some circularity in this kind of setup might not be completely avoidable, it should be minimized where possible, and any remaining circularity must be explicitly stated and discussed. Specific instances of this include:
- SST vs. OSTIA: ERA5's prescribed SST boundary condition is the operational OSTIA product from September 2007 onward (i.e. for the entirety of the 2009-2019 validation period, Hersbach et al., 2018). The model is therefore in part being validated against the same SST fields used to force it.
- SLA and EKE vs. GLORYS: GLORYS12V1 assimilates along-track sea-level anomalies from satellite altimetry, making it a dynamically interpolated product of the same observations used for the "independent" altimetry comparison.
Since GLORYS also supplies TOPAZ5's lateral boundary conditions, validating TOPAZ5's SLA/EKE performance against GLORYS mixes a real model comparison with a comparison against the boundary condition's own data source. - Section 2 states that “surface relaxation of sea surface salinity (SSS) towards the corresponding monthly climatology was applied”. It is not clear from this wording whether the SSS surface relaxation is applied only during 1992 or throughout the entire 1992-2019 simulation (I assume it is the latter). This needs to be stated explicitly, as it directly affects the validity of the SSS results, since WOA 2018 is also used as the initial condition and as the sole observational reference for SSS validation. If the surface SSS relaxation is applied for the full simulation period, the SSS validation would be fully circular, since the model is then being validated against the same field it is restored toward.
- The bathymetry used in this setup is GEBCO_2014, but the authors state in the discussion that “Refining these transports will require higher-resolution grids and improved bathymetric representations.”. Why not make use of the up-to-date bathymetry data GEBCO_2026, which is substantially improved compared to the 12-year-old version, especially in the Arctic region.
Similarly, the authors acknowledge the limited data in the MLD climatology by de Boyer Montégut et al. (2004), but they do not use the newer Holte et al. (2017) climatology that also includes Argo data. - There are several instances when the paper seems to contradict its own findings. For example:
- abstract and conclusions describe "warm and saline" surface biases, but the biases shown in figures 1/2 and summarized in table 1 are cold and fresh.
- The differences in the seasonal cycle between TOPAZ5 and WOA 2018 SSS described as “smallest during winter and increase toward late summer and early autumn”, but figure 4 shows the smallest differences in June.
Detailed comments with direct reference to figures or lines in the text:
Figure 1: It would be helpful to see a map of horizontal resolution. This could possibly be added as figure 1b.
Line 86: "The target densities for the lower 40 hybrid z-layers were adjusted to reflect the water masses in the focus region." Please specify which water masses were targeted, and provide the actual target density values. It would also be useful to state whether this included Arctic water masses (e.g. Pacific Water, the Atlantic Water core, cold halocline density range) or was focused on the Nordic Seas / North Atlantic. How well are these water masses simulated in TOPAZ5?
Line 87: I mentioned the bathymetry dataset in my major comments. GEBCO_2026 is available, including the updated International Bathymetric Chart of the Arctic Ocean (IBCAO), improving the important Arctic Ocean bathymetry.
Line 100: WOA 2018 is used for initialization in 1992. Is there a specific reason for not using PHC3 or another Arctic Ocean-focused climatology? Please also specify which WOA 2018 product was used for initialization (e.g. a specific decadal average, or the multi-decadal mean). If it is the same 2005-2017 climatology used later for validation, this would introduce a possible warm-biased state compared to 1992.
Line 120: Please see my first major comment. Excluding the ice covered Arctic Ocean from validating an Arctic Ocean model is not acceptable.
Line 141: Standard deviation of SSH is used as a proxy for EKE, which is a valid approach. However, in several instances the metric is simply referred to as "EKE," which is not what is actually being computed. Please either compute true geostrophic EKE from SSH gradients, or full EKE from model velocities, or consistently label the metric as SSH variability/standard deviation throughout the text, figures, and abstract.
Figure 2: There are biases in the under ice SST when comparing TOPAZ5 to OSTIA, indicating that the calculation of Tfreeze in the model does not match Tfreeze in the observations.
Figure 4: The differences that are referred to in the text are very small compared to the full seasonal range and therefore hard to see in the plots. As mentioned in major comment 4, the description in the text (and the conclusions drawn from that) don’t match the results shown.
Figures 5/6 (and corresponding appendix figures): The x-axis is unlabeled - please specify what it represents (e.g. distance along section, longitude, or station number). Additionally, the axis range and tick values differ between the observation and model panels, making it unclear whether the two panels share a common horizontal reference frame; this should be corrected so that features can be directly compared across panels. It would also be helpful to add a third panel showing the model-observation difference, as in the earlier bias figures (e.g. Fig. 2).
Figure 7: As stated in major comment 2, GLORYS assimilates along-track satellite altimetry, so the "independent" altimetry and GLORYS comparisons are not independent of each other.
Line 311: “TOPAZ5 also underestimates positive trends in the Barents and Greenland Seas relative to both observations and GLORYS.” - This understates what Fig. 7a actually shows: in part of the Greenland Sea TOPAZ5 produces a negative SLA trend, whereas both altimetry (Fig. 7b) and GLORYS (Fig. 7c) show consistently positive trends there. This is a discrepancy in trend sign, not merely magnitude, and should be described and discussed as such.
Line 322: “differences in peak amplitude are evident, with altimetry recording higher values, suggesting limitations in the resolution of TOPAZ5” - Please clarify the mechanism being proposed here and why it is influenced by model resolution. Fig. 8 shows regionally averaged, monthly-mean SLA time series. This should largely remove any resolution-dependent (i.e. mesoscale/submesoscale) signals. It is not clear how horizontal grid resolution would produce a difference in regionally averaged, monthly-mean peak amplitude. Please justify this explanation or offer an alternative.
Figures 7/9: Why is the Arctic Ocean basin masked out entirely in these figures (likely due to unavailable/unreliable altimetry under sea ice), while the same region is shown with data in Figs. 2 and 3, despite being excluded from the corresponding statistics there? Please state the reason for masking in the captions, and consider using a consistent treatment (e.g. showing the ice mask boundary) across all spatial figures in the paper.
Since altimetry and reanalysis are not fully independent data sets, the authors could consider removing one of the two and instead showing the biases, like in figures 2 and 3.Line 331: "TOPAZ5 successfully captures the large-scale spatial distribution of mesoscale variability across the North Atlantic and Arctic Ocean." This is not well supported by Fig. 9. Beyond the systematic amplitude underestimation described in the following sentences, there are real differences in spatial pattern.
Line 333: “western North Atlantic” should be rephrased here. The "western North Atlantic" conventionally refers to the US/Canadian seaboard region, including the Gulf Stream separation zone, none of which is part of the model domain.
Line 348: The MLD climatology by de Boyer Montégut et al. (2004) is used for validation, but this data set predates the Argo era and has comparatively sparse coverage in the subpolar North Atlantic, Nordic Seas, and Arctic. The authors reference Holte et al. (2017), but it does not seem like the 2017 climatology is used. Additionally, while the constant density threshold of 0.03 kg/m3 is widely used, I recommend using the more recent, Argo-based Holte & Talley (2009) hybrid algorithm and/or the Holte et al. (2017) climatology as the observational reference, given both the more recent data coverage and its more robust handling of convective regions.
Figure 10: Same masking issue as for figures 7/9. Additionally, the colour scale in the winter panels (0-800 m) appears to be saturated in many regions and it is not possible to tell whether these values are close to 800 m or well in excess of it. What is the range of the data? Please choose a color scale that does not hide the actual range of the data.
Line 359: “>700m”, but the color scale in figure 10 shows that TOPAZ5 MLD at least exceeds 800m if not more. Again, what is the actual range of the data?
Line 364: The authors acknowledge that the climatology's "limited availability of hydrological profiles... particularly in the subpolar North Atlantic and Arctic-adjacent seas" affects its representativeness in exactly the regions this study focuses on. Given the authors have identified this limitation themselves, why not use a more recent, better-sampled climatology that addresses it (e.g. Holte et al. (2017) which is already cited elsewhere in the manuscript, see also the comment on line 348)?
Separately, note that this section states TOPAZ5 MLD exceeds 1000 m in the Labrador Sea, which confirms that the Fig. 10a colour scale (capped at 800 m) and the ">700 m" figure text (line 359) both understate the actual model values.Line 368: "spatial patterns align": this is not supported by Fig. 10. TOPAZ5 has the deepest winter convection in the Irminger Sea and a large patch of deep MLD in the Labrador Sea, whereas the climatology shows the deepest winter MLD in the Nordic Seas. Similarly, the summer patterns (including the locations of local minima and maxima) are not well reproduced.
Tables 2/3: These could be very easily combined. The last column in 2 is identical to the second column in 3. The section position is not very helpful here, and if included should also carry units (like °E).
Line 425: "Despite general agreement at the basin scale" is not well supported by Fig. 11 or Fig. 12. The thickness bias covers nearly the entire Arctic basin in positive and negative regions with substantial magnitudes. The September concentration bias likewise extends over a large fraction of the central basin. Additionally, while the text does mention both observational uncertainty and model bias as possible causes of the thickness overestimation, this mismatch is primarily attributed to observational uncertainty.
Line 436: “Comparisons between seasonal sea ice concentration (Fig. 12) further support the model’s ability to reproduce observed patterns, although with varying levels of accuracy.”: A large region corresponding to the Beaufort, Chukchi, and East Siberian Seas shows bias of around 1, i.e. close to the physically maximum possible bias. Given the magnitude, describing this as "varying levels of accuracy" considerably understates the problem.
Line 461: “TOPAZ5 effectively captures the pan-Arctic and North Atlantic hydrographic structure,”: this is a strong statement, given that no subsurface hydrographic validation was presented for the Arctic Basin itself (see major comment 1). Please clarify what this statement is based on, or qualify it accordingly.
Line 467: “Refining these transports will require higher-resolution grids and improved bathymetric representations.”: better bathymetry, especially in the study region of the paper, is available. See main comment.
Line 470: “>700m”, see comment above
Line 472: I don’t understand the reference to de Boyer Montégut et al. (2004). That paper describes the observational climatology, but here it is cited in the context of overly diffuse mixing schemes in ocean models.
The second part of the paragraph makes claims about the stratification (“weak thermocline/halocline gradients reduce vertical stratification, particularly during melt seasons”), but this has not been shown anywhere for the Labrador, Irminger, or Greenland Seas (the regions under discussion in this section). Thermocline/halocline smoothing has only been documented at the hydrographic sections located on the Norwegian coast.Line 514: “TOPAZ5 realistically captures large-scale ocean circulation, sea ice variability, and key seasonal hydrographic patterns, reproducing observed trends including surface warming, upper-ocean freshening, and changes in volume transport and circulation.”: I disagree with this characterization. Large-scale ocean circulation was not analyzed, sea ice variability is not captured realistically, and upper-ocean freshening is not validated (table 1 reports no observed SSS trend to compare against). Given the number and magnitude of the issues raised throughout this review, "realistically captures" considerably overstates the model's demonstrated performance.
References:
de Boyer Montégut, C., Madec, G., Fischer, A. S., Lazar, A., & Iudicone, D. (2004). Mixed layer depth over the global ocean: An examination of profile data and a profile-based climatology. Journal of Geophysical Research C: Oceans, 109(12), 1–20. https://doi.org/10.1029/2004JC002378
Hersbach, H., Bell, B., Berrisford, P., Biavati, G., Horányi, A., Sabater, M., et al. (2018). ERA5 hourly data on single levels from 1959 to present. https://doi.org/10.24381/cds.adbb2d47
Holte, J., & Talley, L. (2009). A new algorithm for finding mixed layer depths with applications to argo data and subantarctic mode water formation. Journal of Atmospheric and Oceanic Technology, 26(9), 1920–1939. https://doi.org/10.1175/2009JTECHO543.1
Holte, J., Talley, L. D., Gilson, J., & Roemmich, D. (2017). An Argo mixed layer climatology and database. Geophysical Research Letters, (May), 1–9. https://doi.org/10.1002/2017GL073426
Khosravi, N., Wang, Q., Koldunov, N., Hinrichs, C., Semmler, T., Danilov, S., & Jung, T. (2022). The Arctic Ocean in CMIP6 Models: Biases and Projected Changes in Temperature and Salinity. Earth’s Future, 10(2). https://doi.org/10.1029/2021EF002282
Shu, Q., Wang, Q., Guo, C., Song, Z., Wang, S., He, Y., & Qiao, F. (2023). Arctic Ocean simulations in the CMIP6 Ocean Model Intercomparison Project (OMIP). Geoscientific Model Development, 16(9), 2539–2563. https://doi.org/10.5194/gmd-16-2539-2023
Wang, Q., Shu, Q., Bozec, A., Chassignet, E. P., Fogli, P. G., Fox-Kemper, B., et al. (2024). Impact of increased resolution on Arctic Ocean simulations in Ocean Model Intercomparison Project phase 2 (OMIP-2). Geoscientific Model Development, 17(1), 347–379. https://doi.org/10.5194/gmd-17-347-2024
Citation: https://doi.org/10.5194/egusphere-2026-1520-RC2
Viewed
| HTML | XML | Total | BibTeX | EndNote | |
|---|---|---|---|---|---|
| 313 | 110 | 15 | 438 | 29 | 32 |
- HTML: 313
- PDF: 110
- XML: 15
- Total: 438
- BibTeX: 29
- EndNote: 32
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
This paper presents the latest version of the TOPAZ modelling system, v5, and evaluates it in free running mode (i.e. no data assimilation) across the pan-Arctic region. The authors focus the evaluation on a 10 year period (2009-2019) and consider fields such as sea surface temperature and salinity, sea surface height, sea ice thickness and concentration, mixed layer depth, and eddy kinetic energy. The authors also examine the 3D structure at a number of key sections, where model transports are calculated and evaluated. The paper shows where the model compares well with the observations as well as where the biases are larger.
This type of work needs to be published, as it provides the baseline to be references in all future papers that use TOPAZ5. GMD is an appropriate venue. That said, I think there are ways the manuscript can be improved, to make it more valuable to the broader community (i.e. not just those using TOPAZ). Additionally, I think some additional technical details, and some additional observational comparisons would help in providing a more complete observational baseline for this TOPAZ evaluation. I would thus recommend major revisions, with the specific details provided below.
A model evaluation paper can, in principle, be used by two research communities. One is users of the given model/configuration, to be able to refer to in future studies to show that the model is a good tool for their specific paper. In general, the manuscript does this well (comments on potentially improving this aspect are given below). However, the other community that can benefit from a model evaluation paper are those setting up their own model simulations, who can learn from the development of this given system. And this aspect is completely missing here. Yet, I doubt the PIs pulled the model out of the box, and ran just one experiment to get the results presented here. I would suspect that they ran multiple sensitivity experiments, to get the set of ocean and sea-ice parameters that they use. At least a sub-section on the steps and development would allow others to learn from the development work done here and increase the value and utility of this manuscript.
For either purpose, the summary of the model and its parameters is also missing relevant information. There is nothing on mixing schemes, parameters for viscosity and diffusion, relevant sea-ice variables such as ice strength, albedo, drag coefficients, etc. If these are all the same as the previous version of TOPAZ, this needs to be stated. But at the very least a table with all the relevant information about the model specifics is needed. As well, does the model include tidal forcing? Icebergs? All these details are needed.
Missing evaluation that would help the manuscript:
I think some needed model evaluation is missing. In terms of looking at the exchange out of the Arctic, the authors only look at Fram Strait. Yet, one question for models is how well they partition the Arctic transport each side of Greenland, which can significantly impact the sub-polar North Atlantic. I think it would be good to add a comparison section west of Greenland, such as Davis Strait.
I think the statement about good representation of deep water formation based on the MLDs is a bit simplified. It looks like the model has too deep deep convection, occurring over too broad a region – which is common for models of this resolution. Some more discussion of the implications of those biases would help. Plotting maximum MLD might also help with the evaluation and show if the model has convection to the bottom in some regions like the Labrador Sea.
Why not freshwater transport as well as volume transport to help evaluate the properties of the water carried through the studied sections? As above, would like to see a section west of Greenland. Also, for Bering Strait, the authors imply a close comparison with the 0.8 Sv of Woodgate et al. Yet the Woodgate et al. papers show an increasing trend that isn’t discussed.
The authors state credible AMOC dynamics. Yet no estimates of the model AMOC is presented. It would be good to see the modelling overturning in depth and density space along the OSNAP section, to compare with those estimates.
Something more on the deeper circulation, such as Atlantic Water circulation within the Arctic Ocean, and/or the overflows from the Nordic Seas would be useful evaluation.
Additional Items:
Is it more appropriate to discuss this paper as a model validation, or a model evaluation? I feel this manuscript is more focused on an observational evaluation, rather than a detailed validation of the code and numerical schemes.
L35: Note sure what “use in their majority advanced…” is trying to say.
Given the comment “horizontal resolution varying from 6.4 km at the Pacific boundary to 10.2 km at the Atlantic boundary”, might it be worth adding a panel to figure 1 showing a spatial map of the model resolution?
The model uses ERA5 forcing. Yet ERA5 is known to have a warm bias in the Arctic as well as having issues like occasional extremely strong and unrealistic wind events. Are any corrections applied to the ERA5 forcing?
L97: The authors mention estuaries along the Greenland coast. Do they mean fjords? In any case, I think it would be good to provide more detail about how this freshwater is distributed? For example, is any of it moved to the shelf for fjord systems that are not well resolved?
Why was the length of the spin-up decided to be the first 17 years? Any objective measure used to decide on that length?
Is the summer SST bias related to sea-ice biases?
Looking at the large SSS biases along the Arctic shelves, it looks as if there might be a river runoff issue?
For the section figures, they use discrete color intervals, which I feel enhances the figure and makes them easier to follow. Yet the spatial contour plots usen discrete color intervals, which blend together and are harder to read. I would suggest changing all of them to discrete color intervals as well.
Figure 4: The fonts are too small (comparing with figure 8 for example), and the lines not thick enough.
L296: The text states “large-scale altimetry pattern with high fidelity”. Is that a surprise if GLORYS is assimilating altimetric fields?
Figure 8: Font size is good but lines could be thicker.
L344: Given the comment about the limitations with small scale eddies – did the PIs want to then consider some eddy parameterization?