the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
Full parallelization of the finite-element Lagrangian sea ice model neXtSIM for kilometer-scale simulations
Abstract. Accurate modeling of sea ice dynamics is a major challenge, but also a key requirement for forecasting its future evolution and assessing its impact on climate change. The sea-ice model neXtSIM is specifically designed for this purpose, relying on a Lagrangian framework to accurately capture highly localized features such as leads and ridges, which likely play an important role in controlling the energy exchanges at the interfaces of the atmosphere-ice-ocean coupled system. Despite a parallelisation effort a few years ago, several components of the code, which are intrinsically required by a Lagrangian framework, such as the remeshing procedure, remained sequential, which created significant bottlenecks that limited the model's overall capabilities. To tackle this problem, we further developed the parallel anisotropic mesh adaptation tool ParMMG2D, which demonstrates excellent scalability up to 512 processors with 20 million elements. This tool has been integrated within neXtSIM to enable parallel remeshing, currently limited to homogeneous isotropic meshes. In addition, parallelization of subsequent interpolation, output writing, and drifting buoy tracking has also been achieved. Together, these improvements now enable kilometer-scale simulations within a reasonable timeframe: a one-year simulation with 2 km elements takes only a few days, while a full season with 1 km elements can be simulated in about a week on 128 processors. Therefore, neXtSIM now stands out as cutting-edge software for simulating highly resolved sea ice dynamics, combining the accuracy of a Lagrangian framework with high computational efficiency.
- Preprint
(53558 KB) - Metadata XML
- BibTeX
- EndNote
Status: final response (author comments only)
- RC1: 'Comment on egusphere-2026-1869', Anonymous Referee #1, 10 Aug 2026
-
RC2: 'Comment on egusphere-2026-1869', Anonymous Referee #2, 02 Sep 2026
The manuscript presents a description of several parallelization instances that were missing in the previous implementation of neXtSIM and were limiting its scalability. As such, the manuscript can be of interest for the journal, yet its present form is far from being optimal. It gives only cursory presentation of many places that could be of interest. Instead, it spends several pages of section 2.1 on the equations of the dynamical part of the model which are never used in the remainder and can be safely omitted, especially because the description is just published in GMD. Some figures and test cases are redundant, as performance can only be attested in real-world simulations. No explanation is given to what is advected and to what an extent the use of the Lagrangian scheme is critical. It is clear that it might help to preserve localized features, but all discrete differential operators are very imprecise on the grid scale the scheme may just emphasize numerical noise.
My specific comments follow the text of the manuscript.
Line 3, 4 -- I would question the need for Lagrangian framework for the reasons just mentioned.
Line 8 This gives about 4000 elements per core, which is a rather large partition. Please comment.
10 - 15 This is still extremely slow, so it cannot be a cutting-edge approach.
Section 2: I do not see the need for 2.1. Instead, I would present some detail of what is advected, how often remeshing is needed, how this depends on resolution, etc., i.e. things that might help to understand the rest (this is only mentioned in passing in 2.2.1).
162 Check the formula - its lhs is a matrix entry, rhs is a matrix
Section 3.2. From the description in this section I cannot understand what precisely is done. In particular, the description of the advancing-front algorithm and conditions under which it is invoked is insufficient. For example: On line 181: 'if a vertex ...' -- such vertices should be present all the time on the periphery of partitioning if it is element-based. Please explain. Figure 1 is useless, as it tells nothing about when and why the advancing front is invoked. Line 191 'Following each remeshing ...' An advancing front was discussed, is it meant as remeshing? Further, the need for remeshing: please be more precise, and say what is finally used, how often, what are the criteria.
Algorithm 1 -- Section 3.2 should introduce and explain all steps mentioned here, also informing the reader about what is n_iter, and what criteria are used. Otherwise it only conveys an approximate idea, but is not reproducible.
Section 3.3. One test case will do, as all they convey a similar message, and neither of them is informative enough as concerns real-world applications.
A similar remark for section 3.4 -- both tests lead to similar conclusions, so leave one of them, and say that the other one is similar. From the left panels in either Fig. 7 or 8 it follows that all time is spent in remeshing, all the rest does not matter, so speedup of redistribution and interpolation is of no interest.
Then, it should be clear that the parallel scaling (speedup) depends largely on the partition size (mesh cells per partition), which does not exclude some sensitivity to the mesh size. It would be therefore much more enlightening to combine the cases with meshes of different size to a single plot of speedup tor total time, instead of 6 separate plots, as there is only one relevant point -- the limit of strong scalability.259 ... around eight partitions -- I would guess it is hardware dependent.
268 ... Scalability is more favorable ... -- Both are just test cases, so there is no point in this statement.314 ... to be profitable ? In which sense?
323 ... Conversely, for coarse meshes ... Why? It should depend on the number of mesh elements per core.
334 ... The same Cartesian ... What is meant under Cartesian?
Please explain everything clearer in this section. There is no description of interpolation, and information provided does not cover the topic at all. Figure 9 hardly explains anything.
4.3 Please provide more detail, there is almost no information allowing one to understand what is done. Which reordering is meant?
Fig. 10 Once again, it would be more relevant to show the dependence on the partition size (the number of mesh elements per partition)
Fig. 11 Why the saturation is seen so early? Such meshes should be run with a much larger number of partitions, so what happens if 512 cores are used?
376 Please provide more details and better incorporate Fig. 12. Why such a selection of bounding boxes is used? Is this technology optimal? In Fig. 12 I see the boxes of very different sizes, which naturally leads to this question.
403 ... are broadcast to every processor -- but this is a global operation with a limited parallel scalability.
Why Fig. 15 is needed? There is no other message than the one on lines 430,431. Please remove.
434, 435 Should the parallel remeshing introduce them? Remove the last sentence.
Figure 16. Same remarks as earlier -- one is generally interested in strong scalability limit, which is commonly related to the partition size. Please provide this information. In this sense only panel b provides this information in a reliable way, and d and f require to increase the number of partitions. Also the information of the speedup of remeshing is redundant in the right plots, as the left plots already tell that the time spent on parallel remeshing becomes irrelevant.
What is the intention of Fig. 17? It carries no message in the context of the manuscript and can be safely omitted.
465 High accuracy can be achieved in many ways, and the Lagrangian framework is not necessarily the best.
482 .... Despite -- your results show that remeshing becomes a secondary factor. Why then 'despite'? The performance stated as feasible is still rather slow. Where it is lost? And what sets strong scalability limit of the fully parallelized version?
Citation: https://doi.org/10.5194/egusphere-2026-1869-RC2
Viewed
| HTML | XML | Total | BibTeX | EndNote | |
|---|---|---|---|---|---|
| 135 | 77 | 15 | 227 | 14 | 13 |
- HTML: 135
- PDF: 77
- XML: 15
- Total: 227
- BibTeX: 14
- EndNote: 13
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
General comments
The presented manuscript contributes to the well-established neXtSIM model development framework. The primary development introduced is the full parallelization of the Lagrangian sea ice model, which enables kilometer-scale simulations within a reasonable runtime. This full parallelization encompasses a set of improvements to the current version of neXtSIM, most notably the implementation of a parallel anisotropic mesh adaptation tool that addresses a major structural bottleneck in the remeshing procedure of previous parallel version. The authors first validate these tools using controlled test cases before applying them to the full neXtSIM model. They demonstrate that the output of the fully parallelized model is consistent with that of the previous version, while achieving a substantial reduction in simulation time.
Overall, the manuscript is clear and well structured, with a title and abstract that reflect the scope and content of the paper. The introduction clearly outlines the limitations of the current neXtSIM parallel implementation and frames the novel development. However, some parts of the introduction need clarification and details (see Specific Comments). The mathematical formulations are generally clear, though a few additional details noted below could further enhance readability. The methodological advancements are thoroughly described, with the main development (ParMMG2D) reasonably placed in a dedicated section of the paper. A few specific comments on the validation methodology may help strengthen the presentation. Finally, the results are sufficiently clear to prove that the parallelization strategy is effective.
In summary, the developments presented in the manuscript are worthy of consideration for publication in GMD, since the authors address questions within the scope of the journal. The novel contribution of this work lies in the application of tools previously unused within the neXtSIM framework.
Specific comments
Lines 17-18: “...shift in the dynamical regime, characterized by an increase in extreme fracturing events and an acceleration of sea ice drift”. Please, provide at least one reference for both the mentioned characteristics, it is not straightforward that in the upcoming Arctic regime the fracturing events will increase.
Lines 35-42: In this paragraph, the authors do not discuss why they parallelized an open-source library (MMG), which is good, instead of adopting an already parallelized remeshing tool. It can be that the latter doesn’t exist, but please specify.
Equations (6)-(7): Pmax is defined but its meaning is not specified. It would be nice for the reader to have few lines describing why it is important.
Equation (10-11): dcrit is used but not defined.
Line 135-138: Does each process read the full-domain forcing files or just a portion defined in the pre-processing stage? From the last paragraph of Conclusions I guess the first option, however, is not clearly stated.
Line 205: Please, elaborate the meaning of “the meshes must look clean”.
Algorithm 1: It is not referenced within the text.
Line 237-238 and figures 4-5: The validation test results in Fig. 4 display low quality in the final mesh from the function f1, with close to 20% of elements having quality smaller than 0.8. Could you elaborate on these results further? What are the reasons behind and is this result acceptable?
Line 266-267: I find it difficult to understand if there are any reasons why scalability performance is better for the shock wave test case than for the circle test case? Do you have any ideas?
Figure 9: Please, indicate the meaning of the numbers in the caption.
Line 372-373 and section 4.4: Please, clarify what you mean with “each point must be handled individually”. I might have lost something, but I don’t understand why you define a new rectangular partitioning instead of directly printing information from the nodes of the domain decomposition with MPI I/O routines. A mapping from local to global domain and a condition for avoiding simultaneous writing would allow you to efficiently store the information in a binary file (I don’t know if MPI I/O supports NetCDF).
Line 419: Is there any reason why you decided to start the validation process from February?
Figure 15: I think that the line is too thick compared to the dots, this may lead to not easy evaluations, especially in panel (c).
Line 430: For the sake of completeness, please also report the maximum discrepancy between the two models.
Line 433 and Figure 15(c): Do you have any idea of why larger differences between the models arise at local minima and maxima?
Technical corrections
Line 72: ice density does not appear in equation (1) where it is stated, but in (12).
Line 73-74: f (Coriolis factor), 𝜂 (SSH), and g (gravity constant) are not specified.
Line 119: Space missing after “Boost.”
Line 176: not uniform citation style.
References: last three items are not in alphabetical order.