the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
PSyclone 3: A source-to-source Fortran compiler for developing maintainable and performance-portable HPC applications
Abstract. PSyclone is a source-to-source Fortran compiler designed to programmatically optimize, parallelize, and instrument HPC applications via user-provided transformation scripts. These scripts allow for a clear separation of concerns between the scientific model, described in Fortran, and the optimization choices. This separation improves the maintainability of complex scientific applications by enabling independent exploration and development of each aspect. In addition, it provides a solution to achieve better performance portability for Fortran applications by encoding the transformations that are beneficial to each platform and compiler in different transformation scripts.
PSyclone supports two modes of operation. The first mode optimizes existing source code, including MPI-based applications, by making the necessary code transformations to effectively use the capabilities and programming models supported by each CPU and GPU vendor. This approach is demonstrated using a benchmark extracted from the NEMO ocean model. The second mode defines a kernel-based parallelism model with domain-specific Fortran-embedded metadata. This approach enables a stricter separation of concerns, thereby allowing PSyclone to take full control of the data dependencies and iteration spaces in order to improve its capabilities and generate distributed and shared-memory parallelism. This approach has been co-designed with the Met Office and is used in the Finite-Element based dynamical core of the LFRic atmospheric model.
Competing interests: David Ham (editor) was an investigator on the original Gung-Ho project that saw the inception of the PSyclone software described in this work.
Publisher's note: Copernicus Publications remains neutral with regard to jurisdictional claims made in the text, published maps, institutional affiliations, or any other geographical representation in this paper. While Copernicus Publications makes every effort to include appropriate place names, the final responsibility lies with the authors. Views expressed in the text are those of the authors and do not necessarily reflect the views of the publisher.- Preprint
(706 KB) - Metadata XML
- BibTeX
- EndNote
Status: open (until 17 Oct 2026)
- RC1: 'Comment on egusphere-2026-3844', Anonymous Referee #1, 16 Sep 2026 reply
-
RC2: 'Comment on egusphere-2026-3844', Anonymous Referee #2, 07 Oct 2026
reply
The paper describes the PSyclone compiler and workflow for optimizing atmospheric models written in Fortran. PSyclone is a source-to-source compiler that can accept general Fortran code and DSLs such as the LFRic DSL. The authors describe the Intermediate Representation and transformations from a bird’s eye view, with novel transformations that exchange communication for computation and introduce concurrency. When combined with a DSL, PSyclone demonstrates modest (2-8%) performance improvements and a 31% reduction in the number of communication calls on the LFRic dynamical core.
It is not an easy feat to write a compiler infrastructure that works at the scale of a dynamical core (compiling the entire codebase, including I/O using Codeblocks) and attain performance improvements over existing codes. It demonstrates that introducing redundant computations and having a halo-exchange-aware system is beneficial, which is a novel contribution over other prior works. However, the manuscript in its current state could be improved on two fronts: motivation and evaluation.
The first front is the reasoning for choosing the specific PSyclone approach and a description of its components that would motivate its use. The manuscript contains a brief high-level description of the IR, but it does not explain what the IR nodes are. Not necessarily as a table of operations (though that might be helpful), but in the sense of what computational patterns can or cannot be supported. The same applies for the PSyIR transformations. The descriptions of LFRic transformations in Section 5.2 are great and could be a basis for that. Particularly the Codeblock idea is commendable and novel, and has the power to potentially ingest any HPC Fortran application, but it is unclear what is in the analyzable regime. Moreover, for the “HaloExchange” node described in line 165, does “extensible” mean that it is currently implemented or subject to future work? A full read of the paper suggests that a version of it exists in the DSL level, but it is not adequately explained in the PsyIR part of the work.
The question of the extent of support extends to PSyKAl: is code outside the dynamical core sufficiently analyzed? Other physics modules are notoriously challenging to represent and optimize.
In terms of motivating the approach and compiler framework, I believe the work is well-positioned among the other competitors, but that should be explained better in the manuscript. Section 2 outlines the other approaches, but does not provide a comparative analysis. One crucial question that arises is what is the benefit of using a source-to-source Fortran compiler rather than writing MLIR passes? With the latest developments in flang, could PSyIR not be a part of HLFIR/FIR? How does the work relate to the MLIR stencil dialect? (Gysi et al., “Domain-specific multi-level IR rewriting for GPU: The Open Earth compiler for GPU-accelerated climate simulation”, ACM Transactions on Architecture and Code Optimization, Vol. 18, No. 4) For example, an equivalent or variant of “ArrayAssignment2LoopTrans” should already be performed during HLFIR-to-FIR lowering. While I agree with some of the software development claims in Section 2, they should be qualified with a specific example.
As for the other DSLs and Fortran parsers, what transformations or code-ingestion capabilities make PSyclone unique over those other approaches?
The second front in which the paper can improve is evaluation. The manuscript currently only makes self-comparisons with PSyclone. While difficult to compare with other DSLs such as GT4Py, any benchmark with both a DSL and Fortran source code will shed light onto the efficacy of the approach. If comparisons with other DSLs are impossible, what about the other Fortran compiler frameworks, such as Loki and the DACE Fortran parser? More importantly, Section 4.1 is missing comparisons with the reference NEMO codebase and OpenMP code (or GPU/OpenACC, if available).
As for portability evaluation, the central point that Section 5 makes is reducing performance portability fragility with DSLs. This merits some presentation of results over GPUs. However, only CPU results are presented for the GungHo dynamical core.
If these two aspects of the papers can be improved, it would make a much stronger argument as a whole.
Specific comments follow:
Section 4.1: You might be overcounting the memory bandwidth, as multiple contiguous accesses (as occur in stencils) may already be cached. Please use performance counters to measure the actual memory transactions. Also, would STREAM Copy rather than Triad give a more accurate bound on measured memory bandwidth? Figure 3 could also benefit from raw runtime numbers.
Section 4.1: Based on the comments in Section 4.2, there may be stark differences in performance based on the structure (e.g., loop schedule) of the code. Are the schedules different between the CPU and GPU versions, or is the source code exactly the same? Namely, is it an unmodified version of the original NEMO code?
Section 5.2: MatMulToCodeTrans- will you not gain more from vendor BLAS libraries such as MKL/OpenBLAS/cuBLAS/rocBLAS?
Section 6: “The application of PSyclone to transform existing models such as NEMO will be the subject of a future publication.” - Is this not the goal of Section 4?
Technical corrections:line 151: “given” -> “giving”
line 323: “discreet H100” -> “discrete H100”
line 324: Shouldn’t the H100 have a compute capability of 90 rather than 70?
Runtimes should be relatively stable but please include error bars.
Typographical: the paper combines both American and British English spelling.Citation: https://doi.org/10.5194/egusphere-2026-3844-RC2
Viewed
| HTML | XML | Total | BibTeX | EndNote | |
|---|---|---|---|---|---|
| 129 | 61 | 76 | 266 | 82 | 70 |
- HTML: 129
- PDF: 61
- XML: 76
- Total: 266
- BibTeX: 82
- EndNote: 70
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
This is a well written paper exploring a production tool in use across several codes. It presents an interesting approach to DSLs and code transformation, and flexible way of working across two different paradigms (direct transformation, and DSLs built around its PSyKAl abstraction). Whilst the paper is fairly heavy on the engineering side, this is worthwhile and a useful tool by the community that others will learn from and leverage.
One suggestion here, although I appreciate it would be more difficult to do, would be to compare and contrast performance generated by PSyclone compared to some of the other tools described in Section 2. The evaluation only compares PSyclone against itself, and this would strengthen things, for instance in Section 4.1 do a version of the NEMO Tracer-Advection Benchmark which manually has OpenMP target offload and/or OpenACC and compare the two (maybe performance and LoC between them). But I think this is fine for a follow on paper.
A few minor suggestions:
- Mention of MLIR in section 2 is a bit naive, although that framework is written in C++ (like PSyclone is written in Python) it is compatible with Fortran, indeed Flang is built atop it.
- It would be good to present absolute numbers (e.g. runtime or GPt/s) in section 4.1, as it is currently quite difficult to interpret Figure 3.
- PSyclone generates OpenMP and OpenACC rather than CUDA or ROCm. It would be useful to add a sentence to explain why this is the case. It links back to my second point, as the reader I am left wondering if there is any performance left on the table for the GPU due to this decision.