Search engine for discovering works of Art, research articles, and books related to Art and Culture
ShareThis
Javascript must be enabled to continue!

Julia for Geophysical Fluid Dynamics: Performance Comparisons between CPU, GPU, and Fortran-MPI

View through CrossRef
Some programming languages are easy to develop at the cost of slow execution, while others are lightning fast at run time but are much more difficult to write. Julia is a programming language that aims to be the best of both worlds—a development and production language at the same time. To test Julia’s utility in scientific high-performance computing (HPC), we built an unstructured-mesh shallow water model in Julia and compared it against an established Fortran-MPI ocean model, MPAS-Ocean, as well as a Python shallow water code. Three versions of the Julia shallow water code were created, for: single-core CPU; graphics processing unit (GPU); and Message Passing Interface (MPI) CPU clusters. Comparing identical simulations revealed that our first version of the single-core CPU Julia model was 13 times faster than Python. Further Julia optimizations, including static typing and removing implicit memory allocations, provided an additional 10–20x speed-up of the single-core CPU Julia model. The GPU-accelerated Julia code is extremely fast, with a speed-up of 230-380x compared to the single-core CPU Julia code if communication with the GPU occurs every 10 time steps. Parallelized Julia-MPI performance was identical to Fortran-MPI MPAS-Ocean for low processor counts, and ranges from 2x faster to 2x slower for higher processor counts. Our experience is that Julia development is fast and convenient for prototyping, but that Julia requires further investment and expertise to be competitive with compiled codes. We provide advice on Julia code optimization for HPC systems.
Title: Julia for Geophysical Fluid Dynamics: Performance Comparisons between CPU, GPU, and Fortran-MPI
Description:
Some programming languages are easy to develop at the cost of slow execution, while others are lightning fast at run time but are much more difficult to write.
Julia is a programming language that aims to be the best of both worlds—a development and production language at the same time.
To test Julia’s utility in scientific high-performance computing (HPC), we built an unstructured-mesh shallow water model in Julia and compared it against an established Fortran-MPI ocean model, MPAS-Ocean, as well as a Python shallow water code.
Three versions of the Julia shallow water code were created, for: single-core CPU; graphics processing unit (GPU); and Message Passing Interface (MPI) CPU clusters.
Comparing identical simulations revealed that our first version of the single-core CPU Julia model was 13 times faster than Python.
Further Julia optimizations, including static typing and removing implicit memory allocations, provided an additional 10–20x speed-up of the single-core CPU Julia model.
The GPU-accelerated Julia code is extremely fast, with a speed-up of 230-380x compared to the single-core CPU Julia code if communication with the GPU occurs every 10 time steps.
Parallelized Julia-MPI performance was identical to Fortran-MPI MPAS-Ocean for low processor counts, and ranges from 2x faster to 2x slower for higher processor counts.
Our experience is that Julia development is fast and convenient for prototyping, but that Julia requires further investment and expertise to be competitive with compiled codes.
We provide advice on Julia code optimization for HPC systems.

Related Results

On the programmability of multi-GPU computing systems
On the programmability of multi-GPU computing systems
Multi-GPU systems are widely used in High Performance Computing environments to accelerate scientific computations. This trend is expected to continue as integrated GPUs will be i...
New approaches for resource management and job scheduling for HEP grid computing
New approaches for resource management and job scheduling for HEP grid computing
(English) The Large Hadron Collider (LHC) ALICE (A Large Ion Collider Experiment) experiment uses grid computing for its extensive data processing and analysis. The ALICE Grid is c...
Porting NEMO diagnostics to GPU accelerators
Porting NEMO diagnostics to GPU accelerators
<p>This work makes part of an effort to make NEMO capable of taking advantage of modern accelerators. To achieve this objective we focus on port routines in NEMO that...
GPU atau singkatan dari Graphical Processing Unit merupakan mikroprosesor khusus yang berfungsi memercepat proses rendering grafik 2 dimensi atau 3 dimensi. GPU telah digunakan di ...
The MPI/OmpSs parallel programming model
The MPI/OmpSs parallel programming model
Even today supercomputing systems have already reached millions of cores for a single machine, which are connected by using a complex network interconnection. Reducing communicatio...
CPU AND GPU (CUDA) TEMPLATE MATCHING COMPARISON / CPU IR GPU (CUDA) PALYGINIMAS VYKDANT ŠABLONŲ ATITIKTIES ALGORITMĄ
CPU AND GPU (CUDA) TEMPLATE MATCHING COMPARISON / CPU IR GPU (CUDA) PALYGINIMAS VYKDANT ŠABLONŲ ATITIKTIES ALGORITMĄ
Image processing, computer vision or other complicated opticalinformation processing algorithms require large resources. It isoften desired to execute algorithms in real time. It i...
Low-power architectures for automatic speech recognition
Low-power architectures for automatic speech recognition
Automatic Speech Recognition (ASR) is one of the most important applications in the area of cognitive computing. Fast and accurate ASR is emerging as a key application for mobile a...
Multidimensional Prognostic Index (MPI) in elderly patients with acute myocardial infarction
Multidimensional Prognostic Index (MPI) in elderly patients with acute myocardial infarction
Abstract Background Management of elderly patients with acute myocardial infarction (AMI) is challenging due to lack of knowledge about the link bet...

Back to Top