Javascript must be enabled to continue!
Exploring the Offload Execution Model in the Intel Xeon Phi via Matrix Inversion
View through CrossRef
The explicit inversion of dense matrices appears in a numerous key scientific and engineering applications such as model reduction or optimal control, asking for the exploitation of high performance computing techniques and architectures when the problem dimension is large. Gauss-Jordan elimination (GJE) is an efficient in-place method for matrix inversion that exposes large amounts of dataparallelism, making it very convenient for hardware accelerators such as graphics processors (GPUs) or the Intel Xeon Phi. In this paper, we present and evaluate several practical implementations of GJE, with partial row pivoting, that especially exploit the off-load execution model available on the Intel Xeon Phi to carry out a significant fraction of the computations on the accelerator. Numerical experiments on a system with two Intel Xeon E5-2640v3 processors and an Intel Xeon Phi 7120P compare the efficiency of these implementations, with the most efficient case delivering about 700 billions double-precision floating-point operations per second.
Title: Exploring the Offload Execution Model in the Intel Xeon Phi via Matrix Inversion
Description:
The explicit inversion of dense matrices appears in a numerous key scientific and engineering applications such as model reduction or optimal control, asking for the exploitation of high performance computing techniques and architectures when the problem dimension is large.
Gauss-Jordan elimination (GJE) is an efficient in-place method for matrix inversion that exposes large amounts of dataparallelism, making it very convenient for hardware accelerators such as graphics processors (GPUs) or the Intel Xeon Phi.
In this paper, we present and evaluate several practical implementations of GJE, with partial row pivoting, that especially exploit the off-load execution model available on the Intel Xeon Phi to carry out a significant fraction of the computations on the accelerator.
Numerical experiments on a system with two Intel Xeon E5-2640v3 processors and an Intel Xeon Phi 7120P compare the efficiency of these implementations, with the most efficient case delivering about 700 billions double-precision floating-point operations per second.
Related Results
Parallel approaches of genetic algorithm in the MIC architecture of the Intel Xeon Phi
Parallel approaches of genetic algorithm in the MIC architecture of the Intel Xeon Phi
Today, genetic algorithms are widely used in many fields such as bioinformatics, computer science, artificial intelligence, finance ... Genetic algorithms are applied to create hig...
Evaluación de rendimiento y eficiencia energética de sistemas heterogéneos para bioinformática
Evaluación de rendimiento y eficiencia energética de sistemas heterogéneos para bioinformática
El problema del consumo energético se presenta como uno de los mayores obstáculos para el diseño de sistemas que sean capaces de alcanzar la escala de los Exaflops. Por lo tanto, l...
HPC-BLAST: Distributed BLAST for Modern HPC Clusters.
HPC-BLAST: Distributed BLAST for Modern HPC Clusters.
The near exponential growth in sequence data available to bioinformaticists, and the emergence of new fields of biological research, continue to fuel an incessant need for in- crea...
Сравнение стратегий распараллеливания векторизованного римановского решателя с помощью OpenMP для микропроцессора Intel Xeon Phi KNL
Сравнение стратегий распараллеливания векторизованного римановского решателя с помощью OpenMP для микропроцессора Intel Xeon Phi KNL
Римановские решатели широко используются в численных методах, при решении задач газовой динамики. При этом во время проведения вычислений требуется решать задачу Римана о распаде п...
Un manoscritto equivocato del copista santo Theophilos († 1548)
Un manoscritto equivocato del copista santo Theophilos († 1548)
<p><font size="3"><span class="A1"><span style="font-family: 'Times New Roman','serif'">ΕΝΑ ΛΑΝ&...
LU Factorisation on Xeon and Xeon Phi Processors
LU Factorisation on Xeon and Xeon Phi Processors
This paper outlines the parallelisation and vectorisation methods we have used to port a LU decomposition library to the Xeon Phi co-processor. We ported a LU factorisation algorit...
Abstract 1627: PHI-501, a novel and potent pan-RAF inhibitor in metastatic melanoma
Abstract 1627: PHI-501, a novel and potent pan-RAF inhibitor in metastatic melanoma
Abstract
Background: PHI-501 has been developed as a novel inhibitor of NRAS mutated acute myeloid leukemia. Big data and artificial intelligence (AI)-based drug dis...
Improving decision tree and neural network learning for evolving data-streams
Improving decision tree and neural network learning for evolving data-streams
High-throughput real-time Big Data stream processing requires fast incremental algorithms that keep models consistent with most recent data. In this scenario, Hoeffding Trees are c...

