Javascript must be enabled to continue!
Locality-Aware Task Scheduling and Data Distribution for OpenMP Programs on NUMA Systems and Manycore Processors
View through CrossRef
Performance degradation due to nonuniform data access latencies has worsened on NUMA systems and can now be felt on-chip in manycore processors. Distributing data across NUMA nodes and manycore processor caches is necessary to reduce the impact of nonuniform latencies. However, techniques for distributing data are error-prone and fragile and require low-level architectural knowledge. Existing task scheduling policies favor quick load-balancing at the expense of locality and ignore NUMA node/manycore cache access latencies while scheduling. Locality-aware scheduling, in conjunction with or as a replacement for existing scheduling, is necessary to minimize NUMA effects and sustain performance. We present a data distribution and locality-aware scheduling technique for task-based OpenMP programs executing on NUMA systems and manycore processors. Our technique relieves the programmer from thinking of NUMA system/manycore processor architecture details by delegating data distribution to the runtime system and uses task data dependence information to guide the scheduling of OpenMP tasks to reduce data stall times. We demonstrate our technique on a four-socket AMD Opteron machine with eight NUMA nodes and on the TILEPro64 processor and identify that data distribution and locality-aware task scheduling improve performance up to 69% for scientific benchmarks compared to default policies and yet provide an architecture-oblivious approach for programmers.
Title: Locality-Aware Task Scheduling and Data Distribution for OpenMP Programs on NUMA Systems and Manycore Processors
Description:
Performance degradation due to nonuniform data access latencies has worsened on NUMA systems and can now be felt on-chip in manycore processors.
Distributing data across NUMA nodes and manycore processor caches is necessary to reduce the impact of nonuniform latencies.
However, techniques for distributing data are error-prone and fragile and require low-level architectural knowledge.
Existing task scheduling policies favor quick load-balancing at the expense of locality and ignore NUMA node/manycore cache access latencies while scheduling.
Locality-aware scheduling, in conjunction with or as a replacement for existing scheduling, is necessary to minimize NUMA effects and sustain performance.
We present a data distribution and locality-aware scheduling technique for task-based OpenMP programs executing on NUMA systems and manycore processors.
Our technique relieves the programmer from thinking of NUMA system/manycore processor architecture details by delegating data distribution to the runtime system and uses task data dependence information to guide the scheduling of OpenMP tasks to reduce data stall times.
We demonstrate our technique on a four-socket AMD Opteron machine with eight NUMA nodes and on the TILEPro64 processor and identify that data distribution and locality-aware task scheduling improve performance up to 69% for scientific benchmarks compared to default policies and yet provide an architecture-oblivious approach for programmers.
Related Results
High-level compiler analysis for OpenMP
High-level compiler analysis for OpenMP
Nowadays, applications from dissimilar domains, such as high-performance computing and high-integrity systems, require levels of performance that can only be achieved by means of s...
Towards a safe and efficient OpenMP
Towards a safe and efficient OpenMP
(English) The growing complexity of contemporary multi-core and heterogeneous architectures necessitates parallel programming models capable of efficiently leveraging the available...
Applying a user-centered approach to evaluate the usability of a mobile application for health professionals in home care services (Preprint)
Applying a user-centered approach to evaluate the usability of a mobile application for health professionals in home care services (Preprint)
BACKGROUND
Mobile health (mHealth), or the use of mobile devices in medicine and health, is a sub-category of e-health. mHealth interventions are designed t...
Automatic Parallelization for Heterogeneous Embedded Systems
Automatic Parallelization for Heterogeneous Embedded Systems
Parallélisation automatique pour systèmes hétérogènes embarqués
L'utilisation d'architectures hétérogènes, combinant des processeurs multicoeurs avec des accélérate...
Probabilistically time-analyzable complex processor designs
Probabilistically time-analyzable complex processor designs
Industry developing Critical Real-Time Embedded Systems (CRTES), such as Aerospace, Space, Automotive and Railways, faces relentless demands for increased guaranteed processor perf...
Grain graphs
Grain graphs
Average programmers struggle to solve performance problems in OpenMP programs with tasks and parallel for-loops. Existing performance analysis tools visualize OpenMP task performan...
Classification, natural history, and evolution of Epiphloeinae (Coleoptera: Cleridae). Part VII. The genera Hapsidopteris Opitz, Iontoclerus Opitz, Katamyurus Opitz, Megatrachys Opitz, Opitzia Nemesio, Pennasolis Opitz, new genus, Pericales Opitz, new gen
Classification, natural history, and evolution of Epiphloeinae (Coleoptera: Cleridae). Part VII. The genera Hapsidopteris Opitz, Iontoclerus Opitz, Katamyurus Opitz, Megatrachys Opitz, Opitzia Nemesio, Pennasolis Opitz, new genus, Pericales Opitz, new gen
This study deals with minimally speciose epiphloeine genera. Hapsidopteris, based on H. diastenus Opitz, (type locality: México: Jalapa), is the presumed sister taxon of Opitzia Ne...
Heterogeneous Uncore Architectures in Future Chip-Multi Processors
Heterogeneous Uncore Architectures in Future Chip-Multi Processors
<p>Uncore components including on-chip memory systems and interconnects consume a significant portion of overall energy consumption in emerging embedded applications. Proposi...

