Xavier Garbet

dblp:119/5058 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
0since 2021 · last 2017
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
GPUs and heterogeneous computing · 36% High-performance computing · 28% Hardware accelerators and domain-specific architectures · 28%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Hardware accelerators and domain-specific architectures
accelerator optimization
0.312017
Optimization of Fusion Kernels on Accelerators with Indirect or Strided Memory Access Patterns · IEEE Trans. Parallel Distributed Syst. 2017
GPUs and heterogeneous computing › GPU memory management
GPU memory access optimization
0.312017
Optimization of Fusion Kernels on Accelerators with Indirect or Strided Memory Access Patterns · IEEE Trans. Parallel Distributed Syst. 2017
High-performance computing
stencil computation
0.312017
Optimization of Fusion Kernels on Accelerators with Indirect or Strided Memory Access Patterns · IEEE Trans. Parallel Distributed Syst. 2017
Storage systems
data layout
0.112017
Optimization of Fusion Kernels on Accelerators with Indirect or Strided Memory Access Patterns · IEEE Trans. Parallel Distributed Syst. 2017
GPUs and heterogeneous computing
GPU computing
0.112017
Optimization of Fusion Kernels on Accelerators with Indirect or Strided Memory Access Patterns · IEEE Trans. Parallel Distributed Syst. 2017

Methods — techniques the papers use, named apart from their topics

temporal blocking · 0.3register reuse · 0.3data layout transformation · 0.3
YearPublicationVenuePosition
2017 Optimization of Fusion Kernels on Accelerators with Indirect or Strided Memory Access Patterns
abstract
This paper describes optimization for high-dimensional stencil computations on accelerators involving complex memory access patterns, which appear in five dimensional fusion plasma turbulence codes, GYSELA and GT5D. They include different types of memory access patterns, the indirect memory access in GYSELA with a Semi-Lagrangian scheme and the strided memory access in GT5D with a Finite-Difference scheme. We focus on the affinity of the memory access patterns to accelerators such as GPGPUs and Xeon Phi coprocessors. On both devices, the Array of Structure of Array (AoSoA) data layout is preferable for contiguous memory accesses. It is shown that the effective local cache usage by improving spatial and temporal data locality is critical on Xeon Phi. On GPGPU, the texture memory usage improves the performance of the indirect memory accesses in the Semi-Lagrangian scheme. The reuse of registers by taking account of the physical symmetry of the Finite-Difference scheme reduces the amount of memory accesses. Through these optimizations, we achieve acceleration of 3.9 (8.1) on Xeon Phi (GPGPU) for the Semi-Lagrangian scheme and of 1.4 (3.9) on Xeon Phi (GPGPU) for the Finite-Different scheme with respect to the fully optimized codes on Sandy Bridge.
Yuuichi Asahi, Guillaume Latu, Takuya Ina, Yasuhiro Idomura, Virginie Grandgirard, Xavier Garbet
IEEE Trans. Parallel Distributed Syst.6