Ali Charara 0001

dblp:41/1192-1 · DBLP profile ↗
← Back
10ranked-venue papers
5as first author
0since 2021 · last 2019
0000-0002-9509-7794ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 4 first-authorTheory of computation · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
High-performance computing · 43% Parallel and multicore computing · 36% GPUs and heterogeneous computing · 10%

Topics — the 8 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Parallel and multicore computing
task scheduling
0.622019
SLATE: design of a modern distributed and accelerated linear algebra library · SC 2019
Pipelining Computational Stages of the Tomographic Reconstructor for Multi-Object Adaptive Optics on a Multi-GPU System · SC 2014
High-performance computing › parallel numerical algorithms
communication-avoiding algorithms
0.412019
SLATE: design of a modern distributed and accelerated linear algebra library · SC 2019
High-performance computing › numerical linear algebra
dense linear algebra
0.412019
SLATE: design of a modern distributed and accelerated linear algebra library · SC 2019
GPUs and heterogeneous computing
multi-GPU computing
0.212014
Pipelining Computational Stages of the Tomographic Reconstructor for Multi-Object Adaptive Optics on a Multi-GPU System · SC 2014
Processor architecture and microarchitecture
pipelining
0.212014
Pipelining Computational Stages of the Tomographic Reconstructor for Multi-Object Adaptive Optics on a Multi-GPU System · SC 2014
Parallel and multicore computing › task scheduling
dynamic scheduling
0.112014
Pipelining Computational Stages of the Tomographic Reconstructor for Multi-Object Adaptive Optics on a Multi-GPU System · SC 2014
Parallel and multicore computing
parallel programming runtimes
0.112014
Pipelining Computational Stages of the Tomographic Reconstructor for Multi-Object Adaptive Optics on a Multi-GPU System · SC 2014
High-performance computing
scientific computing systems
0.112014
Pipelining Computational Stages of the Tomographic Reconstructor for Multi-Object Adaptive Optics on a Multi-GPU System · SC 2014

Methods — techniques the papers use, named apart from their topics

task-based scheduling · 0.4lookahead panels · 0.4task pipelining · 0.2dynamic scheduling · 0.2
YearPublicationVenuePosition
2019 Linear Systems Solvers for Distributed-Memory Machines with GPU Accelerators
Jakub Kurzak, Mark Gates, Ali Charara 0001, Asim YarKhan, Ichitaro Yamazaki, Jack J. Dongarra
Euro-Par3
2019 Least squares solvers for distributed-memory machines with GPU accelerators
abstract
This work presents an implementation of a linear least squares solver for distributed-memory machines with GPU accelerators, developed as part of the Software for Linear Algebra Targeting Exascale (SLATE) package. From the algorithmic standpoint, the work leverages recent advances in dense linear algebra, specifically the communication-avoiding QR factorization. From the implementation standpoint, the work represents a sharp departure from the traditional conventions established by legacy packages, such as LAPACK and ScaLAPACK. It is based on representing the matrix as a collection of individual tiles, and using batch operations for offloading work to accelerators. The article lays out the principles of the new approach, discusses the implementation details and presents the performance results.
Jakub Kurzak, Mark Gates, Ali Charara 0001, Asim YarKhan, Jack J. Dongarra
ICS3
2019 SLATE: design of a modern distributed and accelerated linear algebra library
abstract
The SLATE (Software for Linear Algebra Targeting Exascale) library is being developed to provide fundamental dense linear algebra capabilities for current and upcoming distributed high-performance systems, both accelerated CPU-GPU based and CPU based. SLATE will provide coverage of existing ScaLAPACK functionality, including the parallel BLAS; linear systems using LU and Cholesky; least squares problems using QR; and eigenvalue and singular value problems. In this respect, it will serve as a replacement for ScaLAPACK, which after two decades of operation, cannot adequately be retrofitted for modern accelerated architectures. SLATE uses modern techniques such as communication-avoiding algorithms, lookahead panels to overlap communication and computation, and task-based scheduling, along with a modern C++ framework. Here we present the design of SLATE and initial reports of several of its components.
Mark Gates, Jakub Kurzak, Ali Charara 0001, Asim YarKhan, Jack J. Dongarra
SC3
2019 Batched Triangular Dense Linear Algebra Kernels for Very Small Matrix Sizes on GPUs
abstract
Batched dense linear algebra kernels are becoming ubiquitous in scientific applications, ranging from tensor contractions in deep learning to data compression in hierarchical low-rank matrix approximation. Within a single API call, these kernels are capable of simultaneously launching up to thousands of similar matrix computations, removing the expensive overhead of multiple API calls while increasing the occupancy of the underlying hardware. A challenge is that for the existing hardware landscape (x86, GPUs, etc.), only a subset of the required batched operations is implemented by the vendors, with limited support for very small problem sizes. We describe the design and performance of a new class of batched triangular dense linear algebra kernels on very small data sizes (up to 256) using single and multiple GPUs. By deploying recursive formulations, stressing the register usage, maintaining data locality, reducing threads synchronization, and fusing successive kernel calls, the new batched kernels outperform existing state-of-the-art implementations.
Ali Charara 0001, David E. Keyes, Hatem Ltaief
ACM Trans. Math. Softw.1
2018 Exploiting Data Sparsity for Large-Scale Matrix Computations
Kadir Akbudak, Hatem Ltaief, Aleksandr Mikhalev, Ali Charara 0001, Aniello Esposito, David E. Keyes
Euro-Par4
2018 Tile Low-Rank GEMM Using Batched Operations on GPUs
Ali Charara 0001, David E. Keyes, Hatem Ltaief
Euro-Par1
2018 Real-Time Massively Distributed Multi-object Adaptive Optics Simulations for the European Extremely Large Telescope
abstract
The European Extremely Large Telescope (E-ELT) is one of today's most challenging projects in ground-based astronomy. Addressing one of the key science cases for the E-ELT, the study of the early Universe, requires the implementation of multi-object adaptive optics (MOAO), a dedicated concept relying on turbulence tomography. We use a novel pseudo-analytical approach to simulate the performance of tomographic reconstruction of the atmospheric turbulence in an MOAO system on real datasets. We simulate simultaneously 4K galaxies in a common field of view on massively parallel supercomputers during a single night of observations. We are able to generate a first-ever high-resolution galaxy map at almost a real-time throughput. This simulation scale opens new research horizons in numerical methods for experimental astronomy, some core components of the pipeline standing as pathfinders toward actual operations and future astronomic discoveries on the E-ELT.
Hatem Ltaief, Ali Charara 0001, Damien Gratadour, Nicolas Doucet, Bilel Hadri, Eric Gendron, Saber Feki, David E. Keyes
IPDPS2
2017 A framework for dense triangular matrix kernels on various manycore architectures
abstract
Summary We present a new high‐performance framework for dense triangular Basic Linear Algebra Subroutines (BLAS) kernels, ie, triangular matrix‐matrix multiplication (TRMM) and triangular solve (TRSM), on various manycore architectures. This is an extension of a previous work on a single GPU by the same authors, presented at the EuroPar'16 conference, in which we demonstrated the effectiveness of recursive formulations in enhancing the performance of these kernels. In this paper, the performance of triangular BLAS kernels on a single GPU is further enhanced by implementing customized in‐place CUDA kernels for TRMM and TRSM, which are called at the bottom of the recursion. In addition, a multi‐GPU implementation of TRMM and TRSM is proposed and we show an almost linear performance scaling, as the number of GPUs increases. Finally, the algorithmic recursive formulation of these triangular BLAS kernels is in fact oblivious to the targeted hardware architecture. We, therefore, port these recursive kernels to homogeneous x86 hardware architectures by relying on the vendor optimized BLAS implementations. Results reported on various hardware architectures highlight a significant performance improvement against state‐of‐the‐art implementations. These new kernels are freely available in the KAUST BLAS (KBLAS) open‐source library at https://github.com/ecrc/kblas .
Ali Charara 0001, David E. Keyes, Hatem Ltaief
Concurr. Comput. Pract. Exp.1
2016 Redesigning Triangular Dense Matrix Computations on GPUs
Ali Charara 0001, Hatem Ltaief, David E. Keyes
Euro-Par1
2014 Pipelining Computational Stages of the Tomographic Reconstructor for Multi-Object Adaptive Optics on a Multi-GPU System
abstract
The European Extremely Large Telescope project (E-ELT) is one of Europe's highest priorities in ground-based astronomy. ELTs are built on top of a variety of highly sensitive and critical astronomical instruments. In particular, a new instrument called MOSAIC has been proposed to perform multi-object spectroscopy using the Multi-Object Adaptive Optics (MOAO) technique. The core implementation of the simulation lies in the intensive computation of a tomographic reconstruct or (TR), which is used to drive the deformable mirror in real time from the measurements. A new numerical algorithm is proposed (1) to capture the actual experimental noise and (2) to substantially speed up previous implementations by exposing more concurrency, while reducing the number of floating-point operations. Based on the Matrices Over Runtime System at Exascale numerical library (MORSE), a dynamic scheduler drives all computational stages of the tomographic reconstruct or simulation and allows to pipeline and to run tasks out-of order across different stages on heterogeneous systems, while ensuring data coherency and dependencies. The proposed TR simulation outperforms asymptotically previous state-of-the-art implementations up to 13-fold speedup. At more than 50000 unknowns, this appears to be the largest-scale AO problem submitted to computation, to date, and opens new research directions for extreme scale AO simulations.
Ali Charara 0001, Hatem Ltaief, Damien Gratadour, David E. Keyes, Arnaud Sevin, Ahmad Abdelfattah, Eric Gendron, Carine Morel, Fabrice Vidal
SC1