Arash Ashari

dblp:146/3559 · DBLP profile ↗
← Back
4ranked-venue papers
4as first author
0since 2021 · last 2015
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 4 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
GPUs and heterogeneous computing · 82% High-performance computing · 18%
Theoretical computer science
1 paper
Graph algorithms and graph theory · 100%

Topics — the 5 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
GPUs and heterogeneous computing
GPU computing
0.212015
On optimizing machine learning workloads via kernel fusion · PPoPP 2015
GPUs and heterogeneous computing
GPU kernel optimization
0.212015
On optimizing machine learning workloads via kernel fusion · PPoPP 2015
GPUs and heterogeneous computing › GPU kernel optimization
kernel fusion
0.212015
On optimizing machine learning workloads via kernel fusion · PPoPP 2015
GPUs and heterogeneous computing › GPU computing › GPU sparse computation
GPU sparse linear algebra
0.212014
Fast Sparse Matrix-Vector Multiplication on GPUs for Graph Applications · SC 2014
High-performance computing › sparse linear algebra › sparse matrix computation
sparse matrix-vector multiplication
0.212014
Fast Sparse Matrix-Vector Multiplication on GPUs for Graph Applications · SC 2014

Methods — techniques the papers use, named apart from their topics

sparse matrix computation · 0.2analytical modeling · 0.2BLAS · 0.2
YearPublicationVenuePosition
2015 On optimizing machine learning workloads via kernel fusion
abstract
Exploitation of parallel architectures has become critical to scalable machine learning (ML). Since a wide range of ML algorithms employ linear algebraic operators, GPUs with BLAS libraries are a natural choice for such an exploitation. Two approaches are commonly pursued: (i) developing specific GPU accelerated implementations of complete ML algorithms; and (ii) developing GPU kernels for primitive linear algebraic operators like matrix-vector multiplication, which are then used in developing ML algorithms. This paper extends the latter approach by developing fused kernels for a combination of primitive operators that are commonly found in popular ML algorithms. We identify the generic pattern of computation (alpha * X^T (v * (X * y)) + beta * z) and its various instantiations. We develop a fused kernel to optimize this computation on GPUs -- with specialized techniques to handle both sparse and dense matrices. This approach not only reduces the cost of data loads due to improved temporal locality but also enables other optimizations like coarsening and hierarchical aggregation of partial results. We also present an analytical model that considers input data characteristics and available GPU resources to estimate near-optimal settings for kernel launch parameters. The proposed approach provides speedups ranging from 2 to 67 for different instances of the generic pattern compared to launching multiple operator-level kernels using GPU accelerated libraries. We conclude by demonstrating the effectiveness of the approach in improving end-to-end performance on an entire ML algorithm.
Arash Ashari, Shirish Tatikonda, Matthias Boehm 0001, Berthold Reinwald, Keith Campbell, John Keenleyside, P. Sadayappan
PPoPP1
2015 A model-driven blocking strategy for load balanced sparse matrix-vector multiplication on GPUs
Arash Ashari, Naser Sedaghati, John Eisenlohr, P. Sadayappan
J. Parallel Distributed Comput.1
2014 An efficient two-dimensional blocking strategy for sparse matrix-vector multiplication on GPUs
abstract
Sparse matrix-vector multiplication (SpMV) is one of the key operations in linear algebra. Overcoming thread divergence, load imbalance and non-coalesced and indirect memory access due to sparsity and irregularity are challenges to optimizing SpMV on GPUs.
Arash Ashari, Naser Sedaghati, John Eisenlohr, P. Sadayappan
ICS1
2014 Fast Sparse Matrix-Vector Multiplication on GPUs for Graph Applications
abstract
Sparse matrix-vector multiplication (SpMV) is a widely used computational kernel. The most commonly used format for a sparse matrix is CSR (Compressed Sparse Row), but a number of other representations have recently been developed that achieve higher SpMV performance. However, the alternative representations typically impose a significant preprocessing overhead. While a high preprocessing overhead can be amortized for applications requiring many iterative invocations of SpMV that use the same matrix, it is not always feasible -- for instance when analyzing large dynamically evolving graphs. This paper presents ACSR, an adaptive SpMV algorithm that uses the standard CSR format but reduces thread divergence by combining rows into groups (bins) which have a similar number of non-zero elements. Further, for rows in bins that span a wide range of non zero counts, dynamic parallelism is leveraged. A significant benefit of ACSR over other proposed SpMV approaches is that it works directly with the standard CSR format, and thus avoids significant preprocessing overheads. A CUDA implementation of ACSR is shown to outperform SpMV implementations in the NVIDIA CUSP and cuSPARSE libraries on a set of sparse matrices representing power-law graphs. We also demonstrate the use of ACSR for the analysis of dynamic graphs, where the improvement over extant approaches is even higher.
Arash Ashari, Naser Sedaghati, John Eisenlohr, Srinivasan Parthasarathy 0001, P. Sadayappan
SC1