Mahdi Soltan Mohammadi

dblp:189/1211 · DBLP profile ↗
← Back
2ranked-venue papers
1as first author
0since 2021 · last 2019
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 1Software engineering, systems software and programming languages · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Parallel and multicore computing · 76% High-performance computing · 24%
Software engineering, system software, and programming languages
2 papers
Compilers and program optimization · 100%

Topics — the 9 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Compilers and program optimization
dependence analysis
0.622019
Sparse computation data dependence simplification for efficient compiler-generated inspectors · PLDI 2019
Automating wavefront parallelization for sparse matrix computations · SC 2016
Parallel and multicore computing › parallel algorithms › parallel algorithm design
wavefront parallelism
0.622019
Sparse computation data dependence simplification for efficient compiler-generated inspectors · PLDI 2019
Automating wavefront parallelization for sparse matrix computations · SC 2016
Parallel and multicore computing › parallel programming models
automatic parallelization
0.412019
Sparse computation data dependence simplification for efficient compiler-generated inspectors · PLDI 2019
Compilers and program optimization › parallelization
inspector-executor
0.212016
Automating wavefront parallelization for sparse matrix computations · SC 2016
Compilers and program optimization › loop transformation
polyhedral compilation
0.212016
Automating wavefront parallelization for sparse matrix computations · SC 2016
Parallel and multicore computing
parallelizing compiler
0.212016
Automating wavefront parallelization for sparse matrix computations · SC 2016
High-performance computing › sparse linear algebra
sparse matrix computation
0.212016
Automating wavefront parallelization for sparse matrix computations · SC 2016
High-performance computing › iterative methods
preconditioned conjugate gradient
0.112016
Automating wavefront parallelization for sparse matrix computations · SC 2016
High-performance computing
sparse linear algebra
0.112016
Automating wavefront parallelization for sparse matrix computations · SC 2016

Methods — techniques the papers use, named apart from their topics

run-time dependence testing · 0.8compile-time analysis · 0.8runtime dependence inspection · 0.5polyhedral compilation · 0.5
YearPublicationVenuePosition
2019 Sparse computation data dependence simplification for efficient compiler-generated inspectors
abstract
This paper presents a combined compile-time and runtime loop-carried dependence analysis of sparse matrix codes and evaluates its performance in the context of wavefront parallellism. Sparse computations incorporate indirect memory accesses such as x[col[j]] whose memory locations cannot be determined until runtime. The key contributions of this paper are two compile-time techniques for significantly reducing the overhead of runtime dependence testing: (1) identifying new equality constraints that result in more efficient runtime inspectors, and (2) identifying subset relations between dependence constraints such that one dependence test subsumes another one that is therefore eliminated. New equality constraints discovery is enabled by taking advantage of domain-specific knowledge about index arrays, such as col[j]. These simplifications lead to automatically-generated inspectors that make it practical to parallelize such computations. We analyze our simplification methods for a collection of seven sparse computations. The evaluation shows our methods reduce the complexity of the runtime inspectors significantly. Experimental results for a collection of five large matrices show parallel speedups ranging from 2x to more than 8x running on a 8-core CPU.
Mahdi Soltan Mohammadi, Tomofumi Yuki, Kazem Cheshmi, Eddie C. Davis, Mary W. Hall, Maryam Mehri Dehnavi, Payal Nandy, Catherine Mills Olschanowsky, Anand Venkat, Michelle Mills Strout
PLDI1
2016 Automating wavefront parallelization for sparse matrix computations
abstract
This paper presents a compiler and runtime framework for parallelizing sparse matrix computations that have loop-carried dependences. Our approach automatically generates a runtime inspector to collect data dependence information and achieves wavefront parallelization of the computation, where iterations within a wavefront execute in parallel, and synchronization is required across wavefronts. A key contribution of this paper involves dependence simplification, which reduces the time and space overhead of the inspector. This is implemented within a polyhedral compiler framework, extended for sparse matrix codes. Results demonstrate the feasibility of using automatically-generated inspectors and executors to optimize ILU factorization and symmetric Gauss-Seidel relaxations, which are part of the Preconditioned Conjugate Gradient (PCG) computation. Our implementation achieves a median speedup of 2.97× on 12 cores over the reference sequential PCG implementation, significantly outperforms PCG parallelized using Intel's Math Kernel Library (MKL), and is within 6% of the median performance of manually-parallelized PCG.
Anand Venkat, Mahdi Soltan Mohammadi, Jongsoo Park, Hongbo Rong, Rajkishore Barik, Michelle Mills Strout, Mary W. Hall
SC2