Anand Venkat

dblp:141/4387 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
0since 2021 · last 2020
0000-0002-4167-4525ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 2 first-authorSoftware engineering, systems software and programming languages · 3 · 2 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Software engineering, system software, and programming languages
3 papers
Compilers and program optimization · 100%
Computer architecture, parallel and distributed computing, and storage systems
3 papers
Parallel and multicore computing · 73% High-performance computing · 27%

Topics — the 11 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Compilers and program optimization
dependence analysis
0.622019
Sparse computation data dependence simplification for efficient compiler-generated inspectors · PLDI 2019
Automating wavefront parallelization for sparse matrix computations · SC 2016
Parallel and multicore computing › parallel algorithms › parallel algorithm design
wavefront parallelism
0.622019
Sparse computation data dependence simplification for efficient compiler-generated inspectors · PLDI 2019
Automating wavefront parallelization for sparse matrix computations · SC 2016
Compilers and program optimization › loop transformation
polyhedral compilation
0.522016
Automating wavefront parallelization for sparse matrix computations · SC 2016
Loop and data transformations for sparse matrix code · PLDI 2015
Parallel and multicore computing › parallel programming models
automatic parallelization
0.412019
Sparse computation data dependence simplification for efficient compiler-generated inspectors · PLDI 2019
High-performance computing › sparse linear algebra
sparse matrix computation
0.322016
Automating wavefront parallelization for sparse matrix computations · SC 2016
Loop and data transformations for sparse matrix code · PLDI 2015
Compilers and program optimization › parallelization
inspector-executor
0.212016
Automating wavefront parallelization for sparse matrix computations · SC 2016
Parallel and multicore computing
parallelizing compiler
0.212016
Automating wavefront parallelization for sparse matrix computations · SC 2016
Compilers and program optimization › memory optimization
data layout transformation
0.212015
Loop and data transformations for sparse matrix code · PLDI 2015
Compilers and program optimization
loop transformation
0.212015
Loop and data transformations for sparse matrix code · PLDI 2015
High-performance computing › iterative methods
preconditioned conjugate gradient
0.112016
Automating wavefront parallelization for sparse matrix computations · SC 2016
High-performance computing
sparse linear algebra
0.112016
Automating wavefront parallelization for sparse matrix computations · SC 2016

Methods — techniques the papers use, named apart from their topics

run-time dependence testing · 0.8compile-time analysis · 0.8runtime dependence inspection · 0.5polyhedral compilation · 0.5runtime inspection · 0.4polyhedral transformation · 0.4
YearPublicationVenuePosition
2020 Harnessing Deep Learning via a Single Building Block
abstract
Deep learning (DL) is one of the most prominent branches of machine learning. Due to the immense computational cost of DL workloads, industry and academia have developed DL libraries with highly-specialized kernels for each workload/architecture, leading to numerous, complex code-bases that strive for performance, yet they are hard to maintain and do not generalize. In this work, we introduce the batch-reduce GEMM kernel and show how the most popular DL algorithms can be formulated with this kernel as the basic building-block. Consequently, the DL library-development degenerates to mere (potentially automatic) tuning of loops around this sole optimized kernel. By exploiting our new kernel we implement Recurrent Neural Networks, Convolution Neural Networks and Multilayer Perceptron training and inference primitives in just 3K lines of high-level code. Our primitives outperform vendor-optimized libraries on multi-node CPU clusters, and we also provide proof-of-concept CNN kernels targeting GPUs. Finally, we demonstrate that the batch-reduce GEMM kernel within a tensor compiler yields high-performance CNN primitives, further amplifying the viability of our approach.
Evangelos Georganas, Kunal Banerjee 0001, Dhiraj D. Kalamkar, Sasikanth Avancha, Anand Venkat, Michael J. Anderson, Greg Henry, Hans Pabst, Alexander Heinecke
IPDPS5
2019 ISA mapper: a compute and hardware agnostic deep learning compiler
abstract
Domain specific accelerators present new challenges for code generation onto novel instruction sets, communication fabrics, and memory architectures. We introduce a shared intermediate representation to describe both deep learning programs and hardware capabilities, then formulate and apply instruction mapping to determine how a computation can be performed on a hardware system. Our scheduler chooses a specific mapping and determines data movement and computation order.
Matthew Sotoudeh, Anand Venkat, Michael J. Anderson, Evangelos Georganas, Alexander Heinecke, Jason Knight
CF2
2019 Sparse computation data dependence simplification for efficient compiler-generated inspectors
abstract
This paper presents a combined compile-time and runtime loop-carried dependence analysis of sparse matrix codes and evaluates its performance in the context of wavefront parallellism. Sparse computations incorporate indirect memory accesses such as x[col[j]] whose memory locations cannot be determined until runtime. The key contributions of this paper are two compile-time techniques for significantly reducing the overhead of runtime dependence testing: (1) identifying new equality constraints that result in more efficient runtime inspectors, and (2) identifying subset relations between dependence constraints such that one dependence test subsumes another one that is therefore eliminated. New equality constraints discovery is enabled by taking advantage of domain-specific knowledge about index arrays, such as col[j]. These simplifications lead to automatically-generated inspectors that make it practical to parallelize such computations. We analyze our simplification methods for a collection of seven sparse computations. The evaluation shows our methods reduce the complexity of the runtime inspectors significantly. Experimental results for a collection of five large matrices show parallel speedups ranging from 2x to more than 8x running on a 8-core CPU.
Mahdi Soltan Mohammadi, Tomofumi Yuki, Kazem Cheshmi, Eddie C. Davis, Mary W. Hall, Maryam Mehri Dehnavi, Payal Nandy, Catherine Mills Olschanowsky, Anand Venkat, Michelle Mills Strout
PLDI9
2016 Synchronization Trade-Offs in GPU Implementations of Graph Algorithms
abstract
Although there is an extensive literature on GPU implementations of graph algorithms, we do not yet have a clear understanding of how implementation choices impact performance. As a step towards this goal, we studied how the choice of synchronization mechanism affects the end-to-end performance of complex graph algorithms, using stochastic gradient descent (SGD) as an exemplar. We implemented seven synchronization strategies for this application and evaluated them on two GPU platforms, using both road networks and social network graphs as inputs. Our experiments showed that although none of the seven strategies dominates the rest, it is possible to use properties of the platform and input graph to predict the best strategy.
Rashid Kaleem, Anand Venkat, Sreepathi Pai, Mary W. Hall, Keshav Pingali
IPDPS2
2016 Automating wavefront parallelization for sparse matrix computations
abstract
This paper presents a compiler and runtime framework for parallelizing sparse matrix computations that have loop-carried dependences. Our approach automatically generates a runtime inspector to collect data dependence information and achieves wavefront parallelization of the computation, where iterations within a wavefront execute in parallel, and synchronization is required across wavefronts. A key contribution of this paper involves dependence simplification, which reduces the time and space overhead of the inspector. This is implemented within a polyhedral compiler framework, extended for sparse matrix codes. Results demonstrate the feasibility of using automatically-generated inspectors and executors to optimize ILU factorization and symmetric Gauss-Seidel relaxations, which are part of the Preconditioned Conjugate Gradient (PCG) computation. Our implementation achieves a median speedup of 2.97× on 12 cores over the reference sequential PCG implementation, significantly outperforms PCG parallelized using Intel's Math Kernel Library (MKL), and is within 6% of the median performance of manually-parallelized PCG.
Anand Venkat, Mahdi Soltan Mohammadi, Jongsoo Park, Hongbo Rong, Rajkishore Barik, Michelle Mills Strout, Mary W. Hall
SC1
2015 Loop and data transformations for sparse matrix code
abstract
This paper introduces three new compiler transformations for representing and transforming sparse matrix computations and their data representations. In cooperation with run-time inspection, our compiler derives transformed matrix representations and associated transformed code to implement a variety of representations targeting different architecture platforms. This systematic approach to combining code and data transformations on sparse computations, which extends a polyhedral transformation and code generation framework, permits the compiler to compose these transformations with other transformations to generate code that is on average within 5% and often exceeds manually-tuned, high-performance sparse matrix libraries CUSP and OSKI. Additionally, the compiler-generated inspector codes are on average 1.5 faster than OSKI and perform comparably to CUSP, respectively.
Anand Venkat, Mary W. Hall, Michelle Mills Strout
PLDI1
2014 Non-affine Extensions to Polyhedral Code Generation
Anand Venkat, Manu Shantharam, Mary W. Hall, Michelle Mills Strout
CGO1
2013 Compiler generation and autotuning of communication-avoiding operators for geometric multigrid
abstract
This paper describes a compiler approach to introducing communication-avoiding optimizations in geometric multigrid (GMG), one of the most popular methods for solving partial differential equations. Communication-avoiding optimizations reduce vertical communication through the memory hierarchy and horizontal communication across processes or threads, usually at the expense of introducing redundant computation. We focus on applying these optimizations to the smooth operator, which successively reduces the error and accounts for the largest fraction of the GMG execution time. Our compiler technology applies both novel and known transformations to derive an implementation comparable to manually-tuned code. To make the approach portable, an underlying autotuning system explores the tradeoff between reduced communication and increased computation, as well as tradeoffs in threading schemes, to automatically identify the best implementation for a particular architecture and at each computation phase. Results show that we are able to quadruple the performance of the smooth operation on the finest grids while attaining performance within 94% of manually-tuned code. Overall we improve the overall multigrid solve time by 2.5× without sacrificing programer productivity.
Protonu Basu, Anand Venkat, Mary W. Hall, Samuel Williams 0001, Brian van Straalen, Leonid Oliker
HiPC2