EDBT 2026 Demo / reviewers in the wild / expert
Anand Venkat
dblp:141/4387
· DBLP profile ↗
8ranked-venue papers
3as first author
0since 2021 · last 2020
0000-0002-4167-4525ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 2 first-authorSoftware engineering, systems software and programming languages · 3 · 2 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Software engineering, system software, and programming languages
3 papers |
Compilers and program optimization · 100% | |
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Parallel and multicore computing · 73% High-performance computing · 27% |
Topics — the 11 heaviest of 12, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Compilers and program optimization
dependence analysis |
0.6 | 2 | 2019 | Sparse computation data dependence simplification for efficient compiler-generated inspectors · PLDI 2019 Automating wavefront parallelization for sparse matrix computations · SC 2016 |
Parallel and multicore computing › parallel algorithms › parallel algorithm design
wavefront parallelism |
0.6 | 2 | 2019 | Sparse computation data dependence simplification for efficient compiler-generated inspectors · PLDI 2019 Automating wavefront parallelization for sparse matrix computations · SC 2016 |
Compilers and program optimization › loop transformation
polyhedral compilation |
0.5 | 2 | 2016 | Automating wavefront parallelization for sparse matrix computations · SC 2016 Loop and data transformations for sparse matrix code · PLDI 2015 |
Parallel and multicore computing › parallel programming models
automatic parallelization |
0.4 | 1 | 2019 | Sparse computation data dependence simplification for efficient compiler-generated inspectors · PLDI 2019 |
High-performance computing › sparse linear algebra
sparse matrix computation |
0.3 | 2 | 2016 | Automating wavefront parallelization for sparse matrix computations · SC 2016 Loop and data transformations for sparse matrix code · PLDI 2015 |
Compilers and program optimization › parallelization
inspector-executor |
0.2 | 1 | 2016 | Automating wavefront parallelization for sparse matrix computations · SC 2016 |
Parallel and multicore computing
parallelizing compiler |
0.2 | 1 | 2016 | Automating wavefront parallelization for sparse matrix computations · SC 2016 |
Compilers and program optimization › memory optimization
data layout transformation |
0.2 | 1 | 2015 | Loop and data transformations for sparse matrix code · PLDI 2015 |
Compilers and program optimization
loop transformation |
0.2 | 1 | 2015 | Loop and data transformations for sparse matrix code · PLDI 2015 |
High-performance computing › iterative methods
preconditioned conjugate gradient |
0.1 | 1 | 2016 | Automating wavefront parallelization for sparse matrix computations · SC 2016 |
High-performance computing
sparse linear algebra |
0.1 | 1 | 2016 | Automating wavefront parallelization for sparse matrix computations · SC 2016 |
Methods — techniques the papers use, named apart from their topics
run-time dependence testing · 0.8compile-time analysis · 0.8runtime dependence inspection · 0.5polyhedral compilation · 0.5runtime inspection · 0.4polyhedral transformation · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2020 | Harnessing Deep Learning via a Single Building BlockabstractDeep learning (DL) is one of the most prominent branches of machine learning. Due to the immense computational cost of DL workloads, industry and academia have developed DL libraries with highly-specialized kernels for each workload/architecture, leading to numerous, complex code-bases that strive for performance, yet they are hard to maintain and do not generalize. In this work, we introduce the batch-reduce GEMM kernel and show how the most popular DL algorithms can be formulated with this kernel as the basic building-block. Consequently, the DL library-development degenerates to mere (potentially automatic) tuning of loops around this sole optimized kernel. By exploiting our new kernel we implement Recurrent Neural Networks, Convolution Neural Networks and Multilayer Perceptron training and inference primitives in just 3K lines of high-level code. Our primitives outperform vendor-optimized libraries on multi-node CPU clusters, and we also provide proof-of-concept CNN kernels targeting GPUs. Finally, we demonstrate that the batch-reduce GEMM kernel within a tensor compiler yields high-performance CNN primitives, further amplifying the viability of our approach. Evangelos Georganas, Kunal Banerjee 0001, Dhiraj D. Kalamkar, Sasikanth Avancha, Anand Venkat, Michael J. Anderson, Greg Henry, Hans Pabst, Alexander Heinecke |
IPDPS | 5 |
| 2019 | ISA mapper: a compute and hardware agnostic deep learning compilerabstractDomain specific accelerators present new challenges for code generation onto novel instruction sets, communication fabrics, and memory architectures. We introduce a shared intermediate representation to describe both deep learning programs and hardware capabilities, then formulate and apply instruction mapping to determine how a computation can be performed on a hardware system. Our scheduler chooses a specific mapping and determines data movement and computation order. Matthew Sotoudeh, Anand Venkat, Michael J. Anderson, Evangelos Georganas, Alexander Heinecke, Jason Knight |
CF | 2 |
| 2019 | Sparse computation data dependence simplification for efficient compiler-generated inspectorsabstractThis paper presents a combined compile-time and runtime loop-carried dependence analysis of sparse matrix codes and evaluates its performance in the context of wavefront parallellism. Sparse computations incorporate indirect memory accesses such as x[col[j]] whose memory locations cannot be determined until runtime. The key contributions of this paper are two compile-time techniques for significantly reducing the overhead of runtime dependence testing: (1) identifying new equality constraints that result in more efficient runtime inspectors, and (2) identifying subset relations between dependence constraints such that one dependence test subsumes another one that is therefore eliminated. New equality constraints discovery is enabled by taking advantage of domain-specific knowledge about index arrays, such as col[j]. These simplifications lead to automatically-generated inspectors that make it practical to parallelize such computations. We analyze our simplification methods for a collection of seven sparse computations. The evaluation shows our methods reduce the complexity of the runtime inspectors significantly. Experimental results for a collection of five large matrices show parallel speedups ranging from 2x to more than 8x running on a 8-core CPU. Mahdi Soltan Mohammadi, Tomofumi Yuki, Kazem Cheshmi, Eddie C. Davis, Mary W. Hall, Maryam Mehri Dehnavi, Payal Nandy, Catherine Mills Olschanowsky, Anand Venkat, Michelle Mills Strout |
PLDI | 9 |
| 2016 | Synchronization Trade-Offs in GPU Implementations of Graph AlgorithmsabstractAlthough there is an extensive literature on GPU implementations of graph algorithms, we do not yet have a clear understanding of how implementation choices impact performance. As a step towards this goal, we studied how the choice of synchronization mechanism affects the end-to-end performance of complex graph algorithms, using stochastic gradient descent (SGD) as an exemplar. We implemented seven synchronization strategies for this application and evaluated them on two GPU platforms, using both road networks and social network graphs as inputs. Our experiments showed that although none of the seven strategies dominates the rest, it is possible to use properties of the platform and input graph to predict the best strategy. Rashid Kaleem, Anand Venkat, Sreepathi Pai, Mary W. Hall, Keshav Pingali |
IPDPS | 2 |
| 2016 | Automating wavefront parallelization for sparse matrix computationsabstractThis paper presents a compiler and runtime framework for parallelizing sparse matrix computations that have loop-carried dependences. Our approach automatically generates a runtime inspector to collect data dependence information and achieves wavefront parallelization of the computation, where iterations within a wavefront execute in parallel, and synchronization is required across wavefronts. A key contribution of this paper involves dependence simplification, which reduces the time and space overhead of the inspector. This is implemented within a polyhedral compiler framework, extended for sparse matrix codes. Results demonstrate the feasibility of using automatically-generated inspectors and executors to optimize ILU factorization and symmetric Gauss-Seidel relaxations, which are part of the Preconditioned Conjugate Gradient (PCG) computation. Our implementation achieves a median speedup of 2.97× on 12 cores over the reference sequential PCG implementation, significantly outperforms PCG parallelized using Intel's Math Kernel Library (MKL), and is within 6% of the median performance of manually-parallelized PCG. Anand Venkat, Mahdi Soltan Mohammadi, Jongsoo Park, Hongbo Rong, Rajkishore Barik, Michelle Mills Strout, Mary W. Hall |
SC | 1 |
| 2015 | Loop and data transformations for sparse matrix codeabstractThis paper introduces three new compiler transformations for representing and transforming sparse matrix computations and their data representations. In cooperation with run-time inspection, our compiler derives transformed matrix representations and associated transformed code to implement a variety of representations targeting different architecture platforms. This systematic approach to combining code and data transformations on sparse computations, which extends a polyhedral transformation and code generation framework, permits the compiler to compose these transformations with other transformations to generate code that is on average within 5% and often exceeds manually-tuned, high-performance sparse matrix libraries CUSP and OSKI. Additionally, the compiler-generated inspector codes are on average 1.5 faster than OSKI and perform comparably to CUSP, respectively. Anand Venkat, Mary W. Hall, Michelle Mills Strout |
PLDI | 1 |
| 2014 | Non-affine Extensions to Polyhedral Code Generation
Anand Venkat, Manu Shantharam, Mary W. Hall, Michelle Mills Strout |
CGO | 1 |
| 2013 | Compiler generation and autotuning of communication-avoiding operators for geometric multigridabstractThis paper describes a compiler approach to introducing communication-avoiding optimizations in geometric multigrid (GMG), one of the most popular methods for solving partial differential equations. Communication-avoiding optimizations reduce vertical communication through the memory hierarchy and horizontal communication across processes or threads, usually at the expense of introducing redundant computation. We focus on applying these optimizations to the smooth operator, which successively reduces the error and accounts for the largest fraction of the GMG execution time. Our compiler technology applies both novel and known transformations to derive an implementation comparable to manually-tuned code. To make the approach portable, an underlying autotuning system explores the tradeoff between reduced communication and increased computation, as well as tradeoffs in threading schemes, to automatically identify the best implementation for a particular architecture and at each computation phase. Results show that we are able to quadruple the performance of the smooth operation on the finest grids while attaining performance within 94% of manually-tuned code. Overall we improve the overall multigrid solve time by 2.5× without sacrificing programer productivity. Protonu Basu, Anand Venkat, Mary W. Hall, Samuel Williams 0001, Brian van Straalen, Leonid Oliker |
HiPC | 2 |