VLDB 2026 Research / reviewers in the wild / expert
Venmugil Elango
dblp:140/7204
· DBLP profile ↗
8ranked-venue papers
4as first author
2since 2021 · last 2023
0000-0002-7031-9020ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 3 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
5 papers |
Hardware accelerators and domain-specific architectures · 52% Performance modeling and evaluation · 32% Memory systems · 11% | |
| Software engineering, system software, and programming languages
2 papers |
Compilers and program optimization · 100% | |
| Artificial intelligence
1 paper |
Deep learning architectures and training · 100% |
Topics — the 14 heaviest of 17, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Hardware accelerators and domain-specific architectures › machine learning accelerator
low-precision arithmetic |
0.7 | 1 | 2023 | With Shared Microexponents, A Little Shifting Goes a Long Way · ISCA 2023 |
Hardware accelerators and domain-specific architectures
machine learning accelerator |
0.7 | 1 | 2023 | With Shared Microexponents, A Little Shifting Goes a Long Way · ISCA 2023 |
Compilers and program optimization › code generation › parallel code generation
distributed-memory code generation |
0.2 | 1 | 2015 | Distributed memory code generation for mixed Irregular/Regular computations · PPoPP 2015 |
Compilers and program optimization
parallelizing compiler |
0.2 | 1 | 2015 | Distributed memory code generation for mixed Irregular/Regular computations · PPoPP 2015 |
Compilers and program optimization › loop transformation
polyhedral compilation |
0.2 | 1 | 2015 | Distributed memory code generation for mixed Irregular/Regular computations · PPoPP 2015 |
Performance modeling and evaluation
bottleneck analysis |
0.2 | 1 | 2014 | On Using the Roofline Model with Lower Bounds on Data Movement · ACM Trans. Archit. Code Optim. 2014 |
Performance modeling and evaluation › analytical modeling
roofline model |
0.2 | 1 | 2014 | On Using the Roofline Model with Lower Bounds on Data Movement · ACM Trans. Archit. Code Optim. 2014 |
Compilers and program optimization › memory optimization
data locality optimization |
0.2 | 1 | 2013 | Beyond reuse distance analysis: Dynamic analysis for characterization of data locality potential · ACM Trans. Archit. Code Optim. 2013 |
Memory systems
data locality |
0.2 | 1 | 2013 | Beyond reuse distance analysis: Dynamic analysis for characterization of data locality potential · ACM Trans. Archit. Code Optim. 2013 |
Performance modeling and evaluation › cache performance modeling
reuse distance analysis |
0.2 | 1 | 2013 | Beyond reuse distance analysis: Dynamic analysis for characterization of data locality potential · ACM Trans. Archit. Code Optim. 2013 |
High-performance computing › scientific computing systems
adaptive mesh refinement |
0.1 | 1 | 2015 | Distributed memory code generation for mixed Irregular/Regular computations · PPoPP 2015 |
Memory systems
cache |
0.1 | 1 | 2015 | On Characterizing the Data Access Complexity of Programs · POPL 2015 |
High-performance computing
scientific computing systems |
0.1 | 1 | 2015 | Distributed memory code generation for mixed Irregular/Regular computations · PPoPP 2015 |
Performance modeling and evaluation
workload characterization |
0.0 | 1 | 2013 | Beyond reuse distance analysis: Dynamic analysis for characterization of data locality potential · ACM Trans. Archit. Code Optim. 2013 |
Methods — techniques the papers use, named apart from their topics
quantization · 1.3block floating point · 1.3polyhedral framework · 0.4loop transformation · 0.4static analysis · 0.2red-blue pebble game · 0.2graph decomposition · 0.2computational directed acyclic graphs · 0.2lower bounds on data movement · 0.2cache capacity analysis · 0.2dynamic analysis · 0.2convex partitioning · 0.2CDAG · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | With Shared Microexponents, A Little Shifting Goes a Long WayabstractThis paper introduces Block Data Representations (BDR), a framework for exploring and evaluating a wide spectrum of narrow-precision formats for deep learning. It enables comparison of popular quantization standards, and through BDR, new formats based on shared microexponents (MX) are identified, which outperform other state-of-the-art quantization approaches, including narrow-precision floating-point and block floating-point. MX utilizes multiple levels of quantization scaling with ultra-fine scaling factors based on shared microexponents in the hardware. The effectiveness of MX is demonstrated on real-world models including large-scale generative pretraining and inferencing, and production-scale recommendation systems. Bita Darvish Rouhani, Ritchie Zhao, Venmugil Elango, Rasoul Shafipour, Mathew Hall, Maral Mesmakhosroshahi, Ankit More, Levi Melnick, Maximilian Golub, Girish Varatkar, Lai Shao, Gaurav Kolhe, Dimitry Melts, Jasmine Klar, Renee L'Heureux, Matt Perry, Doug Burger, Eric S. Chung, Zhaoxia Deng, Sam Naghshineh, Jongsoo Park, Maxim Naumov |
ISCA | 3 |
| 2021 | Pase: Parallelization Strategies for Efficient DNN TrainingabstractTraining a deep neural network (DNN) requires substantial computational and memory requirements. It is common to use multiple devices to train a DNN to reduce the overall training time. There are several choices to parallelize each layer in a DNN. Exhaustively searching this list to find an optimal parallelization strategy is prohibitively time consuming and impractical. The standard practice is to use data parallelism because of its simplicity. However, data parallelism is often suboptimal, and suffers from poor performance and high memory requirement. Expert-designed strategies have been proposed on a case-by-case basis using domain specific knowledge. These expert-designed strategies do not generalize well to DNNs other than the ones for which they were designed, and are not always necessarily the best choice. In this paper, we propose an approach to automatically find efficient parallelization strategies for DNNs from their computation graphs. We present an efficient algorithm to compute these strategies within a reasonable time in practice. We evaluate the effectiveness of our approach on various DNNs. We also compare the performance of the strategies identified by our approach against data parallelism, expert-designed strategies, and the state-of-the-art approaches. Our results show that the strategies found using our approach outperform the baseline data parallelism strategy in all the cases. In addition, our strategies achieve better performance than the expert-designed strategies and the state-of-the-art approaches. Venmugil Elango |
IPDPS | 1 |
| 2015 | On Characterizing the Data Access Complexity of ProgramsabstractTechnology trends will cause data movement to account for the majority of energy expenditure and execution time on emerging computers. Therefore, computational complexity will no longer be a sufficient metric for comparing algorithms, and a fundamental characterization of data access complexity will be increasingly important. The problem of developing lower bounds for data access complexity has been modeled using the formalism of Hong and Kung's red/blue pebble game for computational directed acyclic graphs (CDAGs). However, previously developed approaches to lower bounds analysis for the red/blue pebble game are very limited in effectiveness when applied to CDAGs of real programs, with computations comprised of multiple sub-computations with differing DAG structure. We address this problem by developing an approach for effectively composing lower bounds based on graph decomposition. We also develop a static analysis algorithm to derive the asymptotic data-access lower bounds of programs, as a function of the problem size and cache size. Venmugil Elango, Fabrice Rastello, Louis-Noël Pouchet, J. Ramanujam, P. Sadayappan |
POPL | 1 |
| 2015 | Distributed memory code generation for mixed Irregular/Regular computationsabstractMany applications feature a mix of irregular and regular computational structures. For example, codes using adaptive mesh refinement (AMR) typically use a collection of regular blocks, where the number of blocks and the relationship between blocks is irregular. The computational structure in such applications generally involves regular (affine) loop computations within some number of innermost loops, while outer loops exhibit irregularity due to data-dependent control flow and indirect array access patterns. Prior approaches to distributed memory parallelization do not handle such computations effectively. They either target loop nests that are completely affine using polyhedral frameworks, or treat all loops as irregular. Consequently, the generated distributed memory code contains artifacts that disrupt the regular nature of previously affine innermost loops of the computation. This hampers subsequent optimizations to improve on-node performance. We propose a code generation framework that can effectively transform such applications for execution on distributed memory systems. Our approach generates distributed memory code which preserves program properties that enable subsequent polyhederal optimizations. Simultaneously, it addresses a major memory bottleneck of prior techniques that limits the scalability of the generated code. The effectiveness of the proposed framework is demonstrated on computations that are mixed regular/irregular, completely regular, and completely irregular. Mahesh Ravishankar, Roshan Dathathri, Venmugil Elango, Louis-Noël Pouchet, J. Ramanujam, Atanas Rountev, P. Sadayappan |
PPoPP | 3 |
| 2014 | On characterizing the data movement complexity of computational DAGs for parallel executionabstractTechnology trends are making the cost of data movement increasingly dominant, both in terms of energy and time, over the cost of performing arithmetic operations in computer systems. The fundamental ratio of aggregate data movement bandwidth to the total computational power (also referred to the machine balance parameter) in parallel computer systems is decreasing. It is therefore of considerable importance to characterize the inherent data movement requirements of parallel algorithms, so that the minimal architectural balance parameters required to support it on future systems can be well understood. Venmugil Elango, Fabrice Rastello, Louis-Noël Pouchet, J. Ramanujam, P. Sadayappan |
SPAA | 1 |
| 2014 | On Using the Roofline Model with Lower Bounds on Data MovementabstractThe roofline model is a popular approach for “bound and bottleneck” performance analysis. It focuses on the limits to the performance of processors because of limited bandwidth to off-chip memory. It models upper bounds on performance as a function of operational intensity, the ratio of computational operations per byte of data moved from/to memory. While operational intensity can be directly measured for a specific implementation of an algorithm on a particular target platform, it is of interest to obtain broader insights on bottlenecks, where various semantically equivalent implementations of an algorithm are considered, along with analysis for variations in architectural parameters. This is currently very cumbersome and requires performance modeling and analysis of many variants. In this article, we address this problem by using the roofline model in conjunction with upper bounds on the operational intensity of computations as a function of cache capacity, derived from lower bounds on data movement. This enables bottleneck analysis that holds across all dependence-preserving semantically equivalent implementations of an algorithm. We demonstrate the utility of the approach in assessing fundamental limits to performance and energy efficiency for several benchmark algorithms across a design space of architectural variations. Venmugil Elango, Naser Sedaghati, Fabrice Rastello, Louis-Noël Pouchet, J. Ramanujam, Radu Teodorescu, P. Sadayappan |
ACM Trans. Archit. Code Optim. | 1 |
| 2013 | Accelerating Strassen-Winograd's matrix multiplication algorithm on GPUsabstractIn this paper, we report on the development of an efficient GPU implementation of the Strassen-Winograd matrix multiplication algorithm for matrices of arbitrary sizes. We utilize multi-kernel streaming to exploit concurrency across sub-matrix operations in addition to intra-operation parallelism. We evaluate the performance of the implementation in comparison with CUBLAS-5.0 on Fermi and Kepler GPUs. The experimental results demonstrate the usefulness of Strassen's algorithm for practically relevant matrix sizes on GPUs, with up to 1.27X speedup for single-precision and 1.42X speedup for double-precision floating point computation. Pai-Wei Lai, Humayun Arafat, Venmugil Elango, P. Sadayappan |
HiPC | 3 |
| 2013 | Beyond reuse distance analysis: Dynamic analysis for characterization of data locality potentialabstractEmerging computer architectures will feature drastically decreased flops/byte (ratio of peak processing rate to memory bandwidth) as highlighted by recent studies on Exascale architectural trends. Further, flops are getting cheaper, while the energy cost of data movement is increasingly dominant. The understanding and characterization of data locality properties of computations is critical in order to guide efforts to enhance data locality. Reuse distance analysis of memory address traces is a valuable tool to perform data locality characterization of programs. A single reuse distance analysis can be used to estimate the number of cache misses in a fully associative LRU cache of any size, thereby providing estimates on the minimum bandwidth requirements at different levels of the memory hierarchy to avoid being bandwidth bound. However, such an analysis only holds for the particular execution order that produced the trace. It cannot estimate potential improvement in data locality through dependence-preserving transformations that change the execution schedule of the operations in the computation. In this article, we develop a novel dynamic analysis approach to characterize the inherent locality properties of a computation and thereby assess the potential for data locality enhancement via dependence-preserving transformations. The execution trace of a code is analyzed to extract a Computational-Directed Acyclic Graph (CDAG) of the data dependences. The CDAG is then partitioned into convex subsets, and the convex partitioning is used to reorder the operations in the execution trace to enhance data locality. The approach enables us to go beyond reuse distance analysis of a single specific order of execution of the operations of a computation in characterization of its data locality properties. It can serve a valuable role in identifying promising code regions for manual transformation, as well as assessing the effectiveness of compiler transformations for data locality enhancement. We demonstrate the effectiveness of the approach using a number of benchmarks, including case studies where the potential shown by the analysis is exploited to achieve lower data movement costs and better performance. Naznin Fauzia, Venmugil Elango, Mahesh Ravishankar, J. Ramanujam, Fabrice Rastello, Atanas Rountev, Louis-Noël Pouchet, P. Sadayappan |
ACM Trans. Archit. Code Optim. | 2 |