VLDB 2026 Research / reviewers in the wild / expert
Prashant Singh Rawat
dblp:19/10672
· DBLP profile ↗
10ranked-venue papers
6as first author
0since 2021 · last 2019
0000-0003-3189-3983ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 5 first-authorSoftware engineering, systems software and programming languages · 2Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
6 papers |
GPUs and heterogeneous computing · 32% Processor architecture and microarchitecture · 26% Performance modeling and evaluation · 20% | |
| Software engineering, system software, and programming languages
2 papers |
Compilers and program optimization · 75% Program synthesis and code generation · 25% |
Topics — the 16 heaviest of 17, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Hardware accelerators and domain-specific architectures › sparsity exploitation
sparse tensor computation |
0.4 | 1 | 2019 | An efficient mixed-mode representation of sparse tensors · SC 2019 |
Program synthesis and code generation
domain-specific code generation |
0.3 | 1 | 2018 | Domain-Specific Optimization and Generation of High-Performance GPU Code for Stencil Computations · Proc. IEEE 2018 |
Compilers and program optimization
register allocation |
0.3 | 1 | 2018 | Register optimizations for stencils on GPUs · PPoPP 2018 |
Compilers and program optimization › register allocation
register pressure reduction |
0.3 | 1 | 2018 | Register optimizations for stencils on GPUs · PPoPP 2018 |
Compilers and program optimization
stencil computation |
0.3 | 1 | 2018 | Domain-Specific Optimization and Generation of High-Performance GPU Code for Stencil Computations · Proc. IEEE 2018 |
Performance modeling and evaluation
bottleneck analysis |
0.3 | 1 | 2018 | GPU code optimization using abstract kernel emulation and sensitivity analysis · PLDI 2018 |
GPUs and heterogeneous computing › GPU performance optimization
GPU code optimization |
0.3 | 1 | 2018 | Domain-Specific Optimization and Generation of High-Performance GPU Code for Stencil Computations · Proc. IEEE 2018 |
GPUs and heterogeneous computing
GPU compilation |
0.3 | 1 | 2018 | Register optimizations for stencils on GPUs · PPoPP 2018 |
GPUs and heterogeneous computing
GPU kernel optimization |
0.3 | 1 | 2018 | GPU code optimization using abstract kernel emulation and sensitivity analysis · PLDI 2018 |
Performance modeling and evaluation › processor performance modeling › accelerator performance modeling
GPU performance modeling |
0.3 | 1 | 2018 | Performance modeling for GPUs using abstract kernel emulation · PPoPP 2018 |
Processor architecture and microarchitecture
instruction reordering |
0.3 | 1 | 2018 | Associative instruction reordering to alleviate register pressure · SC 2018 |
Processor architecture and microarchitecture
instruction scheduling |
0.3 | 1 | 2018 | Associative instruction reordering to alleviate register pressure · SC 2018 |
Processor architecture and microarchitecture › register management
register pressure |
0.3 | 1 | 2018 | Associative instruction reordering to alleviate register pressure · SC 2018 |
High-performance computing › stencil computation
stencil computation optimization |
0.3 | 1 | 2018 | Register optimizations for stencils on GPUs · PPoPP 2018 |
Performance modeling and evaluation › statistical analysis
sensitivity analysis |
0.1 | 1 | 2018 | GPU code optimization using abstract kernel emulation and sensitivity analysis · PLDI 2018 |
High-performance computing
stencil computation |
0.1 | 1 | 2018 | Domain-Specific Optimization and Generation of High-Performance GPU Code for Stencil Computations · Proc. IEEE 2018 |
Methods — techniques the papers use, named apart from their topics
tiling · 0.7statement reordering · 0.7loop unrolling · 0.7domain-specific language · 0.7DAG scheduling · 0.7mixed-mode representation · 0.4compressed sparse fiber · 0.4latency modeling · 0.3gap modeling · 0.3associative instruction reordering · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2019 | On Optimizing Complex Stencils on GPUsabstractStencil computations are often the compute-intensive kernel in many scientific applications. With the increasing demand for computational accuracy, and the emergence of massively data-parallel high-bandwidth architectures like GPUs, stencils have steadily become more complex in terms of the stencil order, data accesses, and reuse patterns. Many prior efforts have focused on optimizing simpler stencil computations on various platforms. However, existing stencil code generators face challenges in optimizing such complex multi-statement stencil DAGs. This paper addresses the challenges in optimizing high-order stencil DAGs on GPUs by focusing on two key considerations: (1) enabling the domain expert to guide the code optimization, which may otherwise be extremely challenging for complex stencils; and (2) using bottleneck analysis via runtime profiling to guide the application of optimizations, and the tuning of various code generation parameters. We implement these abstractions in a prototype code generation framework termed ARTEMIS, and evaluate its efficacy over multiple stencil kernels with varying complexity and operational intensity on an NVIDIA P100 GPU. Prashant Singh Rawat, Miheer Vaidya, Aravind Sukumaran-Rajam, Atanas Rountev, Louis-Noël Pouchet, P. Sadayappan |
IPDPS | 1 |
| 2019 | An efficient mixed-mode representation of sparse tensorsabstractThe Compressed Sparse Fiber (CSF) representation for sparse tensors is a generalization of the Compressed Sparse Row (CSR) format for sparse matrices. For a tensor with d modes, typical tensor methods such as CANDECOMP/PARAFAC decomposition (CPD) require a sequence of d tensor computations, where efficient memory access with respect to different modes is required for each of them. The straightforward solution is to use d distinct representations of the tensor, with each one being efficient for one of the d computations. However, a d-fold space overhead is often unacceptable in practice, especially with memory-constrained GPUs. In this paper, we present a mixed-mode tensor representation that partitions the tensor's nonzero elements into disjoint sections, each of which is compressed to create fibers along a different mode. Experimental results demonstrate that better performance can be achieved while utilizing only a small fraction of the space required to keep d distinct CSF representations. Israt Nisa, Jiajia Li 0001, Aravind Sukumaran-Rajam, Prashant Singh Rawat, Sriram Krishnamoorthy, P. Sadayappan |
SC | 4 |
| 2018 | GPU code optimization using abstract kernel emulation and sensitivity analysisabstractIn this paper, we develop an approach to GPU kernel optimization by focusing on identification of bottleneck resources and determining optimization parameters that can alleviate the bottleneck. Performance modeling for GPUs is done by abstract kernel emulation along with latency/gap modeling of resources. Sensitivity analysis with respect to resource latency/gap parameters is used to predict the bottleneck resource for a given kernel's execution. The utility of the bottleneck analysis is demonstrated in two contexts: 1) Coupling the new bottleneck-driven optimization strategy with the OpenTuner auto-tuner: experimental results on all kernels from the Rodinia suite and GPU tensor contraction kernels from the NWChem computational chemistry suite demonstrate effectiveness. 2) Manual code optimization: two case studies illustrate the use of the bottleneck analysis to iteratively improve the performance of code from state-of-the-art domain-specific code generators. Changwan Hong, Aravind Sukumaran-Rajam, Prashant Singh Rawat, Sriram Krishnamoorthy, Louis-Noël Pouchet, Fabrice Rastello, P. Sadayappan |
PLDI | 4 |
| 2018 | Performance modeling for GPUs using abstract kernel emulationabstractPerformance modeling of GPU kernels is a significant challenge. In this paper, we develop a novel approach to performance modeling for GPUs through abstract kernel emulation along with latency/gap modeling of resources. Experimental results on all benchmarks from the Rodinia suite demonstrate good accuracy in predicting execution time on multiple GPU platforms. Changwan Hong, Aravind Sukumaran-Rajam, Prashant Singh Rawat, Sriram Krishnamoorthy, Louis-Noël Pouchet, Fabrice Rastello, P. Sadayappan |
PPoPP | 4 |
| 2018 | Register optimizations for stencils on GPUsabstractThe recent advent of compute-intensive GPU architecture has allowed application developers to explore high-order 3D stencils for better computational accuracy. A common optimization strategy for such stencils is to expose sufficient data reuse by means such as loop unrolling, with the expectation of register-level reuse. However, the resulting code is often highly constrained by register pressure. While current state-of-the-art register allocators are satisfactory for most applications, they are unable to effectively manage register pressure for such complex high-order stencils, resulting in sub-optimal code with a large number of register spills. In this paper, we develop a statement reordering framework that models stencil computations as a DAG of trees with shared leaves, and adapts an optimal scheduling algorithm for minimizing register usage for expression trees. The effectiveness of the approach is demonstrated through experimental results on a range of stencils extracted from application codes. Prashant Singh Rawat, Fabrice Rastello, Aravind Sukumaran-Rajam, Louis-Noël Pouchet, Atanas Rountev, P. Sadayappan |
PPoPP | 1 |
| 2018 | Associative instruction reordering to alleviate register pressure
Prashant Singh Rawat, Aravind Sukumaran-Rajam, Atanas Rountev, Fabrice Rastello, Louis-Noël Pouchet, P. Sadayappan |
SC | 1 |
| 2018 | Domain-Specific Optimization and Generation of High-Performance GPU Code for Stencil ComputationsabstractStencil computations arise in a number of computational domains. They exhibit significant data parallelism and are thus well suited for execution on graphical processing units (GPUs), but can be memory-bandwidth limited unless temporal locality is utilized via tiling. This paper describes how effective tiled code can be generated for GPUs from a domain-specific language (DSL) for stencils. Experimental results demonstrate the benefits of such a domain-specific optimization approach over state-of-the-art general-purpose compiler optimizations. Prashant Singh Rawat, Miheer Vaidya, Aravind Sukumaran-Rajam, Mahesh Ravishankar, Vinod Grover, Atanas Rountev, Louis-Noël Pouchet, P. Sadayappan |
Proc. IEEE | 1 |
| 2017 | POSTER: Statement Reordering to Alleviate Register Pressure for Stencils on GPUsabstractCompute-intensive GPU architectures allow the use of high-order 3D stencils for better computational accuracy. These stencils are usually compute-bound. While current state-of-the-art register allocators are satisfactory for most applications, they are unable to effectively manage register pressure for such complex high-order stencils, resulting in a sub-optimal code with a large number of register spills. We develop an optimization framework that models stencils as a forest of trees and performs statement reordering to reduce register use. The effectiveness of the approach is demonstrated through experimental results on several high-order stencils. Prashant Singh Rawat, Aravind Sukumaran-Rajam, Atanas Rountev, Fabrice Rastello, Louis-Noël Pouchet, P. Sadayappan |
PACT | 1 |
| 2016 | Resource Conscious Reuse-Driven Tiling for GPUsabstractComputations involving successive application of 3D stencil operators are widely used in many application domains, such as image processing, computational electromagnetics, seismic processing, and climate modeling. Enhancement of temporal and spatial locality via tiling is generally required in order to overcome performance bottlenecks due to limited bandwidth to global memory on GPUs. However, the low shared memory capacity on current GPU architectures makes effective tiling for 3D stencils very challenging -- several previous domain-specific compilers for stencils have demonstrated very high performance for 2D stencils, but much lower performance on 3D stencils. Prashant Singh Rawat, Changwan Hong, Mahesh Ravishankar, Vinod Grover, Louis-Noël Pouchet, Atanas Rountev, P. Sadayappan |
PACT | 1 |
| 2012 | Liveness-Based Pointer Analysis
Uday P. Khedker, Alan Mycroft, Prashant Singh Rawat |
SAS | 3 |