Prashant Singh Rawat

dblp:19/10672 · DBLP profile ↗
← Back
10ranked-venue papers
6as first author
0since 2021 · last 2019
0000-0003-3189-3983ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 5 first-authorSoftware engineering, systems software and programming languages · 2Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
6 papers
GPUs and heterogeneous computing · 32% Processor architecture and microarchitecture · 26% Performance modeling and evaluation · 20%
Software engineering, system software, and programming languages
2 papers
Compilers and program optimization · 75% Program synthesis and code generation · 25%

Topics — the 16 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Hardware accelerators and domain-specific architectures › sparsity exploitation
sparse tensor computation
0.412019
An efficient mixed-mode representation of sparse tensors · SC 2019
Program synthesis and code generation
domain-specific code generation
0.312018
Domain-Specific Optimization and Generation of High-Performance GPU Code for Stencil Computations · Proc. IEEE 2018
Compilers and program optimization
register allocation
0.312018
Register optimizations for stencils on GPUs · PPoPP 2018
Compilers and program optimization › register allocation
register pressure reduction
0.312018
Register optimizations for stencils on GPUs · PPoPP 2018
Compilers and program optimization
stencil computation
0.312018
Domain-Specific Optimization and Generation of High-Performance GPU Code for Stencil Computations · Proc. IEEE 2018
Performance modeling and evaluation
bottleneck analysis
0.312018
GPU code optimization using abstract kernel emulation and sensitivity analysis · PLDI 2018
GPUs and heterogeneous computing › GPU performance optimization
GPU code optimization
0.312018
Domain-Specific Optimization and Generation of High-Performance GPU Code for Stencil Computations · Proc. IEEE 2018
GPUs and heterogeneous computing
GPU compilation
0.312018
Register optimizations for stencils on GPUs · PPoPP 2018
GPUs and heterogeneous computing
GPU kernel optimization
0.312018
GPU code optimization using abstract kernel emulation and sensitivity analysis · PLDI 2018
Performance modeling and evaluation › processor performance modeling › accelerator performance modeling
GPU performance modeling
0.312018
Performance modeling for GPUs using abstract kernel emulation · PPoPP 2018
Processor architecture and microarchitecture
instruction reordering
0.312018
Associative instruction reordering to alleviate register pressure · SC 2018
Processor architecture and microarchitecture
instruction scheduling
0.312018
Associative instruction reordering to alleviate register pressure · SC 2018
Processor architecture and microarchitecture › register management
register pressure
0.312018
Associative instruction reordering to alleviate register pressure · SC 2018
High-performance computing › stencil computation
stencil computation optimization
0.312018
Register optimizations for stencils on GPUs · PPoPP 2018
Performance modeling and evaluation › statistical analysis
sensitivity analysis
0.112018
GPU code optimization using abstract kernel emulation and sensitivity analysis · PLDI 2018
High-performance computing
stencil computation
0.112018
Domain-Specific Optimization and Generation of High-Performance GPU Code for Stencil Computations · Proc. IEEE 2018

Methods — techniques the papers use, named apart from their topics

tiling · 0.7statement reordering · 0.7loop unrolling · 0.7domain-specific language · 0.7DAG scheduling · 0.7mixed-mode representation · 0.4compressed sparse fiber · 0.4latency modeling · 0.3gap modeling · 0.3associative instruction reordering · 0.3
YearPublicationVenuePosition
2019 On Optimizing Complex Stencils on GPUs
abstract
Stencil computations are often the compute-intensive kernel in many scientific applications. With the increasing demand for computational accuracy, and the emergence of massively data-parallel high-bandwidth architectures like GPUs, stencils have steadily become more complex in terms of the stencil order, data accesses, and reuse patterns. Many prior efforts have focused on optimizing simpler stencil computations on various platforms. However, existing stencil code generators face challenges in optimizing such complex multi-statement stencil DAGs. This paper addresses the challenges in optimizing high-order stencil DAGs on GPUs by focusing on two key considerations: (1) enabling the domain expert to guide the code optimization, which may otherwise be extremely challenging for complex stencils; and (2) using bottleneck analysis via runtime profiling to guide the application of optimizations, and the tuning of various code generation parameters. We implement these abstractions in a prototype code generation framework termed ARTEMIS, and evaluate its efficacy over multiple stencil kernels with varying complexity and operational intensity on an NVIDIA P100 GPU.
Prashant Singh Rawat, Miheer Vaidya, Aravind Sukumaran-Rajam, Atanas Rountev, Louis-Noël Pouchet, P. Sadayappan
IPDPS1
2019 An efficient mixed-mode representation of sparse tensors
abstract
The Compressed Sparse Fiber (CSF) representation for sparse tensors is a generalization of the Compressed Sparse Row (CSR) format for sparse matrices. For a tensor with d modes, typical tensor methods such as CANDECOMP/PARAFAC decomposition (CPD) require a sequence of d tensor computations, where efficient memory access with respect to different modes is required for each of them. The straightforward solution is to use d distinct representations of the tensor, with each one being efficient for one of the d computations. However, a d-fold space overhead is often unacceptable in practice, especially with memory-constrained GPUs. In this paper, we present a mixed-mode tensor representation that partitions the tensor's nonzero elements into disjoint sections, each of which is compressed to create fibers along a different mode. Experimental results demonstrate that better performance can be achieved while utilizing only a small fraction of the space required to keep d distinct CSF representations.
Israt Nisa, Jiajia Li 0001, Aravind Sukumaran-Rajam, Prashant Singh Rawat, Sriram Krishnamoorthy, P. Sadayappan
SC4
2018 GPU code optimization using abstract kernel emulation and sensitivity analysis
abstract
In this paper, we develop an approach to GPU kernel optimization by focusing on identification of bottleneck resources and determining optimization parameters that can alleviate the bottleneck. Performance modeling for GPUs is done by abstract kernel emulation along with latency/gap modeling of resources. Sensitivity analysis with respect to resource latency/gap parameters is used to predict the bottleneck resource for a given kernel's execution. The utility of the bottleneck analysis is demonstrated in two contexts: 1) Coupling the new bottleneck-driven optimization strategy with the OpenTuner auto-tuner: experimental results on all kernels from the Rodinia suite and GPU tensor contraction kernels from the NWChem computational chemistry suite demonstrate effectiveness. 2) Manual code optimization: two case studies illustrate the use of the bottleneck analysis to iteratively improve the performance of code from state-of-the-art domain-specific code generators.
Changwan Hong, Aravind Sukumaran-Rajam, Prashant Singh Rawat, Sriram Krishnamoorthy, Louis-Noël Pouchet, Fabrice Rastello, P. Sadayappan
PLDI4
2018 Performance modeling for GPUs using abstract kernel emulation
abstract
Performance modeling of GPU kernels is a significant challenge. In this paper, we develop a novel approach to performance modeling for GPUs through abstract kernel emulation along with latency/gap modeling of resources. Experimental results on all benchmarks from the Rodinia suite demonstrate good accuracy in predicting execution time on multiple GPU platforms.
Changwan Hong, Aravind Sukumaran-Rajam, Prashant Singh Rawat, Sriram Krishnamoorthy, Louis-Noël Pouchet, Fabrice Rastello, P. Sadayappan
PPoPP4
2018 Register optimizations for stencils on GPUs
abstract
The recent advent of compute-intensive GPU architecture has allowed application developers to explore high-order 3D stencils for better computational accuracy. A common optimization strategy for such stencils is to expose sufficient data reuse by means such as loop unrolling, with the expectation of register-level reuse. However, the resulting code is often highly constrained by register pressure. While current state-of-the-art register allocators are satisfactory for most applications, they are unable to effectively manage register pressure for such complex high-order stencils, resulting in sub-optimal code with a large number of register spills. In this paper, we develop a statement reordering framework that models stencil computations as a DAG of trees with shared leaves, and adapts an optimal scheduling algorithm for minimizing register usage for expression trees. The effectiveness of the approach is demonstrated through experimental results on a range of stencils extracted from application codes.
Prashant Singh Rawat, Fabrice Rastello, Aravind Sukumaran-Rajam, Louis-Noël Pouchet, Atanas Rountev, P. Sadayappan
PPoPP1
2018 Associative instruction reordering to alleviate register pressure
Prashant Singh Rawat, Aravind Sukumaran-Rajam, Atanas Rountev, Fabrice Rastello, Louis-Noël Pouchet, P. Sadayappan
SC1
2018 Domain-Specific Optimization and Generation of High-Performance GPU Code for Stencil Computations
abstract
Stencil computations arise in a number of computational domains. They exhibit significant data parallelism and are thus well suited for execution on graphical processing units (GPUs), but can be memory-bandwidth limited unless temporal locality is utilized via tiling. This paper describes how effective tiled code can be generated for GPUs from a domain-specific language (DSL) for stencils. Experimental results demonstrate the benefits of such a domain-specific optimization approach over state-of-the-art general-purpose compiler optimizations.
Prashant Singh Rawat, Miheer Vaidya, Aravind Sukumaran-Rajam, Mahesh Ravishankar, Vinod Grover, Atanas Rountev, Louis-Noël Pouchet, P. Sadayappan
Proc. IEEE1
2017 POSTER: Statement Reordering to Alleviate Register Pressure for Stencils on GPUs
abstract
Compute-intensive GPU architectures allow the use of high-order 3D stencils for better computational accuracy. These stencils are usually compute-bound. While current state-of-the-art register allocators are satisfactory for most applications, they are unable to effectively manage register pressure for such complex high-order stencils, resulting in a sub-optimal code with a large number of register spills. We develop an optimization framework that models stencils as a forest of trees and performs statement reordering to reduce register use. The effectiveness of the approach is demonstrated through experimental results on several high-order stencils.
Prashant Singh Rawat, Aravind Sukumaran-Rajam, Atanas Rountev, Fabrice Rastello, Louis-Noël Pouchet, P. Sadayappan
PACT1
2016 Resource Conscious Reuse-Driven Tiling for GPUs
abstract
Computations involving successive application of 3D stencil operators are widely used in many application domains, such as image processing, computational electromagnetics, seismic processing, and climate modeling. Enhancement of temporal and spatial locality via tiling is generally required in order to overcome performance bottlenecks due to limited bandwidth to global memory on GPUs. However, the low shared memory capacity on current GPU architectures makes effective tiling for 3D stencils very challenging -- several previous domain-specific compilers for stencils have demonstrated very high performance for 2D stencils, but much lower performance on 3D stencils.
Prashant Singh Rawat, Changwan Hong, Mahesh Ravishankar, Vinod Grover, Louis-Noël Pouchet, Atanas Rountev, P. Sadayappan
PACT1
2012 Liveness-Based Pointer Analysis
Uday P. Khedker, Alan Mycroft, Prashant Singh Rawat
SAS3