Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Mahesh Ravishankar

dblp:15/9083 · DBLP profile ↗
← Back
8ranked-venue papers
2as first author
1since 2021 · last 2025
0000-0003-4782-6638ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 2 first-author · 1 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Software engineering, system software, and programming languages
5 papers
Compilers and program optimization · 83% Program synthesis and code generation · 17%
Computer architecture, parallel and distributed computing, and storage systems
5 papers
High-performance computing · 34% GPUs and heterogeneous computing · 24% Performance modeling and evaluation · 16%

Topics — the 21 heaviest of 21, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Program synthesis and code generation
domain-specific code generation
0.312018
Domain-Specific Optimization and Generation of High-Performance GPU Code for Stencil Computations · Proc. IEEE 2018
Compilers and program optimization
stencil computation
0.312018
Domain-Specific Optimization and Generation of High-Performance GPU Code for Stencil Computations · Proc. IEEE 2018
GPUs and heterogeneous computing › GPU performance optimization
GPU code optimization
0.312018
Domain-Specific Optimization and Generation of High-Performance GPU Code for Stencil Computations · Proc. IEEE 2018
Compilers and program optimization › code generation › parallel code generation
distributed-memory code generation
0.212015
Distributed memory code generation for mixed Irregular/Regular computations · PPoPP 2015
Compilers and program optimization
parallelizing compiler
0.212015
Distributed memory code generation for mixed Irregular/Regular computations · PPoPP 2015
Compilers and program optimization › loop transformation
polyhedral compilation
0.212015
Distributed memory code generation for mixed Irregular/Regular computations · PPoPP 2015
Compilers and program optimization › memory optimization
data locality optimization
0.212013
Beyond reuse distance analysis: Dynamic analysis for characterization of data locality potential · ACM Trans. Archit. Code Optim. 2013
Memory systems
data locality
0.212013
Beyond reuse distance analysis: Dynamic analysis for characterization of data locality potential · ACM Trans. Archit. Code Optim. 2013
Performance modeling and evaluation › cache performance modeling
reuse distance analysis
0.212013
Beyond reuse distance analysis: Dynamic analysis for characterization of data locality potential · ACM Trans. Archit. Code Optim. 2013
Compilers and program optimization › parallelization
inspector-executor
0.112012
Code generation for parallel execution of a class of irregular loops on distributed memory systems · SC 2012
Compilers and program optimization › code generation
parallel code generation
0.112012
Code generation for parallel execution of a class of irregular loops on distributed memory systems · SC 2012
Compilers and program optimization
vectorization
0.112012
Dynamic trace-based analysis of vectorization potential of applications · PLDI 2012
Parallel and multicore computing › parallelization strategies
distributed-memory parallelization
0.112012
Code generation for parallel execution of a class of irregular loops on distributed memory systems · SC 2012
High-performance computing
performance optimization at scale
0.112012
Dynamic trace-based analysis of vectorization potential of applications · PLDI 2012
High-performance computing
stencil computation
0.112018
Domain-Specific Optimization and Generation of High-Performance GPU Code for Stencil Computations · Proc. IEEE 2018
High-performance computing › scientific computing systems
adaptive mesh refinement
0.112015
Distributed memory code generation for mixed Irregular/Regular computations · PPoPP 2015
High-performance computing
scientific computing systems
0.112015
Distributed memory code generation for mixed Irregular/Regular computations · PPoPP 2015
Performance modeling and evaluation
workload characterization
0.012013
Beyond reuse distance analysis: Dynamic analysis for characterization of data locality potential · ACM Trans. Archit. Code Optim. 2013
High-performance computing › scientific computing systems
climate modeling
0.012012
Code generation for parallel execution of a class of irregular loops on distributed memory systems · SC 2012
Processor architecture and microarchitecture
instruction set architecture
0.012012
Dynamic trace-based analysis of vectorization potential of applications · PLDI 2012
High-performance computing
scientific computing
0.012012
Code generation for parallel execution of a class of irregular loops on distributed memory systems · SC 2012

Methods — techniques the papers use, named apart from their topics

tiling · 0.7domain-specific language · 0.7polyhedral framework · 0.4loop transformation · 0.4dynamic analysis · 0.3convex partitioning · 0.3CDAG · 0.3static and runtime analysis · 0.3sensitivity analysis · 0.3dynamic trace-based analysis · 0.3
YearPublicationVenuePosition
2025 Pattern Matching in AI Compilers and Its Formalization
abstract
PyPM is a Python-based domain specific language (DSL) for building rewrite-based optimization passes on machine learning computation graphs. Users define individual optimizations by writing (a) patterns that match subgraphs of a computation graph and (b) corresponding rules which replace a matched subgraph with an optimized kernel. PyPM is distinguished from the many other DSLs for defining rewriting passes by its complex and novel pattern language which borrows concepts from logic programming. PyPM patterns can be recursive, nondeterminstic, and can require checking domain-specific constraints such as the shapes of tensors. The PyPM implementation is thus similarly complicated, consisting of thousands of lines of C++ code. In this paper, we present our work on building PyPM, as well as formalizing and distilling and this complexity to an understandable mathematical core. We have developed a formal core calculus expressing the main operations of the PyPM pattern language. We define both a declarative semantics – describing which patterns match which terms – and an algorithmic semantics – an idealized version of the PyPM pattern interpreter – and prove their equivalence. The development is fully mechanized in the Coq proof assistant.
Joseph W. Cutler, Alex Collins, Mahesh Ravishankar, Vinod Grover
CGO4
2018 Domain-Specific Optimization and Generation of High-Performance GPU Code for Stencil Computations
abstract
Stencil computations arise in a number of computational domains. They exhibit significant data parallelism and are thus well suited for execution on graphical processing units (GPUs), but can be memory-bandwidth limited unless temporal locality is utilized via tiling. This paper describes how effective tiled code can be generated for GPUs from a domain-specific language (DSL) for stencils. Experimental results demonstrate the benefits of such a domain-specific optimization approach over state-of-the-art general-purpose compiler optimizations.
Prashant Singh Rawat, Miheer Vaidya, Aravind Sukumaran-Rajam, Mahesh Ravishankar, Vinod Grover, Atanas Rountev, Louis-Noël Pouchet, P. Sadayappan
Proc. IEEE4
2016 Resource Conscious Reuse-Driven Tiling for GPUs
abstract
Computations involving successive application of 3D stencil operators are widely used in many application domains, such as image processing, computational electromagnetics, seismic processing, and climate modeling. Enhancement of temporal and spatial locality via tiling is generally required in order to overcome performance bottlenecks due to limited bandwidth to global memory on GPUs. However, the low shared memory capacity on current GPU architectures makes effective tiling for 3D stencils very challenging -- several previous domain-specific compilers for stencils have demonstrated very high performance for 2D stencils, but much lower performance on 3D stencils.
Prashant Singh Rawat, Changwan Hong, Mahesh Ravishankar, Vinod Grover, Louis-Noël Pouchet, Atanas Rountev, P. Sadayappan
PACT3
2015 Distributed memory code generation for mixed Irregular/Regular computations
abstract
Many applications feature a mix of irregular and regular computational structures. For example, codes using adaptive mesh refinement (AMR) typically use a collection of regular blocks, where the number of blocks and the relationship between blocks is irregular. The computational structure in such applications generally involves regular (affine) loop computations within some number of innermost loops, while outer loops exhibit irregularity due to data-dependent control flow and indirect array access patterns. Prior approaches to distributed memory parallelization do not handle such computations effectively. They either target loop nests that are completely affine using polyhedral frameworks, or treat all loops as irregular. Consequently, the generated distributed memory code contains artifacts that disrupt the regular nature of previously affine innermost loops of the computation. This hampers subsequent optimizations to improve on-node performance. We propose a code generation framework that can effectively transform such applications for execution on distributed memory systems. Our approach generates distributed memory code which preserves program properties that enable subsequent polyhederal optimizations. Simultaneously, it addresses a major memory bottleneck of prior techniques that limits the scalability of the generated code. The effectiveness of the proposed framework is demonstrated on computations that are mixed regular/irregular, completely regular, and completely irregular.
Mahesh Ravishankar, Roshan Dathathri, Venmugil Elango, Louis-Noël Pouchet, J. Ramanujam, Atanas Rountev, P. Sadayappan
PPoPP1
2013 Beyond reuse distance analysis: Dynamic analysis for characterization of data locality potential
abstract
Emerging computer architectures will feature drastically decreased flops/byte (ratio of peak processing rate to memory bandwidth) as highlighted by recent studies on Exascale architectural trends. Further, flops are getting cheaper, while the energy cost of data movement is increasingly dominant. The understanding and characterization of data locality properties of computations is critical in order to guide efforts to enhance data locality. Reuse distance analysis of memory address traces is a valuable tool to perform data locality characterization of programs. A single reuse distance analysis can be used to estimate the number of cache misses in a fully associative LRU cache of any size, thereby providing estimates on the minimum bandwidth requirements at different levels of the memory hierarchy to avoid being bandwidth bound. However, such an analysis only holds for the particular execution order that produced the trace. It cannot estimate potential improvement in data locality through dependence-preserving transformations that change the execution schedule of the operations in the computation. In this article, we develop a novel dynamic analysis approach to characterize the inherent locality properties of a computation and thereby assess the potential for data locality enhancement via dependence-preserving transformations. The execution trace of a code is analyzed to extract a Computational-Directed Acyclic Graph (CDAG) of the data dependences. The CDAG is then partitioned into convex subsets, and the convex partitioning is used to reorder the operations in the execution trace to enhance data locality. The approach enables us to go beyond reuse distance analysis of a single specific order of execution of the operations of a computation in characterization of its data locality properties. It can serve a valuable role in identifying promising code regions for manual transformation, as well as assessing the effectiveness of compiler transformations for data locality enhancement. We demonstrate the effectiveness of the approach using a number of benchmarks, including case studies where the potential shown by the analysis is exploited to achieve lower data movement costs and better performance.
Naznin Fauzia, Venmugil Elango, Mahesh Ravishankar, J. Ramanujam, Fabrice Rastello, Atanas Rountev, Louis-Noël Pouchet, P. Sadayappan
ACM Trans. Archit. Code Optim.3
2012 Dynamic trace-based analysis of vectorization potential of applications
abstract
Recent hardware trends with GPUs and the increasing vector lengths of SSE-like ISA extensions for multicore CPUs imply that effective exploitation of SIMD parallelism is critical for achieving high performance on emerging and future architectures. A vast majority of existing applications were developed without any attention by their developers towards effective vectorizability of the codes. While developers of production compilers such as GNU gcc, Intel icc, PGI pgcc, and IBM xlc have invested considerable effort and made significant advances in enhancing automatic vectorization capabilities, these compilers still cannot effectively vectorize many existing scientific and engineering codes. It is therefore of considerable interest to analyze existing applications to assess the inherent latent potential for SIMD parallelism, exploitable through further compiler advances and/or via manual code changes.
Justin Holewinski, Ragavendar Ramamurthi, Mahesh Ravishankar, Naznin Fauzia, Louis-Noël Pouchet, Atanas Rountev, P. Sadayappan
PLDI3
2012 Code generation for parallel execution of a class of irregular loops on distributed memory systems
abstract
Parallelization and locality optimization of affine loop nests has been successfully addressed for shared-memory machines. However, many large-scale simulation applications must be executed in a distributed-memory environment, and use irregular/sparse computations where the control-flow and array-access patterns are data-dependent. In this paper, we propose an approach for effective parallel execution of a class of irregular loop computations in a distributed-memory environment, using a combination of static and runtime analysis. We discuss algorithms that analyze sequential code to generate an inspector and an executor. The inspector captures the data-dependent behavior of the computation in parallel and without requiring complete replication of any of the data structures used in the original computation. The executor performs the computation in parallel. The effectiveness of the framework is demonstrated on several benchmarks and a climate modeling application.
Mahesh Ravishankar, John Eisenlohr, Louis-Noël Pouchet, J. Ramanujam, Atanas Rountev, P. Sadayappan
SC1
2010 Optimal loop unrolling for GPGPU programs
abstract
Graphics Processing Units (GPUs) are massively parallel, many-core processors with tremendous computational power and very high memory bandwidth. With the advent of general purpose programming models such as NVIDIA's CUDA and the new standard OpenCL, general purpose programming using GPUs (GPGPU) has become very popular. However, the GPU architecture and programming model have brought along with it many new challenges and opportunities for compiler optimizations. One such classical optimization is loop unrolling. Current GPU compilers perform limited loop unrolling. In this paper, we attempt to understand the impact of loop unrolling on GPGPU programs. We develop a semi-automatic, compile-time approach for identifying optimal unroll factors for suitable loops in GPGPU programs. In addition, we propose techniques for reducing the number of unroll factors evaluated, based on the characteristics of the program being compiled and the device being compiled to. We use these techniques to evaluate the effect of loop unrolling on a range of GPGPU programs and show that we correctly identify the optimal unroll factors. The optimized versions run up to 70 percent faster than the unoptimized versions.
Giridhar Sreenivasa Murthy, Mahesh Ravishankar, Muthu Manikandan Baskaran, P. Sadayappan
IPDPS2