EDBT 2026 Demo / reviewers in the wild / expert
Mahesh Ravishankar
dblp:15/9083
· DBLP profile ↗
8ranked-venue papers
2as first author
1since 2021 · last 2025
0000-0003-4782-6638ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 2 first-author · 1 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Software engineering, system software, and programming languages
5 papers |
Compilers and program optimization · 83% Program synthesis and code generation · 17% | |
| Computer architecture, parallel and distributed computing, and storage systems
5 papers |
High-performance computing · 34% GPUs and heterogeneous computing · 24% Performance modeling and evaluation · 16% |
Topics — the 21 heaviest of 21, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Program synthesis and code generation
domain-specific code generation |
0.3 | 1 | 2018 | Domain-Specific Optimization and Generation of High-Performance GPU Code for Stencil Computations · Proc. IEEE 2018 |
Compilers and program optimization
stencil computation |
0.3 | 1 | 2018 | Domain-Specific Optimization and Generation of High-Performance GPU Code for Stencil Computations · Proc. IEEE 2018 |
GPUs and heterogeneous computing › GPU performance optimization
GPU code optimization |
0.3 | 1 | 2018 | Domain-Specific Optimization and Generation of High-Performance GPU Code for Stencil Computations · Proc. IEEE 2018 |
Compilers and program optimization › code generation › parallel code generation
distributed-memory code generation |
0.2 | 1 | 2015 | Distributed memory code generation for mixed Irregular/Regular computations · PPoPP 2015 |
Compilers and program optimization
parallelizing compiler |
0.2 | 1 | 2015 | Distributed memory code generation for mixed Irregular/Regular computations · PPoPP 2015 |
Compilers and program optimization › loop transformation
polyhedral compilation |
0.2 | 1 | 2015 | Distributed memory code generation for mixed Irregular/Regular computations · PPoPP 2015 |
Compilers and program optimization › memory optimization
data locality optimization |
0.2 | 1 | 2013 | Beyond reuse distance analysis: Dynamic analysis for characterization of data locality potential · ACM Trans. Archit. Code Optim. 2013 |
Memory systems
data locality |
0.2 | 1 | 2013 | Beyond reuse distance analysis: Dynamic analysis for characterization of data locality potential · ACM Trans. Archit. Code Optim. 2013 |
Performance modeling and evaluation › cache performance modeling
reuse distance analysis |
0.2 | 1 | 2013 | Beyond reuse distance analysis: Dynamic analysis for characterization of data locality potential · ACM Trans. Archit. Code Optim. 2013 |
Compilers and program optimization › parallelization
inspector-executor |
0.1 | 1 | 2012 | Code generation for parallel execution of a class of irregular loops on distributed memory systems · SC 2012 |
Compilers and program optimization › code generation
parallel code generation |
0.1 | 1 | 2012 | Code generation for parallel execution of a class of irregular loops on distributed memory systems · SC 2012 |
Compilers and program optimization
vectorization |
0.1 | 1 | 2012 | Dynamic trace-based analysis of vectorization potential of applications · PLDI 2012 |
Parallel and multicore computing › parallelization strategies
distributed-memory parallelization |
0.1 | 1 | 2012 | Code generation for parallel execution of a class of irregular loops on distributed memory systems · SC 2012 |
High-performance computing
performance optimization at scale |
0.1 | 1 | 2012 | Dynamic trace-based analysis of vectorization potential of applications · PLDI 2012 |
High-performance computing
stencil computation |
0.1 | 1 | 2018 | Domain-Specific Optimization and Generation of High-Performance GPU Code for Stencil Computations · Proc. IEEE 2018 |
High-performance computing › scientific computing systems
adaptive mesh refinement |
0.1 | 1 | 2015 | Distributed memory code generation for mixed Irregular/Regular computations · PPoPP 2015 |
High-performance computing
scientific computing systems |
0.1 | 1 | 2015 | Distributed memory code generation for mixed Irregular/Regular computations · PPoPP 2015 |
Performance modeling and evaluation
workload characterization |
0.0 | 1 | 2013 | Beyond reuse distance analysis: Dynamic analysis for characterization of data locality potential · ACM Trans. Archit. Code Optim. 2013 |
High-performance computing › scientific computing systems
climate modeling |
0.0 | 1 | 2012 | Code generation for parallel execution of a class of irregular loops on distributed memory systems · SC 2012 |
Processor architecture and microarchitecture
instruction set architecture |
0.0 | 1 | 2012 | Dynamic trace-based analysis of vectorization potential of applications · PLDI 2012 |
High-performance computing
scientific computing |
0.0 | 1 | 2012 | Code generation for parallel execution of a class of irregular loops on distributed memory systems · SC 2012 |
Methods — techniques the papers use, named apart from their topics
tiling · 0.7domain-specific language · 0.7polyhedral framework · 0.4loop transformation · 0.4dynamic analysis · 0.3convex partitioning · 0.3CDAG · 0.3static and runtime analysis · 0.3sensitivity analysis · 0.3dynamic trace-based analysis · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Pattern Matching in AI Compilers and Its FormalizationabstractPyPM is a Python-based domain specific language (DSL) for building rewrite-based optimization passes on machine learning computation graphs. Users define individual optimizations by writing (a) patterns that match subgraphs of a computation graph and (b) corresponding rules which replace a matched subgraph with an optimized kernel. PyPM is distinguished from the many other DSLs for defining rewriting passes by its complex and novel pattern language which borrows concepts from logic programming. PyPM patterns can be recursive, nondeterminstic, and can require checking domain-specific constraints such as the shapes of tensors. The PyPM implementation is thus similarly complicated, consisting of thousands of lines of C++ code. In this paper, we present our work on building PyPM, as well as formalizing and distilling and this complexity to an understandable mathematical core. We have developed a formal core calculus expressing the main operations of the PyPM pattern language. We define both a declarative semantics – describing which patterns match which terms – and an algorithmic semantics – an idealized version of the PyPM pattern interpreter – and prove their equivalence. The development is fully mechanized in the Coq proof assistant. Joseph W. Cutler, Alex Collins, Mahesh Ravishankar, Vinod Grover |
CGO | 4 |
| 2018 | Domain-Specific Optimization and Generation of High-Performance GPU Code for Stencil ComputationsabstractStencil computations arise in a number of computational domains. They exhibit significant data parallelism and are thus well suited for execution on graphical processing units (GPUs), but can be memory-bandwidth limited unless temporal locality is utilized via tiling. This paper describes how effective tiled code can be generated for GPUs from a domain-specific language (DSL) for stencils. Experimental results demonstrate the benefits of such a domain-specific optimization approach over state-of-the-art general-purpose compiler optimizations. Prashant Singh Rawat, Miheer Vaidya, Aravind Sukumaran-Rajam, Mahesh Ravishankar, Vinod Grover, Atanas Rountev, Louis-Noël Pouchet, P. Sadayappan |
Proc. IEEE | 4 |
| 2016 | Resource Conscious Reuse-Driven Tiling for GPUsabstractComputations involving successive application of 3D stencil operators are widely used in many application domains, such as image processing, computational electromagnetics, seismic processing, and climate modeling. Enhancement of temporal and spatial locality via tiling is generally required in order to overcome performance bottlenecks due to limited bandwidth to global memory on GPUs. However, the low shared memory capacity on current GPU architectures makes effective tiling for 3D stencils very challenging -- several previous domain-specific compilers for stencils have demonstrated very high performance for 2D stencils, but much lower performance on 3D stencils. Prashant Singh Rawat, Changwan Hong, Mahesh Ravishankar, Vinod Grover, Louis-Noël Pouchet, Atanas Rountev, P. Sadayappan |
PACT | 3 |
| 2015 | Distributed memory code generation for mixed Irregular/Regular computationsabstractMany applications feature a mix of irregular and regular computational structures. For example, codes using adaptive mesh refinement (AMR) typically use a collection of regular blocks, where the number of blocks and the relationship between blocks is irregular. The computational structure in such applications generally involves regular (affine) loop computations within some number of innermost loops, while outer loops exhibit irregularity due to data-dependent control flow and indirect array access patterns. Prior approaches to distributed memory parallelization do not handle such computations effectively. They either target loop nests that are completely affine using polyhedral frameworks, or treat all loops as irregular. Consequently, the generated distributed memory code contains artifacts that disrupt the regular nature of previously affine innermost loops of the computation. This hampers subsequent optimizations to improve on-node performance. We propose a code generation framework that can effectively transform such applications for execution on distributed memory systems. Our approach generates distributed memory code which preserves program properties that enable subsequent polyhederal optimizations. Simultaneously, it addresses a major memory bottleneck of prior techniques that limits the scalability of the generated code. The effectiveness of the proposed framework is demonstrated on computations that are mixed regular/irregular, completely regular, and completely irregular. Mahesh Ravishankar, Roshan Dathathri, Venmugil Elango, Louis-Noël Pouchet, J. Ramanujam, Atanas Rountev, P. Sadayappan |
PPoPP | 1 |
| 2013 | Beyond reuse distance analysis: Dynamic analysis for characterization of data locality potentialabstractEmerging computer architectures will feature drastically decreased flops/byte (ratio of peak processing rate to memory bandwidth) as highlighted by recent studies on Exascale architectural trends. Further, flops are getting cheaper, while the energy cost of data movement is increasingly dominant. The understanding and characterization of data locality properties of computations is critical in order to guide efforts to enhance data locality. Reuse distance analysis of memory address traces is a valuable tool to perform data locality characterization of programs. A single reuse distance analysis can be used to estimate the number of cache misses in a fully associative LRU cache of any size, thereby providing estimates on the minimum bandwidth requirements at different levels of the memory hierarchy to avoid being bandwidth bound. However, such an analysis only holds for the particular execution order that produced the trace. It cannot estimate potential improvement in data locality through dependence-preserving transformations that change the execution schedule of the operations in the computation. In this article, we develop a novel dynamic analysis approach to characterize the inherent locality properties of a computation and thereby assess the potential for data locality enhancement via dependence-preserving transformations. The execution trace of a code is analyzed to extract a Computational-Directed Acyclic Graph (CDAG) of the data dependences. The CDAG is then partitioned into convex subsets, and the convex partitioning is used to reorder the operations in the execution trace to enhance data locality. The approach enables us to go beyond reuse distance analysis of a single specific order of execution of the operations of a computation in characterization of its data locality properties. It can serve a valuable role in identifying promising code regions for manual transformation, as well as assessing the effectiveness of compiler transformations for data locality enhancement. We demonstrate the effectiveness of the approach using a number of benchmarks, including case studies where the potential shown by the analysis is exploited to achieve lower data movement costs and better performance. Naznin Fauzia, Venmugil Elango, Mahesh Ravishankar, J. Ramanujam, Fabrice Rastello, Atanas Rountev, Louis-Noël Pouchet, P. Sadayappan |
ACM Trans. Archit. Code Optim. | 3 |
| 2012 | Dynamic trace-based analysis of vectorization potential of applicationsabstractRecent hardware trends with GPUs and the increasing vector lengths of SSE-like ISA extensions for multicore CPUs imply that effective exploitation of SIMD parallelism is critical for achieving high performance on emerging and future architectures. A vast majority of existing applications were developed without any attention by their developers towards effective vectorizability of the codes. While developers of production compilers such as GNU gcc, Intel icc, PGI pgcc, and IBM xlc have invested considerable effort and made significant advances in enhancing automatic vectorization capabilities, these compilers still cannot effectively vectorize many existing scientific and engineering codes. It is therefore of considerable interest to analyze existing applications to assess the inherent latent potential for SIMD parallelism, exploitable through further compiler advances and/or via manual code changes. Justin Holewinski, Ragavendar Ramamurthi, Mahesh Ravishankar, Naznin Fauzia, Louis-Noël Pouchet, Atanas Rountev, P. Sadayappan |
PLDI | 3 |
| 2012 | Code generation for parallel execution of a class of irregular loops on distributed memory systemsabstractParallelization and locality optimization of affine loop nests has been successfully addressed for shared-memory machines. However, many large-scale simulation applications must be executed in a distributed-memory environment, and use irregular/sparse computations where the control-flow and array-access patterns are data-dependent. In this paper, we propose an approach for effective parallel execution of a class of irregular loop computations in a distributed-memory environment, using a combination of static and runtime analysis. We discuss algorithms that analyze sequential code to generate an inspector and an executor. The inspector captures the data-dependent behavior of the computation in parallel and without requiring complete replication of any of the data structures used in the original computation. The executor performs the computation in parallel. The effectiveness of the framework is demonstrated on several benchmarks and a climate modeling application. Mahesh Ravishankar, John Eisenlohr, Louis-Noël Pouchet, J. Ramanujam, Atanas Rountev, P. Sadayappan |
SC | 1 |
| 2010 | Optimal loop unrolling for GPGPU programsabstractGraphics Processing Units (GPUs) are massively parallel, many-core processors with tremendous computational power and very high memory bandwidth. With the advent of general purpose programming models such as NVIDIA's CUDA and the new standard OpenCL, general purpose programming using GPUs (GPGPU) has become very popular. However, the GPU architecture and programming model have brought along with it many new challenges and opportunities for compiler optimizations. One such classical optimization is loop unrolling. Current GPU compilers perform limited loop unrolling. In this paper, we attempt to understand the impact of loop unrolling on GPGPU programs. We develop a semi-automatic, compile-time approach for identifying optimal unroll factors for suitable loops in GPGPU programs. In addition, we propose techniques for reducing the number of unroll factors evaluated, based on the characteristics of the program being compiled and the device being compiled to. We use these techniques to evaluate the effect of loop unrolling on a range of GPGPU programs and show that we correctly identify the optimal unroll factors. The optimized versions run up to 70 percent faster than the unoptimized versions. Giridhar Sreenivasa Murthy, Mahesh Ravishankar, Muthu Manikandan Baskaran, P. Sadayappan |
IPDPS | 2 |