EDBT 2026 Demo / reviewers in the wild / expert
Vijay Menon 0002
dblp:21/1138-2 · also Vijay S. Menon
· DBLP profile ↗
16ranked-venue papers
5as first author
0since 2021 · last 2008
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 3 first-authorSoftware engineering, systems software and programming languages · 7 · 2 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Software engineering, system software, and programming languages
8 papers |
Concurrent programming · 41% Runtime systems and virtual machines · 26% Compilers and program optimization · 15% | |
| Computer architecture, parallel and distributed computing, and storage systems
4 papers |
Parallel and multicore computing · 70% Processor architecture and microarchitecture · 29% High-performance computing · 1% |
Topics — the 29 heaviest of 30, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Concurrent programming › transactional memory
software transactional memory |
0.2 | 3 | 2008 | Practical experiences with Java software transactional memory · PPoPP 2008 Enforcing isolation and ordering in STM · PLDI 2007 Compiler and runtime support for efficient software transactional memory · PLDI 2006 |
Concurrent programming
transactional memory |
0.2 | 3 | 2008 | Concurrent GC leveraging transactional memory · PPoPP 2008 Enforcing isolation and ordering in STM · PLDI 2007 Compiler and runtime support for efficient software transactional memory · PLDI 2006 |
Parallel and multicore computing
transactional memory |
0.1 | 2 | 2008 | Dynamic optimization for efficient strong atomicity · OOPSLA 2008 Open nesting in software transactional memory · PPoPP 2007 |
Runtime systems and virtual machines › garbage collection
concurrent garbage collection |
0.1 | 1 | 2008 | Concurrent GC leveraging transactional memory · PPoPP 2008 |
Runtime systems and virtual machines
garbage collection |
0.1 | 1 | 2008 | Concurrent GC leveraging transactional memory · PPoPP 2008 |
Runtime systems and virtual machines › managed runtime
java runtime |
0.1 | 1 | 2008 | Practical experiences with Java software transactional memory · PPoPP 2008 |
Processor architecture and microarchitecture
dynamic optimization |
0.1 | 1 | 2008 | Dynamic optimization for efficient strong atomicity · OOPSLA 2008 |
Runtime systems and virtual machines
language runtime |
0.1 | 1 | 2007 | Enabling scalability and performance in a large scale CMP environment · EuroSys 2007 |
Concurrent programming
memory models |
0.1 | 1 | 2007 | Enforcing isolation and ordering in STM · PLDI 2007 |
Processor architecture and microarchitecture
chip multiprocessor |
0.1 | 1 | 2007 | Enabling scalability and performance in a large scale CMP environment · EuroSys 2007 |
Parallel and multicore computing
concurrent programming |
0.1 | 1 | 2007 | Open nesting in software transactional memory · PPoPP 2007 |
Parallel and multicore computing
parallel programming runtimes |
0.1 | 1 | 2007 | Enabling scalability and performance in a large scale CMP environment · EuroSys 2007 |
Parallel and multicore computing › transactional memory
software transactional memory |
0.1 | 1 | 2007 | Open nesting in software transactional memory · PPoPP 2007 |
Program verification › security property verification
memory safety verification |
0.1 | 1 | 2006 | A verifiable SSA program representation for aggressive compiler optimization · POPL 2006 |
Concurrent programming › transactional memory
nested transactions |
0.1 | 1 | 2006 | Compiler and runtime support for efficient software transactional memory · PLDI 2006 |
Program analysis
program representation |
0.1 | 1 | 2006 | A verifiable SSA program representation for aggressive compiler optimization · POPL 2006 |
Compilers and program optimization › intermediate representation
static single assignment form |
0.1 | 1 | 2006 | A verifiable SSA program representation for aggressive compiler optimization · POPL 2006 |
Programming languages and type systems › type systems
type soundness |
0.1 | 1 | 2006 | A verifiable SSA program representation for aggressive compiler optimization · POPL 2006 |
Compilers and program optimization › memory optimization
cache optimization |
0.0 | 1 | 2003 | Fractal symbolic analysis · ACM Trans. Program. Lang. Syst. 2003 |
Compilers and program optimization
dependence analysis |
0.0 | 1 | 2003 | Fractal symbolic analysis · ACM Trans. Program. Lang. Syst. 2003 |
Compilers and program optimization
loop transformation |
0.0 | 1 | 2003 | Fractal symbolic analysis · ACM Trans. Program. Lang. Syst. 2003 |
Program analysis
symbolic execution |
0.0 | 1 | 2003 | Fractal symbolic analysis · ACM Trans. Program. Lang. Syst. 2003 |
Runtime systems and virtual machines
dynamic compilation |
0.0 | 1 | 2008 | Dynamic optimization for efficient strong atomicity · OOPSLA 2008 |
Compilers and program optimization › compiler optimization
speculative optimization |
0.0 | 1 | 2008 | Dynamic optimization for efficient strong atomicity · OOPSLA 2008 |
Runtime systems and virtual machines › dynamic compilation
just-in-time compilation |
0.0 | 1 | 2006 | Compiler and runtime support for efficient software transactional memory · PLDI 2006 |
Parallel and multicore computing › multiprocessor system
distributed-memory multiprocessor |
0.0 | 1 | 1997 | MultiMATLAB Integrating MATLAB with High Performance Parallel Computing · SC 1997 |
Parallel and multicore computing › parallel computing
parallel computing environments |
0.0 | 1 | 1997 | MultiMATLAB Integrating MATLAB with High Performance Parallel Computing · SC 1997 |
Parallel and multicore computing
parallel programming models |
0.0 | 1 | 1997 | MultiMATLAB Integrating MATLAB with High Performance Parallel Computing · SC 1997 |
Software maintenance and evolution › software reengineering
program restructuring |
0.0 | 1 | 2003 | Fractal symbolic analysis · ACM Trans. Program. Lang. Syst. 2003 |
Methods — techniques the papers use, named apart from their topics
speculation and recovery · 0.2dynamic analysis · 0.2experimental evaluation · 0.1transactional memory · 0.1software transactional memory · 0.1lock-based synchronization · 0.1type system · 0.1runtime optimization · 0.1just-in-time compilation · 0.1SSA · 0.1MPI communication · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2008 | Fault-safe code motion for type-safe languagesabstractCompilers for Java and other type-safe languages have historically worked to overcome overheads and constraints imposed by runtime safety checks and precise exception semantics. We instead exploit these safety properties to perform code motion optimizations that are even more aggressive than those possible in unsafe languages such as C++. Brian R. Murphy, Vijay Menon 0002, Florian T. Schneider, Tatiana Shpeisman, Ali-Reza Adl-Tabatabai |
CGO | 2 |
| 2008 | A Uniform Transactional Execution Environment for Java
Lukasz Ziarek, Adam Welc, Ali-Reza Adl-Tabatabai, Vijay Menon 0002, Tatiana Shpeisman, Suresh Jagannathan |
ECOOP | 4 |
| 2008 | Dynamic optimization for efficient strong atomicityabstractTransactional memory (TM) is a promising concurrency control alternative to locks. Recent work has highlighted important memory model issues regarding TM semantics and exposed problems in existing TM implementations. For safe, managed languages such as Java, there is a growing consensus towards strong atomicity semantics as a sound, scalable solution. Strong atomicity has presented a challenge to implement efficiently because it requires instrumentation of non-transactional memory accesses, incurring significant overhead even when a program makes minimal or no use of transactions. To minimize overhead, existing solutions require either a sophisticated type system, specialized hardware, or static whole-program analysis. These techniques do not translate easily into a production setting on existing hardware. In this paper, we present novel dynamic optimizations that significantly reduce strong atomicity overheads and make strong atomicity practical for dynamic language environments. We introduce analyses that optimistically track which non-transactional memory accesses can avoid strong atomicity instrumentation, and we describe a lightweight speculation and recovery mechanism that applies these analyses to generate speculatively-optimized but safe code for strong atomicity in a dynamically-loaded environment. We show how to implement these mechanisms efficiently by leveraging existing dynamic optimization infrastructure in a Java system. Measurements on a set of transactional and non-transactional Java workloads demonstrate that our techniques substantially reduce the overhead of strong atomicity from a factor of 5x down to 10% or less over an efficient weak atomicity baseline. Florian T. Schneider, Vijay Menon 0002, Tatiana Shpeisman, Ali-Reza Adl-Tabatabai |
OOPSLA | 2 |
| 2008 | Practical experiences with Java software transactional memoryabstractIn this paper, we evaluate the emerging Transactional Memory (TM) area by developing a set of Java transactional memory workloads and studying their performance under a Java Software Transactional Memory (STM) system and comparing them to their lock based counterparts. We provide a detailed performance and memory consumption analysis of the overheads of software transactional memory and transactional workloads within a production quality open source Java Runtime system. Additionally, we detail the impact of the various performance optimizations in both workloads and the underlying runtime system to improving both single thread performance and scalability. Evgueni Brevnov, Yuri Dolgov, Boris Kuznetsov, Dmitry Yershov, Vyacheslav Shakin, Dong-yuan Chen, Vijay Menon 0002, Suresh Srinivas |
PPoPP | 7 |
| 2008 | Concurrent GC leveraging transactional memoryabstractWe predict that the ever-growing number of cores on our desktops will require a re-examination of concurrent programming. Two technologies are likely to become mainstream in response: Transactional memory provides a superior programming model to traditional lock-based concurrency, while Concurrent GC can take advantage of multiple cores to eliminate perceptible pauses in desktop applications such as games or Internet telephony. This paper proposes a combination of the two technologies, producing a synergy that improves scalability while eliminating the annoyance of user-perceivable pauses. Phil McGachey, Ali-Reza Adl-Tabatabai, Richard L. Hudson, Vijay Menon 0002, Bratin Saha, Tatiana Shpeisman |
PPoPP | 4 |
| 2008 | Practical weak-atomicity semantics for java stmabstractAs memory transactions have been proposed as a language-level replacement for locks, there is growing need for well-defined semantics. In contrast to database transactions, transaction memory (TM) semantics are complicated by the fact that programs may access the same memory locations both inside and outside transactions. Strongly atomic semantics, where non transactional accesses are treated as implicit single-operation transactions, remain difficult to provide without specialized hardware support or significant performance overhead. As an alternative, many in the community have informally proposed that a single global lock semantics [18,10], where transaction semantics are mapped to those of regions protected by a single global lock, provide an intuitive and efficiently implementable model for programmers. Vijay Menon 0002, Steven Balensiefer, Tatiana Shpeisman, Ali-Reza Adl-Tabatabai, Richard L. Hudson, Bratin Saha, Adam Welc |
SPAA | 1 |
| 2007 | Enabling scalability and performance in a large scale CMP environmentabstractHardware trends suggest that large-scale CMP architectures, with tens to hundreds of processing cores on a single piece of silicon, are iminent within the next decade. While existing CMP machines have traditionally been handled in the same way as SMPs, this magnitude of parallelism introduces several fundamental challenges at the architectural level and this, in turn, translates to novel challenges in the design of the software stack for these platforms. This paper presents the "Many Core Run Time" (McRT), a software prototype of an integrated language runtime that was designed to explore configurations of the software stack for enabling performance and scalability on large scale CMP platforms. This paper presents the architecture of McRT and discusses our experiences with the system, including experimental evaluation that lead to several interesting, non-intuitive findings, providing key insights about the structure of the system stack at this scale. A key contribution of this paper is to demonstrate how McRT enables near linear improvements in performance and scalability for desktop workloads such as the popular XviD encoder and a set of RMS (recognition, mining, and synthesis) applications. Another key contribution of this work is its use of McRT to explore non-traditional system configurations such as a light-weight executive in which McRT runs on "bare metal" and replaces the traditional OS. Such configurations are becoming an increasingly attractive alternative to leverage heterogeneous computing uints as seen in today's CPU-GPU configurations. Bratin Saha, Ali-Reza Adl-Tabatabai, Anwar M. Ghuloum, Mohan Rajagopalan, Richard L. Hudson, Leaf Petersen, Vijay Menon 0002, Brian R. Murphy, Tatiana Shpeisman, Eric Sprangle, Anwar Rohillah, Doug Carmean, Jesse Fang |
EuroSys | 7 |
| 2007 | Enforcing isolation and ordering in STMabstractTransactional memory provides a new concurrency control mechanism that avoids many of the pitfalls of lock-based synchronization. High-performance software transactional memory (STM) implementations thus far provide weak atomicity: Accessing shared data both inside and outside a transaction can result in unexpected, implementation-dependent behavior. To guarantee isolation and consistent ordering in such a system, programmers are expected to enclose all shared-memory accesses inside transactions. Tatiana Shpeisman, Vijay Menon 0002, Ali-Reza Adl-Tabatabai, Steven Balensiefer, Dan Grossman, Richard L. Hudson, Katherine F. Moore, Bratin Saha |
PLDI | 2 |
| 2007 | Open nesting in software transactional memoryabstractTransactional memory (TM) promises to simplify concurrent programming while providing scalability competitive to fine-grained locking. Language-based constructs allow programmers to denote atomic regions declaratively and to rely on the underlying system to provide transactional guarantees along with concurrency. In contrast with fine-grained locking, TM allows programmers to write simpler programs that are composable and deadlock-free. Vijay Menon 0002, Ali-Reza Adl-Tabatabai, Antony L. Hosking, Richard L. Hudson, J. Eliot B. Moss, Bratin Saha, Tatiana Shpeisman |
PPoPP | 2 |
| 2006 | Compiler and runtime support for efficient software transactional memoryabstractProgrammers have traditionally used locks to synchronize concurrent access to shared data. Lock-based synchronization, however, has well-known pitfalls: using locks for fine-grain synchronization and composing code that already uses locks are both difficult and prone to deadlock. Transactional memory provides an alternate concurrency control mechanism that avoids these pitfalls and significantly eases concurrent programming. Transactional memory language constructs have recently been proposed as extensions to existing languages or included in new concurrent language specifications, opening the door for new compiler optimizations that target the overheads of transactional memory.This paper presents compiler and runtime optimizations for transactional memory language constructs. We present a high-performance software transactional memory system (STM) integrated into a managed runtime environment. Our system efficiently implements nested transactions that support both composition of transactions and partial roll back. Our JIT compiler is the first to optimize the overheads of STM, and we show novel techniques for enabling JIT optimizations on STM operations. We measure the performance of our optimizations on a 16-way SMP running multi-threaded transactional workloads. Our results show that these techniques enable transactional memory's performance to compete with that of well-tuned synchronization. Ali-Reza Adl-Tabatabai, Brian T. Lewis, Vijay Menon 0002, Brian R. Murphy, Bratin Saha, Tatiana Shpeisman |
PLDI | 3 |
| 2006 | A verifiable SSA program representation for aggressive compiler optimizationabstractWe present a verifiable low-level program representation to embed, propagate, and preserve safety information in high perfor-mance compilers for safe languages such as Java and C#. Our representation precisely encodes safety information via static single-assignment (SSA) [11, 3] proof variables that are first-class constructs in the program.We argue that our representation allows a compiler to both (1) express aggressively optimized machine-independent code and (2) leverage existing compiler infrastructure to preserve safety information during optimization. We demonstrate that this approach supports standard compiler optimizations, requires minimal changes to the implementation of those optimizations, and does not artificially impede those optimizations to preserve safety. We also describe a simple type system that formalizes type safety in an SSA-style control-flow graph program representation. Through the types of proof variables, our system enables compositional verification of memory safety in optimized code. Finally, we discuss experiences integrating this representation into the machine-independent global optimizer of STARJIT, a high-performance just-in-time compiler that performs aggressive control-flow, data-flow, and algebraic optimizations and is competitive with top production systems. Vijay Menon 0002, Neal Glew, Brian R. Murphy, Andrew McCreight, Tatiana Shpeisman, Ali-Reza Adl-Tabatabai, Leaf Petersen |
POPL | 1 |
| 2003 | Fractal symbolic analysisabstractModern compilers restructure programs to improve their efficiency. Dependence analysis is the most widely used technique for proving the correctness of such transformations, but it suffers from the limitation that it considers only the memory locations read and written by a statement without considering what is being computed by that statement. Exploiting the semantics of program statements permits more transformations to be proved correct, and is critical for automatic restructuring of codes such as LU with partial pivoting.One approach to exploiting the semantics of program statements is symbolic analysis and comparison of programs.In principle, this technique is very powerful, but in practice, it is intractable for all but the simplest programs.In this paper, we propose a new form of symbolic analysis and comparison of programs which is appropriate for use in restructuring compilers. Fractal symbolic analysis is an approximate symbolic analysis that compares a program and its transformed version by repeatedly simplifying these programs until symbolic analysis becomes tractable while ensuring that equality of the simplified programs is sufficient to guarantee equality of the original programs.Fractal symbolic analysis combines some of the power of symbolic analysis with the tractability of dependence analysis. We discuss a prototype implementation of fractal symbolic analysis, and show how it can be used to solve the long-open problem of verifying the correctness of transformations required to improve the cache performance of LU factorization with partial pivoting. Vijay Menon 0002, Keshav Pingali, Nikolay Mateev |
ACM Trans. Program. Lang. Syst. | 1 |
| 2001 | Fractal symbolic analysisabstractModern compilers perform wholesale restructuring of programs to improve their efficiency. Dependence analysis is the most widely used technique for proving the correctness of such transformations, but it suffers from the limitation that it considers only the memory locations read and written by a statement, and does not assume any particular interpretation for the operations in that statement. Exploiting the semantics of these operations permits more transformations to be proved correct, and is critical for automatic restructuring of codes such as LU with partial pivoting. Nikolay Mateev, Vijay Menon 0002, Keshav Pingali |
ICS | 2 |
| 2000 | Left-Looking to Right-Looking and Vice Versa: An Application of Fractal Symbolic Analysis to Linear Algebra Code Restructuring
Nikolay Mateev, Vijay Menon 0002, Keshav Pingali |
Euro-Par | 2 |
| 1999 | High-level semantic optimization of numerical codesabstractThis paper presents a mathematical framework to exploit the semantic properties of matrix operations in loop-based numerical codes. The heart of this framework is an algebraic language called the Abstract Matrix Form which a compiler can use to reason about matrix computations in terms of loop nests, high-level matrix operations, and intermediate forms. We demonstrate how this framework may be used to detect and exploit matrix products in loop-based languages such as FORTRAN and MATLAB, and discuss the resulting performance benefits. 1 Introduction Algebraic properties of scalar integer and floating point operations are used by most compilers to optimize programs. These properties enable compilers to reduce of the strength of expressions, enhance the power of common subexpression elimination, and verify the legality of certain loop transformations [2]. Although matrices are also endowed with a rich algebra, it is less common for compilers to exploit matrix algebra to optimize program... Vijay Menon 0002, Keshav Pingali |
International Conference on Supercomputing | 1 |
| 1997 | MultiMATLAB Integrating MATLAB with High Performance Parallel ComputingabstractMATLAB is the most popular scientific computing environment available on uniprocessors today. Unfortunately, no such environment is currently available for multiprocessors. MultiMATLAB [1] is a general extension of the MATLAB environment to any distributed memory multiprocessors. This paper presents a new MultiMATLAB system designed to provide high-performance on multiprocessors while maintaining the functionality and usability of the MATLAB environment. This system will enable users to access high-performance parallel routines from within the MATLAB environment, to extend the environment with new parallel routines, and to use these routines to develop parallel applications with the MATLAB language. We discuss a general MultiMATLAB architecture, present two implementations based upon the MPI communication standard [2], and demonstrate the use of this system. Preliminary results indicate that the MultiMATLAB system can offer the full performance of the underlying multiprocessor to the MATLAB environment. Vijay Menon 0002, Anne E. Trefethen |
SC | 1 |