Vijay Menon 0002

dblp:21/1138-2 · also Vijay S. Menon · DBLP profile ↗
← Back
16ranked-venue papers
5as first author
0since 2021 · last 2008
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 3 first-authorSoftware engineering, systems software and programming languages · 7 · 2 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Software engineering, system software, and programming languages
8 papers
Concurrent programming · 41% Runtime systems and virtual machines · 26% Compilers and program optimization · 15%
Computer architecture, parallel and distributed computing, and storage systems
4 papers
Parallel and multicore computing · 70% Processor architecture and microarchitecture · 29% High-performance computing · 1%

Topics — the 29 heaviest of 30, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Concurrent programming › transactional memory
software transactional memory
0.232008
Practical experiences with Java software transactional memory · PPoPP 2008
Enforcing isolation and ordering in STM · PLDI 2007
Compiler and runtime support for efficient software transactional memory · PLDI 2006
Concurrent programming
transactional memory
0.232008
Concurrent GC leveraging transactional memory · PPoPP 2008
Enforcing isolation and ordering in STM · PLDI 2007
Compiler and runtime support for efficient software transactional memory · PLDI 2006
Parallel and multicore computing
transactional memory
0.122008
Dynamic optimization for efficient strong atomicity · OOPSLA 2008
Open nesting in software transactional memory · PPoPP 2007
Runtime systems and virtual machines › garbage collection
concurrent garbage collection
0.112008
Concurrent GC leveraging transactional memory · PPoPP 2008
Runtime systems and virtual machines
garbage collection
0.112008
Concurrent GC leveraging transactional memory · PPoPP 2008
Runtime systems and virtual machines › managed runtime
java runtime
0.112008
Practical experiences with Java software transactional memory · PPoPP 2008
Processor architecture and microarchitecture
dynamic optimization
0.112008
Dynamic optimization for efficient strong atomicity · OOPSLA 2008
Runtime systems and virtual machines
language runtime
0.112007
Enabling scalability and performance in a large scale CMP environment · EuroSys 2007
Concurrent programming
memory models
0.112007
Enforcing isolation and ordering in STM · PLDI 2007
Processor architecture and microarchitecture
chip multiprocessor
0.112007
Enabling scalability and performance in a large scale CMP environment · EuroSys 2007
Parallel and multicore computing
concurrent programming
0.112007
Open nesting in software transactional memory · PPoPP 2007
Parallel and multicore computing
parallel programming runtimes
0.112007
Enabling scalability and performance in a large scale CMP environment · EuroSys 2007
Parallel and multicore computing › transactional memory
software transactional memory
0.112007
Open nesting in software transactional memory · PPoPP 2007
Program verification › security property verification
memory safety verification
0.112006
A verifiable SSA program representation for aggressive compiler optimization · POPL 2006
Concurrent programming › transactional memory
nested transactions
0.112006
Compiler and runtime support for efficient software transactional memory · PLDI 2006
Program analysis
program representation
0.112006
A verifiable SSA program representation for aggressive compiler optimization · POPL 2006
Compilers and program optimization › intermediate representation
static single assignment form
0.112006
A verifiable SSA program representation for aggressive compiler optimization · POPL 2006
Programming languages and type systems › type systems
type soundness
0.112006
A verifiable SSA program representation for aggressive compiler optimization · POPL 2006
Compilers and program optimization › memory optimization
cache optimization
0.012003
Fractal symbolic analysis · ACM Trans. Program. Lang. Syst. 2003
Compilers and program optimization
dependence analysis
0.012003
Fractal symbolic analysis · ACM Trans. Program. Lang. Syst. 2003
Compilers and program optimization
loop transformation
0.012003
Fractal symbolic analysis · ACM Trans. Program. Lang. Syst. 2003
Program analysis
symbolic execution
0.012003
Fractal symbolic analysis · ACM Trans. Program. Lang. Syst. 2003
Runtime systems and virtual machines
dynamic compilation
0.012008
Dynamic optimization for efficient strong atomicity · OOPSLA 2008
Compilers and program optimization › compiler optimization
speculative optimization
0.012008
Dynamic optimization for efficient strong atomicity · OOPSLA 2008
Runtime systems and virtual machines › dynamic compilation
just-in-time compilation
0.012006
Compiler and runtime support for efficient software transactional memory · PLDI 2006
Parallel and multicore computing › multiprocessor system
distributed-memory multiprocessor
0.011997
MultiMATLAB Integrating MATLAB with High Performance Parallel Computing · SC 1997
Parallel and multicore computing › parallel computing
parallel computing environments
0.011997
MultiMATLAB Integrating MATLAB with High Performance Parallel Computing · SC 1997
Parallel and multicore computing
parallel programming models
0.011997
MultiMATLAB Integrating MATLAB with High Performance Parallel Computing · SC 1997
Software maintenance and evolution › software reengineering
program restructuring
0.012003
Fractal symbolic analysis · ACM Trans. Program. Lang. Syst. 2003

Methods — techniques the papers use, named apart from their topics

speculation and recovery · 0.2dynamic analysis · 0.2experimental evaluation · 0.1transactional memory · 0.1software transactional memory · 0.1lock-based synchronization · 0.1type system · 0.1runtime optimization · 0.1just-in-time compilation · 0.1SSA · 0.1MPI communication · 0.0
YearPublicationVenuePosition
2008 Fault-safe code motion for type-safe languages
abstract
Compilers for Java and other type-safe languages have historically worked to overcome overheads and constraints imposed by runtime safety checks and precise exception semantics. We instead exploit these safety properties to perform code motion optimizations that are even more aggressive than those possible in unsafe languages such as C++.
Brian R. Murphy, Vijay Menon 0002, Florian T. Schneider, Tatiana Shpeisman, Ali-Reza Adl-Tabatabai
CGO2
2008 A Uniform Transactional Execution Environment for Java
Lukasz Ziarek, Adam Welc, Ali-Reza Adl-Tabatabai, Vijay Menon 0002, Tatiana Shpeisman, Suresh Jagannathan
ECOOP4
2008 Dynamic optimization for efficient strong atomicity
abstract
Transactional memory (TM) is a promising concurrency control alternative to locks. Recent work has highlighted important memory model issues regarding TM semantics and exposed problems in existing TM implementations. For safe, managed languages such as Java, there is a growing consensus towards strong atomicity semantics as a sound, scalable solution. Strong atomicity has presented a challenge to implement efficiently because it requires instrumentation of non-transactional memory accesses, incurring significant overhead even when a program makes minimal or no use of transactions. To minimize overhead, existing solutions require either a sophisticated type system, specialized hardware, or static whole-program analysis. These techniques do not translate easily into a production setting on existing hardware. In this paper, we present novel dynamic optimizations that significantly reduce strong atomicity overheads and make strong atomicity practical for dynamic language environments. We introduce analyses that optimistically track which non-transactional memory accesses can avoid strong atomicity instrumentation, and we describe a lightweight speculation and recovery mechanism that applies these analyses to generate speculatively-optimized but safe code for strong atomicity in a dynamically-loaded environment. We show how to implement these mechanisms efficiently by leveraging existing dynamic optimization infrastructure in a Java system. Measurements on a set of transactional and non-transactional Java workloads demonstrate that our techniques substantially reduce the overhead of strong atomicity from a factor of 5x down to 10% or less over an efficient weak atomicity baseline.
Florian T. Schneider, Vijay Menon 0002, Tatiana Shpeisman, Ali-Reza Adl-Tabatabai
OOPSLA2
2008 Practical experiences with Java software transactional memory
abstract
In this paper, we evaluate the emerging Transactional Memory (TM) area by developing a set of Java transactional memory workloads and studying their performance under a Java Software Transactional Memory (STM) system and comparing them to their lock based counterparts. We provide a detailed performance and memory consumption analysis of the overheads of software transactional memory and transactional workloads within a production quality open source Java Runtime system. Additionally, we detail the impact of the various performance optimizations in both workloads and the underlying runtime system to improving both single thread performance and scalability.
Evgueni Brevnov, Yuri Dolgov, Boris Kuznetsov, Dmitry Yershov, Vyacheslav Shakin, Dong-yuan Chen, Vijay Menon 0002, Suresh Srinivas
PPoPP7
2008 Concurrent GC leveraging transactional memory
abstract
We predict that the ever-growing number of cores on our desktops will require a re-examination of concurrent programming. Two technologies are likely to become mainstream in response: Transactional memory provides a superior programming model to traditional lock-based concurrency, while Concurrent GC can take advantage of multiple cores to eliminate perceptible pauses in desktop applications such as games or Internet telephony. This paper proposes a combination of the two technologies, producing a synergy that improves scalability while eliminating the annoyance of user-perceivable pauses.
Phil McGachey, Ali-Reza Adl-Tabatabai, Richard L. Hudson, Vijay Menon 0002, Bratin Saha, Tatiana Shpeisman
PPoPP4
2008 Practical weak-atomicity semantics for java stm
abstract
As memory transactions have been proposed as a language-level replacement for locks, there is growing need for well-defined semantics. In contrast to database transactions, transaction memory (TM) semantics are complicated by the fact that programs may access the same memory locations both inside and outside transactions. Strongly atomic semantics, where non transactional accesses are treated as implicit single-operation transactions, remain difficult to provide without specialized hardware support or significant performance overhead. As an alternative, many in the community have informally proposed that a single global lock semantics [18,10], where transaction semantics are mapped to those of regions protected by a single global lock, provide an intuitive and efficiently implementable model for programmers.
Vijay Menon 0002, Steven Balensiefer, Tatiana Shpeisman, Ali-Reza Adl-Tabatabai, Richard L. Hudson, Bratin Saha, Adam Welc
SPAA1
2007 Enabling scalability and performance in a large scale CMP environment
abstract
Hardware trends suggest that large-scale CMP architectures, with tens to hundreds of processing cores on a single piece of silicon, are iminent within the next decade. While existing CMP machines have traditionally been handled in the same way as SMPs, this magnitude of parallelism introduces several fundamental challenges at the architectural level and this, in turn, translates to novel challenges in the design of the software stack for these platforms. This paper presents the "Many Core Run Time" (McRT), a software prototype of an integrated language runtime that was designed to explore configurations of the software stack for enabling performance and scalability on large scale CMP platforms. This paper presents the architecture of McRT and discusses our experiences with the system, including experimental evaluation that lead to several interesting, non-intuitive findings, providing key insights about the structure of the system stack at this scale. A key contribution of this paper is to demonstrate how McRT enables near linear improvements in performance and scalability for desktop workloads such as the popular XviD encoder and a set of RMS (recognition, mining, and synthesis) applications. Another key contribution of this work is its use of McRT to explore non-traditional system configurations such as a light-weight executive in which McRT runs on "bare metal" and replaces the traditional OS. Such configurations are becoming an increasingly attractive alternative to leverage heterogeneous computing uints as seen in today's CPU-GPU configurations.
Bratin Saha, Ali-Reza Adl-Tabatabai, Anwar M. Ghuloum, Mohan Rajagopalan, Richard L. Hudson, Leaf Petersen, Vijay Menon 0002, Brian R. Murphy, Tatiana Shpeisman, Eric Sprangle, Anwar Rohillah, Doug Carmean, Jesse Fang
EuroSys7
2007 Enforcing isolation and ordering in STM
abstract
Transactional memory provides a new concurrency control mechanism that avoids many of the pitfalls of lock-based synchronization. High-performance software transactional memory (STM) implementations thus far provide weak atomicity: Accessing shared data both inside and outside a transaction can result in unexpected, implementation-dependent behavior. To guarantee isolation and consistent ordering in such a system, programmers are expected to enclose all shared-memory accesses inside transactions.
Tatiana Shpeisman, Vijay Menon 0002, Ali-Reza Adl-Tabatabai, Steven Balensiefer, Dan Grossman, Richard L. Hudson, Katherine F. Moore, Bratin Saha
PLDI2
2007 Open nesting in software transactional memory
abstract
Transactional memory (TM) promises to simplify concurrent programming while providing scalability competitive to fine-grained locking. Language-based constructs allow programmers to denote atomic regions declaratively and to rely on the underlying system to provide transactional guarantees along with concurrency. In contrast with fine-grained locking, TM allows programmers to write simpler programs that are composable and deadlock-free.
Vijay Menon 0002, Ali-Reza Adl-Tabatabai, Antony L. Hosking, Richard L. Hudson, J. Eliot B. Moss, Bratin Saha, Tatiana Shpeisman
PPoPP2
2006 Compiler and runtime support for efficient software transactional memory
abstract
Programmers have traditionally used locks to synchronize concurrent access to shared data. Lock-based synchronization, however, has well-known pitfalls: using locks for fine-grain synchronization and composing code that already uses locks are both difficult and prone to deadlock. Transactional memory provides an alternate concurrency control mechanism that avoids these pitfalls and significantly eases concurrent programming. Transactional memory language constructs have recently been proposed as extensions to existing languages or included in new concurrent language specifications, opening the door for new compiler optimizations that target the overheads of transactional memory.This paper presents compiler and runtime optimizations for transactional memory language constructs. We present a high-performance software transactional memory system (STM) integrated into a managed runtime environment. Our system efficiently implements nested transactions that support both composition of transactions and partial roll back. Our JIT compiler is the first to optimize the overheads of STM, and we show novel techniques for enabling JIT optimizations on STM operations. We measure the performance of our optimizations on a 16-way SMP running multi-threaded transactional workloads. Our results show that these techniques enable transactional memory's performance to compete with that of well-tuned synchronization.
Ali-Reza Adl-Tabatabai, Brian T. Lewis, Vijay Menon 0002, Brian R. Murphy, Bratin Saha, Tatiana Shpeisman
PLDI3
2006 A verifiable SSA program representation for aggressive compiler optimization
abstract
We present a verifiable low-level program representation to embed, propagate, and preserve safety information in high perfor-mance compilers for safe languages such as Java and C#. Our representation precisely encodes safety information via static single-assignment (SSA) [11, 3] proof variables that are first-class constructs in the program.We argue that our representation allows a compiler to both (1) express aggressively optimized machine-independent code and (2) leverage existing compiler infrastructure to preserve safety information during optimization. We demonstrate that this approach supports standard compiler optimizations, requires minimal changes to the implementation of those optimizations, and does not artificially impede those optimizations to preserve safety. We also describe a simple type system that formalizes type safety in an SSA-style control-flow graph program representation. Through the types of proof variables, our system enables compositional verification of memory safety in optimized code. Finally, we discuss experiences integrating this representation into the machine-independent global optimizer of STARJIT, a high-performance just-in-time compiler that performs aggressive control-flow, data-flow, and algebraic optimizations and is competitive with top production systems.
Vijay Menon 0002, Neal Glew, Brian R. Murphy, Andrew McCreight, Tatiana Shpeisman, Ali-Reza Adl-Tabatabai, Leaf Petersen
POPL1
2003 Fractal symbolic analysis
abstract
Modern compilers restructure programs to improve their efficiency. Dependence analysis is the most widely used technique for proving the correctness of such transformations, but it suffers from the limitation that it considers only the memory locations read and written by a statement without considering what is being computed by that statement. Exploiting the semantics of program statements permits more transformations to be proved correct, and is critical for automatic restructuring of codes such as LU with partial pivoting.One approach to exploiting the semantics of program statements is symbolic analysis and comparison of programs.In principle, this technique is very powerful, but in practice, it is intractable for all but the simplest programs.In this paper, we propose a new form of symbolic analysis and comparison of programs which is appropriate for use in restructuring compilers. Fractal symbolic analysis is an approximate symbolic analysis that compares a program and its transformed version by repeatedly simplifying these programs until symbolic analysis becomes tractable while ensuring that equality of the simplified programs is sufficient to guarantee equality of the original programs.Fractal symbolic analysis combines some of the power of symbolic analysis with the tractability of dependence analysis. We discuss a prototype implementation of fractal symbolic analysis, and show how it can be used to solve the long-open problem of verifying the correctness of transformations required to improve the cache performance of LU factorization with partial pivoting.
Vijay Menon 0002, Keshav Pingali, Nikolay Mateev
ACM Trans. Program. Lang. Syst.1
2001 Fractal symbolic analysis
abstract
Modern compilers perform wholesale restructuring of programs to improve their efficiency. Dependence analysis is the most widely used technique for proving the correctness of such transformations, but it suffers from the limitation that it considers only the memory locations read and written by a statement, and does not assume any particular interpretation for the operations in that statement. Exploiting the semantics of these operations permits more transformations to be proved correct, and is critical for automatic restructuring of codes such as LU with partial pivoting.
Nikolay Mateev, Vijay Menon 0002, Keshav Pingali
ICS2
2000 Left-Looking to Right-Looking and Vice Versa: An Application of Fractal Symbolic Analysis to Linear Algebra Code Restructuring
Nikolay Mateev, Vijay Menon 0002, Keshav Pingali
Euro-Par2
1999 High-level semantic optimization of numerical codes
abstract
This paper presents a mathematical framework to exploit the semantic properties of matrix operations in loop-based numerical codes. The heart of this framework is an algebraic language called the Abstract Matrix Form which a compiler can use to reason about matrix computations in terms of loop nests, high-level matrix operations, and intermediate forms. We demonstrate how this framework may be used to detect and exploit matrix products in loop-based languages such as FORTRAN and MATLAB, and discuss the resulting performance benefits. 1 Introduction Algebraic properties of scalar integer and floating point operations are used by most compilers to optimize programs. These properties enable compilers to reduce of the strength of expressions, enhance the power of common subexpression elimination, and verify the legality of certain loop transformations [2]. Although matrices are also endowed with a rich algebra, it is less common for compilers to exploit matrix algebra to optimize program...
Vijay Menon 0002, Keshav Pingali
International Conference on Supercomputing1
1997 MultiMATLAB Integrating MATLAB with High Performance Parallel Computing
abstract
MATLAB is the most popular scientific computing environment available on uniprocessors today. Unfortunately, no such environment is currently available for multiprocessors. MultiMATLAB [1] is a general extension of the MATLAB environment to any distributed memory multiprocessors. This paper presents a new MultiMATLAB system designed to provide high-performance on multiprocessors while maintaining the functionality and usability of the MATLAB environment. This system will enable users to access high-performance parallel routines from within the MATLAB environment, to extend the environment with new parallel routines, and to use these routines to develop parallel applications with the MATLAB language. We discuss a general MultiMATLAB architecture, present two implementations based upon the MPI communication standard [2], and demonstrate the use of this system. Preliminary results indicate that the MultiMATLAB system can offer the full performance of the underlying multiprocessor to the MATLAB environment.
Vijay Menon 0002, Anne E. Trefethen
SC1