VLDB 2026 Research / reviewers in the wild / expert
Easwaran Raman
dblp:33/4037
· DBLP profile ↗
11ranked-venue papers
3as first author
0since 2021 · last 2014
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 3 first-authorSoftware engineering, systems software and programming languages · 8 · 3 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Software engineering, system software, and programming languages
2 papers |
Compilers and program optimization · 52% Program analysis · 37% Operating systems · 11% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Performance modeling and evaluation · 100% |
Topics — the 7 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Program analysis › error detection
memory leak detection |
0.2 | 1 | 2014 | Automated memory leak detection for production use · ICSE 2014 |
Performance modeling and evaluation
profiling |
0.2 | 1 | 2014 | Automated memory leak detection for production use · ICSE 2014 |
Compilers and program optimization › interprocedural optimization
inlining |
0.1 | 1 | 2006 | A framework for unrestricted whole-program optimization · PLDI 2006 |
Compilers and program optimization
interprocedural optimization |
0.1 | 1 | 2006 | A framework for unrestricted whole-program optimization · PLDI 2006 |
Compilers and program optimization
program specialization |
0.1 | 1 | 2006 | A framework for unrestricted whole-program optimization · PLDI 2006 |
Compilers and program optimization › interprocedural optimization
whole-program optimization |
0.1 | 1 | 2006 | A framework for unrestricted whole-program optimization · PLDI 2006 |
Operating systems › resource management
memory management |
0.1 | 1 | 2014 | Automated memory leak detection for production use · ICSE 2014 |
Methods — techniques the papers use, named apart from their topics
statistical analysis · 0.4instruction sampling · 0.4anomaly detection · 0.4performance monitoring units · 0.2performance monitoring unit · 0.2region formation · 0.1encapsulation · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2014 | Automated memory leak detection for production useabstractThis paper presents Sniper, an automated memory leak detection tool for C/C++ production software. To track the staleness of allocated memory (which is a clue to potential leaks) with little overhead (mostly <3%), Sniper leverages instruction sampling using performance monitoring units available in commodity processors. It also offloads the time- and space-consuming analyses, and works on the original software without modifying the underlying memory allocator; it neither perturbs the application execution nor increases the heap size. The Sniper can even deal with multithreaded applications with very low overhead. In particular, it performs a statistical analysis, which views memory leaks as anomalies, for automated and systematic leak determination. Consequently, it accurately detected real-world memory leaks with no false positive, and achieved an F-measure of 81% on average for 17 benchmarks stress-tested with various memory leaks. Changhee Jung, Sangho Lee 0005, Easwaran Raman, Santosh Pande |
ICSE | 3 |
| 2011 | MAO - An extensible micro-architectural optimizerabstractPerformance matters, and so does repeatability and predictability. Today's processors' micro-architectures have become so complex as to now contain many undocumented, not understood, and even puzzling performance cliffs. Small changes in the instruction stream, such as the insertion of a single NOP instruction, can lead to significant performance deltas, with the effect of exposing compiler and performance optimization efforts to perceived unwanted randomness. This paper presents MAO, an extensible micro-architectural assembly to assembly optimizer, which seeks to address this problem for x86/64 processors. In essence, MAO is a thin wrapper around a common open source assembler infrastructure. It offers basic operations, such as creation or modification of instructions, simple data-flow analysis, and advanced infra-structure, such as loop recognition, and a repeated relaxation algorithm to compute instruction addresses and lengths. This infrastructure enables a plethora of passes for pattern matching, alignment specific optimizations, peep-holes, experiments (such as random insertion of NOPs), and fast prototyping of more sophisticated optimizations. MAO can be integrated into any compiler that emits assembly code, or can be used standalone. MAO can be used to discover micro-architectural details semi-automatically. Initial performance results are encouraging. Robert Hundt, Easwaran Raman, Martin Thuresson, Neil Vachharajani |
CGO | 2 |
| 2008 | Parallel-stage decoupled software pipeliningabstractIn recent years, the microprocessor industry has embraced chip multiprocessors (CMPs), also known as multi-core architectures, as the dominant design paradigm. For existing and new applications to make effective use of CMPs, it is desirable that compilers automatically extract thread-level parallelism from single-threaded applications. DOALL is a popular automatic technique for loop-level parallelization employed successfully in the domains of scientific and numeric computing. While DOALL generally scales well with the number of iterations of the loop, its applicability is limited by the presence of loop-carried dependences. A parallelization technique with greater applicability is decoupled software pipelining (DSWP), which parallelizes loops even in the presence of loop-carried dependences. However, the scalability of DSWP is limited by the size of the loop body and the number of recurrences it contains, which are usually smaller than the loop iteration count. Easwaran Raman, Guilherme Ottoni, Arun Raman, Matthew J. Bridges, David I. August |
CGO | 1 |
| 2008 | Spice: speculative parallel iteration chunk executionabstractThe recent trend in the processor industry of packing multiple processor cores in a chip has increased the importance of automatic techniques for extracting thread level parallelism. A promising approach for extracting thread level parallelism in general purpose applications is to apply memory alias or value speculation to break dependences amongst threads and executes them concurrently. Easwaran Raman, Neil Vachharajani, Ram Rangan, David I. August |
CGO | 1 |
| 2007 | Speculative Decoupled Software Pipelining
Neil Vachharajani, Ram Rangan, Easwaran Raman, Matthew J. Bridges, Guilherme Ottoni, David I. August |
PACT | 3 |
| 2007 | Structure Layout Optimization for Multithreaded ProgramsabstractStructure layout optimizations seek to improve runtime performance by improving data locality and reuse. The structure layout heuristics for single-threaded benchmarks differ from those for multi-threaded applications running on multiprocessor machines, where the effects of false sharing need to be taken into account. In this paper we propose a technique for structure layout transformations for multithreaded applications that optimizes both for improved spatial locality and reduced false sharing, simultaneously. We develop a semi-automatic tool that produces actual structure layouts for multi-threaded programs and outputs the key factors contributing to the layout decisions. We apply this tool on the HP-UX kernel and demonstrate the effects of these transformations for a variety of already highly hand-tuned key structures with different set of properties. We show that naive heuristics can result in massive performance degradations on such a highly tuned application, while our technique generally avoids those pitfalls. The improved structures produced by our tool improve performance by up to 3.2% over a highly tuned baseline Easwaran Raman, Robert Hundt, Sandya Mannarswamy |
CGO | 1 |
| 2006 | A framework for unrestricted whole-program optimizationabstractProcedures have long been the basic units of compilation in conventional optimization frameworks. However, procedures are typically formed to serve software engineering rather than optimization goals, arbitrarily constraining code transformations. Techniques, such as aggressive inlining and interprocedural optimization, have been developed to alleviate this problem, but, due to code growth and compile time issues, these can be applied only sparingly.This paper introduces the Procedure Boundary Elimination (PBE) compilation framework, which allows unrestricted whole-program optimization. PBE allows all intra-procedural optimizations and analyses to operate on arbitrary subgraphs of the program, regardless of the original procedure boundaries and without resorting to inlining. In order to control compilation time, PBE also introduces novel extensions of region formation and encapsulation. PBE enables targeted code specialization, which recovers the specialization benefits of inlining while keeping code growth in check. This paper shows that PBE attains better performance than inlining with half the code growth. Spyridon Triantafyllis, Matthew J. Bridges, Easwaran Raman, Guilherme Ottoni, David I. August |
PLDI | 3 |
| 2005 | Practical and Accurate Low-Level Pointer AnalysisabstractPointer analysis is traditionally performed once, early in the compilation process, upon an intermediate representation (IR) with source-code semantics. However, performing pointer analysis only once at this level imposes a phase-ordering constraint, causing alias information to become stale after subsequent code transformations. Moreover, high-level pointer analysis cannot be used at link time or run time, where the source code is unavailable. This paper advocates performing pointer analysis on a low-level intermediate representation. We present the first context-sensitive and partially flow-sensitive points-to analysis designed to operate at the assembly level. As we will demonstrate, low-level pointer analysis can be as accurate as high-level analysis. Additionally, our low-level pointer analysis also enables a quantitative comparison of propagating high-level pointer analysis results through subsequent code transformations, versus recomputing them at the low level. We show that, for C programs, the former practice is considerably less accurate than the latter. Bolei Guo, Matthew J. Bridges, Spyridon Triantafyllis, Guilherme Ottoni, Easwaran Raman, David I. August |
CGO | 5 |
| 2005 | Integrating a New Cluster Assignment and Scheduling Algorithm into an Experimental Retargetable Code Generation Framework
K. Vasanta Lakshmi, Deepak Sreedhar, Easwaran Raman, Priti Shankar |
HiPC | 3 |
| 2004 | Exposing Memory Access Regularities Using Object-Relative Memory ProfilingabstractMemory profiling is the process of characterizing a program's memory behavior by observing and recording its response to specific input sets. Relevant aspects of the program's memory behavior may then be used to guide memory optimizations in an aggressively optimizing compiler. In general, memory access behavior has eluded meaningful characterization because of confounding artifacts from memory allocators, linker data layout, and OS memory management. Since these artifacts may change from run to run, memory access patterns may appear different in each run even for the same input set. Worse, regular memory access behavior such as linked list traversals appear to have no structure. We present object-relative translation and decomposition techniques to eliminate these artifacts and to expose previously obscured memory access patterns. To demonstrate the potential of these ideas, we implement two different memory profilers targeted at different sets of applications. These profilers outperform the existing ones in terms of profile size and useful information per byte of data. The first profiler is a lossless profiler, called WHOMP, which uses object-relativity to achieve a 22% better compression than the previously best known scheme. The second profiler, called LEAP, uses lossy compression to get highly compact profiles while providing useful information to the targeted applications. LEAP correctly characterizes the memory alias rates for 56% more instruction pairs than the previously best known scheme with a practical running time. Artem Pyatakov, Alexey Spiridonov, Easwaran Raman, Douglas W. Clark, David I. August |
CGO | 4 |
| 2001 | A Reconfigurable Co-Processor for Variable Long Precision Arithmetic Using Indian Algorithms
Ranjani Parthasarathi, Easwaran Raman, Karthik Sankaranarayanan, Lakshmi N. Chakrapani |
FCCM | 2 |