EDBT 2026 Demo / reviewers in the wild / expert
Vijayaraghavan Soundararajan
dblp:24/53
· DBLP profile ↗
6ranked-venue papers
3as first author
0since 2021 · last 2014
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 2 first-authorSoftware engineering, systems software and programming languages · 3 · 2 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
4 papers |
Cloud and datacenter computing · 49% Memory systems · 31% Performance modeling and evaluation · 8% |
Topics — the 16 heaviest of 16, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Cloud and datacenter computing
virtualization |
0.1 | 1 | 2010 | The impact of management operations on the virtualized datacenter · ISCA 2010 |
Cloud and datacenter computing › datacenter architecture
virtualized datacenter |
0.1 | 1 | 2010 | The impact of management operations on the virtualized datacenter · ISCA 2010 |
Memory systems
cache coherence |
0.0 | 2 | 1999 | A Quantitative Analysis of the Performance and Scalability of Distributed Shared Memory · IEEE Trans. Computers 1999 Flexible Use of Memory for Replication/Migration in Cache-Coherent DSM Multiprocessors · ISCA 1998 |
Cloud and datacenter computing
datacenter architecture |
0.0 | 1 | 2010 | The impact of management operations on the virtualized datacenter · ISCA 2010 |
Memory systems
cache |
0.0 | 1 | 2000 | MemorIES: A Programmable, Real-Time Hardware Emulation Tool for Multiprocessor Server Design · ASPLOS 2000 |
Memory systems › cache design
cache configuration |
0.0 | 1 | 2000 | MemorIES: A Programmable, Real-Time Hardware Emulation Tool for Multiprocessor Server Design · ASPLOS 2000 |
Electronic design automation
hardware emulation |
0.0 | 1 | 2000 | MemorIES: A Programmable, Real-Time Hardware Emulation Tool for Multiprocessor Server Design · ASPLOS 2000 |
Performance modeling and evaluation
simulation |
0.0 | 1 | 2000 | MemorIES: A Programmable, Real-Time Hardware Emulation Tool for Multiprocessor Server Design · ASPLOS 2000 |
Memory systems › cache coherence
cache coherence protocol |
0.0 | 1 | 1999 | A Quantitative Analysis of the Performance and Scalability of Distributed Shared Memory · IEEE Trans. Computers 1999 |
Memory systems
data locality |
0.0 | 1 | 1998 | Flexible Use of Memory for Replication/Migration in Cache-Coherent DSM Multiprocessors · ISCA 1998 |
Distributed systems
data replication and migration |
0.0 | 1 | 1998 | Flexible Use of Memory for Replication/Migration in Cache-Coherent DSM Multiprocessors · ISCA 1998 |
Parallel and multicore computing › multiprocessor system
shared-memory multiprocessor |
0.0 | 2 | 1999 | A Quantitative Analysis of the Performance and Scalability of Distributed Shared Memory · IEEE Trans. Computers 1999 Flexible Use of Memory for Replication/Migration in Cache-Coherent DSM Multiprocessors · ISCA 1998 |
Performance modeling and evaluation
trace analysis |
0.0 | 1 | 2000 | MemorIES: A Programmable, Real-Time Hardware Emulation Tool for Multiprocessor Server Design · ASPLOS 2000 |
Performance modeling and evaluation
workload characterization |
0.0 | 1 | 2000 | MemorIES: A Programmable, Real-Time Hardware Emulation Tool for Multiprocessor Server Design · ASPLOS 2000 |
Memory systems › shared memory
distributed shared memory |
0.0 | 1 | 1999 | A Quantitative Analysis of the Performance and Scalability of Distributed Shared Memory · IEEE Trans. Computers 1999 |
Memory systems › non-uniform memory access
CC-NUMA |
0.0 | 1 | 1998 | Flexible Use of Memory for Replication/Migration in Cache-Coherent DSM Multiprocessors · ISCA 1998 |
Methods — techniques the papers use, named apart from their topics
workload characterization · 0.1FPGA-based emulation · 0.0simulation · 0.0protocol processors · 0.0programmable memory controller · 0.0directory protocol comparison · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2014 | Applying Graph Databases to Cloud Management: An ExplorationabstractGraph databases have become increasingly popular for a variety of uses ranging from modeling online code repositories to tracking software engineering dependencies. These areas use graph databases because many of their problems can be expressed in terms of graph traversals. Recent work has applied graph databases to virtualization management, noting that many IT questions can also be expressed as graph traversals. In this paper, we study another area in which graphs are valuable: reporting and auditing in cloud infrastructure. We first examine cloud infrastructure and map its data model to a graph. Building upon this model, we recast a number of reporting queries in terms of graph traversals. We then modify the model both for performance and for accommodating additional use cases related to cloud computing, including migration from private to hybrid clouds. Our results show that while a graph backend makes it straightforward to formulate certain kinds of queries, a naive mapping of graphs to a graph database can result in poor performance. Utilizing knowledge of the problem domain and restructuring the graph can provide dramatic gains in performance and make a graph database feasible for such queries. Vijayaraghavan Soundararajan, Shishir Kakaraddi |
IC2E | 1 |
| 2010 | The impact of management operations on the virtualized datacenterabstractVirtualization has the potential to dramatically reduce the total cost of ownership of datacenters and increase the flexibility of deployments for general-purpose workloads. If present trends continue, the datacenter of the future will be largely virtualized. The base platform in such a datacenter will consist of physical hosts that run hypervisors, and workloads will run within virtual machines on these platforms. From a system management perspective, the virtualized environment enables a number of new workflows in the datacenter. These workflows involve operations on the physical hosts themselves, such as upgrading the hypervisor, as well as operations on the virtual machines, such as reconfiguration or reverting from snapshots. While traditional datacenter design has focused on the cost vs. capability tradeoffs for the end-user applications running in the datacenter, we argue that the management workload from these workflows must be factored into the design of the virtualized datacenter. Vijayaraghavan Soundararajan, Jennifer M. Anderson |
ISCA | 1 |
| 2000 | MemorIES: A Programmable, Real-Time Hardware Emulation Tool for Multiprocessor Server DesignabstractModern system design often requires multiple levels of simulation for design validation and performance debugging. However, while machines have gotten faster, and simulators have become more detailed, simulation speeds have not tracked machine speeds, As a result, it is difficult to simulate realistic problem sizes and hardware configurations for a target machine. Instead, researchers have focussed on developing sealing methodologies and running smaller problem sizes and configurations that attempt to represent the behavior of the real problem. Given the increasing size of problems today, it is unclear whether such an approach yields accurate results. Moreover, although commercial workloads are prevalent and important in today's marketplace, many simulation tools are unable to adequately profile such applications, let alone for realistic sizes.In this paper we present a hardware-based emulation tool that can be used to aid memory system designers. Our focus is on the memory system because the ever-widening gap between processor and memory speeds means that optimizing the memory subsystem is critical for performance. We present the design of the Memory Instrumentation and Emulation System (MemoriES). MemoriES is a programmable tool designed using FPGAs and SDRAMs. It plugs into an SMP bus to perform on-line emulation of several cache configurations, structures and protocols while the system is running real-life workloads in real-time, without any slowdown in application execution speed. We demonstrate its usefulness in several case studies, and find several important results. First, using traces to perform system evaluation can lead to incorrect results (off by 100% or more in some cases) if the trace size is not sufficiently large. Second. MemoriES is able to detect performance problems by profiling miss behavior over the entire course of a run, rather than relying on a small interval of time. Finally, we observe that previous studies of SPLASH2 applications using scaled application sizes can result in optimistic miss rates relative to real sizes on real machines, providing potentially misleading data when used for design evaluation. Ashwini K. Nanda, Kwok-Ken Mak, Krishnan Sugavanam, Ramendra K. Sahoo, Vijayaraghavan Soundararajan, T. Basil Smith |
ASPLOS | 5 |
| 1999 | A Quantitative Analysis of the Performance and Scalability of Distributed Shared MemoryabstractScalable cache coherence protocols have become the key technology for creating moderate to large-scale shared-memory multiprocessors. Although the performance of such multiprocessors depends critically on the performance of the cache coherence protocol, little comparative performance data is available. Existing commercial implementations use a variety of different protocols, including bit-vector/coarse-vector protocols, SCI-based protocols, and COMA protocols. Using the programmable protocol processor of the Stanford FLASH multiprocessor, we provide a detailed, implementation-oriented evaluation of four popular cache coherence protocols. In addition to measurements of the characteristics of protocol execution (e.g., memory overhead, protocol execution time, and message count) and of overall performance, we examine the effects of scaling the processor count from 1 to 128 processors. Surprisingly, the optimal protocol changes for different applications and can change with processor count even within the same application. These results help identify the strengths of specific protocols and illustrate the benefits of providing flexibility in the choice of cache coherence protocol. Mark A. Heinrich, Vijayaraghavan Soundararajan, John L. Hennessy, Anoop Gupta |
IEEE Trans. Computers | 2 |
| 1998 | Flexible Use of Memory for Replication/Migration in Cache-Coherent DSM MultiprocessorsabstractGiven the limitations of bus-based multiprocessors, CC-NUMA is the scalable architecture of choice for shared-memory machines. The most important characteristic of the CC-NUMA architecture is that the latency to access data on a remote node is considerably larger than the latency to access local memory. On such machines, good data locality can reduce memory stall time and is therefore a critical factor in application performance. In this paper we study the various options available to system designers to transparently decrease the fraction of data misses serviced remotely. This work is done in the context of the Stanford FLASH multiprocessor. FLASH is unique in that each node has a single pool of DRAM that can be used in a variety of ways by the programmable memory controller. We use the programmability of FLASH to explore different options for cache-coherence and data-locality in compute-server workloads. First, we consider two protocols for providing base cache-coherence, one with centralized directory information (dynamic pointer allocation) and another with distributed directory information (SCI). While several commercial systems are based on SCI, we find that a centralized scheme has superior performance. Next, we consider different hardware and software techniques that use some or all of the local memory in a node to improve data locality. Finally, we propose a hybrid scheme that combines hardware and software techniques. These schemes work on the same base platform with both user and kernel references from the workloads. The paper thus offers a realistic and fair comparison of replication/migration techniques that has not previously been feasible. Vijayaraghavan Soundararajan, Mark A. Heinrich, Ben Verghese, Kourosh Gharachorloo, Anoop Gupta, John L. Hennessy |
ISCA | 1 |
| 1994 | Performance evaluation of hybrid hardware and software distributed shared memory protocolsabstractHardware distributed shared memory (DSM) systems efficiently support fine grain sharing of data by maintaining coherence at the level of individual cache lines and providing automatic replication in processor caches. Software DSM systems, on the other hand, amortize high communication costs by maintaining coherence at coarser granularities and replicating data at the level of local main memories. Even though software DSM systems have traditionally been targeted towards loosely coupled environments, some of the techniques are potentially useful in the context of tightly coupled multiprocessors. In particular, communicating data at a coarse grain can sometimes be more efficient than transferring the data as individual cache lines. Furthermore, replication in local memories can accommodate applications with larger working sets as compared to replication in processor caches only. Therefore, combining the two techniques in a hybrid protocol can potentially exploit the benefits of each approach. Rohit Chandra, Kourosh Gharachorloo, Vijayaraghavan Soundararajan, Anoop Gupta |
International Conference on Supercomputing | 3 |