Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Srinidhi Varadarajan

dblp:39/2498 · DBLP profile ↗
← Back
18ranked-venue papers
6as first author
0since 2021 · last 2012
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 11 · 2 first-authorComputer networks · 6 · 4 first-authorSoftware engineering, systems software and programming languages · 2

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
5 papers
Parallel and multicore computing · 46% Distributed systems · 44% Embedded and real-time systems · 3%
Computer networks
3 papers
Internet architecture and protocols · 69% Network management and operations · 24% Datacenter networks · 7%

Topics — the 20 heaviest of 22, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Parallel and multicore computing
parallel programming models
0.112011
Exploiting coarse-grain speculative parallelism · OOPSLA 2011
Parallel and multicore computing
speculative parallelization
0.112011
Exploiting coarse-grain speculative parallelism · OOPSLA 2011
Distributed systems
distributed coordination
0.112006
Poster reception - DejaVu: transparent user-level checkpointing, migration and recovery for distributed systems · SC 2006
Distributed systems
fault tolerance
0.112006
Poster reception - DejaVu: transparent user-level checkpointing, migration and recovery for distributed systems · SC 2006
Distributed systems › distributed system architecture › distributed operating systems
process migration
0.112006
Poster reception - DejaVu: transparent user-level checkpointing, migration and recovery for distributed systems · SC 2006
Distributed systems › fault tolerance
rollback recovery
0.112006
Poster reception - DejaVu: transparent user-level checkpointing, migration and recovery for distributed systems · SC 2006
Distributed systems › fault tolerance › checkpointing
transparent checkpointing
0.112006
Poster reception - DejaVu: transparent user-level checkpointing, migration and recovery for distributed systems · SC 2006
Parallel and multicore computing › parallel programming models › directive-based programming
OpenMP extensions
0.012011
Exploiting coarse-grain speculative parallelism · OOPSLA 2011
Network management and operations › fault management
fault detection and recovery
0.011999
Automatic Fault Detection and Recovery in Real Time Switched Ethernet Networks · INFOCOM 1999
Embedded and real-time systems › real-time communication
real-time networks
0.011999
Automatic Fault Detection and Recovery in Real Time Switched Ethernet Networks · INFOCOM 1999
Internet architecture and protocols › local area network
switched ethernet
0.011998
EtheReal: A Host-Transparent Real-Time Fast Ethernet Switch · ICNP 1998
Cloud and datacenter computing › quality of service
bandwidth guarantee
0.011998
EtheReal: A Host-Transparent Real-Time Fast Ethernet Switch · ICNP 1998
Parallel and multicore computing
MPI
0.012006
Poster reception - DejaVu: transparent user-level checkpointing, migration and recovery for distributed systems · SC 2006
Parallel and multicore computing
parallel programming models and runtimes
0.012006
Poster reception - DejaVu: transparent user-level checkpointing, migration and recovery for distributed systems · SC 2006
Internet architecture and protocols
ATM networks
0.011997
Design and Evaluation of a DRAM-based Shared Memory ATM Switch · SIGMETRICS 1997
Internet architecture and protocols › ATM networks
ATM switch architecture
0.011997
Design and Evaluation of a DRAM-based Shared Memory ATM Switch · SIGMETRICS 1997
Internet architecture and protocols
quality of service
0.021999
Automatic Fault Detection and Recovery in Real Time Switched Ethernet Networks · INFOCOM 1999
EtheReal: A Host-Transparent Real-Time Fast Ethernet Switch · ICNP 1998
Datacenter networks
bandwidth guarantee
0.011999
Automatic Fault Detection and Recovery in Real Time Switched Ethernet Networks · INFOCOM 1999
Performance modeling and evaluation
simulation
0.011997
Design and Evaluation of a DRAM-based Shared Memory ATM Switch · SIGMETRICS 1997
Interconnection networks and networks-on-chip
switching network
0.011997
Design and Evaluation of a DRAM-based Shared Memory ATM Switch · SIGMETRICS 1997

Methods — techniques the papers use, named apart from their topics

surrogate code blocks · 0.1runtime system design · 0.1state capture instrumentation · 0.1incremental checkpointing · 0.1prototype implementation · 0.0performance evaluation · 0.0simulation · 0.0
YearPublicationVenuePosition
2012 Transparent runtime deadlock elimination
abstract
Thread based concurrent programming is hard due to the potential of concurrency bugs (e.g., data races, atomicity violations, deadlocks, and order violations). While data races and atomicity violations can be ameliorated with appropriate synchronization (a non-trivial problem in itself !), deadlocks require fairly complex avoidance techniques which may fail when the order of lock acquisition is not known apriori [2]. The goal of this research is to present an efficient and practical system that transparently detects and eliminates deadlocks in real-world multi-threaded applications.
Hari K. Pyla, Srinidhi Varadarajan
PACT2
2011 Is It Time to Rethink Distributed Shared Memory Systems?
abstract
We present core elements of Samhita, a new user level software distributed shared memory (DSM) system. Our work is motivated by two observations. First, the rise of many-core architectures is producing a growing emphasis on threaded codes to achieve performance. Second, architectural trends, especially in high performance interconnects, suggest a new look at overcoming the bottlenecks that have hindered DSM performance. Samhita leverages the capabilities of remote direct memory access (RDMA) interconnects, and views the problem of providing a shared global address space as a cache management problem. Performance results on two 256 processor clusters demonstrate scalability on micro benchmarks and two real applications. The results are the largest scale tests and achieve the highest performance of any DSM system reported to date.
Bharath Ramesh 0004, Calvin J. Ribbens, Srinidhi Varadarajan
ICPADS3
2011 Exploiting coarse-grain speculative parallelism
abstract
Speculative execution at coarse granularities (e.g., code-blocks, methods, algorithms) offers a promising programming model for exploiting parallelism on modern architectures. In this paper we present Anumita, a framework that includes programming constructs and a supporting runtime system to enable the use of coarse-grain speculation to improve program performance, without burdening the programmer with the complexity of creating, managing and retiring speculations. Speculations may be composed by specifying surrogate code blocks at any arbitrary granularity, which are then executed concurrently, with a single winner ultimately modifying program state. Anumita provides expressive semantics for winner selection that go beyond time to solution to include user-defined notions of quality of solution. Anumita can be used to improve the performance of hard to parallelize algorithms whose performance is highly dependent on input data. Anumita is implemented as a user-level runtime with programming interfaces to C, C++, Fortran and as an OpenMP extension. Performance results from several applications show the efficacy of using coarse-grain speculation to achieve (a) robustness when surrogates fail and (b) significant speedup over static algorithm choices.
Hari K. Pyla, Calvin J. Ribbens, Srinidhi Varadarajan
OOPSLA3
2010 Avoiding deadlock avoidance
abstract
The evolution of processor architectures from single core designs with increasing clock frequencies to multi-core designs with relatively stable clock frequencies has fundamentally altered application design. Since application programmers can no longer rely on clock frequency increases to boost performance, over the last several years, there has been significant emphasis on application level threading to achieve performance gains. A core problem with concurrent programming using threads is the potential for deadlocks. Even well-written codes that spend an inordinate amount of effort in deadlock avoidance cannot always avoid deadlocks, particularly when the order of lock acquisitions is not known a priori. Furthermore, arbitrarily composing lock based codes may result in deadlock - one of the primary motivations for transactional memory. In this paper, we present a language independent runtime system called Sammati that provides automatic deadlock detection and recovery for threaded applications that use the POSIX threads (pthreads) interface - the de facto standard for UNIX systems. The runtime is implemented as a pre-loadable library and does not require either the application source code or recompiling/relinking phases, enabling its use for existing applications with arbitrary multi-threading models. Performance evaluation of the runtime with unmodified SPLASH, Phoenix and synthetic benchmark suites shows that it is scalable, with speedup comparable to baseline execution with modest memory overhead.
Hari K. Pyla, Srinidhi Varadarajan
PACT2
2007 Tempest: A portable tool to identify hot spots in parallel code
abstract
Compute clusters are consuming more power at higher densities than ever before. This results in increased thermal dissipation, the need for powerful cooling systems, and ultimately a reduction in system reliability as temperatures increase. Over the past several years, the research community has reacted to this problem by producing software tools such as HotSpot and Mercury to estimate system thermal characteristics and validate thermal-management techniques. While these tools are flexible and useful, they suffer several limitations. For the average user such simulation tools can be cumbersome to use. These tools may take significant time and expertise to port to different systems. Lastly, such tools produce significant detail and accuracy at the expense of execution time enough to prohibit iterative testing. We propose a fast, easy to use, accurate, portable software tool called Tempest (for temperature estimator) that leverages emergent thermal sensors to enable user profiling, evaluating, and reducing the thermal characteristics of systems and applications. In this paper, we illustrate the use of Tempest to analyze the thermal effects of various parallel benchmarks in clusters.
Kirk W. Cameron, Hari K. Pyla, Srinidhi Varadarajan
ICPP3
2007 The Adaptive Code Kitchen: Flexible Tools for Dynamic Application Composition
abstract
Driven by the increasing componentization of scientific codes, the deployment of high-end system infrastructures such as the grid, and the desire to support high level problem solving primitives, application composition systems have become prevalent in computational science practice. We present the adaptive code kitchen which, as the name connotes, is a loose collection of capabilities to help realize complex adaptive composition scenarios. These include function interception, continuation modification, dynamic process checkpointing and rollback, and runtime recommendation. Using these broad primitives, a computational scientist can specify many 'recipes' of adaptivity as complete control systems around native object codes. Runtime systems support then enables loading and linking of native code components, monitoring of performance indicators, consulting a recommender system for algorithmic decisions, and dynamically updating application components in response to the recommendations. We present the architecture of the adaptive code kitchen and the key enabling technologies with brief mention of the applications that will be investigated henceforth during the course of the project.
Pilsung Kang 0002, Mike Heffner, Joy Mukherjee, Naren Ramakrishnan, Srinidhi Varadarajan, Calvin J. Ribbens, Danesh K. Tafti
IPDPS5
2007 DejaVu: Transparent User-Level Checkpointing, Migration, and Recovery for Distributed Systems
abstract
In this paper, we present a new fault tolerance system called DejaVu for transparent and automatic checkpointing, migration, and recovery of parallel and distributed applications. DejaVu provides a transparent parallel checkpointing and recovery mechanism that recovers from any combination of systems failures without any modification to parallel applications or the OS. It uses a new runtime mechanism for transparent incremental checkpointing that captures the least amount of state needed to maintain global consistency and provides a novel communication architecture that enables transparent migration of existing MPI codes, without source-code modifications. Performance results from the production-ready implementation show less than 5% overhead in real-world parallel applications with large memory footprints.
Joseph F. Ruscio, Michael A. Heffner, Srinidhi Varadarajan
IPDPS3
2006 Poster reception - DejaVu: transparent user-level checkpointing, migration and recovery for distributed systems
abstract
We present a new fault tolerance system, DejaVu, for transparent and automatic checkpointing, migration and recovery of parallel and distributed applications. DejaVu has several novel features. First, it provides a transparent parallel checkpointing and recovery mechanism that recovers from any combination of systems failures without modification to parallel applications or the underlying operating system. Second, it uses a novel instrumentation and state capture mechanism that transparently captures application state. Third, it uses a new runtime mechanism for transparent incremental checkpointing, capturing the least amount of state needed to maintain global consistency. Finally, it provides a novel communication architecture that enables transparent migration of existing MPI codes, without source-code modifications. DejaVu has been implemented for 32 bit and 64 bit Linux platforms on x86 processors interconnected over Infiniband or Gigabit Ethernet networks. Performance results from the production-ready implementation shows less than 5% overhead with real-world parallel applications with large memory footprints.
Joseph F. Ruscio, Michael A. Heffner, Srinidhi Varadarajan
SC3
2005 Novel runtime systems support for adaptive compositional modeling in PSEs
Srinidhi Varadarajan, Naren Ramakrishnan
Future Gener. Comput. Syst.1
2004 System X: Building the Virginia Tech Supercomputer
abstract
Provides an abstract of the keynote presentation and a brief professional biography of the presenter. The complete presentation was not made available for publication as part of the conference proceedings.
Srinidhi Varadarajan
ICCCN1
2004 Admission control by implicit signaling in support of voice over IP over ADSL
Abhishek Ram, Luiz A. DaSilva, Srinidhi Varadarajan
Comput. Networks3
2003 Reinforcing reachable route
Srinidhi Varadarajan, Naren Ramakrishnan, Muthukumar Thirunavukkarasu
Comput. Networks1
2002 Assessment of voice over IP as a solution for voice over ADSL
abstract
The paper compares VoATM and VoIP in terms of their suitability for carrying voice traffic over DSL. ATM is currently the preferred protocol due to its built-in QoS mechanisms. Through simulations, we show that IP QoS mechanisms can also be used to achieve comparable performance for voice traffic over DSL. Our performance metrics are the end-to-end delay of voice packets across the DSL access network and the bandwidth requirements of a voice call. We also propose an implicit signaling mechanism to provide admission control for individual voice calls over DSL. We implement a simulation model that uses our mechanism and perform simulations to verify its effectiveness We conclude that by incorporating appropriate QoS mechanisms and implicit signaling, it is possible to achieve performance for VoIP comparable to that provided by ATM. In this case, the ubiquity of IP makes it a very attractive candidate for future deployments of VoDSL.
Abhishek Ram, Luiz A. DaSilva, Srinidhi Varadarajan
GLOBECOM3
2001 Experiences with EtheReal: a fault-tolerant real-time Ethernet switch
abstract
We present our experiences with the implementation of a real-time Ethernet switch called EtheReal. EtheReal provides three innovations for real-time traffic over switched Ethernet networks. First, EtheReal delivers connection oriented hard bandwidth guarantees without requiring any changes to the end host operating system and network hardware/software. For ease of deployment by commercial vendors, EtheReal is implemented in software over Ethernet switches, with no special hardware requirements. QoS support is contained within two modules, switches and end-host user level libraries that expose a socket like API to real time applications. Secondly, EtheReal provides automatic fault detection and recovery mechanisms that operate within the constraints of a real-time network. Finally EtheReal supports server-side push applications with a guaranteed bandwidth link-layer multicast scheme. Performance results from the implementation show that EtheReal switches deliver bandwidth guarantees to real time-applications within 0.6% of the contracted value, even in the presence of interfering best-effort traffic between the same pair of communicating hosts.
Srinidhi Varadarajan
ETFA (1)1
1999 Automatic Fault Detection and Recovery in Real Time Switched Ethernet Networks
abstract
EtheReal is a real-time fast Ethernet switch architecture that provides bandwidth guarantees to distributed multimedia applications without OS or hardware modifications on the host machines. It implements true link-layer multicast, and offers a natural match to support network-layer QoS protocols such as RSVP. Because real-time performance guarantees fundamentally require state to be installed inside the network, link/switch failures could lead to significant disruption to the QoS promised to the user applications. This paper describes the fault detection and recovery mechanism supported by the EtheReal architecture, and reports on the performance measurements of the initial prototype implementation. The heart of EtheReal's fault detection and recovery mechanism is a fast spanning tree reconfiguration algorithm to reduce the total fault recovery time, and a delayed link inactivation scheme that allows real-time connections which are not affected by the failed links/switches to continue to exist, even though some of the links are marked as "blocked" in the new spanning tree topology. Measurements on the prototype show that the fault detection and recovery time on a network whose diameter is 10 hops are 220 ms and 31 ms, respectively. This combined delay corresponds to a minor jitter in real-time audio/video communication, and is a significant improvement over the standard IEEE 802.1d implementation, which takes on the order of 30 sec.
Srinidhi Varadarajan, Tzi-cker Chiueh
INFOCOM1
1998 EtheReal: A Host-Transparent Real-Time Fast Ethernet Switch
abstract
Distributed multimedia applications require guaranteed quality of service (QoS) from the underlying networks. This paper describes the design, implementation, and evaluation of a real-time Fast Ethernet switch called EtheReal that provides bandwidth guarantees to real-time applications running on Ethernet without modification to the hardware and operating system on the host machines. At the heart of the EtheReal switch architecture is a novel real-time connection setup protocol that is completely transparent to the host machine's OS, and thus allows the switch to be deployed in a network of heterogeneous machines running different OS platforms. The only dependency of EtheReal on the hosts is their support for TCP/UDP/IP. In addition, the EtheReal switch uses an Ethernet address swapping technique for real-time packets, similar to ATM to avoid the complexity of maintaining a global connection ID space. Because of the inherent CRC support built into Ethernet hardware, this technique significantly reduces the total processing overhead compared to address swapping at higher network layers. The current EtheReal switch prototype is fully operational; and is built with off-the-shelf PC hardware. It is capable of supporting up to 640 Mbits/sec across 4 ports, and has a per-hop connection establishment overhead of 0.1-0.6 msec and a 10 /spl mu/s non-real-time packet latency.
Srinidhi Varadarajan, Tzi-cker Chiueh
ICNP1
1998 Design, Implementation, and Evaluation of a Parallel Index Server for Shape Image Database
abstract
Describes the design, implementation and evaluation of a parallel indexer called PAMIS (PArallel Multimedia Index Server) for a polygonal 2D shape image database. PAMIS is based on a shape representation scheme called the turning function, which exhibits the desirable properties of position, scale and rotation invariance, and has a similarity metric function that satisfies the triangular inequality, which is required for efficient database indexing. Because the goal of the PAMIS project is to support "like-this" image queries, the indexing scheme we chose, the vantage-point tree (VPT), uses relative rather than absolute distance values to organize the database elements for efficient nearest-neighbor searching. We have successfully implemented PAMIS on a network of workstations to exploit the I/O and computation parallelism inherent in the VPT algorithm. We found that it is preferable to make the VPT node size as small as possible in order to have a lean and deep VPT structure, and the best-case scheduling strategy performs the best among the scheduling strategies considered. Overall, the performance of the VPT algorithm scales very well with the number of processors, and the indexing efficiency (defined as the percentage of database elements touched by the search) of PAMIS is 6% and 39% for "good" queries that ask for 1 and 50 nearest neighbors, respectively.
Tzi-cker Chiueh, Srinidhi Varadarajan
ICPADS3
1997 Design and Evaluation of a DRAM-based Shared Memory ATM Switch
abstract
Beluga is a single-chip switch architecture specifically targeted at local area ATM networks, and it features three architectural innovations. First, an interconnection hierarchy composed of multiple switching fabrics is built into the chip to provide both low-latency cell transfer when the traffic is light and low cell drop rate under heavy load. Secondly, to improve silicon efficiency, Beluga is based on shared memory architecture, and the buffers are implemented using DRAM rather than SRAM technology. Heavy interleaving and selective invalidation are used to address long latency and periodic refreshing problems, respectively. Thirdly, Beluga supports multicast with minimal physical bit replication. It also separates support for unicast and multicast cells to optimize for the common case, where multicast cells occur infrequently. This paper describes the design details of Beluga and the results of a comprehensive simulation study to quantify the performance impact of each of its architectural features. The most important result from this research is that DRAM-based buffer implementation significantly reduces the cell-drop rate during heavy while exhibiting almost identical cell latency to SRAM-based implementation during light load. Therefore, we believe DRAM makes an attractive alternative for switch buffer implementation, especially for single-chip architecture such as Beluga.
Tzi-cker Chiueh, Srinidhi Varadarajan
SIGMETRICS2