Umakishore Ramachandran

dblp:r/URamachandran · DBLP profile ↗
← Back
81ranked-venue papers
19as first author
11since 2021 · last 2026
0000-0003-4071-684XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 44 · 12 first-author · 5 since 2021Software engineering, systems software and programming languages · 15 · 4 first-author · 2 since 2021Computer networks · 11 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 first-authorArtificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Harvesting Spare CPU Resources in Container Systems
Adam Hall, Anirudh Sarma, Esha Choukse, Umakishore Ramachandran, Sameh Elnikety
NSDI4
2025 Bringing Context to the Underserved: Rethinking Context-Aware Design to Bridge the Digital Divide
Summit Shrestha, Josiah D. Hester, Ashutosh Dhekne, Umakishore Ramachandran, Alex Cabral
COMPASS4
2025 DroEdgeEM: A Drone-Edge Collaborative Emulation Platform for Emerging Situation-aware Applications
abstract
Unmanned Aerial Vehicles (UAVs) are increasingly used for surveillance, mapping, delivery, and search and rescue. These applications require continuous data streams processed through low-latency pipelines to remain situation-aware. Despite advances in hardware and software, UAV uptime is constrained by limited battery or fuel. Adding onboard compute further shortens flight time; where extra payload increases propulsion demand, and the compute itself draws power, so either UAVs carry limited compute or are treated as mobile sensors. Emerging 5G/edge deployments hold promise for supporting such compute limited drones by enabling drone-to-edge task offloading and collaboration for latency-critical processing. However, such infrastructure support is not yet widespread, complicating design and evaluation. Practical and privacy constraints also limit real-world testing, pushing much current research toward theoretical models with uncertain practicality. A key gap is the absence of realistic emulator that couples mobile UAVs and edge nodes within a dynamic virtual world (energy, location, mobility). This work envisions addressing that gap by: 1)adding emulation support on top of existing simulators to emulate necessary infrastructure (mobile drones, edge servers, charging stations) that support data streaming, processing, and control while managing time-varying state; 2)integrating user-defined plug-and-play models (e.g. energy and FPS) to realize realistic emulation behaviors of heterogeneous infrastructure with varying capabilities; and 3)exposing high-level APIs for integrations with application processing pipelines. Using a city-scale vehicle-tracking scenario, we demonstrate the emulator's usability and ease of application integration. We present this as an ongoing step toward an end-to-end platform for developing and assessing drone-edge applications, motivating and easing further research in the domain.
Summit Shrestha, Rit Muliashia, Rohit Rao, Anirudh Sarma, Alan Nussbaum, Myungjin Lee, Umakishore Ramachandran
SEC7
2025 A Hybrid Runtime for Function-as-a-Service at the Edge
abstract
Contemporary Function-as-a-Service (FaaS) platforms leverage lightweight container-based virtualization to host instances of applications on-demand. This on-demand model is a good fit for computing environments where resource efficiency is desirable (such as the Cloud) or necessary (such as the Edge). However, the use of containers introduces significant delays to end-to-end service times in the form of cold start latency. As a workaround, FaaS platforms employ techniques like pre-warming containers or keeping idle containers alive to reduce the number of cold starts required. These techniques are effective, but greatly reduce platform resource efficiency, in turn making them less suitable for target computing deployments. Alternatives to containers like WebAssembly offer significantly lower cold start time, but have slower execution speeds than native code and cannot serve as a replacement to containers. We argue that combining the strengths of both containers and WebAssembly into a single hybrid runtime for FaaS platforms can improve both the end-to-end response time latencies and resource efficiency. We demonstrate this principle through proof-by-construction and introduce RUNE: a hybrid RUNtime for FaaS at the Edge. We evaluate RUNE empirically and show that its hybrid approach reduces tail latency response times at the 99th percentile by more than half while also improving memory utilization of the underlying platform by reducing the number of idle containers.
Adam Hall, Umakishore Ramachandran
Middleware2
2022 Opportunities for Optimizing the Container Runtime
abstract
Container-based virtualization provides lightweight mechanisms for process isolation and resource control that are essential for maintaining a high degree of multi-tenancy in Function-as-a-Service (FaaS) platforms, where compute functions are instantiated on-demand and exist only as long as their exe-cution is active. This model is especially advantageous for Edge computing environments, where hardware resources are limited due to physical space constraints. Despite their many advantages, state-of-the-art container runtimes still suffer from startup delays of several hundred milliseconds. This delay adversely impacts user experience for existing human-in-the-loop applications and quickly erodes the low latency response times required by emerging machine-in-the-loop IoT and Edge computing applications utilizing FaaS. In turn, it causes developers of these applications to employ unsanctioned workarounds that artificially extend the lifetime of their functions, resulting in wasted platform resources. In this paper, we provide an exploration of the cause of this startup delay and insight on how container-based virtualization might be made more efficient for FaaS scenarios at the Edge. Our results show that a small number of container startup operations account for the majority of cold start time, that several of these operations have room for improvement, and that startup time is largely bound by the underlying operating system mechanisms that are the building blocks for containers. We draw on our detailed analysis to provide guidance toward developing a container runtime for Edge computing environments and demonstrate how making a few key improvements to the container creation process can lead to a 20 % reduction in cold start time.
Adam Hall, Umakishore Ramachandran
SEC2
2022 ClairvoyantEdge: Prescient Prefetching of On-demand Video at the Edge of the Network
abstract
On-demand video contributes a large fraction of the data traffic on mobile networks. This share is expected to increase even more drastically in the coming years. While the cellular infrastructure is continuously evolving to keep pace with this increasing demand, it is necessary to ensure that sufficient bandwidth is reserved for other latency-sensitive realtime applications like video conferencing and multiplayer video games. A tangible approach involves reducing on-demand video load on cellular networks, especially from users on the move. We see an opportunity for cellular load reduction using edge nodes based on two observations: (1) video streaming is mostly a download-only operation with sequential data access; and (2) short-range mmWave links can deliver an extremely high throughput for nearby recipients of data. The knowledge of the user's planned travel route creates opportunities for prescient prefetching and delivering the content as the vehicle passes through just in time, using mmWave devices on en route edge nodes. ClairvoyantEdge is a novel networked system infrastructure that leverages inter-edge node communication and the knowledge of users' trajectories to plan and deliver buffered video segments to the vehicles passing by. To evaluate ClairvoyantEdge, we built a comprehensive end-to-end emulation-based workflow that incorporates in situ field measurements of mmWave links into our own homegrown emulation framework. With a minuscule 0.12% coverage of a 46km2geographical area employing 20 edge nodes distributed in that area providing short-range mmWave access to passing vehicles, we achieve an average reduction of up to 21% in cellular bandwidth usage for video downloads, using a real-world workload comprising 758 vehicles. Our results validate the promise of ClairvoyantEdge for incorporation in future edge infrastructure evolution.
Manasvini Sethuraman, Anirudh Sarma, Adwait Bauskar, Ashutosh Dhekne, Umakishore Ramachandran
SEC5
2022 MicroEdge: a multi-tenant edge cluster system architecture for scalable camera processing
abstract
With the proliferation of high bandwidth cameras and AR/VR devices, and their increasing use in situation awareness applications, edge computing is gaining prominence to meet the throughput requirements of such applications. This work focuses on camera applications that perform real-time Machine Learning inferences on camera frames. We find that Machine Learning based camera applications suffer from hardware resource fragmentation due to models under-utilizing or over-utilizing the accelerator. Meanwhile, it is challenging to support fine-grained resource sharing for accelerators such as TPUs because they can only process requests sequentially in a run to completion fashion. We present MicroEdge, a multi-tenant low-cost edge cluster for camera processing applications running at the edge. MicroEdge provides multi-tenancy support for Coral TPUs by extending K3s, an edge-specific distribution of Kubernetes. Through an admission control algorithm, it allows for fractional assignment of TPU resources commensurate with the application pipeline requirements to ensure that the TPUs are fully utilized. Using real-time camera processing applications and a real-world trace, we show that MicroEdge can support up to 2.8x camera streams for a given hardware configuration compared to vanilla K3s, while maintaining scalability and performance requirements.
Difei Cao, Jinsun Yoo, Zhuangdi Xu, Enrique Saurez, Tushar Krishna, Umakishore Ramachandran
Middleware7
2022 EVA: A Symbolic Approach to Accelerating Exploratory Video Analytics with Materialized Views
abstract
Advances in deep learning have led to a resurgence of interest in video analytics. In an exploratory video analytics pipeline, a data scientist often starts by searching for a global trend and then iteratively refines the query until they identify the desired local trend. These queries tend to have overlapping computation and often differ in their predicates. However, these predicates are computationally expensive to evaluate since they contain user-defined functions (UDFs) that wrap around deep learning models.
Zhuangdi Xu, Gaurav Tarlok Kakkar, Joy Arulraj, Umakishore Ramachandran
SIGMOD Conference4
2021 OneEdge: An Efficient Control Plane for Geo-Distributed Infrastructures
abstract
Resource management for geo-distributed infrastructures is challenging due to the scarcity and non-uniformity of edge resources, as well as the high client mobility and workload surges inherent to situation awareness applications. Due to their centralized nature, state-of-the-art schedulers that work well in datacenters lack the performance and feature requirements of such applications. We present OneEdge, a hybrid control plane that enables autonomous decision-making at edge sites for localized, rapid single-site application deployment. Edge sites handle mobility, churn, and load spikes, by cooperating with a centralized controller that allows coordinated multi-site scheduling and dynamic reconfiguration.
Enrique Saurez, Alexandros Daglis, Umakishore Ramachandran
SoCC4
2021 ePulsar: Control Plane for Publish-Subscribe Systems on Geo-Distributed Edge Infrastructure
Tyler C. Landle, Umakishore Ramachandran
SEC3
2021 Foresight: planning for spatial and temporal variations in bandwidth for streaming services on mobile devices
abstract
Spatiotemporal variation in cellular bandwidth availability is well-known and could affect a mobile user's quality of experience (QoE), especially while using bandwidth intensive streaming applications such as movies, podcasts, and music videos during commute. If such variations are made available to a streaming service in advance it could perhaps plan better to avoid sub-optimal performance while the user travels through regions of low bandwidth availability. The intuition is that such future knowledge could be used to buffer additional content in regions of higher bandwidth availability to tide over the deficits in regions of low bandwidth availability. Foresight is a service designed to provide this future knowledge for client apps running on a mobile device. It comprises three components: (a) a crowd-sourced bandwidth estimate reporting facility, (b) an on-cloud bandwidth service that records the spatiotemporal variations in bandwidth and serves queries for bandwidth availability from mobile users, and (c) an on-device bandwidth manager that caters to the bandwidth requirements from client apps by providing them with bandwidth allocation schedules. Foresight is implemented in the Android framework. As a proof of concept for using this service, we have modified an open-source video player---Exoplayer---to use the results of Foresight in its video buffer management. Our performance evaluation shows Foresight's scalability. We also showcase the opportunity that Foresight offers to ExoPlayer to enhance video quality of experience (QoE) despite spatiotemporal bandwidth variations for metrics such as overall higher bitrate of playback, reduction in number of bitrate switches, and reduction in the number of stalls during video playback.
Manasvini Sethuraman, Anirudh Sarma, Ashutosh Dhekne, Umakishore Ramachandran
MMSys4
2020 Coral-Pie: A Geo-Distributed Edge-compute Solution for Space-Time Vehicle Tracking
abstract
We present a distributed system architecture which is scalable by design for cross-camera vehicle tracking at video ingestion time dubbed Coral-Pie. To meet the latency bounds for timely processing of every frame at each camera, we associate dedicated low-cost computational resource for each camera, which consists of two Raspberry Pi 3B+'s and one Coral Accelerator (EdgeTpu). The end-to-end system generates and stores the tracks in a graph database for easy querying. We use the Cloud-Edge-Device continuum to appropriately place the components of the distributed system architecture. Using the timing profiles of the sub-tasks involved in the continuous processing that needs to happen on every frame in each camera, we map the elements of the processing onto the computational resource associated with each camera. Performance evaluation of the proof-of-concept system is conducted using live streams from five campus cameras. The evaluation includes microbenchmarks as well as application level studies. The controlled experiments using live cameras are augmented with a simulation-based study to show the self-healing property of the system and the system scalability.
Zhuangdi Xu, Harshil S. Shah, Umakishore Ramachandran
Middleware3
2020 A Drop-in Middleware for Serializable DB Clustering across Geo-distributed Sites
abstract
Many geo-distributed services at web-scale companies still rely on databases (DBs) primarily optimized for single-site performance. At AT&T this is exemplified by services in the network control plane that rely on third-party software that uses DBs like MariaDB and PostgreSQL, which do not provide strict serializability across sites without a significant performance impact. Moreover, it is often impractical for these services to re-purpose their code to use newer DBs optimized for geo-distribution. In this paper, a novel drop-in solution for DB clustering across sites called Metric is presented that can be used by services without changing a single line of code. Metric leverages the single-site performance of an existing service's DB and combines it with a cross-site clustering solution based on an entry-consistent redo log that is specifically tailored for geo-distribution. Detailed correctness arguments are presented and extensive evaluations with various benchmarks show that Metric outperforms other solutions for the access patterns in our production use-cases where service replicas access different tables on different sites. In particular, Metric achieves up to 56% less latency and 5.2x higher throughput than MariaDB and PostgreSQL clustering, and up to 90% less latency and 26x higher throughput than CockroachDB and TiDB, systems that are designed to support geo-distribution.
Enrique Saurez, Bharath Balasubramanian, Richard D. Schlichting, Brendan Tschaen, Shankaranarayanan Puzhavakath Narayanan, Zhe Huang 0001, Umakishore Ramachandran
Proc. VLDB Endow.7
2019 Elevating the Edge to Be a Peer of the Cloud
abstract
Enabling next generation technologies such as self-driving cars or smart cities requires us to rethink the way we support their applications. The emergence of these technologies is fueled by the proliferation of a large number of devices in the Internet of Things. These devices have the potential to generate massive amounts of data, and applications supporting them often require this data to be processed in a timely manner. Because of these requirements, we must augment and extend the Cloud computing model to better serve such applications. The backhaul links connecting clients to Cloud data centers could quickly become overwhelmed by such data, and the physical distance of these data centers from clients prevents low-latency response times. To meet the challenges posed by emerging IoT applications, we must provide Cloudlike functionality closer to the edge of the network, where clients and their data live. We propose to elevate the Edge to be a peer of the Cloud for addressing these challenges.
Umakishore Ramachandran, Adam Hall, Enrique Saurez, Zhuangdi Xu
CLOUD1
2019 Extension Framework for File Systems in User space
Ashish Bijlani, Umakishore Ramachandran
USENIX ATC2
2017 SPECTRE: supporting consumption policies in window-based parallel complex event processing
abstract
Distributed Complex Event Processing (DCEP) is a paradigm to infer the occurrence of complex situations in the surrounding world from basic events like sensor readings. In doing so, DCEP operators detect event patterns on their incoming event streams. To yield high operator throughput, data parallelization frameworks divide the incoming event streams of an operator into overlapping windows that are processed in parallel by a number of operator instances. In doing so, the basic assumption is that the different windows can be processed independently from each other. However, consumption policies enforce that events can only be part of one pattern instance; then, they are consumed, i.e., removed from further pattern detection. That implies that the constituent events of a pattern instance detected in one window are excluded from all other windows as well, which breaks the data parallelism between different windows. In this paper, we tackle this problem by means of speculation: Based on the likelihood of an event's consumption in a window, subsequent windows may speculatively suppress that event. We propose the SPECTRE framework for speculative processing of multiple dependent windows in parallel. Our evaluations show an up to linear scalability of SPECTRE with the number of CPU cores.
Ruben Mayer, Ahmad Slo, Muhammad Adnan Tariq, Kurt Rothermel, Manuel Gräber, Umakishore Ramachandran
Middleware6
2014 MCEP: A Mobility-Aware Complex Event Processing System
abstract
With the proliferation of mobile devices and sensors, complex event proceesing (CEP) is becoming increasingly important to scalably detect situations in real time. Current CEP systems are not capable of dealing efficiently with highly dynamic mobile consumers whose interests change with their location. We introduce the distributed mobile CEP (MCEP) system which automatically adapts the processing of events according to a consumer's location. MCEP significantly reduces latency, network utilization, and processing overhead by providing on-demand and opportunistic adaptation algorithms to dynamically assign event streams and computing resources to operators of the MCEP system.
Beate Ottenwälder, Boris Koldehofe, Kurt Rothermel, Kirak Hong, David J. Lillethun, Umakishore Ramachandran
ACM Trans. Internet Techn.6
2013 FlashStream: a multi-tiered storage architecture for adaptive HTTP streaming
abstract
Video streaming on the Internet is popular and the need to store and stream video content using CDNs is continually on the rise thanks to services such as Hulu and Netflix. Adaptive HTTP streaming using the deployed CDN infrastructure has become the de facto standard for meeting the increasing demand for video streaming on the Internet. The storage architecture that is used for storing and streaming the video content is the focus of this study. Hard-disk as the storage medium has been the norm for enterprise-class storage servers for the longest time. More recently, multi-tiered storage servers (incorporating SSDs) such as Sun's ZFS and Facebook's flashcache offer an alternative to disk-based storage servers for enterprise applications. Both these systems use the SSD as a cache between the DRAM and the hard disk. The thesis of our work is that the current-state-of-the art in multi-tiered storage systems, architected for general-purpose enterprise workloads, do not cater to the unique needs of adaptive HTTP streaming. We present FlashStream, a multi-tiered storage architecture that addresses the unique needs of adaptive HTTP streaming. Like ZFS and flashcache, it also incorporates SSDs as a cache between the DRAM and the hard disk. The key architectural elements of FlashStream include optimal write granularity to overcome the write amplification effect of flash memory SSDs and a QoS-sensitive caching strategy that monitors the activity of the flash memory SSDs to ensure that video streaming performance is not hampered by the caching activity. We have implemented FlashStream and experimentally compared it with ZFS and flashcache for adaptive HTTP streaming workloads. We show that FlashStream outperforms both these systems for the same hardware configuration. Specifically, it is better by a factor of two compared to its nearest competitor, namely ZFS. In addition, we have compared FlashStream with a traditional two-level storage architecture (DRAM + HDDs), and have shown that, for the same investment cost, FlashStream provides 33% better performance and 94% better energy efficiency.
Moonkyung Ryu, Umakishore Ramachandran
ACM Multimedia2
2013 Fjord: Informed storage management for smartphones
abstract
Smartphone applications are becoming more sophisticated and require high storage performance. Unfortunately, the OS storage software stack is not well engineered to support flash-based storage used in smartphones. On top of that, storage software stack is configured to be too conservative due to the fear of sudden power failures. We believe that this conservatism with respect to data reliability is misplaced considering that many of the popular apps (e.g., Web browsing, Facebook, Gmail) that run on today's smartphones are cloud-backed, and the local storage on the smartphone is often used as a cache for cloud data. In this paper, we propose Informed Storage Management framework, named Fjord, for mobile platforms. The key insight is to use system-wide dynamic context information to improve the storage performance on mobile platforms. We implement a set of mechanisms (write buffering, logging, and fine-grained reliability control), and through judicious use of these mechanisms based on system context, we show how we can achieve significant improvement in storage performance. As proof of concept, we implement Fjord in two Android smartphones and experimentally validate the performance advantage of informed storage management with multiple smartphone applications.
Hyojun Kim, Umakishore Ramachandran
MSST2
2012 Community Membership Management for Transient Social Networks
abstract
Confluence of technologies represented by geo-location, geo-sensing, context and activity recognition, and smart phones with rich sensors is opening up new avenues for media-rich social interactions for a spectrum of applications from entertainment, to commerce, to emergency response. This paper addresses the challenges in membership management of a transient social network (TSN), a community of users with mobile devices engaging in social activities of common interest within specific temporal and geographical locality (e.g., flea market, emergency response, local auctions, etc.). Creating and maintaining viable transient social communities requires solving a number of significant challenges including managing the dynamically created social graph, maintaining connectivity across heterogeneous nodes and interfaces and efficient message delivery among the nodes in a community. Micrograph is a middleware for managing community membership in TSNs. It helps nodes to discover and participate with other nodes based on device-level, or application-level attributes. It allows a node to simultaneously participate in multiple distinct social communities by overlaying multiple TSNs on top of the available nodes in the physical network. Micrograph gives complete isolation for the activities of a node in each of the TSNs that a node may be participating in simultaneously and it also gives complete transparency to a participant as to the membership of a TSN in which he/she is involved in. In this paper, we present the design and implementation of Micrograph, a proof-of-concept implementation of the middleware using Android platforms, four applications to show the feasibility of Micrograph, and simulation-based evaluation of the implemented prototype.
Lateef Yusuf, Umakishore Ramachandran
ICCCN2
2012 Why are state-of-the-art flash-based multi-tiered storage systems performing poorly for HTTP video streaming?
abstract
MLC flash memory is a promising technology for building a high-performance and cost-effective video streaming system when it is used as an intermediate level cache in a multi-tiered storage hierarchy. Therefore, we were quite surprised when through extensive measurements we found that two state-of-the-art flash-based multi-tiered storage systems (namely, flashcache and ZFS) have quite disappointing performance for HTTP video streaming using the DASH protocol. We have conducted a thorough analysis to understand the reasons for the poor performance of these two systems. In a nutshell, unless attention is paid to the unique performance characteristics of flash memory-based SSDs, we could end up with suboptimal or even poor performance as we discovered through experimentation with these two systems. Based on the analysis, we present design guidelines for building a cost-effective high-performance HTTP video streaming server.
Moonkyung Ryu, Hyojun Kim, Umakishore Ramachandran
NOSSDAV3
2012 What is a good buffer cache replacement scheme for mobile flash storage?
abstract
Smartphones are becoming ubiquitous and powerful. The Achilles' heel in such devices that limits performance is the storage. Low-end flash memory is the storage technology of choice in such devices due to energy, size, and cost considerations. In this paper, we take a critical look at the performance of flash on smartphones for mobile applications. Specifically, we ask the question whether the state-of-the-art buffer cache replacement schemes proposed thus far (both flash-agnostic and flash-aware ones) are the right ones for mobile flash storage. To answer this question, we first expose the limitations of current buffer cache performance evaluation methods, and propose a novel evaluation framework that is a hybrid between trace-driven simulation and real implementation of such schemes inside an operating system. Such an evaluation reveals some unexpected and surprising insights on the performance of buffer management schemes that contradicts conventional wisdom. Armed with this knowledge, we propose a new buffer cache replacement scheme called SpatialClock.
Hyojun Kim, Moonkyung Ryu, Umakishore Ramachandran
SIGMETRICS3
2012 Large-Scale Situation Awareness With Camera Networks and Multimodal Sensing
abstract
Sensors of various modalities and capabilities, especially cameras, have become ubiquitous in our environment. Their intended use is wide ranging and encompasses surveillance, transportation, entertainment, education, healthcare, emergency response, disaster recovery, and the like. Technological advances and the low cost of such sensors enable deployment of large-scale camera networks in large metropolises such as London and New York. Multimedia algorithms for analyzing and drawing inferences from video and audio have also matured tremendously in recent times. Despite all these advances, large-scale reliable systems for media-rich sensor-based applications, often classified as situation-awareness applications, are yet to become commonplace. Why is that? There are several forces at work here. First, the system abstractions are just not at the right level for quickly prototyping such applications on a large scale. Second, while Moore's law has held true for predicting the growth of processing power, the volume of data that applications are called upon to handle is growing similarly, if not faster. Enormous amount of sensing data is continually generated for real-time analysis in such applications. Further, due to the very nature of the application domain, there are dynamic and demanding resource requirements for such analyses. The lack of right set of abstractions for programing such applications coupled with their data-intensive nature have hitherto made realizing reliable large-scale situation-awareness applications difficult. Incidentally, situation awareness is a very popular but ill-defined research area that has attracted researchers from many different fields. In this paper, we adopt a strong systems perspective and consider the components that are essential in realizing a fully functional situation-awareness system.
Umakishore Ramachandran, Kirak Hong, Liviu Iftode, Ramesh Jain 0001, Kurt Rothermel, JunSuk Shin, Raghupathy Sivakumar
Proc. IEEE1
2011 Impact of flash memory on video-on-demand storage: analysis of tradeoffs
abstract
There is no doubt that video-on-demand (VoD) services are very popular these days. However, disk storage is a serious bottleneck limiting the scalability of a VoD server. Disk throughput degrades dramatically due to seek time overhead when the server is called upon to serve a large number of simultaneous video streams. To address the performance problem of disk, buffer cache algorithms that utilize RAM have been proposed. Interval caching is a state-of-the-art caching algorithm for a VoD server. Flash Memory Solid-State Drive (SSD) is a relatively new storage technology. Its excellent random read performance, low power consumption, and sharply dropping cost per gigabyte are opening new opportunities to efficiently use the device for enterprise systems. On the other hand, it has deficiencies such as poor small random write performance and limited number of erase operations. In this paper, we analyze tradeoffs and potential impact that flash memory SSD can have for a VoD server. Performance of various commercially available flash memory SSD models is studied. We find that low-end flash memory SSD provides better performance than the high-end one while costing less than the high-end one when the I/O request size is large, which is typical for a VoD server. Because of the wear problem and asymmetric read/write performance of flash memory SSD, we claim that interval caching cannot be used with it. Instead, we propose using file-level Least Frequently Used (LFU) due to the highly skewed video access pattern of the VoD workload. We compare the performance of interval caching with RAM and file-level LFU with flash memory by simulation experiments. In addition, from the cost-effectiveness analysis of three different storage configurations, we find that flash memory with hard disk drive is the most cost-effective solution compared to DRAM with hard disk drive or hard disk drive only.
Moonkyung Ryu, Hyojun Kim, Umakishore Ramachandran
MMSys3
2010 DynaStream: Adaptive Overlay Management for Peer-to-Peer Video Streaming
abstract
A number of mesh-based peer-to-peer video streaming systems have been proposed, however, they have not paid careful attention to the impact of the number of data streams on streaming quality. In this paper, we first explore the effect of the choice of the number of data streams to the streaming quality. We propose a simple, practical, and fully distributed mechanism, Loss Rate Window, that maintains the overlay topology adaptive to peers' uplink bandwidth and network fluctuation. The mechanism allows each node to make an entirely local decision to adaptively control the right number of data streams to support on its uplink without explicitly knowing or measuring its uplink bandwidth. This mechanism is general, and can be applied to any kind of mesh-based peer-to-peer video streaming system to improve the quality of the video streaming service.
Moonkyung Ryu, Umakishore Ramachandran
ICCCN2
2009 FlashLite: A User-Level Library to Enhance Durability of SSD for P2P File Sharing
abstract
Peer-to-peer file sharing is popular, but it generates random write traffic to storage due to the nature of swarming. NAND flash memory based Solid-State Drive (SSD) technology is available as an alternative to hard drives for notebook and tablet PCs. As it turns out, random write is extremely detrimental to the lifetime of SSD drives. This paper focuses on the following problem, namely, P2P file downloading when the target of the download is an SSD drive. We make three contributions: first, analysis of write patterns of downloading program to establish the premise of the problem; second, development of a simple yet powerful technique called FlashLite to combat this problem, by automatically converting the random writes to sequential writes; third, showing through performance evaluation using modified eMule file downloading program that FlashLite does change random writes to sequential, and most importantly eliminates about 94% of erase operations of the original eMule program.
Hyojun Kim, Umakishore Ramachandran
ICDCS2
2009 Persistent Temporal Streams
David Hilley, Umakishore Ramachandran
Middleware2
2007 Use of Dependency Information for Memory Optimizations in Distributed Streaming Applications
abstract
In this paper we explore the potential of using application data dependency information to reduce the average memory consumption in distributed streaming applications. By analyzing data dependencies during the application runtime, we can infer which data items are not going to influence the application's output. This information is then incorporated into the garbage collector, extending the garbage identification problem to include not only data items that are not reachable, but also those data items that are not fully processed and dropped. We present three garbage collection algorithms. Each of the algorithms uses different data dependency information. We implement the algorithms and compare their performance for a color tracker application. Our results show that these algorithms not only succeed in substantially reducing the average memory usage but also improve the overall performance of the application. The results also indicate that the garbage identification algorithms that achieve a low memory footprint perform their garbage identification decisions locally; however, they base these decisions on best-effort global information. The results also indicate that the garbage identification algorithms perform best when they base their decisions on best-effort global information obtained from other components of the distributed application.
Nissim Harel, Hasnain A. Mandviwala, Umakishore Ramachandran, Kathleen Knobe
ICCCN3
2007 On Improving the Reliability of Packet Delivery in Dense Wireless Sensor Networks
abstract
Wireless sensor networks (WSN) built using current Berkeley Mica motes exhibit low reliability for packet delivery. There is anecdotal evidence of poor packet delivery rates from several field trials of WSN deployment. All-to-one communication pattern is a dominant one in many such deployments. As we scale up the size of the network and the traffic density in this communication pattern, improving the reliability of packet delivery performance becomes very important. This study is aimed at two things. Firstly, it aims to understand the factors limiting reliable packet delivery for all-to-one communication pattern in dense wireless sensor networks. Secondly, it aims to suggest enhancements to well-known protocols that may help boost the performance to acceptable levels. We first postulate the potential reasons hampering packet delivery rates with current CSMA-based MAC layer used by the radios deployed in WSN. We then propose a set of enhancements that are aimed to mitigate the ill-effects of these factors. We pick three protocols, namely, Flooding, AODV, and Geographic routing as candidates for this study. Using TOSSIM, we perform a detailed study of these protocols and the proposed enhancements. This study serves several purposes. First, it helps us to quantify the detrimental effects of these factors. Second, it helps us to quantify the extent to which our proposed enhancements improves packet delivery performance. Concretely, we show that using Geographic routing in a WSN with 225 nodes spread over 150 feet times 150 feet, the proposed enhancements yield a 23-fold improvement in packet delivery performance over the baseline. Further, the enhancements result in fairness (measured by the number of messages received from each node at the destination). Lastly, we show that the overhead (in terms of retransmissions, acknowledgement messages, and control messages) is reasonable.
JunSuk Shin, Umakishore Ramachandran, Mostafa H. Ammar
ICCCN2
2007 Stampede RT: Programming Abstractions for Live Streaming Applications
abstract
We present StampedeRT, middleware designed to provide a natural programming model appropriate for live streaming applications. Such applications require pervasive access to multiple streaming data sources for distributed online analysis. One motivating example is a distributed robotics application which analyzes live camera feeds for control and planning. Most existing middlewares for streaming data focus on media streams and low-level transport characteristics such as delivery latency and efficient transfer, but do not define a programming model to succinctly express applications that manipulate and analyze the streaming content. StampedeRTprovides for straightforward transport and manipulation of temporally-ordered data streams, enabling simple synchronization and correlation of data sources. We present an abstract programming model to support the aforementioned class of applications and then describe a concrete realization of the model as a distributed middleware architecture. We also evaluate our implementation of the architecture and present several motivating applications StampedeRTis designed to support.
David Hilley, Umakishore Ramachandran
ICDCS2
2007 Methods of Memory Optimizations in Streaming Applications
abstract
Streaming applications are often distributed, manage large quantities of data and, as a result, have large memory requirements. Therefore, efficient garbage collection (GC) is crucial for their performance. On the other hand, not all data items affect the application output due to differences in the processing rates of various application threads. In this paper we propose extending the definition of the garbage identification problem for streaming applications and include not only data items that are not "reachable " but also data items that have no effect on the final outcome of the application. We present four optimizations to an existing GC algorithm in Stampede, a parallel programming system to support interactive multimedia applications. We ask the question how far off these algorithms are from an ideal garbage collector, one in which the memory usage exactly equals the amount required for buffering only the relevant data items. This oracle, while unimplementable, serves as an empirical lower-bound for memory usage. We then propose optimizations that will help us get closer to this lower- bound. Using an elaborate measurement and post-mortem analysis infrastructure, we simulate the performance potential for these optimizations and implement the most promising ones. A color-based people tracking application is used for the performance evaluation. Our results show that these optimizations reduce the memory usage by up to 60%.
Nissim Harel, Hasnain A. Mandviwala, Kathleen Knobe, Umakishore Ramachandran
ICPP4
2007 RF2ID: A Reliable Middleware Framework for RFID Deployment
abstract
The reliability of RFID systems depends on a number of factors including: RF interference, deployment environment, configuration of the readers, and placement of readers and tags. While RFID technology is improving rapidly, a reliable deployment of this technology is still a significant challenge impeding widespread adoption. This paper investigates system software solutions for achieving a highly reliable deployment that mitigates inherent unreliability in RFID technology. We have developed (1) a virtual reader abstraction to improve the potentially error-prone nature of reader generated data (2) a novel path abstraction to capture the logical flow of information among virtual readers. We have designed and implemented an RFID middleware: RF2ID (reliable framework for radio frequency identification) to organize and support queries over data streams in an efficient manner. Prototype implementation using both RFID readers and simulated readers using an empirical model of RFID readers show that RF2ID is able to provide high reliability and support path-based object detection.
Nova Ahmed, Robert Steven French, Umakishore Ramachandran
IPDPS4
2007 ASAP: A Camera Sensor Network for Situation Awareness
JunSuk Shin, Dushmanta Mohapatra, Umakishore Ramachandran, Mostafa H. Ammar
OPODIS4
2007 MB++: An Integrated Architecture for Pervasive Computing and High-Performance Computing
abstract
MB++ is a system that caters to the dynamic needs of applications in a distributed, pervasive computing environment that has a wide variety of devices that act as producers and consumers of stream data. The architecture encompasses several elements: The type server allows clients to dynamically inject transformation code that operates on data streams. The transformation engine executes dataflow graphs of transformations on high-performance computing resources. The stream server manages all data streams in the system and dispatches new dataflow graphs to the transformation engine. We have implemented the architecture and show performance results that demonstrate that our implementation scales well with increasing workload, commensurate with the available HPC resources. Further, we show that our implementation can exploit opportunities for parallelism in dataflow graphs, as well as efficiently sharing common subgraphs between dataflow graphs.
David J. Lillethun, David Hilley, Seth Horrigan, Umakishore Ramachandran
RTCSA4
2007 MobiGo: A Middleware for Seamless Mobility
abstract
Nominally, one can expect any user of modern technology to carry a handheld device such an iPAQ or cellphone and utilize resources in the environment to remain connected and enjoy continuous services while travelling. We present a middleware infrastructure, called MobiGo, that provides seamless mobility of services among these environments. We identify three different kinds of environments (spaces) - self-owned, familiar, and totally-new and three axes for supporting mobility, namely, hard state, soft state, and I/O state in these spaces. MobiGo provides the architectural elements for efficiently managing these different states in the different spaces. Focusing on a specific demanding video service, we describe an implementation and performance results that show that MobiGo enhances user experience for seamless mobility.
Umakishore Ramachandran
RTCSA2
2007 Streamline: scheduling streaming applications in a wide area environment
Bikash Agarwalla, Nova Ahmed, David Hilley, Umakishore Ramachandran
Multim. Syst.4
2006 A Bridging Framework for Universal Interoperability in Pervasive Systems
abstract
We explore the design patterns and architectural tradeoffs for achieving interoperability across communication middleware platforms, and describe uMiddle, a bridging framework for universal interoperability that enables seamless device interaction over diverse platforms. The proliferation of middleware platforms that cater to specific devices has created isolated islands of devices with no uniform protocol for interoperability across these islands. This void makes it difficult to rapidly prototype pervasive computing applications spanning a wide variety of devices. We discuss the design space of architectural solutions that can address this void, and detail the trade-offs that must be faced when trying to achieve cross-platform interoperability. uMiddle is a framework for achieving such interoperability, and serves as a powerful platform for creating applications that are independent of specific underlying communication platforms.
Jin Nakazawa, Hideyuki Tokuda, W. Keith Edwards, Umakishore Ramachandran
ICDCS4
2006 On Improving Wireless Broadcast Reliability of Sensor Networks Using Erasure Codes
Arnab Paul, Umakishore Ramachandran, David Kotz
MSN3
2006 System Support for Cross-Layering in Sensor Network Stack
Santashil PalChaudhuri, Charles Reiss, Umakishore Ramachandran
MSN4
2006 UbiqStack: a taxonomy for a ubiquitous computing software stack
Martin Modahl, Bikash Agarwalla, T. Scott Saponas, Gregory D. Abowd, Umakishore Ramachandran
Pers. Ubiquitous Comput.5
2006 Dynamic data fusion for future sensor networks
abstract
DFuse is an architectural framework for dynamic application-specified data fusion in sensor networks. It bridges an important abstraction gap for developing advanced fusion applications that takes into account the dynamic nature of applications and sensor networks. Elements of the DFuse architecture include a fusion API, a distributed role assignment algorithm that dynamically adapts the placement of the application task graph on the network, and an abstraction migration facility that aids such dynamic role assignment. Experimental evaluations show that the API has low overhead, and simulation results show that the role assignment algorithm significantly increases the network lifetime over static placement.
Umakishore Ramachandran, Matthew Wolenetz, Brian Cooper, Bikash Agarwalla, JunSuk Shin, Phillip W. Hutto, Arnab Paul
ACM Trans. Sens. Networks1
2006 Distributed Garbage Collection Algorithms for Timestamped Data
abstract
There is an important class of interactive multimedia applications that deals with stream data from distributed sources. Indexing the data temporally facilitates ordering individual streams as well as correlating items from different streams. The Stampede programming system organizes stream data into channels that are distributed and synchronized data structures that contain timestamped items. A stampede program is a data flow graph of threads and channels. Stampede semantics for channels allow concurrent access from multiple threads for input and output. While a channel holds timestamped items, the semantics do not place any restriction on either the production or consumption order of these items. Furthermore, timestamps of items in a channel need not be contiguous. These flexibilities are required due to the dynamic and parallel structure of stream-oriented applications targeted by the stampede system. Under such circumstances, a key issue is the "garbage collection" (GC) of channel items. In this paper, we present and compare three different GC algorithms: 1) REF is a simple algorithm that keeps a reference count on individual items; 2) TGC is a distributed algorithm for computing a global low watermark for timestamp values of interest in the entire application; 3) DGC is another distributed algorithm that uses information about the dependencies between the producers and consumers of data streams to compute a low water mark local to each node of the data flow graph. DGC can simultaneously eliminate garbage from channels and unneeded computations from threads, in tests performed using an interactive application, DGC enjoys nearly 30 percent reduction in the application memory footprint, compared, to TGC and REF. DGC and REF are also shown to be more scalable compared to TGC
Umakishore Ramachandran, Kathleen Knobe, Nissim Harel, Hasnain A. Mandviwala
IEEE Trans. Parallel Distributed Syst.1
2005 LAWN: A Protocol for Remote Authentication overWireless Networks
abstract
Remote authentication over a long range wireless network using large signature keys such as biometric samples (fingerprint, retinal scans etc.) is soon going to become an integral feature of various kinds of transactions. In the domain of mobile and ad hoc networking, this become even more relevant due to the intrinsic dynamism in the applications. Because of the large size of the authentication keys, and continual need for authentication, considerable power and bandwidth are consumed by such a process. Authentication being only a background process supporting other transactions, should not take away too much of resources, especially bandwidth and power that are quite critical for small mobile devices. We present LAWN, a light-weight authentication protocol for wireless networks that trades computation for communication and can be tuned for any desired security guarantee. For an authentication token of length n, LAWN prepares a small sketch of length O(log n) (adding very low computational overhead), and transmits the sketch over the network. Under a reasonable energy consumption model, we show that this technique results in 70% to 80% saving in power for long-range wireless applications
Arnab Paul, Umakishore Ramachandran
NCA3
2005 Experiences with optimizing two stream-based applications for cluster execution
Yavor Angelov, Umakishore Ramachandran, Kenneth M. Mackenzie, James M. Rehg, Irfan A. Essa
J. Parallel Distributed Comput.2
2005 MediaBroker: A pervasive computing infrastructure for adaptive transformation and sharing of stream data
Umakishore Ramachandran, Martin Modahl, Ilya Bagrak, Matthew Wolenetz, David J. Lillethun, Bin Liu 0010, James Kim, Phillip W. Hutto, Ramesh Jain 0001
Pervasive Mob. Comput.1
2004 MediaBroker: An Architecture for Pervasive Computing
abstract
MediaBroker is a distributed framework designed to support pervasive computing applications. Specifically, the architecture consists of a transport engine and peripheral clients and addresses issues in scalability, data sharing, data transformation and platform heterogeneity. Key features of MediaBroker are a type-aware data transport that is capable of dynamically transforming data en route from source to sinks; an extensible system for describing types of streaming data; and the interaction between the transformation engine and the type system. Details of the MediaBroker architecture and implementation are presented in this paper. Through experimental study, we show reasonable performance for selected streaming media-intensive applications. For example, relative to baseline TCP performance, MediaBroker incurs under 11% latency overhead and achieves roughly 80% of the TCP throughput when streaming items larger than 100 KB across our infrastructure.
Martin Modahl, Ilya Bagrak, Matthew Wolenetz, Phillip W. Hutto, Umakishore Ramachandran
PerCom5
2003 Performance study of a cluster runtime system for dynamic interactive stream-oriented applications
abstract
Emerging application domains such as interactive vision, animation, and multimedia collaboration need specialized runtime systems that provide support mechanisms to enable plumbing, cross module data transfer, buffer management, synchronization and so on. Using Stampede, a cluster programming system that is designed to meet the requirements of such applications, we quantify the performance of such mechanisms. We have developed a timing infrastructure that helps tease out the time spent by an application in different layers of software, viz., the main algorithmic component, the support mechanisms, and the raw messaging. Several interesting insights have surfaced from this study. First, memory allocation does not take up a significant amount of the execution time despite the interactive and dynamic nature of the application domain. Second, the Stampede runtime adds a minimal overhead over raw messaging for structuring such applications. Third, the results suggest that the thread scheduler on Linux may be more responsive than the one on Solaris. Fourth, the messaging layer spends quite a bit of time in synchronization operations. Perhaps the most interesting result of this study is that general-purpose operating systems such as Linux and Solaris are quite adequate to meet the requirements of emerging dynamic interactive stream-oriented applications.
Arnab Paul, Nissim Harel, Sameer Adhikari, Bikash Agarwalla, Umakishore Ramachandran, Kenneth M. Mackenzie
ISPASS5
2003 DFuse: a framework for distributed data fusion
abstract
Simple in-network data aggregation (or fusion) techniques for sensor networks have been the focus of several recent research efforts, but they are insufficient to support advanced fusion applications. We extend these techniques to future sensor networks and ask two related questions: (a) what is the appropriate set of data fusion techniques, and (b) how do we dynamically assign aggregation roles to the nodes of a sensor network. We have developed an architectural framework, DFuse, for answering these two questions. It consists of a data fusion API and a distributed algorithm for energy-aware role assignment. The fusion API enables an application to be specified as a coarse-grained dataflow graph, and eases application development and deployment. The role assignment algorithm maps the graph onto the network, and optimally adapts the mapping at run-time using role migration. Experiments on an iPAQ farm show that, the fusion API has low-overhead, and the role assignment algorithm with role migration significantly increases the network lifetime compared to any static assignment.
Matthew Wolenetz, Bikash Agarwalla, JunSuk Shin, Phillip W. Hutto, Arnab Paul, Umakishore Ramachandran
SenSys7
2003 Stampede: A Cluster Programming Middleware for Interactive Stream-Oriented Applications
abstract
Emerging application domains such as interactive vision, animation, and multimedia collaboration display dynamic scalable parallelism and high-computational requirements, making them good candidates for executing on parallel architectures such as SMPs and clusters of SMPs. Stampede is a programming system that has many of the needed functionalities such as high-level data sharing, dynamic cluster-wide threads and their synchronization, support for task and data parallelism, handling of time-sequenced data items, and automatic buffer management. We present an overview of Stampede, the primary data abstractions, the algorithmic basis of garbage collection, and the issues in implementing these abstractions on a cluster of SMPs. We also present a set of micromeasurements along with two multimedia applications implemented on top of Stampede, through which we demonstrate the low overhead of this runtime and that it is suitable for the streaming multimedia applications.
Umakishore Ramachandran, Rishiyur S. Nikhil, James M. Rehg, Yavor Angelov, Arnab Paul, Sameer Adhikari, Kenneth M. Mackenzie, Nissim Harel, Kathleen Knobe
IEEE Trans. Parallel Distributed Syst.1
2002 D-Stampede: Distributed Programming System for Ubiquitous Computing
abstract
We focus on an important problem in ubiquitous computing, namely, programming support for the distributed heterogeneous computing elements that make up this environment. We address the interactive, dynamic, and stream-oriented nature of this application class and develop appropriate computational abstractions in the D-Stampede distributed programming system. The key features of D-Stampede include indexing data streams temporally, correlating different data streams temporally, performing automatic distributed garbage collection of unnecessary stream data, supporting high performance by exploiting hardware parallelism where available, supporting platform and language heterogeneity, and dealing with application level dynamism. We discuss the features of D-Stampede, the programming ease it affords, and its performance.
Sameer Adhikari, Arnab Paul, Umakishore Ramachandran
ICDCS3
2002 Dead Timestamp Identification in Stampede
abstract
Stampede is a parallel programming system to support computationally demanding applications including interactive vision, speech and multimedia collaboration. The system alleviates concerns such as communication, synchronization, and buffer management in programming such real-time stream-oriented applications. Threads are loosely connected by channels that hold timestamped data items. There are two performance concerns when programming with Stampede. The first is space, namely, ensuring that memory is not wasted on items that are not fully processed. The second is time, namely, ensuring that processing resource is not wasted on a timestamp that is not fully processed. In this paper we introduce a single unifying framework, dead timestamp identification, that addresses both the space and time concerns simultaneously. Dead timestamps on a channel represent garbage. Dead timestamps at a thread represent computations that need not be performed. This framework has been implemented in the Stampede system. Experimental results showing the space advantage of this framework are presented. Using a color-based people tracker application, we show that the space advantage can be significant (up to 40%) compared to the previous garbage collection techniques in Stampede.
Nissim Harel, Hasnain A. Mandviwala, Kathleen Knobe, Umakishore Ramachandran
ICPP4
2000 Garbage collection of timestamped data in Stampede
abstract
Stampede is a parallel programming system to facilitate the programming of interactive multimedia applications on clusters of SMPs.
Rishiyur S. Nikhil, Umakishore Ramachandran
PODC2
1999 Space-Time Memory: A Parallel Programming Abstraction for Interactive Multimedia Applications
abstract
Realistic interactive multimedia involving vision, animation, and multimedia collaboration is likely to become an important aspect of future computer applications. The scalable parallelism inherent in such applications coupled with their computational demands make them ideal candidates for SMPs and clusters of SMPs. These applications have novel requirements that offer new kinds of challenges for parallel system design.We have designed a programming system called Stampede that offers many functionalities needed to simplify development of such applications (such as high-level data sharing abstractions, dynamic cluster-wide threads, and multiple address spaces). We have built Stampede and it runs on clusters of SMPs. To date we have implemented two applications on Stampede, one of which is discussed herein.In this paper we describe a part of Stampede called Space-Time Memory (STM). It is a novel data sharing abstraction that enables interactive multimedia applications to manage a collection of time-sequenced data items simply, efficiently, and transparently across a cluster. STM relieves the application programmer from low level synchronization and data communication by providing a high level interface that subsumes buffer management, inter-thread synchronization, and location transparency for data produced and accessed anywhere in the cluster. STM also automatically handles garbage collection of data items that will no longer be accessed by any of the application threads. We discuss ease of use issues for developing applications using STM, and present preliminary performance results to show that STM's overhead is low.
Umakishore Ramachandran, Rishiyur S. Nikhil, Nissim Harel, James M. Rehg, Kathleen Knobe
PPoPP1
1999 Scheduling Constrained Dynamic Applications on Clusters
abstract
There is an emerging class of computationally demanding multimedia applications involving vision, speech and interaction with the real world (e.g., CRL's Smart Kiosk). These applications are highly parallel and require low latencies for good performance. They are well-suited for implementation on clusters of SMP's, but they require efficient scheduling of application tasks. General purpose schedulers produce high latencies because they lack knowledge of the dependencies between tasks. Previous research in optimal scheduling has been limited to static problems. In contrast, our application is highly dynamic as the optimal schedule depends upon the behavior of the kiosk's customers. We observe that the dynamism of our application class is constrained, in that there are a small number of operating regimes which are determined by the state of the application. We present a framework for optimal scheduling of constrained dynamic applications. The results of an experimental compariso...
Kathleen Knobe, James M. Rehg, Arun Chauhan 0001, Rishiyur S. Nikhil, Umakishore Ramachandran
SC5
1999 An Application-Driven Study of Parallel System Overheads and Network Bandwidth Requirements
abstract
Evaluating and analyzing the performance of a parallel application on an architecture to explain the disparity between projected and delivered performance is an important aspect of parallel systems research. However, conducting such a study is hard due to the vast design space of these systems. We study two important aspects related to the performance of parallel applications on shared memory parallel architectures. First, we quantify overheads observed during the execution of these applications on three different simulated architectures. We next use these results to synthesize the bandwidth requirements for the applications with respect to different network topologies. This study is performed using an execution-driven simulation tool called SPASM, which provides a way of isolating and quantifying the different parallel system overheads in a nonintrusive manner. The first exercise shows that in shared memory machines with private caches, as long as the applications are well-structured to exploit locality, the key determinant that impacts performance is network connection. The second exercise quantifies the network bandwidth needed to minimize the effect of network connection. Specifically, it is shown that for the applications considered, as long as the problem sizes are increased commensurate with the system size, current network technologies supporting 200-300 MBytes/sec link bandwidth are sufficient to keep the network overheads (such as latency and contention) within acceptable bounds.
Anand Sivasubramaniam, Aman Singla, Umakishore Ramachandran, H. Venkateswaran
IEEE Trans. Parallel Distributed Syst.3
1997 Temporal Notions of Synchronization and Consistency in Beehive
abstract
An important attribute in the specification of many compute-intensive applications is "time".Simulation of interactive virtual environments is one such domain.There is a mismatch between the synchronization and consistency guarantees needed by such applications (which are temporal in nature) and the guarantees offered by current shared memory systems.Consequently, programming such applications using standard shared memory style synchronization and communication is cumbersome.Furthermore, such applications offer opportunities for relaxing both the synchronization and consistency requirements along the temporal dimension.In this work, we develop a temporal programming model that is more intuitive for the development of applications that need temporal correctness guarantees.This model embodies two mechanisms: "delta consistency" -a novel time-based correctness criterion to govern the shared memory access guarantees, and a companion "temporal synchronization" -a mechanism for thread synchronization along the time axis.These mechanisms are particularly appropriate for expressing the requirements in interactive application domains. in addition to the temporal programming model, we develop efficient explicit communication mechanisms that aggressively push the data out to "future" consumers to hide the read miss latency at the receiving end.We implement these mechanisms on a cluster of workstations in a software distributed shared memory architecture called "Beehive?Using a virtual environment application as the driver, we show the efficacy of the proposed mechanisms in meeting the real time requirements of such applications.
Aman Singla, Umakishore Ramachandran, Jessica K. Hodgins
SPAA2
1997 Toward a More Realistic Performance Evaluation of Interconnection Networks
abstract
lnterconnection network design plays a central role in the design of parallel systems. Most of the previous research has evaluated the performance of interconnection networks in isolation. In this study, we investigate the relationship between application program characteristics and interconnection network performance using an execution driven simulation test bed: the Reconfigurable Architecture Workbench (RAW). We simulate five topological configurations of a k-ary n-cube interconnect and four different network link models for a 4,096 node SIMD machine, and quantify the impact of the network on two application programs. We provide experimental evidence that such "in-context" simulation provides a better view of the impact of network design variables on system performance. We show that recent results, indicating that low-dimensional designs provide better ICN performance, ignore application requirements that may favor high-dimensional designs. Furthermore, applications that would appear to favor low dimensional designs may not, in fact, be significantly impacted by the network's dimensionality. We experimentally test the results of published performance models comparing the use of a synthetic load to that of a load generated by a typical application program.
Walter B. Ligon III, Umakishore Ramachandran
IEEE Trans. Parallel Distributed Syst.2
1996 Relaxed Index Consistency for a Client-Server Database
abstract
Client-server systems cache data in client buffers to deliver good performance. Several efficient protocols have been proposed to maintain the coherence of the cached data. However, none of the protocols distinguish between index pages and data pages. We propose a new coherence protocol, called relaxed index consistency, that exploits the inherent differences in the coherence and concurrency control (C&CC) requirements for index and data pages. The key idea is to incur a small increase in computation time at the clients to gain a significant reduction in the number of messages exchanged between the clients and the servers. The protocol uses concurrency control on data pages to maintain the coherence of index pages. A performance-conscious implementation of the protocol that makes judicious use of version numbers is proposed. We show, through both qualitative and quantitative analyses, the performance benefits of making the distinction between index pages and data pages for the purposes of C&CC. Our simulation studies show that the relaxed index consistency protocol improves system throughput by as much as 15% to 88%, based on the workload.
Vibby Gottemukkala, Edward Omiecinski, Umakishore Ramachandran
ICDE3
1996 Cache-Based Synchronization in Shared Memory Multiprocessors
Umakishore Ramachandran, Joonwon Lee
J. Parallel Distributed Comput.1
1996 Scalability Study of the KSR-1
Umakishore Ramachandran, Gautam Shah, Ravi Kumar 0001, Jeyakumar Muthukumarasamy
Parallel Comput.1
1995 Abstracting Network Characteristics and Locality Properties of Parallel Systems
abstract
Abstracting features of parallel systems is a technique that has been traditionally used in theoretical and analytical models for program development and performance evaluation. We explore the use of abstractions in execution-driven simulators in order to speed up simulation. In particular, we evaluate abstractions for the interconnection network and locality, properties of parallel systems in the context of simulating cache-coherent shared memory (CC-NUMA) multiprocessors. We use the recently proposed LogP model to abstract the network. We abstract locality by modeling a cache at each processing node in the system which is maintained coherent, without modeling the overheads associated with coherence maintenance. Such an abstraction tries to capture the true communication characteristics of the application without modeling any hardware induced artifacts. Using a suite of applications and three network topologies simulated on a novel simulation platform, we show that the latency overhead modeled by LogP is fairly accurate. On the other hand, the contention overhead can become pessimistic when the applications display sufficient communication locality. Our abstraction for data locality closely models the behavior of the target system over the chosen range of applications. The simulation model which incorporated these abstractions was around 250-300% faster than the simulation of the target machine.>
Anand Sivasubramaniam, Aman Singla, Umakishore Ramachandran, H. Venkateswaran
HPCA3
1995 The Quest for a Zero Overhead Shared Memory Parallel Machine
Gautam Shah, Aman Singla, Umakishore Ramachandran
ICPP (1)3
1995 Architectural Mechanisms for Explicit Communication in Shared Memory Multiprocessors
abstract
The goal of this work is to explore architectural mechanisms for supporting explicit communication in cache-coherent shared memory multiprocessors. The motivation stems from the observation that applications display wide diversity in terms of sharing characteristics and hence impose different communication requirements on the system. Explicit communication mechanisms would allow tailoring the coherence management under software control to match these differing needs and strive to provide a close approximation to a zero overhead machine from the application perspective. Toward achieving these goals, we first analyze the characteristics of sharing observed in certain specific applications. We then use these characteristics to synthesize explicit communication primitives. The proposed primitives allow selectively updating a set of processors, or requesting a stream of data ahead of its intended use. These primitives are essentially generalizations of prefetch and poststore, with the ability to specify the sharer set for poststore either statically or dynamically. The proposed primitives are to be used in conjunction with an underlying invalidation based protocol. Used in this manner, the resulting memory system can dynamically adapt itself to performing either invalidations or updates to match the communication needs. Through application driven performance study we show the utility of these mechanisms in being able to reduce and tolerate communication latencies.
Umakishore Ramachandran, Gautam Shah, Anand Sivasubramaniam, Aman Singla, Ivan Yanasak
SC1
1995 Timepatch: A Novel Technique for the Parallel Simulation of Multiprocessor Caches
abstract
No abstract available.
Gautam Shah, Umakishore Ramachandran, Richard M. Fujimoto
SIGMETRICS2
1995 On Characterizing Bandwidth Requirements of Parallel Applications
abstract
Synthesizing architectural requirements from an application viewpoint can help in making important architectural design decisions towards building large scale parallel machines. In this paper, we quantify the link bandwidth requirement on a binary hypercube topology for a set of five parallel applications. We use an execution-driven simulator called SPASM to collect data points for system sizes that are feasible to be simulated. These data points are then used in a regression analysis for projecting the link bandwidth requirements for larger systems. The requirements are projected as a function of the following system parameters: number of processors, CPU clock speed, and problem size. These results are also used to project the link bandwidths for other network topologies. Our study quantifies the link bandwidth that has to be made available to limit the network overhead in an application to a specified tolerance level. The results show that typical link bandwidths (200-300 MBytes/sec) found in current commercial parallel architectures (such as Intel Paragon and Cray T3D) would have fairly low network overhead for the applications considered in this study. For two of the applications, this overhead is negligible. For the other applications, this overhead can be limited to about 30% of the execution time provided the problem sizes are increased commensurate with the processor clock speed. The technique presented can be useful to a system architect to synthesize the bandwidth requirements for realizing well-balanced parallel architectures.
Anand Sivasubramaniam, Aman Singla, Umakishore Ramachandran, H. Venkateswaran
SIGMETRICS3
1994 An Approach to Scalability Study of Shared Memory Parallel Systems
abstract
The overheads in a parallel system that limit its scalability need to be identified and separated in order to enable parallel algorithm design and the development of parallel machines. Such overheads may be broadly classified into two components. The first one is intrinsic to the algorithm and arises due to factors such as the work-imbalance and the serial fraction. The second one is due to the interaction between the algorithm and the architecture and arises due to latency and contention in the network. A top-down approach to scalability study of shared memory parallel systems is proposed in this research. We define the notion of overhead functions associated with the different algorithmic and architectural characteristics to quantify the scalability of parallel systems; we isolate the algorithmic overhead and the overheads due to network latency and contention from the overall execution time of an application; we design and implement an execution-driven simulation platform that incorporates these methods for quantifying the overhead functions; and we use this simulator to study the scalability characteristics of five applications on shared memory platforms with different communication topologies.
Anand Sivasubramaniam, Aman Singla, Umakishore Ramachandran, H. Venkateswaran
SIGMETRICS3
1994 Evaluating Multigauge Architectures for Computer Vision
Walter B. Ligon III, Umakishore Ramachandran
J. Parallel Distributed Comput.2
1994 A Simulation-Based Scalability Study of Parallel Systems
Anand Sivasubramaniam, Aman Singla, Umakishore Ramachandran, H. Venkateswaran
J. Parallel Distributed Comput.3
1993 Scalability Study of the KSR-1
abstract
There has been concern in the architectural community regarding the scalability of shared memory parallel architectures owing to the potential for large latencies for remote memory accesses. KSR-1 is a re cently introduced commercial shared memory parallel architecture, and the scalability of KSR-1 is the focus of this research. Our key conclusions are as follows: The communication network of KSR-1 is fairly re silient in supporting simultaneous remote memory ac cesses from several processors. The multiple communi cation paths realized through this pipelining help in the efficient implementation of tournament-style barrier synchronization algorithms. The architectural features of KSR-1 such as the poststore and prefetch are useful for boosting the performance of parallel applications. The network does saturate when there are simultane ous remote memory accesses from a fully populated (32 node) ring.
Umakishore Ramachandran, Gautam Shah, Ravi Kumar 0001, Jeyakumar Muthukumarasamy
ICPP (1)1
1993 An Empirical Methodology for Exploring Reconfigurable Architecutures
Walter B. Ligon III, Umakishore Ramachandran
J. Parallel Distributed Comput.2
1992 A Distributed Hardware Barrier in an Optical Bus-Based Distributed Shared Memory Multiprocessor
Martin H. Davis Jr., Umakishore Ramachandran
ICPP (1)2
1991 Architectural Primitives for a Scalable Shared Memory Multiprocessor
abstract
Article Free Access Share on Architectural primitives for a scalable shared memory multiprocessor Authors: Joonwon Lee College of Computing, Georgia Institute of Technology, Atlanta, Georgia College of Computing, Georgia Institute of Technology, Atlanta, GeorgiaView Profile , Umakishore Ramachandran College of Computing, Georgia Institute of Technology, Atlanta, Georgia College of Computing, Georgia Institute of Technology, Atlanta, GeorgiaView Profile Authors Info & Claims SPAA '91: Proceedings of the third annual ACM symposium on Parallel algorithms and architecturesJune 1991 Pages 103–114https://doi.org/10.1145/113379.113389Published:01 June 1991Publication History 4citation145DownloadsMetricsTotal Citations4Total Downloads145Last 12 Months4Last 6 weeks1 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF
Joonwon Lee, Umakishore Ramachandran
SPAA2
1991 An Implementation of Distributed Shared Memory
abstract
Abstract Shared memory is a simple yet powerful paradigm for structuring systems. Recently, there has been an interest in extending this paradigm to non‐shared memory architectures as well. For example, the virtual address spaces for all objects in a distributed object‐based system could be viewed as constituting a globaldistributed shared memory. We propose a set of primitives for managing distributed shared memory. We present an implementation of these primitives in the context of an object‐based operating system as well as on top of Unix.
Umakishore Ramachandran, Yousef Y. A. Khalidi
Softw. Pract. Exp.1
1990 Synchronization with Multiprocessor Caches
abstract
Introducing private caches in bus-based shared memory multiprocessors leads to the cache consistency problem since there may be multiple copies of shared data. However, the ability to snoop on the bus coupled with the fast broadcast capability allows the design of special hardware support for synchronization. We present a new lock-based cache scheme which incorporates synchronization into the cache coherency mechanism. With this scheme high-level synchronization primitives as well as low-level ones can be implemented without excessive overhead. Cost functions for well-known synchronization methods are derived for invalidation schemes, write update schemes, and our lock-based scheme. To accurately predict the performance implications of the new scheme, a new simulation model is developed embodying a widely accepted paradigm of parallel programming. It is shown that our lock-based protocol outperforms existing cache protocols.
Joonwon Lee, Umakishore Ramachandran
ISCA2
1990 Hardware Support for Interprocess Communication
abstract
The use of a special-purpose coprocessor for supporting message passing is proposed. An actual message-based operating system is partitioned into computation and communication parts, executing, respectively, on a host and a message coprocessor which interact through shared queues. Its performance is measured on a multiprocessor. Hardware support in the form of a special-purpose smart bus and smart shared memory is designed. The benefits of these components are demonstrated through analytical modeling using generalized timed Petri nets. The analysis shows good agreement with experimental results and indicates that substantial benefits may be obtained when the software is partitioned between host and the message coprocessor and when a small amount of special-purpose hardware is added.>
Umakishore Ramachandran, Marvin H. Solomon, Mary K. Vernon
IEEE Trans. Parallel Distributed Syst.1
1989 Programming with distributed shared memory
abstract
In a distributed system, remote services may be provided either by a remote procedure call (RPC) mechanism or by paging in the required memory segments and performing the services locally. The latter approach, termed distributed shared memory (DSM) has several benefits given the current trend of structuring computing systems using diskless computational servers (workstations) and data servers (file servers). A set of distributed shared memory mechanisms that handle networkwide memory management for an object-based system is proposed. The implementation of these mechanisms is discussed and examples of their use in implementing the programming language Linda, process migration, two-phase commit and a distributed game are provided.>
Umakishore Ramachandran, Yousef Y. A. Khalidi
COMPSAC1
1989 A design of a memory management unit for object-based systems
abstract
Object-based operating systems have the desirable property of separating policies from mechanisms while providing a protected procedure call interface for accessing system services. However, the kernel mechanisms in such systems rely very heavily on efficient memory management. A set of criteria for supporting the kernel mechanisms in object-based systems is presented. The design of a simple MMU (memory management unit) tailored for object-based systems is presented. The proposed design is an engineering solution combining the features available in commercial MMUs.>
Umakishore Ramachandran, Yousef Y. A. Khalidi
ICCD1
1989 Coherence of Distributed Shared Memory: Unifying Synchronization and Data Transfer
Umakishore Ramachandran, Mustaque Ahamad, Yousef Y. A. Khalidi
ICPP (2)1
1989 A Measurement-based Study of Hardware Support for Object Invocation
Umakishore Ramachandran, Yousef Y. A. Khalidi
Softw. Pract. Exp.1
1987 Techniques for Reducing the Complexity of Large System Models
Umakishore Ramachandran, Marvin H. Solomon, Mary K. Vernon
ICPP1
1987 Hardware Support for Interprocess Communication
abstract
In recent years there has been increasing interest in message-based operating systems, particularly in distributed environments. Such systems consist of a small message-passing kernel supporting a collection of system server processes that provide such services as resource management, file service, and global communications. For such an architecture to be practical, it is essential that basic messages be fast, since they often replace what would be a simple procedure call or “kernel call” in a more traditional system. Careful study of several operating systems shows that the limiting factor, especially for small messages, is typically not network bandwidth but processing overhead. Therefore, we propose using a special-purpose coprocessor to support message passing. Our research has two parts: First, we partitioned an actual message-based operating system into communication and computation parts interacting through shared queues and measured its performance on a multiprocessor. Second, we designed hardware support in the form of a special-purpose smart bus and smart shared memory and demonstrated the benefits of these components through analytical modeling using Generalized Timed Petri Nets. Our analysis shows good agreement with the experimental results and indicates that substantial benefits may be obtained from both the partitioning of the software and the addition of a small amount of special-purpose hardware.
Umakishore Ramachandran, Marvin H. Solomon, Mary K. Vernon
ISCA1