EDBT 2026 Demo / reviewers in the wild / expert
Prashant J. Shenoy
dblp:s/PrashantJShenoy
· DBLP profile ↗
227ranked-venue papers
12as first author
41since 2021 · last 2026
0000-0002-5435-1901ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 68 · 3 first-author · 24 since 2021Computer networks · 61 · 1 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 32 · 8 first-authorDatabases, data management, data science and information retrieval · 30 · 2 since 2021Software engineering, systems software and programming languages · 22 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 14 · 3 since 2021Human-computer interaction and ubiquitous computing · 10 · 3 since 2021Artificial intelligence and machine learning · 8 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FM-CAC: Carbon-Aware Control for Battery-Buffered Edge AI via Time-Series Foundation Models
Kang Yang 0005, Walid A. Hanafy, Prashant J. Shenoy, Mani Srivastava 0001 |
ISLPED | 3 |
| 2026 | To Offload or Not To Offload: Model-driven Comparison of Edge-native and On-device Processing In the Era of AcceleratorsabstractComputational offloading is a promising approach to overcome client device resource constraints by moving application computations to remote servers. With the advent of specialized hardware accelerators, client devices can now perform fast local processing of tasks such as machine learning inference, reducing the need for offloading. However, edge servers with accelerators also offer faster offloading performance than was previously possible. In this paper, we present an analytic and experimental comparison of on-device processing and edge offloading across accelerator, multi-tenant, and workload scenarios to understand when to use local processing versus offloading. We present models that leverage queuing theory to derive explainable closed-form equations for end-to-end latencies, yielding quantitative performance crossover predictions to guide adaptive offloading. We validate our models across various settings and show that they achieve a mean absolute percentage error of 2.2% compared to observed latencies. We further use these models to develop a resource manager for adaptive offloading and demonstrate its effectiveness in dynamic multi-tenant edge environments. Nathan Ng 0002, David Irwin 0001, Ananthram Swami, Don Towsley, Prashant J. Shenoy |
ICPE | 5 |
| 2026 | CarbonShare: Carbon-Fair Allocation for Shared ClustersabstractComputing's energy demand, and thus its carbon emissions, are rapidly accelerating with the emergence of a wide range of useful, but computationally-intensive, AI-driven applications. At the same time, computing, along with the rest of society, must rapidly reduce its emissions to avoid the worst consequences of climate change, e.g., population displacement, agricultural collapse, mass extinction, extreme weather, etc. To do so, computing and other industries will eventually need to limit their carbon emissions. Such a limit imposes a new allocation problem: how should datacenters allocate their limited carbon emissions across multiple applications? John Thiede, David Irwin 0001, Prashant J. Shenoy |
ICPE | 3 |
| 2026 | Lilou: Resource-aware model-driven latency prediction for GPU-accelerated model serving
Qianlin Liang, Prashant J. Shenoy |
Perform. Evaluation | 3 |
| 2026 | Untangling the Carbon-Cost Tradeoffs and Stampede Effect Challenges in Cloud ComputingabstractAs society explores new computing applications, data centers have seen a sharp rise in capacity and energy use, driving up data centers’ carbon footprint and raising concerns about their sustainability. In response, efforts to reduce emissions now complement the traditional goals of cutting energy costs and boosting performance. Given the spatiotemporal variability of the grid carbon intensity, researchers are increasingly adopting workload shifting strategies to minimize the operational carbon footprint of cloud workloads. However, shifting from cost- and performance-centric operations to carbon-aware operations introduces performance penalties and cost overheads. Moreover, large-scale spatial and temporal shifts can cause stampede effects, with synchronized migration to the same low-carbon regions or times, thereby straining resources and creating imbalances. In this paper, we quantify the carbon-cost-performance trade-offs of carbon-aware scheduling and evaluate the risk of such stampede effects. We provide real-world examples illustrating these tradeoffs for application and cloud providers. Lastly, we offer practical insights to help mitigate these challenges and guide sustainable scheduling decisions. Walid A. Hanafy, Thanathorn Sukprasert, Abel Souza, David Irwin 0001, Prashant J. Shenoy |
IEEE Trans. Computers | 5 |
| 2025 | FailLite: Failure-Resilient Model Serving for Resource-Constrained Edge EnvironmentsabstractModel serving systems have become popular for deploying deep learning models for various latency-sensitive inference tasks. While traditional replication-based methods have been used for failure-resilient model serving in the cloud, such methods are often infeasible in edge environments due to significant resource constraints that preclude full replication. To address this problem, this paper presents FailLite, a failure-resilient model serving system that employs (i) a heterogeneous replication where the failover model is a smaller variant of the original one, (ii) an intelligent approach that uses warm replicas to ensure quick failover for critical applications while using cold replicas, and (iii) progressive failover to provide low mean time to recovery (MTTR) for the remaining applications. We implement a full prototype of our system and demonstrate its efficacy on an experimental edge testbed and large-scale simulations. Our results using 27 models show that FailLite can recover all failed applications with 2× lower MTTR and only a 0.6% reduction in accuracy. Under extreme failure scenarios, where 50% of edge sites fail simultaneously, FailLite improves recovery rate by at least 39.3% compared to the baseline methods. Walid A. Hanafy, Tarek F. Abdelzaher, David Irwin 0001, Jesse Milzman, Prashant J. Shenoy |
SoCC | 6 |
| 2025 | CarbonEdge: Leveraging Mesoscale Spatial Carbon-Intensity Variations for Low Carbon Edge ComputingabstractThe proliferation of latency-critical and compute-intensive edge applications is driving increases in computing demand and carbon emissions at the edge. To better understand carbon emissions at the edge, we analyze granular carbon intensity traces at intermediate "mesoscales," such as within a single US state or among neighboring countries in Europe, and observe significant variations in carbon intensity at these spatial scales. Importantly, our analysis shows that carbon intensity variations, which are known to occur at large continental scales (e.g., cloud regions), also occur at much finer spatial scales, making it feasible to exploit geographic workload shifting in the edge computing context. Motivated by these findings, we propose CarbonEdge, a carbon-aware framework for edge computing that optimizes the placement of edge workloads across mesoscale edge data centers to reduce carbon emissions while meeting latency SLOs. We implement CarbonEdge and evaluate it on a real edge computing testbed and through large-scale simulations for multiple edge workloads and settings. Our experimental results on a real testbed demonstrate that CarbonEdge can reduce emissions by up to 78.7% for a regional edge deployment in central Europe. Moreover, our CDN-scale experiments show potential savings of 49.5% and 67.8% in the US and Europe, respectively, while limiting the one-way latency increase to less than 5.5 ms. Walid A. Hanafy, Abel Souza, Jan Harkes, David Irwin 0001, Mahadev Satyanarayanan, Prashant J. Shenoy |
HPDC | 8 |
| 2025 | Ahead of the Curve: Leveraging Periodicity to Improve Job Placement in Data Centers
Xiaoding Guan, Noman Bashir, David Irwin 0001, Prashant J. Shenoy |
IC2E | 4 |
| 2025 | LLM-Driven Auto Configuration for Transient IoT Device CollaborationabstractToday's Internet of Things (IoT) has evolved from simple sensing and actuation devices to those with embedded processing and intelligent services, enabling rich collaborations between users and their devices. However, enabling such collaboration becomes challenging when transient devices need to interact with host devices in temporarily visited environments. In such cases, fine-grained access control policies are necessary to ensure secure interactions; however, manually implementing them is often impractical for non-expert users. Moreover, at run-time, the system must automatically configure the devices and enforce such fine-grained access control rules. Additionally, the system must address the heterogeneity of devices. Hetvi Shastri, Walid A. Hanafy, David Irwin 0001, Mani Srivastava 0001, Prashant J. Shenoy |
SEC | 6 |
| 2025 | Poster Abstract: LiveDetector: Towards Privacy-Preserving Voice Liveness DetectionabstractVoice-based authentication is widely used for secure access, yet it remains vulnerable to replay attacks. We present a lightweight privacy-aware liveness detection system leveraging acoustic feature engineering to distinguish genuine voices from spoofed attempts. Using the ASVspoof 2017 dataset, our method achieves EER of 16.2% while being faster than existing DNN-based liveness detection models. We show that higher order formant frequencies, reverberation, and group delay play a crucial role in liveness detection. Our approach provides a privacy-conscious method for preventing attacks on voice-based authentication systems, so that security does not come at the cost of privacy. Bhawana Chhaglani, Jeremy Gummeson, Prashant J. Shenoy |
SenSys | 3 |
| 2025 | Poster Abstract: Rethinking Collaboration Among Mobile Devices in IoT EnvironmentsabstractMany emerging IoT devices are mobile, enabling them to visit new environments and networks beyond their home networks. Mobile devices often have to interact and collaborate with users and their devices, which belong to the different administrative environments they are temporarily visiting. In this paper, we envision a system for seamless collaboration among transient devices in IoT environments. The system is based on zero-conf collaboration and allows for fine-grained access control. Our proposed design supports hardware-independent interfaces and supports a large number of devices. Hetvi Shastri, Walid A. Hanafy, David Irwin 0001, Mani Srivastava 0001, Prashant J. Shenoy |
SenSys | 6 |
| 2024 | Going Green for Less Green: Optimizing the Cost of Reducing Cloud Carbon EmissionsabstractThe continued exponential growth of cloud datacenter capacity has increased awareness of the carbon emissions when executing large compute-intensive workloads. To reduce carbon emissions, cloud users often temporally shift their batch workloads to periods with low carbon intensity. While such time shifting can increase job completion times due to their delayed execution, the cost savings from cloud purchase options, such as reserved instances, also decrease when users operate in a carbon-aware manner. This happens because carbon-aware adjustments change the demand pattern by periodically leaving resources idle, which creates a trade-off between carbon emissions and cost. In this paper, we present GAIA, a carbon-aware scheduler that enables users to address the three-way trade-off between carbon, performance, and cost in cloud-based batch schedulers. Our results quantify the carbon-performance-cost trade-off in cloud platforms and show that compared to existing carbon-aware scheduling policies, our proposed policies can double the amount of carbon savings per percentage increase in cost, while decreasing the performance overhead by 26%. Walid A. Hanafy, Qianlin Liang, Noman Bashir, Abel Souza, David Irwin 0001, Prashant J. Shenoy |
ASPLOS (3) | 6 |
| 2024 | SLO-Power: SLO and Power-aware Elastic Scaling for Web ServicesabstractManaging the performance of online web services in cloud data centers while optimizing resource allocation and power consumption is a multifaceted challenge. Often, resource and power management techniques, such as elastic scaling and power capping, are handled independently, leading to conflicts and sub-optimal power-performance trade-offs. To tackle this issue, we introduce SLO-Power, a system that coordinates the resource and power scaling techniques to achieve power savings while adhering to service level objectives (SLOs), such as tail latency constraints. Our approach employs a combination of analytic queuing models and feedback-driven techniques to jointly allocate resources and power to cloud applications in an SLO and power-aware manner. We implement a prototype of our system and evaluate it using realistic workloads to demonstrate its ability to harmonize elastic and power scaling, enabling enhanced resource utilization and reduced power consumption while ensuring the application performance. Our findings indicate that SLO-Power achieves exceptional power and resource efficiency, approaching near-optimal power-efficiency levels at 90%, all while preventing SLO violations. Furthermore, compared to state-of-the-art solutions, SLO-Power demonstrates lower P95 latency, accompanied by a 12% reduction in resource usage. Mehmet Savasci, Abel Souza, David Irwin 0001, Ahmed Ali-Eldin, Prashant J. Shenoy |
CCGrid | 6 |
| 2024 | CDN-Shifter: Leveraging Spatial Workload Shifting to Decarbonize Content Delivery NetworksabstractContent Delivery Networks (CDNs) are Internet-scale systems that deliver streaming and web content to users from many geographically distributed edge data centers. Since large CDNs can comprise hundreds of thousands of servers deployed in thousands of global data centers, they can consume a large amount of energy for their operations and thus are responsible for large amounts of Green House Gas (GHG) emissions. As these networks scale to cope with increased demand for bandwidth-intensive content, their emissions are expected to rise further, making sustainable design and operation an important goal for the future. Since different geographic regions vary in the carbon intensity and cost of their electricity supply, in this paper, we consider spatial shifting as a key technique to jointly optimize the carbon emissions and energy costs of a CDN. We present two forms of shifting: spatial load shifting, which operates within the time scale of minutes, and VM capacity shifting, which operates at a coarse time scale of days or weeks. The proposed techniques jointly reduce carbon and electricity costs while considering the performance impact of increased request latency from such optimizations. Using real-world traces from a large CDN and carbon intensity and energy prices data from electric grids in different regions, we show that increasing the latency by 60ms can reduce carbon emissions by up to 35.5%, 78.6%, and 61.7% across the US, Europe, and worldwide, respectively. In addition, we show that capacity shifting can increase carbon savings by up to 61.2%. Finally, we analyze the benefits of spatial shifting and show that it increases carbon savings from added solar energy by 68% and 130% in the US and Europe, respectively. Jorge Murillo, Walid A. Hanafy, David Irwin 0001, Ramesh K. Sitaraman, Prashant J. Shenoy |
SoCC | 5 |
| 2024 | TailClipper: Reducing Tail Response Time of Distributed Services Through System-Wide SchedulingabstractReducing tail latency has become a crucial issue for optimizing the performance of online cloud services and distributed applications. In distributed applications, there are many causes of high end-to-end tail latency, including operating system delays, request re-ordering due to fan-out/fanin, and network congestion. Although recent research has focused on reducing tail latency for individual application components, such as by replicating requests and scheduling, in this paper, we argue for a holistic approach for reducing the end-to-end tail latency across application components. We propose TailClipper, a distributed scheduler that tags each arriving request with an arrival timestamp, and propagates it across the microservices' call chain. TailClipper then uses arrival timestamps to implement an oldest request first scheduler that combines global first-come first serve with a limited form of processor sharing to reduce end-to-end tail latency. In doing so, TailClipper can counter the performance degradation caused by request reordering in multi-tiered and microservices-based applications. We implement TailClipper as a userspace Linux scheduler and evaluate it using cloud workload traces and a real-world microservices application. Compared to state-of-the-art schedulers, our experiments reveal that TailClipper improves the 99th percentile response time by up to 81%, while also improving the mean response time and the system throughput by up to 54% and 29% respectively under high loads. Nathan Ng 0002, Abel Souza, Ahmed Ali-Eldin, David Irwin 0001, Don Towsley, Prashant J. Shenoy |
SoCC | 6 |
| 2024 | On the Limitations of Carbon-Aware Temporal and Spatial Workload Shifting in the CloudabstractCloud platforms have been focusing on reducing their carbon emissions by shifting workloads across time and locations to when and where low-carbon energy is available. Despite the prominence of this idea, prior work has only quantified the potential of spatiotemporal workload shifting in narrow settings, i.e., for specific workloads in select regions. In particular, there has been limited work on quantifying an upper bound on the ideal and practical benefits of carbon-aware spatiotemporal workload shifting for a wide range of cloud workloads. To address the problem, we conduct a detailed data-driven analysis to understand the benefits and limitations of carbon-aware spatiotemporal scheduling for cloud workloads. We utilize carbon intensity data from 123 regions, encompassing most major cloud sites, to analyze two broad classes of workloads---batch and interactive---and their various characteristics, e.g., job duration, deadlines, and SLOs. Our findings show that while spatiotemporal workload shifting can reduce workloads' carbon emissions, the practical upper bounds of these carbon reductions are currently limited and far from ideal. We also show that simple scheduling policies often yield most of these reductions, with more sophisticated techniques yielding little additional benefit. Notably, we also find that the benefit of carbon-aware workload scheduling relative to carbon-agnostic scheduling will decrease as the energy supply becomes "greener." Thanathorn Sukprasert, Abel Souza, Noman Bashir, David Irwin 0001, Prashant J. Shenoy |
EuroSys | 5 |
| 2024 | Acies-OS: A Content-Centric Platform for Edge AI Twinning and OrchestrationabstractThis paper describes Acies-OS, a content-centric platform for edge AI twinning and orchestration that allows easy deployment, re-configuration, and control of edge AI services, augmented by a digital twin. The work is motivated by the proliferation of edge AI in a plethora of IoT applications, ranging from home automation to military defense, and the emergence of digital twins that go beyond monitoring and emulation into configuration management and optimization of edge capabilities. While past work focused on either the edge capabilities themselves or the digital twin, this work focuses on their seamless interactions, offering abstractions that enable the digital twin to manage and optimize an increasingly diverse edge AI system. Acies-OS features a structured namespace, a thin client library with flexible pub/sub-based communication, health monitoring support, and a control plane for twin-based value-added analysis and optimization. To illustrate the use of Acies-OS, we implemented a multi-node multi-modality vehicle classification application and used Acies-OS to interface it to a digital twin. We then deployed the system in the field to showcase run-time twin-based optimizations of inference latency, classification accuracy, and robustness to failures in noisy and challenging conditions. Jinyang Li 0004, Yizhuo Chen, Tomoyoshi Kimura, Tianshi Wang 0002, Ruijie Wang 0004, Denizhan Kara, Yigong Hu, Walid A. Hanafy, Abel Souza, Prashant J. Shenoy, Maggie B. Wigness, Joydeep Bhattacharyya, Jae Kim, Guijun Wang, Greg Kimberly, Josh D. Eckhardt, Denis Osipychev, Tarek F. Abdelzaher |
ICCCN | 11 |
| 2024 | Chasing Convex Functions with Long-term ConstraintsabstractWe introduce and study a family of online metric problems with long-term constraints. In these problems, an online player makes decisions $\mathbf{x}_t$ in a metric space $(X,d)$ to simultaneously minimize their hitting cost $f_t(\mathbf{x}_t)$ and switching cost as determined by the metric. Over the time horizon $T$, the player must satisfy a long-term demand constraint $\sum_t c(\mathbf{x}_t) \geq 1$, where $c(\mathbf{x}_t)$ denotes the fraction of demand satisfied at time $t$. Such problems can find a wide array of applications to online resource allocation in sustainable energy/computing systems. We devise optimal competitive and learning-augmented algorithms for the case of bounded hitting cost gradients and weighted $\ell_1$ metrics, and further show that our proposed algorithms perform well in numerical experiments. Adam Lechowicz, Nicolas Christianson, Bo Sun 0004, Noman Bashir, Mohammad Hajiesmaili, Adam Wierman, Prashant J. Shenoy |
ICML | 7 |
| 2024 | INVAR: Inversion Aware Resource Provisioning and Workload Scheduling for Edge ComputingabstractEdge computing is emerging as a complementary architecture to cloud computing to address some of its associated issues. One of the major advantages of edge computing is that edge data centers are usually much closer to users compared to traditional cloud data centers. Therefore, it is commonly believed that for developers of latency-sensitive applications, they can effectively reduce the overall end-to-end latency by simply transitioning from a cloud deployment to an edge deployment. However, as recent work has shown, the performance of an edge deployment is vulnerable to a couple of factors which under many practical scenarios can lead to edge servers providing worse end-to-end response time than cloud servers. This phenomenon is referred to as edge performance inversion. In this paper, we propose resource allocation and workload scheduling algorithms that actively prevent edge performance inversion. Our algorithms, named INVAR, are based on queueing theory results and optimization techniques. Evaluation results show that INVAR can find a near-optimal solution that outperforms the performance of a cloud deployment by an adjustable margin. Simulation results based on production workloads from Akamai data centers show that INVAR can outperform common heuristic-based edge deployment by 11% to 24% in real-world scenarios. David Irwin 0001, Prashant J. Shenoy, Don Towsley |
INFOCOM | 3 |
| 2024 | W4-Groups: Modeling the Who, What, When and Where of Group Behavior via Mobility SensingabstractHuman social interactions occur in group settings of varying sizes and locations, depending on the type of social activity. The ability to distinguish group formations based on their purposes transforms how group detection mechanisms function. Not only should such tools support the effective detection of serendipitous encounters, but they can derive categories of relation types among users. Determining who is involved, what activity is performed, and when and where the activity occurs are critical to understanding group processes in greater depth, including supporting goal-oriented applications (e.g., performance, productivity, and mental health) that require sensing social factors. In this work, we propose W4-Groups that captures the functional perspective of variability and repeatability when automatically constructing short-term and long-term groups via multiple data sources (e.g., WiFi and location check-in data). We design and implement W4-Groups to detect and extract all four group features who-what-when-where from the user's daily mobility patterns. We empirically evaluate the framework using two real-world WiFi datasets and a location check-in dataset, yielding an average of 92% overall accuracy, 96% precision, and 94% recall. Further, we supplement two case studies to demonstrate the application of W4-Groups for next-group activity prediction and analyzing changes in group behavior at a longitudinal scale, exemplifying short-term and long-term occurrences. Akanksha Atrey, Camellia Zakaria, Rajesh Krishna Balan, Prashant J. Shenoy |
Proc. ACM Hum. Comput. Interact. | 4 |
| 2023 | Ecovisor: A Virtual Energy System for Carbon-Efficient ApplicationsabstractCloud platforms' rapid growth is raising significant concerns about their carbon emissions. To reduce carbon emissions, future cloud platforms will need to increase their reliance on renewable energy sources, such as solar and wind, which have zero emissions but are highly unreliable. Unfortunately, today's energy systems effectively mask this unreliability in hardware, which prevents applications from optimizing their carbon-efficiency, or work done per kilogram of carbon emitted. To address the problem, we design an "ecovisor", which virtualizes the energy system and exposes software-defined control of it to applications. An ecovisor enables each application to handle clean energy's unreliability in software based on its own specific requirements. We implement a small-scale ecovisor prototype that virtualizes a physical energy system to enable software-based application-level i) visibility into variable grid carbon-intensity and local renewable generation and ii) control of server power usage and battery charging and discharging. We evaluate the ecovisor approach by showing how multiple applications can concurrently exercise their virtual energy system in different ways to better optimize carbon-efficiency based on their specific requirements compared to general system-wide policies. Abel Souza, Noman Bashir, Jorge Murillo, Walid A. Hanafy, Qianlin Liang, David Irwin 0001, Prashant J. Shenoy |
ASPLOS (2) | 7 |
| 2023 | Carbon Containers: A System-level Facility for Managing Application-level Carbon EmissionsabstractTo reduce their environmental impact, cloud datacenters' are increasingly focused on optimizing applications' carbon-efficiency, or work done per mass of carbon emitted. To facilitate such optimizations, we present Carbon Containers, a simple system-level facility, which extends prior work on power containers, that automatically regulates applications' carbon emissions in response to variations in both their work-load's intensity and their energy's carbon-intensity. Specifically, Carbon Containers enable applications to specify a maximum carbon emissions rate (in g.CO2e/hr), and then transparently enforce this rate via a combination of vertical scaling, container migration, and suspend/resume while maximizing either energy-efficiency or performance. John Thiede, Noman Bashir, David Irwin 0001, Prashant J. Shenoy |
SoCC | 4 |
| 2023 | SODA: Protecting Proprietary Information in On-Device Machine Learning ModelsabstractThe growth of low-end hardware has led to a proliferation of machine learning-based services in edge applications. These applications gather contextual information about users and provide some services, such as personalized offers, through a machine learning (ML) model. A growing practice has been to deploy such ML models on the user's device to reduce latency, maintain user privacy, and minimize continuous reliance on a centralized source. However, deploying ML models on the user's edge device can leak proprietary information about the service provider. In this work, we investigate on-device ML models that are used to provide mobile services and demonstrate how simple attacks can leak proprietary information of the service provider. We show that different adversaries can easily exploit such models to maximize their profit and accomplish content theft. Motivated by the need to thwart such attacks, we present an end-to-end framework, SODA, for deploying and serving on edge devices while defending against adversarial usage. Our results demonstrate that SODA can detect adversarial usage with 89% accuracy in less than 50 queries with minimal impact on service performance, latency, and storage. Akanksha Atrey, Ritwik Sinha, Saayan Mitra, Prashant J. Shenoy |
SEC | 4 |
| 2023 | Energy Time Fairness: Balancing Fair Allocation of Energy and Time for GPU WorkloadsabstractTraditionally, multi-tenant cloud and edge platforms use fair-share schedulers to fairly multiplex resources across applications. These schedulers ensure applications receive processing time proportional to a configurable share of the total time. Unfortunately, enforcing time-fairness across applications often violates energy-fairness, such that some applications consume more than their fair share of energy. This occurs because applications either do not fully utilize their resources or operate at a reduced frequency/voltage during their time-slice. The problem is particularly acute for machine learning (ML) applications using GPUs, where model size largely dictates utilization and energy usage. Enforcing energy-fairness is also important since energy is a costly and limited resource. For example, in cloud platforms, energy dominates operating costs and is limited by the power delivery infrastructure, while in edge platforms, energy is often scarce and limited by energy harvesting and battery constraints. Qianlin Liang, Walid A. Hanafy, Noman Bashir, David Irwin 0001, Prashant J. Shenoy |
SEC | 5 |
| 2023 | RAVAS: Interference-Aware Model Selection and Resource Allocation for Live Edge Video AnalyticsabstractNumerous edge applications that rely on video analytics demand precise, low-latency processing of multiple video streams from cameras. When these cameras are mobile, such as when mounted on a car or a robot, the processing load on the shared edge GPU can vary considerably. Provisioning the edge with GPUs for the worst-case load can be expensive and, for many applications, not feasible. Ali Rahmanian, Ahmed Ali-Eldin, Selome Kostentinos Tesfatsion, Björn Skubic, Harald Gustafsson, Prashant J. Shenoy, Erik Elmroth |
SEC | 6 |
| 2023 | Understanding the Benefits of Hardware-Accelerated Communication in Model-Serving ApplicationsabstractIt is commonly assumed that the end-to-end networking performance of edge offloading is purely dictated by that of the network connectivity between end devices and edge computing facilities, where ongoing innovation in 5G/6G networking can help. However, with the growing complexity of edge-offloaded computation and dynamic load balancing requirements, an offloaded task often goes through a multi-stage pipeline that spans across multiple compute nodes and proxies interconnected via a dedicated network fabric within a given edge computing facility. As the latest hardware-accelerated transport technologies such as RDMA and GPUDirect RDMA are adopted to build such network fabric, there is a need for good understanding of the full potential of these technologies in the context of computation offload and the effect of different factors such as GPU scheduling and characteristics of computation on the net performance gain achievable by these technologies. This paper unveils detailed insights into the latency overhead in typical machine learning (ML)-based computation pipelines and analyzes the potential benefits of adopting hardware-accelerated communication. To this end, we build a model-serving framework that supports various communication mechanisms. Using the framework, we identify performance bottlenecks in state-of-the-art model-serving pipelines and show how hardware-accelerated communication can alleviate them. For example, we show that GPUDirect RDMA can save 15-50% of model-serving latency, which amounts to 70–160 ms. Walid A. Hanafy, Limin Wang 0010, Hyunseok Chang, Sarit Mukherjee, T. V. Lakshman, Prashant J. Shenoy |
IWQoS | 6 |
| 2023 | DDPC: Automated Data-Driven Power-Performance Controller Design on-the-fly for Latency-sensitive Web ServicesabstractTraditional power reduction techniques such as DVFS or RAPL are challenging to use with web services because they significantly affect the services’ latency and throughput. Previous work suggested the use of controllers based on control theory or machine learning to reduce performance degradation under constrained power. However, generating these controllers is challenging as every web service applications running in a data center requires a power-performance model and a fine-tuned controller. In this paper, we present DDPC, a system for autonomic data-driven controller generation for power-latency management. DDPC automates the process of designing and deploying controllers for dynamic power allocation to manage the power-performance trade-offs for latency-sensitive web applications such as a social network. For each application, DDPC uses system identification techniques to learn an adaptive power-performance model that captures the application’s power-latency trade-offs which is then used to generate and deploy a Proportional-Integral (PI) power controller with gain-scheduling to dynamically manage the power allocation to the server running application using RAPL. We evaluate DDPC with two realistic latency-sensitive web applications under varying load scenarios. Our results show that DDPC is capable of autonomically generating and deploying controllers within a few minutes reducing the active power allocation of a web-server by more than 50% compared to state-of-the-art techniques while maintaining the latency well below the target of the application. Mehmet Savasci, Ahmed Ali-Eldin, Johan Eker, Anders Robertsson, Prashant J. Shenoy |
WWW | 5 |
| 2023 | WattScope: Non-intrusive application-level power disaggregation in datacenters
Xiaoding Guan, Noman Bashir, David Irwin 0001, Prashant J. Shenoy |
Perform. Evaluation | 4 |
| 2023 | Model-driven Cluster Resource Management for AI Workloads in Edge CloudsabstractSince emerging edge applications such as Internet of Things (IoT) analytics and augmented reality have tight latency constraints, hardware AI accelerators have been recently proposed to speed up deep neural network (DNN) inference run by these applications. Resource-constrained edge servers and accelerators tend to be multiplexed across multiple IoT applications, introducing the potential for performance interference between latency-sensitive workloads. In this article, we design analytic models to capture the performance of DNN inference workloads on shared edge accelerators, such as GPU and edgeTPU, under different multiplexing and concurrency behaviors. After validating our models using extensive experiments, we use them to design various cluster resource management algorithms to intelligently manage multiple applications on edge accelerators while respecting their latency constraints. We implement a prototype of our system in Kubernetes and show that our system can host 2.3× more DNN applications in heterogeneous multi-tenant edge clusters with no latency violations when compared to traditional knapsack hosting algorithms. Qianlin Liang, Walid A. Hanafy, Ahmed Ali-Eldin, Prashant J. Shenoy |
ACM Trans. Auton. Adapt. Syst. | 4 |
| 2022 | PeakTK: An Open Source Toolkit for Peak Forecasting in Energy SystemsabstractAs the electric grid undergoes the transition to a carbon free future, many new techniques for optimizing the grid’s energy usage and carbon footprint are being designed. A common technique used by many approaches is to reduce the energy usage of the grid’s peak demand periods since doing so is beneficial for reducing the carbon usage of the grid. Consequently, the design of peak forecasting methods that predict when and how much peak demand will be seen is at the heart of many energy optimization approaches. In this paper, we present PeakTK, an open-source toolkit and reference datasets for peak forecasting in energy systems. PeakTK implements a range of peak forecasting methods that have been proposed recently and exposes them through well-defined interfaces and library modules. Our goal is to improve reproducibility of energy systems research by providing a common framework for evaluating and comparing new peak forecasting algorithms. Further, PeakTK provides libraries to enable researchers and practitioners to easily incorporate peak forecasting methods into their research when implementing higher level grid optimizations. We discuss the design and implementation of PeakTK and present case studies to demonstrate how PeakTK can be used for forecasting or quantitative comparisons of energy optimization methods. Phuthipong Bovornkeeratiroj, John Wamburu, David Irwin 0001, Prashant J. Shenoy |
COMPASS | 4 |
| 2022 | CAVE: caching 360° videos at the edgeabstractWhile 360° videos are gaining popularity due to the emergence of VR technologies, storing and streaming such videos can incur up to 20X higher overheads than traditional HD content. Edge caching, which involves caching and serving 360° videos from edge servers, is one possible approach for addressing these overheads. Prior work on 360° video caching has been based on using past history to cache tiles that are likely to be in a viewer's field of view and has not considered methods to intelligently share a limited edge cache across a set of videos that exhibit large variations in their popularity, size, content, and user abandonment patterns. Towards this end, we present CAVE, an adaptive edge caching framework that intelligently optimizes cache allocation across a set of videos taking into account video content, size, and popularity. Our experiments using realistic video workloads shows CAVE improves cache hit-rates, and thus network saving, by up to 50% over state-of-the-art approaches, while also scaling to up to two thousand videos per edge cache. In addition, in terms of scalability, our developed algorithm is embarrassingly parallel, allowing CAVE to scale beyond state-of-the-art solutions that typically do not support parallelization. Ahmed Ali-Eldin, Chirag Goel, Mayank Jha, Bo Chen 0025, Klara Nahrstedt, Prashant J. Shenoy |
NOSSDAV | 6 |
| 2021 | Good Things Come to Those Who Wait: Optimizing Job Waiting in the CloudabstractCloud-enabled schedulers execute jobs on either fixed resources or those acquired on demand from cloud platforms. Thus, these schedulers must define not only a scheduling policy, which selects which jobs run when fixed resources become available, but also a waiting policy, which selects which jobs wait for fixed resources when they are not available, rather than run on on-demand resources. As with scheduling policies, optimizing waiting policies requires a priori knowledge of job runtime. Unfortunately, prior work has shown that accurately predicting job runtime is challenging. In this paper, we show that optimizing job waiting in the cloud is possible without accurate job runtime predictions. To do so, we i) speculatively execute jobs on on-demand resources for a small time and cost to learn more about job runtime, and ii) develop a ML model to predict wait time from cluster state, which is more accurate and has less overhead than prior approaches that use job runtime predictions. We evaluate our approach on a year-long batch workload consisting of 14 million jobs, and show that it yields a cost and average wait time within 4% and 13%, respectively, of the optimal. Lurdh Pradeep Reddy Ambati, Noman Bashir, David Irwin 0001, Prashant J. Shenoy |
SoCC | 4 |
| 2021 | Enabling Sustainable Clouds: The Case for Virtualizing the Energy SystemabstractCloud platforms' growing energy demand and carbon emissions are raising concern about their environmental sustainability. The current approach to enabling sustainable clouds focuses on improving energy-efficiency and purchasing carbon offsets. These approaches have limits: many cloud data centers already operate near peak efficiency, and carbon offsets cannot scale to near zero carbon where there is little carbon left to offset. Instead, enabling sustainable clouds will require applications to adapt to when and where unreliable low-carbon energy is available. Applications cannot do this today because their energy use and carbon emissions are not visible to them, as the energy system provides the rigid abstraction of a continuous, reliable energy supply. This vision paper instead advocates for a "carbon first" approach to cloud design that elevates carbon-efficiency to a firs--class metric. To do so, we argue that cloud platforms should virtualize the energy system by exposing visibility into, and software-defined control of, it to applications, enabling them to define their own abstractions for managing energy and carbon emissions based on their own requirements. Noman Bashir, Tian Guo 0001, Mohammad Hajiesmaili, David Irwin 0001, Prashant J. Shenoy, Ramesh K. Sitaraman, Abel Souza, Adam Wierman |
SoCC | 5 |
| 2021 | WiFiMod: Transformer-based Indoor Human Mobility Modeling using Passive SensingabstractModeling human mobility has a wide range of applications from urban planning to simulations of disease spread. It is well known that humans spend 80% of their time indoors but modeling indoor human mobility is challenging due to three main reasons: (i) the absence of easily acquirable, reliable, low-cost indoor mobility datasets, (ii) high prediction space in modeling the frequent indoor mobility, and (iii) multi-scalar periodicity and correlations in mobility. To deal with all these challenges, we propose WiFiMod, a Transformer-based, data-driven approach that models indoor human mobility at multiple spatial scales using WiFi system logs. WiFiMod takes as input enterprise WiFi system logs to extract human mobility trajectories from smartphone digital traces. Next, for each extracted trajectory, we identify the mobility features at multiple spatial scales, macro and micro, to design a multi-modal embedding Transformer that predicts user mobility for several hours to an entire day across multiple spatial granularities. Multi-modal embedding captures the mobility periodicity and correlations across various scales while Transformers capture long term mobility dependencies boosting model prediction performance. This approach significantly reduces the prediction space by first predicting macro mobility, then modeling indoor scale mobility, micro mobility, conditioned on the estimated macro mobility distribution, thereby using the topological constraint of the macro-scale. Experimental results show that WiFiMod achieves a prediction accuracy of at least 10% points higher than the current state-of-art models. Additionally, we present 3 real-world applications of WiFiMod - (i) predict high density hot pockets and space utilization for policy making decisions for COVID19 or ILI, (ii) generate a realistic simulation of indoor mobility data to simulate spread of diseases, (iii) design personal assistants. Amee Trivedi, Kate Silverstein, Emma Strubell, Prashant J. Shenoy, Mohit Iyyer |
COMPASS | 4 |
| 2021 | LaSS: Running Latency Sensitive Serverless Computations at the EdgeabstractServerless computing has emerged as a new paradigm for running short-lived computations in the cloud. Due to its ability to handle IoT workloads, there has been considerable interest in running serverless functions at the edge. However, the constrained nature of the edge and the latency sensitive nature of workloads result in many challenges for serverless platforms. In this paper, we present LaSS, a platform that uses model-driven approaches for running latency-sensitive serverless computations on edge resources. LaSS uses principled queuing-based methods to determine an appropriate allocation for each hosted function and auto-scales the allocated resources in response to workload dynamics. LaSS uses a fair-share allocation approach to guarantee a minimum of allocated resources to each function in the presence of overload. In addition, it utilizes resource reclamation methods based on container deflation and termination to reassign resources from over-provisioned functions to under-provisioned ones. We implement a prototype of our approach on an OpenWhisk serverless edge cluster and conduct a detailed experimental evaluation. Our results show that LaSS can accurately predict the resources needed for serverless functions in the presence of highly dynamic workloads, and reprovision container capacity within hundreds of milliseconds while maintaining fair share allocation guarantees. Ahmed Ali-Eldin, Prashant J. Shenoy |
HPDC | 3 |
| 2021 | Preserving Privacy in Personalized Models for Distributed Mobile ServicesabstractThe ubiquity of mobile devices has led to the proliferation of mobile services that provide personalized and context-aware content to their users. Modern mobile services are distributed between end-devices, such as smartphones, and remote servers that reside in the cloud. Such services thrive on their ability to predict future contexts to pre-fetch content or make context-specific recommendations. An increasingly common method to predict future contexts, such as location, is via machine learning (ML) models. Recent work in context prediction has focused on ML model personalization where a personalized model is learned for each individual user in order to tailor predictions or recommendations to a user's mobile behavior. While the use of personalized models increases efficacy of the mobile service, we argue that it increases privacy risk since a personalized model encodes contextual behavior unique to each user. To demonstrate these privacy risks, we present several attribute inference-based privacy attacks and show that such attacks can leak privacy with up to 78% efficacy for top-3 predictions. We present Pelican, a privacy-preserving personalization system for context-aware mobile services that leverages both device and cloud resources to personalize ML models while minimizing the risk of privacy leakage for users. We evaluate Pelican using real world traces for location-aware mobile services and show that Pelican can substantially reduce privacy leakage by up to 75%. Akanksha Atrey, Prashant J. Shenoy, David D. Jensen |
ICDCS | 2 |
| 2021 | Spark-based Cloud Data Analytics using Multi-Objective OptimizationabstractData analytics in the cloud has become an integral part of enterprise businesses. Big data analytics systems, however, still lack the ability to take task objectives such as user performance goals and budgetary constraints and automatically configure an analytic job to achieve these objectives. This paper presents UDAO, a Spark-based Unified Data Analytics Optimizer that can automatically determine a cluster configuration with a suitable number of cores as well as other system parameters that best meet the task objectives. At a core of our work is a principled multi-objective optimization (MOO) approach that computes a Pareto optimal set of configurations to reveal tradeoffs between different objectives, recommends a new Spark configuration that best explores such tradeoffs, and employs novel optimizations to enable such recommendations within a few seconds. Detailed experiments using benchmark workloads show that our MOO techniques provide a 2-50× speedup over existing MOO methods, while offering good coverage of the Pareto frontier. Compared to Ottertune, a state-of-the-art performance tuning system, UDAO recommends Spark configurations that yield 26%-49% reduction of running time of the TPCx-BB benchmark while adapting to different user preferences on multiple objectives. Khaled Zaouk, Chenghao Lyu, Yanlei Diao, Prashant J. Shenoy |
ICDE | 7 |
| 2021 | The hidden cost of the edge: a performance comparison of edge and cloud latenciesabstractEdge computing has emerged as a popular paradigm for running latency-sensitive applications due to its ability to offer lower network latencies to end-users. In this paper, we argue that despite its lower network latency, the resource-constrained nature of the edge can result in higher end-to-end latency, especially at higher utilizations, when compared to cloud data centers. We study this edge performance inversion problem through an analytic comparison of edge and cloud latencies and analyze conditions under which the edge can yield worse performance than the cloud. To verify our analytic results, we conduct a detailed experimental comparison of the edge and the cloud latencies using a realistic application and real cloud workloads. Both our analytical and experimental results show that even at moderate utilizations, the edge queuing delays can offset the benefits of lower network latencies, and even result in performance inversion where running in the cloud would provide superior latencies. We finally discuss practical implications of our results and provide insights into how application designers and service providers should design edge applications and systems to avoid these pitfalls. Ahmed Ali-Eldin, Prashant J. Shenoy |
SC | 3 |
| 2021 | Deep Contextualized Compressive Offloading for ImagesabstractRecent years have witnessed sensors becoming an indispensable part of our life with the camera being one of the most popular and widely deployed sensors. The camera gives rise to numerous vision-based IoT applications that generate high-level understandings of a live video stream by performing analysis on end devices like mobile or embedded devices. Typically, these applications are built with deep learning (DL) models to conduct complex vision tasks, e.g., image classification and object detection. Due to the prohibitive cost of running DL models on end devices close to the camera and with limited computation capabilities, it is widely adopted to offload the computation to a nearby powerful edge server. However, there is a gap between the restricted offloading bandwidth of the end device and the large volume of image data incurred by the live video stream. In this paper, we present Deep Contextualized Compressive Offloading for Images (DCCOI), a lightweight, context-aware, and bandwidth-efficient offloading framework for images. DCCOI consists of the spatial-adaptive encoder, a lightweight neural network, to spatial-adaptively compress the image, and the generative decoder for reconstructing the image from the compressed data. In contrast to existing DL-based encoders, the spatial-adaptive encoder allows an image region to be encoded into different numbers of feature values based on the information in it. This offers a variable-length coding method for image compression, which is a more optimal way for compression than the fix-length coding method took by existing DL-based compression approaches and demonstrates superior accuracy-compression rate trade-offs. We evaluate DCCOI against several baseline compression techniques while serving an object detection-based application. The results show that DCCOI roughly reduces the offloading size of JPEG by a factor of 9 and DeepCOD, the state-of-the-art offloading approach, by 20% with similar accuracy and a compression overhead less than 50ms. Bo Chen 0025, Zhisheng Yan, Hongpeng Guo, Zhe Yang 0010, Ahmed Ali-Eldin, Prashant J. Shenoy, Klara Nahrstedt |
SenSys | 6 |
| 2021 | Model-driven Per-panel Solar Anomaly Detection for Residential ArraysabstractThere has been significant growth in both utility-scale and residential-scale solar installations in recent years, driven by rapid technology improvements and falling prices. Unlike utility-scale solar farms that are professionally managed and maintained, smaller residential-scale installations often lack sensing and instrumentation for performance monitoring and fault detection. As a result, faults may go undetected for long periods of time, resulting in generation and revenue losses for the homeowner. In this article, we present SunDown, a sensorless approach designed to detect per-panel faults in residential solar arrays. SunDown does not require any new sensors for its fault detection and instead uses a model-driven approach that leverages correlations between the power produced by adjacent panels to detect deviations from expected behavior. SunDown can handle concurrent faults in multiple panels and perform anomaly classification to determine probable causes. Using two years of solar generation data from a real home and a manually generated dataset of multiple solar faults, we show that SunDown has a Mean Absolute Percentage Error of 2.98% when predicting per-panel output. Our results show that SunDown is able to detect and classify faults, including from snow cover, leaves and debris, and electrical failures with 99.13% accuracy, and can detect multiple concurrent faults with 97.2% accuracy. Menghong Feng, Noman Bashir, Prashant J. Shenoy, David Irwin 0001, Beka Kosanovic |
ACM Trans. Cyber Phys. Syst. | 3 |
| 2021 | Modeling and Analyzing Waiting Policies for Cloud-Enabled SchedulersabstractCloud platforms have popularized the Infrastructure-as-a-Service (IaaS) purchasing model, which enables users to rent computing resources on demand to execute their jobs. However, buying fixed resources is still much cheaper than renting if their resource utilization is high. Thus, to optimize cost, users must decide how many fixed resources to provision versus rent “on demand” based on their workload. In this article, we introduce the concept of a waiting policy for cloud-enabled schedulers and show that the optimal cost depends on it. The waiting policy explicitly controls how long jobs wait for resources, as jobs never need to wait, since cloud platforms provide the illusion of infinite scalability. A waiting policy is the dual of a scheduling policy: while a scheduling policy determines which jobs should run when fixed resources are available, a waiting policy determines which jobs should wait when fixed resources are not available. We define multiple waiting policies and develop simple and general analytical models to reveal their tradeoff between fixed resource provisioning, cost, and job waiting time. We evaluate the impact of different waiting policies on a real year-long batch workload consisting of 14M jobs run on a 14.3k-core cluster. We show that a compound waiting policy, which forces jobs with long running times or short waiting times to wait for fixed resources, offers the best tradeoff. The policy decreases both the cost (by 5 percent) and mean job waiting time (by 7×) compared to the current cluster, and also decreases the cost (by 43 percent) compared to renting on-demand resources for a modest increase in mean job waiting time (at 1.74 hours). Lurdh Pradeep Reddy Ambati, Noman Bashir, David Irwin 0001, Prashant J. Shenoy |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2020 | SunDown: Model-driven Per-Panel Solar Anomaly Detection for Residential ArraysabstractSolar arrays often experience faults that go undetected for long periods of time, resulting in generation and revenue losses. In this paper, we present SunDown, a sensorless approach for detecting per-panel faults in solar arrays. SunDown's model-driven approach leverages correlations between the power produced by adjacent panels to detect deviations from expected behavior, can handle concurrent faults in multiple panels, and performs anomaly classification to determine probable causes. Using two years of solar data from a real home and a manually generated dataset of solar faults, we show that our approach is able to detect and classify faults, including from snow, leaves and debris, and electrical failures with 99.13% accuracy, and can detect concurrent faults with 97.2% accuracy. Menghong Feng, Noman Bashir, Prashant J. Shenoy, David Irwin 0001, Dragoljub Kosanovic |
COMPASS | 3 |
| 2020 | Cloud-scale VM-deflation for Running Interactive Applications On Transient ServersabstractTransient computing has become popular in public cloud environments for running delay-insensitive batch and data processing applications at low cost. Since transient cloud servers can be revoked at any time by the cloud provider, they are considered unsuitable for running interactive application such as web services. In this paper, we present VM deflation as an alternative mechanism to server preemption for reclaiming resources from transient cloud servers under resource pressure. Using real traces from top-tier cloud providers, we show the feasibility of using VM deflation as a resource reclamation mechanism for interactive applications in public clouds. We show how current hypervisor mechanisms can be used to implement VM deflation and present cluster deflation policies for resource management of transient and on-demand cloud VMs. Experimental evaluation of our deflation system on a Linux cluster shows that microservice-based applications can be deflated by up to 50% with negligible performance overhead. Our cluster-level deflation policies allow overcommitment levels as high as 50%, with less than a 1% decrease in application throughput, and can enable cloud platforms to increase revenue by 30%. Alexander Fuerst, Ahmed Ali-Eldin, Prashant J. Shenoy, Prateek Sharma 0001 |
HPDC | 3 |
| 2020 | Hedge Your Bets: Optimizing Long-term Cloud Costs by Mixing VM Purchasing OptionsabstractCloud platforms offer the same VMs under many purchasing options that specify different costs and time commitments, such as on-demand, reserved, sustained-use, scheduled reserve, transient, and spot block. In general, the stronger the commitment, i.e., longer and less flexible, the lower the price. However, longer and less flexible time commitments can increase cloud costs for users if future workloads cannot utilize the VMs they committed to buying. Large cloud customers often find it challenging to choose the right mix of purchasing options to reduce their long-term costs, while retaining the ability to adjust capacity up and down in response to workload variations.To address the problem, we design policies to optimize long-term cloud costs by selecting a mix of VM purchasing options based on short- and long-term expectations of workload utilization. We consider a batch trace spanning 4 years from a large shared cluster for a major state University system that includes 14k cores and 60 million job submissions, and evaluate how these jobs could be judiciously executed using cloud servers using our approach. Our results show that our policies incur a cost within 41% of an optimistic optimal offline approach, and 50% less than solely using on-demand VMs. Lurdh Pradeep Reddy Ambati, Noman Bashir, David Irwin 0001, Mohammad Hajiesmaili, Prashant J. Shenoy |
IC2E | 5 |
| 2020 | Real-time Spatio-Temporal Action Localization in 360 VideosabstractSpatio-temporal action localization of human actions in a video has been a popular topic over the past few years. It tries to localize the bounding boxes, the time span and the class of one action, which summarizes information in the video and helps humans understand it. Though many approaches have been proposed to solve this problem, these efforts have only focused on perspective videos. Unfortunately, perspective videos only cover a small field-of-view (FOV), which limits the capability of action localization. In this paper, we develop a comprehensive approach to real-time spatio-temporal localization that can be used to detect actions in 360 videos. We create two datasets named UCF-101-24-360 and JHMDB-21-360 for our evaluation. Our experiments show that our method consistently outperforms other competing approaches and achieves a real-time processing speed of 15fps for 360 videos. Bo Chen 0025, Ahmed Ali-Eldin, Prashant J. Shenoy, Klara Nahrstedt |
ISM | 3 |
| 2020 | Waiting game: optimally provisioning fixed resources for cloud-enabled schedulersabstractWhile cloud platforms enable users to rent computing resources on demand to execute their jobs, buying fixed resources is still much cheaper than renting if their utilization is high. Thus, optimizing cloud costs requires users to determine how many fixed resources to buy versus rent based on their workload. In this paper, we introduce the concept of a waiting policy for cloud-enabled schedulers, which is the dual of a scheduling policy, and show that the optimal cost depends on it. We define multiple waiting policies and develop simple analytical models to reveal their tradeoff between fixed resource provisioning, cost, and job waiting time. We evaluate the impact of these waiting policies on a year-long production batch workload consisting of 14Mjobs run on a 14.3k-core cluster, and show that a compound waiting policy decreases the cost (by 5%) and mean job waiting time (by 7×) compared to a fixed cluster of the current size. Lurdh Pradeep Reddy Ambati, Noman Bashir, David Irwin 0001, Prashant J. Shenoy |
SC | 4 |
| 2020 | WiFiMon: a mobility analytics platform for building occupancy monitoring and contact tracing using wifi sensing: poster abstractabstractWith the current COVID-19 pandemic, contact tracing and building occupancy tracking are key components of re-opening policies and quickly containing virus outbreaks. WiFiMon is a network-centric contact tracing method that uses enterprise WiFi networks logs for tracking devices and inferring building occupancy and building contact tracing reports in office and campus settings. Emmanuel Cecchet, Amrita Acharya, Tergel Molom-Ochir, Amee Trivedi, Prashant J. Shenoy |
SenSys | 5 |
| 2020 | New Frontiers in IoT: Networking, Systems, Reliability, and Security ChallengesabstractThe field of IoT has blossomed and is positively influencing many application domains. In this article, we bring out the unique challenges this field poses to research in computer systems and networking. The unique challenges arise from the unique characteristics of IoT systems such as the diversity of application domains where they are used and the increasingly demanding protocols they are being called upon to run (such as video and LIDAR processing) on constrained resources (on-node and network). We show how these open challenges can benefit from foundations laid in other areas, such as fifth-generation network cellular protocols, machine learning model reduction, and device-edge-cloud offloading. We then discuss the unique challenges for reliability, security, and privacy posed by IoT systems due to their salient characteristics which include heterogeneity of devices and protocols, dependence on the physical environment, and the close coupling with humans. We again show how open research challenges benefit from the reliability, security, and privacy advancements in other areas. We conclude by providing a vision for a desirable end state for IoT systems. Saurabh Bagchi, Tarek F. Abdelzaher, Ramesh Govindan, Prashant J. Shenoy, Akanksha Atrey, Pradipta Ghosh, Ran Xu 0003 |
IEEE Internet Things J. | 4 |
| 2020 | Efficient Online Classification and Tracking on Resource-constrained IoT DevicesabstractTimely processing has been increasingly required on smart IoT devices, which leads to directly implementing information processing tasks on an IoT device for bandwidth savings and privacy assurance. Particularly, monitoring and tracking the observed signals in continuous form are common tasks for a variety of near real-time processing IoT devices, such as in smart homes, body-area, and environmental sensing applications. However, these systems are likely low-cost resource-constrained embedded systems, equipped with compact memory space, whereby the ability to store the full information state of continuous signals is limited. Hence, in this article,* we develop solutions of efficient timely processing embedded systems for online classification and tracking of continuous signals with compact memory space. Particularly, we focus on the application of smart plugs that are capable of timely classification of appliance types and tracking of appliance behavior in a standalone manner. We implemented a smart plug prototype using low-cost Arduino platform with small amount of memory space to demonstrate the following timely processing operations: (1) learning and classifying the patterns associated with the continuous power consumption signals and (2) tracking the occurrences of signal patterns using small local memory space. Furthermore, our system designs are also sufficiently generic for timely monitoring and tracking applications in other resource-constrained IoT devices. Muhammad Aftab, Sid Chi-Kin Chau, Prashant J. Shenoy |
ACM Trans. Internet Things | 3 |
| 2019 | Resource Deflation: A New Approach For Transient Resource ReclamationabstractData centers and clouds are increasingly offering low-cost computational resources in the form of transient virtual machines. Whenever demand for computational resources exceeds their availability, transient resources can reclaimed by preempting the transient VMs. Conventionally, these transient VMs are used by low-priority applications that can tolerate the disruption caused by preemptions. Prateek Sharma 0001, Ahmed Ali-Eldin, Prashant J. Shenoy |
EuroSys | 3 |
| 2019 | SpotWeb: Running Latency-sensitive Distributed Web Services on Transient Cloud ServersabstractMany cloud providers offer servers with transient availability at a reduced cost. These servers can be unilaterally revoked by the provider, usually after a warning period to the user. Until recently, it has been thought that these servers are not suitable to run latency-sensitive workloads due to their transient availability. In this paper, we introduce SpotWeb, a framework for running latency-sensitive web workloads on transient computing platforms while maintaining the Quality-of-Service (QoS) of the running applications. SpotWeb is based on three novel concepts; using multi-period optimization---a novel approach developed in finance---for server selection; transiency-aware load-balancing; and using intelligent capacity over-provisioning. We implement SpotWeb and evaluate its performance in both simulations and testbed experiments. Our results show that SpotWeb reduces costs by up to 50% compared to state-of-the-art solutions while being scalable to hundreds of cloud server configurations. Ahmed Ali-Eldin, Jonathan Westin, Prateek Sharma 0001, Prashant J. Shenoy |
HPDC | 5 |
| 2019 | Understanding Synchronization Costs for Distributed ML on Transient Cloud ResourcesabstractCloud platforms often execute parallel batch applications, such as distributed machine learning (ML), that include numerous synchronization barriers. These barriers, which prevent any task from advancing beyond a specified point until all tasks have reached that point, significantly degrade application performance by reducing it to that of the slowest "straggler" task. To address the problem, researchers have proposed numerous straggler mitigation techniques, including speculatively re-executing straggler tasks and various relaxations of strict barrier semantics. While these techniques improve parallel application performance, they incur a cost in terms of the resources wasted re-executing tasks or waiting. Importantly, these costs, which are often implicit in prior work that targets dedicated resources, become explicit in the cloud, which charges for resources at fine-grained intervals. In addition, the cost difference between techniques is exacerbated in cloud platforms, since they charge substantially less for transient resources that effectively yield a probabilistic performance across a wide range. While transient resources' low list price is attractive, revocations increase the frequency and severity of stragglers, which decreases parallel job performance and increases overall execution cost. To better understand the cost of synchronization, we develop simple analytical models of different straggler mitigation techniques and compare their cost and performance on on-demand and transient resources. Our analysis shows that i) transient servers offer complex tradeoffs compared to on-demand servers, and can result in higher overall costs despite their highly discounted price due to their probabilistic performance; ii) common approaches to straggler mitigation, which is a well-studied problem, are less effective using transient servers that cause frequent and severe stragglers; and iii) a recent approach to flexible synchronization offers the best cost and performance. Lurdh Pradeep Reddy Ambati, David Irwin 0001, Prashant J. Shenoy, Lixin Gao 0001, Ahmed Ali-Eldin, Jeannie R. Albrecht |
IC2E | 3 |
| 2019 | The Price Is (Not) Right: Reflections on Pricing for Transient Cloud ServersabstractAmazon introduced spot instances in December 2009, enabling "customers to bid on unused Amazon EC2 capacity and run those instances for as long as their bid exceeds the current Spot Price.'' Amazon's real-time computational spot market was novel in multiple respects. For example, it was the first (and to date only) large-scale public implementation of market-based resource allocation based on dynamic pricing after decades of research, and it provided users with useful information, control knobs, and options for optimizing the cost of running cloud applications. Spot instances also introduced the concept of transient cloud servers derived from variable idle capacity that cloud platforms could revoke at any time. Transient servers have since become central to efficient resource management of modern clusters and clouds. As a result, Amazon's spot market was the motivation for substantial research over the past decade. Yet, in November 2017, Amazon effectively ended its realtime spot market by announcing that users no longer needed to place bids and that spot prices will "...adjust more gradually, based on longer-term trends in supply and demand.'' The changes made spot instances more similar to the fixed-price transient servers offered by other cloud platforms. Unfortunately, while these changes made spot instances less complex, they eliminated many benefits to sophisticated users in optimizing their applications. This paper provides a retrospective on Amazon's real-time spot market, including its advantages and disadvantages for allocating transient servers compared to current fixed-price approaches. We also discuss some fundamental problems with Amazon's spot market, which we identified in prior work (from 2016), that predicted its eventual end. We then discuss potential options for allocating transient servers that combine the advantages of Amazon's real-time spot market, while also addressing the problems that likely led to its elimination. David Irwin 0001, Prashant J. Shenoy, Lurdh Pradeep Reddy Ambati, Prateek Sharma 0001, Supreeth Shastri, Ahmed Ali-Eldin |
ICCCN | 2 |
| 2019 | DeepRoof: A Data-driven Approach For Solar Potential Estimation Using Rooftop ImageryabstractRooftop solar deployments are an excellent source for generating clean energy. As a result, their popularity among homeowners has grown significantly over the years. Unfortunately, estimating the solar potential of a roof requires homeowners to consult solar consultants, who manually evaluate the site. Recently there have been efforts to automatically estimate the solar potential for any roof within a city. However, current methods work only for places where LIDAR data is available, thereby limiting their reach to just a few places in the world. In this paper, we propose DeepRoof, a data-driven approach that uses widely available satellite images to assess the solar potential of a roof. Using satellite images, DeepRoof determines the roof's geometry and leverages publicly available real-estate and solar irradiance data to provide a pixel-level estimate of the solar potential for each planar roof segment. Such estimates can be used to identify ideal locations on the roof for installing solar panels. Further, we evaluate our approach on an annotated roof dataset, validate the results with solar experts and compare it to a LIDAR-based approach. Our results show that DeepRoof can accurately extract the roof geometry such as the planar roof segments and their orientation, achieving a true positive rate of 91.1% in identifying roofs and a low mean orientation error of 9.3 degree. We also show that DeepRoof's median estimate of the available solar installation area is within 11% of a LIDAR-based approach. Stephen Lee, Srinivasan Iyengar, Menghong Feng, Prashant J. Shenoy, Subhransu Maji |
KDD | 4 |
| 2019 | Solar-TK: A Data-Driven Toolkit for Solar PV Performance Modeling and ForecastingabstractSolar energy capacity is continuing to increase. The key challenge with integrating solar into buildings and the electric grid is its high power generation variability, which is a function of many factors, including a site's location, time, weather, and numerous physical attributes. There has been significant prior work on solar performance modeling and forecasting that infers a site's current and future solar generation based on these factors. Accurate solar performance models and forecasts are also a pre-requisite for conducting a wide range of building and grid energy-efficiency research. Unfortunately, much of the prior work is not accessible to researchers, either because it has not been released as open source, is time-consuming to re-implement, or requires access to proprietary data sources. To address the problem, we present Solar-TK, a data-driven toolkit for solar performance modeling and forecasting that is simple, extensible, and publicly accessible. Solar-TK's simple approach models and forecasts a site's solar output given only its location and a small amount of historical generation data. Solar-TK's extensible design includes a small collection of independent modules that connect together to implement basic modeling and forecasting, while also enabling users to implement new energy analytics. We plan to release Solar-TK as open source to enable research that requires realistic solar models and forecasts, and to serve as a baseline for comparing new solar modeling and forecasting techniques. We compare Solar-TK's simple approach with PVlib and show that it yields comparable accuracy. We present three case studies showing how Solar-TK can advance energy-efficiency research. Noman Bashir, Dong Chen 0010, David Irwin 0001, Prashant J. Shenoy |
MASS | 4 |
| 2019 | Performance Evaluation of Multi-Path TCP for Data Center and Cloud WorkloadsabstractToday's cloud data centers host a wide range of applications including data analytics, batch processing, and interactive processing. These applications require high throughput, low latency, and high reliability from the network. Satisfying these requirements in the face of dynamically varying network conditions remains a challenging problem. Multi-Path TCP (MPTCP) is a recently proposed IETF extension to TCP that divides a conventional TCP flow into multiple subflows so as to utilize multiple paths over the network. Despite the theoretical and practical benefits of MPTCP, its effectiveness for cloud applications and environments remains unclear as there has been little work to quantify the benefits of MPTCP for real cloud applications. We present a broad empirical study of the effectiveness and feasibility of MPTCP for data center and cloud applications, under different network conditions. Our results show that while MPTCP provides useful bandwidth aggregation, congestion avoidance, and improved resiliency for some cloud applications, these benefits do not apply uniformly across applications, especially in cloud settings. Lucas Chaufournier, Ahmed Ali-Eldin, Prateek Sharma 0001, Prashant J. Shenoy, Don Towsley |
ICPE | 4 |
| 2019 | UDAO: A Next-Generation Unified Data Analytics OptimizerabstractBig data analytics systems today still lack the ability to take user performance goals and budgetary constraints, collectively referred to as "objectives", and automatically configure an analytic job to achieve the objectives. This paper presents UDAO, a unified data analytics optimizer that can automatically determine the parameters of the runtime system, collectively called a job configuration, for general dataflow programs based on user objectives. UDAO embodies key techniques including in-situ modeling , which learns a model for each user objective in the same computing environment as the job is run, and multi-objective optimization , which computes a Pareto optimal set of job configurations to reveal tradeoffs between different objectives. Using benchmarks developed based on industry needs, our demonstration will allow the user to explore (1) learned models to gain insights into how various parameters affect user objectives; (2) Pareto frontiers to understand interesting tradeoffs between different objectives and how a configuration recommended by the optimizer explores these tradeoffs; (3) end-to-end benefits that UDAO can provide over default configurations or those manually tuned by engineers. Khaled Zaouk, Chenghao Lyu, Yanlei Diao, Prashant J. Shenoy |
Proc. VLDB Endow. | 6 |
| 2019 | Building Virtual Power Meters for Online Load TrackingabstractMany energy optimizations require fine-grained, load-level energy data collected in real time, most typically by a plug-level energy meter. Online load tracking is the problem of monitoring an individual electrical load’s energy usage in software by analyzing the building’s aggregate smart meter data. Load tracking differs from the well-studied problem of load disaggregation in that it emphasizes per-load accuracy and efficient, online operation rather than accurate disaggregation of every building load via offline analysis. In essence, tracking a particular load creates a virtual power meter for it, which mimics having a networked-connected power meter attached to the load, but notably does not require tracking every other load as well. We propose PowerPlay , a model-driven system for performing accurate, high-performance online load tracking. Our results from applying the system to real-world energy data demonstrate that PowerPlay (i) enables efficient online tracking on low-power embedded platforms, (ii) scales to thousands of loads (across many buildings) on server platforms, and (iii) improves per-load accuracy by more than a factor of two compared to a state-of-the-art load disaggregation algorithm. Our results point to the potential of replacing physical energy meters by “virtual” power meters using a system like PowerPlay. Sean Kenneth Barker, Sandeep Kalra, David Irwin 0001, Prashant J. Shenoy |
ACM Trans. Cyber Phys. Syst. | 4 |
| 2019 | Inferring Smart Schedules for Dumb ThermostatsabstractHeating, ventilation, and air conditioning (HVAC) accounts for over 50% of a typical home’s energy usage. A thermostat generally controls HVAC usage in a home to ensure user comfort. In this article, we focus on making existing “dumb” programmable thermostats smart by applying energy analytics on smart meter data to infer home occupancy patterns and compute an optimized thermostat schedule. Utilities with smart meter deployments are capable of immediately applying our approach, called iProgram, to homes across their customer base. iProgram addresses new challenges in inferring home occupancy from smart meter data where (i) training data is not available and (ii) the thermostat schedule may be misaligned with occupancy, frequently resulting in high power usage during unoccupied periods. iProgram translates occupancy patterns inferred from opaque smart meter data into a custom schedule for existing types of programmable thermostats, e.g., 1-day, 7-day, and so on. We implement iProgram as a web service and show that it reduces the mismatch time between the occupancy pattern and the thermostat schedule by a median value of 44.28min (out of 100 homes) when compared to a default 8am-6pm weekday schedule, with a median deviation of 30.76min off the optimal schedule. Further, iProgram yields a daily energy savings of 0.42kWh on average across the 100 homes. Moreover, the schedules generated from iProgram converge to optimal schedules within a couple of weeks for most homes. We also show that homeowners having multiple HVAC zones can utilize iProgram and potentially increase unconditioned times of less occupied parts of their homes by 70%. Utilities may use iProgram to recommend thermostat schedules to customers and provide them estimates of potential energy savings in their energy bills. Srinivasan Iyengar, Sandeep Kalra, Anushree Ghosh, David Irwin 0001, Prashant J. Shenoy, Benjamin M. Marlin |
ACM Trans. Cyber Phys. Syst. | 5 |
| 2018 | SolarClique: Detecting Anomalies in Residential Solar ArraysabstractThe proliferation of solar deployments has significantly increased over the years. Analyzing these deployments can lead to the timely detection of anomalies in power generation, which can maximize the benefits from solar energy. In this paper, we propose SolarClique, a data-driven approach that can flag anomalies in power generation with high accuracy. Unlike prior approaches, our work neither depends on expensive instrumentation nor does it require external inputs such as weather data. Rather our approach exploits correlations in solar power generation from geographically nearby sites to predict the expected output of a site and flag anomalies. We evaluate our approach on 88 solar installations located in Austin, Texas. We show that our algorithm can even work with data from few geographically nearby sites (>5 sites) to produce results with high accuracy. Thus, our approach can scale to sparsely populated regions, where there are few solar installations. Further, among the 88 installations, our approach reported 76 sites with anomalies in power generation. Moreover, our approach is robust enough to distinguish between reduction in power output due to anomalies and other factors such as cloudy conditions. Srinivasan Iyengar, Stephen Lee, Daniel Sheldon, Prashant J. Shenoy |
COMPASS | 4 |
| 2018 | Will Distributed Computing Revolutionize Peace? The Emergence of Battlefield IoTabstractAn upcoming frontier for distributed computing might literally save lives in future military operations. In civilian scenarios, significant efficiencies were gained from interconnecting devices into networked services and applications that automate much of everyday life from smart homes to intelligent transportation. The ecosystem of such applications and services is collectively called the Internet of Things (IoT). Can similar benefits be gained in a military context by developing an IoT for the battlefield? This paper describes unique challenges in such a context as well as potential risks, mitigation strategies, and benefits. Tarek F. Abdelzaher, Nora Ayanian, Tamer Basar, Suhas N. Diggavi, Jana Diesner, Deepak Ganesan, Ramesh Govindan, Susmit Jha, Tancrède Lepoint, Benjamin M. Marlin, Klara Nahrstedt, David M. Nicol, Ragunathan Rajkumar, Stephen Russell 0001, Sanjit A. Seshia, Fei Sha, Prashant J. Shenoy, Mani Srivastava 0001, Gaurav S. Sukhatme, Ananthram Swami, Paulo Tabuada, Don Towsley, Nitin H. Vaidya, Venugopal V. Veeravalli |
ICDCS | 17 |
| 2018 | Private Memoirs of IoT Devices: Safeguarding User Privacy in the IoT EraabstractThe rise of the Internet-of-Things (IoT) holds great promise to transform people's lives by making society more efficient in many areas, including energy, transportation, healthcare, commerce, manufacturing, etc. At their core, IoT devices use sensors to collect data on real-world physical processes and then transmit it over the Internet to cloud servers, which store, process, and learn from the data to better optimize these processes, either directly (by issuing remote commands that actuate IoT devices) or indirectly (by issuing notifications that direct users to take some action). Unfortunately, IoT devices also expose users to multiple new types of privacy attacks. In particular, the sensor data collected from IoT devices can indirectly reveal a variety of sensitive private information. In addition, users generally connect IoT devices to local networks, which they implicitly trust, with little understanding of what the IoT device is doing on the network. In this visionpaper, we discuss recent work on sensor data privacy in the context of smart energy systems to provide examples of i) the surprising types of private information we can glean from seemingly innocuous IoT data and ii) the different types of defenses we have developed to preserve IoT data privacy for smart energy systems. These defenses lie at different discrete points in the tradeoff between user privacy and IoT functionality, which motivates ongoing work on developing defenses that provide a more tunable tradeoff. We also discuss the privacy implications of connecting tens-to-hundreds of untrusted IoT devices to implicitly trusted local networks, and avenues for research to mitigate these concerns. Dong Chen 0010, Phuthipong Bovornkeeratiroj, David Irwin 0001, Prashant J. Shenoy |
ICDCS | 4 |
| 2018 | WattHome: A Data-driven Approach for Energy Efficiency Analytics at City-scaleabstractBuildings consume over 40% of the total energy in modern societies and improving their energy efficiency can significantly reduce our energy footprint. In this paper, we present WattHome, a data-driven approach to identify the least energy efficient buildings from a large population of buildings in a city or a region. Unlike previous approaches such as least squares that use point estimates, WattHome uses Bayesian inference to capture the stochasticity in the daily energy usage by estimating the parameter distribution of a building. Further, it compares them with similar homes in a given population using widely available datasets. WattHome also incorporates a fault detection algorithm to identify the underlying causes of energy inefficiency. We validate our approach using ground truth data from different geographical locations, which showcases its applicability in different settings. Moreover, we present results from a case study from a city containing >10,000 buildings and show that more than half of the buildings are inefficient in one way or another indicating a significant potential from energy improvement measures. Additionally, we provide probable cause of inefficiency and find that 41%, 23.73%, and 0.51% homes have poor building envelope, heating, and cooling system faults respectively. Srinivasan Iyengar, Stephen Lee, David Irwin 0001, Prashant J. Shenoy, Benjamin Weil |
KDD | 4 |
| 2018 | Latency-aware virtual desktops optimization in distributed clouds
Tian Guo 0001, Prashant J. Shenoy, K. K. Ramakrishnan, Vijay Gopalakrishnan |
Multim. Syst. | 2 |
| 2018 | Performance and Cost Considerations for Providing Geo-Elasticity in Database CloudsabstractOnline applications that serve global workload have become a norm and those applications are experiencing not only temporal but also spatial workload variations. In addition, more applications are hosting their backend tiers separately for benefits such as ease of management. To provision for such applications, traditional elasticity approaches that only consider temporal workload dynamics and assume well-provisioned backends are insufficient. Instead, in this article, we propose a new type of provisioning mechanisms—geo-elasticity, by utilizing distributed clouds with different locations. Centered on this idea, we build a system called DBScale that tracks geographic variations in the workload to dynamically provision database replicas at different cloud locations across the globe. Our geo-elastic provisioning approach comprises a regression-based model that infers database query workload from spatially distributed front-end workload, a two-node open queueing network model that estimates the capacity of databases serving both CPU and I/O-intensive query workloads and greedy algorithms for selecting best cloud locations based on latency and cost. We implement a prototype of our DBScale system on Amazon EC2’s distributed cloud. Our experiments with our prototype show up to a 66% improvement in response time when compared to local elasticity approaches. Tian Guo 0001, Prashant J. Shenoy |
ACM Trans. Auton. Adapt. Syst. | 2 |
| 2018 | Providing Geo-Elasticity in Geographically Distributed CloudsabstractGeographically distributed cloud platforms are well suited for serving a geographically diverse user base. However, traditional cloud provisioning mechanisms that make local scaling decisions are not adequate for delivering the best possible performance for modern web applications that observe both temporal and spatial workload fluctuations. We propose GeoScale, a system that provides geo-elasticity by combining model-driven proactive and agile reactive provisioning approaches. GeoScale can dynamically provision server capacity at any location based on workload dynamics. We conduct a detailed evaluation of GeoScale on Amazon’s geo-distributed cloud and show up to 40% improvement in the 95th percentile response time when compared to traditional elasticity techniques. Tian Guo 0001, Prashant J. Shenoy |
ACM Trans. Internet Techn. | 2 |
| 2018 | Mechanisms and Policies for Controlling Distributed Solar CapacityabstractThe rapid expansion of intermittent grid-tied solar capacity is making the job of balancing electricity’s real-time supply and demand increasingly challenging. Recent work proposes mechanisms for actively controlling solar power in the grid at individual sites by enabling software to cap it as a fraction of its time-varying maximum output. However, while enforcing an equal fraction of each solar site’s time-varying maximum output results in “fair” short-term contributions of solar power across all sites, it does not result in “fair” long-term contributions of solar energy. Enforcing fair long-term energy access is important when controlling distributed solar capacity, since limits on solar output impact the compensation users receive for net metering and the battery capacity required to store excess solar energy. This discrepancy arises from fundamental differences in enforcing “fair” access to the grid to contribute solar energy, compared to analogous fair sharing in networks and processors. To address the problem, we first present both a centralized and distributed algorithm to enable control of distributed solar capacity that enforces fair grid energy access. We then present multiple policies that show how utilities can leverage this new distributed rate-limiting mechanism to reduce variations in grid demand from intermittent solar generation. Noman Bashir, David Irwin 0001, Prashant J. Shenoy, Jay Taneja |
ACM Trans. Sens. Networks | 3 |
| 2018 | Managing Risk in a Derivative IaaS CloudabstractInfrastructure-as-a-Service (IaaS) cloud platforms rent computing resources with different cost and availability tradeoffs. For example, users may acquire virtual machines (VMs) in the spot market-that are cheap, but can be unilaterally terminated by the cloud operator. Because of this revocation risk, spot servers have been conventionally used for delay and risk tolerant batch jobs. In this paper, we develop risk mitigation policies which allow even interactive applications to run on spot servers. Our System, SpotCheck is a derivative cloud platform, and provides the illusion of an IaaS platform that offers always-available VMs on demand for a cost near that of spot servers, and supports unmodified applications. SpotCheck's design combines virtualization-based mechanisms for fault-tolerance, and bidding and server selection policies for managing the risk and cost. We implement SpotCheck on EC2 and show that it i) provides nested VMs with 99.9989 percent availability, ii) achieves upto 2-5x cost savings compared to using on-demand VMs, and iii) eliminates any risk of losing VM state. Prateek Sharma 0001, Stephen Lee, Tian Guo 0001, David Irwin 0001, Prashant J. Shenoy |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2017 | The Financialization of Cloud Computing: Opportunities and ChallengesabstractUnder competitive pressure to maximize their infrastructure's utilization and revenue, modern cloud platforms are quickly evolving into server markets that offer increasingly sophisticated contracts beyond simple on-demand servers, such as spot, preemptible, burstable, and reserved servers. In parallel, continuing advances in system and network virtualization are making server-time a more fungible commodity. These trends have motivated calls for open cloud commodity markets akin to other commodity markets, e.g., for oil, gold, corn, etc. However, such open cloud markets have not yet materialized due to key differences between cloud resources and other commodities. In particular, the relationship between applications and their underlying server resources is fundamentally different and more complex than other commodities. Unfortunately, software developers generally do not have the necessary background to effectively manage this complexity as part of their applications. Financial cloud computing is an emerging area that focuses on adapting and extending concepts from economics and finance to explicitly manage applications' tradeoffs between cost, risk, availability, and performance in cloud markets. A key goal of financial cloud computing is to develop systems-level abstractions and mechanisms that manage the market's complexity. This paper introduces this emerging area and its potential benefits, surveys related work, discusses challenges to realizing a cloud commodity market, and then outlines future research directions. David Irwin 0001, Prateek Sharma 0001, Supreeth Shastri, Prashant J. Shenoy |
ICCCN | 4 |
| 2017 | Pervasive Energy Monitoring and Control Through Low-Bandwidth Power Line CommunicationabstractThe Internet of Things (IoT) is growing rapidly, with increasingly sophisticated networking, sensing, and actuation functions embedded into everyday devices. One important IoT application is managing a building's energy usage by monitoring and controlling its electrical devices. Many existing IoT-enabled devices operate through low-cost and convenient power line networks, using protocols such as X10 and Insteon for communication. However, as these technologies have traditionally targeted low-bandwidth device control, they are often not readily suited to higher bandwidth uses such as continuous energy monitoring. In this paper, we consider the challenge of leveraging existing low-bandwidth power line communication networks for energy monitoring, and present several techniques that enable reliable and high-resolution monitoring in such networks. As a case study, we consider the popular Insteon protocol and show that intelligent polling and event detection methods can reduce the bandwidth requirements and undetected power events in a realworld Insteon network by 50% or more versus naive methods. Our techniques have been employed in a real IoT-enabled smart home, which has collected much of the data publicly released in the UMass Smart* energy dataset. Sean Kenneth Barker, David Irwin 0001, Prashant J. Shenoy |
IEEE Internet Things J. | 3 |
| 2017 | Minimizing Transmission Loss in Smart Microgrids by Sharing Renewable EnergyabstractRenewable energy (e.g., solar energy) is an attractive option to provide green energy to homes. Unfortunately, the intermittent nature of renewable energy results in a mismatch between when these sources generate energy and when homes demand it. This mismatch reduces the efficiency of using harvested energy by either (i) requiring batteries to store surplus energy, which typically incurs ∼ 20% energy conversion losses, or (ii) using net metering to transmit surplus energy via the electric grid’s AC lines, which severely limits the maximum percentage of renewable penetration possible. In this article, we propose an alternative structure where nearby homes explicitly share energy with each other to balance local energy harvesting and demand in microgrids. We develop a novel energy sharing approach to determine which homes should share energy, and when to minimize system-wide energy transmission losses in the microgrid. We evaluate our approach in simulation using real traces of solar energy harvesting and home consumption data from a deployment in Amherst, MA. We show that our system (i) reduces the energy loss on the AC line by 64% without requiring large batteries, (ii) performance scales up with larger battery capacities, and (iii) is robust to different energy consumption patterns and energy prediction accuracy in the microgrid. Zhichuan Huang, Ting Zhu 0001, David Irwin 0001, Aditya Kumar Mishra, Daniel Sadoc Menasché, Prashant J. Shenoy |
ACM Trans. Cyber Phys. Syst. | 6 |
| 2017 | Enabling Distributed Energy Storage by Incentivizing Small Load ShiftsabstractReducing peak demands and achieving a high penetration of renewable energy sources are important goals in achieving a smarter grid. To reduce peak demand, utilities are introducing variable rate electricity prices to incentivize consumers to manually shift their demand to low-price periods. Consumers may also use energy storage to automatically shift their demand by storing energy during low-price periods for use during high-price periods. Unfortunately, variable rate pricing provides only a weak incentive for distributed energy storage and does not promote its adoption at large scales. In this article, we present the storage adoption dilemma to capture the problems with incentivizing energy storage using variable rate prices. To address the problem, we propose a simple pricing scheme, called flat-power pricing , which incentivizes consumers to shift small amounts of load to flatten their demand rather than shift as much of their power usage as possible to low-price, off-peak periods. We show that compared to variable rate pricing, flat-power pricing (i) reduces consumers’ upfront capital costs, as it requires significantly less storage capacity per consumer; (ii) increases energy storage’s return on investment, as it mitigates free riding and maintains the incentive to use energy storage at large scales; and (iii) uses aggregate storage capacity within 31% of an optimal centralized approach. In addition, unlike variable rate pricing, we also show that flat-power pricing incentivizes the scheduling of elastic background loads, such as air conditioners and heaters, to reduce peak demand. We evaluate our approach using real smart meter data from 14,000 homes in a small town. David Irwin 0001, Srinivasan Iyengar, Stephen Lee, Aditya Kumar Mishra, Prashant J. Shenoy |
ACM Trans. Cyber Phys. Syst. | 5 |
| 2017 | A Cloud-Based Black-Box Solar Predictor for Smart HomesabstractThe popularity of rooftop solar for homes is rapidly growing. However, accurately forecasting solar generation is critical to fully exploiting the benefits of locally generated solar energy. In this article, we present two machine-learning techniques to predict solar power from publicly available weather forecasts. We use these techniques to develop SolarCast, a cloud-based web service that automatically generates models that provide customized site-specific predictions of solar generation. SolarCast utilizes a “black box” approach that requires only (1) a site’s geographic location and (2) a minimal amount of historical generation data. Since we intend SolarCast for small rooftop deployments, it does not require detailed site- and panel-specific information, which owners may not know, but instead automatically learns these parameters for each site. We evaluate the accuracy of SolarCast’s different algorithms on two publicly available datasets, each containing over 100 rooftop deployments with a variety of attributes (e.g., climate, tilt, orientation, etc.). We show that SolarCast learns a more accurate model using much less data (∼1 month) than prior SVM-based approaches, which require ∼3 months of data. SolarCast also provides a programmatic API, enabling developers to integrate its predictions into energy efficiency applications. Finally, we present two case studies of using SolarCast to demonstrate how real-world applications can leverage its predictions. We first evaluate a “sunny” load scheduler, which schedules a dryer’s energy usage to maximally align with a home’s solar generation. We then evaluate a smart solar-powered charging station, which can optimally charge the maximum number of electric vehicles (EVs) on a given day. Our results indicate that a representative home is capable of reducing its grid demand up to 40% by providing a modest amount of flexibility (of ∼5 hours) in the dryer’s start time with opportunistic load scheduling. Further, our charging station uses SolarCast to provide EV owners the amount of energy they can expect to receive from solar energy sources. Srinivasan Iyengar, Navin Sharma, David Irwin 0001, Prashant J. Shenoy, Krithi Ramamritham |
ACM Trans. Cyber Phys. Syst. | 4 |
| 2016 | Flint: batch-interactive data-intensive processing on transient serversabstractCloud providers now offer transient servers, which they may revoke at anytime, for significantly lower prices than on-demand servers, which they cannot revoke. The low price of transient servers is particularly attractive for executing an emerging class of workload, which we call Batch-Interactive Data-Intensive (BIDI), that is becoming increasingly important for data analytics. BIDI workloads require large sets of servers to cache massive datasets in memory to enable low latency operation. In this paper, we illustrate the challenges of executing BIDI workloads on transient servers, where revocations (akin to failures) are the common case. To address these challenges, we design Flint, which is based on Spark and includes automated checkpointing and server selection policies that i) support batch and interactive applications and ii) dynamically adapt to application characteristics. We evaluate a prototype of Flint using EC2 spot instances, and show that it yields cost savings of up to 90% compared to using on-demand servers, while increasing running time by < 2%. Prateek Sharma 0001, Tian Guo 0001, David Irwin 0001, Prashant J. Shenoy |
EuroSys | 5 |
| 2016 | GeoScale: Providing Geo-Elasticity in Distributed CloudsabstractDistributed cloud platforms are well suited for serving a geographically diverse user base. However traditional cloud provisioning mechanisms that make local scaling decisions are not well suited for temporal and spatial workload fluctuations seen by modern web applications. In this paper, we argue the need of geo-elasticity and present GeoScale, a system to provide geo-elasticity in distributed clouds. We describe GeoScale's model-driven proactive provisioning approach and conduct an initial evaluation of GeoScale on Amazon's distributed EC2 cloud. Our results show up to 31% improvement in the 95th percentile response time when compared to traditional elasticity techniques. Tian Guo 0001, Prashant J. Shenoy, Hakan Hacigümüs |
IC2E | 2 |
| 2016 | Firebird: Network-Aware Task Scheduling for Spark Using SDNsabstractRecently Spark has become a popular cluster computing platform because of its fast in-memory computing which allows users to cache data in servers' memory and query it repeatedly. However, since network I/O is much slower than local I/O, the network can be a bottleneck for data intensive jobs. When data locality is not well balanced or network capacity is limited, data contention can occur, which will significantly slow down task execution. The current delay scheduling method in Spark cannot perfectly solve this problem because it is agnostic to the network status in a cluster. In this paper, we propose a network-aware scheduling method in Spark and design Firebird, a derivative of Spark that runs on top of software-defined network (SDN). By using SDNs, the communication barriers between cluster computing platform and underlying networking are removed and tasks can be scheduled based on network conditions in Spark clusters. We demonstrate the effectiveness of the methods through detailed experiments with different types of jobs on our system. Experimental results show significant improvement of data-intensive jobs, and our system can achieve best scheduling in different cases without tuning Spark. Firebird can be up to 9 times faster than default Spark in some cases. Prashant J. Shenoy |
ICCCN | 2 |
| 2016 | SpotLight: An Information Service for the CloudabstractInfrastructure-as-a-Service cloud platforms are incredibly complex: they rent hundreds of different types of servers across multiple geographical regions under a wide range of contract types that offer varying tradeoffs between risk and cost. Unfortunately, the internal dynamics of cloud platforms are opaque along several dimensions. For example, while the risk of servers not being available when requested is critical in optimizing the cloud's risk-cost tradeoffs, it is not typically made visible to users. Thus, inspired by prior work on Internet bandwidth probing, we propose actively probing cloud platforms to explicitly learn such information, where each "probe" is a request for a particular type of server. We model the relationships between different contracts types to develop a market-based probing policy, which leverages the insight that real-time prices in cloud spot markets loosely correlate with the supply (and availability) of fixed-price on-demand servers. That is, the higher the spot price for a server, the more likely the corresponding fixed-price on-demand server is not available. We incorporate market-based probing into SpotLight, an information service that enables cloud applications to query this and other data, and use it to monitor the availability of more than 4500 distinct server types across 9 geographical regions in Amazon's Elastic Compute Cloud over a 3 month period. We analyze this data to reveal interesting observations about the platform's internal dynamics. We then show how SpotLight enables two recently proposed derivative cloud services to select a better mix of servers to host applications, which improves their availability from ~70-90% to near 100% in practice. David Irwin 0001, Prashant J. Shenoy |
ICDCS | 3 |
| 2016 | Containers and Virtual Machines at Scale: A Comparative Study
Prateek Sharma 0001, Lucas Chaufournier, Prashant J. Shenoy, Y. C. Tay |
Middleware | 3 |
| 2016 | Analyzing the Efficiency of a Green University Data CenterabstractData centers are an indispensable part of today's IT infrastructure. To keep pace with modern computing needs, data centers continue to grow in scale and consume increasing amounts of power. While prior work on data centers has led to significant improvements in their energy-efficiency, detailed measurements from these facilities' operations are not widely available, as data center design is often considered part of a company's competitive advantage. However, such detailed measurements are critical to the research community in motivating and evaluating new energy-efficiency optimizations. In this paper, we present a detailed analysis of a state-of-the-art 15MW green multi-tenant data center that incorporates many of the technological advances used in commercial data centers. We analyze the data center's computing load and its impact on power, water, and carbon usage using standard effectiveness metrics, including PUE, WUE, and CUE. Our results reveal the benefits of optimizations, such as free cooling, and provide insights into how the various effectiveness metrics change with the seasons and increasing capacity usage. More broadly, our PUE, WUE, and CUE analysis validate the green design of this LEED Platinum data center. Patrick Pegus II, Benoy Varghese, Tian Guo 0001, David Irwin 0001, Prashant J. Shenoy, Anirban Mahanti, James Culbert, John Goodhue, Chris Hill |
ICPE | 5 |
| 2016 | Beyond Energy-Efficiency: Evaluating Green Datacenter Applications for Energy-AgilityabstractComputing researchers have long focused on improving energy-efficiency under the implicit assumption that all energy is created equal. Yet, this assumption is actually incorrect: energy's cost and carbon footprint vary substantially over time. As a result, consuming energy inefficiently when it is cheap and clean may sometimes be preferable to consuming it efficiently when it is expensive and dirty. Green datacenters adapt their energy usage to optimize for such variations, as reflected in changing electricity prices or renewable energy output. Thus, we introduce energy-agility as a new metric to evaluate green datacenter applications. To illustrate fundamental tradeoffs in energy-agile design, we develop GreenSort, a distributed sorting system optimized for energy-agility. GreenSort is representative of the long-running, massively-parallel, data-intensive tasks that are common in datacenters and amenable to delays from power variations. Our results demonstrate the importance of energy-agile design when considering the benefits of using variable power. For example, we show that GreenSort requires 31% more time and energy to complete when power varies based on real-time electricity prices versus when it is constant. Thus, in this case, real-time prices should be at least 31% lower than fixed prices to warrant using them. Supreeth Subramanya, Zain Mustafa, David Irwin 0001, Prashant J. Shenoy |
ICPE | 4 |
| 2015 | SpotOn: a batch computing service for the spot marketabstractCloud spot markets enable users to bid for compute resources, such that the cloud platform may revoke them if the market price rises too high. Due to their increased risk, revocable resources in the spot market are often significantly cheaper (by as much as 10×) than the equivalent non-revocable on-demand resources. One way to mitigate spot market risk is to use various fault-tolerance mechanisms, such as checkpointing or replication, to limit the work lost on revocation. However, the additional performance overhead and cost for a particular fault-tolerance mechanism is a complex function of both an application's resource usage and the magnitude and volatility of spot market prices. Supreeth Subramanya, Tian Guo 0001, Prateek Sharma 0001, David Irwin 0001, Prashant J. Shenoy |
SoCC | 5 |
| 2015 | SpotCheck: designing a derivative IaaS cloud on the spot marketabstractInfrastructure-as-a-Service (IaaS) cloud platforms rent resources, in the form of virtual machines (VMs), under a variety of contract terms that offer different levels of risk and cost. For example, users may acquire VMs in the spot market that are often cheap but entail significant risk, since their price varies over time based on market supply and demand and they may terminate at any time if the price rises too high. Currently, users must manage all the risks associated with using spot servers. As a result, conventional wisdom holds that spot servers are only appropriate for delay-tolerant batch applications. In this paper, we propose a derivative cloud platform, called SpotCheck, that transparently manages the risks associated with using spot servers for users. Prateek Sharma 0001, Stephen Lee, Tian Guo 0001, David Irwin 0001, Prashant J. Shenoy |
EuroSys | 5 |
| 2015 | Cutting the Cost of Hosting Online Services Using Cloud Spot MarketsabstractThe use of cloud servers to host modern Internet-based services is becoming increasingly common. Today's cloud platforms offer a choice of server types, including non-revocable on-demand servers and cheaper but revocable spot servers. A service provider requiring servers can bid in the spot market where the price of a spot server changes dynamically according to the current supply and demand for cloud resources. Spot servers are usually cheap, but can be revoked by the cloud provider when the cloud resources are scarce. While it is well-known that spot servers can reduce the cost of performing time-flexible interruption-tolerant tasks, we explore the novel possibility of using spot servers for reducing the cost of hosting an Internet-based service such as an e-commerce site that must {\em always} be on and the penalty for service unavailability is high. Prashant J. Shenoy, Ramesh K. Sitaraman, David Irwin 0001 |
HPDC | 2 |
| 2015 | Video BenchLab demo: an open platform for video realistic streaming benchmarkingabstractIn this demonstration, we present an open, flexible and realistic benchmarking platform named Video BenchLab to measure the performance of streaming media workloads. While Video BenchLab can be used with any existing media server, we provide a set of tools for researchers to experiment with their own platform and protocols. The components include a MediaDrop video server, a suite of tools to bulk insert videos and generate streaming media workloads, a dataset of freely available video and a client runtime to replay videos in the native video players of real Web browsers such as Firefox, Chrome and Internet Explorer. Various metrics are collected to capture the quality of video playback and identify issues that can happen during video replay. Finally, we provide a Dashboard to manage experiments, collect results and perform analytics to compare performance between experiments. Patrick Pegus II, Emmanuel Cecchet, Prashant J. Shenoy |
MMSys | 3 |
| 2015 | Video BenchLab: an open platform for realistic benchmarking of streaming media workloadsabstractIn this paper, we present an open, flexible and realistic benchmarking platform named Video BenchLab to measure the performance of streaming media workloads. While Video BenchLab can be used with any existing media server, we provide a set of tools for researchers to experiment with their own platform and protocols. The components include a MediaDrop video server, a suite of tools to bulk insert videos and generate streaming media workloads, a dataset of freely available video and a client runtime to replay videos in the native video players of real Web browsers such as Firefox, Chrome and Internet Explorer. We define simple metrics that are able to capture the quality of video playback and identify issues that can happen during video replay. Finally, we provide a Dashboard to manage experiments, collect results and perform analytics to compare performance between experiments. Patrick Pegus II, Emmanuel Cecchet, Prashant J. Shenoy |
MMSys | 3 |
| 2015 | Towards Cooling Internet-Scale Distributed Networks on the CheapabstractInternet-scale Distributed Networks (IDNs) are large distributed systems that comprise hundreds of thousands of servers located around the world. IDNs consume significant amounts of energy to power their deployed server infrastructure, and nearly as much energy to cool that infrastructure. We study the potential benefits of using renewable open air cooling (OAC) in an IDN. Our results show that by using OAC, a global IDN can extract 51% cooling energy reducing during summers and a 92% reduction in the winter. Vani Gupta, Stephen Lee, Prashant J. Shenoy, Ramesh K. Sitaraman, Rahul Urgaonkar |
SIGMETRICS | 3 |
| 2015 | Supporting Scalable Analytics with Latency ConstraintsabstractRecently there has been a significant interest in building big data analytics systems that can handle both "big data" and "fast data". Our work is strongly motivated by recent real-world use cases that point to the need for a general, unified data processing framework to support analytical queries with different latency requirements. Toward this goal, we start with an analysis of existing big data systems to understand the causes of high latency. We then propose an extended architecture with mini-batches as granularity for computation and shuffling, and augment it with new model-driven resource allocation and runtime scheduling techniques to meet user latency requirements while maximizing throughput. Results from real-world workloads show that our techniques, implemented in Incremental Hadoop, reduce its latency from tens of seconds to sub-second, with 2x-5x increase in throughput. Our system also outperforms state-of-the-art distributed stream systems, Storm and Spark Streaming, by 1-2 orders of magnitude when combining latency and throughput. Boduo Li, Yanlei Diao, Prashant J. Shenoy |
Proc. VLDB Endow. | 3 |
| 2015 | CloudNet: Dynamic Pooling of Cloud Resources by Live WAN Migration of Virtual MachinesabstractVirtualization technology and the ease with which virtual machines (VMs) can be migrated within the LAN have changed the scope of resource management from allocating resources on a single server to manipulating pools of resources within a data center. We expect WAN migration of virtual machines to likewise transform the scope of provisioning resources from a single data center to multiple data centers spread across the country or around the world. In this paper, we present the CloudNet architecture consisting of cloud computing platforms linked with a virtual private network (VPN)-based network infrastructure to provide seamless and secure connectivity between enterprise and cloud data center sites. To realize our vision of efficiently pooling geographically distributed data center resources, CloudNet provides optimized support for live WAN migration of virtual machines. Specifically, we present a set of optimizations that minimize the cost of transferring storage and virtual machine memory during migrations over low bandwidth and high-latency Internet links. We evaluate our system on an operational cloud platform distributed across the continental US. During simultaneous migrations of four VMs between data centers in Texas and Illinois, CloudNet's optimizations reduce memory migration time by 65% and lower bandwidth consumption for the storage and memory transfer by 19 GB, a 50% reduction. Timothy Wood 0001, K. K. Ramakrishnan, Prashant J. Shenoy, Jacobus E. van der Merwe, Jinho Hwang, Guyue Liu, Lucas Chaufournier |
IEEE/ACM Trans. Netw. | 3 |
| 2014 | VMShadow: optimizing the performance of latency-sensitive virtual desktops in distributed cloudsabstractDistributed clouds offer a choice of data center locations to application providers to host their applications. In this paper we consider distributed clouds that host virtual desktops(VDs) which are then accessed by their users through remote desktop protocols. VDs have different sensitivities to latency, primarily determined by the types of applications running (games or video players are more sensitive to latency) and the end users' locations. We design VMShadow, a system to automatically optimize the location and performance of latency-sensitive VDs in the cloud. VMShadow performs black-box fingerprinting of a VM's network traffic to infer its latency-sensitivity and employs a greedy heuristic based algorithm to move highly latency-sensitive VMs to cloud sites that are closer to their end users. VMShadow employs WAN-based live migration and a new network connection migration protocol to ensure that the VM migration and subsequent changes to the VM's network address are transparent to end-users. We implement a prototype of VMShadow in a nested hypervisor and demonstrate its effectiveness for optimizing the performance of VM-based desktops in the cloud. Our experiments on a private and the public EC2 cloud show that VMShadow is able to discriminate between latency-sensitive and insensitive desktop applications and judiciously move only those VMs that will benefit the most. For desktop VMs with video activity, VMShadow improves VNC's refresh rate by 90%. Further our connection migration proxy, which utilizes dynamic rewriting of packet headers, imposes a rewriting overhead of only 13μs per packet. Trans-continental VM migrations take about 4 minutes. Tian Guo 0001, Vijay Gopalakrishnan, K. K. Ramakrishnan, Prashant J. Shenoy, Arun Venkataramani, Seungjoon Lee |
MMSys | 4 |
| 2014 | Combined heat and privacy: Preventing occupancy detection from smart metersabstractElectric utilities are rapidly deploying smart meters that record and transmit electricity usage in real-time. As prior research shows, smart meter data indirectly leaks sensitive, and potentially valuable, information about a home's activities. An important example of the sensitive information smart meters reveal is occupancy-whether or not someone is home and when. As prior work also shows, occupancy is surprisingly easy to detect, since it highly correlates with simple statistical metrics, such as power's mean, variance, and range. Unfortunately, prior research that uses chemical energy storage, e.g., batteries, to prevent appliance power signature detection is prohibitively expensive when applied to occupancy detection. To address this problem, we propose preventing occupancy detection using the thermal energy storage of large elastic heating loads already present in many homes, such as electric water and space heaters. In essence, our approach, which we call Combined Heat and Privacy (CHPr), controls the power usage of these large loads to make it look like someone is always home. We design a CHPr-enabled water heater that regulates its energy usage to mask occupancy without violating its objective, e.g., to provide hot water on demand, and evaluate it in simulation and using a prototype. Our results show that a 50-gallon CHPr-enabled water heater decreases the Matthews Correlation Coefficient (a standard measure of a binary classifier's performance) of a threshold-based occupancy detection attack in a representative home by 10x (from 0.44 to 0.045), effectively preventing occupancy detection at no extra cost. Dong Chen 0010, David Irwin 0001, Prashant J. Shenoy, Jeannie R. Albrecht |
PerCom | 3 |
| 2014 | Empirical Characterization, Modeling, and Analysis of Smart Meter DataabstractSmart meter deployments are spurring renewed interest in analysis techniques for electricity usage data. However, an important prerequisite for data analysis is characterizing and modeling how electrical loads use power. While prior work has made significant progress in deriving insights from electricity data, one issue that limits accuracy is the use of general and often simplistic load models. Prior models often associate a fixed power level with an “on” state and either no power, or some minimal amount, with an “off” state. This paper's goal is to develop a new methodology for modeling electric loads that is both simple and accurate. Our approach is empirical in nature: we monitor a wide variety of common loads to distill a small number of common usage characteristics, which we then leverage to construct accurate load-specific models. We show that our models are significantly more accurate than binary on-off models, decreasing the root mean square error by as much as 8× for representative loads. Finally, we demonstrate three novel applications that use our empirical load models to analyze and derive insights from smart meter data, including i) generating device-accurate synthetic traces of building electricity usage; ii) filtering out loads that generate rapid and random power variations in smart meter data; and iii) detecting the presence of specific load models in time-series power data. Sean Kenneth Barker, Sandeep Kalra, David Irwin 0001, Prashant J. Shenoy |
IEEE J. Sel. Areas Commun. | 4 |
| 2014 | Cost-Aware Cloud Bursting for Enterprise ApplicationsabstractThe high cost of provisioning resources to meet peak application demands has led to the widespread adoption of pay-as-you-go cloud computing services to handle workload fluctuations. Some enterprises with existing IT infrastructure employ a hybrid cloud model where the enterprise uses its own private resources for the majority of its computing, but then “bursts” into the cloud when local resources are insufficient. However, current commercial tools rely heavily on the system administrator’s knowledge to answer key questions such as when a cloud burst is needed and which applications must be moved to the cloud. In this article, we describe Seagull, a system designed to facilitate cloud bursting by determining which applications should be transitioned into the cloud and automating the movement process at the proper time. Seagull optimizes the bursting of applications using an optimization algorithm as well as a more efficient but approximate greedy heuristic. Seagull also optimizes the overhead of deploying applications into the cloud using an intelligent precopying mechanism that proactively replicates virtualized applications, lowering the bursting time from hours to minutes. Our evaluation shows over 100% improvement compared to naïve solutions but produces more expensive solutions compared to ILP. However, the scalability of our greedy algorithm is dramatically better as the number of VMs increase. Our evaluation illustrates scenarios where our prototype can reduce cloud costs by more than 45% when bursting to the cloud, and that the incremental cost added by precopying applications is offset by a burst time reduction of nearly 95%. Tian Guo 0001, Upendra Sharma, Prashant J. Shenoy, Timothy Wood 0001, Sambit Sahu |
ACM Trans. Internet Techn. | 3 |
| 2013 | VMShadow: optimizing the performance of virtual desktops in distributed cloudsabstractWe present VMShadow, a system that automatically optimizes the location and performance of applications based on their dynamic workloads. We prototype VMShadow and demonstrate its efficacy using VM-based desktops in the cloud as an example application. Our experiments on a private cloud as well as the EC2 cloud, using a nested hypervisor, show that VMShadow is able to discriminate between location-sensitive and location-insensitive desktop VMs and judiciously moves only those that will benefit the most from the migration. For example, VMShadow performs transcontinental VM migrations in ~ 4 mins and can improve VNC's video refresh rate by up to 90%. Tian Guo 0001, Vijay Gopalakrishnan, K. K. Ramakrishnan, Prashant J. Shenoy, Arun Venkataramani, Seungjoon Lee |
SoCC | 4 |
| 2013 | mBenchLab: Measuring QoE of Web applications using mobile devicesabstractIn this paper, we present mBenchLab, a software infrastructure to measure the Quality of Experience (QoE) on tablet and smartphones accessing cloud hosted Web services. mBenchLab does not rely on emulation but uses real phones and tablets with their original software stack and communication interfaces for performance evaluation. We have used mBenchLab to measure the QoE of well-known web sites on various devices (Android tablets and smartphones) and networks (Wifi, 3G, 4G). We present our experimental results and lessons learned measuring QoE on mobile devices with mBenchLab. In our QoE analysis, we were able to discover a new bug in a very popular smartphone that impacts both performance and data usage. We have also made the entire mBenchLab software available as open source to the community to measure QoE on mobile devices that access cloud-hosted Web applications. Emmanuel Cecchet, Robert Sims, Prashant J. Shenoy |
IWQoS | 4 |
| 2013 | GreenCache: augmenting off-the-grid cellular towers with multimedia cachesabstractThe growth of smartphones combined with advances in mobile networking have revolutionized the way people consume multimedia data. In particular, users in developing countries primarily rely on smartphones since they often do not have access to more powerful (and more expensive) computing devices. Unfortunately, cellular networks in developing countries have historically had low reliability, due to grid instability and lack of infrastructure. The situation has led network operators to experiment with running cellular towers "off the grid" using intermittent renewable energy sources. In parallel, network operators are also experimenting with co-locating server caches close to cell towers to reduce access latency and back-haul bandwidth. In this paper, we study techniques for optimizing multimedia caches for intermittent renewable energy sources. Specifically, we examine how to apply a blinking abstraction proposed in prior work, which rapidly transitions servers between an active and inactive state, to improve the performance of a multimedia cache powered by renewables, called GreenCache. Our results show that GreenCache's staggered load-proportional blinking policy, which coordinates when servers are active over brief intervals, results in 3X less buffering (or pause) time by the client compared to an activation blinking policy, which simply activates and deactivates servers over long periods as power fluctuates, for realistic power variations from renewable energy sources. Navin Sharma, Dilip Kumar Krishnappa, David Irwin 0001, Michael Zink, Prashant J. Shenoy |
MMSys | 5 |
| 2013 | Yank: Enabling Green Data Centers to Pull the Plug
David Irwin 0001, Prashant J. Shenoy, K. K. Ramakrishnan |
NSDI | 3 |
| 2013 | Automated modeling of complex data center applicationsabstractAs online services become increasingly common, the complexity of backend distributed server applications in data centers has also continued to grow. At the same time, there is an increasing need to enhance the manageability of these large applications by automating common management tasks, which requires a good understanding of the run-time behavior of the application under different scenarios. However, the rising complexity of these applications makes the tasks of manually modeling and analyzing their run-time behavior increasingly difficult. In this talk, I will argue for the need to automate the modeling of the run-time performance of distributed data center applications. I will present techniques for automatically deriving application models using statistical methods from machine learning. I will describe how we have put these ideas into practice into two systems that we have built, Modellus and Predico, and will present case studies of using these systems for management tasks such as capacity planning and what-if analysis of data center applications. Prashant J. Shenoy |
ICPE | 1 |
| 2013 | GreenCharge: Managing RenewableEnergy in Smart BuildingsabstractDistributed generation (DG) uses many small on-site energy harvesting deployments at individual buildings to generate electricity. DG has the potential to make generation more efficient by reducing transmission and distribution losses, carbon emissions, and demand peaks. However, since renewables are intermittent and uncontrollable, buildings must still rely, in part, on the electric grid for power. While DG deployments today use net metering to offset costs and balance local supply and demand, scaling net metering for intermittent renewables to a large fraction of buildings is challenging. In this paper, we explore an alternative approach that combines market-based electricity pricing models with on-site renewables and modest energy storage (in the form of batteries) to incentivize DG. We propose a system architecture and optimization algorithm, called GreenCharge, to efficiently manage the renewable energy and storage to reduce a building's electric bill. To determine when to charge and discharge the battery each day, the algorithm leverages prediction models for forecasting both future energy demand and future energy harvesting. We evaluate GreenCharge in simulation using a collection of real-world data sets, and compare with an oracle that has perfect knowledge of future energy demand/harvesting and a system that only leverages a battery to lower costs (without any renewables). We show that GreenCharge's savings for a typical home today are near 20%, which are greater than the savings from using only net metering. Aditya Kumar Mishra, David Irwin 0001, Prashant J. Shenoy, James F. Kurose, Ting Zhu 0001 |
IEEE J. Sel. Areas Commun. | 3 |
| 2013 | Multimedia systems research: The first twenty years and lessons for the next twentyabstractThis retrospective article examines the past two decades of multimedia systems research through the lens of three research topics that were in vogue in the early days of the field and offers perspectives on the evolution of these research topics. We discuss the eventual impact of each line of research and offer lessons for future research in the field. Prashant J. Shenoy |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2012 | "Cut me some slack": latency-aware live migration for databasesabstractCloud-based data management platforms often employ multitenant databases, where service providers achieve economies of scale by consolidating multiple tenants on shared servers. In such database systems, a key functionality for service providers is database migration, which is useful for dynamic provisioning, load balancing, and system maintenance. Practical migration solutions have several requirements, including high availability, low performance overhead, and self-management. We present Slacker, an end-to-end database migration system at the middleware level satisfying these requirements. Slacker leverages off-the-shelf hot backup tools to achieve live migration with effectively zero down-time. Additionally, Slacker minimizes the performance impact of migrations on both the migrating tenant and collocated tenants by leveraging 'migration slack', or resources that can be used for migration without excessively impacting query latency. We apply a PID controller to this problem, allowing Slacker to automatically detect and exploit migration slack in real time. Using our prototype, we demonstrate that Slacker effectively controls interference during migrations, maintaining latency within 10% of a given latency target, while still performing migrations rapidly and efficiently. Sean Kenneth Barker, Yun Chi, Hyun Jin Moon, Hakan Hacigümüs, Prashant J. Shenoy |
EDBT | 5 |
| 2012 | eTransform: Transforming Enterprise Data Centers by Automated ConsolidationabstractModern day enterprises have a large IT infrastructure comprising thousands of applications running on servers housed in tens of data centers geographically spread out. These enterprises periodically perform a transformation of their entire IT infrastructure to simplify, decrease operational costs and enable easier management. However, the large number of different kinds of applications and data centers involved and the variety of constraints make the task of data center transformation challenging. The state-of-the-art technique for performing this transformation is simplistic, often unable to account for all but the simplest of constraints. We present eTransform, a system for generating a transformation and consolidation plan for the IT infrastructure of large scale enterprises. We devise a linear programming based approach that simultaneously optimizes all the costs involved in enterprise data centers taking into account the constraints of applications groups. Our algorithm handles the various idiosyncrasies of enterprise data centers like volume discounts in pricing, wide-area network costs, traffic matrices, latency constraints, distribution of users accessing the data etc. We include a disaster recovery (DR) plan, so that eTransform, thus provides an integrated disaster recovery and consolidation plan to transform the enterprise IT infrastructure. We use eTransform to perform case studies based on real data from three different large scale enterprises. In our experiments, eTransform is able to suggest a plan to reduce the operational costs by more than 50% from the "as-is" state of these enterprise to the consolidated enterprise IT environment. Even including the DR capability, eTransform is still able to reduce the operational costs by more than 25% from the simple "as-is" state. In our experiments, eTransform is able to simultaneously optimize multiple parameters and constraints and discover solutions that are 7x cheaper than other solutions. Prashant J. Shenoy, K. K. Ramakrishnan, Rahul Kelkar, Harrick M. Vin |
ICDCS | 2 |
| 2012 | Energy-aware load balancing in content delivery networksabstractInternet-scale distributed systems such as content delivery networks (CDNs) operate hundreds of thousands of servers deployed in thousands of data center locations around the globe. Since the energy costs of operating such a large IT infrastructure are a significant fraction of the total operating costs, we argue for redesigning CDNs to incorporate energy optimizations as a first-order principle. We propose techniques to turn off CDN servers during periods of low load while seeking to balance three key design goals: maximize energy reduction, minimize the impact on client-perceived service availability (SLAs), and limit the frequency of on-off server transitions to reduce wear-and-tear and its impact on hardware reliability. We propose an optimal offline algorithm and an online algorithm to extract energy savings both at the level of local load balancing within a data center and global load balancing across data centers. We evaluate our algorithms using real production workload traces from a large commercial CDN. Our results show that it is possible to reduce the energy consumption of a CDN by 51% while ensuring a high level of availability that meets customer SLA requirements and incurring an average of one on-off transition per server per day. Further, we show that keeping even 10% of the servers as hot spares helps absorb load spikes due to global flash crowds and minimize any impact on availability SLAs. Finally, we show that redistributing load across highly proximal data centers can enhance service availability significantly, but has only a modest impact on energy savings. Vimal Mathew, Ramesh K. Sitaraman, Prashant J. Shenoy |
INFOCOM | 3 |
| 2012 | SmartCap: Flattening peak electricity demand in smart homesabstractFlattening household electricity demand reduces generation costs, since costs are disproportionately affected by peak demands. While the vast majority of household electrical loads are interactive and have little scheduling flexibility (TVs, microwaves, etc.), a substantial fraction of home energy use derives from background loads with some, albeit limited, flexibility. Examples of such devices include A/Cs, refrigerators, and dehumidifiers. In this paper, we study the extent to which a home is able to transparently flatten its electricity demand by scheduling only background loads with such flexibility. We propose a Least Slack First (LSF) scheduling algorithm for household loads, inspired by the well-known Earliest Deadline First algorithm. We then integrate the algorithm into Smart-Cap, a system we have built for monitoring and controlling electric loads in homes. To evaluate LSF, we collected power data at outlets, panels, and switches from a real home for 82 days. We use this data to drive simulations, as well as experiment with a real testbed implementation that uses similar background loads as our home. Our results indicate that LSF is most useful during peak usage periods that exhibit “peaky” behavior, where power deviates frequently and significantly from the average. For example, LSF decreases the average deviation from the mean power by over 20% across all 4-hour periods where the deviation is at least 400 watts. Sean Kenneth Barker, Aditya Kumar Mishra, David Irwin 0001, Prashant J. Shenoy, Jeannie R. Albrecht |
PerCom | 4 |
| 2012 | An Empirical Study of Memory Sharing in Virtual Machines
Sean Kenneth Barker, Timothy Wood 0001, Prashant J. Shenoy, Ramesh K. Sitaraman |
USENIX ATC | 3 |
| 2012 | Seagull: Intelligent Cloud Bursting for Enterprise Applications
Tian Guo 0001, Upendra Sharma, Timothy Wood 0001, Sambit Sahu, Prashant J. Shenoy |
USENIX ATC | 5 |
| 2012 | Enterprise-Ready Virtual Cloud Pools: Vision, Opportunities and ChallengesabstractCloud computing platforms such as Amazon EC2 provide customers with flexible, on demand resources at low cost. However, while existing offerings are useful for providing basic computation and storage resources, they have not provided the transparency, security and network controls that many enterpise customers would like. While cloud computing has a great potential to change how enterprises run and manage their IT systems, a more comprehensive control over network resources and security needs to be provided for such users. Towards this goal, we propose a Virtual Cloud Pool abstraction to logically unify cloud and enterprise data center resources, and present the vision behind CloudNet, a cloud platform architecture which utilizes virtual private networks to securely and seamlessly link cloud and enterprise sites. It also enables the pooling of resources across data centers to provide enterprises the capability to have cloud resources that are dynamic and adaptive to their needs. We describe several usage scenarios for virtual cloud pools and discuss the benefits of using this abstraction in enterprise settings. Timothy Wood 0001, K. K. Ramakrishnan, Prashant J. Shenoy, Jacobus E. van der Merwe |
Comput. J. | 3 |
| 2012 | MultiSense: proportional-share for mechanically steerable sensor networks
Navin Sharma, David Irwin 0001, Michael Zink, Prashant J. Shenoy |
Multim. Syst. | 4 |
| 2012 | SPIRE: Efficient Data Inference and Compression over RFID StreamsabstractDespite its promise, RFID technology presents numerous challenges, including incomplete data, lack of location and containment information, and very high volumes. In this work, we present a novel data inference and compression substrate over RFID streams to address these challenges. Our substrate employs a time-varying graph model to efficiently capture possible object locations and interobject relationships such as containment from raw RFID streams. It then employs a probabilistic algorithm to estimate the most likely location and containment for each object. By performing such online inference, it enables online compression that recognizes and removes redundant information from the output stream of this substrate. We have implemented a prototype of our inference and compression substrate and evaluated it using both real traces from a laboratory warehouse setup and synthetic traces emulating enterprise supply chains. Results of a detailed performance study show that our data inference techniques provide high accuracy while retaining efficiency over RFID data streams, and our compression algorithm yields significant reduction in output data volume. Yanming Nie, Richard Cocci, Zhao Cao, Yanlei Diao, Prashant J. Shenoy |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2012 | SCALLA: A Platform for Scalable One-Pass Analytics Using MapReduceabstractToday’s one-pass analytics applications tend to be data-intensive in nature and require the ability to process high volumes of data efficiently. MapReduce is a popular programming model for processing large datasets using a cluster of machines. However, the traditional MapReduce model is not well-suited for one-pass analytics, since it is geared towards batch processing and requires the dataset to be fully loaded into the cluster before running analytical queries. This article examines, from a systems standpoint, what architectural design changes are necessary to bring the benefits of the MapReduce model to incremental one-pass analytics. Our empirical and theoretical analyses of Hadoop-based MapReduce systems show that the widely used sort-merge implementation for partitioning and parallel processing poses a fundamental barrier to incremental one-pass analytics, despite various optimizations. To address these limitations, we propose a new data analysis platform that employs hash techniques to enable fast in-memory processing, and a new frequent key based technique to extend such processing to workloads that require a large key-state space. Evaluation of our Hadoop-based prototype using real-world workloads shows that our new platform significantly improves the progress of map tasks, allows the reduce progress to keep up with the map progress, with up to 3 orders of magnitude reduction of internal data spills, and enables results to be returned continuously during the job. Boduo Li, Edward Mazur, Yanlei Diao, Andrew McGregor 0001, Prashant J. Shenoy |
ACM Trans. Database Syst. | 5 |
| 2012 | Modellus: Automated modeling of complex internet data center applicationsabstractThe rising complexity of distributed server applications in Internet data centers has made the tasks of modeling and analyzing their behavior increasingly difficult. This article presents Modellus , a novel system for automated modeling of complex web-based data center applications using methods from queuing theory, data mining, and machine learning. Modellus uses queuing theory and statistical methods to automatically derive models to predict the resource usage of an application and the workload it triggers; these models can be composed to capture multiple dependencies between interacting applications. Model accuracy is maintained by fast, distributed testing, automated relearning of models when they change, and methods to bound prediction errors in composite models. We have implemented a prototype of Modellus, deployed it on a data center testbed, and evaluated its efficacy for modeling and analysis of several distributed multitier web applications. Our results show that this feature-based modeling technique is able to make predictions across several data center tiers, and maintain predictive accuracy (typically 95% or better) in the face of significant shifts in workload composition; we also demonstrate practical applications of the Modellus system to prediction and provisioning of real-world data center applications. Peter Desnoyers, Timothy Wood 0001, Prashant J. Shenoy, Sangameshwar Patil, Harrick M. Vin |
ACM Trans. Web | 3 |
| 2011 | Blink: managing server clusters on intermittent powerabstractReducing the energy footprint of data centers continues to receive significant attention due to both its financial and environmental impact. There are numerous methods that limit the impact of both factors, such as expanding the use of renewable energy or participating in automated demand-response programs. To take advantage of these methods, servers and applications must gracefully handle intermittent constraints in their power supply. In this paper, we propose blinking---metered transitions between a high-power active state and a low-power inactive state---as the primary abstraction for conforming to intermittent power constraints. We design Blink, an application-independent hardware-software platform for developing and evaluating blinking applications, and define multiple types of blinking policies. We then use Blink to design BlinkCache, a blinking version of memcached, to demonstrate the effect of blinking on an example application. Our results show that a load-proportional blinking policy combines the advantages of both activation and synchronous blinking for realistic Zipf-like popularity distributions and wind/solar power signals by achieving near optimal hit rates (within 15% of an activation policy), while also providing fairer access to the cache (within 2% of a syn- chronous policy) for equally popular objects. Navin Sharma, Sean Kenneth Barker, David Irwin 0001, Prashant J. Shenoy |
ASPLOS | 4 |
| 2011 | PipeCloud: using causality to overcome speed-of-light delays in cloud-based disaster recoveryabstractDisaster Recovery (DR) is a desirable feature for all enterprises, and a crucial one for many. However, adoption of DR remains limited due to the stark tradeoffs it imposes. To recover an application to the point of crash, one is limited by financial considerations, substantial application overhead, or minimal geographical separation between the primary and recovery sites. In this paper, we argue for cloud-based DR and pipelined synchronous replication as an antidote to these problems. Cloud hosting promises economies of scale and on-demand provisioning that are a perfect fit for the infrequent yet urgent needs of DR. Pipelined synchrony addresses the impact of WAN replication latency on performance, by efficiently overlapping replication with application processing for multi-tier servers. By tracking the consequences of the disk modifications that are persisted to a recovery site all the way to client-directed messages, applications realize forward progress while retaining full consistency guarantees for client-visible state in the event of a disaster. PipeCloud, our prototype, is able to sustain these guarantees for multi-node servers composed of black-box VMs, with no need of application modification, resulting in a perfect fit for the arbitrary nature of VM-based cloud hosting. We demonstrate disaster failover to the Amazon EC2 platform, and show that PipeCloud can increase throughput by an order of magnitude and reduce response times by more than half compared to synchronous replication, all while providing the same zero data loss consistency guarantees. Timothy Wood 0001, H. Andrés Lagar-Cavilla, K. K. Ramakrishnan, Prashant J. Shenoy, Jacobus E. van der Merwe |
SoCC | 4 |
| 2011 | ZZ and the art of practical BFT executionabstractThe high replication cost of Byzantine fault-tolerance (BFT) methods has been a major barrier to their widespread adop-tion in commercial distributed applications. We present ZZ, a new approach that reduces the replication cost of BFT ser-vices from 2f + 1 to practically f + 1. The key insight in ZZ is to use f + 1 execution replicas in the normal case and to activate additional replicas only upon failures. In data cen-ters where multiple applications share a physical server, ZZ reduces the aggregate number of execution replicas running in the data center, improving throughput and response times. ZZ relies on virtualization—a technology already employed in modern data centers—for fast replica activation upon fail-ures, and enables newly activated replicas to immediately be-gin processing requests by fetching state on-demand. A pro-totype implementation of ZZ using the BASE library and Xen shows that, when compared to a system with 2f + 1 repli-cas, our approach yields lower response times and up to 33% higher throughput in a prototype data center with four BFT web applications. We also show that ZZ can handle simulta-neous failures and achieve sub-second recovery. 1 Timothy Wood 0001, Arun Venkataramani, Prashant J. Shenoy, Emmanuel Cecchet |
EuroSys | 4 |
| 2011 | A Cost-Aware Elasticity Provisioning System for the CloudabstractIn this paper we present Kingfisher, a cost-aware system that provides efficient support for elasticity in the cloud by (i) leveraging multiple mechanisms to reduce the time to transition to new configurations, and (ii) optimizing the selection of a virtual server configuration that minimizes the cost. We have implemented a prototype of Kingfisher and have evaluated its efficacy on a laboratory cloud platform. Our experiments with varying application workloads demonstrate that Kingfisher is able to (i) decrease the cost of virtual server resources by as much as 24% compared to the current cost-unaware approach, (ii) reduce by an order of magnitude the time to transition to a new configuration through multiple elasticity mechanisms in the cloud, and (iii), illustrate the opportunity for design alternatives which trade-off the cost of server resources with the time required to scale the application. Upendra Sharma, Prashant J. Shenoy, Sambit Sahu, Anees Shaikh |
ICDCS | 2 |
| 2011 | Kingfisher: Cost-aware elasticity in the cloudabstractIn this paper we present Kingfisher, a cost-aware system that provides efficient support for elasticity in the cloud by (i) leveraging multiple mechanisms to reduce the time to transition to new configurations, and (ii) optimizing the selection of a virtual server configuration that minimizes the cost. We have implemented a prototype of Kingfisher and have evaluated its efficacy on a laboratory cloud platform. Our experiments with varying application workloads demonstrate that Kingfisher is able to (i) decrease the cost of virtual server resources by as much as 24% compared to the current cost-unaware approach, (ii) reduce by an order of magnitude the time to transition to a new configuration through multiple elasticity mechanisms in the cloud, and (iii), illustrate the opportunity for further design alternatives which trade-off the cost of server resources with the time required to scale the application. Upendra Sharma, Prashant J. Shenoy, Sambit Sahu, Anees Shaikh |
INFOCOM | 2 |
| 2011 | Predico: A System for What-if Analysis in Complex Data Center Applications
Prashant J. Shenoy, Maitreya Natu, Vaishali P. Sadaphal, Harrick M. Vin |
Middleware | 2 |
| 2011 | MultiSense: fine-grained multiplexing for steerable camera sensor networksabstractSteerable sensors, such as pan-tilt-zoom video cameras, expose programmable actuators to applications, which steer them in different directions based on their goals. Despite being expensive to deploy and maintain, existing steerable sensor networks allow only a single application to control them due to the slow speed of their mechanical actuators. To address the problem, we design MultiSense to enable fine-grained multiplexing by (i) exposing a virtual sensor to each application and (ii) optimizing the time to context-switch between virtual sensors and satisfy requests. Navin Sharma, David Irwin 0001, Prashant J. Shenoy, Michael Zink |
MMSys | 3 |
| 2011 | A platform for scalable one-pass analytics using MapReduceabstractToday’s one-pass analytics applications tend to be data-intensive in nature and require the ability to process high volumes of data efficiently. MapReduce is a popular programming model for processing large datasets using a cluster of machines. However, the traditional MapReduce model is not well-suited for one-pass analytics, since it is geared towards batch processing and requires the data set to be fully loaded into the cluster before running analytical queries. This paper examines, from a systems standpoint, what architectural design changes are necessary to bring the benefits of the MapReduce model to incremental one-pass analytics. Our empirical and theoretical analyses of Hadoop-based MapReduce systems show that the widely-used sort-merge implementation for partitioning and parallel processing poses a fundamental barrier to incremental one-pass analytics, despite various optimizations. To address these limitations, we propose a new data analysis platform that employs hash techniques to enable fast in-memory processing, and a new frequent key based technique to extend such processing to workloads that require a large key-state space. Evaluation of our Hadoop-based prototype using real-world workloads shows that our new platform significantly improves the progress of map tasks, allows the reduce progress to keep up with the map progress, with up to 3 orders of magnitude reduction of internal data spills, and enables results to be returned continuously during the job. 1. Boduo Li, Edward Mazur, Yanlei Diao, Andrew McGregor 0001, Prashant J. Shenoy |
SIGMOD Conference | 5 |
| 2011 | Sharing-aware algorithms for virtual machine colocationabstractVirtualization technology enables multiple virtual machines (VMs) to run on a single physical server. VMs that run on the same physical server can share memory pages that have identical content, thereby reducing the overall memory requirements on the server. We develop sharing-aware algorithms that can colocate VMs with similar page content on the same physical server to optimize the benefits of inter-VM sharing. We show that inter-VM sharing occurs in a largely hierarchical fashion, where the sharing can be attributed to VM's running the same OS platform, OS version, software libraries, or applications. We propose two hierarchical sharing models: a tree model and a more general cluster-tree model. Using a set of VM traces, we show that up to 67% percent of the inter-VM sharing is captured by the tree model and up to 82% is captured by the cluster-tree model. Next, we study two problem variants of critical interest to a virtualization service provider: the VM Maximization problem that determines the most profitable subset of the VMs that can be packed into the given set of servers, and the VM packing problem that determines the smallest set of servers that can accommodate a set of VMs. While both variants are NP-hard, we show that both admit provably good approximation schemes in the hierarchical sharing models. We show that VM maximization for the tree and cluster-tree models can be approximated in polytime to within a (1 - 1/e) factor of optimal. Further, we show that VM packing can be approximated in polytime to within a factor of O(log n) of optimal for cluster-trees and to within a factor of 3 of optimal for trees, where n is the number of VMs. Finally, we evaluate our VM packing algorithm for the tree sharing model on real-world VM traces and show that our algorithm can exploit most of the available inter-VM sharing to achieve a 32% to 50% reduction in servers and a 25% to 57% reduction in memory footprint compared to sharing-oblivious algorithms. Michael Sindelar, Ramesh K. Sitaraman, Prashant J. Shenoy |
SPAA | 3 |
| 2011 | Dolly: virtualization-driven database provisioning for the cloudabstractCloud computing platforms are becoming increasingly popular for e-commerce applications that can be scaled on-demand in a very cost effective way. Dynamic provisioning is used to autonomously add capacity in multi-tier cloud-based applications that see workload increases. While many solutions exist to provision tiers with little or no state in applications, the database tier remains problematic for dynamic provisioning due to the need to replicate its large disk state. In this paper, we explore virtual machine (VM) cloning techniques to spawn database replicas and address the challenges of provisioning shared-nothing replicated databases in the cloud. We argue that being able to determine state replication time is crucial for provisioning databases and show that VM cloning provides this property. We propose Dolly, a database provisioning system based on VM cloning and cost models to adapt the provisioning policy to the cloud infrastructure specifics and application requirements. We present an implementation of Dolly in a commercial-grade replication middleware and evaluate database provisioning strategies for a TPC-W workload on a private cloud and on Amazon EC2. By being aware of VM-based state replication cost, Dolly can solve the challenge of automated provisioning for replicated databases on cloud platforms. Emmanuel Cecchet, Upendra Sharma, Prashant J. Shenoy |
VEE | 4 |
| 2011 | CloudNet: dynamic pooling of cloud resources by live WAN migration of virtual machinesabstractVirtual machine technology and the ease with which VMs can be migrated within the LAN, has changed the scope of resource management from allocating resources on a single server to manipulating pools of resources within a data center. We expect WAN migration of virtual machines to likewise transform the scope of provisioning compute resources from a single data center to multiple data centers spread across the country or around the world. In this paper we present the CloudNet architecure as a cloud framework consisting of cloud computing platforms linked with a VPN based network infrastructure to provide seamless and secure connectivity between enterprise and cloud data center sites. To realize our vision of efficiently pooling geographically distributed data center resources, CloudNet provides optimized support for live WAN migration of virtual machines. Specifically, we present a set of optimizations that minimize the cost of transferring storage and virtual machine memory during migrations over low bandwidth and high latency Internet links. We evaluate our system on an operational cloud platform distributed across the continental US. During simultaneous migrations of four VMs between data centers in Texas and Illinois, CloudNet's optimizations reduce memory migration time by 65% and lower bandwidth consumption for the storage and memory transfer by 19GB, a 50% reduction. Timothy Wood 0001, K. K. Ramakrishnan, Prashant J. Shenoy, Jacobus E. van der Merwe |
VEE | 3 |
| 2011 | Ferret: An RFID-enabled pervasive multimedia application
Mark D. Corner, Prashant J. Shenoy |
Ad Hoc Networks | 3 |
| 2011 | Distributed inference and query processing for RFID tracking and monitoringabstractIn this paper, we present the design of a scalable, distributed stream processing system for RFID tracking and monitoring. Since RFID data lacks containment and location information that is key to query processing, we propose to combine location and containment inference with stream query processing in a single architecture, with inference as an enabling mechanism for high-level query processing. We further consider challenges in instantiating such a system in large distributed settings and design techniques for distributed inference and query processing. Our experimental results, using both real-world data and large synthetic traces, demonstrate the accuracy, efficiency, and scalability of our proposed techniques. Zhao Cao, Charles Sutton, Yanlei Diao, Prashant J. Shenoy |
Proc. VLDB Endow. | 4 |
| 2010 | Resource management in data-intensive clouds: Opportunities and challengesabstractToday's cloud computing platforms have seen much success in running compute-bound applications with time-varying or one-time needs. In this position paper, we will argue that the cloud paradigm is also well suited for handling data-intensive applications, characterized by the processing and storage of data produced by high-bandwidth sensors or streaming applications. The data rates and the processing demands vary over time for many such applications, making the on-demand cloud paradigm a good match for their needs. However, today's cloud platforms need to evolve to meet the storage, communication, and processing demands of data-intensive applications. We present an ongoing GENI project to connect high-bandwidth radar sensor networks with computational and storage resources in the cloud and use this example to highlight the opportunities and challenges in designing end-to-end data-intensive cloud systems. David Irwin 0001, Prashant J. Shenoy, Emmanuel Cecchet, Michael Zink |
LANMAN | 2 |
| 2010 | Empirical evaluation of latency-sensitive application performance in the cloudabstractCloud computing platforms enable users to rent computing and storage resources on-demand to run their networked applications and employ virtualization to multiplex virtual servers belonging to different customers on a shared set of servers. In this paper, we empirically evaluate the efficacy of cloud platforms for running latency-sensitive multimedia applications. Since multiple virtual machines running disparate applications from independent users may share a physical server, our study focuses on whether dynamically varying background load from such applications can interfere with the performance seen by latency-sensitive tasks. We first conduct a series of experiments on Amazon's EC2 system to quantify the CPU, disk, and network jitter and throughput fluctuations seen over a period of several days. We then turn to a laboratory-based cloud and systematically introduce different levels of background load and study the ability to isolate applications under different settings of the underlying resource control mechanisms. We use a combination of micro-benchmarks and two real-world applications--the Doom 3 game server and Apple's Darwin Streaming Server--for our experimental evaluation. Our results reveal that the jitter and the throughput seen by a latency-sensitive application can indeed degrade due to background load from other virtual machines. The degree of interference varies from resource to resource and is the most pronounced for disk-bound latency-sensitive tasks, which can degrade by nearly 75% under sustained background load. We also find that careful configuration of the resource control mechanisms within the virtualization layer can mitigate, but not eliminate, this interference. Sean Kenneth Barker, Prashant J. Shenoy |
MMSys | 2 |
| 2010 | Exploiting the Interplay between Memory and Flash Storage in Embedded Sensor DevicesabstractAlthough memory is an important constraint in embedded sensor nodes, existing embedded applications and systems are typically designed to work under the memory constraints of a single platform and do not consider the interplay between memory and flash storage. In this paper, we present the design of a memory-adaptive flash-based embedded sensor system that allows an application to exploit the presence of flash and adapt to different amounts of RAM on the embedded device. We describe how such a system can be exploited by data-centric sensor applications. Our design involves several novel features: flash and memory-efficient storage and indexing, techniques for efficient storage reclamation, and intelligent buffer management to maximize write coalescing. Our results show that our system is highly energy-efficient under different workloads, and can be configured for embedded sensor platforms with memory constraints ranging from a few kilobytes to hundreds of kilobytes. Devesh Agrawal, Boduo Li, Zhao Cao, Deepak Ganesan, Yanlei Diao, Prashant J. Shenoy |
RTCSA | 6 |
| 2010 | Cloudy Computing: Leveraging Weather Forecasts in Energy Harvesting Sensor SystemsabstractTo sustain perpetual operation, systems that harvest environmental energy must carefully regulate their usage to satisfy their demand. Regulating energy usage is challenging if a system's demands are not elastic and its hardware components are not energy-proportional, since it cannot precisely scale its usage to match its supply. Instead, the system must choose when to satisfy its energy demands based on its current energy reserves and predictions of its future energy supply. In this paper, we explore the use of weather forecasts to improve a system's ability to satisfy demand by improving its predictions. We analyze weather forecast, observational, and energy harvesting data to formulate a model that translates a weather forecast to a wind or solar energy harvesting prediction, and quantify its accuracy. We evaluate our model for both energy sources in the context of two different energy harvesting sensor systems with inelastic demands: a sensor testbed that leases sensors to external users and a lexicographically fair sensor network that maintains steady node sensing rates. We show that using weather forecasts in both wind- and solar-powered sensor systems increases each system's ability to satisfy its demands compared with existing prediction strategies. Navin Sharma, Jeremy Gummeson, David Irwin 0001, Prashant J. Shenoy |
SECON | 4 |
| 2010 | An adaptive link layer for heterogeneous multi-radio mobile sensor networksabstractAn important challenge in mobile sensor networks is to enable energy-efficient communication over a diversity of distanceYC while being robust to wireless effects caused by node mobility. In this paper, we argue that the pairing of two complementary radios with heterogeneous range characteristics enables greater range and interference diversity at lower energy cost than a single radio. We make three contributions towards the design of such multi-radio mobile sensor systems. First, we present the design of a novel reinforcement learning-based link layer algorithm that continually learns channel characteristics and dynamically decides when to switch between radios. Second, we describe a simple protocol that translates the benefits of the adaptive link layer into practice in an energy-efficient manner. Third, we present the design of Arthropod, a mote-class sensor platform that combines two such heterogneous radios (XE1205 and CC2420) and our implementation of the Q-learning based switching protocol in TinyOS 2.0. Using experiments conducted in a variety of urban and forested environments, we show that our system achieves up to 52% energy gains over a single radio system while handling node mobility. Our results also show that our system can handle short, medium and long-term wireless interference in such environments. Jeremy Gummeson, Deepak Ganesan, Mark D. Corner, Prashant J. Shenoy |
IEEE J. Sel. Areas Commun. | 4 |
| 2009 | SRCP: Simple Remote Control for Perpetual High-Power Sensor Networks
Navin Sharma, Jeremy Gummeson, David Irwin 0001, Prashant J. Shenoy |
EWSN | 4 |
| 2009 | Probabilistic Inference over RFID Streams in Mobile EnvironmentsabstractRecent innovations in RFID technology are enabling large-scale cost-effective deployments in retail, healthcare, pharmaceuticals and supply chain management. The advent of mobile or handheld readers adds significant new challenges to RFID stream processing due to the inherent reader mobility, increased noise, and incomplete data. In this paper, we address the problem of translating noisy, incomplete raw streams from mobile RFID readers into clean, precise event streams with location information. Specifically we propose a probabilistic model to capture the mobility of the reader, object dynamics, and noisy readings. Our model can self-calibrate by automatically estimating key parameters from observed data. Based on this model, we employ a sampling-based technique called particle filtering to infer clean, precise information about object locations from raw streams from mobile RFID readers. Since inference based on standard particle filtering is neither scalable nor efficient in our settings, we propose three enhancements-particle factorization, spatial indexing, and belief compression-for scalable inference over large numbers of objects and high-volume streams. Our experiments show that our approach can offer 49% error reduction over a state-of-the-art data cleaning approach such as SMURF while also being scalable and efficient. Thanh T. L. Tran, Charles Sutton, Richard Cocci, Yanming Nie, Yanlei Diao, Prashant J. Shenoy |
ICDE | 6 |
| 2009 | A Multi-Agent Learning Approach to Online Distributed Resource Allocation
Chongjie Zhang, Victor R. Lesser, Prashant J. Shenoy |
IJCAI | 3 |
| 2009 | An Adaptive Link Layer for Range Diversity in Multi-Radio Mobile Sensor NetworksabstractAn important challenge in mobile sensor networks is to enable energy-efficient communication over a diversity of distances while being robust to wireless effects caused by node mobility. In this paper, we argue that the pairing of two complementary radios with heterogeneous range characteristics enables greater range diversity at lower energy cost than a single radio. We make three contributions towards the design of such multi-radio mobile sensor systems. First, we present the design of a novel reinforcement learning-based link layer algorithm that continually learns channel characteristics and dynamically decides when to switch between radios. Second, we describe a simple protocol that translates the benefits of the adaptive link layer into practice in an energy-efficient manner. Third, we present the design of Arthropod, a mote-class sensor platform that combines two such heterogeneous radios (XE1205 and CC2420) and our implementation of the Q-learning based switching protocol in TinyOS 2.0. Using experiments conducted in a variety of urban and forested environments, we show that our system achieves up to 52% energy gains over a single radio system. Jeremy Gummeson, Deepak Ganesan, Mark D. Corner, Prashant J. Shenoy |
INFOCOM | 4 |
| 2009 | Memory buddies: exploiting page sharing for smart colocation in virtualized data centersabstractMany data center virtualization solutions, such as VMware ESX, employ content-based page sharing to consolidate the resources of multiple servers. Page sharing identifies virtual machine memory pages with identical content and consolidates them into a single shared page. This technique, implemented at the host level, applies only between VMs placed on a given physical host. In a multi-server data center, opportunities for sharing may be lost because the VMs holding identical pages are resident on different hosts. In order to obtain the full benefit of content-based page sharing it is necessary to place virtual machines such that VMs with similar memory content are located on the same hosts. Timothy Wood 0001, Gabriel Tarasuk-Levin, Prashant J. Shenoy, Peter Desnoyers, Emmanuel Cecchet, Mark D. Corner |
VEE | 3 |
| 2009 | Sandpiper: Black-box and gray-box resource management for virtual machines
Timothy Wood 0001, Prashant J. Shenoy, Arun Venkataramani, Mazin S. Yousif |
Comput. Networks | 2 |
| 2009 | Resource overbooking and application profiling in a shared Internet hosting platformabstractIn this article, we present techniques for provisioning CPU and network resources in shared Internet hosting platforms running potentially antagonistic third-party applications. The primary contribution of our work is to demonstrate the feasibility and benefits of overbooking resources in shared Internet platforms. Since an accurate estimate of an application's resource needs is necessary when overbooking resources, we present techniques to profile applications on dedicated nodes, possibly while in service, and use these profiles to guide the placement of application components onto shared nodes. We then propose techniques to overbook cluster resources in a controlled fashion. We outline an empirical appraoch to determine the degree of overbooking that allows a platform to achieve improvements in revenue while providing performance guarantees to Internet applications. We show how our techniques can be combined with commonly used QoS resource allocation mechanisms to provide application isolation and performance guarantees at run-time. We implement our techniques in a Linux cluster and evaluate them using common server applications. We find that the efficiency (and consequently revenue) benefits from controlled overbooking of resources can be dramatic. Specifically, we find that overbooking resources by as little as 1% we can increase the utilization of the cluster by a factor of two, and a 5% overbooking yields a 300--500% improvement, while still providing useful resource guarantees to applications. Bhuvan Urgaonkar, Prashant J. Shenoy, Timothy Roscoe |
ACM Trans. Internet Techn. | 2 |
| 2009 | SEVA: Sensor-enhanced video annotationabstractIn this article, we study how a sensor-rich world can be exploited by digital recording devices such as cameras and camcorders to improve a user's ability to search through a large repository of image and video files. We design and implement a digital recording system that records identities and locations of objects (as advertised by their sensors) along with visual images (as recorded by a camera). The process, which we refer to as Sensor-Enhanced Video Annotation (SEVA) , combines a series of correlation, interpolation, and extrapolation techniques. It produces a tagged stream that later can be used to efficiently search for videos or frames containing particular objects or people. We present detailed experiments with a prototype of our system using both stationary and mobile objects as well as GPS and ultrasound. Our experiments show that: (i) SEVA has zero error rates for static objects, except very close to the boundary of the viewable area; (ii) for moving objects or a moving camera, SEVA only misses objects leaving or entering the viewable area by 1--2 frames; (iii) SEVA can scale to 10 fast-moving objects using current sensor technology; and (iv) SEVA runs online using relatively inexpensive hardware. Mark D. Corner, Prashant J. Shenoy |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2009 | PRESTO: feedback-driven data management in sensor networks
Ming Li 0009, Deepak Ganesan, Prashant J. Shenoy |
IEEE/ACM Trans. Netw. | 3 |
| 2009 | Ultra-low power data storage for sensor networksabstractLocal storage is required in many sensor network applications, both for archival of detailed event information, as well as to overcome sensor platform memory constraints. Recent gains in energy efficiency of new-generation NAND flash storage have strengthened the case for in-network storage by data-centric sensor network applications. We argue that current storage solutions offering a simple file system abstraction are inadequate for sensor applications to exploit storage. Instead, we propose Capsule—a rich, flexible and portable object storage abstraction that offers stream, file, array, queue and index storage objects for data storage and retrieval. Further, Capsule supports checkpointing and rollback of object state for fault tolerance. Our experiments demonstrate that Capsule provides platform independence, greater functionality and greater energy efficiency than existing storage solutions. Gaurav Mathur, Peter Desnoyers, Paul Chukiu, Deepak Ganesan, Prashant J. Shenoy |
ACM Trans. Sens. Networks | 5 |
| 2008 | Efficient Data Interpretation and Compression over RFID StreamsabstractDespite its promise, RFID technology presents numerous challenges, including incomplete data, lack of location and containment information, and very high volumes. In this work, we present a novel data interpretation and compression substrate over RFID streams to address these challenges in enterprise supply-chain environments. Our results show that our inference techniques provide good accuracy while retaining efficiency, and our compression algorithm yields significant reduction in data volume. Richard Cocci, Thanh T. L. Tran, Yanlei Diao, Prashant J. Shenoy |
ICDE | 4 |
| 2008 | Profiling and Modeling Resource Usage of Virtualized Applications
Timothy Wood 0001, Ludmila Cherkasova, Kivanc M. Ozonat, Prashant J. Shenoy |
Middleware | 4 |
| 2008 | Sherlock: automatically locating objects for humansabstractOver the course of a day a human interacts with tens or hundreds of individual objects. Many of these articles are nomadic, relying on human memory to manually index, inventory, organize, search, and locate them. However, Radio Frequency Identification (RFID) tags hold great promise for automating these tasks. While originally envisioned for managing supply chains and store inventories, RFID tags support the properties necessary for helping humans to manage their objects. This paper presents Sherlock, a system that leverages RFID tags for human-object interaction. Sherlock combines concepts from sensors, radar technology, and computer graphics to implement a novel localization and visualization system for everyday objects. At the heart of Sherlock is a new RFID localization technique that uses steerable antennas to sweep a room, discovering, localizing and indexing tagged objects. In response to user queries, Sherlock displays the locations of matching objects using images from a video camera. We have implemented a prototype of Sherlock to conduct experiments in a real office environment. Our results demonstrate the effectiveness of Sherlock in localizing to a volume of less than 0.55 cubic meters for 90% of objects. Aditya Nemmaluri, Mark D. Corner, Prashant J. Shenoy |
MobiSys | 3 |
| 2008 | Data Quality and Query Cost in Pervasive Sensing SystemsabstractThis research is motivated by large-scale pervasive sensing applications. We examine the benefits and costs of caching data for such applications. We propose and evaluate several approaches to querying for, and then caching data in a sensor field data server. We show that for some application requirements (i.e., when delay drives data quality), policies that emulate cache hits by computing and returning approximate values for sensor data yield a simultaneous quality improvement and cost savings. This win-win is because when system delay is sufficiently important, the benefit to both query cost and data quality achieved by using approximate values outweighs the negative impact on quality due to the approximation. In contrast, when data accuracy drives quality, a linear trade-off between query cost and data quality emerges. We also identify caching and lookup policies for which the sensor field query rate is bounded when servicing an arbitrary workload of user queries. This upper bound is achieved by having multiple user queries share the cost of a single sensor field query. Finally, we demonstrate that our results are robust to the manner in which the environment being monitored changes using models for two different sensing systems. David J. Yates, Erich M. Nahum, James F. Kurose, Prashant J. Shenoy |
PerCom | 4 |
| 2008 | Cataclysm: Scalable overload policing for internet applications
Bhuvan Urgaonkar, Prashant J. Shenoy |
J. Netw. Comput. Appl. | 2 |
| 2008 | Data quality and query cost in pervasive sensing systems
David J. Yates, Erich M. Nahum, James F. Kurose, Prashant J. Shenoy |
Pervasive Mob. Comput. | 4 |
| 2008 | Agile dynamic provisioning of multi-tier Internet applicationsabstractDynamic capacity provisioning is a useful technique for handling the multi-time-scale variations seen in Internet workloads. In this article, we propose a novel dynamic provisioning technique for multi-tier Internet applications that employs (1) a flexible queuing model to determine how much of the resources to allocate to each tier of the application, and (2) a combination of predictive and reactive methods that determine when to provision these resources, both at large and small time scales. We propose a novel data center architecture based on virtual machine monitors to reduce provisioning overheads. Our experiments on a forty-machine Xen/Linux-based hosting platform demonstrate the responsiveness of our technique in handling dynamic workloads. In one scenario where a flash crowd caused the workload of a three-tier application to double, our technique was able to double the application capacity within five minutes, thus maintaining response-time targets. Our technique also reduced the overhead of switching servers across applications from several minutes to less than a second, while meeting the performance targets of residual sessions. Bhuvan Urgaonkar, Prashant J. Shenoy, Abhishek Chandra, Pawan Goyal 0001, Timothy Wood 0001 |
ACM Trans. Auton. Adapt. Syst. | 2 |
| 2008 | Chameleon: Application-Level Power ManagementabstractIn this paper, we present Chameleon an application-level power management approach for reducing energy consumption in mobile processors. By using application domain knowledge, as opposed to OS-level or hardware-level inferred knowledge, Chameleon can substantially reduce CPU energy consumption. By exporting the energy management to user-space, designers can design more flexible and easily portable algorithms and systems, and use multiple energy management policies simultaneously. Specifically, we propose a minimal operating system interface that applications use to obtain global knowledge from the kernel in order to make local decisions. We consider three classes of applications soft real-time, interactive and batch and design user level power management strategies for representative applications such as a movie player, a word processor, a web browser, and a batch compiler. Our experiments show that, compared to the traditional system-wide CPU voltage scaling approaches, Chameleon can achieve up to 32-50% energy savings while delivering comparable or better performance to applications. Similarly, Chameleon extracts 9-41% more energy when compared to Grace OS, which uses some application knowledge but operates within the kernel. Further, Chameleon imposes minimal overhead and is effective at scheduling concurrent applications with diverse energy needs. Prashant J. Shenoy, Mark D. Corner |
IEEE Trans. Mob. Comput. | 2 |
| 2008 | Multimedia streaming via TCP: An analytic performance studyabstractTCP is widely used in commercial multimedia streaming systems, with recent measurement studies indicating that a significant fraction of Internet streaming media is currently delivered over HTTP/TCP. These observations motivate us to develop analytic performance models to systematically investigate the performance of TCP for both live and stored-media streaming. We validate our models via ns simulations and experiments conducted over the Internet. Our models provide guidelines indicating the circumstances under which TCP streaming leads to satisfactory performance, showing, for example, that TCP generally provides good streaming performance when the achievable TCP throughput is roughly twice the media bitrate, with only a few seconds of startup delay. Bing Wang 0001, James F. Kurose, Prashant J. Shenoy, Don Towsley |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2008 | Hierarchical Scheduling for Symmetric MultiprocessorsabstractHierarchical scheduling has been proposed as a scheduling technique to achieve aggregate resource partitioning among related groups of threads and applications in uniprocessor and packet scheduling environments. Existing hierarchical schedulers are not easily extensible to multiprocessor environments because 1) they do not incorporate the inherent parallelism of a multiprocessor system while resource partitioning and 2) they can result in unbounded unfairness or starvation if applied to a multiprocessor system in a naive manner. In this paper, we present hierarchical multiprocessor scheduling (H-SMP), a novel hierarchical CPU scheduling algorithm designed for a symmetric multiprocessor (SMP) platform. The novelty of this algorithm lies in its combination of space and time multiplexing to achieve the desired bandwidth partition among the nodes of the hierarchical scheduling tree. This algorithm is also characterized by its ability to incorporate existing proportional-share algorithms as auxiliary schedulers to achieve efficient hierarchical CPU partitioning. In addition, we present a generalized weight feasibility constraint that specifies the limit on the achievable CPU bandwidth partitioning in a multiprocessor hierarchical framework and propose a hierarchical weight readjustment algorithm designed to transparently satisfy this feasibility constraint. We evaluate the properties of H-SMP using hierarchical surplus fair scheduling (H-SFS), an instantiation of H-SMP that employs surplus fair scheduling (SFS) as an auxiliary algorithm. This evaluation is carried out through a simulation study that shows that H-SFS provides better fairness properties in multiprocessor environments as compared to existing algorithms and their naive extensions. Abhishek Chandra, Prashant J. Shenoy |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2007 | Rethinking Data Management for Storage-centric Sensor Networks
Yanlei Diao, Deepak Ganesan, Gaurav Mathur, Prashant J. Shenoy |
CIDR | 4 |
| 2007 | Approximate Initialization of Camera Sensor Networks
Purushottam Kulkarni, Prashant J. Shenoy, Deepak Ganesan |
EWSN | 2 |
| 2007 | Black-box and Gray-box Strategies for Virtual Machine Migration
Timothy Wood 0001, Prashant J. Shenoy, Arun Venkataramani, Mazin S. Yousif |
NSDI | 2 |
| 2007 | Multi-user data sharing in radar sensor networksabstractIn this paper, we focus on a network of rich sensors that are geographically distributed and argue that the design of such networks poses very different challenges from traditional mote-class sensor network design. We identify the need to handle the diverse requirements of multiple users to be a major design challenge, and propose a utility-driven approach to maximize data sharing across users while judiciously using limited network and computational resources. Our utility-driven architecture addresses three key challenges for such rich multi-user sensor networks: how to define utility functions for networks with data sharing among end-users, how to compress and prioritize data transmissions according to its importance to end-users, and how to gracefully degrade end-user utility in the presence of bandwidth fluctuations. We instantiate this architecture in the context of geographically distributed wireless radar sensor networks for weather, and present results from an implementation of our system on a multi-hop wireless mesh network that uses real radar data with real end-user applications. Our results demonstrate that our progressive compression and transmission approach achieves an order of magnitude improvement in application utility over existing utility-agnostic non-progressive approaches, while also scaling better with the number of nodes in the network. Ming Li 0009, Tingxin Yan, Deepak Ganesan, Eric Lyons 0001, Prashant J. Shenoy, Arun Venkataramani, Michael Zink |
SenSys | 5 |
| 2007 | Multi-user data sharing in radar sensor networksabstractThe emerging of rich sensor networks poses very different design challenges from traditional "mote-class" sensor networks. One important challenge is that these networks are designed to handle the diverse requirements of multiple users. In this work, we demonstrate how multiple end user needs are handled in rich sensor networks using a utility-driven architecture. We instantiate this architecture in the context of geographically distributed wireless radar sensor networks for weather, and demonstrate the real-time operation of the prototype on a radar testbed in Okalahoma. Ming Li 0009, Tingxin Yan, Deepak Ganesan, Eric Lyons 0001, Prashant J. Shenoy, Arun Venkataramani, Michael Zink |
SenSys | 5 |
| 2007 | Hyperion: High Volume Stream Archival for Retrospective Querying
Peter Desnoyers, Prashant J. Shenoy |
USENIX ATC | 2 |
| 2007 | Resource management for real-time tasks in mobile robotics
Huan Li 0001, Krithi Ramamritham, Prashant J. Shenoy, Roderic A. Grupen, John Sweeney |
J. Syst. Softw. | 3 |
| 2007 | Design and analysis of a demand adaptive and locality aware streaming media server cluster
Zihui Ge, Ping Ji 0002, Prashant J. Shenoy |
Multim. Syst. | 3 |
| 2007 | Introduction to the special issue on Multimedia Information Systems
Maria Luisa Sapino, Prashant J. Shenoy |
Multim. Tools Appl. | 2 |
| 2007 | Guest editorial
Marco Ajmone Marsan, Christoph Lindemann, Prashant J. Shenoy |
Perform. Evaluation | 3 |
| 2007 | Analytic modeling of multitier Internet applicationsabstractSince many Internet applications employ a multitier architecture, in this article, we focus on the problem of analytically modeling the behavior of such applications. We present a model based on a network of queues where the queues represent different tiers of the application. Our model is sufficiently general to capture (i) the behavior of tiers with significantly different performance characteristics and (ii) application idiosyncrasies such as session-based workloads, tier replication, load imbalances across replicas, and caching at intermediate tiers. We validate our model using real multitier applications running on a Linux server cluster. Our experiments indicate that our model faithfully captures the performance of these applications for a number of workloads and configurations. Furthermore, our model successfully handles a comprehensive range of resource utilization---from 0 to near saturation for the CPU---for two separate tiers. For a variety of scenarios, including those with caching at one of the application tiers, the average response times predicted by our model were within the 95% confidence intervals of the observed average response times. Our experiments also demonstrate the utility of the model for dynamic capacity provisioning, performance prediction, bottleneck identification, and session policing. In one scenario, where the request arrival rate increased from less than 1500 to nearly 4200 requests/minute, a dynamic provisioning technique employing our model was able to maintain response time targets by increasing the capacity of two of the tiers by factors of 2 and 3.5, respectively. Bhuvan Urgaonkar, Giovanni Pacifici, Prashant J. Shenoy, Mike Spreitzer, Asser N. Tantawi |
ACM Trans. Web | 3 |
| 2006 | Snapshot: A Self-Calibration Protocol for Camera Sensor NetworksabstractA camera sensor network is a wireless network of cameras designed for ad-hoc deployment. The camera sensors in such a network need to be properly calibrated by determining their location, orientation, and range. This paper presents Snapshot, an automated calibration protocol that is explicitly designed and optimized for camera sensor networks. Snapshot uses the inherent imaging abilities of the cameras themselves for calibration and can determine the location and orientation of a camera sensor using only four reference points. Our techniques draw upon principles from computer vision, optics, and geometry and are designed to work with low-fidelity, low-power camera sensors that are typical in sensor networks. An experimental evaluation of our prototype implementation shows that Snapshot yields an error of 1-2.5 degrees when determining the camera orientation and 5-10cm when determining the camera location. We show that this is a tolerable error in practice since a Snapshot-calibrated sensor network can track moving objects to within 11cm of their actual locations. Finally, our measurements indicate that Snapshot can calibrate a camera sensor within 20 seconds, enabling it to calibrate a sensor network containing tens of cameras within minutes. Purushottam Kulkarni, Prashant J. Shenoy, Deepak Ganesan |
BROADNETS | 3 |
| 2006 | Ferret: RFID Localization for Pervasive Multimedia
Mark D. Corner, Prashant J. Shenoy |
UbiComp | 3 |
| 2006 | Ultra-low power data storage for sensor networksabstractLocal storage is required in many sensor network applications, both for archival of detailed event information, as well as to overcome sensor platform memory constraints. While extensive measurement studies have been performed to highlight the trade-off between computation and communication in sensor networks, the role of storage has received little attention. The storage subsystems on currently available sensor platforms have not exploited technology trends, and consequently the energy cost of storage on these platforms is as high as that of communication. Current flash memories, however, offer a low-priced, high-capacity and extremely energy-efficient storage solution.In this paper, we perform a comprehensive evaluation of the active and sleep-mode energy consumption of available flash-based storage options for sensor platforms. Our results demonstrate more than a 100-fold decrease in per-byte energy consumption for surface-mount parallel NAND flash in comparison with the MicaZ on-board serial flash. In addition, this dramatically reduces storage energy costs relative to communication, introducing a new dimension in traditional computation vs communication trade-offs. Our results have significant ramifications on the design of sensor platforms as well as on the energy consumption of sensing applications. We quantify the potential energy gains for two commonly used sensor network services: communication and in-network data aggregation. Our measurements show significant improvements in each service: 50-fold and up to 10-fold reductions in energy for communication and data aggregation respectively. Gaurav Mathur, Peter Desnoyers, Deepak Ganesan, Prashant J. Shenoy |
IPSN | 4 |
| 2006 | PRESTO: Feedback-driven Data Management in Sensor Networks
Ming Li 0009, Deepak Ganesan, Prashant J. Shenoy |
NSDI | 3 |
| 2006 | A storage-centric camera sensor networkabstractImproved energy-efficiency and storage capacity of new-generation NAND flash memory makes a compelling case for storage-centric sensor networks. Such a storage-centric sensor network emphasizes the use of platforms with larger storage and more extensive use of the storage capacities on sensors. We demonstrate the feasibility of storage-centric sensor networks using an instance of a storage-centric camera sensor network that is more energy-efficient in comparison to a traditional camera sensor network. We demonstrate multiple camera sensors, each consisting of a Cyclops camera attached to a MicaZ mote, using motion-triggered image capturing. The captured images are archived locally on flash storage and summaries of detected events are transmitted to the base-station. The base-station picks the events of interest from the summaries and requests the original captured image from the sensor as required. The use of high-capacity energy-efficient flash storage at the sensor allows us to trade-off expensive radio communication for cheaper local storage, improving the life-time of the battery and consequently, the life of the storage-centric camera sensor network. Gaurav Mathur, Paul Chukiu, Peter Desnoyers, Deepak Ganesan, Prashant J. Shenoy |
SenSys | 5 |
| 2006 | Capsule: an energy-optimized object storage system for memory-constrained sensor devicesabstractRecent gains in energy-efficiency of new-generation NAND flash storage have strengthened the case for in-network storage by data-centric sensor network applications. This paper argues that a simple file system abstraction is inadequate for realizing the full benefits of high-capacity lowpower NAND flash storage in data-centric applications. Instead we advocate a rich object storage abstraction to support flexible use of the storage system for a variety of application needs and one that is specifically optimized for memory and energy-constrained sensor platforms. We propose Capsule, an energy-optimized log-structured object storage system for flash memories that enables sensor applications to exploit storage resources in a multitude of ways. Capsule employs a hardware abstraction layer that hides the vagaries of flash memories for the application and supports energy-optimized implementations of commonly used storage objects such as streams, files, arrays, queues and lists. Further, Capsule supports checkpointing and rollback of object states to tolerate software faults in sensor applications running on inexpensive, unreliable hardware. Our experiments demonstrate that Capsule provides platform-independence, greater functionality, more tunability, and greater energy-efficiency than existing sensor storage solutions, while operating even within the memory constraints of the Mica2 Mote. Our experiments not only demonstrate the energy and memory-efficiency of I/O operations in Capsule but also shows that Capsule consumes less than 15% of the total energy cost in a typical sensor application. Gaurav Mathur, Peter Desnoyers, Deepak Ganesan, Prashant J. Shenoy |
SenSys | 4 |
| 2006 | Consistency maintenance in dynamic peer-to-peer overlay networks
Jiang Lan, Prashant J. Shenoy, Krithi Ramamritham |
Comput. Networks | 3 |
| 2006 | An observation-based approach towards self-managing web servers
Abhishek Chandra, Prashant Pradhan, Renu Tewari, Sambit Sahu, Prashant J. Shenoy |
Comput. Commun. | 5 |
| 2006 | Dynamic cache reconfiguration strategies for cluster-based streaming proxy
Yang Guo 0001, Zihui Ge, Bhuvan Urgaonkar, Prashant J. Shenoy, Don Towsley |
Comput. Commun. | 4 |
| 2005 | PRESTO: A Predictive Storage Architecture for Sensor Networks
Peter Desnoyers, Deepak Ganesan, Huan Li 0001, Ming Li 0009, Prashant J. Shenoy |
HotOS | 5 |
| 2005 | SensEye: a multi-tier camera sensor networkabstractThis paper argues that a camera sensor network containing heterogeneous elements provides numerous benefits over traditional homogeneous sensor networks. We present the design and implementation of senseye---a multi-tier network of heterogeneous wireless nodes and cameras. To demonstrate its benefits, we implement a surveillance application using senseye comprising three tasks: object detection, recognition and tracking. We propose novel mechanisms for low-power low-latency detection, low-latency wakeups, efficient recognition and tracking. Our techniques show that a multi-tier sensor network can reconcile the traditionally conflicting systems goals of latency and energy-efficiency. An experimental evaluation of our prototype shows that, when compared to a single-tier prototype, our multi-tier senseye can achieve an order of magnitude reduction in energy usage while providing comparable surveillance accuracy. Purushottam Kulkarni, Deepak Ganesan, Prashant J. Shenoy, Qifeng Lu |
ACM Multimedia | 3 |
| 2005 | SEVA: sensor-enhanced video annotationabstractIn this paper, we study how a sensor-rich world can be exploited by digital recording devices such as cameras and camcorders to improve a user's ability to search through a large repository of image and video files. We design and implement a digital recording system that records identities and locations of objects (as advertised by their sensors) along with visual images (as recorded by a camera). The process, which we refer to as sensor-enhanced video annotation (SEVA), combines a series of correlation, interpolation, and extrapolation techniques. It produces a tagged stream that later can be used to efficiently search for videos or frames containing particular objects or people. We present detailed experiments with a prototype of our system using both stationary and mobile objects as well as GPS and ultrasound. Our experiments show that: (i) SEVA has zero error rates for static objects, except very close to the boundary of the viewable area; (ii) for moving objects or a moving camera, SEVA only misses objects leaving or entering the viewable area by 1-2 frames; (iii) SEVA can scale to 10 fast moving objects using current sensor technology; and (iv) SEVA runs online using relatively inexpensive hardware. Mark D. Corner, Prashant J. Shenoy |
ACM Multimedia | 3 |
| 2005 | Chameleon: application level power management with performance isolationabstractIn this paper, we present Chameleon---an application-level power management approach for reducing energy consumption in mobile processors. Our approach exports the entire responsibility of power management decisions to the application level. We propose an operating system interface that can be used by applications to achieve energy savings. We consider three classes of applications---soft real-time, interactive and batch---and design user-level power management strategies for representative applications such as a movie player, a word processor, a web browser, and a batch compiler. We also design a user level power manager based on GraceOS using Chameleon. We implement our approach in the Linux kernel running on a Sony Transmeta laptop. Our experiments show that, compared to the traditional system-wide CPU voltage scaling approaches, Chameleon can achieve up to 32-50% energy savings while delivering comparable or better performance to applications. Further, Chameleon imposes small overheads and is very effective at scheduling concurrent applications with diverse energy needs. Prashant J. Shenoy, Mark D. Corner |
ACM Multimedia | 2 |
| 2005 | Online scheduling in modular multimedia systems with stream reuseabstractWhen properly constructed, a modular multimedia system can satisfy a client's request in multiple ways by using different sequences of modules and by reusing existing streams within the system. Such flexibility in a multimedia server or proxy can provide a rich set of services to clients while efficiently utilizing system resources by choosing the best way to schedule (i.e. allocate resources for) new clients. However, it is difficult to optimally schedule clients in an online fashion as the problem is NP-complete. In this paper, we provide an efficient online algorithm to schedule client requests using a modular multimedia platform. Michael K. Bradshaw, James F. Kurose, Prashant J. Shenoy, Don Towsley |
NOSSDAV | 3 |
| 2005 | The case for multi-tier camera sensor networksabstractIn this position paper, we examine recent technology trends that have resulted in a broad spectrum of camera sensors, wireless radio technologies, and embedded sensor platforms with varying capabilities. We argue that future sensor applications will be hierarchical with multiple tiers, where each tier employs sensors with different characteristics. We argue that multi-tier networks are not only scalable, they offer a number of advantages over simpler, single-tier unimodal networks: lower cost, better coverage, higher functionality, and better reliability. However, the design of such mixed networks raises a number of new challenges that are not adequately addressed by current research. We discuss several of these challenges and illustrate how they can be addressed in the context of SensEye, a multi-tier video surveillance application that we are designing in our research group. Purushottam Kulkarni, Deepak Ganesan, Prashant J. Shenoy |
NOSSDAV | 3 |
| 2005 | Scheduling Messages with Deadlines in Multi-Hop Real-Time Sensor NetworksabstractConsider a team of robots equipped with sensors that collaborate with one another to achieve a common goal. Sensors on robots produce periodic updates that must be transmitted to other robots and processed in real-time to enable such collaboration. Since the robots communicate with one another over an ad-hoc wireless network, we consider the problem of providing timeliness guarantees for multihop message transmissions in such a network. We derive the effective deadline and the latest start time for per-hop message transmissions from the validity intervals of the sensor data and the constraints imposed by the consuming task at the destination. Our technique schedules messages by carefully exploiting spatial channel reuse for each per-hop transmission to avoid MAC layer collisions, so that deadline misses are minimized. Extensive simulations show the effectiveness of our channel reuse-based SLF (smallest latest-start-time first) technique when compared to a simple per-hop SLF technique, especially at moderate to high channel utilization or when the probability of collisions is high. Huan Li 0001, Prashant J. Shenoy, Krithi Ramamritham |
IEEE Real-Time and Embedded Technology and Applications Symposium | 2 |
| 2005 | TSAR: a two tier sensor storage architecture using interval skip graphsabstractArchival storage of sensor data is necessary for applications that query, mine, and analyze such data for interesting features and trends. We argue that existing storage systems are designed primarily for flat hierarchies of homogeneous sensor nodes and do not fully exploit the multi-tier nature of emerging sensor networks, where an application can comprise tens of tethered proxies, each managing tens to hundreds of untethered sensors. We present TSAR, a fundamentally different storage architecture that envisions separation of data from metadata by employing local archiving at the sensors and distributed indexing at the proxies. At the proxy tier, TSAR employs a novel multi-resolution ordered distributed index structure, the Interval Skip Graph, for efficiently supporting spatio-temporal and value queries. At the sensor tier,TSAR supports energy-aware adaptive summarization that can trade off the cost of transmitting metadata to the proxies against the overhead of false hits resulting from querying a coarse-grain index. We implement TSAR in a two-tier sensor testbed comprising Stargate-based proxies and Mote-based sensors. Our experiments demonstrate the benefits and feasibility of using our energy-efficient storage architecture in multi-tier sensor networks. Peter Desnoyers, Deepak Ganesan, Prashant J. Shenoy |
SenSys | 3 |
| 2005 | An analytical model for multi-tier internet services and its applicationsabstractSince many Internet applications employ a multi-tier architecture, in this paper, we focus on the problem of analytically modeling the behavior of such applications. We present a model based on a network of queues, where the queues represent different tiers of the application. Our model is sufficiently general to capture (i) the behavior of tiers with significantly different performance characteristics and (ii) application idiosyncrasies such as session-based workloads, concurrency limits, and caching at intermediate tiers. We validate our model using real multi-tier applications running on a Linux server cluster. Our experiments indicate that our model faithfully captures the performance of these applications for a number of workloads and configurations. For a variety of scenarios, including those with caching at one of the application tiers, the average response times predicted by our model were within the 95% confidence intervals of the observed average response times. Our experiments also demonstrate the utility of the model for dynamic capacity provisioning, performance prediction, bottleneck identification, and session policing. In one scenario, where the request arrival rate increased from less than 1500 to nearly 4200 requests/min, a dynamic provisioning technique employing our model was able to maintain response time targets by increasing the capacity of two of the application tiers by factors of 2 and 3.5, respectively. Bhuvan Urgaonkar, Giovanni Pacifici, Prashant J. Shenoy, Mike Spreitzer, Asser N. Tantawi |
SIGMETRICS | 3 |
| 2005 | Cataclysm: policing extreme overloads in internet applicationsabstractIn this paper we present the Cataclysm server platform for handling extreme overloads in hosted Internet applications. The primary contribution of our work is to develop a low overhead, highly scalable admission control technique for Internet applications. Cataclysm provides several desirable features, such as guarantees on response time by conducting accurate size-based admission control, revenue maximization at multiple time-scales via preferential admission of important requests and dynamic capacity provisioning, and the ability to be operational even under extreme overloads. Cataclysm can transparently trade-off the accuracy of its decision making with the intensity of the workload allowing it to handle incoming rates of several tens of thousands of requests/second. We implement a prototype Cataclysm hosting platform on a Linux cluster and demonstrate the benefits of our integrated approach using a variety of workloads. Bhuvan Urgaonkar, Prashant J. Shenoy |
WWW | 2 |
| 2005 | Selected papers from the ACM multimedia conference 2003abstractNo abstract available. Thomas Plagemann, Prashant J. Shenoy, John R. Smith |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2004 | A Performance Comparison of NFS and iSCSI for IP-Networked Storage
Peter Radkov, Pawan Goyal 0001, Prasenjit Sarkar, Prashant J. Shenoy |
FAST | 5 |
| 2004 | Multimedia streaming via TCP: an analytic performance studyabstractTCP is widely used in commercial media streaming systems, with recent measurement studies indicating that a significant fraction of Internet streaming media is currently delivered over HTTP/TCP. These observations motivate us to develop analytic performance models to systematically investigate the performance of TCP for both live and stored media streaming. We validate our models via ns simulations and experiments conducted over the Internet. Our models provide guidelines indicating the circumstances under which TCP streaming leads to satisfactory performance, showing, for example, that TCP generally provides good streaming performance when the achievable TCP throughput is roughly twice the media bitrate, with only a few seconds of startup delay. Bing Wang 0001, James F. Kurose, Prashant J. Shenoy, Don Towsley |
ACM Multimedia | 3 |
| 2004 | A time series-based approach for power management in mobile processors and disksabstractIn this paper, we present a time series-based approach for managing power in mobile processors and disks that see multimedia workloads. Since multimedia applications impose soft real-time constraints, a key goal of our approach is to reduce energy consumption of multimedia applications without degrading performance. We present simple statistical techniques based on time series to dynamically compute the processor and I/O demands of multimedia applications and present techniques to dynamically vary the voltage settings and rotational speeds of mobile processors and disks, respectively. We implement our approaches in the Linux kernel running on a Sony Transmeta laptop and in a trace-driven simulator. Our experiments show that, compared to the traditional system-wide CPU voltage scaling approaches, our technique can achieve up to a 38.6% energy saving while delivering good performance to applications. Simulation results for our disk power management technique show a 20.3% reduction in energy consumption without any significant performance loss when compared to a traditional disk power management scheme. Prashant J. Shenoy, Weibo Gong |
NOSSDAV | 2 |
| 2004 | AMPS: a flexible, scalable proxy testbed for implementing streaming servicesabstractWe present the design, implementation, and performance evaluation of AMPS --- a flexible, scalable proxy testbed that supports a wide and extensible set of next-generation proxy streaming services. AMPS employs a modular architecture and is built on top of a commodity Linux system. We study the performance of AMPS proxy using a server-proxy-client configuration in a switched-Gigabit LAN environment. We identify the CPU to be the system bottleneck. Through profiling study, we further identify the kernel network protocol processing and the Network Reception Module inside the proxy to be the most CPU-intensive components. We also quantify the maximum achievable throughput for two of the principal components of the proxy - the control plane and data plane, and characterize the end-to-end performance along the server-to-proxy-to-client path. We discuss lessons learned and the various optimizations made in the course of our study to improve system performance. Xiaolan Zhang 0003, Michael K. Bradshaw, Yang Guo 0001, Bing Wang 0001, James F. Kurose, Prashant J. Shenoy, Don Towsley |
NOSSDAV | 6 |
| 2004 | Brief announcement: Cataclysm: handling extreme overloads in internet servicesabstractIn this paper we present Cataclysm, a comprehensive approach for handling extreme overloads in hosted Internet applications. The primary contribution of our work is to develop an overload control approach that brings together admission control, dynamic provisioning of platform resources, and adaptive degradation of QoS into one integrated system. We implement a prototype Cataclysm hosting platform on a Linux cluster and demonstrate the benefits of our integrated approach using a variety of workloads. Bhuvan Urgaonkar, Prashant J. Shenoy |
PODC | 2 |
| 2004 | Scheduling Communication in Real-Time Sensor ApplicationsabstractWe consider a class of wireless sensor applications - such as mobile robotics - that impose timeliness constraints. We assume that these applications are built using commodity 802.11 wireless networks and focus on the problem of providing qualitatively-better QoS during network transmission of sensor data. Our techniques are designed to explicitly avoid network collisions and minimize the completion time to transmit a set of sensor messages. We argue that this problem is NP-complete and present three heuristics, based on edge coloring, to achieve these goals. Our simulations results show that the minimum weight color heuristic is robust to increases in communication density and yields results that are close to the optimal solution. Huan Li 0001, Prashant J. Shenoy, Krithi Ramamritham |
IEEE Real-Time and Embedded Technology and Applications Symposium | 2 |
| 2004 | Multimedia streaming via TCP: an analytic performance studyabstractTCP is widely used in commercial media streaming systems, with recent measurement studies indicating that a significant fraction of Internet streaming media is currently delivered over HTTP/TCP. These observations motivate us to develop analytic performance models to systematically investigate the performance of TCP for both live and stored media streaming. We validate our models via ns simulations and experiments conducted over the Internet. Our models provide guidelines indicating the circumstances under which TCP streaming leads to satisfactory performance, showing, for example, that TCP generally provides good streaming performance when the achievable TCP throughput is roughly twice the media bitrate, with only a few seconds of startup delay. Bing Wang 0001, James F. Kurose, Prashant J. Shenoy, Don Towsley |
SIGMETRICS | 3 |
| 2004 | Resilient and Coherence Preserving Dissemination of Dynamic Data Using Cooperating PeersabstractThe focus of our work is to design and build a dynamic data distribution system that is coherence-preserving, i.e. the delivered data must preserve associated coherence requirements (the user-specified bound on tolerable imprecision) and resilient to failures. To this end, we consider a system in which a set of repositories cooperate with each other and the sources, forming a peer-to-peer network. In this system, necessary changes are pushed to the users so that they are automatically informed, about changes of interest. We present techniques 1) to determine when to push an update from one repository to another for coherence maintenance, 2) to construct an efficient dissemination tree for propagating changes from sources to cooperating repositories, and 3) to make the system resilient to failures. An experimental evaluation using real world traces of dynamically changing data demonstrates that 1) careful dissemination of updates through a network of cooperating repositories can substantially lower the cost of coherence maintenance, 2) unless designed carefully, even push-based systems experience considerable loss in fidelity due to message delays and processing costs, 3) the computational and communication cost of achieving resiliency is made to be low, and 4) surprisingly, adding resiliency actually improve fidelity even in the absence of failures. Shetal Shah, Krithi Ramamritham, Prashant J. Shenoy |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2004 | Sharc: Managing CPU and Network Bandwidth in Shared ClustersabstractWe argue the need for effective resource management mechanisms for sharing resources in commodity clusters. To address this issue, we present the design of Sharc-a system that enables resource sharing among applications in such clusters. Sharc depends on single node resource management mechanisms such as reservations or shares, and extends the benefits of such mechanisms to clustered environments. We present techniques for managing two important resources-CPU and network interface bandwidth-on a cluster-wide basis. Our techniques allow Sharc to 1) support reservation of CPU and network interface bandwidth for distributed applications, 2) dynamically allocate resources based on past usage, and 3) provide performance isolation to applications. Our experimental evaluation has shown that Sharc can scale to 256 node clusters running 100,000 applications. These results demonstrate that Sharc can be an effective approach for sharing resources among competing applications in moderate size clusters. Bhuvan Urgaonkar, Prashant J. Shenoy |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2004 | PTC: Proxies that Transcode and Cache in Heterogeneous Web Client Environments
Aameek Singh, Abhishek Trivedi, Krithi Ramamritham, Prashant J. Shenoy |
World Wide Web | 4 |
| 2003 | Dynamic Resource Allocation for Shared Data Centers Using Online Measurements
Abhishek Chandra, Weibo Gong, Prashant J. Shenoy |
IWQoS | 3 |
| 2003 | A Practical Learning-Based Approach for Dynamic Storage Bandwidth Allocation
Vijay Sundaram, Prashant J. Shenoy |
IWQoS | 2 |
| 2003 | Handling Client Mobility and Intermittent Connectivity in Mobile Web Accesses
Purushottam Kulkarni, Prashant J. Shenoy, Krithi Ramamritham |
Mobile Data Management | 2 |
| 2003 | Dynamic resource allocation for shared data centers using online measurementsabstractNo abstract available. Abhishek Chandra, Weibo Gong, Prashant J. Shenoy |
SIGMETRICS | 3 |
| 2003 | Scalable techniques for memory-efficient CDN simulationsabstractSince CDN simulations are known to be highly memory-intensive, in this paper, we argue the need for reducing the memory requirements of such simulations. We propose a novel memory-efficient data structure that stores cache state for a small subset of popular objects accurately and uses approximations for storing the state for the remaining objects. Since popular objects receive a large fraction of the requests while less frequently accessed objects consume much of the memory space, this approach yields large memory savings and reduces errors. We use bloom filters to store approximate state and show that careful choice of parameters can substantially reduce the probability of errors due to approximations. We implement our techniques into a user library for constructing proxy caches in CDN simulators. Our experimental results show up to an order of magnitude reduction in memory requirements of CDN simulations, while incurring a 5-10% error. Purushottam Kulkarni, Prashant J. Shenoy, Weibo Gong |
WWW | 2 |
| 2003 | Periodic broadcast and patching services - implementation, measurement and analysis in an internet streaming video testbed
Michael K. Bradshaw, Bing Wang 0001, Subhabrata Sen, Lixin Gao 0001, James F. Kurose, Prashant J. Shenoy, Don Towsley |
Multim. Syst. | 6 |
| 2003 | Design considerations for the symphony integrated multimedia file system
Prashant J. Shenoy, Pawan Goyal 0001, Sriram Rao, Harrick M. Vin |
Multim. Syst. | 1 |
| 2003 | Introduction to selected papers from MMCN 2002
Prashant J. Shenoy, Ketan Mayer-Patel |
Multim. Syst. | 1 |
| 2003 | Adaptive Leases: A Strong Consistency Mechanism for the World Wide WebabstractWe argue that weak cache consistency mechanisms supported by existing Web proxy caches must be augmented by strong consistency mechanisms to support the growing diversity in application requirements. Existing strong consistency mechanisms are not appealing for Web environments due to their large state space or control message overhead. We focus on the lease approach that balances these trade-offs and present analytical models and policies for determining the optimal lease duration. We present extensions to the HTTP protocol to incorporate leases and then implement our techniques in the Squid proxy cache and the Apache Web server. Our experimental evaluation of the leases approach shows that: 1) our techniques impose modest overheads even for long leases (a lease duration of 1 hour requires state to be maintained for 1030 leases and imposes an per-object overhead of a control message every 33 minutes), 2) leases yields a 138-425 percent improvement over existing strong consistency mechanisms, and 3) the implementation overhead of leases is comparable to existing weak consistency mechanisms. Venkata Duvvuri, Prashant J. Shenoy, Renu Tewari |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2003 | Scalable Consistency Maintenance in Content Distribution Networks Using Cooperative LeasesabstractWe argue that cache consistency mechanisms designed for stand-alone proxies do not scale to the large number of proxies in a content distribution network and are not flexible enough to allow consistency guarantees to be tailored to object needs. To meet the twin challenges of scalability and flexibility, we introduce the notion of cooperative consistency along with a mechanism, called cooperative leases, to achieve it. By supporting /spl Delta/-consistency semantics and by using a single lease for multiple proxies, cooperative leases allow the notion of leases to be applied in a flexible, scalable manner to CDNs. Further, the approach employs application-level multicast to propagate server notifications to proxies in a scalable manner. We implement our approach in the Apache Web server and the Squid proxy cache and demonstrate its efficacy using a detailed experimental evaluation. Our results show a factor of 2.5 reduction in server message overhead and a 20 percent reduction in server state space overhead when compared to original leases albeit at an increased interproxy communication overhead. Anoop George Ninan, Purushottam Kulkarni, Prashant J. Shenoy, Krithi Ramamritham, Renu Tewari |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2002 | A Demand Adaptive and Locality Aware (DALA) streaming media server cluster architectureabstractThe wide availability of broadband networking technologies such as cable modems and DSL coupled with the growing popularity of the Internet has led to a dramatic increase in the availability and the use of online streaming media. With the "last mile" network bandwidth no longer a constraint, the bottleneck for video streaming has been pushed closer to the server. Streaming high quality audio and video to a myriad of clients imposes significant resource demands on the server. In this work, we propose a demand adaptive and locality aware (DALA) clustered media server architecture that can dynamically allocate resources to adapt to changing demand and also maximize the number of clients serviced by the server cluster. Moreover, our design exploits temporal locality among requests by dispatching newly arriving requests to servers that are already servicing prior requests for those objects, thereby extracting the benefits of locality. We explore the efficacy of the DALA clustered architecture using simulations. Our simulation results show that DALA is highly adaptive, exhibits significant performance gains when compared to static schemes, and has a low system overhead. Our results demonstrate that DALA is a simple, yet effective approach for designing clustered media servers. Zihui Ge, Ping Ji 0002, Prashant J. Shenoy |
NOSSDAV | 3 |
| 2002 | Resource Overbooking and Application Profiling in Shared Hosting Platforms
Bhuvan Urgaonkar, Prashant J. Shenoy, Timothy Roscoe |
OSDI | 2 |
| 2002 | Maintaining Coherency of Dynamic Data in Cooperating Repositories
Shetal Shah, Krithi Ramamritham, Prashant J. Shenoy |
VLDB | 3 |
| 2002 | PTC : Proxies that Transcode and Cache in Heterogeneous Web Client EnvironmentsabstractAdvances in computing and communication technologies have resulted in a wide variety of networked mobile devices that access data over the Internet. We argue that servers by themselves may not be able to handle this diversity in client characteristics and intermediate proxies should be employed to handle the mismatch between server-supplied data and client capabilities. Since existing proxies are primarily designed to handle traditional wired hosts, such proxy architectures will need to be enhanced to handle mobile devices. We propose such an enhanced proxy architecture that is capable of handling the heterogeneity in client needs - specifically the variations in client bandwidth and display capabilities. Our architecture combines transcoding (which is used to match the fidelity of the requested object to client capabilities) and caching (which is used to reduce the latency for accessing popular objects). Our proxies can intelligently adapt to prevailing system conditions using learning techniques to intelligently decide whether to transcode locally or fetch an appropriate version from the server. Our experimental results indicate that such strategies produce significant improvements in client response times. Further we find that even simple learning techniques can lead to significant performance improvements. Aameek Singh, Abhishek Trivedi, Krithi Ramamritham, Prashant J. Shenoy |
WISE | 4 |
| 2002 | Cooperative leases: scalable consistency maintenance in content distribution networksabstractIn this paper, we argue that cache consistency mechanisms designed for stand-alone proxies do not scale to the large number of proxies in a content distribution network and are not flexible enough to allow consistency guarantees to be tailored to object needs. To meet the twin challenges of scalability and flexibility, we introduce the notion of cooperative consistency along with a mechanism, called cooperative leases, to achieve it. By supporting Δ-consistency semantics and by using a single lease for multiple proxies, cooperative leases allows the notion of leases to be applied in a flexible, scalable manner to CDNs. Further, the approach employs application-level multicast to propagate server notifications to proxies in a scalable manner. We implement our approach in the Apache web server and the Squid proxy cache and demonstrate its efficacy using a detailed experimental evaluation. Our results show a factor of 2.5 reduction in server message overhead and a 20% reduction in server state space overhead when compared to original leases albeit at an increased inter-proxy communication overhead. Anoop George Ninan, Purushottam Kulkarni, Prashant J. Shenoy, Krithi Ramamritham, Renu Tewari |
WWW | 3 |
| 2002 | Implications of proxy caching for provisioning networks and serversabstractIn this paper, we examine the potential benefits of Web proxy caches in improving the effective capacity of servers and networks. Since networks and servers are typically provisioned based on a high percentile of the load, we focus on the effects of proxy caching on the tail of the load distribution. We find that, unlike their substantial impact on the average load, proxies have a diminished impact on the tail of the load distribution. The exact reduction in the tail and the corresponding capacity savings depend on the nature of the workload and the percentile of the load distribution chosen for provisioning networks and servers-the higher the percentile, the smaller the savings. For workloads considered in this study, compared with over a 50% reduction in the average load, the savings in network and server capacity was only 20%-35% for the 99th percentile of the load distribution. We also find that while proxies can be somewhat useful in smoothing out some of the burstiness in Web workloads; the resulting workload continues, however, to exhibit substantial burstiness and a heavy-tailed nature. We identify one-time requests for large objects to be the limiting factor that diminishes the impact of proxies on the tail of load distribution. We conclude that, while proxies are immensely useful to users due to the reduction in the average response time, they are less effective in improving the capacities of networks and servers. M. S. Raunak 0001, Prashant J. Shenoy, Pawan Goyal 0001, Krithi Ramamritham, Purushottam Kulkarni |
IEEE J. Sel. Areas Commun. | 2 |
| 2002 | Architectural considerations for next-generation file systems
Prashant J. Shenoy, Pawan Goyal 0001, Harrick M. Vin |
Multim. Syst. | 1 |
| 2002 | Cello: A Disk Scheduling Framework for Next Generation Operating Systems
Prashant J. Shenoy, Harrick M. Vin |
Real Time Syst. | 1 |
| 2002 | Adaptive Push-Pull: Disseminating Dynamic Web DataabstractAn important issue in the dissemination of time-varying Web data such as sports scores and stock prices is the maintenance of temporal coherency. In the case of servers adhering to the HTTP protocol, clients need to frequently pull the data based on the dynamics of the data and a user's coherency requirements. In contrast, servers that possess push capability maintain state information pertaining to clients and push only those changes that are of interest to a user. These two canonical techniques have complementary properties with respect to the level of temporal coherency maintained, communication overheads, state space overheads, and loss of coherency due to (server) failures. In this paper, we show how to combine push and pull-based techniques to achieve the best features of both approaches. Our combined technique tailors the dissemination of data from servers to clients based on 1) the capabilities and load at servers and proxies and 2) clients' coherency requirements. Our experimental results demonstrate that such adaptive data dissemination is essential to meet diverse temporal coherency requirements, to be resilient to failures, and for the efficient and scalable utilization of server and network resources. Manish Bhide, Pavan Deolasee, Amol Katkar, Ankur Panchbudhe, Krithi Ramamritham, Prashant J. Shenoy |
IEEE Trans. Computers | 6 |
| 2001 | Maintaining Mutual Consistency for Cached Web ObjectsabstractExisting Web proxy caches employ cache consistency mechanisms to ensure that locally cached data is consistent with that at the server. We argue that techniques for maintaining consistency of individual objects are not sufficient; a proxy should employ additional mechanisms to ensure that related Web objects are mutually consistent with one another. We formally define the notion of mutual consistency and the semantics provided by a mutual consistency mechanism to end users. We then present techniques for maintaining mutual consistency in the temporal and value domains. A novel aspect of our techniques is that they can adapt to the variations in the rate of change of the source data, resulting in judicious use of proxy and network resources. We evaluate our approaches using real-world Web traces and show that: (i) careful tuning can result in substantial savings in the network overhead incurred without any substantial loss in fidelity, of the consistency guarantees, and (ii) the incremental cost of providing mutual consistency guarantees over mechanisms to provide individual consistency guarantees is small. Bhuvan Urgaonkar, Anoop George Ninan, M. S. Raunak 0001, Prashant J. Shenoy, Krithi Ramamritham |
ICDCS | 4 |
| 2001 | Periodic broadcast and patching services: implementation, measurement, and analysis in an internet streaming video testbedabstractMultimedia streaming applications can consume a significant amount of server and network resources. Periodic broadcast and patching are two approaches that use multicast transmission and client buffering in innovative ways to reduce server and network load, while at the same time allowing asynchronous access to multimedia steams by a large number of clients. Current research in this area has focussed primarily on the algorithmic aspects of these approaches, with evaluation performed via analysis or simulation. In this paper, we describe the design and implementation of a flexible streaming video server and client testbed that implements both periodic broadcast and patching, and explore the issues that arise when implementing these algorithms. We present measurements detailing the overheads associated with the various server components (signaling, transmission schedule computation, data retrieval and transmission), the interactions between the various components of the architecture, and the overall end-to-end performance. We also discuss the importance of an appropriate server video segment caching policy. We conclude with a discussion of the insights gained from our implementation and experimental evaluation. Michael K. Bradshaw, Bing Wang 0001, Lixin Gao 0001, James F. Kurose, Prashant J. Shenoy, Don Towsley, Subhabrata Sen |
ACM Multimedia | 5 |
| 2001 | Periodic broadcast and patching services: implementation, measurement, and analysis in an internet streaming video testbedabstractNo abstract available. Michael K. Bradshaw, Bing Wang 0001, Subhabrata Sen, Lixin Gao 0001, James F. Kurose, Prashant J. Shenoy, Don Towsley |
ACM Multimedia | 6 |
| 2001 | Bandwidth allocation in a self-managing multimedia file serverabstractIn this paper, we argue that manageability of file servers is just as important, if not more, as performance. We focus on the design of a self-managing file server and address the specific problem of automating bandwidth allocation to application classes in single-disk and multi-disk servers. The bandwidth allocation techniques that we propose consists of two key components: a workload monitoring module that efficiently monitors the load in each application class and a bandwidth manager that uses these workload statistics to dynamically determine the allocation of each class. We evaluate the efficacy of our techniques via a simulation study and demonstrate that our techniques (i) exploit the semantics of each application class while determining their allocations, (ii) provide control over the time-scale of monitoring and allocation, and (iii) provide stable behavior even during transient overloads. Our comparison with a static allocation technique shows that dynamic bandwidth allocation can yield queue lengths that are 59% smaller during overloads and admit a larger number of soft real-time clients into the system. Vijay Sundaram, Prashant J. Shenoy |
ACM Multimedia | 2 |
| 2001 | Dissemination of Dynamic DataabstractNo abstract available. Pavan Deolasee, Amol Katkar, Ankur Panchbudhe, Krithi Ramamritham, Prashant J. Shenoy |
SIGMOD Conference | 5 |
| 2001 | Adaptive push-pull: disseminating dynamic web dataabstractArticle Share on Adaptive push-pull: disseminating dynamic web data Authors: Pavan Deolasee Department of Computer Science and Engineering., Indian Institute of Technology, Bombay, Mumbai, India 400076 Department of Computer Science and Engineering., Indian Institute of Technology, Bombay, Mumbai, India 400076View Profile , Amol Katkar Department of Computer Science and Engineering., Indian Institute of Technology Bombay, Mumbai, India 400076 Department of Computer Science and Engineering., Indian Institute of Technology Bombay, Mumbai, India 400076View Profile , Ankur Panchbudhe Department of Computer Science and Engineering., Indian Institute of Technology Bombay, Mumbai, India 400076 Department of Computer Science and Engineering., Indian Institute of Technology Bombay, Mumbai, India 400076View Profile , Krithi Ramamritham Department of Computer Science, University of Massachusetts, Amherst, MA Department of Computer Science, University of Massachusetts, Amherst, MAView Profile , Prashant Shenoy Department of Computer Science, University of Massachusetts, Amherst, MA Department of Computer Science, University of Massachusetts, Amherst, MAView Profile Authors Info & Claims WWW '01: Proceedings of the 10th international conference on World Wide WebMay 2001Pages 265–274https://doi.org/10.1145/371920.372066Published:01 April 2001Publication History 88citation1,366DownloadsMetricsTotal Citations88Total Downloads1,366Last 12 Months25Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Pavan Deolasee, Amol Katkar, Ankur Panchbudhe, Krithi Ramamritham, Prashant J. Shenoy |
WWW | 5 |
| 2000 | Rules of Thumb in Data EngineeringabstractThis paper reexamines the rules of thumb for the design of data storage systems. Briefly, it looks at storage, processing, and networking costs, ratios, and trends with a particular focus on performance and price/performance. Amdahl's ratio laws for system design need only slight revision after 35 years-the major change being the increased use of RAM. An analysis also indicates storage should be used to cache both database and Web data to save disk bandwidth, network bandwidth, and people's time. Surprisingly, the 5-minute rule for disk caching becomes a cache-everything rule for Web caching. Jim Gray 0001, Prashant J. Shenoy |
ICDE | 2 |
| 2000 | Adaptive Leases: A Strong Consistency Mechanism for the World Wide WebabstractIn this paper, we argue that weak cache consistency mechanisms supported by existing Web proxy caches must be augmented by strong consistency mechanisms to support the growing diversity in application requirements. Existing strong consistency mechanisms are not appealing for Web environments due to their large state space or control message overhead. We focus on the lease approach that balances these tradeoffs and present analytical models and policies for determining the optimal lease duration. We present extensions to HTTP to incorporate leases and then implement our techniques in the Squid proxy cache and the Apache Web server. Our experimental evaluation of the leases approach shows that: (i) our techniques impose modest overheads even for long leases (a lease duration of 1 hour requires state to be maintained for 1030 leases and imposes a per object overhead of a control message every 33 minutes); (ii) leases yield a 138-425% improvement over existing strong consistency mechanisms; and (iii) the implementation overhead of leases is comparable to existing weak consistency mechanisms. Venkata Duvvuri, Prashant J. Shenoy, Renu Tewari |
INFOCOM | 2 |
| 2000 | Application performance in the QLinux multimedia operating systemabstractIn this paper, we argue that conventional operating systems need to be enhanced with predictable resource management mechanisms to meet the diverse performance requirements of emerging multimedia and web applications. We present QLinux—a multimedia operating system based on the Linux kernel that meets this requirement. QLinux employs hierarchical schedulers for fair, predictable allocation of processor, disk and network bandwidth, and accounting mechanisms for appropriate charging of resource usage. We experimentally evaluate the efficacy of these mechanisms using benchmarks and real-world applications. Our experimental results show that (i) emerging applications can indeed benefit from predictable allocation of resources, and (ii) the overheads imposed by the resource allocation mechanisms in QLinux are small. For instance, we show that the QLinux CPU scheduler can provide predictable performance guarantees to applications such as web servers and MPEG players, albeit at the expense of increasing the scheduling overhead. We conclude from our experiments that the benefits due to the resource management mechanisms in QLinux outweigh their increased overheads, making them a practical choice for conventional operating systems. Vijay Sundaram, Abhishek Chandra, Pawan Goyal 0001, Prashant J. Shenoy, Jasleen Sahni, Harrick M. Vin |
ACM Multimedia | 4 |
| 2000 | Surplus Fair Scheduling: A Proportional-Share CPU Scheduling Algorithm for Symmetric Multiprocessors
Abhishek Chandra, Micah Adler, Pawan Goyal 0001, Prashant J. Shenoy |
OSDI | 4 |
| 2000 | Implications of proxy caching for provisioning networks and serversabstractIn this paper, we examine the potential benefits of web proxy caches in improving the effective capacity of servers and networks. Since networks and servers are typically provisioned based on a high percentile of the load, we focus on the effects of proxy caching on the tail of the load distribution. We find that, unlike their substantial impact on the average load, proxies have a diminished impact on the tail of the load distribution. The exact reduction in the tail and the corresponding capacity savings depend on the percentile of the load distribution chosen for provisioning networks and servers—the higher the percentile, the smaller the savings. In particular, compared to over a 50% reduction in the average load, the savings in network and server capacity is only 20-35% for the 99th percentile of the load distribution. We also find that while proxies can be somewhat useful in smoothing out some of the burstiness in web workloads; the resulting workload continues, however, to exhibit substantial burstiness and a heavy-tailed nature. We identify large objects with poor locality to be the limiting factor that diminishes the impact of proxies on the tail of load distribution. We conclude that, while proxies are immensely useful to users due to the reduction in the average response time, they are less effective in improving the capacities of networks and servers. M. S. Raunak 0001, Prashant J. Shenoy, Pawan Goyal 0001, Krithi Ramamritham |
SIGMETRICS | 2 |
| 2000 | Failure Recovery Algorithms for Multimedia Servers
Prashant J. Shenoy, Harrick M. Vin |
Multim. Syst. | 1 |
| 1999 | Architectural considerations for next generation file systemsabstractWe evaluate two architectural alternatives—partitioned and integrated—for designing next generation file systems. Whereas a partitioned server employs a separate file system for each application class, an integrated file server multiplexes its resources among all application classes; we evaluate the performance of the two architectures with respect to sharing of disk bandwidth among the application classes. We show that although the problem of sharing disk bandwidth in integrated file systems is conceptually similar to that of sharing network link bandwidth in integrated services networks, the arguments that demonstrate the superiority of integrated services networks over separate networks are not applicable to file systems. Furthermore, we show that: (i) an integrated server outperforms the partitioned server in a large operating region and has slightly worse performance in the remaining region, (ii) the capacity of an integrated server is larger than that of the partitioned server, and (iii) an integrated server outperforms the partitioned server by up to a factor of 6 in the presence of bursty workloads. Prashant J. Shenoy, Pawan Goyal 0001, Harrick M. Vin |
ACM Multimedia (1) | 1 |
| 1999 | Efficient Support for Interactive Operations in Multi-Resolution Video Servers
Prashant J. Shenoy, Harrick M. Vin |
Multim. Syst. | 1 |
| 1999 | Efficient Striping Techniques for Variable Bit Rate Continuous Media File Servers
Prashant J. Shenoy, Harrick M. Vin |
Perform. Evaluation | 1 |
| 1998 | Cello: A Disk Scheduling Framework for Bext Generation Operating SystemsabstractIn this paper, we present the Cello disk scheduling framework for meeting the diverse service requirements of applications. Cello employs a two-level disk scheduling architecture, consisting of a class-independent scheduler and a set of class-specific schedulers. The two levels of the framework allocate disk bandwidth at two time-scales: the class-independent scheduler governs the coarse-grain allocation of bandwidth to application classes, while the class-specific schedulers control the fine-grain interleaving of requests. The two levels of the architecture separate application-independent mechanisms from application-specific scheduling policies, and thereby facilitate the co-existence of multiple class-specific schedulers. We demonstrate that Cello is suitable for next generation operating systems since: (i) it aligns the service provided with the application requirements, (ii) it protects application classes from one another, (iii) it is work-conserving and can adapt to changes in work-load, (iv) it minimizes the seek time and rotational latency overhead incurred during access, and (v) it is computationally efficient. Prashant J. Shenoy, Harrick M. Vin |
SIGMETRICS | 1 |
| 1996 | A Reliable, Adaptive Network Protocol for Video TransportabstractWe present an adaptive network layer protocol for VBR video transport. It: (1) minimizes the buffer requirement in the network while guaranteeing that packets of VBR encoded video flows will not be lost, and (2) minimizes the end-to-end delay and jitter of frames. To achieve the former objective, we utilize a receiver-oriented adaptive credit-based flow control algorithm, and derive the necessary and sufficient number of buffers that should be reserved for ensuring its reliability. To minimize the end-to-end delay and jitter for VBR encoded video streams, we: (1) present bandwidth estimation techniques which exploit the structure of the video traffic, and (2) define a new fairness criteria for buffer allocation and then present a fair buffer/bandwidth allocation algorithm. We experimentally evaluate this protocol for a wide range of parameters and many network configurations, and demonstrate its adaptability. We also compare the performance of the protocol with numerous other schemes and demonstrate its suitability for video transport. Pawan Goyal 0001, Harrick M. Vin, Chia Shen, Prashant J. Shenoy |
INFOCOM | 4 |
| 1996 | Fault-tolerant Architectures for Continuous Media ServersabstractContinuous media servers that provide support for the storage and retrieval of continuous media data (e.g., video, audio) at guaranteed rates are becoming increasingly important. Such servers, typically, rely on several disks to service a large number of clients, and are thus highly susceptible to disk failures. We have developed two fault-tolerant approaches that rely on admission control in order to meet rate guarantees for continuous media requests. The schemes enable data to be retrieved from disks at the required rate even if a certain disk were to fail. For both approaches, we present data placement strategies and admission control algorithms. We also present design techniques for maximizing the number of clients that can be supported by a continuous media server. Finally, through extensive simulations, we demonstrate the effectiveness of our schemes. 1 Introduction Rapid advances in computing and communication technologies have fueled an explosive growth in the multimedia indus... Banu Özden, Rajeev Rastogi, Prashant J. Shenoy, Avi Silberschatz |
SIGMOD Conference | 3 |
| 1995 | Efficient Support for Scan Operations in Video ServersabstractIn this paper, we present an algorithm that integrates scalable compression techniques with placement algorithms for disk arrays to efficiently support interactive scan operations (i.e., fast-forward and rewind) in video servers. We demonstrate that by suitably exploiting the characteristics of video streams and human perceptual tolerances, the overhead of such interactive operations can be substantially reduced. We present an analytical model for evaluating the impact of the fast-forward operation on the performance of disk-array-based servers. We validate the model through extensive simulations and analyze our results. 1 Introduction Recent advances in computing and communication technologies promise to create an infrastructure in which computer systems will support a wide range of interactive multimedia services in a variety of commercial and entertainment domains (e.g., advertising, online news, customer support, video on-demand, etc.). In its simplest configuration, such services... Prashant J. Shenoy, Harrick M. Vin |
ACM Multimedia | 1 |