EDBT 2026 Demo / reviewers in the wild / expert
Bhuvan Urgaonkar
dblp:61/6430
· DBLP profile ↗
76ranked-venue papers
11as first author
12since 2021 · last 2025
0000-0003-3495-6345ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 46 · 6 first-author · 5 since 2021Software engineering, systems software and programming languages · 16 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 8 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 2 since 2021Computer networks · 7 · 2 first-authorSecurity and privacy · 2 · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | PBench: Workload Synthesizer with Real Statistics for Cloud Analytics BenchmarkingabstractCloud service providers commonly use standard benchmarks like TPC-H and TPC-DS to evaluate and optimize cloud data analytics systems. However, these benchmarks rely on fixed query patterns and fail to capture real execution statistics of production cloud workloads. Although some cloud database vendors have recently released real workload traces, these traces alone do not qualify as benchmarks, as they typically lack essential components (i.e., queries and databases). To overcome this limitation, this paper studies a new problem of workload synthesis with real statistics , which generates synthetic workloads that closely approximate real execution statistics, including key performance metrics and operator distributions. To address this problem, we propose PBench, a novel workload synthesizer that constructs synthetic workloads by (1) selecting and combining workload components from existing benchmarks and (2) augmenting new workload components. This paper studies the key challenges in PBench. First, we address the challenge of balancing performance metrics and operator distributions by introducing a multi-objective optimization-based component selection method. Second, to capture the temporal dynamics of real workloads, we design a timestamp assignment method that progressively reines workload timestamps. Third, to handle the disparity between the original workload and the candidate workload, we propose a component augmentation approach that leverages large language models (LLMs) to generate additional workload components while maintaining statistical idelity. Experimental results show that PBench reduces approximation error by up to 6X compared to state-of-the-art methods. Chunwei Liu, Bhuvan Urgaonkar, Zhengle Wang, Magnus Mueller, Chao Zhang 0034, Songyue Zhang, Pascal Pfeil, Dominik Horn, Zhengchun Liu, Davide Pagano, Tim Kraska, Samuel Madden 0001, Ju Fan |
Proc. VLDB Endow. | 3 |
| 2025 | TraceScaler: A Framework for Scaling Load in Real-World Traces for System EvaluationabstractTrace replay is a common approach for evaluating systems by rerunning historical traffic patterns, but it’s not always possible to find suitable real-world traces at the desired level of system load. To experiment with different loads, one needs to downscale a trace to decrease the load or upscale a trace to artificially increase the load. This article expands upon our work, TraceUpscaler [ 92 ], by considering the interaction of upscaling and downscaling. In addition to evaluating upscaling with traces collected from a subset of the cluster, we also evaluate upscaling with traces that were downscaled with the state-of-the-art downscaling tool, TraceSplitter [ 91 ], to demonstrate that the upscaling and downscaling techniques are compatible and do not introduce unexpected artifacts in the scaling. In addition to comparing against prior approaches, we develop a novel upscaling technique, TraceOverlap , based on the idea of overlapping different time periods in a trace, where we identify the most similar time periods to overlap. Our evaluation demonstrates that TraceUpscaler and TraceOverlap are both more accurate in maintaining latency characteristics than prior approaches, with TraceUpscaler matching the original trace latency more closely. Finally, we provide a unified framework, TraceScaler , that combines TraceUpscaler with TraceSplitter to provide experimenters a common tool for their trace scaling needs. Sultan Mahmud Sajal, Salman Estyak, Rubaba Hasan, Timothy Zhu, Bhuvan Urgaonkar, Siddhartha Sen 0001 |
ACM Trans. Comput. Syst. | 5 |
| 2024 | AutoBurst: Autoscaling Burstable Instances for Cost-effective Latency SLOsabstractBurstable instances provide a low-cost option for consumers using the public cloud, but they come with significant resource limitations. They can be viewed as "fractional instances" where one receives a fraction of the compute and memory capacity at a fraction of the cost of regular instances. The fractional compute is achieved via rate limiting, where a unique characteristic of the rate limiting is that it allows for the CPU to burst to 100% utilization for limited periods of time. Prior research has shown how this ability to burst can be used to serve specific roles such as a cache backup and handling flash crowds. Our work provides a general-purpose approach to meeting latency SLOs via this burst capability while optimizing for cost. AutoBurst is able to achieve this by controlling both the number of burstable and regular instances along with how/when they are used. Evaluations show that our system is able to reduce cost by up to 25% over the state-of-the-art while maintaining latency SLOs. Rubaba Hasan, Timothy Zhu, Bhuvan Urgaonkar |
SoCC | 3 |
| 2024 | TraceUpscaler: Upscaling Traces to Evaluate Systems at High LoadabstractTrace replay is a common approach for evaluating systems by rerunning historical traffic patterns, but it is not always possible to find suitable real-world traces at the desired level of system load. Experimenting with higher traffic loads requires upscaling a trace to artificially increase the load. Unfortunately, most prior research has adopted ad-hoc approaches for upscaling, and there has not been a systematic study of how the upscaling approach impacts the results. One common approach is to count the arrivals in a predefined time-interval and multiply these counts by a factor, but this requires generating new requests/jobs according to some model (e.g., a Poisson process), which may not be realistic. Another common approach is to divide all the timestamps in the trace by an upscaling factor to squeeze the requests into a shorter time period. However, this can distort temporal patterns within the input trace. This paper evaluates the pros and cons of existing trace upscaling techniques and introduces a new approach, TraceUpscaler, that avoids the drawbacks of existing methods. The key idea behind TraceUpscaler is to decouple the arrival timestamps from the request parameters/data and upscale just the arrival timestamps in a way that preserves temporal patterns within the input trace. Our work applies to open-loop traffic where requests have arrival timestamps that aren't dependent on previous request completions. We evaluate TraceUpscaler under multiple experimental settings using both real-world and synthetic traces. Through our study, we identify the trace characteristics that affect the quality of upscaling in existing approaches and show how TraceUpscaler avoids these pitfalls. We also present a case study demonstrating how inaccurate trace upscaling can lead to incorrect conclusions about a system's ability to handle high load. Sultan Mahmud Sajal, Timothy Zhu, Bhuvan Urgaonkar, Siddhartha Sen 0001 |
EuroSys | 3 |
| 2024 | Resource Management in Aurora ServerlessabstractAmazon Aurora Serverless is an on-demand, autoscaling configuration for Amazon Aurora with full MySQL and PostgreSQL compatibility. It automatically offers capacity scale-up/down (i.e., vertical scaling) based on a customer database application's needs. For customers with time-varying workloads, it offers cost savings compared to provisioned Aurora or other alternatives due to its agile and granular scaling and its usage-based charging model. This paper describes the key ideas underlying Aurora Serverless's resource management. To help meet its goals, Aurora Serverless adapts and fine tunes well-established ideas related to resource over-subscription; reactive control informed by recent measurements; distributed & hierarchical decision-making; and innovations in the DB engine, OS, and hypervisor for efficiency. Perhaps the most challenging goal is to offer a consistent resource elasticity experience while operating hosts at high degrees of utilization. Aurora Serverless implements several novel ideas for striking a balance between these opposing needs. Its technique for mapping workloads to hosts ensures that, in the common case, there is adequate spare capacity within a host to support fast scale-up for a workload. In the rare event this is not so, it live migrates workloads to ensure seamless scale-up. Its load distribution strategy is characterized by "unbalancing" of load across hosts to enable agile live migrations. Finally, it employs a token bucket-based rate regulation mechanism to prevent a growing workload from saturating its host faster than live migration-based remedial actions. Bradley Barnhart, Marc Brooker, Daniil Chinenkov, Tony Hooper, Jihoun Im, Prakash Chandra Jha 0003, Tim Kraska, Ashok Kurakula, Grant Mcalister, Arjun Muthukrishnan, Aravinthan Narayanan, Douglas Terry, Bhuvan Urgaonkar, Jiaming Yan |
Proc. VLDB Endow. | 14 |
| 2023 | SCOOP: A Scalable Object-Oriented Serverless PlatformabstractFunction-as-a-Service (FaaS) has been the primary component to drive the movement toward serverless computing. These lightweight and scalable components, though attractive, are non-trivial to accommodate the needs of long-running stateful applications. In this paper, we highlight the drawbacks of existing stateful FaaS proposals, in turn motivating the need to rethink the stateful serverless model for building general-purpose applications, while maintaining its benefits such as auto-scaling and pay-per-use cost model. We present a novel serverless model based on the object-oriented (OO) programming paradigm, with Object-as-a-Service (OaaS), acting as the only component of the serverless design. Through our experimental evaluations, we demonstrate that the proposed architecture, named SCOOP, can improve the end-to-end latency of applications by 52% and 58%, compared to the state-of-the-art stateless and stateful FaaS implementations, respectively, while reducing the SLO violations by up to 14% by scaling resources based on the traffic fluctuations in the WITS and Berkeley traces. Narges Shahidi, Jashwant Raj Gunasekaran, Mahmut T. Kandemir, Bhuvan Urgaonkar |
CLOUD | 4 |
| 2022 | Splice: An Automated Framework for Cost-and Performance-Aware Blending of Cloud ServicesabstractWith the rapid growth of users adopting public clouds to run their applications, the types of resources procured from the different public cloud resource offerings are critical in simultaneously achieving satisfactory performance and reducing deployment costs. Typically, no one resource type can meet all application requirements, and thus combining different resource offerings is known to considerably reduce the performance-cost problem. However, it is non-trivial to use blended resources, due to the manual overhead of designing and implementing such blended approaches. Specifically, it necessitates rewriting the application code to suit a given resource and scaling it on demand. In order to overcome this manual hurdle, we take the first step by proposing Splice, an automated framework for cost-and performance-aware blending of IaaS and FaaS services. The three major goals of Splice are: (1) while cost-saving opportunities exist from blending resources, we aim to largely automate the blending process for public cloud services through a compiler-driven approach; (2) more specifically, we focus on automated blending of VMs and serverless functions; and (3) for serverless applications which contain multiple chained functions, we unearth the potential choices in determining a portion of the services to be blended cost-efficiently. We implement Splice on Amazon Web Services (AWS) using an Abstract Syntax Tree (AST), and extensively evaluate its effectiveness using several ap-plications with real-world traces. Our experiments demonstrate that, through automated blending, Splice is able to reduce SLO violations by 31 % compared to VM - based resource procurement schemes, while simultaneously minimizing costs by up to 32 %. Myungjun Son, Shruti Mohanty, Jashwant Raj Gunasekaran, Aman Jain, Mahmut T. Kandemir, George Kesidis, Bhuvan Urgaonkar |
CCGRID | 7 |
| 2022 | Multi-resource fair allocation for consolidated flash-based caching systemsabstractUsing a flash-based layer to serve the caching and buffering needs of multiple workloads has become a common practice. In such settings, resource demands will inevitably exceed available capacity sometimes. "Fair" resource allocation may offer a systematic way of partitioning resources across competing workloads during such periods of scarcity. Existing works only offer fair allocation strategies for a single resource (capacity or bandwidth) within a flash device in isolation. However, since there exist multiple critical resources that need to be partitioned within a flash device and they are correlated to each other, fair allocation of a single resource may result in a waste of other resource(s) or performance degradation of workload(s). To this end, we make a case for multi-resource fair allocation solutions for flash-based caches that consolidate multiple workloads. Furthermore, we argue that device lifetime, which depends on the behavior of running workloads, should also be considered as a first-class resource on par with capacity and bandwidth. Specifically, we build upon existing ideas related to dominant resource fairness (DRF) to devise flash-specific multi-resource fair algorithms: (i) nDRF, that jointly allocates capacity and bandwidth taking their non-linear relationship into account; (ii) ℓDRF, that explicitly considers lifetime as well in its allocation; and (iii) several variants of these. Our experimental evaluation offers important findings: (i) both nDRF and ℓDRF result in superior performance fairness compared to the state-of-the-art techniques that partition capacity in isolation; (ii) ℓDRF additionally offers improved device "wear" behavior; and (iii) our algorithms combined with reasonable demand prediction work very well in online settings with workload dynamism and uncertainty. Wonil Choi, Bhuvan Urgaonkar, Mahmut T. Kandemir, George Kesidis |
Middleware | 2 |
| 2022 | Invited Paper: Towards Practical Atomic Distributed Shared Memory: An Experimental Evaluation
Andria Trigeorgi, Nicolas C. Nicolaou, Chryssis Georgiou, Theophanis Hadjistasi, Efstathios Stavrakis, Viveck R. Cadambe, Bhuvan Urgaonkar |
SSS | 7 |
| 2022 | LEGOStore: A Linearizable Geo-Distributed Store Combining Replication and Erasure CodingabstractWe design and implement LEGOStore, an erasure coding (EC) based linearizable data store over geo-distributed public cloud data centers (DCs). For such a data store, the confluence of the following factors opens up opportunities for EC to be latency-competitive with replication: (a) the necessity of communicating with remote DCs to tolerate entire DC failures and implement linearizability; and (b) the emergence of DCs near most large population centers. LEGOStore employs an optimization framework that, for a given object, carefully chooses among replication and EC, as well as among various DC placements to minimize overall costs. To handle workload dynamism, LEGOStore employs a novel agile reconfiguration protocol. Our evaluation using a LEGOStore prototype spanning 9 Google Cloud Platform DCs demonstrates the efficacy of our ideas. We observe cost savings ranging from moderate (5-20%) to significant (60%) over baselines representing the state of the art while meeting tail latency SLOs. Our reconfiguration protocol is able to transition key placements in 3 to 4 inter-DC RTTs (< 1s in our experiments), allowing for agile adaptation to dynamic conditions. Hamidreza Zare, Viveck R. Cadambe, Bhuvan Urgaonkar, Nader Alfares, Praneet Soni, Arif Merchant |
Proc. VLDB Endow. | 3 |
| 2021 | Building an Accessible, Usable, Scalable, and Sustainable Service for Scholarly Big DataabstractSince the emergence of scholarly big data, there have been several efforts for web-based services such as digital library search engines (DLSEs). However, much of the design and specifications of an accessible, usable, scalable, and sustainable DLSE have not been well represented and discussed in the literature. We argue that these four characteristics are essential to providing a high-quality service for scholarly big data from both the user and developer’s perspectives. This paper reviews the design, implementation, and operation experiences, and lessons of CiteSeerX, a real-world digital library search engine. We analyze the strengths and weaknesses of the current design, and proposed a new design with a revised architecture, enhanced hardware, and software infrastructure. The Alpha version of the new design has been implemented and tested. The new system replaces MySQL and Apache Solr with a single instance of Elasticsearch, which plays a dual role of data storage and search. Another major improvement is the integration of extraction and ingestion, which significantly boosts document ingestion speed. The web application is re-engineered to enhance the user experience by applying a learning-to-rank model and offering more refined search tools. The system is also improved in many other aspects. We believe the design considerations and experience can benefit researchers and engineers who plan, design, and upgrade future systems with comparable scales and functionalities. Jian Wu 0006, Shaurya Rohatgi, Sai Raghav Reddy Keesara, Jason Chhay, Kevin Kuo, Arjun Manoj Menon, Sean Parsons, Bhuvan Urgaonkar, C. Lee Giles |
IEEE BigData | 8 |
| 2021 | TraceSplitter: a new paradigm for downscaling tracesabstractRealistic experimentation is a key component of systems research and industry prototyping, but experimental clusters are often too small to replay the high traffic rates found in production traces. Thus, it is often necessary to downscale traces to lower their arrival rate, and researchers/practitioners generally do this in an ad-hoc manner. For example, one practice is to multiply all arrival timestamps in a trace by a scaling factor to spread the load across a longer timespan. However, temporal patterns are skewed by this approach, which may lead to inappropriate conclusions about some system properties (e.g., the agility of auto-scaling). Another popular approach is to count the number of arrivals in fixed-sized time intervals and scale it according to some modeling assumptions. However, such approaches can eliminate or exaggerate the fine-grained burstiness in the trace depending on the time interval length. Sultan Mahmud Sajal, Rubaba Hasan, Timothy Zhu, Bhuvan Urgaonkar, Siddhartha Sen 0001 |
EuroSys | 4 |
| 2020 | Fair Write Attribution and Allocation for Consolidated Flash CacheabstractConsolidating multiple workloads on a single flash-based storage device is now a common practice. We identify a new problem related to lifetime management in such settings: how should one partition device resources among consolidated workloads such that their allowed contributions to the device's wear (resulting from their writes including hidden writes due to garbage collection) may be deemed fairly assigned? When flash is used as a cache/buffer, such fairness is important because it impacts what and how much traffic from various workloads may be serviced using flash which in turn affects their performance. We first clarify why the write attribution problem (i.e., which workload contributed how many writes) is non-trivial. We then present a technique for it inspired by the Shapley value, a classical concept from cooperative game theory, and demonstrate that it is accurate, fair, and feasible. We next consider how to treat an overall "write budget" (i.e., total allowable writes during a given time period) for the device as a first-class resource worthy of explicit management. Towards this, we propose a novel write budget allocation technique. Finally, we construct a dynamic lifetime management framework for consolidated devices by putting the above elements together. Our experiments using real-world workloads demonstrate that our write allocation and attribution techniques lead to performance fairness across consolidated workloads. Wonil Choi, Bhuvan Urgaonkar, Mahmut T. Kandemir, Myoungsoo Jung |
ASPLOS | 2 |
| 2020 | SplitServe: Efficiently Splitting Apache Spark Jobs Across FaaS and IaaSabstractDue to their lower startup latencies and finer-grain pricing than virtual machines (VMs), Amazon Lambdas and other cloud functions (CFs) have been identified as ideal candidates for handling unexpected spikes in simple, stateless workloads. However, it is not immediately clear if CFs would be similarly effective in autoscaling complex workloads involving significant state transfer across distributed application components. We have found that, through careful design, currently available CFs can indeed be useful even for complex workloads. To demonstrate this, we design and implement SplitServe, an enhancement of Apache Spark. If not enough executors on existing VMs are available for a newly arriving latency-sensitive job, SplitServe is able to use CFs to quickly bridge this shortfall in VMs, so avoiding the startup latencies of newly requested VMs. If desirable in terms of performance or cost, when newly requested VMs, or executors on existing VMs, do become available, SplitServe is able to move ongoing work from CFs to them. Our experimental evaluation of SplitServe using four different workloads (either on a mixture of VM-based executors and CFs or just CFs) shows that it improves execution time by up to (a) 55% for workloads with small to modest amount of shuffling, and (b) 31% in workloads with large amounts of shuffling, when compared to only VM-based autoscaling. Aman Jain, Ataollah Fatahi Baarzi, George Kesidis, Bhuvan Urgaonkar, Nader Alfares, Mahmut T. Kandemir |
Middleware | 4 |
| 2019 | Spock: Exploiting Serverless Functions for SLO and Cost Aware Resource Procurement in Public CloudabstractWe are witnessing the emergence of elastic web services which are hosted in public cloud infrastructures. For reasons of cost-effectiveness, it is crucial for the elasticity of these web services to match the dynamically-evolving user demand. Traditional approaches employ clusters of virtual machines (VMs) to dynamically scale resources based on application demand. However, they still face challenges such as higher cost due to over-provisioning or incur service level objective (SLO) violations due to under-provisioning. Motivated by this observation, we propose Spock, a new scalable and elastic control system that exploits both VMs and serverless functions to reduce cost and ensure SLO for elastic web services. We show that under two different scaling policies, Spock reduces SLO violations of queries by up to 74% when compared to VM-based resource procurement schemes. Further, Spock yields significant cost savings, by up to 33% compared to traditional approaches which use only VMs. Jashwant Raj Gunasekaran, Prashanth Thinakaran, Mahmut T. Kandemir, Bhuvan Urgaonkar, George Kesidis, Chita R. Das |
CLOUD | 4 |
| 2019 | BurScale: Using Burstable Instances for Cost-Effective Autoscaling in the Public CloudabstractCloud providers have recently introduced burstable instances - virtual machines whose CPU capacity is rate limited by token-bucket mechanisms. A user of a burstable instance is able to burst to a much higher resource capacity ("peak rate") than the instance's long-term average capacity ("sustained rate"), provided the bursts are short and infrequent. A burstable instance tends to be much cheaper than a conventional instance that is always provisioned for the peak rate. Consequently, cloud providers advertise burstable instances as cost-effective options for customers with intermittent needs and small (e.g., single VM) clusters. By contrast, this paper presents two novel usage scenarios for burstable instances in larger clusters with sustained usage. We demonstrate (i) how burstable instances can be utilized alongside conventional instances to handle the transient queueing arising from variability in traffic, and (ii) how burstable instances can mask the VM startup/warmup time when autoscaling to handle flash crowds. We implement our ideas in a system called BurScale and use it to demonstrate cost-effective autoscaling for two important workloads: (i) a stateless web server cluster, and (ii) a stateful Memcached caching cluster. Results from our prototype system show that via its careful combination of burstable and regular instances, BurScale can ensure similar application performance as traditional autoscaling systems that use all regular instances while reducing cost by up to 50%. Ataollah Fatahi Baarzi, Timothy Zhu, Bhuvan Urgaonkar |
SoCC | 3 |
| 2019 | SpIitServe: Efficiently Splitting Complex Workloads Across FaaS and IaaSabstractAmazon Web Services (AWS) Lambdas and other "cloud functions" (CFs) offer much lower startup latencies than virtual machines (VMs) (tens/hundreds of milliseconds vs. a few/several minutes) with lower minimum cost. This makes it appealing to use them for handling unexpected spikes in simple, stateless workloads [2, 3, 5]. If the spike persists, additional VMs may be launched and CFs can be decommissioned when the VMs are ready (VMs are cheaper per unit resource procured than CFs). However, it is not immediately clear if using CFs for complex workloads - those involving significant state exchange among components - is similarly effective. Current CFs have several restrictions that may limit their efficacy: (i) relatively limited resource capacity, especially main memory (e.g., an AWS Lambda may only have up to 3GB memory), (ii) limited lifetime (e.g., Lambdas are terminated after 15 minutes), and (iii) limited support for sharing of intermediate state (e.g., Lambdas must employ an external storage system such as AWS S3). Contrary to conventional wisdom, we show that it is possible to exploit the faster startup times of CFs to improve cost and performance of autoscaling even for complex workloads. Aman Jain, Ataollah Fatahi Baarzi, Nader Alfares, George Kesidis, Bhuvan Urgaonkar, Mahmut T. Kandemir |
SoCC | 5 |
| 2019 | Fair Resource Allocation in Consolidated Flash Systems
Wonil Choi, Bhuvan Urgaonkar, Mahmut T. Kandemir, Myoungsoo Jung |
HotStorage | 2 |
| 2018 | A Cost-Efficient and Fair Multi-Resource Allocation Mechanism for Self-Organizing ServersabstractIn this paper, we study cost-efficient and fair allocation of multiple types of resources in an environment of heterogeneous and self-organizing servers. To address this problem, we formulate an optimization problem which aims at minimizing the operational costs for all servers, while providing fairness across different users. We propose a fully distributed implementation to solve this problem. The proposed mechanism is shown to achieve envy-freeness among different users. Furthermore, we show how it captures the trade-off between cost-efficiency and fairness. We employ numerical experiments to show the effectiveness of our proposed mechanism in reducing operational costs for a geo-distributed data-center. Jalal Khamse-Ashari, Ioannis Lambadaris, George Kesidis, Bhuvan Urgaonkar, Yiqiang Q. Zhao |
GLOBECOM | 4 |
| 2018 | Scheduling Distributed Resources in Heterogeneous Private CloudsabstractWe first consider the static problem of allocating resources to (i.e., scheduling) multiple distributed application frameworks, possibly with different priorities and server preferences, in a private cloud with heterogeneous servers. Several fair scheduling mechanisms have been proposed for this purpose. We extend prior results on max-min fair (MMF) and proportional fair (PF) scheduling to this constrained multiresource and multiserver case for generic fair scheduling criteria. The task efficiencies (a metric related to proportional fairness) of max-min fair allocations found by progressive filling are compared by illustrative examples. In the second part of this paper, we consider the online problem (with framework churn) by implementing variants of these schedulers in Apache Mesos using progressive filling to dynamically approximate max-min fair allocations. We evaluate the implemented schedulers in terms of overall execution time of realistic distributed Spark workloads. Our experiments show that resource efficiency is improved and execution times are reduced when the scheduler is "server specific" or when it leverages characterized required resources of the workloads (when known). George Kesidis, Yuquan Shan, Aman Jain, Bhuvan Urgaonkar, Jalal Khamse-Ashari, Ioannis Lambadaris |
MASCOTS | 4 |
| 2018 | Effective Capacity Modulation as an Explicit Control Knob for Public Cloud ProfitabilityabstractIn this article, we explore the efficacy of dynamic effective capacity modulation (i.e., using virtualization techniques to offer lower resource capacity than that advertised by the cloud provider) as a control knob for a cloud provider’s profit maximization complementing the more well-studied approach of dynamic pricing. In particular, our focus is on emerging cloud ecosystems wherein we expect tenants to modify their demands strategically in response to such modulation in effective capacity and prices. Toward this, we consider a simple model of a cloud provider that offers a single type of virtual machine to its tenants and devise a leader/follower game-based cloud control framework to capture the interactions between the provider and its tenants. We assume both parties employ myopic control and short-term predictions to reflect their operation under the high dynamism and poor predictability in such environments. Our evaluation using a combination of real data center traces and real-world benchmarks hosted on a prototype OpenStack-based cloud shows 10% to 30% profit improvement for a cloud provider compared with baselines that use static pricing and/or static effective capacity. Cheng Wang 0014, Bhuvan Urgaonkar, George Kesidis, Lydia Y. Chen, Robert Birke |
ACM Trans. Auton. Adapt. Syst. | 2 |
| 2018 | An Efficient and Fair Multi-Resource Allocation Mechanism for Heterogeneous ServersabstractEfficient and fair allocation of multiple types of resources is a crucial objective in a cloud/distributed computing cluster. Users may have diverse resource needs. Furthermore, diversity in server properties/capabilities may mean that only a subset of servers may be usable by a given user. In platforms with such heterogeneity, we identify important limitations in existing multi-resource fair allocation mechanisms, notably Dominant Resource Fairness and its follow-up work. To overcome such limitations, we propose a new server-based approach; each server allocates resources by maximizing a per-server utility function. We propose a specific class of utility functions which, when appropriately parameterized, adjusts the trade-off between efficiency and fairness, and captures a variety of fairness measures (such as our recently proposed Per-Server Dominant Share Fairness ). We establish conditions for the proposed mechanism to satisfy certain properties that are generally deemed desirable, e.g., envy-freeness, sharing incentive, bottleneck fairness, and Pareto optimality. To implement our resource allocation mechanism, we develop an iterative algorithm which is shown to be globally convergent. Subsequently, we show how the proposed mechanism could be implemented in a distributed fashion. Finally, we carry out extensive trace-driven simulations to show the enhanced performance of our proposed mechanism over the existing ones. Jalal Khamse-Ashari, Ioannis Lambadaris, George Kesidis, Bhuvan Urgaonkar, Yiqiang Q. Zhao |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2017 | Exploiting Spot and Burstable Instances for Improving the Cost-efficacy of In-Memory Caches on the Public CloudabstractIn order to keep the costs of operating in-memory storage on the public cloud low, we devise novel ideas and enabling modeling and optimization techniques for combining conventional Amazon EC2 instances with the cheaper spot and burstable instances. Whereas a naturally appealing way of using failure-prone spot instances is to selectively store unpopular ("cold") content, we show that a form of "hot-cold mixing" across regular and spot instances might be more cost-effective. To overcome performance degradation resulting from spot instance revocations, we employ a highly available passive backup using the recently emergent burstable instances. We show how the idiosyncratic resource allocations of burstable instances make them ideal candidates for such a backup. We implement all our ideas in an EC2-based memcached prototype. Using simulations and live experiments on our prototype, we show that (i) our hot-cold mixing, informed by our modeling of spot prices, helps improve cost savings by 50-80% compared to only using regular instances, and (ii) our burstable-based backup helps reduce performance degradation during spot revocation, e.g., the 95% latency during failure recovery improves by 25% compared to a backup based on regular instances. Cheng Wang 0014, Bhuvan Urgaonkar, George Kesidis, Qianlin Liang |
EuroSys | 2 |
| 2017 | Per-Server Dominant-Share Fairness (PS-DSF): A multi-resource fair allocation mechanism for heterogeneous serversabstractUsers of cloud computing platforms pose different types of demands for multiple resources on servers (physical or virtual machines). Besides differences in their resource capacities, servers may be additionally heterogeneous in their ability to service users - certain users' tasks may only be serviced by a subset of the servers. We identify important shortcomings in existing multi-resource fair allocation mechanisms - Dominant Resource Fairness (DRF) and its follow up work - when used in such environments. We develop a new fair allocation mechanism called Per-Server Dominant-Share Fairness (PS-DSF) which we show offers all desirable sharing properties that DRF is able to offer in the case of a single “resource pool” (i.e., if the resources of all servers were pooled together into one hypothetical server). We evaluate the performance of PS-DSF through simulations. Our evaluation shows the enhanced efficiency of PS-DSF compared to the existing allocation mechanisms. We argue how our proposed allocation mechanism is applicable in cloud computing networks and especially large scale data-centers. Jalal Khamse-Ashari, Ioannis Lambadaris, George Kesidis, Bhuvan Urgaonkar, Yiqiang Q. Zhao |
ICC | 4 |
| 2017 | Competition and Peak-Demand Pricing in Clouds Under Tenants' Demand ResponseabstractA significant fraction of the operational expenditures incurred by cloud service providers relates to their networking (Internet access) and electricity consumption. Both depend on the peak-demand over the billing interval. In the future, cloud services providers may in turn recoup these costs from their long-term customers through peak-based pricing. We explore two different methods for the cloud provider to recoup this charge: (i) equal allocation and (ii) proportional to usage allocation. Furthermore, we consider multiple strategic tenants whose active demand response to cloud price settings jointly depends on job responsiveness (modeled as queueing delay of admitted jobs) and lost/shed workload (due to excessive delay). Under certain conditions, we prove existence and uniqueness of Nash equilibria for regimes (i) and (ii). Due to nonconvexity in the utility (or cost) functions, existence statements require leveraging potentiality arguments while uniqueness statements rely on imposing further convexityrequirements. The resulting Nash equilibrium is parametrized by the price per unit demand, which may be strategically set by the cloud to maximize its revenue subject to tenants reaching a Nash equilibrium. We model the resulting interactions as a Stackelberg game between the cloud and a set of tenants. A relatively general existence statement is provided for the Stackelberg equilibrium under regime (i). For a special case of regime (ii), the unique Stackelberg equilibrium is characterized. Finally, we provide a numerical study for such a framework using real-world peak-based prices from an electric utility and demands given by Google workload traces". George Kesidis, Uday V. Shanbhag, Neda Nasiriani, Bhuvan Urgaonkar |
MASCOTS | 4 |
| 2017 | An Empirical Analysis of Amazon EC2 Spot Instance Features Affecting Cost-effective Resource ProcurementabstractMany cost-conscious public cloud workloads ("tenants") are turning to Amazon EC2's spot instances because, on average, these instances offer significantly lower prices (up to 10 times lower) than on-demand and reserved instances of comparable advertized resource capacities. To use spot instances effectively, a tenant must carefully weigh the lower costs of these instances against their poorer availability. Towards this, we empirically study four features of EC2 spot instance operation that a cost-conscious tenant may find useful to model. Using extensive evaluation based on both historical and current spot instance data, we show shortcomings in the state-of-the-art modeling of these features that we overcome. Our analysis reveals many novel properties of spot instance operation some of which offer predictive value while others do not. Using these insights, we design predictors for our features that offer a balance between computational efficiency (allowing for online resource procurement) and cost-efficacy. We explore "case studies" wherein we implement prototypes of dynamic spot instance procurement advised by our predictors for two types of workloads. Compared to the state-of-the-art, our approach achieves (i) comparable cost but much better performance (fewer bid failures) for a latency-sensitive in-memory Memcached cache, and (ii) an additional 18% cost-savings with comparable (if not better than) performance for a delay-tolerant batch workload. Cheng Wang 0014, Qianlin Liang, Bhuvan Urgaonkar |
ICPE | 3 |
| 2017 | Resource Accounting of Shared IT Resources in Multi-Tenant CloudsabstractIn today's IT platforms, the capability to accurately account overall resource usage among applications is crucial for variety of management actions (e.g., capacity planning, dynamic resource reallocation and/or load balancing). However, in the environments where small number of shared services cater to a large number of distinct entities' requests, resource accounting becomes significantly challenging. First, the overall resource consumption at the shared service is the aggregate of the resource consumption for multiple remote entities whose identities are not visible to the shared service. Second, even if such information becomes available, common monitoring tools (e.g., top, iostat) are unable to deliver accurate break-down of resource consumption since sharing occurs at sub-instance level (i.e., service instances are not exclusive). We study inherent challenges of performing resource accounting of shared resource. We compare two nonintrusive approaches having different balance between local monitoring and collective inference - (i) LR that uses easily-available tools which provide aggregate measurement and applying well-known linear regression as inference, and (ii) Rameter that puts more emphasis on gathering fine-grained per-thread information from within the hypervisor and applying light inference on the data. Evaluation shows that Rameter offers less than 1% error in accounting whereas LR's error fluctuates between 5-150%. Byung-Chul Tak, Youngjin Kwon, Bhuvan Urgaonkar |
IEEE Trans. Serv. Comput. | 3 |
| 2016 | Towards Performance Modeling as a Service by Exploiting Resource Diversity in the Public CloudabstractCloud computing platforms such as Amazon EC2, Google Computing Engine, and Microsoft Azure offer dozens of virtual machine (VM) types with a wide range of resource capacity vs. price trade-offs, requiring a customer to consider numerous resource configurations when evaluating service needs. This report investigates the possibility of using the diversity of VM types to predict the performance of new VM types using black box modeling. The performance model used is a multiple linear regression of the average server response time, server load (throughput in requests per second), the number of CPU cores, and the memory in the procured VM. For three commonly used database servers - Redis (key-value stores), Apache Cassandra (NoSQL) and MySQL - the model accuracy increases for larger sets of VMs. e.g., for Redis, the measure of model efficacy improves from 0.4-0.5 with 2 VM types for training and 0.7 for 3 VM types to 0.8 for 4 VM types. These results suggest further interesting research challenges, such as the possibility of automating the process of calibrating performance models using diverse resource types on a public cloud leading to "performance modeling as a service". Mark Meredith, Bhuvan Urgaonkar |
CLOUD | 2 |
| 2016 | Fine-Grained Resource Scaling in a Public Cloud: A Tenant's PerspectiveabstractGrowing tenant workload needs and an increasingly competitive market will force cloud providers to operate their data centers at significantly higher utilization levels than seen today. We argue that a key enabler of such cloud ecosystems would be facilities for tenants to engage in fine-grained resource scaling in addition to those offered by current providers. The basic unit of resource scaling exposed by current cloud providers is the canonical interface of virtual machines (VMs) with relatively static resource capacities. This paper describes opportunities and challenges in augmenting this interface to also include fine-grained scaling of CPU and memory within an already procured VM. Qualitative arguments for why this would offer cost benefits for both the provider and its tenants are presented. We focus on the cost-effective operation of a tenant in such an environment via the design of a feedback controller. The efficacy of our ideas is illustrated by implementing a case study in a Memcached tenant workload. Our results are promising and point to an interesting and broad area for further research - e.g., with the real-world workload in our evaluation, up to 50% utility improvement can be achieved by just applying memory scaling, a further 66% improvement can be achieved by coordinating fine-grained CPU and memory scaling. Cheng Wang 0014, Bhuvan Urgaonkar |
CLOUD | 3 |
| 2016 | Constrained Max-Min Fair Scheduling of Variable-Length Packet-Flows to Multiple ServersabstractWe describe a scheduler for multiple servers shared among different packet-flows, where each packet-flow may be served by only a subset of available (preferred) servers. The scheduler allocates tokens to flows in a round-by-round manner, where token allocation to flows at the beginning of each round is weighted max-min fair. We present a packet scheduling scheme where when a server becomes free, it is allocated to serve the HOL packet of an eligible flow with the maximum remaining tokens. The scheduling algorithm is applicable even when the capacity of servers are not known a priori and may vary over duration of a round. Numerical examples are given to illustrate that the scheduler itself is weighted max-min fair. Jalal Khamse-Ashari, George Kesidis, Ioannis Lambadaris, Bhuvan Urgaonkar, Yiqiang Q. Zhao |
GLOBECOM | 4 |
| 2015 | On Fair Attribution of Costs under Peak-Based Pricing to Cloud TenantsabstractThe costs incurred by cloud providers towards operating their data centers are often determined in large part by their peak demands. The pricing schemes currently used by cloud providers to recoup these costs from their tenants, however, do not distinguish tenants based on their contributions to the cloud's overall peak demand. Using the concrete example of peak-based pricing as employed by many electric utility companies, we show that this "gap" may lead to unfair attribution of costs to the tenants. Simple enhancements of existing cloud pricing (e.g., analogous to the coincident peak pricing (CPP) used by some electric utilities) do not adequately address these shortcomings and suffer from short-term unfairness and undesirable oscillatory price vs. demand relationship offered to tenants. To overcome these shortcomings, we define an alternative pricing scheme to more fairly distribute a cloud's costs among its tenants. Our approach to fair attribution of cloud's costs is inspired by the concept of Shapley values used to fairly divide revenue among participants of a financial coalition. We demonstrate the efficacy of our scheme under price-sensitive tenant demand response using a combination of (i) extensive empirical evaluation with recent workloads from commercial data centers operated by IBM, and (ii) analytical modeling through non-cooperative game theory for a special case of tenant demand model. Neda Nasiriani, Cheng Wang 0014, George Kesidis, Bhuvan Urgaonkar, Lydia Y. Chen, Robert Birke |
MASCOTS | 4 |
| 2014 | A Hierarchical Demand Response Framework for Data Center Power Cost Optimization under Real-World Electricity PricingabstractWe study the problem of optimizing data center electric utility bill under uncertainty in workloads and real-world pricing schemes. Our focus is on using control knobs that modulate the power consumption of IT equipment. To overcome the difficulty of casting/updating such control problems and the computational intractability they suffer from in general, we propose and evaluate a hierarchical optimization framework wherein an upper layer uses (i) temporal aggregation to restrict the number of decision instants during a billing cycle to computationally feasible values, and (ii) spatial (i.e., control knob) aggregation whereby it models the large and diverse set of power control knobs with two abstract knobs labeled demand dropping and demand delaying. These abstract knobs operate upon a fluid power demand. The key insight underlying our modeling is that the power modulation effects of most IT control knobs can be succinctly captured as dropping and/or delaying a portion of the power demand. These decisions are passed onto a lower layer that leverages existing research to translate them into decisions for real IT knobs. We develop a suite of algorithms for our upper layer that deal with different forms of input uncertainty. An experimental evaluation of the proposed approach offers promising results: e.g., it offers net cost savings of about 25% and 18% to a streaming media server and a MapReduce-based batch workload, respectively. Cheng Wang 0014, Bhuvan Urgaonkar, Qian Wang 0029, George Kesidis |
MASCOTS | 2 |
| 2014 | HybridPlan: a capacity planning technique for projecting storage requirements in hybrid storage systems
Youngjae Kim 0001, Bhuvan Urgaonkar, Piotr Berman, Anand Sivasubramaniam |
J. Supercomput. | 3 |
| 2013 | Using Dark Fiber to Displace Diesel Generators
Aman Kansal, Bhuvan Urgaonkar, Sriram Govindan |
HotOS | 2 |
| 2013 | ACE: abstracting, characterizing and exploiting peaks and valleys in datacenter power consumptionabstractPeak power management of datacenters has tremendous cost implications. While numerous mechanisms have been proposed to cap power consumption, real datacenter power consumption data is scarce. To address this gap, we collect power demands at multiple spatial and fine-grained temporal resolutions from the load of geo-distributed datacenters of Microsoft over 6 months. We conduct aggregate analysis of this data, to study its statistical properties. With workload characterization a key ingredient for systems design and evaluation, we note the importance of better abstractions for capturing power demands, in the form of peaks and valleys. We identify and characterize attributes for peaks and valleys, and important correlations across these attributes that can influence the choice and effectiveness of different power capping techniques. With the wide scope of exploitability of such characteristics for power provisioning and optimizations, we illustrate its benefits with two specific case studies. Di Wang 0003, Chuangang Ren, Sriram Govindan, Anand Sivasubramaniam, Bhuvan Urgaonkar, Aman Kansal, Kushagra Vaid |
SIGMETRICS | 5 |
| 2013 | A Temporal Locality-Aware Page-Mapped Flash Translation Layer
Youngjae Kim 0001, Bhuvan Urgaonkar |
J. Comput. Sci. Technol. | 3 |
| 2013 | Aggressive Datacenter Power Provisioning with BatteriesabstractDatacenters spend $10--25 per watt in provisioning their power infrastructure, regardless of the watts actually consumed. Since peak power needs arise rarely, provisioning power infrastructure for them can be expensive. One can, thus, aggressively underprovision infrastructure assuming that simultaneous peak draw across all equipment will happen rarely. The resulting nonzero probability of emergency events where power needs exceed provisioned capacity, however small, mandates graceful reaction mechanisms to cap the power draw instead of leaving it to disruptive circuit breakers/fuses. Existing strategies for power capping use temporal knobs local to a server that throttle the rate of execution (using power modes), and/or spatial knobs that redirect/migrate excess load to regions of the datacenter with more power headroom. We show these mechanisms to have performance degrading ramifications, and propose an entirely orthogonal solution that leverages existing UPS batteries to temporarily augment the utility supply during emergencies. We build an experimental prototype to demonstrate such power capping on a cluster of 8 servers, each with an individual battery, and implement several online heuristics in the context of different datacenter workloads to evaluate their effectiveness in handling power emergencies. We show that our battery-based solution can: (i) handle emergencies of short durations on its own, (ii) supplement existing reaction mechanisms to enhance their efficacy for longer emergencies, and (iii) create more slack for shifting applications temporarily to nonpeak durations. Sriram Govindan, Di Wang 0003, Anand Sivasubramaniam, Bhuvan Urgaonkar |
ACM Trans. Comput. Syst. | 4 |
| 2013 | Cloudy with a Chance of Cost SavingsabstractCloud-based hosting is claimed to possess many advantages over traditional in-house (on-premise) hosting such as better scalability, ease of management, and cost savings. It is not difficult to understand how cloud-based hosting can be used to address some of the existing limitations and extend the capabilities of many types of applications. However, one of the most important questions is whether cloud-based hosting will be economically feasible for my application if migrated into the cloud. It is not straightforward to answer this question because it is not clear how my application will benefit from the claimed advantages, and, in turn, be able to convert them into tangible cost savings. Within cloud-based hosting offerings, there is a wide range of hosting options one can choose from, each impacting the cost in a different way. Answering these questions requires an in-depth understanding of the cost implications of all the possible choices specific to my circumstances. In this study, we identify a diverse set of key factors affecting the costs of deployment choices. Using benchmarks representing two different applications (TPC-W and TPC-E) we investigate the evolution of costs for different deployment choices. We consider important application characteristics such as workload intensity, growth rate, traffic size, storage, and software license to understand their impact on the overall costs. We also discuss the impact of workload variance and cloud elasticity, and certain cost factors that are subjective in nature. Byung-Chul Tak, Bhuvan Urgaonkar, Anand Sivasubramaniam |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2012 | Leveraging stored energy for handling power emergencies in aggressively provisioned datacentersabstractDatacenters spend $10-25 per watt in provisioning their power infrastructure, regardless of the watts actually consumed. Since peak power needs arise rarely, provisioning power infrastructure for them can be expensive. One can, thus, aggressively under-provision infrastructure assuming that simultaneous peak draw across all equipment will happen rarely. The resulting non-zero probability of emergency events where power needs exceed provisioned capacity, however small, mandates graceful reaction mechanisms to cap the power draw instead of leaving it to disruptive circuit breakers/fuses. Existing strategies for power capping use temporal knobs local to a server that throttle the rate of execution (using power modes), and/or spatial knobs that redirect/migrate excess load to regions of the datacenter with more power headroom. We show these mechanisms to have performance degrading ramifications, and propose an entirely orthogonal solution that leverages existing UPS batteries to temporarily augment the utility supply during emergencies. We build an experimental prototype to demonstrate such power capping on a cluster of 8 servers, each with an individual battery, and implement several online heuristics in the context of different datacenter workloads to evaluate their effectiveness in handling power emergencies. We show that: (i) our battery-based solution can handle emergencies of short duration on its own, (ii) supplement existing reaction mechanisms to enhance their efficacy for longer emergencies, and (iii) battery even provide feasible options when other knobs do not suffice. Sriram Govindan, Di Wang 0003, Anand Sivasubramaniam, Bhuvan Urgaonkar |
ASPLOS | 4 |
| 2012 | Using batteries to reduce the power costs of internet-scale distributed networksabstractModern Internet-scale distributed networks have hundreds of thousands of servers deployed in hundreds of locations and networks around the world. Canonical examples of such networks are content delivery networks (called CDNs) that we study in this paper. The operating expenses of large distributed networks are increasingly driven by the cost of supplying power to their servers. Typically, CDNs procure power through long-term contracts from co-location providers and pay on the basis of the power (KWs) provisioned for them, rather than on the basis of the energy (KWHs) actually consumed. We propose the use of batteries to reduce both the required power supply and the incurred power cost of a CDN. We provide a theoretical model and an algorithmic framework for provisioning batteries to minimize the total power supply and the total power costs of a CDN. We evaluate our battery provisioning algorithms using extensive load traces derived from Akamai's CDN to empirically study the achievable benefits. We show that batteries can provide up to 14% power savings, that would increase to 22% for more power-proportional next-generation servers, and would increase even more to 35.3% for perfectly power-proportional servers. Likewise, the cost savings, inclusive of the additional battery costs, range from 13.26% to 33.8% as servers become more power-proportional. Further, much of these savings can be achieved with a small cycle rate of one full discharge/charge cycle every three days that is conducive to satisfactory battery lifetimes. In summary, we show that a CDN can utilize batteries to significantly reduce both the total supplied power and the total power costs, thereby establishing batteries as a key element in future distributed network architecture. While we use the canonical example of a CDN, our results also apply to other similar Internet-scale distributed networks. Darshan S. Palasamudram, Ramesh K. Sitaraman, Bhuvan Urgaonkar, Rahul Urgaonkar |
SoCC | 3 |
| 2012 | Carbon-Aware Energy Capacity Planning for DatacentersabstractDatacenters are facing increasing pressure to cap their carbon footprints at low cost. Recent work has shown the significant environmental benefits of using renewable energy for datacenters by supply-following techniques (workload scheduling, geographical load balancing, etc.) However, all such prior work has only considered on-site renewable generation when numerous other options also exist, which may be superior to on-site renewables for many datacenters. Alternative ways for datacenters to incorporate renewable energy into their overall energy portfolio include: construction of or investment into off-site renewable farms at locations with more abundant renewable energy potential, indirect purchase of renewable energy through buying renewable energy certificates (RECs), purchase of renewable energy products such as power purchase agreements (PPAs) or through third-party renewable providers. We propose a general, optimization-based framework to minimize datacenter costs in the presence of different carbon footprint reduction goals, renewable energy characteristics, policies, utility tariff, and energy storage devices (ESDs). We expect that our work can help datacenter operators make informed decisions about sustainable, renewable-energy-powered IT system design. Chuangang Ren, Di Wang 0003, Bhuvan Urgaonkar, Anand Sivasubramaniam |
MASCOTS | 3 |
| 2012 | Energy storage in datacenters: what, where, and how much?abstractEnergy storage - in the form of UPS units - in a datacenter has been primarily used to fail-over to diesel generators upon power outages. There has been recent interest in using these Energy Storage Devices (ESDs) for demand-response (DR) to either shift peak demand away from high tariff periods, or to shave demand allowing aggressive under-provisioning of the power infrastructure. All such prior work has only considered a single/specific type of ESD (typically re-chargeable lead-acid batteries), and has only employed them at a single level of the power delivery network. Continuing technological advances have provided us a plethora of competitive ESD options ranging from ultra-capacitors, to different kinds of batteries, flywheels and even compressed air-based storage. These ESDs offer very different trade-offs between their power and energy costs, densities, lifetimes, and energy efficiency, among other factors, suggesting that employing hybrid combinations of these may allow more effective DR than with a single technology. Furthermore, ESDs can be placed at different, and possibly multiple, levels of the power delivery hierarchy with different associated trade-offs. To our knowledge, no prior work has studied the extensive design space involving multiple ESD technology provisioning and placement options. This paper intends to fill this critical void, by presenting a theoretical framework for capturing important characteristics of different ESD technologies, the trade-offs of placing them at different levels of the power hierarchy, and quantifying the resulting cost-benefit trade-offs as a function of workload properties. Di Wang 0003, Chuangang Ren, Anand Sivasubramaniam, Bhuvan Urgaonkar, Hosam K. Fathy |
SIGMETRICS | 4 |
| 2011 | Leveraging Value Locality in Optimizing NAND Flash-based SSDs
Raghav Pisolkar, Bhuvan Urgaonkar, Anand Sivasubramaniam |
FAST | 3 |
| 2011 | Benefits and limitations of tapping into stored energy for datacentersabstractDatacenter power consumption has a significant impact on both its recurring electricity bill (Op-ex) and one-time construction costs (Cap-ex). Existing work optimizing these costs has relied primarily on throttling devices or workload shaping, both with performance degrading implications. In this paper, we present a novel knob of energy buffer (eBuff) available in the form of UPS batteries in datacenters for this cost optimization. Intuitively, eBuff stores energy in UPS batteries during "valleys" - periods of lower demand, which can be drained during "peaks" - periods of higher demand. UPS batteries are normally used as a fail-over mechanism to transition to captive power sources upon utility failure. Furthermore, frequent discharges can cause UPS batteries to fail prematurely. We conduct detailed analysis of battery operation to figure out feasible operating regions given such battery lifetime and datacenter availability concerns. Using insights learned from this analysis, we develop peak reduction algorithms that combine the UPS battery knob with existing throttling based techniques for minimizing datacenter power costs. Using an experimental platform, we offer insights about Op-ex savings offered by eBuff for a wide range of workload peaks/valleys, UPS provisioning, and application SLA constraints. We find that eBuff can be used to realize 15-45% peak power reduction, corresponding to 6-18% savings in Op-ex across this spectrum. eBuff can also play a role in reducing Cap-ex costs by allowing tighter overbooking of power infrastructure components and we quantify the extent of such Cap-ex savings. To our knowledge, this is the first paper to exploit stored energy - typically lying untapped in the datacenter - to address the peak power draw problem. Sriram Govindan, Anand Sivasubramaniam, Bhuvan Urgaonkar |
ISCA | 3 |
| 2011 | HybridStore: A Cost-Efficient, High-Performance Storage System Combining SSDs and HDDsabstractUnlike the use of DRAM for caching or buffering, certain idiosyncrasies of SSDs make their integration into existing systems non-trivial. Flash memory suffers from limits on its reliability, is an order of magnitude more expensive than the HDD, and can sometimes be as slow as the HDD (due to excessive garbage collection (GC) induced by high intensity of random writes). Given these trade-offs between HDDs and SSDs in terms of cost, performance, and lifetime, the current consensus among several storage experts is to view SSDs not as a replacement for HDD but rather as a complementary device within the high performance storage hierarchy. We design and evaluate such a hybrid system called Hybrid Store to provide: (a) Hybrid Plan: improved capacity planning technique to administrators with the overall goal of operating within cost-budgets and (b) HybridDyn: improved performance/lifetime guarantees during episodes of deviations from expected workloads through two novel mechanisms: write-regulation and fragmentation busting. As an illustrative example of HybridStore's efficacy, Hybrid Plan is able to find the most cost-effective storage configuration for a large scale workload of Microsoft Research and suggest one MLC SSD with ten 7.2K RPM HDDs instead of fourteen 7.2K RPM HDDs only. HybridDyn is able to reduce the average response time for an enterprise scale random-write dominant workload by about 71%as compared to a HDD-based system. Youngjae Kim 0001, Bhuvan Urgaonkar, Piotr Berman, Anand Sivasubramaniam |
MASCOTS | 3 |
| 2011 | Optimal power cost management using stored energy in data centersabstractSince the electricity bill of a data center constitutes a significant portion of its overall operational costs, reducing this has become important. We investigate cost reduction opportunities that arise by the use of uninterrupted power supply (UPS) units as energy storage devices. This represents a deviation from the usual use of these devices as mere transitional fail-over mechanisms between utility and captive sources such as diesel generators. We consider the problem of opportunistically using these devices to reduce the time average electric utility bill in a data center. Using the technique of Lyapunov optimization, we develop an online control algorithm that can optimally exploit these devices to minimize the time average cost. This algorithm operates without any knowledge of the statistics of the workload or electricity cost processes, making it attractive in the presence of workload and pricing uncertainties. An interesting feature of our algorithm is that its deviation from optimality reduces as the storage capacity is increased. Our work opens up a new area in data center power management. Rahul Urgaonkar, Bhuvan Urgaonkar, Michael J. Neely, Anand Sivasubramaniam |
SIGMETRICS | 2 |
| 2011 | A comprehensive study of energy efficiency and performance of flash-based SSD
Seon-Yeong Park, Youngjae Kim 0001, Bhuvan Urgaonkar, Joonwon Lee, Euiseong Seo |
J. Syst. Archit. | 3 |
| 2011 | Bandwidth provisioning in infrastructure-based wireless networks employing directional antennas
Shiva Prasad Kasiviswanathan, Bo Zhao 0009, Sudarshan Vasudevan, Bhuvan Urgaonkar |
Pervasive Mob. Comput. | 4 |
| 2010 | Cloud Computing: A Digital Libraries PerspectiveabstractProvisioning and maintenance of infrastructure for Web based digital library search engines such as CiteSeerxpresent several challenges. CiteSeerxprovides autonomous citation indexing, full text indexing, and extensive document metadata from document scrawled from the web across computer and information sciences and related fields. Infrastructure virtualization and cloud computing are particularly attractive choices for CiteSeerx, which is challenged by both growth in the size of the indexed document collection, new features and most prominently usage. In this paper, we discuss constraints and choices faced by information retrieval systems like CiteSeerxby exploring in detail aspects of placing CiteSeerxinto current cloud infrastructure offerings. We also implement an ad-hoc virtualized storage system for experimenting with adoption of cloud infrastructure services. Our results show that a cloud implementation of CiteSeerxmay be a feasible alternative for its continued operation and growth. Pradeep B. Teregowda, Bhuvan Urgaonkar, C. Lee Giles |
IEEE CLOUD | 2 |
| 2010 | Multi-level Crypto Disk: Secondary Storage with Flexible Performance Versus Security Trade-offsabstractSecondary storage devices have become increasingly vulnerable to security attacks as they are now accessed remotely, attached to mobile devices, or used in other previously unanticipated operating environments. Storage vendors have responded to this by offering solutions that encrypt data on the fly-in software or device firmware-before recording. The performance versus security trade-off offered by these secure devices is limited due to their use of only a single level of data encryption. To address these limitations, we propose the Multi-level Crypto Disk (MLCD), a generic storage device with multiple crypto levels for encoding data. Using stochastic modeling, we derive optimal policies to dynamically select crypto levels for data in an MLCD to achieve desired performance versus security trade-offs. Shiva Chaitanya, Bhuvan Urgaonkar, Anand Sivasubramaniam |
MASCOTS | 2 |
| 2010 | Middleware for a Re-configurable Distributed Archival Store Based on Secret Sharing
Shiva Chaitanya, Dharani Vijayakumar, Bhuvan Urgaonkar, Anand Sivasubramaniam |
Middleware | 3 |
| 2010 | Power Consumption Prediction and Power-Aware Packing in Consolidated EnvironmentsabstractConsolidation of workloads has emerged as a key mechanism to dampen the rapidly growing energy expenditure within enterprise-scale data centers. To gainfully utilize consolidation-based techniques, we must be able to characterize the power consumption of groups of colocated applications. Such characterization is crucial for effective prediction and enforcement of appropriate limits on power consumption-power budgets-within the data center. We identify two kinds of power budgets: 1) an average budget to capture an upper bound on long-term energy consumption within that level and 2) a sustained budget to capture any restrictions on sustained draw of current above a certain threshold. Using a simple measurement infrastructure, we derive power profiles-statistical descriptions of the power consumption of applications. Based on insights gained from detailed profiling of several applications-both individual and consolidated-we develop models for predicting average and sustained power consumption of consolidated applications. We conduct an experimental evaluation of our techniques on a Xen-based server that consolidates applications drawn from a diverse pool. For a variety of consolidation scenarios, we are able to predict average power consumption within five percent error margin and sustained power within 10 percent error margin. Using prediction techniques allows us to ensure safe yet efficient system operation-in a representative case, we are able to improve the number of applications consolidated on a server from two to three (compared to existing baseline techniques) by choosing the appropriate power state that satisfies the power budgets associated with the server. Jeonghwan Choi, Sriram Govindan, Jinkyu Jeong, Bhuvan Urgaonkar, Anand Sivasubramaniam |
IEEE Trans. Computers | 4 |
| 2009 | DFTL: a flash translation layer employing demand-based selective caching of page-level address mappingsabstractRecent technological advances in the development of flash-memory based devices have consolidated their leadership position as the preferred storage media in the embedded systems market and opened new vistas for deployment in enterprise-scale storage systems. Unlike hard disks, flash devices are free from any mechanical moving parts, have no seek or rotational delays and consume lower power. However, the internal idiosyncrasies of flash technology make its performance highly dependent on workload characteristics. The poor performance of random writes has been a cause of major concern, which needs to be addressed to better utilize the potential of flash in enterprise-scale environments. We examine one of the important causes of this poor performance: the design of the Flash Translation Layer (FTL), which performs the virtual-to-physical address translations and hides the erase-before-write characteristics of flash. We propose a complete paradigm shift in the design of the core FTL engine from the existing techniques with our Demand-based Flash Translation Layer (DFTL), which selectively caches page-level address mappings. We develop a flash simulation framework called FlashSim. Our experimental evaluation with realistic enterprise-scale workloads endorses the utility of DFTL in enterprise-scale storage systems by demonstrating: (i) improved performance, (ii) reduced garbage collection overhead and (iii) better overload behavior compared to state-of-the-art FTL schemes. For example, a predominantly random-write dominant I/O trace from an OLTP application running at a large financial institution shows a 78% improvement in average response time (due to a 3-fold reduction in operations of the garbage collector), compared to a state-of-the-art FTL scheme. Even for the well-known read-dominant TPC-H benchmark, for which DFTL introduces additional overheads, we improve system response time by 56%. Youngjae Kim 0001, Bhuvan Urgaonkar |
ASPLOS | 3 |
| 2009 | Statistical profiling-based techniques for effective power provisioning in data centersabstractCurrent capacity planning practices based on heavy over-provisioning of power infrastructure hurt (i) the operational costs of data centers as well as (ii) the computational work they can support. We explore a combination of statistical multiplexing techniques to improve the utilization of the power hierarchy within a data center. At the highest level of the power hierarchy, we employ controlled underprovisioning and over-booking of power needs of hosted workloads. At the lower levels, we introduce the novel notion of soft fuses to flexibly distribute provisioned power among hosted workloads based on their needs. Our techniques are built upon a measurement-driven profiling and prediction framework to characterize key statistical properties of the power needs of hosted workloads and their aggregates. We characterize the gains in terms of the amount of computational work (CPU cycles) per provisioned unit of power Computation per Provisioned Watt (CPW). Our technique is able to double the CPWoffered by a Power Distribution Unit (PDU) running the e-commerce benchmark TPC-W compared to conventional provisioning practices. Over-booking the PDU by 10% based on tails of power profiles yields a further improvement of 20%. Reactive techniques implemented on our Xen VMM-based servers dynamically modulate CPU DVFS states to ensure power draw below the limits imposed by soft fuses. Finally, information captured in our profiles also provide ways of controlling application performance degradation despite overbooking. The 95th percentile of TPC-W session response time only grew from 1.59 sec to 1.78 sec--a degradation of 12%. Sriram Govindan, Jeonghwan Choi, Bhuvan Urgaonkar, Anand Sivasubramaniam, Andrea Baldini |
EuroSys | 3 |
| 2009 | vPath: Precise Discovery of Request Processing Paths from Black-Box Observations of Thread and Network Activities
Byung-Chul Tak, Chunqiang Tang, Sriram Govindan, Bhuvan Urgaonkar, Rong Chang 0001 |
USENIX ATC | 5 |
| 2009 | Virtualized Data Centers
Krishna Kant 0001, Bhuvan Urgaonkar |
Comput. Networks | 2 |
| 2009 | Xen and Co.: Communication-Aware CPU Management in Consolidated Xen-Based Hosting PlatformsabstractRecent advances in software and architectural support for server virtualization have created interest in using this technology in the design of consolidated hosting platforms. Since virtualization enables easier and faster application migration as well as secure colocation of antagonistic applications, higher degrees of server consolidation are likely to result in such virtualization-based hosting platforms (VHPs). We identify two shortcomings in existing virtual machine monitors (VMMs) that prove to be obstacles in operating hosting platforms, such as Internet data centers, under conditions of such high consolidation: 1) CPU schedulers that are agnostic to the communication behavior of modern, multitier applications and 2) inadequate or inaccurate mechanisms for accounting the CPU overheads of I/O virtualization. We develop a new communication-aware CPU scheduling algorithm and a CPU usage accounting mechanism. We implement our algorithms in the Xen VMM and build a prototype VHP on a cluster of 36 servers. Our experimental evaluation with realistic Internet server applications and benchmarks demonstrates the performance/cost benefits and the wide applicability of our algorithms. For example, the TPC-W benchmark exhibited improvements in average response times between 20 percent and 35 percent for a variety of consolidation scenarios. A streaming media server hosted on our prototype VHP was able to satisfactorily service up to 3.5 times as many clients as one running on the default Xen. Sriram Govindan, Jeonghwan Choi, Arjun R. Nath, Amitayu Das, Bhuvan Urgaonkar, Anand Sivasubramaniam |
IEEE Trans. Computers | 5 |
| 2009 | Resource overbooking and application profiling in a shared Internet hosting platformabstractIn this article, we present techniques for provisioning CPU and network resources in shared Internet hosting platforms running potentially antagonistic third-party applications. The primary contribution of our work is to demonstrate the feasibility and benefits of overbooking resources in shared Internet platforms. Since an accurate estimate of an application's resource needs is necessary when overbooking resources, we present techniques to profile applications on dedicated nodes, possibly while in service, and use these profiles to guide the placement of application components onto shared nodes. We then propose techniques to overbook cluster resources in a controlled fashion. We outline an empirical appraoch to determine the degree of overbooking that allows a platform to achieve improvements in revenue while providing performance guarantees to Internet applications. We show how our techniques can be combined with commonly used QoS resource allocation mechanisms to provide application isolation and performance guarantees at run-time. We implement our techniques in a Linux cluster and evaluate them using common server applications. We find that the efficiency (and consequently revenue) benefits from controlled overbooking of resources can be dramatic. Specifically, we find that overbooking resources by as little as 1% we can increase the utilization of the cluster by a factor of two, and a 5% overbooking yields a 300--500% improvement, while still providing useful resource guarantees to applications. Bhuvan Urgaonkar, Prashant J. Shenoy, Timothy Roscoe |
ACM Trans. Internet Techn. | 1 |
| 2008 | Evaluating the usefulness of content addressable storage for high-performance data intensive applicationsabstractContent Addressable Storage (CAS) is a data representation technique that operates by partitioning a given data-set into non-intersecting units called chunks and then employing techniques to efficiently recognize chunks occurring multiple times. This allows CAS to eliminate duplicate instances of such chunks, resulting in reduced storage space compared to conventional representations of data. CAS is an attractive technique for reducing the storage and network bandwidth needs of performance-sensitive, data-intensive applications in a variety of domains. These include enterprise applications, Web-based e-commerce or entertainment services and highly parallel scientific/engineering applications and simulations, to name a few. Partho Nath, Bhuvan Urgaonkar, Anand Sivasubramaniam |
HPDC | 2 |
| 2008 | Project status: RIVER: Resource management infrastructure for consolidated hosting in virtualized data centersabstractThis paper describes the status of the NSF-funded project titled "River. Resource Management Infrastructure for Consolidated Hosting in Virtualized Data Centers" CNS-0720456. Bhuvan Urgaonkar, Anand Sivasubramaniam |
IPDPS | 1 |
| 2008 | Profiling, Prediction, and Capping of Power Consumption in Consolidated Environments
Jeonghwan Choi, Sriram Govindan, Bhuvan Urgaonkar, Anand Sivasubramaniam |
MASCOTS | 3 |
| 2008 | QDSL: a queuing model for systems with differential service levelsabstractA feature exhibited by many modern computing systems is their ability to improve the quality of output they generate for a given input by spending more computing resources on processing it. Often this improvement comes at the price of degraded performance in the form of reduced throughput or increased response time. We formulate QDSL, a class of constrained optimization problems defined in the context of a queueing server equipped with multiple levels of service. Solutions to QDSL provide rules for dynamically varying the service level to achieve desired trade-offs between output quality and performance. Our approach involves reducing restricted versions of such systems to Markov Decision Processes. We find two variants of such systems worth studying: (i) VarSL, in which a single request may be serviced using a combination of multiple levels during its lifetime and (ii) FixSL in which the service level may not change during the lifetime of a request. Our modeling indicates that optimal service level selection policies in these systems correspond to very simple rules that can be implemented very efficiently in realistic, online systems. We find our policies to be useful in two response-time-sensitive real-world systems: (i) qSecStore, an iSCSI-based secure storage system that has access to multiple encryption functions, and (ii) qPowServer, a server with DVFS-capable processor. As a representative result, in an instance of qSecStore serving disk requests derived from the well-regarded TPC-H traces, we are able to improve the fraction of requests using more reliable encryption functions by 40-60%, while meeting performance targets. In a simulation of qPowServer employing realistic DVFS parameters, we are able to improve response times significantly while only violating specified server-wide power budgets by less than 5W. Shiva Chaitanya, Bhuvan Urgaonkar, Anand Sivasubramaniam |
SIGMETRICS | 2 |
| 2008 | Towards event source unobservability with minimum network traffic in sensor networksabstractSensors deployed to monitor the surrounding environment report such information as event type, location, and time when a real event of interest is detected. An adversary may identify the real event source through eavesdropping and traffic analysis. Previous work has studied the source location privacy problem under a local adversary model. In this work, we aim to provide a stronger notion: event source unobservability, which promises that a global adversary cannot know whether a real event has ever occurred even if he is capable of collecting and analyzing all the messages in the network at all the time. Clearly, event source unobservability is a desirable and critical security property for event monitoring applications, but unfortunately it is also very difficult and expensive to achieve for resource-constrained sensor network. Yi Yang 0002, Sencun Zhu, Bhuvan Urgaonkar, Guohong Cao |
WISEC | 4 |
| 2008 | Cataclysm: Scalable overload policing for internet applications
Bhuvan Urgaonkar, Prashant J. Shenoy |
J. Netw. Comput. Appl. | 1 |
| 2008 | Agile dynamic provisioning of multi-tier Internet applicationsabstractDynamic capacity provisioning is a useful technique for handling the multi-time-scale variations seen in Internet workloads. In this article, we propose a novel dynamic provisioning technique for multi-tier Internet applications that employs (1) a flexible queuing model to determine how much of the resources to allocate to each tier of the application, and (2) a combination of predictive and reactive methods that determine when to provision these resources, both at large and small time scales. We propose a novel data center architecture based on virtual machine monitors to reduce provisioning overheads. Our experiments on a forty-machine Xen/Linux-based hosting platform demonstrate the responsiveness of our technique in handling dynamic workloads. In one scenario where a flash crowd caused the workload of a three-tier application to double, our technique was able to double the application capacity within five minutes, thus maintaining response-time targets. Our technique also reduced the overhead of switching servers across applications from several minutes to less than a second, while meeting the performance targets of residual sessions. Bhuvan Urgaonkar, Prashant J. Shenoy, Abhishek Chandra, Pawan Goyal 0001, Timothy Wood 0001 |
ACM Trans. Auton. Adapt. Syst. | 1 |
| 2007 | Optimizing Energy-Efficient Query Processing in Wireless Sensor NetworksabstractThis paper studies the issues of energy-efficient query optimization for wireless sensor networks. Different from existing query optimization techniques that consider only query plans for extracting data from sensors at individual nodes, our approach takes into account both of the sensing and communication cost in query plans. Central to our study is a cost-based analysis, based on which the energy cost of candidate plans for a given query are estimated to determine a query plan that is likely to consume the least energy for execution. Simulation results show that the query plan chosen in our approach consumes significantly less energy than an approach that optimizes on sensing cost only. Ross Rosemark, Wang-Chien Lee, Bhuvan Urgaonkar |
MDM | 3 |
| 2007 | Packing to angles and sectorsabstractMotivated by the widespread proliferation of wireless networks employing directional antennas, we study some capacitated covering problems arising in these networks. Geometrically, the area covered by a directional antenna with parameters α,ρ,r is a set of points with polar coordinates (r,θ) such that r ≤ r and α ≤ θ ≤ α + ρ. Given a set of customers, their positions on the plane and their bandwidth demands, the capacitated covering problem considered here is to cover all the customers with the minimum number of directional antennas such that the demands of customers assigned to an antenna stays within a bound. We consider two settings of this capacitated cover problem arising in wireless networks. In the first setting where the antennas have variable angular range, we present an approximation algorithm with ratio 3. In the setting where the angular range of antennas is fixed, we improve this approximation ratio to 1.5. Piotr Berman, Jieun K. Jeong, Shiva Prasad Kasiviswanathan, Bhuvan Urgaonkar |
SPAA | 4 |
| 2007 | Xen and co.: communication-aware CPU scheduling for consolidated xen-based hosting platformsabstractRecent advances in software and architectural support for server virtualization have created interest in using this technology in the design of consolidated hosting platforms. Since virtualization enables easier and faster application migration as well as secure co-location of antagonistic applications, higher degrees of server consolidation are likely to result in such virtualization-based hosting platforms (VHPs). We identify a key shortcoming in existing virtual machine monitors (VMMs) that proves to be an obstacle in operating hosting platforms, such as Internet data centers, under conditions of such high consolidation: CPU schedulers that are agnostic to the communication behavior of modern, multi-tier applications. We develop a new communication-aware CPU scheduling algorithm to alleviate this problem. We implement our algorithm in the Xen VMM and build a prototype VHP on a cluster of servers. Our experimental evaluation with realistic Internet server applications and benchmarks demonstrates the performance/cost benefits and the wide applicability of our algorithms. For example, the TPC-W benchmark exhibited improvements in average response times of up to 35% for a variety of consolidation scenarios. A streaming media server hosted on our prototype VHP was able to satisfactorily service up to 3.5 times as many clients as one running on the default Xen. Sriram Govindan, Arjun R. Nath, Amitayu Das, Bhuvan Urgaonkar, Anand Sivasubramaniam |
VEE | 4 |
| 2007 | Analytic modeling of multitier Internet applicationsabstractSince many Internet applications employ a multitier architecture, in this article, we focus on the problem of analytically modeling the behavior of such applications. We present a model based on a network of queues where the queues represent different tiers of the application. Our model is sufficiently general to capture (i) the behavior of tiers with significantly different performance characteristics and (ii) application idiosyncrasies such as session-based workloads, tier replication, load imbalances across replicas, and caching at intermediate tiers. We validate our model using real multitier applications running on a Linux server cluster. Our experiments indicate that our model faithfully captures the performance of these applications for a number of workloads and configurations. Furthermore, our model successfully handles a comprehensive range of resource utilization---from 0 to near saturation for the CPU---for two separate tiers. For a variety of scenarios, including those with caching at one of the application tiers, the average response times predicted by our model were within the 95% confidence intervals of the observed average response times. Our experiments also demonstrate the utility of the model for dynamic capacity provisioning, performance prediction, bottleneck identification, and session policing. In one scenario, where the request arrival rate increased from less than 1500 to nearly 4200 requests/minute, a dynamic provisioning technique employing our model was able to maintain response time targets by increasing the capacity of two of the tiers by factors of 2 and 3.5, respectively. Bhuvan Urgaonkar, Giovanni Pacifici, Prashant J. Shenoy, Mike Spreitzer, Asser N. Tantawi |
ACM Trans. Web | 1 |
| 2006 | Dynamic cache reconfiguration strategies for cluster-based streaming proxy
Yang Guo 0001, Zihui Ge, Bhuvan Urgaonkar, Prashant J. Shenoy, Don Towsley |
Comput. Commun. | 3 |
| 2005 | An analytical model for multi-tier internet services and its applicationsabstractSince many Internet applications employ a multi-tier architecture, in this paper, we focus on the problem of analytically modeling the behavior of such applications. We present a model based on a network of queues, where the queues represent different tiers of the application. Our model is sufficiently general to capture (i) the behavior of tiers with significantly different performance characteristics and (ii) application idiosyncrasies such as session-based workloads, concurrency limits, and caching at intermediate tiers. We validate our model using real multi-tier applications running on a Linux server cluster. Our experiments indicate that our model faithfully captures the performance of these applications for a number of workloads and configurations. For a variety of scenarios, including those with caching at one of the application tiers, the average response times predicted by our model were within the 95% confidence intervals of the observed average response times. Our experiments also demonstrate the utility of the model for dynamic capacity provisioning, performance prediction, bottleneck identification, and session policing. In one scenario, where the request arrival rate increased from less than 1500 to nearly 4200 requests/min, a dynamic provisioning technique employing our model was able to maintain response time targets by increasing the capacity of two of the application tiers by factors of 2 and 3.5, respectively. Bhuvan Urgaonkar, Giovanni Pacifici, Prashant J. Shenoy, Mike Spreitzer, Asser N. Tantawi |
SIGMETRICS | 1 |
| 2005 | Cataclysm: policing extreme overloads in internet applicationsabstractIn this paper we present the Cataclysm server platform for handling extreme overloads in hosted Internet applications. The primary contribution of our work is to develop a low overhead, highly scalable admission control technique for Internet applications. Cataclysm provides several desirable features, such as guarantees on response time by conducting accurate size-based admission control, revenue maximization at multiple time-scales via preferential admission of important requests and dynamic capacity provisioning, and the ability to be operational even under extreme overloads. Cataclysm can transparently trade-off the accuracy of its decision making with the intensity of the workload allowing it to handle incoming rates of several tens of thousands of requests/second. We implement a prototype Cataclysm hosting platform on a Linux cluster and demonstrate the benefits of our integrated approach using a variety of workloads. Bhuvan Urgaonkar, Prashant J. Shenoy |
WWW | 1 |
| 2004 | Brief announcement: Cataclysm: handling extreme overloads in internet servicesabstractIn this paper we present Cataclysm, a comprehensive approach for handling extreme overloads in hosted Internet applications. The primary contribution of our work is to develop an overload control approach that brings together admission control, dynamic provisioning of platform resources, and adaptive degradation of QoS into one integrated system. We implement a prototype Cataclysm hosting platform on a Linux cluster and demonstrate the benefits of our integrated approach using a variety of workloads. Bhuvan Urgaonkar, Prashant J. Shenoy |
PODC | 1 |
| 2004 | Sharc: Managing CPU and Network Bandwidth in Shared ClustersabstractWe argue the need for effective resource management mechanisms for sharing resources in commodity clusters. To address this issue, we present the design of Sharc-a system that enables resource sharing among applications in such clusters. Sharc depends on single node resource management mechanisms such as reservations or shares, and extends the benefits of such mechanisms to clustered environments. We present techniques for managing two important resources-CPU and network interface bandwidth-on a cluster-wide basis. Our techniques allow Sharc to 1) support reservation of CPU and network interface bandwidth for distributed applications, 2) dynamically allocate resources based on past usage, and 3) provide performance isolation to applications. Our experimental evaluation has shown that Sharc can scale to 256 node clusters running 100,000 applications. These results demonstrate that Sharc can be an effective approach for sharing resources among competing applications in moderate size clusters. Bhuvan Urgaonkar, Prashant J. Shenoy |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2002 | Resource Overbooking and Application Profiling in Shared Hosting Platforms
Bhuvan Urgaonkar, Prashant J. Shenoy, Timothy Roscoe |
OSDI | 1 |
| 2001 | Maintaining Mutual Consistency for Cached Web ObjectsabstractExisting Web proxy caches employ cache consistency mechanisms to ensure that locally cached data is consistent with that at the server. We argue that techniques for maintaining consistency of individual objects are not sufficient; a proxy should employ additional mechanisms to ensure that related Web objects are mutually consistent with one another. We formally define the notion of mutual consistency and the semantics provided by a mutual consistency mechanism to end users. We then present techniques for maintaining mutual consistency in the temporal and value domains. A novel aspect of our techniques is that they can adapt to the variations in the rate of change of the source data, resulting in judicious use of proxy and network resources. We evaluate our approaches using real-world Web traces and show that: (i) careful tuning can result in substantial savings in the network overhead incurred without any substantial loss in fidelity, of the consistency guarantees, and (ii) the incremental cost of providing mutual consistency guarantees over mechanisms to provide individual consistency guarantees is small. Bhuvan Urgaonkar, Anoop George Ninan, M. S. Raunak 0001, Prashant J. Shenoy, Krithi Ramamritham |
ICDCS | 1 |