Sriram Govindan

dblp:18/4286 · DBLP profile ↗
← Back
20ranked-venue papers
7as first author
0since 2021 · last 2020
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 18 · 7 first-authorSoftware engineering, systems software and programming languages · 8 · 2 first-authorSecurity and privacy · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
12 papers
Energy-efficient computing · 58% Cloud and datacenter computing · 19% Storage systems · 7%
Computer networks
1 paper
Network measurement and analytics · 100%

Topics — the 21 heaviest of 27, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Energy-efficient computing
datacenter power management
0.962016
SizeCap: Efficiently handling power surges in fuel cell powered data centers · HPCA 2016
Aggressive Datacenter Power Provisioning with Batteries · ACM Trans. Comput. Syst. 2013
ACE: abstracting, characterizing and exploiting peaks and valleys in datacenter power consumption · SIGMETRICS 2013
Energy-efficient computing
power management
0.842017
Viyojit: Decoupling Battery and DRAM Capacities for Battery-Backed DRAM · ISCA 2017
Underprovisioning backup power infrastructure for datacenters · ASPLOS 2014
Aggressive Datacenter Power Provisioning with Batteries · ACM Trans. Comput. Syst. 2013
Energy-efficient computing › power management
power capping
0.742016
SizeCap: Efficiently handling power surges in fuel cell powered data centers · HPCA 2016
Aggressive Datacenter Power Provisioning with Batteries · ACM Trans. Comput. Syst. 2013
ACE: abstracting, characterizing and exploiting peaks and valleys in datacenter power consumption · SIGMETRICS 2013
Cloud and datacenter computing
cloud storage
0.412020
LeapIO: Efficient and Portable Virtual NVMe Storage on ARM SoCs · ASPLOS 2020
Energy-efficient computing › datacenter power management
power provisioning
0.432013
Aggressive Datacenter Power Provisioning with Batteries · ACM Trans. Comput. Syst. 2013
Leveraging stored energy for handling power emergencies in aggressively provisioned datacenters · ASPLOS 2012
Statistical profiling-based techniques for effective power provisioning in data centers · EuroSys 2009
Memory systems
DRAM
0.312017
Viyojit: Decoupling Battery and DRAM Capacities for Battery-Backed DRAM · ISCA 2017
Cloud and datacenter computing › virtualization › virtual machine management
server consolidation
0.222010
Power Consumption Prediction and Power-Aware Packing in Consolidated Environments · IEEE Trans. Computers 2010
Xen and Co.: Communication-Aware CPU Management in Consolidated Xen-Based Hosting Platforms · IEEE Trans. Computers 2009
Performance modeling and evaluation
workload characterization
0.222013
ACE: abstracting, characterizing and exploiting peaks and valleys in datacenter power consumption · SIGMETRICS 2013
Power Consumption Prediction and Power-Aware Packing in Consolidated Environments · IEEE Trans. Computers 2010
Cloud and datacenter computing
datacenter infrastructure
0.212014
Underprovisioning backup power infrastructure for datacenters · ASPLOS 2014
Energy-efficient computing
datacenter energy consumption
0.112011
Benefits and limitations of tapping into stored energy for datacenters · ISCA 2011
Energy-efficient computing › power management › peak power management
peak power reduction
0.112011
Benefits and limitations of tapping into stored energy for datacenters · ISCA 2011
Cloud and datacenter computing › resource management
datacenter resource management
0.112010
Power Consumption Prediction and Power-Aware Packing in Consolidated Environments · IEEE Trans. Computers 2010
Energy-efficient computing
power prediction
0.112010
Power Consumption Prediction and Power-Aware Packing in Consolidated Environments · IEEE Trans. Computers 2010
Distributed systems › observability › distributed monitoring
distributed tracing
0.112009
vPath: Precise Discovery of Request Processing Paths from Black-Box Observations of Thread and Network Activities · USENIX ATC 2009
Energy-efficient computing › datacenter power management
power oversubscription
0.112009
Statistical profiling-based techniques for effective power provisioning in data centers · EuroSys 2009
Cloud and datacenter computing
virtualization
0.112009
Xen and Co.: Communication-Aware CPU Management in Consolidated Xen-Based Hosting Platforms · IEEE Trans. Computers 2009
Hardware reliability and fault tolerance
uninterruptible power supply
0.012013
Aggressive Datacenter Power Provisioning with Batteries · ACM Trans. Comput. Syst. 2013
Cloud and datacenter computing › datacenter operations
datacenter workload management
0.012012
Leveraging stored energy for handling power emergencies in aggressively provisioned datacenters · ASPLOS 2012
Cloud and datacenter computing › datacenter operations
datacenter cost optimization
0.012011
Benefits and limitations of tapping into stored energy for datacenters · ISCA 2011
Electronic design automation › power analysis
power profiling
0.012010
Power Consumption Prediction and Power-Aware Packing in Consolidated Environments · IEEE Trans. Computers 2010
Operating systems › resource management › process management
CPU scheduling
0.012009
Xen and Co.: Communication-Aware CPU Management in Consolidated Xen-Based Hosting Platforms · IEEE Trans. Computers 2009

Methods — techniques the papers use, named apart from their topics

online heuristics · 0.3trace-driven simulation · 0.2cost modeling · 0.2availability analysis · 0.2statistical analysis · 0.2workload throttling · 0.1peak reduction algorithm · 0.1xen virtualization · 0.1statistical power modeling · 0.1communication-aware scheduling · 0.1DVFS · 0.1CPU usage accounting · 0.1
YearPublicationVenuePosition
2020 LeapIO: Efficient and Portable Virtual NVMe Storage on ARM SoCs
abstract
Today's cloud storage stack is extremely resource hungry, burning 10-20% of datacenter x86 cores, a major "storage tax" that cloud providers must pay. Yet, the complex cloud storage stack is not completely offload-ready to today's IO accelerators. We present LeapIO, a new cloud storage stack that leverages ARM-based co-processors to offload complex storage services. LeapIO addresses many deployment challenges, such as hardware fungibility, software portability, virtualizability, composability, and efficiency. It uses a set of OS/software techniques and new hardware properties that provide a uni- form address space across the x86 and ARM cores and ex- pose virtual NVMe storage to unmodified guest VMs, at a performance that is competitive with bare-metal servers.
Huaicheng Li, Mingzhe Hao, Stanko Novakovic, Vaibhav Gogte, Sriram Govindan, Dan R. K. Ports, Irene Zhang, Ricardo Bianchini, Haryadi S. Gunawi, Anirudh Badam
ASPLOS5
2019 Getting more performance with polymorphism from emerging memory technologies
abstract
Storage-intensive systems in data centers rely heavily on DRAM and SSDs for the performance of reads and persistent writes, respectively. These applications pose a diverse set of requirements, and are limited by fixed capacity, fixed access latency, and fixed function of these resources as either memory or storage. In contrast, emerging memory technologies like 3D-Xpoint, battery-backed DRAM, and ASIC-based fast memory-compression offer capabilities across several dimensions. However, existing proposals to use such technologies can only improve either read or write performance but not both without requiring extensive changes to the application, and the operating system. We present PolyEMT, a system that employs an emerging memory technology based cache to the SSD, and transparently morphs the capabilities of this cache across several dimensions - persistence, capacity, latency - to jointly improve both read and write performance. We demonstrate the benefits of PolyEMT using several large-scale storage-intensive workloads from our datacenters.
Iyswarya Narayanan, Aishwarya Ganesan, Anirudh Badam, Sriram Govindan, Bikash Sharma, Anand Sivasubramaniam
SYSTOR4
2017 Rain or Shine? - Making Sense of Cloudy Reliability Data
abstract
Cloud datacenters must ensure high availability for the hosted applications and failures can be the bane of datacenter operators. Understanding the what, when and why of failures can help tremendously to mitigate their occurrence and impact. Failures can, however, depend on numerous spatial and temporal factors spanning hardware, workloads, support facilities, and even the environment. One has to rely on failure data from the field to quantify the influence of these factors on failures. Towards this goal, we collect failures data along with many parameters that might influence failures from two large production datacenters with very diverse characteristics. We show that multiple factors simultaneously affect failures, and these factors may interact in non-trivial ways. This makes conventional approaches that study aggregate characteristics or single parameter influences, rather inaccurate. Instead, we build a multi-factor analysis framework to systematically identify influencing factors, quantify their relative impact, and help in more accurate decision making for failure mitigation. We demonstrate this approach for three important decisions: spare capacity provisioning, comparing the reliability of hardware for vendor selection, and quantifying flexibility in datacenter climate control for cost-reliability trade-offs.
Iyswarya Narayanan, Bikash Sharma, Di Wang 0003, Sriram Govindan, Laura Caulfield, Anand Sivasubramaniam, Aman Kansal, Jie Liu 0001, Badriddine M. Khessib, Kushagra Vaid
ICDCS4
2017 Viyojit: Decoupling Battery and DRAM Capacities for Battery-Backed DRAM
Rajat Kateja, Anirudh Badam, Sriram Govindan, Bikash Sharma, Gregory R. Ganger
ISCA3
2016 SizeCap: Efficiently handling power surges in fuel cell powered data centers
abstract
Fuel cells are a promising power source for future data centers, offering high energy efficiency, low greenhouse gas emissions, and high reliability. However, due to mechanical limitations related to fuel delivery, fuel cells are slow to adjust to sudden increases in data center power demands, which can result in temporary power shortfalls. To mitigate the impact of power shortfalls, prior work has proposed to either perform power capping by throttling the servers, or to leverage energy storage devices (ESDs) that can temporarily provide enough power to make up for the shortfall while the fuel cells ramp up power generation. Both approaches have disadvantages: power capping conservatively limits server performance and can lead to service level agreement (SLA) violations, while ESD-only solutions must significantly overprovision the energy storage device capacity to tolerate the shortfalls caused by the worst-case (i.e., largest) power surges, which greatly increases the total cost of ownership (TCO). We propose SizeCap, the first ESD sizing framework for fuel cell powered data centers, which coordinates ESD sizing with power capping to enable a cost-effective solution to power shortfalls in data centers. SizeCap sizes the ESD just large enough to cover the majority of power surges, but not the worst-case surges that occur infrequently, to greatly reduce TCO. It then uses the smaller capacity ESD in conjunction with power capping to cover the power shortfalls caused by the worst-case power surges. As part of our new flexible framework, we propose multiple power capping policies with different degrees of awareness of fuel cell and workload behavior, and evaluate their impact on workload performance and ESD size. Using traces from Microsoft's production data center systems, we demonstrate that SizeCap significantly reduces the ESD size (by 85%ofor a workload with infrequent yet large power surges, and by 50% for a workload with frequent power surges) without violating any SLAs.
Yang Li 0183, Di Wang 0003, Saugata Ghose, Jie Liu 0001, Sriram Govindan, Sean James, Eric Peterson, John Siegler, Rachata Ausavarungnirun, Onur Mutlu
HPCA5
2016 X-Mem: A cross-platform and extensible memory characterization tool for the cloud
abstract
Effective use of the memory hierarchy is crucial to cloud computing. Platform memory subsystems must be carefully provisioned and configured to minimize overall cost and energy for cloud providers. For cloud subscribers, the diversity of available platforms complicates comparisons and the optimization of performance. To address these needs, we present X-Mem, a new open-source software tool that characterizes the memory hierarchy for cloud computing.
Mark Gottscho, Sriram Govindan, Bikash Sharma, Mohammed Shoaib, Puneet Gupta 0001
ISPASS2
2014 Underprovisioning backup power infrastructure for datacenters
abstract
While there has been prior work to underprovision the power distribution infrastructure for a datacenter to save costs, the ability to underprovision the backup power infrastructure, which contributes significantly to capital costs, is little explored. There are two main components in the backup infrastructure - Diesel Generators (DGs) and UPS units - which can both be underprovisioned (or even removed) in terms of their power and/or energy capacities. However, embarking on such underprovisioning mandates studying several ramifications - the resulting cost savings, the lower availability, and the performance and state loss consequences on individual applications - concurrently. This paper presents the first such study, considering cost, availability, performance and application consequences of underprovisioning the backup power infrastructure. We present a framework to quantify the cost of backup capacity that is provisioned, and implement techniques leveraging existing software and hardware mechanisms to provide as seamless an operation as possible for an application within the provisioned backup capacity during a power outage. We evaluate the cost-performance-availability trade-offs for different levels of backup underprovisioning for applications with diverse reliance on the backup infrastructure. Our results show that one may be able to completely do away with DGs, compensating for it with additional UPS energy capacities, to significantly cut costs and still be able to handle power outages lasting as high as 40 minutes (which constitute bulk of the outages). Further, we can push the limits of outage duration that can be handled in a cost-effective manner, if applications are willing to tolerate degraded performance during the outage. Our evaluations also show that different applications react differently to the outage handling mechanisms, and that the efficacy of the mechanisms is sensitive to the outage duration. The insights from this paper can spur new opportunities for future work on backup power infrastructure optimization.
Di Wang 0003, Sriram Govindan, Anand Sivasubramaniam, Aman Kansal, Jie Liu 0001, Badriddine M. Khessib
ASPLOS2
2014 Characterizing Application Memory Error Vulnerability to Optimize Datacenter Cost via Heterogeneous-Reliability Memory
abstract
Memory devices represent a key component of datacenter total cost of ownership (TCO), and techniques used to reduce errors that occur on these devices increase this cost. Existing approaches to providing reliability for memory devices pessimistically treat all data as equally vulnerable to memory errors. Our key insight is that there exists a diverse spectrum of tolerance to memory errors in new data-intensive applications, and that traditional one-size-fits-all memory reliability techniques are inefficient in terms of cost. For example, we found that while traditional error protection increases memory system cost by 12.5%, some applications can achieve 99.00% availability on a single server with a large number of memory errors without any error protection. This presents an opportunity to greatly reduce server hardware cost by provisioning the right amount of memory reliability for different applications. Toward this end, in this paper, we make three main contributions to enable highly-reliable servers at low datacenter cost. First, we develop a new methodology to quantify the tolerance of applications to memory errors. Second, using our methodology, we perform a case study of three new dataintensive workloads (an interactive web search application, an in-memory key -- value store, and a graph mining framework) to identify new insights into the nature of application memory error vulnerability. Third, based on our insights, we propose several new hardware/software heterogeneous-reliability memory system designs to lower datacenter cost while achieving high reliability and discuss their trade-off. We show that our new techniques can reduce server hardware cost by 4.7% while achieving 99.90% single server availability.
Sriram Govindan, Bikash Sharma, Mark Santaniello, Justin Meza, Aman Kansal, Jie Liu 0001, Badriddine M. Khessib, Kushagra Vaid, Onur Mutlu
DSN2
2013 Using Dark Fiber to Displace Diesel Generators
Aman Kansal, Bhuvan Urgaonkar, Sriram Govindan
HotOS3
2013 ACE: abstracting, characterizing and exploiting peaks and valleys in datacenter power consumption
abstract
Peak power management of datacenters has tremendous cost implications. While numerous mechanisms have been proposed to cap power consumption, real datacenter power consumption data is scarce. To address this gap, we collect power demands at multiple spatial and fine-grained temporal resolutions from the load of geo-distributed datacenters of Microsoft over 6 months. We conduct aggregate analysis of this data, to study its statistical properties. With workload characterization a key ingredient for systems design and evaluation, we note the importance of better abstractions for capturing power demands, in the form of peaks and valleys. We identify and characterize attributes for peaks and valleys, and important correlations across these attributes that can influence the choice and effectiveness of different power capping techniques. With the wide scope of exploitability of such characteristics for power provisioning and optimizations, we illustrate its benefits with two specific case studies.
Di Wang 0003, Chuangang Ren, Sriram Govindan, Anand Sivasubramaniam, Bhuvan Urgaonkar, Aman Kansal, Kushagra Vaid
SIGMETRICS3
2013 Aggressive Datacenter Power Provisioning with Batteries
abstract
Datacenters spend $10--25 per watt in provisioning their power infrastructure, regardless of the watts actually consumed. Since peak power needs arise rarely, provisioning power infrastructure for them can be expensive. One can, thus, aggressively underprovision infrastructure assuming that simultaneous peak draw across all equipment will happen rarely. The resulting nonzero probability of emergency events where power needs exceed provisioned capacity, however small, mandates graceful reaction mechanisms to cap the power draw instead of leaving it to disruptive circuit breakers/fuses. Existing strategies for power capping use temporal knobs local to a server that throttle the rate of execution (using power modes), and/or spatial knobs that redirect/migrate excess load to regions of the datacenter with more power headroom. We show these mechanisms to have performance degrading ramifications, and propose an entirely orthogonal solution that leverages existing UPS batteries to temporarily augment the utility supply during emergencies. We build an experimental prototype to demonstrate such power capping on a cluster of 8 servers, each with an individual battery, and implement several online heuristics in the context of different datacenter workloads to evaluate their effectiveness in handling power emergencies. We show that our battery-based solution can: (i) handle emergencies of short durations on its own, (ii) supplement existing reaction mechanisms to enhance their efficacy for longer emergencies, and (iii) create more slack for shifting applications temporarily to nonpeak durations.
Sriram Govindan, Di Wang 0003, Anand Sivasubramaniam, Bhuvan Urgaonkar
ACM Trans. Comput. Syst.1
2012 Leveraging stored energy for handling power emergencies in aggressively provisioned datacenters
abstract
Datacenters spend $10-25 per watt in provisioning their power infrastructure, regardless of the watts actually consumed. Since peak power needs arise rarely, provisioning power infrastructure for them can be expensive. One can, thus, aggressively under-provision infrastructure assuming that simultaneous peak draw across all equipment will happen rarely. The resulting non-zero probability of emergency events where power needs exceed provisioned capacity, however small, mandates graceful reaction mechanisms to cap the power draw instead of leaving it to disruptive circuit breakers/fuses. Existing strategies for power capping use temporal knobs local to a server that throttle the rate of execution (using power modes), and/or spatial knobs that redirect/migrate excess load to regions of the datacenter with more power headroom. We show these mechanisms to have performance degrading ramifications, and propose an entirely orthogonal solution that leverages existing UPS batteries to temporarily augment the utility supply during emergencies. We build an experimental prototype to demonstrate such power capping on a cluster of 8 servers, each with an individual battery, and implement several online heuristics in the context of different datacenter workloads to evaluate their effectiveness in handling power emergencies. We show that: (i) our battery-based solution can handle emergencies of short duration on its own, (ii) supplement existing reaction mechanisms to enhance their efficacy for longer emergencies, and (iii) battery even provide feasible options when other knobs do not suffice.
Sriram Govindan, Di Wang 0003, Anand Sivasubramaniam, Bhuvan Urgaonkar
ASPLOS1
2011 Cuanta: quantifying effects of shared on-chip resource interference for consolidated virtual machines
abstract
Workload consolidation is very attractive for cloud platforms due to several reasons including reduced infrastructure costs, lower energy consumption, and ease of management. Advances in virtualization hardware and software continue to improve resource isolation among consolidated workloads but a particular form of resource interference is yet to see a commercially widely adopted solution - the interference due to shared processor caches. Existing solutions for handling cache interference require new hardware features, extensive software changes, or reduce the achieved overall throughput. A crucial requirement for effective consolidation is to be able to predict the impact of cache interference among consolidated workloads. In this paper, we present a practical technique for predicting performance interference due to shared processor cache which works on current processor architectures and requires minimal software changes. While performance degradation can be empirically measured for a given placement of consolidated workloads, the number of possible placements grows exponentially with the number of workloads and actual measurement of degradation is thus not practical for every possible placement. Our technique predicts the degradation for any possible placement using only a linear number of measurements, and can be used to select the most efficient consolidation pattern, for required performance and resource constraints. An average prediction error of less than 4% is achieved across a wide variety of benchmark workloads, using Xen VMM on Intel Core 2 Duo and Nehalem quad-core processor platforms. We also illustrate the usefulness of our prediction technique in realizing better workload placement decisions for given performance and resource cost objectives.
Sriram Govindan, Jie Liu 0001, Aman Kansal, Anand Sivasubramaniam
SoCC1
2011 Benefits and limitations of tapping into stored energy for datacenters
abstract
Datacenter power consumption has a significant impact on both its recurring electricity bill (Op-ex) and one-time construction costs (Cap-ex). Existing work optimizing these costs has relied primarily on throttling devices or workload shaping, both with performance degrading implications. In this paper, we present a novel knob of energy buffer (eBuff) available in the form of UPS batteries in datacenters for this cost optimization. Intuitively, eBuff stores energy in UPS batteries during "valleys" - periods of lower demand, which can be drained during "peaks" - periods of higher demand. UPS batteries are normally used as a fail-over mechanism to transition to captive power sources upon utility failure. Furthermore, frequent discharges can cause UPS batteries to fail prematurely. We conduct detailed analysis of battery operation to figure out feasible operating regions given such battery lifetime and datacenter availability concerns. Using insights learned from this analysis, we develop peak reduction algorithms that combine the UPS battery knob with existing throttling based techniques for minimizing datacenter power costs. Using an experimental platform, we offer insights about Op-ex savings offered by eBuff for a wide range of workload peaks/valleys, UPS provisioning, and application SLA constraints. We find that eBuff can be used to realize 15-45% peak power reduction, corresponding to 6-18% savings in Op-ex across this spectrum. eBuff can also play a role in reducing Cap-ex costs by allowing tighter overbooking of power infrastructure components and we quantify the extent of such Cap-ex savings. To our knowledge, this is the first paper to exploit stored energy - typically lying untapped in the datacenter - to address the peak power draw problem.
Sriram Govindan, Anand Sivasubramaniam, Bhuvan Urgaonkar
ISCA1
2010 Power Consumption Prediction and Power-Aware Packing in Consolidated Environments
abstract
Consolidation of workloads has emerged as a key mechanism to dampen the rapidly growing energy expenditure within enterprise-scale data centers. To gainfully utilize consolidation-based techniques, we must be able to characterize the power consumption of groups of colocated applications. Such characterization is crucial for effective prediction and enforcement of appropriate limits on power consumption-power budgets-within the data center. We identify two kinds of power budgets: 1) an average budget to capture an upper bound on long-term energy consumption within that level and 2) a sustained budget to capture any restrictions on sustained draw of current above a certain threshold. Using a simple measurement infrastructure, we derive power profiles-statistical descriptions of the power consumption of applications. Based on insights gained from detailed profiling of several applications-both individual and consolidated-we develop models for predicting average and sustained power consumption of consolidated applications. We conduct an experimental evaluation of our techniques on a Xen-based server that consolidates applications drawn from a diverse pool. For a variety of consolidation scenarios, we are able to predict average power consumption within five percent error margin and sustained power within 10 percent error margin. Using prediction techniques allows us to ensure safe yet efficient system operation-in a representative case, we are able to improve the number of applications consolidated on a server from two to three (compared to existing baseline techniques) by choosing the appropriate power state that satisfies the power budgets associated with the server.
Jeonghwan Choi, Sriram Govindan, Jinkyu Jeong, Bhuvan Urgaonkar, Anand Sivasubramaniam
IEEE Trans. Computers2
2009 Statistical profiling-based techniques for effective power provisioning in data centers
abstract
Current capacity planning practices based on heavy over-provisioning of power infrastructure hurt (i) the operational costs of data centers as well as (ii) the computational work they can support. We explore a combination of statistical multiplexing techniques to improve the utilization of the power hierarchy within a data center. At the highest level of the power hierarchy, we employ controlled underprovisioning and over-booking of power needs of hosted workloads. At the lower levels, we introduce the novel notion of soft fuses to flexibly distribute provisioned power among hosted workloads based on their needs. Our techniques are built upon a measurement-driven profiling and prediction framework to characterize key statistical properties of the power needs of hosted workloads and their aggregates. We characterize the gains in terms of the amount of computational work (CPU cycles) per provisioned unit of power Computation per Provisioned Watt (CPW). Our technique is able to double the CPWoffered by a Power Distribution Unit (PDU) running the e-commerce benchmark TPC-W compared to conventional provisioning practices. Over-booking the PDU by 10% based on tails of power profiles yields a further improvement of 20%. Reactive techniques implemented on our Xen VMM-based servers dynamically modulate CPU DVFS states to ensure power draw below the limits imposed by soft fuses. Finally, information captured in our profiles also provide ways of controlling application performance degradation despite overbooking. The 95th percentile of TPC-W session response time only grew from 1.59 sec to 1.78 sec--a degradation of 12%.
Sriram Govindan, Jeonghwan Choi, Bhuvan Urgaonkar, Anand Sivasubramaniam, Andrea Baldini
EuroSys1
2009 vPath: Precise Discovery of Request Processing Paths from Black-Box Observations of Thread and Network Activities
Byung-Chul Tak, Chunqiang Tang, Sriram Govindan, Bhuvan Urgaonkar, Rong Chang 0001
USENIX ATC4
2009 Xen and Co.: Communication-Aware CPU Management in Consolidated Xen-Based Hosting Platforms
abstract
Recent advances in software and architectural support for server virtualization have created interest in using this technology in the design of consolidated hosting platforms. Since virtualization enables easier and faster application migration as well as secure colocation of antagonistic applications, higher degrees of server consolidation are likely to result in such virtualization-based hosting platforms (VHPs). We identify two shortcomings in existing virtual machine monitors (VMMs) that prove to be obstacles in operating hosting platforms, such as Internet data centers, under conditions of such high consolidation: 1) CPU schedulers that are agnostic to the communication behavior of modern, multitier applications and 2) inadequate or inaccurate mechanisms for accounting the CPU overheads of I/O virtualization. We develop a new communication-aware CPU scheduling algorithm and a CPU usage accounting mechanism. We implement our algorithms in the Xen VMM and build a prototype VHP on a cluster of 36 servers. Our experimental evaluation with realistic Internet server applications and benchmarks demonstrates the performance/cost benefits and the wide applicability of our algorithms. For example, the TPC-W benchmark exhibited improvements in average response times between 20 percent and 35 percent for a variety of consolidation scenarios. A streaming media server hosted on our prototype VHP was able to satisfactorily service up to 3.5 times as many clients as one running on the default Xen.
Sriram Govindan, Jeonghwan Choi, Arjun R. Nath, Amitayu Das, Bhuvan Urgaonkar, Anand Sivasubramaniam
IEEE Trans. Computers1
2008 Profiling, Prediction, and Capping of Power Consumption in Consolidated Environments
Jeonghwan Choi, Sriram Govindan, Bhuvan Urgaonkar, Anand Sivasubramaniam
MASCOTS2
2007 Xen and co.: communication-aware CPU scheduling for consolidated xen-based hosting platforms
abstract
Recent advances in software and architectural support for server virtualization have created interest in using this technology in the design of consolidated hosting platforms. Since virtualization enables easier and faster application migration as well as secure co-location of antagonistic applications, higher degrees of server consolidation are likely to result in such virtualization-based hosting platforms (VHPs). We identify a key shortcoming in existing virtual machine monitors (VMMs) that proves to be an obstacle in operating hosting platforms, such as Internet data centers, under conditions of such high consolidation: CPU schedulers that are agnostic to the communication behavior of modern, multi-tier applications. We develop a new communication-aware CPU scheduling algorithm to alleviate this problem. We implement our algorithm in the Xen VMM and build a prototype VHP on a cluster of servers. Our experimental evaluation with realistic Internet server applications and benchmarks demonstrates the performance/cost benefits and the wide applicability of our algorithms. For example, the TPC-W benchmark exhibited improvements in average response times of up to 35% for a variety of consolidation scenarios. A streaming media server hosted on our prototype VHP was able to satisfactorily service up to 3.5 times as many clients as one running on the default Xen.
Sriram Govindan, Arjun R. Nath, Amitayu Das, Bhuvan Urgaonkar, Anand Sivasubramaniam
VEE1