Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Kushagra Vaid

dblp:57/7456 · DBLP profile ↗
← Back
14ranked-venue papers
0as first author
0since 2021 · last 2018
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10Software engineering, systems software and programming languages · 6Security and privacy · 2Computer networks · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
6 papers
Cloud and datacenter computing · 43% Storage systems · 31% Energy-efficient computing · 14%
Databases, data mining, and information retrieval
2 papers
Information retrieval · 100%

Topics — the 17 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Cloud and datacenter computing › virtualization
network virtualization
0.312018
Azure Accelerated Networking: SmartNICs in the Public Cloud · NSDI 2018
Cloud and datacenter computing › computation offloading › network function offloading
SmartNIC offload
0.312018
Azure Accelerated Networking: SmartNICs in the Public Cloud · NSDI 2018
Cloud and datacenter computing
virtualization
0.312018
Azure Accelerated Networking: SmartNICs in the Public Cloud · NSDI 2018
Storage systems › storage reliability
failure characterization
0.212016
SSD Failures in Datacenters: What, When and Why? · SIGMETRICS 2016
Storage systems
flash and SSD
0.212016
SSD Failures in Datacenters: What, When and Why? · SIGMETRICS 2016
Storage systems › flash and SSD
SSD reliability
0.212016
SSD Failures in Datacenters: What, When and Why? · SIGMETRICS 2016
Cloud and datacenter computing
datacenter operations
0.212013
Datacenter Scale Evaluation of the Impact of Temperature on Hard Disk Drive Failures · ACM Trans. Storage 2013
Energy-efficient computing
datacenter power management
0.212013
ACE: abstracting, characterizing and exploiting peaks and valleys in datacenter power consumption · SIGMETRICS 2013
Energy-efficient computing › power management
power capping
0.212013
ACE: abstracting, characterizing and exploiting peaks and valleys in datacenter power consumption · SIGMETRICS 2013
Storage systems
storage reliability
0.212013
Datacenter Scale Evaluation of the Impact of Temperature on Hard Disk Drive Failures · ACM Trans. Storage 2013
Performance modeling and evaluation
workload characterization
0.212013
ACE: abstracting, characterizing and exploiting peaks and valleys in datacenter power consumption · SIGMETRICS 2013
Embedded and real-time systems › embedded processor
mobile processor
0.112011
Mobile processors for energy-efficient web search · ACM Trans. Comput. Syst. 2011
Hardware accelerators and domain-specific architectures
network accelerator
0.112018
Azure Accelerated Networking: SmartNICs in the Public Cloud · NSDI 2018
Storage systems › storage reliability
disk failure
0.012013
Datacenter Scale Evaluation of the Impact of Temperature on Hard Disk Drive Failures · ACM Trans. Storage 2013
Storage systems › magnetic storage
hard disk drive
0.012013
Datacenter Scale Evaluation of the Impact of Temperature on Hard Disk Drive Failures · ACM Trans. Storage 2013
Information retrieval
web search
0.012011
Mobile processors for energy-efficient web search · ACM Trans. Comput. Syst. 2011
Information retrieval › search engines
web search engine
0.012010
Web search using mobile cores: quantifying and mitigating the price of efficiency · ISCA 2010

Methods — techniques the papers use, named apart from their topics

workload characterization · 0.2statistical characterization · 0.2performance-per-joule analysis · 0.2field data analysis · 0.2microarchitecture-level analysis · 0.2measurement · 0.2statistical correlation analysis · 0.2statistical analysis · 0.2
YearPublicationVenuePosition
2018 Azure Accelerated Networking: SmartNICs in the Public Cloud
Daniel Firestone, Andrew Putnam, Sambrama Mundkur, Derek Chiou, Alireza Dabagh, Mike Andrewartha, Hari Angepat, Vivek Bhanu, Adrian M. Caulfield, Eric S. Chung, Harish Kumar Chandrappa, Somesh Chaturmohta, Matt Humphrey, Jack Lavier, Norman Lam, Fengfen Liu, Kalin Ovtcharov, Jitendra Padhye, Gautham Popuri, Shachar Raindel, Tejas Sapre, Mark Shaw 0001, Gabriel Silva, Madhan Sivakumar, Nisheeth Srivastava, Anshuman Verma, Qasim Zuhair, Deepak Bansal, Doug Burger, Kushagra Vaid, David A. Maltz, Albert G. Greenberg
NSDI30
2017 Rain or Shine? - Making Sense of Cloudy Reliability Data
abstract
Cloud datacenters must ensure high availability for the hosted applications and failures can be the bane of datacenter operators. Understanding the what, when and why of failures can help tremendously to mitigate their occurrence and impact. Failures can, however, depend on numerous spatial and temporal factors spanning hardware, workloads, support facilities, and even the environment. One has to rely on failure data from the field to quantify the influence of these factors on failures. Towards this goal, we collect failures data along with many parameters that might influence failures from two large production datacenters with very diverse characteristics. We show that multiple factors simultaneously affect failures, and these factors may interact in non-trivial ways. This makes conventional approaches that study aggregate characteristics or single parameter influences, rather inaccurate. Instead, we build a multi-factor analysis framework to systematically identify influencing factors, quantify their relative impact, and help in more accurate decision making for failure mitigation. We demonstrate this approach for three important decisions: spare capacity provisioning, comparing the reliability of hardware for vendor selection, and quantifying flexibility in datacenter climate control for cost-reliability trade-offs.
Iyswarya Narayanan, Bikash Sharma, Di Wang 0003, Sriram Govindan, Laura Caulfield, Anand Sivasubramaniam, Aman Kansal, Jie Liu 0001, Badriddine M. Khessib, Kushagra Vaid
ICDCS10
2016 SSD Failures in Datacenters: What, When and Why?
abstract
Despite the growing popularity of Solid State Disks (SSDs) in the datacenter, little is known about their reliability characteristics in the field. The little knowledge is mainly vendor supplied, which cannot really help understand how SSD failures can manifest and impact production systems, in order to take appropriate actions. Besides failure data, a detailed characterization requires wide spectrum of data about factors influencing SSD failures, right from provisioning (what models' where and when deployed' etc.) to the operational ones (workloads, read-write intensities, write amplification, etc.). We analyze over half a million SSDs that span multiple generations spread across several datacenters which host a wide range of workloads over nearly 3 years. By studying the diverse set of factors on SSD failures, and their symptoms, our work provides the first look at the what, when and why characteristics of SSD failures in production datacenters.
Iyswarya Narayanan, Di Wang 0003, Myeongjae Jeon, Bikash Sharma, Laura Caulfield, Anand Sivasubramaniam, Ben Cutler, Jie Liu 0001, Badriddine M. Khessib, Kushagra Vaid
SIGMETRICS10
2016 SSD Failures in Datacenters: What? When? and Why?
abstract
Despite the growing popularity of Solid State Disks (SSDs) in the datacenter, little is known about their reliability characteristics in the field. The little knowledge is mainly vendor supplied, and such information cannot really help understand how SSD failures can manifest and impact the operation of production systems, in order to take appropriate remedial measures. Besides actual failure data and the symptoms exhibited by SSDs before failing, a detailed characterization effort requires wide set of data about factors influencing SSD failures, right from provisioning factors to the operational ones. This paper presents an extensive SSD failure characterization by analyzing a wide spectrum of data from over half a million SSDs that span multiple generations spread across several datacenters which host a wide spectrum of workloads over nearly 3 years. By studying the diverse set of design, provisioning and operational factors on failures, and their symptoms, our work provides the first comprehensive analysis of the what, when and why characteristics of SSD failures in production datacenters.
Iyswarya Narayanan, Di Wang 0003, Myeongjae Jeon, Bikash Sharma, Laura Caulfield, Anand Sivasubramaniam, Ben Cutler, Jie Liu 0001, Badriddine M. Khessib, Kushagra Vaid
SYSTOR10
2014 Characterizing Application Memory Error Vulnerability to Optimize Datacenter Cost via Heterogeneous-Reliability Memory
abstract
Memory devices represent a key component of datacenter total cost of ownership (TCO), and techniques used to reduce errors that occur on these devices increase this cost. Existing approaches to providing reliability for memory devices pessimistically treat all data as equally vulnerable to memory errors. Our key insight is that there exists a diverse spectrum of tolerance to memory errors in new data-intensive applications, and that traditional one-size-fits-all memory reliability techniques are inefficient in terms of cost. For example, we found that while traditional error protection increases memory system cost by 12.5%, some applications can achieve 99.00% availability on a single server with a large number of memory errors without any error protection. This presents an opportunity to greatly reduce server hardware cost by provisioning the right amount of memory reliability for different applications. Toward this end, in this paper, we make three main contributions to enable highly-reliable servers at low datacenter cost. First, we develop a new methodology to quantify the tolerance of applications to memory errors. Second, using our methodology, we perform a case study of three new dataintensive workloads (an interactive web search application, an in-memory key -- value store, and a graph mining framework) to identify new insights into the nature of application memory error vulnerability. Third, based on our insights, we propose several new hardware/software heterogeneous-reliability memory system designs to lower datacenter cost while achieving high reliability and discuss their trade-off. We show that our new techniques can reduce server hardware cost by 4.7% while achieving 99.90% single server availability.
Sriram Govindan, Bikash Sharma, Mark Santaniello, Justin Meza, Aman Kansal, Jie Liu 0001, Badriddine M. Khessib, Kushagra Vaid, Onur Mutlu
DSN9
2013 ACE: abstracting, characterizing and exploiting peaks and valleys in datacenter power consumption
abstract
Peak power management of datacenters has tremendous cost implications. While numerous mechanisms have been proposed to cap power consumption, real datacenter power consumption data is scarce. To address this gap, we collect power demands at multiple spatial and fine-grained temporal resolutions from the load of geo-distributed datacenters of Microsoft over 6 months. We conduct aggregate analysis of this data, to study its statistical properties. With workload characterization a key ingredient for systems design and evaluation, we note the importance of better abstractions for capturing power demands, in the form of peaks and valleys. We identify and characterize attributes for peaks and valleys, and important correlations across these attributes that can influence the choice and effectiveness of different power capping techniques. With the wide scope of exploitability of such characteristics for power provisioning and optimizations, we illustrate its benefits with two specific case studies.
Di Wang 0003, Chuangang Ren, Sriram Govindan, Anand Sivasubramaniam, Bhuvan Urgaonkar, Aman Kansal, Kushagra Vaid
SIGMETRICS7
2013 Innovative practices session 5C: Cloud atlas - Unreliability through massive connectivity
abstract
The rapid pace of integration, emergence of low power, low cost computing elements, and ubiquitous and ever-increasing bandwidth of connectivity have given rise to data center and cloud infrastructures. These infrastructures are beginning to be used on a massive scale across vast geographic boundaries to provide commercial services to businesses such as banking, enterprise computing, online sales, and data mining and processing for targeted marketing to name a few. Such an infrastructure comprises of thousands of compute and storage nodes that are interconnected by massive network fabrics, each of them having their own hardware and firmware stacks, with layers of software stacks for operating systems, network protocols, schedulers and application programs. The scale of such an infrastructure has made possible service that has been unimaginable only a few years ago, but has the downside of severe losses in case of failure. A system of such scale and risk necessitates methods to (a) proactively anticipate and protect against impending failures, (b) efficiently, transparently and quickly detect, diagnose and correct failures in any software or hardware layer, and (c) be able to automatically adapt itself based on prior failures to prevent future occurrences. Addressing the above reliability challenges is inherently different from the traditional reliability techniques. First, there is a great amount of redundant resources available in the cloud from networking to computing and storage nodes, which opens up many reliability approaches by harvesting these available redundancies. Second, due to the large scale of the system, techniques with high overheads, especially in power, are not acceptable. Consequently, cross layer approaches to optimize the availability and power have gained traction recently. This session will address these challenges in maintaining reliable service with solutions across the hardware/software stacks. The currently available commercial data-center and cloud infrastructures will be reviewed and the relative occurrences of different causalities of failures, the level to which they are anticipated and diagnosed in practice, and their impact on the quality of service and infrastructure design will be discussed. A study on real-time analytics to proactively address failures in a private, secure cloud engaged in domain-specific computations, with streaming inputs received from embedded computing platforms (such as airborne image sources, data streams, or sensors) will be presented next. The session concludes with a discussion on the increased relevance of resiliency features built inside individual systems and components (private cloud) and how the macro public cloud absorbs innovations from this realm.
Helia Naeimi, Suriyaprakash Natarajan, Kushagra Vaid, Prabhakar Kudva, Mahesh Natu
VTS3
2013 Datacenter Scale Evaluation of the Impact of Temperature on Hard Disk Drive Failures
abstract
With the advent of cloud computing and online services, large enterprises rely heavily on their datacenters to serve end users. A large datacenter facility incurs increased maintenance costs in addition to service unavailability when there are increased failures. Among different server components, hard disk drives are known to contribute significantly to server failures; however, there is very little understanding of the major determinants of disk failures in datacenters. In this work, we focus on the interrelationship between temperature, workload, and hard disk drive failures in a large scale datacenter. We present a dense storage case study from a population housing thousands of servers and tens of thousands of disk drives, hosting a large-scale online service at Microsoft. We specifically establish correlation between temperatures and failures observed at different location granularities: (a) inside drive locations in a server chassis, (b) across server locations in a rack, and (c) across multiple racks in a datacenter. We show that temperature exhibits a stronger correlation to failures than the correlation of disk utilization with drive failures. We establish that variations in temperature are not significant in datacenters and have little impact on failures. We also explore workload impacts on temperature and disk failures and show that the impact of workload is not significant. We then experimentally evaluate knobs that control disk drive temperature, including workload and chassis design knobs. We corroborate our findings from the real data study and show that workload knobs show minimal impact on temperature. Chassis knobs like disk placement and fan speeds have a larger impact on temperature. Finally, we also show the proposed cost benefit of temperature optimizations that increase hard disk drive reliability.
Sriram Sankar, Mark Shaw 0001, Kushagra Vaid, Sudhanva Gurumurthi
ACM Trans. Storage3
2011 Impact of temperature on hard disk drive reliability in large datacenters
abstract
When datacenters are pushed to their limits of operational efficiency, reducing failure rates becomes critical for maintaining high levels of healthy server operation. In this experience report, we present a dense storage case study from a large population of servers housing tens of thousands of disk drives. Previous studies have presented divergent results concerning correlation between temperature and hard disk drive failures. In our paper, we specifically establish correlation between temperatures and failures observed at different location granularities: a) inside drive locations in a server chassis, b) across server locations in a rack and c) across multiple racks in a datacenter. We also establish that temperature exhibits a stronger correlation to failures compared to the correlation of disk utilization with drive failures. Thus, we show that temperature-aware server and datacenter design plays a pivotal role in datacenter reliability. Following our case study, we present a reliability model for estimating hard disk drive failures correlated with the datacenter operating temperature. We use a physical Arrhenius model with empirically derived coefficients for our model. We show an application of the model for selecting the datacenter inlet temperature setpoint for two different server storage configurations. Finally, with the help of a datacenter cost discussion, we highlight the need to incorporate reliability-aware datacenter design for increased efficiency in large scale datacenters.
Sriram Sankar, Mark Shaw 0001, Kushagra Vaid
DSN3
2011 Storage I/O generation and replay for datacenter applications
abstract
With the advent of social networking and cloud data-stores, user data is increasingly being stored in large capacity and high performance storage systems, which account for a significant portion of the total cost of ownership of a datacenter (DC) [3]. One of the main challenges when trying to evaluate storage system options is the difficulty in replaying the entire application in all possible system configurations. Furthermore, code and datasets of DC applications are rarely available to storage system designers. This makes the development of a representative model that captures key aspects of the workload's storage profile, even more appealing. Once such a model is available, the next step is to create a tool that convincingly reproduces the application's storage behavior via a synthetic I/O access pattern.
Christina Delimitrou, Sriram Sankar, Kushagra Vaid, Christoforos E. Kozyrakis
ISPASS3
2011 Optimizing benchmark configurations for energy efficiency
abstract
Historically compute server performance has been the most important pillar in the evaluation of datacenter efficiency, which can be measured using a variety of industry standard benchmarks. With the introduction of industry standard servers, price-performance became the second pillar in the 'efficiencyequation'. Today with an increased awareness in the industry for power optimized designs and corporate initiatives to reduce carbon emissions, data center efficiency needs to incorporate yet another key element in this equation: energy efficiency. Initial models based on 'name-plate' power consumption have been used to estimate energy efficiency while recently industry standard consortia like SPEC, TPC and SPC have started amalgating new energy metrics with their traditional performance metrics. TPC-Energy, enables the measuring and reporting of energy efficiency for transaction processing systems and decision support systems. In this paper we analyze TPC C benchmark configurations that may achieve leadership results in TPC-Energy using existing, more energy efficient technologies, such as solid states drives for storage subsystems, low power processors and high density DRAM in back end server and middle tier systems. Even though the study is based on TPC-C configurations these configuration optimizations are applicable to other benchmarks and production systems alike. We envision that the energy efficiency metrics and related optimizations to claim benchmark leadership will accelerate development and qualifications of energy efficient component and solutions.
Meikel Pöss, Raghunath Othayoth Nambiar, Kushagra Vaid
ICPE3
2011 Energy-delay based provisioning for large datacenters: an energy-efficient and cost optimal approach
abstract
It is challenging to determine the optimum number of servers required to provision for large online applications because of the conflicting mandates of (a) achieving peak performance needs and (b) minimizing unused datacenter power and capacity. Since Online Services application loads are unpredictable, datacenter operators often conservatively provision for maximum power utilization by characterizing workloads for peak load performance. In contrast, we aim to optimize the service capacity per total cost of ownership (TCO) of an Online Service datacenter deployment by characterizing the energy-delay properties of large scale datacenter workloads. We show that the peak load performance is not the energy efficient point of operation for most applications. We choose two industry-strength workloads (Internet Search and D-Process) and analyze their energy-delay behavior under varying loads. We then calculate the optimal operating point for the specific large-scale application and provision datacenter energy and capacity based on the energy-delay curves. In contrast to workload-based peak power provisioning, we show a 7% benefit in Service Capacity-per-TCO-dollar for the energy-delay characterization methodology in our cost analysis for Online Services Applications.
Sriram Sankar, Kushagra Vaid, Harry Rogers
ICPE2
2011 Mobile processors for energy-efficient web search
abstract
As cloud and utility computing spreads, computer architects must ensure continued capability growth for the data centers that comprise the cloud. Given megawatt scale power budgets, increasing data center capability requires increasing computing hardware energy efficiency. To increase the data center's capability for work, the work done per Joule must increase. We pursue this efficiency even as the nature of data center applications evolves. Unlike traditional enterprise workloads, which are typically memory or I/O bound, big data computation and analytics exhibit greater compute intensity. This article examines the efficiency of mobile processors as a means for data center capability. In particular, we compare and contrast the performance and efficiency of the Microsoft Bing search engine executing on the mobile-class Atom processor and the server-class Xeon processor. Bing implements statistical machine learning to dynamically rank pages, producing sophisticated search results but also increasing computational intensity. While mobile processors are energy-efficient, they exact a price for that efficiency. The Atom is 5× more energy-efficient than the Xeon when comparing queries per Joule. However, search queries on Atom encounter higher latencies, different page results, and diminished robustness for complex queries. Despite these challenges, quality-of-service is maintained for most, common queries. Moreover, as different computational phases of the search engine encounter different bottlenecks, we describe implications for future architectural enhancements, application tuning, and system architectures. After optimizing the Atom server platform, a large share of power and cost go toward processor capability. With optimized Atoms, more servers can fit in a given data center power budget. For a data center with 15MW critical load, Atom-based servers increase capability by 3.2× for Bing.
Vijay Janapa Reddi, Benjamin C. Lee, Trishul M. Chilimbi, Kushagra Vaid
ACM Trans. Comput. Syst.4
2010 Web search using mobile cores: quantifying and mitigating the price of efficiency
abstract
The commoditization of hardware, data center economies of scale, and Internet-scale workload growth all demand greater power efficiency to sustain scalability. Traditional enterprise workloads, which are typically memory and I/O bound, have been well served by chip multiprocessors com- prising of small, power-efficient cores. Recent advances in mobile computing have led to modern small cores capable of delivering even better power efficiency. While these cores can deliver performance-per-Watt efficiency for data center workloads, small cores impact application quality-of-service robustness, and flexibility, as these workloads increasingly invoke computationally intensive kernels. These challenges constitute the price of efficiency. We quantify efficiency for an industry-strength online web search engine in production at both the microarchitecture- and system-level, evaluating search on server and mobile-class architectures using Xeon and Atom processors.
Vijay Janapa Reddi, Benjamin C. Lee, Trishul M. Chilimbi, Kushagra Vaid
ISCA4