EDBT 2026 Demo / reviewers in the wild / expert
Badriddine M. Khessib
dblp:85/9055 · also Badriddine Khessib
· DBLP profile ↗
6ranked-venue papers
0as first author
0since 2021 · last 2017
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5Software engineering, systems software and programming languages · 2Security and privacy · 1Databases, data management, data science and information retrieval · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Storage systems · 55% Cloud and datacenter computing · 14% Energy-efficient computing · 14% |
Topics — the 5 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Storage systems › storage reliability
failure characterization |
0.2 | 1 | 2016 | SSD Failures in Datacenters: What, When and Why? · SIGMETRICS 2016 |
Storage systems
flash and SSD |
0.2 | 1 | 2016 | SSD Failures in Datacenters: What, When and Why? · SIGMETRICS 2016 |
Storage systems › flash and SSD
SSD reliability |
0.2 | 1 | 2016 | SSD Failures in Datacenters: What, When and Why? · SIGMETRICS 2016 |
Cloud and datacenter computing
datacenter infrastructure |
0.2 | 1 | 2014 | Underprovisioning backup power infrastructure for datacenters · ASPLOS 2014 |
Energy-efficient computing
power management |
0.2 | 1 | 2014 | Underprovisioning backup power infrastructure for datacenters · ASPLOS 2014 |
Methods — techniques the papers use, named apart from their topics
statistical characterization · 0.2field data analysis · 0.2cost modeling · 0.2availability analysis · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2017 | Rain or Shine? - Making Sense of Cloudy Reliability DataabstractCloud datacenters must ensure high availability for the hosted applications and failures can be the bane of datacenter operators. Understanding the what, when and why of failures can help tremendously to mitigate their occurrence and impact. Failures can, however, depend on numerous spatial and temporal factors spanning hardware, workloads, support facilities, and even the environment. One has to rely on failure data from the field to quantify the influence of these factors on failures. Towards this goal, we collect failures data along with many parameters that might influence failures from two large production datacenters with very diverse characteristics. We show that multiple factors simultaneously affect failures, and these factors may interact in non-trivial ways. This makes conventional approaches that study aggregate characteristics or single parameter influences, rather inaccurate. Instead, we build a multi-factor analysis framework to systematically identify influencing factors, quantify their relative impact, and help in more accurate decision making for failure mitigation. We demonstrate this approach for three important decisions: spare capacity provisioning, comparing the reliability of hardware for vendor selection, and quantifying flexibility in datacenter climate control for cost-reliability trade-offs. Iyswarya Narayanan, Bikash Sharma, Di Wang 0003, Sriram Govindan, Laura Caulfield, Anand Sivasubramaniam, Aman Kansal, Jie Liu 0001, Badriddine M. Khessib, Kushagra Vaid |
ICDCS | 9 |
| 2016 | SSD Failures in Datacenters: What, When and Why?abstractDespite the growing popularity of Solid State Disks (SSDs) in the datacenter, little is known about their reliability characteristics in the field. The little knowledge is mainly vendor supplied, which cannot really help understand how SSD failures can manifest and impact production systems, in order to take appropriate actions. Besides failure data, a detailed characterization requires wide spectrum of data about factors influencing SSD failures, right from provisioning (what models' where and when deployed' etc.) to the operational ones (workloads, read-write intensities, write amplification, etc.). We analyze over half a million SSDs that span multiple generations spread across several datacenters which host a wide range of workloads over nearly 3 years. By studying the diverse set of factors on SSD failures, and their symptoms, our work provides the first look at the what, when and why characteristics of SSD failures in production datacenters. Iyswarya Narayanan, Di Wang 0003, Myeongjae Jeon, Bikash Sharma, Laura Caulfield, Anand Sivasubramaniam, Ben Cutler, Jie Liu 0001, Badriddine M. Khessib, Kushagra Vaid |
SIGMETRICS | 9 |
| 2016 | SSD Failures in Datacenters: What? When? and Why?abstractDespite the growing popularity of Solid State Disks (SSDs) in the datacenter, little is known about their reliability characteristics in the field. The little knowledge is mainly vendor supplied, and such information cannot really help understand how SSD failures can manifest and impact the operation of production systems, in order to take appropriate remedial measures. Besides actual failure data and the symptoms exhibited by SSDs before failing, a detailed characterization effort requires wide set of data about factors influencing SSD failures, right from provisioning factors to the operational ones. This paper presents an extensive SSD failure characterization by analyzing a wide spectrum of data from over half a million SSDs that span multiple generations spread across several datacenters which host a wide spectrum of workloads over nearly 3 years. By studying the diverse set of design, provisioning and operational factors on failures, and their symptoms, our work provides the first comprehensive analysis of the what, when and why characteristics of SSD failures in production datacenters. Iyswarya Narayanan, Di Wang 0003, Myeongjae Jeon, Bikash Sharma, Laura Caulfield, Anand Sivasubramaniam, Ben Cutler, Jie Liu 0001, Badriddine M. Khessib, Kushagra Vaid |
SYSTOR | 9 |
| 2014 | Underprovisioning backup power infrastructure for datacentersabstractWhile there has been prior work to underprovision the power distribution infrastructure for a datacenter to save costs, the ability to underprovision the backup power infrastructure, which contributes significantly to capital costs, is little explored. There are two main components in the backup infrastructure - Diesel Generators (DGs) and UPS units - which can both be underprovisioned (or even removed) in terms of their power and/or energy capacities. However, embarking on such underprovisioning mandates studying several ramifications - the resulting cost savings, the lower availability, and the performance and state loss consequences on individual applications - concurrently. This paper presents the first such study, considering cost, availability, performance and application consequences of underprovisioning the backup power infrastructure. We present a framework to quantify the cost of backup capacity that is provisioned, and implement techniques leveraging existing software and hardware mechanisms to provide as seamless an operation as possible for an application within the provisioned backup capacity during a power outage. We evaluate the cost-performance-availability trade-offs for different levels of backup underprovisioning for applications with diverse reliance on the backup infrastructure. Our results show that one may be able to completely do away with DGs, compensating for it with additional UPS energy capacities, to significantly cut costs and still be able to handle power outages lasting as high as 40 minutes (which constitute bulk of the outages). Further, we can push the limits of outage duration that can be handled in a cost-effective manner, if applications are willing to tolerate degraded performance during the outage. Our evaluations also show that different applications react differently to the outage handling mechanisms, and that the efficacy of the mechanisms is sensitive to the outage duration. The insights from this paper can spur new opportunities for future work on backup power infrastructure optimization. Di Wang 0003, Sriram Govindan, Anand Sivasubramaniam, Aman Kansal, Jie Liu 0001, Badriddine M. Khessib |
ASPLOS | 6 |
| 2014 | Characterizing Application Memory Error Vulnerability to Optimize Datacenter Cost via Heterogeneous-Reliability MemoryabstractMemory devices represent a key component of datacenter total cost of ownership (TCO), and techniques used to reduce errors that occur on these devices increase this cost. Existing approaches to providing reliability for memory devices pessimistically treat all data as equally vulnerable to memory errors. Our key insight is that there exists a diverse spectrum of tolerance to memory errors in new data-intensive applications, and that traditional one-size-fits-all memory reliability techniques are inefficient in terms of cost. For example, we found that while traditional error protection increases memory system cost by 12.5%, some applications can achieve 99.00% availability on a single server with a large number of memory errors without any error protection. This presents an opportunity to greatly reduce server hardware cost by provisioning the right amount of memory reliability for different applications. Toward this end, in this paper, we make three main contributions to enable highly-reliable servers at low datacenter cost. First, we develop a new methodology to quantify the tolerance of applications to memory errors. Second, using our methodology, we perform a case study of three new dataintensive workloads (an interactive web search application, an in-memory key -- value store, and a graph mining framework) to identify new insights into the nature of application memory error vulnerability. Third, based on our insights, we propose several new hardware/software heterogeneous-reliability memory system designs to lower datacenter cost while achieving high reliability and discuss their trade-off. We show that our new techniques can reduce server hardware cost by 4.7% while achieving 99.90% single server availability. Sriram Govindan, Bikash Sharma, Mark Santaniello, Justin Meza, Aman Kansal, Jie Liu 0001, Badriddine M. Khessib, Kushagra Vaid, Onur Mutlu |
DSN | 8 |
| 2006 | Large scale Itanium® 2 processor OLTP workload characterization and optimizationabstractLarge scale OLTP workloads on modern database servers are well understood across the industry. Their runtime performance characterizations serve to drive both server side software features and processor specific design decisions but are not understood outside of the primary industry stakeholders. We provide a rare glimpse into the performance characterizations of processor and platform targeted software optimizations running on a large-scale 32 processor, Intel® Itanium® 2 based, ccNUMA platform. Gerrit Saylor, Badriddine M. Khessib |
DaMoN | 2 |