EDBT 2026 Demo / reviewers in the wild / expert
Sriram Sankar
dblp:91/2590
· DBLP profile ↗
29ranked-venue papers
6as first author
11since 2021 · last 2026
0009-0008-4581-8371ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 18 · 5 first-author · 9 since 2021Software engineering, systems software and programming languages · 13 · 2 first-author · 7 since 2021Databases, data management, data science and information retrieval · 2Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Computer networks · 1Security and privacy · 1 · 1 first-authorTheory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SEVI: Silent Data Corruption of Vector Instructions in Hyper-Scale DatacentersabstractSilent Data Corruption (SDC) poses a reliability threat in modern datacenters. These insidious errors evade detections and propagate incorrect results throughout the system. Companies including Google, Meta, and Alibaba have reported SDC incidents affecting their production. In this paper, we present the first comprehensive instruction- and application-level analysis of vector instruction SDCs in hyper-scale datacenters using a two-stage approach. We perform over 78 trillion test rounds in more than 14 billion CPU seconds. Our observations reveal undocumented SDC patterns that provide insights into possible underlying causes and inspire new mitigation strategies. Based on these findings, we propose a low-overhead SDC detection mechanism leveraging in-application algorithm-based fault tolerance. Our method achieves 88% to 100% SDC machine detection rate with a time overhead of only 1.35% even for modestly sized inputs. Yixuan Mei, Shreya Varshini, Harish Dattatraya Dixit, Sriram Sankar, K. V. Rashmi |
ASPLOS (2) | 4 |
| 2026 | PinDrop: Breaking the Silence on SDCs in a Large-Scale FleetabstractSilent Data Corruptions (SDCs) pose a significant and often hidden threat to the reliability of large-scale computing infrastructure, as they can silently compromise data integrity without immediate detection. Detecting such behaviors at hyperscale is often challenging due to their intermittent nature and the vast diversity of hardware and workloads in a large-scale fleet. This work addresses these challenges by introducing PinDrop, a characterization methodology that leverages continuous, high-frequency testing infrastructure across millions of servers to gather information about SDCs at scale. By leveraging extensive test-suites tailored to mimic real-world applications and exercise a wide range of CPU features, we provide the most comprehensive characterization of SDC failures to date, analyzing over 500 million test executions across millions of devices. Our findings reveal that 0.035% of tested machines suffer from at least one SDC failure during their lifetime. Examining years of data (rather than just a testing snapshot in time), we observe SDCs emerging long after initial deployment and persisting over time. Detailed analysis shows that an average of 0.0024% of tested machines begin failing in each quarter they are tested beyond an initial burn-in period, confirming a fundamental need for continuous testing. Our findings also provide further insights into SDC behaviors at-scale, including failure breakdowns across architectures, test families, specific core IDs, and output-level behaviors. Peter W. Deutsch, Harish Dattatraya Dixit, Gautham Vunnam, Carl Moran, Eleanor Ozer, Sriram Sankar |
HPCA | 6 |
| 2026 | Server Life Extensions at Scale
Rhea Dutta, Harish Dattatraya Dixit, Sriram Sankar |
IOLTS | 5 |
| 2025 | Hardware Sentinel: Protecting Software Applications from Hardware Silent Data CorruptionsabstractSilent Data Corruptions (SDCs) pose a significant challenge in large-scale infrastructures, affecting data center applications unpredictably and reducing service reliability. Primarily caused by silicon defects, traditional hardware testing methods are insufficient to prevent SDC propagation. SDCs are influenced by various factors, including data randomization, workload characteristics, environmental conditions, and aging, necessitating top-down approaches from the application layer. In this paper, we introduce Hardware Sentinel, a novel framework that detects SDCs through typical software failure indicators such as segmentation faults, core dumps, application crashes, and logs. We have validated our framework in a large-scale data center fleet, across diverse application, kernel, and hardware configurations, achieving a high success rate of SDC detection. Hardware Sentinel has uncovered novel instances of SDCs, surpassing the detection capabilities of published testing techniques. Our analysis of over 6 years' worth of application and system failure data within a large-scale infrastructure has successfully identified hundreds of defective CPUs that triggered SDCs. Notably, the Hardware Sentinel flow increases effective coverage over existing hardware-testing methods like Fleetscanner (out-of-production testing) by 1.74x and Ripple (in-production testing) by 1.92x. We share the top kernel exceptions with the highest correlation to silent data corruption failures. We present results spanning 7 CPU generations from multiple semiconductor manufacturers, 13 large-scale workloads, and 27 data center regions, providing insights into the trade-offs involved in detection and fleet deployment. Rhea Dutta, Harish Dattatraya Dixit, Rik van Riel, Gautham Vunnam, Sriram Sankar |
ASPLOS (2) | 5 |
| 2025 | From Gates to SDCs: Understanding Fault Propagation Through the Compute StackabstractSilent Data Corruption (SDC) is the most severe effect of a silicon defect in a CPU or other computing chip. The arithmetic units of a CPU are, usually, unprotected and are, thus, the ones that most likely produce SDCs (as well as visible malfunctions of programs such as crashes). In this work, we shed light on the traversal of silicon defects from their point of origin deep inside arithmetic units of complex CPUs towards the program result. We employ microarchitecture-level fault injection enhanced with gate-level designs of the arithmetic units of interest. The hybrid setup combines (i) the accuracy of the hardware and fault modeling and (ii) the speed of program simulation to run long programs to end (thus observing SDC incidents); the analysis that this combination delivers is impossible at other abstraction layers which are either hardware-agnostic (software level) or extremely slow (gate-level). We quantify the effects of faults in two stages and with multiple metrics: (a) how faults propagate to the outputs of the arithmetic units when individual instructions are executed, and (b) how faults eventually affect the outcome of the program generating SDCs, crashes, or being masked. Our fine-grain findings can be utilized for informed fault detection and tolerance strategies at the hardware or the software levels. Odysseas Chatzopoulos, George Papadimitriou 0001, Dimitris Gizopoulos, Harish Dattatraya Dixit, Sriram Sankar |
DATE | 5 |
| 2025 | Veritas - Demystifying Silent Data Corruptions: μArch-Level Modeling and Fleet Data of Modern x86 CPUsabstractHyperscalers have reported unexpectedly high numbers of defective CPU chips, with a defect rate of 1 in a 1000, leading to Silent Data Corruptions (SDCs) in their computing fleets. However, there is no public data on the rate of SDC incidents (corrupted program executions) in large fleets, nor nor any detailed information on which CPU units, microarchitectures, or workloads are more likely to generate SDCs due to silicon defects. While CPU array structures have been studied for fault effects, arithmetic units like integer and floating-point units have not been thoroughly analyzed as potential root causes of SDCs. This paper addresses this critical gap by accurately modeling hardware faults in the arithmetic units of modern x86 CPUs and measuring the probability and rates of SDCs. Using a full-system gem5-based fault injector, the paper examines SDC trends across five recent $x 86$ microarchitectures, various arithmetic units, and instruction classes. By integrating real-world defect rates from large-scale datacenter experiments with early-stage modeling and simulation, the paper provides critical insights into SDC incident rates across different systems. This information is essential for guiding hardware-based or software-based fault protection methods and is the paper’s primary contribution to minimizing the impact of silent data corruptions in computing. Odysseas Chatzopoulos, Nikos Karystinos, George Papadimitriou 0001, Dimitris Gizopoulos, Harish Dattatraya Dixit, Sriram Sankar |
HPCA | 6 |
| 2025 | Understanding Recommendation System Robustness Against Silent Data Corruption: An Empirical StudyabstractModern deep learning-based recommendation system (DRS) is a dominant workload in industrial data centers. However, with continuous transistor scaling and increasing hardware complexity, silent data corruption (SDC) has become a notable threat to the reliability of data center workloads. Multiple industry hyperscalars have reported the difficulty in addressing SDC due to their “stealthy” nature and elusive manifestation. Given the critical role that DRS plays in maintaining the quality of online services, understanding and enhancing its robustness against SDCs are imperative. To the best our knowledge, this paper presents the first empirical study on understanding DRS robustness against SDCs. Specifically, we develop PyTEI, a PyTorch-based, user-friendly, and highly-efficient error injection framework, based on which we perform large-scale error injection experiments to the parameters of five representative DRS models under three datasets. Experimental results reveal that, the sparsity level of input data and feature affect the robustness of DRS, and in particular, MLP modules inside DRS are especially vulnerable to SDCs. Further, we evaluate the effectiveness of three representative error mitigation methods – algorithm based fault tolerance (ABFT), activation clipping, and selective bit protection (SBP), in enhancing DRS robustness. Experimental results reveal that activation clipping obtains the best result by recovering up to 30% of the degraded DRS performance under SDCs. This study provides valuable insights for industry practitioners in developing robust fault-tolerant strategies for DRS workloads. We open source PyTEI at https://github.com/facebookresearch/PyTEI. Dongning Ma, Fan Fred Lin, Sriram Sankar |
ISSRE | 5 |
| 2024 | Dr. DNA: Combating Silent Data Corruptions in Deep Learning using Distribution of Neuron ActivationsabstractDeep neural networks (DNNs) have been widely-adopted in various safety-critical applications such as computer vision and autonomous driving. However, as technology scales and applications diversify, coupled with the increasing heterogeneity of underlying hardware architectures, silent data corruption (SDC) has been emerging as a pronouncing threat to the reliability of DNNs. Recent reports from industry hyperscalars underscore the difficulty in addressing SDC due to their "stealthy" nature and elusive manifestation. In this paper, we propose Dr. DNA, a novel approach to enhance the reliability of DNN systems by detecting and mitigating SDCs. Specifically, we formulate and extract a set of unique SDC signatures from the Distribution of Neuron Activations (DNA), based on which we propose early-stage detection and mitigation of SDCs during DNN inference. We perform an extensive evaluation across 3 vision tasks, 5 different datasets, and 10 different models, under 4 different error models. Results show that Dr. DNA achieves 100% SDC detection rate for most cases, 95% detection rate on average and >90% detection rate across all cases, representing 20% - 70% improvement over baselines. Dr. DNA can also mitigate the impact of SDCs by effectively recovering DNN model performance with <1% memory overhead and <2.5% latency overhead. Dongning Ma, Fan Fred Lin, Alban Desmaison, Joel Coburn, Sriram Sankar, Xun Jiao 0001 |
ASPLOS (3) | 6 |
| 2024 | Silent Data Corruptions in Computing Systems: Early Predictions and Large-Scale MeasurementsabstractSilent Data Corruptions (SDCs) due to defects in computing chips (CPUs, GPUs, AI accelerators) is a critical threat to the quality of large-scale computing in different application domains: cloud computing, high-performance computing, edge computing. Recent public reports by cloud hyperscalers have emphasized that apart from the usual suspects for SDCs (memory, storage, network), the heart of the computations, the processing elements of all types generate an unexpectedly large rate of SDCs which can cause erroneous calculations and severe information loss. We report, in a consolidated form, recent efforts to correlate early microarchitecture-level simulation-based predictions about the likelihood, rates, severity, and root causes of SDCs and large-scale in-field studies in cloud data centers. Early microarchitecture-level prediction of SDC characteristics (susceptible units, workloads, instructions) can shed light to the cryptic problem of SDCs. The findings of a diligent pre-silicon analysis can assist better understanding of SDCs and can thus drive effective protection decisions either at the hardware or at the software levels at deployment stages. Dimitris Gizopoulos, George Papadimitriou 0001, Odysseas Chatzopoulos, Nikos Karystinos, Harish Dattatraya Dixit, Sriram Sankar |
ETS | 6 |
| 2023 | Silent Data Corruptions: The Stealthy Saboteurs of Digital IntegrityabstractSilent Data Corruptions (SDCs) pose a significant threat to the integrity of digital systems. These stealthy saboteurs silently corrupt data, remaining undetected by traditional error handling mechanisms. The silent nature of SDCs makes them challenging to trace at the hardware level, as they evade error reporting systems. Instead, their effects manifest at the application level, potentially causing data loss and system-wide issues. Detecting and measuring SDCs present unique challenges. Their low occurrence rates, dependence on hardware structure and software workloads, and correlation to environmental factors make accurate measurement complex. Addressing SDCs requires proactive measures to prevent data corruption and ensure digital integrity. Software redundancy methods provide a means to tolerate SDCs by introducing duplication or triplication of application resources. However, these methods come with their own limitations, including increased code size, altered execution patterns, and potential vulnerability to other types of failures. Understanding the nature of SDCs and developing effective mitigation strategies are crucial for maintaining digital integrity in large-scale infrastructure services. This paper sheds light on the stealthy saboteurs that silently corrupt data, emphasizes the need for comprehensive measurement techniques, and explores the limitations of existing mitigation approaches. By addressing the challenges posed by SDCs, we can fortify digital systems against these hidden threats and ensure the reliability and integrity of our digital infrastructure. George Papadimitriou 0001, Dimitris Gizopoulos, Harish Dattatraya Dixit, Sriram Sankar |
IOLTS | 4 |
| 2023 | Brief Industry Paper: Evaluating Robustness of Deep Learning-Based Recommendation Systems Against Hardware Errors: A Case StudyabstractDeep learning-based recommendation systems (DL-RMs) are industry-scale recommendation models developed by Meta, designed to make use of both categorical and numerical inputs to make personalized recommendations. To serve billions of users in real-time, DLRMs rely on high-performance hardware and accelerators within our data centers, optimizing for execution latency and recommendation quality. However, continuous technology scaling, expanding workload, and increasing hardware heterogeneity could lead to increased risk of hardware errors. Addressing this risk often involves introducing extra design redundancy, which can pose a non-negligible overhead in performance and latency. In this paper, we present a case study of evaluating DLRM robustness against hardware errors by performing an extensive error injection campaign to DLRM. Our findings unveil that DLRM is notably robust to hardware errors and we further find that embedding tables in DLRM show an especially strong robustness. Additionally, we explore a software-level error mitigation techniques, activation clipping, for mitigating the hardware errors, which improves the DLRM robustness further. This industrial case study of understanding and improving DLRM robustness can enable the system to continue to deliver timely recommendations even in the presence of hardware challenges, or reduce the timing latency overhead posed by design redundancy, enhancing overall recommendation system performance. Fan Fred Lin, Matt Xiao, Alban Desmaison, Sriram Sankar |
RTSS | 6 |
| 2020 | Optimizing Interrupt Handling Performance for Memory Failures in Large Scale Data CentersabstractIntermittent hardware failures are generally non-catastrophic and typical large-scale service infrastructures are designed to tolerate them while still serving user traffic. However, intermittent errors cause performance aberrations if they are not handled appropriately. System error reporting mechanisms send hardware interrupts to the Central Processing Unit (CPU) for handling the hardware errors. This disrupts the CPU's normal operation, which impacts the performance of the server. In this paper, we describe common intermittent hardware errors observed on server systems in a large-scale data center environment. We discuss two methodologies of handling interrupts in server systems - System Management Interrupt (SMI) and Corrected Machine Check Interrupt (CMCI). We characterize the performance of these methods in live environments as compared to prior studies that used error injection to simulate error behavior. Our experience shows that error injection methods are not reflective of production behavior. We also present a hybrid approach for handling error interrupts that achieves better performance, while preserving monitoring granularity, in large scale data center environments. Harish Dattatraya Dixit, Fan Fred Lin, Bill Holland, Matt Beadon, Zhengyu Yang 0003, Sriram Sankar |
ICPE | 6 |
| 2019 | CapNet: Exploiting Wireless Sensor Networks for Data Center Power CappingabstractAs the scale and density of data centers continue to grow, cost-effective data center management (DCM) is becoming a significant challenge for enterprises hosting large-scale online and cloud services. Machines need to be monitored, and the scale of operations mandates an automated management with high reliability and real-time performance. The limitations of today’s typical DCM network are many-fold. Primarily, it is a fixed wired network, and hence scaling it for a large number of servers increases its cost. In addition, with server densities increasing over recent years, this network also has to be cabled correctly and the management of this network parallels the complexity of managing a data network, since it needs to be networked with multiple switches and routers. In this article, we propose a wireless sensor network as a cost-effective networking solution for DCM while satisfying the reliability and latency performance requirements of DCM. We have developed CapNet, a real-time wireless sensor network for power capping, a time-critical DCM function for power management in a cluster of servers. CapNet employs an efficient event-driven protocol that triggers data collection only on the detection of a potential power capping event. We deploy and evaluate CapNet in a data center. Using server power traces, our experimental results on a cluster of 480 servers inside the data center show that CapNet can meet the real-time requirements of power capping. CapNet demonstrates the feasibility and efficacy of wireless sensor networks for time-critical DCM applications. Abusayeed Saifullah, Sriram Sankar, Jie Liu 0001, Chenyang Lu 0001, Ranveer Chandra, Bodhi Priyantha |
ACM Trans. Sens. Networks | 2 |
| 2016 | Environmental Conditions and Disk Reliability in Free-cooled Datacenters
Ioannis Manousakis, Sriram Sankar, Gregg McKnight, Thu D. Nguyen, Ricardo Bianchini |
FAST | 2 |
| 2016 | Environmental Conditions and Disk Reliability in Free-cooled Datacenters
Ioannis Manousakis, Sriram Sankar, Gregg McKnight, Thu D. Nguyen, Ricardo Bianchini |
USENIX ATC | 2 |
| 2015 | CoolProvision: underprovisioning datacenter coolingabstractCloud providers have made significant strides in reducing the cooling capital and operational costs of their datacenters, for example, by leveraging outside air ("free") cooling where possible. Despite these advances, cooling costs still represent a significant expense mainly because cloud providers typically provision their cooling infrastructure for the worst-case scenario (i.e., very high load and outside temperature at the same time). Thus, in this paper, we propose to reduce cooling costs by underprovisioning the cooling infrastructure. When the cooling is underprovisioned, there might be (rare) periods when the cooling infrastructure cannot cool down the IT equipment enough. During these periods, we can either (1) reduce the processing capacity and potentially degrade the quality of service, or (2) let the IT equipment temperature increase in exchange for a controlled degradation in reliability. To determine the ideal amount of underprovisioning, we introduce CoolProvision, an optimization and simulation framework for selecting the cheapest provisioning within performance constraints defined by the provider. CoolProvision leverages an abstract trace of the expected workload, as well as cooling, performance, power, reliability, and cost models to explore the space of potential provisionings. Using data from a real small free-cooled datacenter, our results suggest that CoolProvision can reduce the cost of cooling by up to 55%. We extrapolate our experience and results to larger cloud datacenters as well. Ioannis Manousakis, Íñigo Goiri, Sriram Sankar, Thu D. Nguyen, Ricardo Bianchini |
SoCC | 3 |
| 2014 | CapNet: A Real-Time Wireless Management Network for Data Center Power CappingabstractData center management (DCM) is increasingly becoming a significant challenge for enterprises hosting large scale online and cloud services. Machines need to be monitored, and the scale of operations mandates an automated management with high reliability and real-time performance. Existing wired networking solutions for DCM come with high cost. In this paper, we propose a wireless sensor network as a cost-effective networking solution for DCM while satisfying the reliability and latency performance requirements of DCM. We have developed Cap Net, a real-time wireless sensor network for power capping, a time-critical DCM function for power management in a cluster of servers. Cap Net employs an efficient event-driven protocol that triggers data collection only upon the detection of a potential power capping event. We deploy and evaluate Cap Net in a data center. Using server power traces, our experimental results on a cluster of 480 servers inside the data center show that Cap Net can meet the real-time requirements of power capping. Cap Net demonstrates the feasibility and efficacy of wireless sensor networks for time-critical DCM applications. Abusayeed Saifullah, Sriram Sankar, Jie Liu 0001, Chenyang Lu 0001, Ranveer Chandra, Bodhi Priyantha |
RTSS | 2 |
| 2013 | Unicorn: A System for Searching the Social GraphabstractUnicorn is an online, in-memory social graph-aware indexing system designed to search trillions of edges between tens of billions of users and entities on thousands of commodity servers. Unicorn is based on standard concepts in information retrieval, but it includes features to promote results with good social proximity. It also supports queries that require multiple round-trips to leaves in order to retrieve objects that are more than one edge away from source nodes. Unicorn is designed to answer billions of queries per day at latencies in the hundreds of milliseconds, and it serves as an infrastructural building block for Facebook's Graph Search product. In this paper, we describe the data model and query language supported by Unicorn. We also describe its evolution as it became the primary backend for Facebook's search offerings. Michael Curtiss, Iain Becker, Tudor Bosman, Sergey Doroshenko, Lucian Grijincu, Sandhya Kunnatur, Søren B. Lassen, Philip Pronin, Sriram Sankar, Guanghao Shen, Gintaras Woss, Ning Zhang 0002 |
Proc. VLDB Endow. | 10 |
| 2013 | Datacenter Scale Evaluation of the Impact of Temperature on Hard Disk Drive FailuresabstractWith the advent of cloud computing and online services, large enterprises rely heavily on their datacenters to serve end users. A large datacenter facility incurs increased maintenance costs in addition to service unavailability when there are increased failures. Among different server components, hard disk drives are known to contribute significantly to server failures; however, there is very little understanding of the major determinants of disk failures in datacenters. In this work, we focus on the interrelationship between temperature, workload, and hard disk drive failures in a large scale datacenter. We present a dense storage case study from a population housing thousands of servers and tens of thousands of disk drives, hosting a large-scale online service at Microsoft. We specifically establish correlation between temperatures and failures observed at different location granularities: (a) inside drive locations in a server chassis, (b) across server locations in a rack, and (c) across multiple racks in a datacenter. We show that temperature exhibits a stronger correlation to failures than the correlation of disk utilization with drive failures. We establish that variations in temperature are not significant in datacenters and have little impact on failures. We also explore workload impacts on temperature and disk failures and show that the impact of workload is not significant. We then experimentally evaluate knobs that control disk drive temperature, including workload and chassis design knobs. We corroborate our findings from the real data study and show that workload knobs show minimal impact on temperature. Chassis knobs like disk placement and fan speeds have a larger impact on temperature. Finally, we also show the proposed cost benefit of temperature optimizations that increase hard disk drive reliability. Sriram Sankar, Mark Shaw 0001, Kushagra Vaid, Sudhanva Gurumurthi |
ACM Trans. Storage | 1 |
| 2011 | Impact of temperature on hard disk drive reliability in large datacentersabstractWhen datacenters are pushed to their limits of operational efficiency, reducing failure rates becomes critical for maintaining high levels of healthy server operation. In this experience report, we present a dense storage case study from a large population of servers housing tens of thousands of disk drives. Previous studies have presented divergent results concerning correlation between temperature and hard disk drive failures. In our paper, we specifically establish correlation between temperatures and failures observed at different location granularities: a) inside drive locations in a server chassis, b) across server locations in a rack and c) across multiple racks in a datacenter. We also establish that temperature exhibits a stronger correlation to failures compared to the correlation of disk utilization with drive failures. Thus, we show that temperature-aware server and datacenter design plays a pivotal role in datacenter reliability. Following our case study, we present a reliability model for estimating hard disk drive failures correlated with the datacenter operating temperature. We use a physical Arrhenius model with empirically derived coefficients for our model. We show an application of the model for selecting the datacenter inlet temperature setpoint for two different server storage configurations. Finally, with the help of a datacenter cost discussion, we highlight the need to incorporate reliability-aware datacenter design for increased efficiency in large scale datacenters. Sriram Sankar, Mark Shaw 0001, Kushagra Vaid |
DSN | 1 |
| 2011 | Storage I/O generation and replay for datacenter applicationsabstractWith the advent of social networking and cloud data-stores, user data is increasingly being stored in large capacity and high performance storage systems, which account for a significant portion of the total cost of ownership of a datacenter (DC) [3]. One of the main challenges when trying to evaluate storage system options is the difficulty in replaying the entire application in all possible system configurations. Furthermore, code and datasets of DC applications are rarely available to storage system designers. This makes the development of a representative model that captures key aspects of the workload's storage profile, even more appealing. Once such a model is available, the next step is to create a tool that convincingly reproduces the application's storage behavior via a synthetic I/O access pattern. Christina Delimitrou, Sriram Sankar, Kushagra Vaid, Christoforos E. Kozyrakis |
ISPASS | 2 |
| 2011 | Energy-delay based provisioning for large datacenters: an energy-efficient and cost optimal approachabstractIt is challenging to determine the optimum number of servers required to provision for large online applications because of the conflicting mandates of (a) achieving peak performance needs and (b) minimizing unused datacenter power and capacity. Since Online Services application loads are unpredictable, datacenter operators often conservatively provision for maximum power utilization by characterizing workloads for peak load performance. In contrast, we aim to optimize the service capacity per total cost of ownership (TCO) of an Online Service datacenter deployment by characterizing the energy-delay properties of large scale datacenter workloads. We show that the peak load performance is not the energy efficient point of operation for most applications. We choose two industry-strength workloads (Internet Search and D-Process) and analyze their energy-delay behavior under varying loads. We then calculate the optimal operating point for the specific large-scale application and provision datacenter energy and capacity based on the energy-delay curves. In contrast to workload-based peak power provisioning, we show a 7% benefit in Service Capacity-per-TCO-dollar for the energy-delay characterization methodology in our cost analysis for Online Services Applications. Sriram Sankar, Kushagra Vaid, Harry Rogers |
ICPE | 1 |
| 2009 | Sensitivity-Based Optimization of Disk ArchitectureabstractStorage plays a pivotal role in the performance of many applications. Many applications, especially those that run on servers, are I/O intensive and therefore require high performance storage systems. These high-end storage systems consume a large amount of power, the bulk of which is due to the disk drives. Optimizing disk architectures is a design time as well as a run time issue and requires balancing between performance and power. There are different figures of merit, such as performance and energy, and a large space of design and runtime "knobs" that can be used to optimize disk drive behavior. Given such a large space, it is desirable to have a systematic methodology to optimally set these knobs to satisfy our figures of merit as efficiently as possible. In this paper we present the sensitivity-based optimization methodology for disk architectures (SODA), which leverages results previously obtained in digital circuit design optimization scenarios. Using detailed models of the electro-mechanical behavior of disk drives and a suite of realistic workloads, we show how SODA can aid in design and runtime optimization of disk drive architectures. Sriram Sankar, Yan Zhang 0028, Sudhanva Gurumurthi, Mircea R. Stan |
IEEE Trans. Computers | 1 |
| 2008 | Intra-disk Parallelism: An Idea Whose Time Has ComeabstractServer storage systems use a large number of disks to achieve high performance, thereby consuming a significant amount of power. In this paper, we propose to significantly reduce the power consumed by such storage systems via intra-disk parallelism, wherein disk drives can exploit parallelism in the I/O request stream. Intra-disk parallelism can facilitate replacing a large disk array with a smaller one, using the minimum number of disk drives needed to satisfy the capacity requirements. We show that the design space of intra-disk parallelism is large and present a taxonomy to formulate specific implementations within this space. Using a set of commercial workloads, we perform a limit study to identify the key performance bottlenecks that arise when we replace a storage array that is tuned to provide high performance with a single high-capacity disk drive. We show that it is possible to match, and even surpass, the performance of a storage array for these workloads by using a single disk drive of sufficient capacity that exploits intra-disk parallelism, while significantly reducing the power consumed by the storage system. We evaluate the performance and power consumption of disk arrays composed of intra-disk parallel drives, and discuss engineering and cost issues related to the implementation and deployment of such disk drives. Sriram Sankar, Sudhanva Gurumurthi, Mircea R. Stan |
ISCA | 1 |
| 2008 | Sensitivity Based Power Management of Enterprise Storage Systems
Sriram Sankar, Sudhanva Gurumurthi, Mircea R. Stan |
MASCOTS | 1 |
| 1996 | Structural Specification-Based Testing with ADLabstractThis paper describes a specification-based black-box technique for testing program units. The main contribution is the method that we have developed to derive test conditions, which are descriptions of test cases, from the formal specification of each program unit. The derived test conditions are used to guide test selection and to measure comprehensiveness of existing test suites. Our technique complements traditional code-based techniques such as statement coverage and branch coverage. It allows the tester to quickly develop a black-box test suite.In particular, this paper presents techniques for deriving test conditions from specifications written in the Assertion Definition Language (ADL) [SH94], a predicate logic-based language that is used to describe the relationships between inputs and outputs of a program unit. Our technique is fully automatable, and we are currently implementing a tool based on the techniques presented in this paper. Juei Chang, Debra J. Richardson, Sriram Sankar |
ISSTA | 3 |
| 1991 | Exploiting Locality in Maintaining Potential CausalityabstractIn distributed systems it is often important to be able to determine the temporal relationships between events generated by differentprocesses.An algorithm Sigurd Meldal, Sriram Sankar, James Vera |
PODC | 2 |
| 1990 | Application of formal specification to software maintenanceabstractThe authors describe the use of formal specifications and associated tools in addressing various aspects of software maintenance-corrective, perfective, and adaptive. They also address the refinement of the software development process to build programs that are easily maintainable. The task of software maintenance in this case includes the task of maintaining the specification, as well as the program. The authors focus on the use of Anna, a specification language for formally specifying Ada programs, to aid in maintaining Ada programs. The techniques are applicable to most other specification language and programming language environments. The tools of interest are (1) the Anna Specification Analyzer, which permits analysis of the specification for correctness with respect to the informal understanding of program behavior; and (2) the Anna Consistency Checking System, which monitors the Ada program at run time on the basis of the Anna specification.> Neel Madhav, Sriram Sankar |
ICSM | 2 |
| 1986 | Concurrent Runtime Checking of Annotated Ada Programs
David S. Rosenblum, Sriram Sankar, David C. Luckham |
FSTTCS | 2 |