Lakshmi N. Bairavasundaram

dblp:94/6213 · DBLP profile ↗
← Back
15ranked-venue papers
8as first author
0since 2021 · last 2013
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 11 · 8 first-authorSoftware engineering, systems software and programming languages · 6 · 2 first-authorDatabases, data management, data science and information retrieval · 4 · 1 first-authorSecurity and privacy · 2 · 2 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
10 papers
Storage systems · 77% Memory systems · 11% Hardware reliability and fault tolerance · 7%
Software engineering, system software, and programming languages
3 papers
Debugging and program repair · 28% Software testing · 28% Software maintenance and evolution · 28%

Topics — the 17 heaviest of 21, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Storage systems
storage reliability
0.452008
An analysis of data corruption in the storage stack · ACM Trans. Storage 2008
Parity Lost and Parity Regained · FAST 2008
An Analysis of Data Corruption in the Storage Stack · FAST 2008
Storage systems
file systems
0.122009
Tolerating File-System Mistakes with EnvyFS · USENIX ATC 2009
IRON file systems · SOSP 2005
Debugging and program repair
program repair
0.112011
How do fixes become bugs? · SIGSOFT FSE 2011
Software testing
software reliability
0.112011
How do fixes become bugs? · SIGSOFT FSE 2011
Storage systems › storage reliability
disk failure
0.122005
IRON file systems · SOSP 2005
Life or Death at Block-Level · OSDI 2004
Storage systems › storage reliability
data corruption
0.112008
An Analysis of Data Corruption in the Storage Stack · FAST 2008
Hardware reliability and fault tolerance › soft errors
silent data corruption
0.112008
An analysis of data corruption in the storage stack · ACM Trans. Storage 2008
Performance modeling and evaluation
workload characterization
0.112008
An Analysis of Data Corruption in the Storage Stack · FAST 2008
Storage systems › storage reliability
RAID
0.122008
X-RAY: A Non-Invasive Exclusive Caching Mechanism for RAIDs · ISCA 2004
Parity Lost and Parity Regained · FAST 2008
Storage systems › storage reliability › disk reliability
latent sector errors
0.112007
An analysis of latent sector errors in disk drives · SIGMETRICS 2007
Storage systems › storage reliability
file system reliability
0.112005
IRON file systems · SOSP 2005
Memory systems
cache
0.012013
Warming up storage-level caches with bonfire · FAST 2013
Memory systems › cache management
cache replacement
0.012004
X-RAY: A Non-Invasive Exclusive Caching Mechanism for RAIDs · ISCA 2004
Memory systems › memory hierarchy › cache hierarchy management
exclusive caching
0.012004
X-RAY: A Non-Invasive Exclusive Caching Mechanism for RAIDs · ISCA 2004
Empirical software engineering
mining software repositories
0.012011
An empirical study on configuration errors in commercial and open source systems · SOSP 2011
Operating systems › resource management › storage management › file systems
file system reliability
0.012009
Tolerating File-System Mistakes with EnvyFS · USENIX ATC 2009
Memory systems › cache management › storage caching
file cache
0.012004
X-RAY: A Non-Invasive Exclusive Caching Mechanism for RAIDs · ISCA 2004

Methods — techniques the papers use, named apart from their topics

semantic information extraction · 0.1large-scale field data analysis · 0.1empirical analysis · 0.1replication · 0.1parity · 0.1checksumming · 0.1trace-based simulation · 0.0machine learning · 0.0gray-box inference · 0.0
YearPublicationVenuePosition
2013 Warming up storage-level caches with bonfire
Yiying Zhang 0005, Gokul Soundararajan, Mark W. Storer, Lakshmi N. Bairavasundaram, Sethuraman Subbiah, Andrea C. Arpaci-Dusseau, Remzi H. Arpaci-Dusseau
FAST4
2011 Italian for Beginners: The Next Steps for SLO-Based Management
Lakshmi N. Bairavasundaram, Gokul Soundararajan, Vipul Mathur, Kaladhar Voruganti, Steve R. Kleiman
HotStorage1
2011 How do fixes become bugs?
abstract
Software bugs affect system reliability. When a bug is exposed in the field, developers need to fix them. Unfortunately, the bug-fixing process can also introduce errors, which leads to buggy patches that further aggravate the damage to end users and erode software vendors' reputation.
Zuoning Yin, Ding Yuan 0004, Yuanyuan Zhou 0001, Shankar Pasupathy, Lakshmi N. Bairavasundaram
SIGSOFT FSE5
2011 An empirical study on configuration errors in commercial and open source systems
abstract
Configuration errors (i.e., misconfigurations) are among the dominant causes of system failures. Their importance has inspired many research efforts on detecting, diagnosing, and fixing misconfigurations; such research would benefit greatly from a real-world characteristic study on misconfigurations. Unfortunately, few such studies have been conducted in the past, primarily because historical misconfigurations usually have not been recorded rigorously in databases.
Zuoning Yin, Xiao Ma 0014, Yuanyuan Zhou 0001, Lakshmi N. Bairavasundaram, Shankar Pasupathy
SOSP5
2009 Tolerating File-System Mistakes with EnvyFS
Lakshmi N. Bairavasundaram, Swaminathan Sundararaman, Andrea C. Arpaci-Dusseau, Remzi H. Arpaci-Dusseau
USENIX ATC1
2008 Analyzing the effects of disk-pointer corruption
abstract
The long-term availability of data stored in a file system depends on how well it safeguards on-disk pointers used to access the data. Ideally, a system would correct all pointer errors. In this paper, we examine how well corruption-handling techniques work in reality. We develop a new technique called type-aware pointer corruption to systematically explore how a file system reacts to corrupt pointers. This approach reduces the exploration space for corruption experiments and works without source code. We use type-aware pointer corruption to examine Windows NTFS and Linux ext3. We find that they rely on type and sanity checks to detect corruption, and NTFS recovers using replication in some instances. However, NTFS and ext3 do not recover from most corruptions, including many scenarios for which they possess sufficient redundant information, leading to further corruption, crashes, and unmountable file systems. We use our study to identify important lessons for handling corrupt pointers.
Lakshmi N. Bairavasundaram, Meenali Rungta, Nitin Agrawal 0001, Andrea C. Arpaci-Dusseau, Remzi H. Arpaci-Dusseau, Michael M. Swift
DSN1
2008 An Analysis of Data Corruption in the Storage Stack
Lakshmi N. Bairavasundaram, Garth R. Goodson, Bianca Schroeder, Andrea C. Arpaci-Dusseau, Remzi H. Arpaci-Dusseau
FAST1
2008 Parity Lost and Parity Regained
Andrew Krioukov, Lakshmi N. Bairavasundaram, Garth R. Goodson, Kiran Srinivasan, Randy Thelen, Andrea C. Arpaci-Dusseau, Remzi H. Arpaci-Dusseau
FAST2
2008 An analysis of data corruption in the storage stack
abstract
An important threat to reliable storage of data is silent data corruption. In order to develop suitable protection mechanisms against data corruption, it is essential to understand its characteristics. In this article, we present the first large-scale study of data corruption. We analyze corruption instances recorded in production storage systems containing a total of 1.53 million disk drives, over a period of 41 months. We study three classes of corruption: checksum mismatches, identity discrepancies, and parity inconsistencies. We focus on checksum mismatches since they occur the most. We find more than 400,000 instances of checksum mismatches over the 41-month period. We find many interesting trends among these instances, including: (i) nearline disks (and their adapters) develop checksum mismatches an order of magnitude more often than enterprise-class disk drives, (ii) checksum mismatches within the same disk are not independent events and they show high spatial and temporal locality, and (iii) checksum mismatches across different disks in the same storage system are not independent. We use our observations to derive lessons for corruption-proof system design.
Lakshmi N. Bairavasundaram, Andrea C. Arpaci-Dusseau, Remzi H. Arpaci-Dusseau, Garth R. Goodson, Bianca Schroeder
ACM Trans. Storage1
2007 An analysis of latent sector errors in disk drives
abstract
The reliability measures in today's disk drive-based storage systems focus predominantly on protecting against complete disk failures. Previous disk reliability studies have analyzed empirical data in an attempt to better understand and predict disk failure rates. Yet, very little is known about the incidence of latent sector errors i.e., errors that go undetected until the corresponding disk sectors are accessed.
Lakshmi N. Bairavasundaram, Garth R. Goodson, Shankar Pasupathy, Jiri Schindler
SIGMETRICS1
2006 Dependability Analysis of Virtual Memory Systems
abstract
Recent research has shown that even modern hard disks have complex failure modes that do not conform to "fail-stop" operation. Disks exhibit partial failures like block access errors and block corruption. Commodity operating systems are required to deal with such failures as commodity hard disks are known to be failure-prone. An important operating system component that is exposed to disk failures is the virtual memory system. In this paper, we examine the failure handling policies of different virtual memory systems for different classes of partial disk errors. We use type and context aware fault injection to explore as many of the internal code paths as possible. From experiments, we find that failure handling policies in current virtual memory systems are at best simplistic, and often inconsistent or even absent. Our fault injection technique also identifies bugs in the failure handling code in these systems. The study identifies possible reasons for poor failure handling, which can help in the design of a failure-aware virtual memory system
Lakshmi N. Bairavasundaram, Andrea C. Arpaci-Dusseau, Remzi H. Arpaci-Dusseau
DSN1
2005 Database-Aware Semantically-Smart Storage
Muthian Sivathanu, Lakshmi N. Bairavasundaram, Andrea C. Arpaci-Dusseau, Remzi H. Arpaci-Dusseau
FAST2
2005 IRON file systems
abstract
Commodity file systems trust disks to either work or fail completely, yet modern disks exhibit more complex failure modes. We suggest a new fail-partial failure model for disks, which incorporates realistic localized faults such as latent sector errors and block corruption. We then develop and apply a novel failure-policy fingerprinting framework, to investigate how commodity file systems react to a range of more realistic disk failures. We classify their failure policies in a new taxonomy that measures their Internal RObustNess (IRON), which includes both failure detection and recovery techniques. We show that commodity file system failure policies are often inconsistent, sometimes buggy, and generally inadequate in their ability to recover from partial disk failures. Finally, we design, implement, and evaluate a prototype IRON file system, Linux ixt3, showing that techniques such as in-disk checksumming, replication, and parity greatly enhance file system robustness while incurring minimal time and space overheads.
Vijayan Prabhakaran, Lakshmi N. Bairavasundaram, Nitin Agrawal 0001, Haryadi S. Gunawi, Andrea C. Arpaci-Dusseau, Remzi H. Arpaci-Dusseau
SOSP2
2004 X-RAY: A Non-Invasive Exclusive Caching Mechanism for RAIDs
abstract
RAID storage arrays often possess gigabytes of RAM for caching disk blocks. Currently, most RAID systems use LRU or LRU-like policies to manage these caches. Since these array caches do not recognize the presence of file system buffer caches, they redundantly retain many of the same blocks as those cached by the file system, thereby wasting precious cache space. In this paper, we introduce X-RAY, an exclusive RAID array caching mechanism. X-RAY achieves a high degree of (but not perfect) exclusivity through gray-box methods: by observing which files have been accessed through updates to file system meta-data, X-RAY constructs an approximate image of the contents of the file system cache and uses that information to determine the exclusive set of blocks that should be cached by the array. We use microbenchmarks to demonstrate that X-RAY's prediction of the file system buffer cache contents is highly accurate, and trace-based simulation to show that X-RAY considerably outperforms LRU and performs as well as other more invasive approaches. The main strength of the X-RAY approach is that it is easy to deploy - all performance gains are achieved without changes to the SCSI protocol or the file system above.
Lakshmi N. Bairavasundaram, Muthian Sivathanu, Andrea C. Arpaci-Dusseau, Remzi H. Arpaci-Dusseau
ISCA1
2004 Life or Death at Block-Level
Muthian Sivathanu, Lakshmi N. Bairavasundaram, Andrea C. Arpaci-Dusseau, Remzi H. Arpaci-Dusseau
OSDI2