EDBT 2026 Demo / reviewers in the wild / expert
Andy A. Hwang
dblp:27/11029
· DBLP profile ↗
4ranked-venue papers
1as first author
0since 2021 · last 2020
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 1 first-authorSoftware engineering, systems software and programming languages · 2 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
4 papers |
Storage systems · 69% Hardware reliability and fault tolerance · 23% Memory systems · 4% |
Topics — the 15 heaviest of 16, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Storage systems
crash consistency |
0.4 | 1 | 2020 | The Reliability of Modern File Systems in the face of SSD Errors · ACM Trans. Storage 2020 |
Hardware reliability and fault tolerance
fault injection |
0.4 | 1 | 2020 | The Reliability of Modern File Systems in the face of SSD Errors · ACM Trans. Storage 2020 |
Storage systems
file systems |
0.4 | 1 | 2020 | The Reliability of Modern File Systems in the face of SSD Errors · ACM Trans. Storage 2020 |
Storage systems › file systems
journaling file system |
0.4 | 1 | 2020 | The Reliability of Modern File Systems in the face of SSD Errors · ACM Trans. Storage 2020 |
Storage systems › flash and SSD
SSD reliability |
0.4 | 1 | 2020 | The Reliability of Modern File Systems in the face of SSD Errors · ACM Trans. Storage 2020 |
Storage systems › storage reliability
file system reliability |
0.4 | 1 | 2019 | Evaluating File System Reliability on Solid State Drives · USENIX ATC 2019 |
Storage systems
storage reliability |
0.4 | 1 | 2019 | Evaluating File System Reliability on Solid State Drives · USENIX ATC 2019 |
Memory systems
DRAM |
0.1 | 1 | 2012 | Cosmic rays don't strike twice: understanding the nature of DRAM errors and the implications for system design · ASPLOS 2012 |
Hardware reliability and fault tolerance › memory reliability
DRAM errors |
0.1 | 1 | 2012 | Cosmic rays don't strike twice: understanding the nature of DRAM errors and the implications for system design · ASPLOS 2012 |
Hardware reliability and fault tolerance › memory reliability
DRAM error characterization |
0.1 | 1 | 2012 | Cosmic rays don't strike twice: understanding the nature of DRAM errors and the implications for system design · ASPLOS 2012 |
Hardware reliability and fault tolerance
soft errors |
0.1 | 1 | 2012 | Cosmic rays don't strike twice: understanding the nature of DRAM errors and the implications for system design · ASPLOS 2012 |
Energy-efficient computing
thermal management |
0.1 | 1 | 2012 | Temperature management in data centers: why some (might) like it hot · SIGMETRICS 2012 |
Storage systems › flash and SSD
solid-state drive |
0.1 | 1 | 2020 | The Reliability of Modern File Systems in the face of SSD Errors · ACM Trans. Storage 2020 |
Storage systems
flash and SSD |
0.1 | 1 | 2019 | Evaluating File System Reliability on Solid State Drives · USENIX ATC 2019 |
Cloud and datacenter computing › datacenter operations
datacenter reliability |
0.0 | 1 | 2012 | Cosmic rays don't strike twice: understanding the nature of DRAM errors and the implications for system design · ASPLOS 2012 |
Methods — techniques the papers use, named apart from their topics
fault injection framework · 0.4production measurement study · 0.1page retirement policy analysis · 0.1field data analysis · 0.1experimental testbed · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2020 | The Reliability of Modern File Systems in the face of SSD ErrorsabstractAs solid state drives (SSDs) are increasingly replacing hard disk drives, the reliability of storage systems depends on the failure modes of SSDs and the ability of the file system layered on top to handle these failure modes. While the classical paper on IRON File Systems provides a thorough study of the failure policies of three file systems common at the time, we argue that 13 years later it is time to revisit file system reliability with SSDs and their reliability characteristics in mind, based on modern file systems that incorporate journaling, copy-on-write, and log-structured approaches and are optimized for flash. This article presents a detailed study, spanning ext4, Btrfs, and F2FS, and covering a number of different SSD error modes. We develop our own fault injection framework and explore over 1,000 error cases. Our results indicate that 16% of these cases result in a file system that cannot be mounted or even repaired by its system checker. We also identify the key file system metadata structures that can cause such failures, and, finally, we recommend some design guidelines for file systems that are deployed on top of SSDs. Shehbaz Jaffer, Stathis Maneas, Andy A. Hwang, Bianca Schroeder |
ACM Trans. Storage | 3 |
| 2019 | Evaluating File System Reliability on Solid State Drives
Shehbaz Jaffer, Stathis Maneas, Andy A. Hwang, Bianca Schroeder |
USENIX ATC | 3 |
| 2012 | Cosmic rays don't strike twice: understanding the nature of DRAM errors and the implications for system designabstractMain memory is one of the leading hardware causes for machine crashes in today's datacenters. Designing, evaluating and modeling systems that are resilient against memory errors requires a good understanding of the underlying characteristics of errors in DRAM in the field. While there have recently been a few first studies on DRAM errors in production systems, these have been too limited in either the size of the data set or the granularity of the data to conclusively answer many of the open questions on DRAM errors. Such questions include, for example, the prevalence of soft errors compared to hard errors, or the analysis of typical patterns of hard errors. In this paper, we study data on DRAM errors collected on a diverse range of production systems in total covering nearly 300 terabyte-years of main memory. As a first contribution, we provide a detailed analytical study of DRAM error characteristics, including both hard and soft errors. We find that a large fraction of DRAM errors in the field can be attributed to hard errors and we provide a detailed analytical study of their characteristics. As a second contribution, the paper uses the results from the measurement study to identify a number of promising directions for designing more resilient systems and evaluates the potential of different protection mechanisms in the light of realistic error patterns. One of our findings is that simple page retirement policies might be able to mask a large number of DRAM errors in production systems, while sacrificing only a negligible fraction of the total DRAM in the system. Andy A. Hwang, Ioan A. Stefanovici, Bianca Schroeder |
ASPLOS | 1 |
| 2012 | Temperature management in data centers: why some (might) like it hotabstractThe energy consumed by data centers is starting to make up a significant fraction of the world's energy consumption and carbon emissions. A large fraction of the consumed energy is spent on data center cooling, which has motivated a large body of work on temperature management in data centers. Interestingly, a key aspect of temperature management has not been well understood: controlling the setpoint temperature at which to run a data center's cooling system. Most data centers set their thermostat based on (conservative) suggestions by manufacturers, as there is limited understanding of how higher temperatures will affect the system. At the same time, studies suggest that increasing the temperature setpoint by just one degree could save 2-5% of the energy consumption. This paper provides a multi-faceted study of temperature management in data centers. We use a large collection of field data from different production environments to study the impact of temperature on hardware reliability, including the reliability of the storage subsystem, the memory subsystem and server reliability as a whole. We also use an experimental testbed based on a thermal chamber and a large array of benchmarks to study two other potential issues with higher data center temperatures: the effect on server performance and power. Based on our findings, we make recommendations for temperature management in data centers, that create the potential for saving energy, while limiting negative effects on system reliability and performance. Nosayba El-Sayed, Ioan A. Stefanovici, George Amvrosiadis, Andy A. Hwang, Bianca Schroeder |
SIGMETRICS | 4 |