EDBT 2026 Demo / reviewers in the wild / expert
Mai Zheng
dblp:47/3537
· DBLP profile ↗
37ranked-venue papers
6as first author
20since 2021 · last 2025
0000-0002-0741-3436ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 32 · 5 first-author · 17 since 2021Computer networks · 3 · 3 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Be Aware of Metadata Corruption in Parallel File System: It can be Silent and CatastrophicabstractHigh-Performance Computing (HPC) systems rely on parallel file systems (PFS) like Lustre to reliably manage large-scale data. Unfortunately, such data stored in PFS are subjected to corruption due to hardware failures, power outages, on-disk corruption, in-memory corruption, and software bugs. Previous work has shown the impacts of general data corruptions on scientific applications, but barely investigated the impacts of corruptions on metadata, which is a special type of data for many critical PFS functionalities. Since the metadata are accessed and updated frequently, the metadata corruption is non-negligible, and the impact on both HPC systems and applications is often complicated. In this study, we systematically studied the effects of PFS metadata corruption on representative scientific applications and workflows through fault injections. We observed various abnormal behaviors of applications against metadata corruptions, such as failing to finish execution or successfully finishing execution but generating wrong outputs silently. We also showed that neither the defensive programming practice from the application developers nor existing metadata checking mechanism deployed in PFS can resolve the metadata corruption issues or guild applications to react correctly. To address this issue, we further propose a metadata checksum mechanism as a system-level mitigation strategy to detect and handle metadata corruptions at runtime. We implement a prototype of the checksum mechanism on Lustre file system using FUSE interface and evaluate its functionality and performance overhead. Our results show the proposed checksum mechanism can effectively detect metadata corruptions during data accesses with extremely low overhead and stop the applications from continuous execution and generating wrong results. Saisha Kamat, Mai Zheng, Bo Fang 0002, Dong Dai 0001 |
IPDPS | 2 |
| 2025 | Design and implementation of ARA wireless living lab for rural broadband and applications
Taimoor Ul Islam, Joshua Ofori Boateng, Md Nadim, Guoying Zu, Mukaram Shahid, Tianyi Zhang 0016, Salil Reddy, Wei Xu 0056, Ataberk Atalar, Vincent Lee, Yung-fu Chen, Evan Gossling, Elisabeth Permatasari, Christ Somiah, Owen Perrin, Zhibo Meng, Reshal Afzal, Sarath Babu 0001, Mohammed Soliman, Ali Hussain, Daji Qiao, Mai Zheng, Ozdal Boyraz, Anish Arora, Mohamed Y. Selim, Arsalan Ahmad, Myra B. Cohen, Mike Luby, Ranveer Chandra, James Gross, Kate Keahey, Hongwei Zhang 0001 |
Comput. Networks | 23 |
| 2025 | Analyzing Configuration Dependencies of File SystemsabstractFile systems play an essential role in modern society for managing precious data. To meet diverse needs, they often support many configuration parameters. Such flexibility comes at the price of additional complexity which can lead to subtle configuration-related issues. To address this challenge, we study the configuration-related issues of two major file systems (i.e., Ext4 and XFS) in depth, and identify a prevalent pattern called multilevel configuration dependencies. Based on the study, we build an extensible tool called ConfD to extract the dependencies automatically, and create a set of plugins to address different configuration-related issues. Our experiments on Ext4, XFS and a modern copy-on-write file system (i.e., ZFS) show that ConfD was able to extract 160 configuration dependencies for the file systems with a low false positive rate. Moreover, the dependency-guided plugins can identify various configuration issues (e.g., mishandling of configurations, regression test failures induced by valid configurations). In addition, we also explore the applicability of ConfD on a popular storage engine (i.e., WiredTiger). We hope that this comprehensive analysis of configuration dependencies of storage systems can shed light on addressing configuration-related challenges for the system community in general. Tabassum Mahmud, Om Rameshwar Gatla, Carson Love, Ryan Bumann, Varun S. Girimaji, Mai Zheng |
ACM Trans. Comput. Syst. | 7 |
| 2024 | Revisiting Erasure Codes: A Configuration PerspectiveabstractErasure coding (EC) plays a crucial role in the fault tolerance of modern distributed storage systems (DSS). Inspired by recent research on storage configuration, we study the configuration sensitivity of EC in real DSS in this paper. We systematically inject faults to trigger EC recovery under various configurations, and measure the impact on recovery time and storage overhead quantitatively. Our results show that configurations may affect the EC recovery time significantly (e.g., up to 426%). More interestingly, theoretically superior codes may perform worse in DSS under certain configurations. Also, there is a system checking period before EC recovery that accounts for 41% to 58% of the overall system recovery time, which has been largely ignored in previous studies. Finally, in terms of storage overhead, EC may introduce 32.3% to 72.0% more write amplification (WA) than the theoretical expectation, and we derive a formula to help estimate WA more precisely. Our work suggests the importance of considering the context of real DSS for EC research, and we hope the methodology and findings can contribute to a firmer footing for EC optimization in practice. Runzhou Han, Tabassum Mahmud, Zeren Yang, Vladislav Esaulov, Lipeng Wan 0001, Yong Chen 0001, Jim Wayda, Matthew Wolf, Mai Zheng |
HotStorage | 10 |
| 2024 | Demo: Ara Pawr Wireless Living Lab for Smart and Connected Rural CommunitiesabstractARA is an at-scale Platform for Advanced Wireless Research (PAWR), specifically tailored to the unique community, application, and economic context of rural regions. It features the first-of-its-kind real-world implementation of long-distance, high-capacity wireless backhaul and access systems spanning over 30 km in diameter. Leveraging both software-defined radios and programmable Commercial Off-The-Shelf (COTS) systems, ARA orchestrates the wireless resources alongside the networking and compute resources for enabling end-to-end experiments involving user equipment, base stations, edge computing, and cloud infrastructure. Such an integration facilitates the coevolution of rural-focused wireless innovation and applications, while helping to advance the frontiers of advanced Next-G wireless systems such as Open RAN. As of summer 2024, ARA is publicly accessible with 7 base stations (BSes) and over 30 user equipment (UEs). In this demo, we share advanced wireless research experiments enabled by ARA, involving MU-MIMO in TV White Space (TVWS) bands, long-range mmWave and microwave backhaul communications, and open-source 5G NR protocol stacks such as srsRAN and OpenAirInterface (OAI). Taimoor Ul Islam, Joshua Ofori Boateng, Md Nadim, Guoying Zu, Mukaram Shahid, Tianyi Zhang 0016, Salil Reddy, Wei Xu 0056, Ataberk Atalar, Vincent Lee, Evan Gossling, Elisabeth Permatasari, Zhibo Meng, Sarath Babu 0001, Mohammed Soliman, Ali Hussain, Daji Qiao, Mai Zheng, Ozdal Boyraz, Anish Arora, Mohamed Y. Selim, Arsalan Ahmad, Myra B. Cohen, Hongwei Zhang 0001 |
ICNP | 19 |
| 2024 | PROV-IO$^+$+: A Cross-Platform Provenance Framework for Scientific Data on HPC SystemsabstractData provenance, or data lineage, describes the life cycle of data. In scientific workflows on HPC systems, scientists often seek diverse provenance (e.g., origins of data products, usage patterns of datasets). Unfortunately, existing provenance solutions cannot address the challenges due to their incompatible provenance models and/or system implementations. In this paper, we analyze four representative scientific workflows in collaboration with the domain scientists to identify concrete provenance needs. Based on the first-hand analysis, we propose a provenance framework called PROV-IO$^+$, which includes an I/O-centric provenance model for describing scientific data and the associated I/O operations and environments precisely. Moreover, we build a prototype of PROV-IO$^+$to enable end-to-end provenance support on real HPC systems with little manual effort. The PROV-IO$^+$framework can support both containerized and non-containerized workflows on different HPC platforms with flexibility in selecting various classes of provenance. Our experiments with realistic workflows show that PROV-IO$^+$can address the provenance needs of the domain scientists effectively with reasonable performance (e.g., less than 3.5% tracking overhead for most experiments). Moreover, PROV-IO$^+$outperforms a state-of-the-art system (i.e., ProvLake) in our experiments. Runzhou Han, Mai Zheng, Surendra Byna, Houjun Tang, Bin Dong 0002, Dong Dai 0001, Yong Chen 0001, Dongkyun Kim, Joseph Hassoun, David Thorsley |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2023 | λFS: A Scalable and Elastic Distributed File System Metadata Service using Serverless FunctionsabstractThe metadata service (MDS) sits on the critical path for distributed file system (DFS) operations, and therefore it is key to the overall performance of a large-scale DFS. Common "serverful" MDS architectures, such as a single server or cluster of servers, have a significant shortcoming: either they are not scalable, or they make it difficult to achieve an optimal balance of performance, resource utilization, and cost. A modern MDS requires a novel architecture that addresses this shortcoming. Benjamin Carver, Runzhou Han, Mai Zheng, Yue Cheng 0001 |
ASPLOS (4) | 4 |
| 2023 | ConfD: Analyzing Configuration Dependencies of File Systems for Fun and Profit
Tabassum Mahmud, Om Rameshwar Gatla, Carson Love, Ryan Bumann, Mai Zheng |
FAST | 6 |
| 2023 | FaultyRank: A Graph-based Parallel File System CheckerabstractSimilar to local file system checkers such as e2fsck for Ext4, a parallel file system (PFS) checker ensures the file system's correctness. The basic idea of file system checkers is straightforward: important metadata are stored redundantly in separate places for cross-checking; inconsistent metadata will be repaired or overwritten by its ‘more correct' counterpart, which is defined by the developers. Unfortunately, implementing the idea for PFSes is non-trivial due to the system complexity. Although many popular parallel file systems already contain dedicated checkers (e.g., LFSCK for Lustre, BeeGFS-FSCK for BeeGFS, mmfsck for GPFS), the existing checkers often cannot detect or repair inconsistencies accurately due to one fundamental limitation: they rely on a fixed set of consistency rules predefined by developers, which cannot cover the various failure scenarios that may occur in practice.In this study, we propose a new graph-based method to build PFS checkers. Specifically, we model important PFS metadata into graphs, then generalize the logic of cross-checking and repairing into graph analytic tasks. We design a new graph algorithm, FaultyRank, to quantitatively calculate the correctness of each metadata object. By leveraging the calculated correctness, we are able to recommend the most promising repairs to users. Based on the idea, we implement a prototype of FaultyRank on Lustre, one of the most widely used parallel file systems, and compare it with Lustre's default file system checker LFSCK. Our experiments show that FaultyRank can achieve the same checking and repairing logic as LFSCK. Moreover, it is capable of detecting and repairing complicated PFS consistency issues that LFSCK can not handle. We also show the performance advantage of FaultyRank compared with LFSCK. Through this study, we believe FaultyRank opens a new opportunity for building PFS checkers effectively and efficiently. Saisha Kamat, Abdullah Al Raqibul Islam, Mai Zheng, Dong Dai 0001 |
IPDPS | 3 |
| 2023 | Drill: Log-based Anomaly Detection for Large-scale Storage Systems Using Source Code AnalysisabstractLarge-scale storage systems, a critical part of modern computing systems, are subject to various runtime bugs, failures, and anomalies in production. Identifying their anomalies at runtime is thus critical for users and administrators. Since runtime logs record the important status of the systems, log-based anomaly detection has been studied extensively for timely identifying system malfunctions. However, existing log-based anomaly detection solutions share common limitations in representing log entries accurately and robustly, hence can not effectively handle log entries that were not seen in the historical logs, which is a common real-world scenario due to logs' inherent rarity and the continuous evolution of the systems. To address the issues of existing methods, we propose Drill, a new log pre-processing method to generate high-quality vector representation of runtime logs by leveraging both storage system-specific sentiment-classifying language models and log contexts built from the source code. Through extensive evaluations of two representative distributed storage systems (Apache HDFS and Lustre), we show that Drill can achieve up to 41% improvement when compared with state-of-the-art anomaly detection solutions, showing it is a promising solution for general anomaly detection. Di Zhang 0015, Chris Egersdoerfer, Tabassum Mahmud, Mai Zheng, Dong Dai 0001 |
IPDPS | 4 |
| 2023 | ARA PAWR: Wireless Living Lab for Smart and Connected Rural CommunitiesabstractAs the Platform for Advanced Wireless Research (PAWR) in rural broadband, the ARA wireless living lab features the deployment of first-of-its-kind wireless access and backhaul platforms in real-world agriculture and rural settings, and preliminary experiments have demonstrated very promising results, e.g., up to 3.2 Gbps wireless access throughput and more than 10 Gbps throughput across a wireless backhaul link of over 10 km. ARA is expected to be publicly released for broad community use starting in September 2023. Through this demo, we plan to share, for the first time, with the wireless research community the transformative research experiments enabled by ARA. To stimulate discussion and community participation, we will demonstrate a few example experiments ranging from MU-MIMO in TV White Space (TVWS) bands to long-range mmWave and microwave backhaul communications, as well as open-source 5G NR protocol stacks such as srsRAN and OpenAirInterface. Taimoor Ul Islam, Joshua Ofori Boateng, Guoying Zu, Mukaram Shahid, Md Nadim, Wei Xu 0056, Tianyi Zhang 0016, Salil Reddy, Ataberk Atalar, Yung-fu Chen, Sarath Babu 0001, Hongwei Zhang 0001, Daji Qiao, Mai Zheng, Ozdal Boyraz, Anish Arora, Mohamed Y. Selim, Myra B. Cohen |
MobiCom | 15 |
| 2023 | Data Distribution for Heterogeneous Storage SystemsabstractThe exponential growth of data in many science and engineering domains poses significant challenges to storage systems. Data distribution is a critical component in large-scale distributed storage systems and plays a vital role in placing petabytes of data and beyond, among tens to hundreds of thousands of storage devices. Meantime, heterogeneous storage systems, such as those having devices with hard disk drives (HDDs) and storage class memories (SCMs), have become increasingly popular for massive data storage due to their distinct and complement characteristics. This paper presents a new data distribution algorithm called SUORA (Scalable and Uniform storage via Optimally-adaptive and Random number Addressing) specifically for heterogeneous devices to maximize the benefits of them. SUORA provides a fully symmetric, highly efficient methodology to distribute data across a hybrid and tiered storage cluster. It divides heterogeneous devices into different buckets and segments, and adopts pseudo-random functions to map data onto them with the balanced consideration of capacity, performance and life-time. By analyzing hotness and access patterns, SUORA gradually moves hot data from HDDs to SCMs to optimize the throughput, and moves cold data reversely for load balance. It combines data replication with migration to significantly reduce movement overhead while making data placement more adaptive to different workloads. Extensive evaluations on simulation and Sheepdog storage system show that, with considering distinct characteristics of various devices thoroughly, SUORA improves the overall performance efficiency of heterogeneous storage systems. Yong Chen 0001, Mai Zheng, Weiping Wang 0005 |
IEEE Trans. Computers | 3 |
| 2023 | Understanding Persistent-memory-related Issues in the Linux KernelabstractPersistent memory (PM) technologies have inspired a wide range of PM-based system optimizations. However, building correct PM-based systems is difficult due to the unique characteristics of PM hardware. To better understand the challenges as well as the opportunities to address them, this article presents a comprehensive study of PM-related issues in the Linux kernel. By analyzing 1,553 PM-related kernel patches in depth and conducting experiments on reproducibility and tool extension, we derive multiple insights in terms of PM patch categories, PM bug patterns, consequences, fix strategies, triggering conditions, and remedy solutions. We hope our results could contribute to the development of robust PM-based storage systems. Om Rameshwar Gatla, Wei Xu 0056, Mai Zheng |
ACM Trans. Storage | 4 |
| 2022 | Understanding configuration dependencies of file systemsabstractFile systems have many configuration parameters. Such flexibility comes at the price of additional complexity which could lead to subtle configuration-related issues. To address the challenge, we study the potential configuration dependencies of a representative file system (i.e., Ext4), and identify a prevalent pattern called multi-level configuration dependencies. We build a static analyzer to extract the dependencies and leverage the information to address different configuration issues. Our preliminary prototype is able to extract 64 multi-level dependencies with a low false positive rate. Additionally, we can identify multiple configuration issues effectively. Tabassum Mahmud, Om Rameshwar Gatla, Mai Zheng |
HotStorage | 4 |
| 2022 | PROV-IO: An I/O-Centric Provenance Framework for Scientific Data on HPC SystemsabstractcData provenance, or data lineage, describes the life cycle of data. In scientific workflows on HPC systems, scientists often seek diverse provenance (e.g., origins of data products, usage patterns of datasets). Unfortunately, existing provenance solutions cannot address the challenges due to their incompatible provenance models and/or system implementations. Runzhou Han, Surendra Byna, Houjun Tang, Bin Dong 0002, Mai Zheng |
HPDC | 5 |
| 2022 | On the Reproducibility of Bugs in File-System Aware Storage ApplicationsabstractMany storage applications such as file system checkers, defragmentation tools, etc. require a detailed understanding of file systems. Such file-system aware applications play an essential role today, but unfortunately they are error-prone. To better understand the challenges as well as the opportunities to address the issues, this paper presents an empirical study of real world bugs in file-system aware storage applications. By analyzing 59 bug cases from 4 representative applications in depth, we derive multiple insights in terms of general bug patterns, triggering conditions, and implications for building effective tools to address the issues. We hope that our study and the resulting dataset could contribute to the development of reliability tools for building robust file-system aware storage applications in general. Tabassum Mahmud, Om Rameshwar Gatla, Runzhou Han, Yong Chen 0001, Mai Zheng |
NAS | 6 |
| 2022 | Ensuring high reliability and performance with low space overhead for deduplicated and delta-compressed storage systemsabstractAbstract Data deduplication is a widely used technique to remove duplicate data to reduce the storage overhead. However, deduplication typically cannot eliminate the redundancy among nonidentical but similar data chunks. To reduce the storage overhead further, delta compression is often applied to compress the post‐deduplication data. While the two techniques are effective in saving storage space, they introduce complex references among data chunks, which inevitably undermines the system reliability and introduces fragmentation that may degrade the restore performance. In this paper, we observe that the delta compressed chunks (DCCs) are much smaller than regular chunks (non‐DCCs). Also, most fragmentation caused by the base chunk of DCCs remain fragmented in consecutive backups. Based on these observations, we introduce a framework called , which combines replication and erasure coding and uses History‐aware Delta Selection to ensure high reliability and restore performance. Specifically, uses a delta‐utilization‐aware filter and a cooperative cache scheme (CCS) to maintain cache locality and avoid unnecessary container reads, respectively. Moreover, the system selectively performs delta compression by historical information to avoid cyclic fragmentation in consecutive backups. Experimental results based on four real‐world datasets demonstrate that significantly improves the restore performance by 58.3%–76.7% with a low storage overhead. Chunxue Zuo, Fang Wang 0001, Mai Zheng, Yuchong Hu, Dan Feng 0001 |
Concurr. Comput. Pract. Exp. | 3 |
| 2022 | A Study of Failure Recovery and Logging of High-Performance Parallel File SystemsabstractLarge-scale parallel file systems (PFSs) play an essential role in high-performance computing (HPC). However, despite their importance, their reliability is much less studied or understood compared with that of local storage systems or cloud storage systems. Recent failure incidents at real HPC centers have exposed the latent defects in PFS clusters as well as the urgent need for a systematic analysis. To address the challenge, we perform a study of the failure recovery and logging mechanisms of PFSs in this article. First, to trigger the failure recovery and logging operations of the target PFS, we introduce a black-box fault injection tool called PFault , which is transparent to PFSs and easy to deploy in practice. PFault emulates the failure state of individual storage nodes in the PFS based on a set of pre-defined fault models and enables examining the PFS behavior under fault systematically. Next, we apply PFault to study two widely used PFSs: Lustre and BeeGFS. Our analysis reveals the unique failure recovery and logging patterns of the target PFSs and identifies multiple cases where the PFSs are imperfect in terms of failure handling. For example, Lustre includes a recovery component called LFSCK to detect and fix PFS-level inconsistencies, but we find that LFSCK itself may hang or trigger kernel panics when scanning a corrupted Lustre. Even after the recovery attempt of LFSCK, the subsequent workloads applied to Lustre may still behave abnormally (e.g., hang or report I/O errors). Similar issues have also been observed in BeeGFS and its recovery component BeeGFS-FSCK. We analyze the root causes of the abnormal symptoms observed in depth, which has led to a new patch set to be merged into the coming Lustre release. In addition, we characterize the extensive logs generated in the experiments in detail and identify the unique patterns and limitations of PFSs in terms of failure logging. We hope this study and the resulting tool and dataset can facilitate follow-up research in the communities and help improve PFSs for reliable high-performance computing. Runzhou Han, Om Rameshwar Gatla, Mai Zheng, Jinrui Cao, Di Zhang 0015, Dong Dai 0001, Yong Chen 0001, Jonathan E. Cook 0001 |
ACM Trans. Storage | 3 |
| 2021 | SentiLog: Anomaly Detecting on Parallel File Systems via Log-based Sentiment AnalysisabstractAs core components of High-performance computing (HPC) platforms, parallel file systems (PFSes) grow quickly in scale and complexity, hence are subject to various failures and anomalies. Identifying their anomalies in runtime is critically helpful for HPC operators and administrators. Analyzing the runtime logs to detect the anomalies of large-scale systems has been proven effective in many recent studies. However, applying them to parallel file systems logs faces significant challenges due to the large volume and irregularity of PFSes logs. This study proposes SentiLog, a new approach to analyzing PFSes system logs for detecting anomalies. Unlike existing solutions, SentiLog works by training a general sentimental, natural language model based on the logging-relevant source code collected from a set of PFSes. In this way, SentiLog learns information embedded by developers from the source code. Our preliminary results show SentiLog is able to accurately predict anomalies and performs better than state-of-the-art log analysis solutions on two representative PFSes (Lustre and BeeGFS). This preliminary study shows sentiment analysis could be a promising method to analyze complex and irregular system logs. Di Zhang 0015, Dong Dai 0001, Runzhou Han, Mai Zheng |
HotStorage | 4 |
| 2021 | A study of persistent memory bugs in the Linux kernelabstractPersistent memory (PM) technologies have inspired a wide range of PM-based system optimizations. However, building correct PM-based systems is difficult due to the unique characteristics of PM hardware. To better understand the challenges as well as the opportunities to address them, this paper presents a comprehensive study of PM-related bugs in the Linux kernel. By analyzing 1,350 PM-related kernel patches in depth, we derive multiple insights in terms of PM patch categories, PM bug patterns, consequences, and fix strategies. We hope our results could contribute to the development of effective PM bug detectors and robust PM-based systems. Om Rameshwar Gatla, Mai Zheng |
SYSTOR | 4 |
| 2020 | Position: On Failure Diagnosis of the Storage Stack
Om Rameshwar Gatla, Runzhou Han, Mai Zheng |
HotStorage | 4 |
| 2019 | A Performance Study of Lustre File System Checker: Bottlenecks and PotentialsabstractLustre, as one of the most popular parallel file systems in high-performance computing (HPC), provides POSIX interface and maintains a large set of POSIX-related metadata, which could be corrupted due to hardware failures, software bugs, configuration errors, etc. The Lustre file system checker (LFSCK) is the remedy tool to detect metadata inconsistencies and to restore a corrupted Lustre to a valid state, hence is critical for reliable HPC. Unfortunately, in practice, LFSCK runs slow in large deployment, making system administrators reluctant to use it as a routine maintenance tool. Consequently, cascading errors may lead to unrecoverable failures, resulting in significant downtime or even data loss. Given the fact that HPC is rapidly marching to Exascale and much larger Lustre file systems are being deployed, it is critical to understand the performance of LFSCK. In this paper, we study the performance of LFSCK to identify its bottlenecks and analyze its performance potentials. Specifically, we design an aging method based on real-world HPC workloads to age Lustre to representative states, and then systematically evaluate and analyze how LFSCK runs on such an aged Lustre via monitoring the utilization of various resources. From our experiments, we find out that the design and implementation of LFSCK is sub-optimal. It consists of scalability bottleneck on the metadata server (MDS), relatively high fan-out ratio in network utilization, and unnecessary blocking among internal components. Based on these observations, we discussed potential optimization and present some preliminary results. Dong Dai 0001, Om Rameshwar Gatla, Mai Zheng |
MSST | 3 |
| 2019 | Lessons and Actions: What We Learned from 10K SSD-Related Storage System Failures
Erci Xu, Mai Zheng, Yikang Xu, Jiesheng Wu |
USENIX ATC | 2 |
| 2018 | Towards Robust File System Checkers
Om Rameshwar Gatla, Muhammad Hameed, Mai Zheng, Viacheslav Dubeyko, Adam Manzanares, Filip Blagojevic, Cyril Guyot, Robert Mateescu |
FAST | 3 |
| 2018 | PFault: A General Framework for Analyzing the Reliability of High-Performance Parallel File SystemsabstractHigh-performance parallel file systems (PFSes) are of prime importance today. However, despite the importance, their reliability is much less studied compared with that of local storage systems, largely due to the lack of an effective analysis methodology. Jinrui Cao, Om Rameshwar Gatla, Mai Zheng, Dong Dai 0001, Vidya Eswarappa, Yan Mu, Yong Chen 0001 |
ICS | 3 |
| 2018 | Towards Robust File System CheckersabstractFile systems may become corrupted for many reasons despite various protection techniques. Therefore, most file systems come with a checker to recover the file system to a consistent state. However, existing checkers are commonly assumed to be able to complete the repair without interruption, which may not be true in practice. In this work, we demonstrate via fault injection experiments that checkers of widely used file systems (EXT4, XFS, BtrFS, and F2FS) may leave the file system in an uncorrectable state if the repair procedure is interrupted unexpectedly. To address the problem, we first fix the ordering issue in the undo logging of e2fsck and then build a general logging library (i.e., rfsck-lib) for strengthening checkers. To demonstrate the practicality, we integrate rfsck-lib with existing checkers and create two new checkers: rfsck-ext, a robust checker for Ext-family file systems, and rfsck-xfs, a robust checker for XFS file systems, both of which require only tens of lines of modification to the original versions. Both rfsck-ext and rfsck-xfs are resilient to faults in our experiments. Also, both checkers incur reasonable performance overhead (i.e., up to 12%) compared to the original unreliable versions. Moreover, rfsck-ext outperforms the patched e2fsck by up to nine times while achieving the same level of robustness. Om Rameshwar Gatla, Mai Zheng, Muhammad Hameed, Viacheslav Dubeyko, Adam Manzanares, Filip Blagojevic, Cyril Guyot, Robert Mateescu |
ACM Trans. Storage | 2 |
| 2017 | Understanding the Fault Resilience of File System Checkers
Om Rameshwar Gatla, Mai Zheng |
HotStorage | 2 |
| 2017 | Reliability Analysis of SSDs Under Power FaultabstractModern storage technology (solid-state disks (SSDs), NoSQL databases, commoditized RAID hardware, etc.) brings new reliability challenges to the already-complicated storage stack. Among other things, the behavior of these new components during power faults—which happen relatively frequently in data centers—is an important yet mostly ignored issue in this dependability-critical area. Understanding how new storage components behave under power fault is the first step towards designing new robust storage systems. In this article, we propose a new methodology to expose reliability issues in block devices under power faults. Our framework includes specially designed hardware to inject power faults directly to devices, workloads to stress storage components, and techniques to detect various types of failures. Applying our testing framework, we test 17 commodity SSDs from six different vendors using more than three thousand fault injection cycles in total. Our experimental results reveal that 14 of the 17 tested SSD devices exhibit surprising failure behaviors under power faults, including bit corruption, shorn writes, unserializable writes, metadata corruption, and total device failure. Mai Zheng, Joseph A. Tucek, Mark Lillibridge, Bill W. Zhao, Elizabeth S. Yang |
ACM Trans. Comput. Syst. | 1 |
| 2016 | Emulating Realistic Flash Device Errors with High FidelityabstractModern storage software is designed to guarantee data integrity and consistency based on decades of experience with the foibles of hard disk drives. However, recent research shows that flash-based SSDs may fail in different and surprising ways, breaking their contract with the software above them. This raises the question of whether the software stack's guarantees to users still hold when SSDs are substituted. In this position paper, we propose a framework to emulate the erroneous behaviors of SSDs for understanding the failure resilience of the storage software stack. We first model the device behaviors reported in previous work and create a database of realistic patterns of device errors. Based on the patterns, the framework manipulates the I/O commands at the driver level and emulates the device errors with minimal disturbance to the target software. Preliminary results show that the framework is able to emulate the device errors with high fidelity, which provides a solid foundation for further studying the failure handling of the storage software stack. Simeng Wang, Jinrui Cao, Danny V. Murillo, Yiliang Shi, Mai Zheng |
NAS | 5 |
| 2016 | An adaptive algorithm for scheduling parallel jobs in meteorological Cloud
Yongsheng Hao, Mai Zheng |
Knowl. Based Syst. | 3 |
| 2014 | Torturing Databases for Fun and Profit
Mai Zheng, Joseph A. Tucek, Dachuan Huang, Mark Lillibridge, Elizabeth S. Yang, Bill W. Zhao, Shashank Singh 0003 |
OSDI | 1 |
| 2014 | GMRace: Detecting Data Races in GPU Programs via a Low-Overhead SchemeabstractIn recent years, GPUs have emerged as an extremely cost-effective means for achieving high performance. While languages like CUDA and OpenCL have eased GPU programming for nongraphical applications, they are still explicitly parallel languages. All parallel programmers, particularly the novices, need tools that can help ensuring the correctness of their programs. Like any multithreaded environment, data races on GPUs can severely affect the program reliability. In this paper, we propose GMRace, a new mechanism for detecting races in GPU programs. GMRace combines static analysis with a carefully designed dynamic checker for logging and analyzing information at runtime. Our design utilizes GPUs memory hierarchy to log runtime data accesses efficiently. To improve the performance, GMRace leverages static analysis to reduce the number of statements that need to be instrumented. Additionally, by exploiting the knowledge of thread scheduling and the execution model in the underlying GPUs, GMRace can accurately detect data races with no false positives reported. Our experimental results show that comparing to previous approaches, GMRace is more effective in detecting races in the evaluated cases, and incurs much less runtime and space overhead. Mai Zheng, Vignesh T. Ravi, Gagan Agrawal |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2013 | Understanding the robustness of SSDS under power fault
Mai Zheng, Joseph A. Tucek, Mark Lillibridge |
FAST | 1 |
| 2013 | LiU: Hiding Disk Access Latency for HPC Applications with a New SSD-Enabled Data LayoutabstractUnlike in the consumer electronics and personal computing areas, in the HPC environment hard disks can hardly be replaced by SSDs. The reasons include hard disk's large capacity, very low price, and decent peak throughput. However, when latency dominates the I/O performance (e.g., when accessing random data), the hard disk's performance can be compromised. If the issue of high latency could be effectively solved, the HPC community would enjoy a large, affordable and fast storage without having to replace disks completely with expensive SSDs. In this paper, we propose an almost latency-free hard-disk dominated storage system called LiU for HPC. The key technique is leveraging limited amount of SSD storage for its low-latency access, and changing data layout in a hybrid storage hierarchy with low-latency SSD at the top and high-latency hard disk at the bottom. If a segment of data would be randomly accessed, we lift its top part (the head) up in the hierarchy to the SSD and leave the remaining part (the body) untouched on the disk. As a result, the latency of accessing this whole segment can be removed because access latency of the body can be hidden by the access time of the head on the SSD. Combined with the effect of prefetching a large segment, LiU (Lift it Up) can effectively remove disk access latency so disk's high peak throughput can now be fully exploited for data-intensive HPC applications. We have implemented a prototype of LiU in the PVFS parallel file system and evaluated it with representative MPI-IO micro benchmarks, including mpi-io-test, mpi-tile-io, and ior-mpi-io, and one macro-benchmark BTIO. Our experimental results show that LiU can effectively improve the I/O performance for HPC applications, with the throughput improvement ratio up to 5.8. Furthermore, LiU can bring much more benefits to sequential-I/O MPI applications when the applications are interfered by other workloads. For example, LiU improves the I/O throughput of mpi-io-test, which is under interference, by 1.1-3.4 times, while improving the same workload without interference by 15%. Dachuan Huang, Xuechen Zhang 0001, Wei Shi 0001, Mai Zheng, Song Jiang 0001 |
MASCOTS | 4 |
| 2012 | GMProf: A low-overhead, fine-grained profiling approach for GPU programsabstractDriven by the cost-effectiveness and the power-efficiency, GPUs are being increasingly used to accelerate computations in many domains. However, developing highly efficient GPU implementations requires a lot of expertise and effort. Thus, tool support for tuning GPU programs is urgently needed, and more specifically, low-overhead mechanisms for collecting fine-grained runtime information are critically required. Unfortunately, profiling tools and mechanisms available today either collect very coarse-grained information, or have prohibitive overheads. This paper presents a low-overhead and fine-grained profiling technique developed specifically for GPUs, which we refer to as GMProf. GMProf uses two ideas to help reduce the overheads of collecting fine-grained information. The first idea involves exploiting a number of GPU architectural features to collect reasonably accurate information very efficiently, and the second idea is to use simple static analysis methods to reduce the overhead of runtime profiling. The specific implementation of GMProf we report in this paper focuses on shared memory usage. Particularly, we help programmers understand (1) which locations in shared memory are infrequently accessed? and (2) which data elements in device memory are frequently accessed? We have evaluated GMProf using six popular GPU kernels with different characteristics. Our experimental results show that GMProf, with all optimizations, incurs a moderate overhead, e.g., 1.36 times on average for shared memory profiling. Furthermore, for three of the six evaluated kernels, GMProf verified that shared memory is effectively used, and for the remaining three kernels, it not only helped accurately identify the inefficient use of shared memory, but also helped tune the implementations. The resulting tuned implementations had a speedup of 15.18 times on average. Mai Zheng, Vignesh T. Ravi, Wenjing Ma, Gagan Agrawal |
HiPC | 1 |
| 2011 | 2ndStrike: toward manifesting hidden concurrency typestate bugsabstractConcurrency bugs are becoming increasingly prevalent in the multi-core era. Recently, much research has focused on data races and atomicity violation bugs, which are related to low-level memory accesses. However, a large number of concurrency typestate bugs such as "invalid reads to a closed file from a different thread" are under-studied. These concurrency typestate bugs are important yet challenging to study since they are mostly relevant to high-level program semantics. Qi Gao 0004, Wenbin Zhang 0005, Zhezhe Chen, Mai Zheng |
ASPLOS | 4 |
| 2011 | GRace: a low-overhead mechanism for detecting data races in GPU programsabstractIn recent years, GPUs have emerged as an extremely cost-effective means for achieving high performance. Many application developers, including those with no prior parallel programming experience, are now trying to scale their applications using GPUs. While languages like CUDA and OpenCL have eased GPU programming for non-graphical applications, they are still explicitly parallel languages. All parallel programmers, particularly the novices, need tools that can help ensuring the correctness of their programs. Like any multithreaded environment, data races on GPUs can severely affect the program reliability. Thus, tool support for detecting race conditions can significantly benefit GPU application developers. Existing approaches for detecting data races on CPUs or GPUs have one or more of the following limitations: 1) being illsuited for handling non-lock synchronization primitives on GPUs; 2) lacking of scalability due to the state explosion problem; 3) reporting many false positives because of simplified modeling; and/or 4) incurring prohibitive runtime and space overhead. In this paper, we propose GRace, a new mechanism for detecting races in GPU programs that combines static analysis with a carefully designed dynamic checker for logging and analyzing information at runtime. Our design utilizes GPUs memory hierarchy to log runtime data accesses efficiently. To improve the performance, GRace leverages static analysis to reduce the number of statements that need to be instrumented. Additionally, by exploiting the knowledge of thread scheduling and the execution model in the underlying GPUs, GRace can accurately detect data races with no false positives reported. Based on the above idea, we have built a prototype of GRace with two schemes, i.e., GRace-stmt and GRace-addr, for NVIDIA GPUs. Both schemes are integrated with the same static analysis. We have evaluated GRace-stmt and GRace-addr with three data race bugs in three GPU kernel functions and also have compared them with the existing approach, referred to as B-tool. Our experimental results show that both schemes of GRace are effective in detecting all evaluated cases with no false positives, whereas Btool reports many false positives for one evaluated case. On the one hand, GRace-addr incurs low runtime overhead, i.e., 22-116%, and low space overhead, i.e., 9-18MB, for the evaluated kernels. On the other hand, GRace-stmt offers more help in diagnosing data races with larger overhead. Mai Zheng, Vignesh T. Ravi, Gagan Agrawal |
PPoPP | 1 |