VLDB 2026 Research / reviewers in the wild / expert
Runzhou Han
dblp:264/5688
· DBLP profile ↗
8ranked-venue papers
4as first author
7since 2021 · last 2024
0000-0003-1440-7568ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 4 first-author · 7 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Revisiting Erasure Codes: A Configuration PerspectiveabstractErasure coding (EC) plays a crucial role in the fault tolerance of modern distributed storage systems (DSS). Inspired by recent research on storage configuration, we study the configuration sensitivity of EC in real DSS in this paper. We systematically inject faults to trigger EC recovery under various configurations, and measure the impact on recovery time and storage overhead quantitatively. Our results show that configurations may affect the EC recovery time significantly (e.g., up to 426%). More interestingly, theoretically superior codes may perform worse in DSS under certain configurations. Also, there is a system checking period before EC recovery that accounts for 41% to 58% of the overall system recovery time, which has been largely ignored in previous studies. Finally, in terms of storage overhead, EC may introduce 32.3% to 72.0% more write amplification (WA) than the theoretical expectation, and we derive a formula to help estimate WA more precisely. Our work suggests the importance of considering the context of real DSS for EC research, and we hope the methodology and findings can contribute to a firmer footing for EC optimization in practice. Runzhou Han, Tabassum Mahmud, Zeren Yang, Vladislav Esaulov, Lipeng Wan 0001, Yong Chen 0001, Jim Wayda, Matthew Wolf, Mai Zheng |
HotStorage | 1 |
| 2024 | PROV-IO$^+$+: A Cross-Platform Provenance Framework for Scientific Data on HPC SystemsabstractData provenance, or data lineage, describes the life cycle of data. In scientific workflows on HPC systems, scientists often seek diverse provenance (e.g., origins of data products, usage patterns of datasets). Unfortunately, existing provenance solutions cannot address the challenges due to their incompatible provenance models and/or system implementations. In this paper, we analyze four representative scientific workflows in collaboration with the domain scientists to identify concrete provenance needs. Based on the first-hand analysis, we propose a provenance framework called PROV-IO$^+$, which includes an I/O-centric provenance model for describing scientific data and the associated I/O operations and environments precisely. Moreover, we build a prototype of PROV-IO$^+$to enable end-to-end provenance support on real HPC systems with little manual effort. The PROV-IO$^+$framework can support both containerized and non-containerized workflows on different HPC platforms with flexibility in selecting various classes of provenance. Our experiments with realistic workflows show that PROV-IO$^+$can address the provenance needs of the domain scientists effectively with reasonable performance (e.g., less than 3.5% tracking overhead for most experiments). Moreover, PROV-IO$^+$outperforms a state-of-the-art system (i.e., ProvLake) in our experiments. Runzhou Han, Mai Zheng, Surendra Byna, Houjun Tang, Bin Dong 0002, Dong Dai 0001, Yong Chen 0001, Dongkyun Kim, Joseph Hassoun, David Thorsley |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2023 | λFS: A Scalable and Elastic Distributed File System Metadata Service using Serverless FunctionsabstractThe metadata service (MDS) sits on the critical path for distributed file system (DFS) operations, and therefore it is key to the overall performance of a large-scale DFS. Common "serverful" MDS architectures, such as a single server or cluster of servers, have a significant shortcoming: either they are not scalable, or they make it difficult to achieve an optimal balance of performance, resource utilization, and cost. A modern MDS requires a novel architecture that addresses this shortcoming. Benjamin Carver, Runzhou Han, Mai Zheng, Yue Cheng 0001 |
ASPLOS (4) | 2 |
| 2022 | PROV-IO: An I/O-Centric Provenance Framework for Scientific Data on HPC SystemsabstractcData provenance, or data lineage, describes the life cycle of data. In scientific workflows on HPC systems, scientists often seek diverse provenance (e.g., origins of data products, usage patterns of datasets). Unfortunately, existing provenance solutions cannot address the challenges due to their incompatible provenance models and/or system implementations. Runzhou Han, Surendra Byna, Houjun Tang, Bin Dong 0002, Mai Zheng |
HPDC | 1 |
| 2022 | On the Reproducibility of Bugs in File-System Aware Storage ApplicationsabstractMany storage applications such as file system checkers, defragmentation tools, etc. require a detailed understanding of file systems. Such file-system aware applications play an essential role today, but unfortunately they are error-prone. To better understand the challenges as well as the opportunities to address the issues, this paper presents an empirical study of real world bugs in file-system aware storage applications. By analyzing 59 bug cases from 4 representative applications in depth, we derive multiple insights in terms of general bug patterns, triggering conditions, and implications for building effective tools to address the issues. We hope that our study and the resulting dataset could contribute to the development of reliability tools for building robust file-system aware storage applications in general. Tabassum Mahmud, Om Rameshwar Gatla, Runzhou Han, Yong Chen 0001, Mai Zheng |
NAS | 4 |
| 2022 | A Study of Failure Recovery and Logging of High-Performance Parallel File SystemsabstractLarge-scale parallel file systems (PFSs) play an essential role in high-performance computing (HPC). However, despite their importance, their reliability is much less studied or understood compared with that of local storage systems or cloud storage systems. Recent failure incidents at real HPC centers have exposed the latent defects in PFS clusters as well as the urgent need for a systematic analysis. To address the challenge, we perform a study of the failure recovery and logging mechanisms of PFSs in this article. First, to trigger the failure recovery and logging operations of the target PFS, we introduce a black-box fault injection tool called PFault , which is transparent to PFSs and easy to deploy in practice. PFault emulates the failure state of individual storage nodes in the PFS based on a set of pre-defined fault models and enables examining the PFS behavior under fault systematically. Next, we apply PFault to study two widely used PFSs: Lustre and BeeGFS. Our analysis reveals the unique failure recovery and logging patterns of the target PFSs and identifies multiple cases where the PFSs are imperfect in terms of failure handling. For example, Lustre includes a recovery component called LFSCK to detect and fix PFS-level inconsistencies, but we find that LFSCK itself may hang or trigger kernel panics when scanning a corrupted Lustre. Even after the recovery attempt of LFSCK, the subsequent workloads applied to Lustre may still behave abnormally (e.g., hang or report I/O errors). Similar issues have also been observed in BeeGFS and its recovery component BeeGFS-FSCK. We analyze the root causes of the abnormal symptoms observed in depth, which has led to a new patch set to be merged into the coming Lustre release. In addition, we characterize the extensive logs generated in the experiments in detail and identify the unique patterns and limitations of PFSs in terms of failure logging. We hope this study and the resulting tool and dataset can facilitate follow-up research in the communities and help improve PFSs for reliable high-performance computing. Runzhou Han, Om Rameshwar Gatla, Mai Zheng, Jinrui Cao, Di Zhang 0015, Dong Dai 0001, Yong Chen 0001, Jonathan E. Cook 0001 |
ACM Trans. Storage | 1 |
| 2021 | SentiLog: Anomaly Detecting on Parallel File Systems via Log-based Sentiment AnalysisabstractAs core components of High-performance computing (HPC) platforms, parallel file systems (PFSes) grow quickly in scale and complexity, hence are subject to various failures and anomalies. Identifying their anomalies in runtime is critically helpful for HPC operators and administrators. Analyzing the runtime logs to detect the anomalies of large-scale systems has been proven effective in many recent studies. However, applying them to parallel file systems logs faces significant challenges due to the large volume and irregularity of PFSes logs. This study proposes SentiLog, a new approach to analyzing PFSes system logs for detecting anomalies. Unlike existing solutions, SentiLog works by training a general sentimental, natural language model based on the logging-relevant source code collected from a set of PFSes. In this way, SentiLog learns information embedded by developers from the source code. Our preliminary results show SentiLog is able to accurately predict anomalies and performs better than state-of-the-art log analysis solutions on two representative PFSes (Lustre and BeeGFS). This preliminary study shows sentiment analysis could be a promising method to analyze complex and irregular system logs. Di Zhang 0015, Dong Dai 0001, Runzhou Han, Mai Zheng |
HotStorage | 3 |
| 2020 | Position: On Failure Diagnosis of the Storage Stack
Om Rameshwar Gatla, Runzhou Han, Mai Zheng |
HotStorage | 3 |