EDBT 2026 Demo / reviewers in the wild / expert
Di Zhang 0015
dblp:80/3482-15
· DBLP profile ↗
8ranked-venue papers
5as first author
7since 2021 · last 2024
0009-0005-3115-0276ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 5 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Cross-System Analysis of Job Characterization and Scheduling in Large-Scale Computing ClustersabstractAmid the growing prevalence of artificial intelligence (AI) and deep learning (DL) across industries and science disciplines, high-performance computing (HPC) clusters are increasingly used for DL tasks, in addition to their traditional role in numerical simulations. This shared use of HPC systems for both DL tasks and numerical codes is altering the characteristics of their workloads, leaving many previously observed and well-accepted workload characteristics unchecked, potentially outdated and imprecise, for these new mixed workloads. Thus, to understand these changes and their implications for job scheduling, we conduct a cross-system analysis of job characterization and its scheduling across a range of representative clusters, including two classic HPC clusters (Mira, Theta), two classic DL clusters (Philly, Helios), and a hybrid cluster (Blue Waters).Our cross-system analyses focus on three key aspects: 1) job geometries (job size, run time, arrival interval) and their impacts on scheduling results, 2) job failure patterns and their correlations to job geometries, 3) per-user behaviors and their indication on job scheduling. Through these comparisons, we confirm notable disparity and similarity among different systems (summarized as 8 takeaways), which would help design more efficient job schedulers for the future HPC systems. We further introduce two use case studies (job runtime prediction and adaptive relaxed backfilling) that leverage these new observations and show the improved job scheduling results. In summary, we expect our observations, insights, and systematic analysis approach can be useful for the community in building efficient HPC scheduling for the upcoming hybrid workloads. Di Zhang 0015, Monish Soundar Raj, Sheng Di, Dong Dai 0001 |
IPDPS | 1 |
| 2024 | An Empirical Study of Machine Learning-Based Synthetic Job Trace Generation Methods
Monish Soundar Raj, Thomas MacDougall, Di Zhang 0015, Dong Dai 0001 |
JSSPP | 3 |
| 2023 | Early Exploration of Using ChatGPT for Log-based Anomaly Detection on Parallel File Systems LogsabstractLog-based anomaly detection has been extensively studied to help detect complex runtime anomalies in production systems. However, existing techniques exhibit several common issues. First, they rely heavily on expert-labeled logs to discern anomalous behavior patterns. But labelling enough log data manually to effectively train deep neural networks may take too long. Second, they rely on numeric model prediction based on numeric vector input which causes model decisions to be largely non-interpretable by humans which further rules out targeted error correction. Chris Egersdoerfer, Di Zhang 0015, Dong Dai 0001 |
HPDC | 2 |
| 2023 | Drill: Log-based Anomaly Detection for Large-scale Storage Systems Using Source Code AnalysisabstractLarge-scale storage systems, a critical part of modern computing systems, are subject to various runtime bugs, failures, and anomalies in production. Identifying their anomalies at runtime is thus critical for users and administrators. Since runtime logs record the important status of the systems, log-based anomaly detection has been studied extensively for timely identifying system malfunctions. However, existing log-based anomaly detection solutions share common limitations in representing log entries accurately and robustly, hence can not effectively handle log entries that were not seen in the historical logs, which is a common real-world scenario due to logs' inherent rarity and the continuous evolution of the systems. To address the issues of existing methods, we propose Drill, a new log pre-processing method to generate high-quality vector representation of runtime logs by leveraging both storage system-specific sentiment-classifying language models and log contexts built from the source code. Through extensive evaluations of two representative distributed storage systems (Apache HDFS and Lustre), we show that Drill can achieve up to 41% improvement when compared with state-of-the-art anomaly detection solutions, showing it is a promising solution for general anomaly detection. Di Zhang 0015, Chris Egersdoerfer, Tabassum Mahmud, Mai Zheng, Dong Dai 0001 |
IPDPS | 1 |
| 2022 | SchedInspector: A Batch Job Scheduling Inspector Using Reinforcement LearningabstractImproving the performance of job executions is an important goal of HPC batch job schedulers, such as minimizing job waiting time, slowdown, or completion time. Such a goal is often accomplished using carefully designed heuristics based on job features, such as job size and job duration. However, these heuristics overlook important runtime factors (e.g., cluster availability and waiting job patterns), which may vary across time and make a previously sound scheduling decision not hold any longer. In this study, we propose a new approach to incorporate runtime factors into batch job scheduling for better job execution performance. The key idea is to add a scheduling inspector on top of the base job scheduler to scrutinize its scheduling decisions. The inspector will take the runtime factors into consideration and accordingly determine the fitness of the scheduled job. It then either accepts the scheduled job or rejects it and asks the base schedulers to try again later. We realize such an inspector, namely SchedInspector, by leveraging the intelligence of reinforcement learning. Through extensive experiments, we show SchedInspector can intelligently integrate the runtime factors into various batch job scheduling policies, including the state-of-the-art one, to gain better job execution performance, such as smaller average bounded job slowdown (up to 69% better) or average job waiting time (up to 52% better), across various real-world workloads. We also show that although rejecting scheduling decisions may leave the resources idle hence affect the system utilization, SchedInspector is able to achieve the job execution performance improvement with marginal impact on the system utilization (typically less than 1%). We consider one key advantage of SchedInspector is it automatically learns to work with and improve existing job scheduling policies without changing them, which makes it promising to serve as a generic enhancer for various batch job scheduling policies. Di Zhang 0015, Dong Dai 0001 |
HPDC | 1 |
| 2022 | A Study of Failure Recovery and Logging of High-Performance Parallel File SystemsabstractLarge-scale parallel file systems (PFSs) play an essential role in high-performance computing (HPC). However, despite their importance, their reliability is much less studied or understood compared with that of local storage systems or cloud storage systems. Recent failure incidents at real HPC centers have exposed the latent defects in PFS clusters as well as the urgent need for a systematic analysis. To address the challenge, we perform a study of the failure recovery and logging mechanisms of PFSs in this article. First, to trigger the failure recovery and logging operations of the target PFS, we introduce a black-box fault injection tool called PFault , which is transparent to PFSs and easy to deploy in practice. PFault emulates the failure state of individual storage nodes in the PFS based on a set of pre-defined fault models and enables examining the PFS behavior under fault systematically. Next, we apply PFault to study two widely used PFSs: Lustre and BeeGFS. Our analysis reveals the unique failure recovery and logging patterns of the target PFSs and identifies multiple cases where the PFSs are imperfect in terms of failure handling. For example, Lustre includes a recovery component called LFSCK to detect and fix PFS-level inconsistencies, but we find that LFSCK itself may hang or trigger kernel panics when scanning a corrupted Lustre. Even after the recovery attempt of LFSCK, the subsequent workloads applied to Lustre may still behave abnormally (e.g., hang or report I/O errors). Similar issues have also been observed in BeeGFS and its recovery component BeeGFS-FSCK. We analyze the root causes of the abnormal symptoms observed in depth, which has led to a new patch set to be merged into the coming Lustre release. In addition, we characterize the extensive logs generated in the experiments in detail and identify the unique patterns and limitations of PFSs in terms of failure logging. We hope this study and the resulting tool and dataset can facilitate follow-up research in the communities and help improve PFSs for reliable high-performance computing. Runzhou Han, Om Rameshwar Gatla, Mai Zheng, Jinrui Cao, Di Zhang 0015, Dong Dai 0001, Yong Chen 0001, Jonathan E. Cook 0001 |
ACM Trans. Storage | 5 |
| 2021 | SentiLog: Anomaly Detecting on Parallel File Systems via Log-based Sentiment AnalysisabstractAs core components of High-performance computing (HPC) platforms, parallel file systems (PFSes) grow quickly in scale and complexity, hence are subject to various failures and anomalies. Identifying their anomalies in runtime is critically helpful for HPC operators and administrators. Analyzing the runtime logs to detect the anomalies of large-scale systems has been proven effective in many recent studies. However, applying them to parallel file systems logs faces significant challenges due to the large volume and irregularity of PFSes logs. This study proposes SentiLog, a new approach to analyzing PFSes system logs for detecting anomalies. Unlike existing solutions, SentiLog works by training a general sentimental, natural language model based on the logging-relevant source code collected from a set of PFSes. In this way, SentiLog learns information embedded by developers from the source code. Our preliminary results show SentiLog is able to accurately predict anomalies and performs better than state-of-the-art log analysis solutions on two representative PFSes (Lustre and BeeGFS). This preliminary study shows sentiment analysis could be a promising method to analyze complex and irregular system logs. Di Zhang 0015, Dong Dai 0001, Runzhou Han, Mai Zheng |
HotStorage | 1 |
| 2020 | RLScheduler: an automated HPC batch job scheduler using reinforcement learningabstractToday's high-performance computing (HPC) platforms are still dominated by batch jobs. Accordingly, effective batch job scheduling is crucial to obtain high system efficiency. Existing HPC batch job schedulers typically leverage heuristic priority functions to prioritize and schedule jobs. But, once configured and deployed by the experts, such priority functions can hardly adapt to the changes of job loads, optimization goals, or system settings, potentially leading to degraded system efficiency when changes occur. To address this fundamental issue, we present RLScheduler, an automated HPC batch job scheduler built on reinforcement learning. RLScheduler relies on minimal manual interventions or expert knowledge, but can learn high-quality scheduling policies via its own continuous `trial and error'. We introduce a new kernel-based neural network structure and trajectory filtering mechanism in RLScheduler to improve and stabilize the learning process. Through extensive evaluations, we confirm that RLScheduler can learn high-quality scheduling policies towards various workloads and various optimization goals with relatively low computation cost. Moreover, we show that the learned models perform stably even when applied to unseen workloads, making them practical for production use. Di Zhang 0015, Dong Dai 0001, Youbiao He, Forrest Sheng Bao |
SC | 1 |