EDBT 2026 Demo / reviewers in the wild / expert
Zhisheng Ye 0002
dblp:12/8051-2
· DBLP profile ↗
9ranked-venue papers
1as first author
9since 2021 · last 2026
0000-0002-4568-0244ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 1 first-author · 7 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FlowGPU: Transparent and Efficient GPU Checkpointing and Restore
Zehua Yang, Yonghao Zou, Junyang Zhang 0003, Zhisheng Ye 0002, Xiaolin Wang 0001, Yingwei Luo, Zhenlin Wang 0003, Diyu Zhou |
Euro-Par (2) | 5 |
| 2026 | ResiHP: Taming LLM Training Failures with Dynamic Hybrid ParallelismabstractHybrid parallelism underpins large-scale LLM training across tens of thousands of GPUs. At such scale, hardware failures on individual devices lead to performance skew across devices, diminishing overall training efficiency. Existing resilient systems overlook sequence length variability in datasets and device performance skew under hybrid parallelism. As a result, (1) iteration time fluctuations induced by sequence length variability can trigger spurious fail-slow detections, and (2) failures are mitigated through individual adaptations in hybrid parallelism, leading to unnecessary detection overhead and inefficient resilient training. Tenghui Ma, Jihu Guo, Wei Gao 0064, Sitian Lu, Zhisheng Ye 0002, Hanjing Wang, Dahua Lin |
HPDC | 5 |
| 2026 | Latency-SLO-Aware Memory Offloading for Large Language Model InferenceabstractOffloading large language models (LLMs) states to host memory during inference promises to reduce operational costs by supporting larger models, longer prompts, and larger batch sizes. However, the design of existing memory offloading mechanisms does not take latency service-level objectives (SLOs) into consideration. As a result, they either lead to frequent SLO violations or underutilize host memory, thereby incurring economic loss and thus defeating the purpose of memory offloading. Chenxiang Ma, Zhisheng Ye 0002, Zehua Yang, Tianhao Fu, Jiaxun Han, Jie Zhang 0048, Yingwei Luo, Xiaolin Wang 0001, Zhenlin Wang 0003, Yong Li 0045, Diyu Zhou |
ICS | 3 |
| 2024 | Characterization of Large Language Model Development in the Datacenter
Qinghao Hu 0004, Zhisheng Ye 0002, Zerui Wang, Guoteng Wang, Meng Zhang 0045, Qiaoling Chen, Peng Sun 0006, Dahua Lin, Xiaolin Wang 0001, Yingwei Luo, Yonggang Wen 0001, Tianwei Zhang 0004 |
NSDI | 2 |
| 2024 | UniSched: A Unified Scheduler for Deep Learning Training Jobs With Different User DemandsabstractThe growth of deep learning training (DLT) jobs in modern GPU clusters calls for efficient deep learning (DL) scheduler designs. Due to the extensive applications of DL technology, developers may have different demands for their DLT jobs. It is important for a GPU cluster to support all these demands and efficiently execute those DLT jobs. Unfortunately, existing DL schedulers mainly focus on part of those demands, and cannot provide comprehensive scheduling services.In this work, we present UniSched, a unified scheduler to optimize different types of scheduling objectives (e.g., guaranteeing the deadlines of SLO jobs, minimizing the latency of best-effort jobs). Meanwhile, UniSchedsupports different job stopping criteria (e.g., iteration-based, performance-based). UniSched includes two key components: Estimator for estimating the job duration, and Selector for selecting jobs and allocating resources. We perform large-scale simulations over the job traces from the production clusters. Compared to state-of-the-art schedulers, UniSchedcan significantly decrease the deadline miss rate of SLO jobs by up to 6.84×, and the latency of best-effort jobs by up to 4.02×, To demonstrate the practicality of UniSched, we implement and deploy a prototype on Kubernetes in a physical cluster consisting of 64 GPUs. Wei Gao 0064, Zhisheng Ye 0002, Peng Sun 0006, Tianwei Zhang 0004, Yonggang Wen 0001 |
IEEE Trans. Computers | 2 |
| 2023 | Hydro: Surrogate-Based Hyperparameter Tuning Service in Datacenters
Qinghao Hu 0004, Zhisheng Ye 0002, Meng Zhang 0045, Qiaoling Chen, Peng Sun 0006, Yonggang Wen 0001, Tianwei Zhang 0004 |
OSDI | 2 |
| 2022 | Tear Up the Bubble Boom: Lessons Learned From a Deep Learning Research and Development ClusterabstractWith the proliferation of deep learning, there exists a strong need to efficiently operate GPU clusters for deep learning production in giant AI companies, as well as for research and development (R&D) in small-sized research institutes and universities. Existing works have performed thorough trace analysis on large-scale production-level clusters in giant companies, which discloses the characteristics of deep learning production jobs and motivates the design of scheduling frameworks. However, R&D clusters significantly differ from production-level clusters in both job properties and user behaviors, calling for a different scheduling mechanism. In this paper, we present a detailed workload characterization of an R&D cluster, CloudBrain-I, in a research institute, Peng Cheng Laboratory. After analyzing the fine-grained resource utilization, we discover a severe problem for R&D clusters, resource underutilization, which is especially important in R&D clusters while not characterised by existing works. We further investigate two specific underutilization phenomena and conclude several implications and lessons on R&D cluster scheduling. The traces will be open-sourced to motivate further studies in the community. Zehua Yang, Zhisheng Ye 0002, Tianhao Fu, Yingwei Luo, Xiaolin Wang 0001, Zhenlin Wang 0003, Tianwei Zhang 0004 |
ICCD | 2 |
| 2022 | Astraea: A Fair Deep Learning Scheduler for Multi-Tenant GPU ClustersabstractModern GPU clusters are designed to support distributed Deep Learning jobs from multiple tenants concurrently. Each tenant may have varied and dynamic resource demands. Unfortunately, existing GPU schedulers fail to thoroughly consider the fairness among the tenants and jobs, which can result in unbalanced resource allocation and unfair user experience. In this article, we present an efficient solution to provide strong fairness while maintaining high scheduling effectiveness in multi-tenant GPU clusters. First, we introduce a novel Long-Term GPU-time Fairness metric, which can comprehensively evaluate the fairness at both the tenant and job levels, based on both the temporal and spatial impacts of resource allocation. Second, we design a new and practical GPU scheduler,Astraea, to enforce the desired fairness among tenants and jobs. Large-scale evaluations show thatAstraeacan improve tenant fairness by up to 9.42× compared to state-of-the-art schedulers, without sacrificing the average job completion time. Zhisheng Ye 0002, Peng Sun 0006, Wei Gao 0064, Tianwei Zhang 0004, Xiaolin Wang 0001, Shengen Yan, Yingwei Luo |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2021 | Chronus: A Novel Deadline-aware Scheduler for Deep Learning Training JobsabstractModern GPU clusters support Deep Learning training (DLT) jobs in a distributed manner. Job scheduling is the key to improve the training performance, resource utilization and fairness across users. Different training jobs may require various objectives and demands in terms of completion time. How to efficiently satisfy all these requirements is not extensively studied. Wei Gao 0064, Zhisheng Ye 0002, Peng Sun 0006, Yonggang Wen 0001, Tianwei Zhang 0004 |
SoCC | 2 |