EDBT 2026 Demo / reviewers in the wild / expert
Zeyu Zhang 0005
dblp:44/8352-5
· DBLP profile ↗
6ranked-venue papers
2as first author
6since 2021 · last 2026
0009-0005-7853-6854ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021Computer networks · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PEACE: Preemptive and Efficient Cluster Scheduling for LLM Inference with Mixed Prompts
Zeyu Zhang 0005, Haiying Shen |
IPDPS | 1 |
| 2025 | Deep Learning Training Job Scheduling for Proactive Straggler ReductionabstractIn this work, from our trace-driven experimental measurements, we observed that despite employing homogeneous GPUs for distributed deep learning (DDL) training, stragglers persistently emerge, significantly prolonging time-to-accuracy (TTA) and squandering GPU resources. Previous approaches typically react to stragglers as they occur or are imminent during execution after scheduling, resulting in training delays before they are addressed and introducing additional overhead. To reduce the number of stragglers, this paper introduces a novel DDL training job scheduler for proactive straggler reduction (STRN), the first effort in mitigating stragglers during scheduling. STRN is devised based on our findings that various DDL jobs exhibit distinct sensitivities to straggling and a particular resource's overload, and the optimal synchronization strategy for a job hinges on its specific characteristics and operational environment. Thus, STRN assesses each job's sensitivity to each resource type and straggling. It aims to minimize the likelihood of resource overload for jobs with higher sensitivity to that resource and to ensure that jobs with higher sensitivity to straggling encounter less straggling. STRN initially runs a heuristic method and then transitions to a Reinforcement Learning (RL)-based method once trained, facilitating expedited and optimal job scheduling. STRN also optimizes synchronization strategies for each DDL job to reduce overall TTA. Trace-driven real experiments demonstrate that STRN reduces average TTA by up to 59% and enhances average accuracy by up to 91% compared to state-of-the-art methods. We have made the source code available for distribution. Haiying Shen, Zeyu Zhang 0005 |
CCGrid | 2 |
| 2025 | ModServe: Modality- and Stage-Aware Resource Disaggregation for Scalable Multimodal Model ServingabstractLarge multimodal models (LMMs) demonstrate impressive capabilities in understanding images, videos, and audio beyond text. However, efficiently serving LMMs in production environments poses significant challenges due to their complex model architectures and heterogeneous characteristics across their multi-stage inference pipelines and modalities. Haoran Qiu, Anish Biswas, Jayashree Mohan, Alind Khare, Esha Choukse, Íñigo Goiri, Zeyu Zhang 0005, Haiying Shen, Chetan Bansal, Ramachandran Ramjee, Rodrigo Fonseca |
SoCC | 8 |
| 2025 | Revisiting the Straggling Problem in GPU-based Distributed Deep Learning TrainingabstractThe straggler problem has been extensively studied in CPU-based distributed deep learning (DL) training but has not received significant attention in homogeneous GPU-based distributed training, possibly because GPUs do not typically become bottlenecks in this scenario. In this paper, we conduct experiment measurements and find that the straggler problems persist in this scenario, primarily stemming from communication hurdles, compounded by computation delays, and stragglers substantially inflate resource consumption and training time by ∼50%. Existing straggler mitigation methods do not directly address the communication stragglers in this scenario, and they suffer from drawbacks such as prolonged latency in straggler removal, high resource consumption, or compromised training accuracy. To tackle these limitations, based on the insights derived from thorough measurements, we propose a Straggler-aware Time and Resource Efficient distributed DL Training system (STRET). STRET is tailored for both homogeneous and heterogeneous GPU-based distributed training, encompassing both the parameter server (PS) and all-reduce architectures. It creates a hybrid architecture that connects a straggler to a non-straggler possessing high communication bandwidth with it to reduce communication delay. If this method fails to eliminate stragglers, it runs two complementary methods in sequence to remove the stragglers. First, it further reduces communication overhead by withholding reporting gradients when the accuracy increment is marginal. Second, it conducts one-time batch size tuning to reduce iteration time. Real experimental results on TensorFlow show that STRET can reduce up to 56% and 41% training time and save up to 94% and 96% resources in the heterogeneous and homogeneous scenarios, respectively, compared to state-of-the-art approaches while preserving accuracy. Suraiya Tairin, Zeyu Zhang 0005, Haiying Shen |
ICCCN | 2 |
| 2025 | HACK: Homomorphic Acceleration via Compression of the Key-Value Cache for Disaggregated LLM InferenceabstractDisaggregated Large Language Model (LLM) inference decouples the compute-intensive prefill stage from the memory-intensive decode stage, allowing low-end, compute-focused GPUs for prefill and high-end, memory-rich GPUs for decode, which reduces cost while maintaining high throughput. However, transmitting Key-Value (KV) data between the two stages can be a bottleneck, especially for long prompts. Additionally, the computational overhead in the two stages is key for optimizing Job Completion Time (JCT), and KV data size can become prohibitive for long prompts and sequences. Existing KV quantization methods can alleviate transmission and memory bottlenecks, but they introduce significant dequantization overhead, exacerbating the computation time. Zeyu Zhang 0005, Haiying Shen, Shay Vargaftik, Ran Ben-Basat, Michael Mitzenmacher, Minlan Yu |
SIGCOMM | 1 |
| 2023 | Embracing Uncertainty for Equity in Resource Allocation in ML TrainingabstractTo reduce the Deep Learning (DL) model training time and hence resource consumption, it is critical to avoid stragglers. However, the dynamics and uncertainty features of resource availability pose a challenge to avoiding stragglers caused. To handle this challenge, we propose a Straggler-Avoiding job Scheduling approach (SAS), which smartly ensures that the tasks of a job receive resources with similar dynamics and uncertainty so that the tasks can complete at approximately the same time. Specifically, SAS uses an ML method to predict available resource amounts with probability in future times, groups nodes with similar available resource amounts and probabilities, and then assigns each job to one node group with the objective of minimizing job completion time (JCT). To reduce the decision making time, we also propose a reinforcement learning (RL) based scheduling approach (SAS-RL) that assigns each job to a node group. In addition, we propose a distributed parameter server (PS) load reassignment method to handle PS stragglers. Our trace-driven real experiments show that SAS reduce up to 45% JCT and 63% stragglers compared with existing job schedulers, and our PS load reassignment reduces up to 48% JCT compared with the previous PS load distribution scheme. Suraiya Tairin, Haiying Shen, Zeyu Zhang 0005 |
ICPP | 3 |