Qizhen Weng 0001

dblp:222/1610-1 · DBLP profile ↗
← Back
13ranked-venue papers
2as first author
9since 2021 · last 2026
0000-0001-9195-6443ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 1 first-author · 6 since 2021Computer networks · 3 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2026 Towards Efficient Serving of Network-intensive LLM Inferences
abstract
Prefix caching has become a key technique for LLM serving, and nowadays the reusable KVCache contents are often hosted on distributed servers. For long-context LLM inferences with high cache hit ratio, cross-server KVCache transmission has become an emerging performance bottleneck; such network-intensive LLM inferences are increasingly prevalent in the coming era of agentic AI. However, existing LLM inference engines are essentially compute-centric; we find that they are highly inefficient when serving such workloads due to compute-stage service blocking and ignorance of KVCache-transfer cost.
Chen Chen 0067, Junxue Zhang 0001, Zhusheng Wang, Zixuan Guan, Qizhen Weng 0001, Minyi Guo
APNet8
2026 Suika: Efficient and High-quality Re-scheduling of 3D-parallelized LLM Training Jobs in Shared Clusters
abstract
Large Language Models (LLMs) are usually trained with 3D (data, tensor, and pipeline) parallelism—in shared GPU clusters where the available resources are highly dynamic. Rescheduling the idle resources to ongoing jobs can help improve cluster utilization, but doing so for 3D-parallelized training jobs suffers large overheads in performance modeling, decision making, and redeployment. We present Suika, a cluster training system that supports efficient and high-quality resource rescheduling for 3D-parallelized LLM training jobs. Suika holistically addresses the complexity challenges by exploiting the incremental nature of rescheduling. For performance modeling, it builds an accurate performance estimator with non-disruptive online profiling. For decision-making, it employs topology-aware sorting and an expand-and-balance algorithm to reduce the complexity of resource allocation and job parallelization, without compromising decision quality. Suika further integrates a device-to-device redeployment method to leverage the overlapping nature of incremental reconfiguration for overhead reduction. Experiments on 64-GPU physical cluster and 1024-GPU simulated cluster show that, Suika achieves 1.29 ~ 1.31× reduction in average JCT compared to state-of-the-art schedulers.
Chen Chen 0067, Chunyu Xue, Qizhen Weng 0001, Zeren Li, Xuqi Zhu, Yongqiang Yang, Quan Chen 0002, Minyi Guo
EuroSys5
2026 Efficient Data Passing for Serverless Inference Workflows: A GPU-Centric Approach
abstract
Serverless computing offers a compelling paradigm for deploying machine learning inference workflows composed of heterogeneous CPU and GPU functions. However, existing data-passing solutions in serverless systems primarily rely on host memory for data exchange (host-centric), leading to substantial data movement and salient I/O overhead. Moreover, modern GPU communication libraries (e.g., NCCL, NVSHMEM, UCX) are ill-suited to serverless environments, suffering from redundant data copies, underutilized transfer bandwidth, and inefficient temporary GPU storage.
Hao Wu 0032, Yaochen Liu, Minchen Yu, Qizhen Weng 0001, Junxiao Deng, Hao Fan 0006, Song Wu 0001, Wei Wang 0030, Hai Jin 0001
EuroSys4
2025 GPU-Disaggregated Serving for Deep Learning Recommendation Models at Scale
Lingyun Yang, Yongchen Wang, Yinghao Yu, Qizhen Weng 0001, Jianbo Dong, Chi Zhang 0005, Yanyi Zi, Zechao Zhang, Menglei Zheng, Lanlan Xi, Binzhang Fu, Tao Lan, Liping Zhang 0013, Lin Qu, Wei Wang 0030
NSDI4
2025 Toppings: CPU-Assisted, Rank-Aware Adapter Serving for LLM Inference
Suyi Li 0002, Hanfeng Lu, Tianyuan Wu, Minchen Yu, Qizhen Weng 0001, Xusheng Chen, Yizhou Shan, Binhang Yuan, Wei Wang 0030
USENIX ATC5
2023 Beware of Fragmentation: Scheduling GPU-Sharing Workloads with Fragmentation Gradient Descent
Qizhen Weng 0001, Lingyun Yang, Yinghao Yu, Wei Wang 0030, Xiaochuan Tang, Liping Zhang 0013
USENIX ATC1
2023 Accelerating Distributed Learning in Non-Dedicated Environments
abstract
Machine learning (ML) models are increasingly trained with distributed workers possessing heterogeneous resources. In such scenarios, model training efficiency may be negatively affected bystragglers—workers that run much slower than others. Efficient model training requires eliminating such stragglers, yet for modern ML workloads, existing load balancing strategies are inefficient and even infeasible. In this article, we propose a novel strategy, calledsemi-dynamic load balancing, to eliminate stragglers of distributed ML workloads. The key insight is that ML workers shall be load-balanced atiteration boundaries, being non-intrusive to intra-iteration execution. Based on it we further develop LB-BSP, an integrated worker coordination mechanism that adapts workers’ load to their instantaneous processing capabilities—by right-sizing the sample batches at the synchronization barriers. We have designed distinct load tuning algorithms for ML in CPU clusters, in GPU clusters as well as in federated learning setups, based on their respective characteristics. LB-BSP has been implemented as a Python module for ML frameworks like TensorFlow and PyTorch. Our EC2 deployment confirms that LB-BSP is practical, effective and light-weight, and is able to accelerating distributed training by up to 54 percent.
Chen Chen 0067, Qizhen Weng 0001, Wei Wang 0030, Baochun Li, Bo Li 0001
IEEE Trans. Cloud Comput.2
2022 Workload consolidation in alibaba clusters: the good, the bad, and the ugly
abstract
Web companies typically run latency-critical long-running services and resource-intensive, throughput-hungry batch jobs in a shared cluster for improved utilization and reduced cost. Despite many recent studies on workload consolidation, the production practice remains largely unknown. This paper describes our efforts to efficiently consolidate the two types of workloads in Alibaba clusters to support the company's e-commerce businesses.
Yongkang Zhang 0003, Yinghao Yu, Wei Wang 0030, Qiukai Chen, Tianchen Ding, Qizhen Weng 0001, Lingyun Yang, Jian He 0004, Liping Zhang 0013
SoCC9
2022 MLaaS in the Wild: Workload Analysis and Scheduling in Large-Scale Heterogeneous GPU Clusters
Qizhen Weng 0001, Wencong Xiao, Yinghao Yu, Wei Wang 0030, Jian He 0004, Yong Li 0045, Liping Zhang 0013, Wei Lin 0016
NSDI1
2020 Semi-dynamic load balancing: efficient distributed learning in non-dedicated environments
abstract
Machine learning (ML) models are increasingly trained in clusters with non-dedicated workers possessing heterogeneous resources. In such scenarios, model training efficiency can be negatively affected by stragglers---workers that run much slower than others. Efficient model training requires eliminating such stragglers, yet for modern ML workloads, existing load balancing strategies are inefficient and even infeasible. In this paper, we propose a novel strategy called semi-dynamic load balancing to eliminate stragglers of distributed ML workloads. The key insight is that ML workers shall be load-balanced at iteration boundaries, being non-intrusive to intra-iteration execution. We develop LB-BSP based on such an insight, which is an integrated worker coordination mechanism that adapts workers' load to their instantaneous processing capabilities by right-sizing the sample batches at the synchronization barriers. We have custom-designed the batch sizing algorithm respectively for CPU and GPU clusters based on their own characteristics. LB-BSP has been implemented as a Python module for ML frameworks like TensorFlow and PyTorch. Our EC2 deployment confirms that LB-BSP is practical, effective and light-weight, and is able to accelerating distributed training by up to 54%.
Chen Chen 0067, Qizhen Weng 0001, Wei Wang 0030, Baochun Li, Bo Li 0001
SoCC2
2020 Metis: learning to schedule long-running applications in shared container clusters at scale
abstract
Online cloud services are increasingly deployed as long-running applications (LRAs) in containers. Placing LRA containers is known to be difficult as they often have sophisticated resource interferences and I/O dependencies. Existing schedulers rely on operators to manually express the container scheduling requirements as placement constraints and strive to satisfy as many constraints as possible. Such schedulers, however, fall short in performance as placement constraints only provide qualitative scheduling guidelines and minimizing constraint violations does not necessarily result in the optimal performance.In this work, we present Metis, a general-purpose scheduler that learns to optimally place LRA containers using deep reinforcement learning (RL) techniques. This eliminates the complex manual specification of placement constraints and offers, for the first time, concrete quantitative scheduling criteria. As directly training an RL agent does not scale, we develop a novel hierarchical learning technique that decomposes a complex container placement problem into a hierarchy of subproblems with significantly reduced state and action space. We show that many subproblems have similar structures and can hence be solved by training a unified RL agent offline. Large-scale EC2 deployment shows that compared with the traditional constraint-based schedulers, Metis improves the throughput by up to 61%, optimizes various performance metrics, and easily scales to a large cluster where 3K containers run on over 700 machines.
Qizhen Weng 0001, Wei Wang 0030, Chen Chen 0067, Bo Li 0001
SC2
2018 Fast Distributed Deep Learning via Worker-adaptive Batch Sizing
abstract
In heterogeneous or shared clusters, distributed learning processes are slowed down by straggling workers. In this work, we propose LB-BSP, a new synchronization scheme that eliminates stragglers by adapting each worker's training load (batch size) to its processing capability. For training in shared production clusters, a prerequisite for deciding the workers' batch sizes is to know their processing speeds before each iteration starts. To this end, we adopt NARX, an extended recurrent neural network that accounts for both the historical speeds and the driving factors such as CPU and memory in prediction.
Chen Chen 0067, Qizhen Weng 0001, Wei Wang 0030, Baochun Li, Bo Li 0001
SoCC2
2018 OpuS: Fair and Efficient Cache Sharing for In-Memory Data Analytics
abstract
We study the fair cache allocation problem in shared cloud environments, where many users and applications contend for the main memory to cache shared datasets or files. Unlike other resources such as CPUs and networks, in-memory caches can be non-exclusively shared across many users, e.g., a cached columnar dataset queried by many Spark SQL jobs. This results in a unique challenge of the "free-riding" problem, where a user lies about its caching preferences to trick other users to cache files for it, using their allocated cache space. We show that existing cache allocation policies either suffer from such manipulations or result in poor efficiency. To address this problem, we propose a new cache allocation algorithm, termed OpuS, or Opportunistic Sharing for high efficiency. We show that OpuS provides performance isolation between users and is strategy-proof against "free-riding" manipulations. We have implemented OpuS as a pluggable cache manager in Alluxio, a popular memory-centric filesystem. Cluster deployment and trace-driven simulations demonstrate that OpuS allocates each user a fair share of caches while achieving near-optimal efficiency in cache utilization.
Yinghao Yu, Wei Wang 0030, Jun Zhang 0004, Qizhen Weng 0001, Khaled Ben Letaief
ICDCS4