Zizhao Mo

dblp:356/8799 · DBLP profile ↗
← Back
6ranked-venue papers
4as first author
6since 2021 · last 2025
0000-0002-3590-4400ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 3 first-author · 5 since 2021Software engineering, systems software and programming languages · 3 · 2 first-author · 3 since 2021
YearPublicationVenuePosition
2025 Fast and Fair Training for Deep Learning in Heterogeneous GPU Clusters
abstract
The GPU device heterogeneity in accelerating deep learning training workloads poses significant challenges for job scheduling in datacenters.Existing heterogeneity-aware job schedulers, however, cannot effectively reduce the overall job completion time (JCT) or provide fairness guarantees due to their coarse-grained resource allocation and poor integration of conflicting objectives.This paper presents FFT, a novel scheduling system designed for Fast and Fair deep learning Training in heterogeneous GPU clusters.FFT incorporates two key designs.First, it incorporates a resource allocation scheme in each round to enable fine-grained control over resource utilization.Second, it seamlessly integrates a fairness compensation mechanism that dynamically evaluates fairness in real-time.Building upon these designs, FFT formulates a cost minimization problem to determine the optimal schedule, striking a delicate balance between efficiency and fairness.Extensive experiments conducted in physical clusters as well as large-scale testbed demonstrate that FFT can significantly accelerate the overall JCT by up to 5.2× while improving job finishtime-fairness by more than 2.2× compared to state-of-the-art heterogeneity-aware solutions.
Zizhao Mo, Huanle Xu, Wing Cheong Lau
ICS1
2025 Hetis: Serving LLMs in Heterogeneous GPU Clusters with Fine-grained and Dynamic Parallelism
abstract
The significant resource demands in LLM serving prompts production clusters to fully utilize heterogeneous hardware by partitioning LLM models across a mix of high-end and low-end GPUs. However, existing parallelization approaches often struggle to scale efficiently in heterogeneous environments due to their coarse-grained and static parallelization strategies.
Zizhao Mo, Jianxiong Liao, Huanle Xu, Zhi Zhou 0006, Cheng-Zhong Xu 0001
SC1
2024 Heet: Accelerating Elastic Training in Heterogeneous Deep Learning Clusters
abstract
Modern GPU clusters inherently exhibit heterogeneity, encompassing various aspects such as computation and communication. This heterogeneity poses a significant challenge for the elastic scheduling of deep learning workloads. Unfortunately, existing elastic schedulers often overlook the impact of heterogeneity on scaling efficiency, resulting in considerably prolonged job completion times.
Zizhao Mo, Huanle Xu, Cheng-Zhong Xu 0001
ASPLOS (2)1
2024 Derm: SLA-aware Resource Management for Highly Dynamic Microservices
abstract
Ensuring efficient resource allocation while providing service level agreement (SLA) guarantees for end-to-end (E2E) latency is crucial for microservice applications. Although existing studies have made significant contributions towards achieving this objective, they primarily concentrate on static graphs. However, microservice graphs are inherently dynamic during runtime in production environments, necessitating more effective and scalable resource management solutions.In this paper, we present Derm, a new resource management system designed for microservice applications with highly dynamic graphs. Our principal finding is that prioritizing different microservice graphs can lead to a substantial reduction in resource allocation. To take advantage of this opportunity, we develop three main components. The first is a performance model that describes uncertainties of microservice latency through a conditional exponential distribution. The second is a probabilistic quantification of the dynamics of microservice graphs. The third is an optimization method for adjusting the resource allocation of microservices to minimize resource usage. We evaluate Derm in our cluster using real microservice benchmarks and production traces. The results highlight that Derm reduces the resource usage by $68.4 \%$ and lowers SLA violation probability by $6.7 \times$, compared to existing approaches.
Liao Chen 0001, Shutian Luo, Chenyu Lin, Zizhao Mo, Huanle Xu, Kejiang Ye, Cheng-Zhong Xu 0001
ISCA4
2024 Optimal Resource Efficiency with Fairness in Heterogeneous GPU Clusters
abstract
Ensuring the highest training throughput to maximize resource efficiency, while maintaining fairness among users, is critical for deep learning (DL) training in heterogeneous GPU clusters. However, current DL schedulers provide only limited fairness properties and suboptimal training throughput, impeding tenants from effectively leveraging heterogeneous resources. The underlying design challenge stems from inherent conflicts between efficiency and fairness properties.
Zizhao Mo, Huanle Xu, Wing Cheong Lau
Middleware1
2023 Interference-aware Multiplexing for Deep Learning in GPU Clusters: A Middleware Approach
abstract
A common strategy for improving efficiency in training deep learning entails multiplexing tasks on a single GPU. To mitigate the interference caused by multiplexing, existing approaches primarily employ kernel-level solutions to regulate GPU kernel execution, or harness hardware-level techniques to explicitly restrict GPU streaming multiprocessors and memory. Nevertheless, none of them perform satisfactorily in optimizing the completion time of tasks.
Wenyan Chen 0001, Zizhao Mo, Huanle Xu, Kejiang Ye, Cheng-Zhong Xu 0001
SC2