Chenzhi Liao

dblp:214/8220 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
1since 2021 · last 2026
0009-0003-5645-4167ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Cloud and datacenter computing · 91% Embedded and real-time systems · 9%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Cloud and datacenter computing › cluster resource management and scheduling
cluster resource management
1.012026
GFS: A Preemption-aware Scheduling Framework for GPU Clusters with Predictive Spot Instance Management · ASPLOS (1) 2026
Cloud and datacenter computing › cluster resource management and scheduling › cluster scheduling
GPU cluster scheduling
1.012026
GFS: A Preemption-aware Scheduling Framework for GPU Clusters with Predictive Spot Instance Management · ASPLOS (1) 2026
Cloud and datacenter computing › resource provisioning
spot instance management
1.012026
GFS: A Preemption-aware Scheduling Framework for GPU Clusters with Predictive Spot Instance Management · ASPLOS (1) 2026
Embedded and real-time systems › real-time scheduling
preemptive scheduling
0.312026
GFS: A Preemption-aware Scheduling Framework for GPU Clusters with Predictive Spot Instance Management · ASPLOS (1) 2026

Methods — techniques the papers use, named apart from their topics

dynamic allocation · 1.0demand forecasting · 1.0
YearPublicationVenuePosition
2026 GFS: A Preemption-aware Scheduling Framework for GPU Clusters with Predictive Spot Instance Management
abstract
The surge in large language models (LLMs) has fundamentally reshaped the landscape of GPU usage patterns, creating an urgent need for more efficient management strategies. While cloud providers employ spot instances to reduce costs for low-priority (LP) tasks, existing schedulers still grapple with high eviction rates and lengthy queuing times. To address these limitations, we present GFS, a novel preemptive scheduling framework that enhances service-level objective (SLO) compliance for high-priority (HP) tasks while minimizing preemptions to LP tasks. Firstly, GFS utilizes a lightweight forecasting model that predicts GPU demand among different tenants, enabling proactive resource management. Secondly, GFS employs a dynamic allocation mechanism to adjust the spot quota for LP tasks with guaranteed durations. Lastly, GFS incorporates a preemptive scheduling policy that prioritizes HP tasks while minimizing the impact on LP tasks. We demonstrate the effectiveness of GFS through both real-world implementation and simulations. The results show that GFS reduces eviction rates by 33.0%, and cuts queuing delays by 44.1% for LP tasks. Furthermore, GFS enhances the GPU allocation rate by up to 22.8% in real production clusters. In a production cluster of more than 10,000 GPUs, GFS yields roughly $459,715 in monthly benefits.
Jiaang Duan, Shenglin Xu, Shiyou Qian, Dingyu Yang, Kangjin Wang, Chenzhi Liao, Yinghao Yu, Qin Hua, Hanwen Hu, Dongqing Bao, Tianyu Lu, Jian Cao 0001, Guangtao Xue, Liping Zhang 0013, Gang Chen 0001
ASPLOS (1)6