Chao Sun 0008

dblp:54/3957-8 · DBLP profile ↗
← Back
14ranked-venue papers
1as first author
6since 2021 · last 2025
0000-0001-6069-441XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 11 · 1 first-author · 5 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021
YearPublicationVenuePosition
2025 Efficiency Optimization Under Spatiotemporal Sharing Fairness for Deep Learning Workloads in Heterogeneous GPU Clusters
abstract
Modern GPU clusters increasingly comprise diverse heterogeneous GPUs, driven by the continuous release of new GPU models. Achieving a balance between fairness and efficiency when scheduling multi-tenant Deep Learning (DL) training jobs on such clusters is inherently challenging. Existing DL training schedulers largely emphasize fairness through GPU temporal sharing, while the spatial dimension of resource allocation is often underexplored. This oversight can lead to GPU fragmentation and suboptimal system performance. In this paper, we propose STS-Fairness, a spatiotemporal sharing fairness scheduler. STS-Fairness partitions each GPU into multiple isolated slots under a novel spatiotemporal fairness constraint and allocates jobs using a round-based allocation mechanism. We guarantee that STS-Fairness achieves overall performance optimality while satisfying spatiotemporal fairness constraints. The scheduling problem is formulated as an integer nonlinear program (INLP) that is solved to optimality in polynomial time via dynamic programming. We deployed the STS-Fairness framework on both physical and simulated heterogeneous clusters and conducted large-scale experiments. These results demonstrate that STS-Fairness reduces average JCT by$1.2 \times$, shortens makespan by$1.24 \times$, and increases throughput by$\mathbf{1. 2 5} \times$compared to state-of-the-art (SoTA) schedulers.
Chunhong Du, Mengyu Shi, Shanjiang Tang, Jianhang Tang, Ce Yu, Jian Xiao 0001, Chao Sun 0008, Bin Yang 0043
ICPADS7
2025 Solving online resource-constrained scheduling for follow-up observation in astronomy: A reinforcement learning approach
Ce Yu, Chao Sun 0008, Jizeng Wei, Junhan Ju, Shanjiang Tang
Future Gener. Comput. Syst.3
2025 Task Scheduling in Geo-Distributed Computing: A Survey
abstract
Geo-distributed computing, a paradigm that assigns computational tasks to globally distributed nodes, has emerged as a promising approach in cloud computing, edge computing, cloud-edge computing, and supercomputer computing (SC). It enables low-latency services, ensures data locality, and handles large-scale applications. As global computing capacity and task demands increase rapidly, scheduling tasks for efficient execution in geo-distributed computing systems has become an increasingly critical research challenge. It arises from the inherent characteristics of geographic distribution, including heterogeneous network conditions, region-specific resource pricing, and varying computational capabilities across locations. Researchers have developed diverse task scheduling methods tailored to geo-distributed scenarios, aiming to achieve objectives such as performance enhancement, fairness assurance, and fault-tolerance improvement. This survey provides a comprehensive and systematic review of task scheduling techniques across four major distributed computing environments, with an in-depth analysis of these approaches based on their core scheduling objectives. Through our analysis, we identify key research challenges and outline promising directions for advancing task scheduling in geo-distributed computing.
Yujian Wu, Shanjiang Tang, Ce Yu, Bin Yang 0043, Chao Sun 0008, Jian Xiao 0001, Hutong Wu, Jinghua Feng
IEEE Trans. Parallel Distributed Syst.5
2024 Fast and accurate novelty detection for large surveillance video
Shanjiang Tang, Ce Yu, Chao Sun 0008, Yusen Li, Jian Xiao 0001
CCF Trans. High Perform. Comput.4
2024 Fairness-Efficiency Scheduling for Pay-as-You-Go Shared Caching Systems With Long-Term Fairness Guarantees
abstract
Pay-as-you-go caching systems are now widely used as storage services in cloud computing. However, users’ data caching requirements not only change over time, but also they are affected by workload characteristics, making it difficult to always ensure high efficient use of cache resources. Cache resource sharing is an effective way to improve the efficiency of cache usage. To incentivize users to share caches, it is essential to ensure long-term fairness among multiple users. However, traditional resource allocation strategies canonly guarantee memory less fairness among users,which is not applicable to long-term cache sharing systems. In this paper, we propose a fair allocation policy named FairCache for Pay-as-you-go cache resources. First, FairCache can satisfy four desirable properties of resource allocation: sharing incentive, pay-as-you-usefairness, strategy proofness, and pare to efficiency. Second, FairCache is an efficiency-fairness resource allocation policy based on the efficiency knob θ. The strategy keeps sensitive to the constantly changing cache demands of multiple users within the system by adjusting the efficiency knob θ, thus ensuring long-term multi-user fairness while maximizing the efficiency of cache usage. In addition, FairCache also has an anti-cheating mechanism to avoid possible free-rider problems when multiple users cache access files. Finally, this paper implements the FairCache policy in Alluxio. The experimental results show that FairCache is a lightweight scheduler and that it can maximize the efficiency usage of cache resources while ensuring the long-term for multiple users in the pay-as-you-go Cache systems fairness.
Shanjiang Tang, Zhongyu Zhou, Jiekai Gou, Ce Yu, Yusen Li, Hao Fu 0021, Chao Sun 0008, Jian Xiao 0001
IEEE Trans. Serv. Comput.7
2022 Long-Term Fairness Scheduler for Pay-as-You-Use Cache Sharing Systems
Zhongyu Zhou, Shanjiang Tang, Hao Fu 0021, Wanqing Chang, Ce Yu, Chao Sun 0008, Yusen Li, Jian Xiao 0001
ICA3PP6
2020 Balancing Fairness and Efficiency for Cache Sharing in Semi-external Memory System
abstract
Data caching and sharing is an effective approach for achieving high performance to many applications in shared platforms such as the cloud. DRAM and SSD are two popular caching devices widely used by many large-scale data application systems such Hadoop and Spark. Due to the limited size of DRAM as well as the large access latency of SSD (relative to DRAM), there is a trend of integrating DRAM and SSD (called semi-external memory) together for large-scale data caching.
Shanjiang Tang, Qifei Chai, Ce Yu, Yusen Li, Chao Sun 0008
ICPP5
2018 GpDL: A Spatially Aggregated Data Layout for Long-Term Astronomical Observation Archive
Ce Yu, Chao Sun 0008, Shanjiang Tang, Xiangfei Meng
ICA3PP (2)3
2018 A Data-Aware Energy-Saving Storage Management Strategy for On-Site Astronomical Observation at Dome A
Xiaoxiao Lu, Chao Sun 0008, Ce Yu, Ming Che, Zijun Xia, Zhaohui Shang
ICA3PP (2)2
2018 QKnober: A Knob-Based Fairness-Efficiency Scheduler for Cloud Computing with QoS Guarantees
Shanjiang Tang, Ce Yu, Chao Sun 0008, Jian Xiao 0001, Yinglong Li
ICSOC3
2017 Optimized Data Layout for Spatio-temporal Data in Time Domain Astronomy
Ce Yu, Chao Sun 0008, Zhaohui Shang, Jinghua Feng, Jian Xiao 0001
ICA3PP3
2017 Survey on Energy-Saving Technologies for Disk-Based Storage Systems
Ce Yu, Chao Sun 0008, Xiaoxiao Lu, Jian Xiao 0001
ICA3PP3
2015 Fast 3-Point Correlation Function Approximation on GPU
Chao Sun 0008, Mujin Yang, Ce Yu
ICA3PP (2)1
2011 Visual Modeling for Parallel Programming Based on DSL
abstract
Parallel Application Visual Modeling (PAVM) is a system that simplifies the development of parallel applications by providing a graphical user interface for visually modeling and generating corresponding source code framework according to the constructed model. The specification of graphical construction blocks and composition rules are proposed to help construct feasible visual models. Model checker is designed to conduct model verification, and code generator is implemented to translate models to source code framework. PAVM is implemented based on DSL tools. With the help of PAVM, programmers are able to focus on algorithm design and obtain source code framework from graphical models.
Ce Yu, Chao Sun 0008
CloudCom4