Ruobing Chen 0002

dblp:123/7895-2 · DBLP profile ↗
← Back
8ranked-venue papers
6as first author
6since 2021 · last 2023
0000-0002-4660-0883ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 6 first-author · 6 since 2021
YearPublicationVenuePosition
2023 OLPart: Online Learning based Resource Partitioning for Colocating Multiple Latency-Critical Jobs on Commodity Computers
abstract
Colocating multiple jobs on the same server has been a commonly used approach for improving resource utilization in cloud environments. However, performance interference due to the contention over shared resources makes resource partitioning an important research problem. Partitioning multiple resources coordinately is particularly challenging when multiple latency-critical (LC) jobs are colocated with best-effort (BE) jobs, since the QoS needs to be protected for all the LC jobs. So far, this problem is not well-addressed in the literatures.
Ruobing Chen 0002, Haosen Shi 0001, Yusen Li, Xiaoguang Liu 0001, Gang Wang 0001
EuroSys1
2023 CoTuner: A Hierarchical Learning Framework for Coordinately Optimizing Resource Partitioning and Parameter Tuning
abstract
The performance of modern multi-core systems is reliant upon two crucial configurations: how the resources are partitioned among the co-located applications to mitigate resource contention, and the setting of the parameters of applications. However, finding the optimal resource partition configuration and parameter setting for the co-located applications is challenging due to the prohibitively large search space and the high interdependency between resource partitioning and parameter tuning.
Tiannuo Yang, Ruobing Chen 0002, Yusen Li, Xiaoguang Liu 0001, Gang Wang 0001
ICPP2
2023 Jointly Optimizing Job Assignment and Resource Partitioning for Improving System Throughput in Cloud Datacenters
abstract
Colocating multiple jobs on the same server has been widely applied for improving resource utilization in cloud datacenters. However, the colocated jobs would contend for the shared resources, which could lead to significant performance degradation. An efficient approach for eliminating performance interference is to partition the shared resources among the colocated jobs. However, this makes the resource management in datacenters very challenging. In this paper, we propose JointOPT, the first resource management framework that optimizes job assignment and resource partitioning jointly for improving the throughput of cloud datacenters. JointOPT uses a local search based algorithm to find the near optimal job assignment configuration, and uses a deep reinforcement learning (DRL) based approach to dynamically partition the shared resources among the colocated jobs. In order to reduce the interaction overhead with real systems, it leverages deep learning to estimate job performance without running them on real servers. We conduct extensive experiments to evaluate JointOPT and the results show that JointOPT significantly outperforms the state-of-the-art baselines, with an advantage from 13.3% to 47.7%.
Ruobing Chen 0002, Haosen Shi 0001, Jinping Wu, Yusen Li, Xiaoguang Liu 0001, Gang Wang 0001
ACM Trans. Archit. Code Optim.1
2023 Orchid: An Online Learning Based Resource Partitioning Framework for Job Colocation With Multiple Objectives
abstract
Colocating multiple throughput-oriented jobs on the same server is a commonly used approach for improving system throughput in modern datacenters. The shared resources of the server are usually partitioned among the colocated jobs in order to prevent performance interference caused by resource contention. However, how to properly partition the shared resources among the colocated jobs is nontrival, because it usually has to trade off between two conflict objectives, as datacenter manager wants to maximize server throughput while job owners hope to experience a fair slowdown. Moreover, a desirable resource partitioning strategy should also be efficient, autonomous and adaptive. So far, this problem is not well-addressed in the literature due to several critical challenges. In this paper, we propose an online learning based framework, namedOrchid, to address the multi-objective resource partitioning problem. Orchid leverages contextual multi-armed bandit (CMAB) to model the resource partitioning problem and uses a light-weight online learning algorithm to learn the optimal partitioning configuration according to some easy-to-collect runtime system status. Orchid has two distinguished properties compared to the existing solutions: first, it has the ability to trade off between the two objectives flexibly according to the aspiration of decision maker; second, it has the awareness about runtime system status, which can help to improve the efficiency of finding the optimal partitioning configuration and the adaptivity to dynamic environment changes. Moreover, Orchid does not require any prior knowledge of jobs, and incurs a very small computational overhead. Our evaluations show that Orchid achieves all the desired properties of a good resource partitioning strategy, which outperforms the state-of-the-art baselines with significant margins.Orchidis publicly available athttps://github.com/OpenSourceOrchid/Orchid.
Ruobing Chen 0002, Wangqi Peng, Yusen Li, Xiaoguang Liu 0001, Gang Wang 0001
IEEE Trans. Computers1
2022 GCNPart: Interference-Aware Resource Partitioning Framework with Graph Convolutional Neural Networks and Deep Reinforcement Learning
Ruobing Chen 0002, Haosen Shi 0001, Jinping Wu, Yusen Li, Xiaoguang Liu 0001, Gang Wang 0001
ICA3PP1
2021 DRLPart: A Deep Reinforcement Learning Framework for Optimally Efficient and Robust Resource Partitioning on Commodity Servers
abstract
Workload consolidation is a commonly used approach for improving resource utilization of commodity servers. However, colocated workloads often suffer from significant performance degradations due to resource contention, which makes resource partitioning an important research problem. Partitioning multiple resources coordinately is particularly challenging due to the complex contention behaviors and huge solution space, which is not well-addressed in the literature.
Ruobing Chen 0002, Jinping Wu, Haosen Shi 0001, Yusen Li, Xiaoguang Liu 0001, Gang Wang 0001
HPDC1
2020 Deep Learning Assisted Resource Partitioning for Improving Performance on Commodity Servers
abstract
In this paper, we introduce a deep reinforcement learning (DRL) framework for solving the problem of partitioning LLC and memory bandwidth coordinately in an end-to-end manner. To this end, we formulate the problem as a markov decision process and utilize DRL algorithm to derive the optimal partition. To avoid the extensive cost of training the policy on physical server, we present a model-based solution, where a reward prediction model is leveraged to train the partitioning policy offline. To construct a precise reward prediction model, we introduce a novel representation for the partitioning scheme, where graph convolutional networks (GCN) is employed to represent the LLC partition as a bipartite graph so that those heterogeneous but identical partitions could result in the same representations and thus eases the prediction task.
Ruobing Chen 0002, Jinping Wu, Haosen Shi 0001, Yusen Li, Haiyan Yin, Shanjiang Tang, Xiaoguang Liu 0001, Gang Wang 0001
PACT1
2019 GAugur: Quantifying Performance Interference of Colocated Games for Improving Resource Utilization in Cloud Gaming
abstract
Cloud gaming has been very popular recently, but providing satisfactory gaming experiences to players at a modest cost is still challenging. Colocating several games onto one server could improve server utilization. To enable efficient colocations while providing Quality of Service (QoS) guarantees, a precise quantification of performance interference among colocated games is required. However, achieving such precise interference prediction is very challenging for games due to the complexity introduced by the contention on many shared resources across CPU and GPU. Moreover, the distinctive properties of cloud gaming require that the prediction model should be constructed beforehand and the prediction should be made instantaneously at request arrivals, which further increases the difficulty. The existing solutions are either not applicable or not effective due to many limitations. In this paper, we present GAugur, a novel methodology that enables highly accurate prediction of the performance interference among games arbitrarily colocated. By leveraging machine learning technologies, GAugur is able to capture the complex relationship between the interference and the contention features of colocated games. We evaluate GAugur through extensive experiments using a large number of real popular games. The results show that GAugur is able to identify whether a colocated game satisfies QoS requirement within an average error of 5%, and is able to quantify the performance degradation of a colocated game within an average error of 7.9%, which significantly outperforms the alternatives. Moreover, GAugur incurs an offline profiling cost linear to the number of games, and negligible overhead for online prediction. We apply GAugur to guiding efficient game colocations for cloud gaming. Experimental results show that GAugur is able to increase the resource utilization by 20% to 60%, and improve the overall performance by up to 15%, compared to the state-of-the-art solutions.
Yusen Li, Chuxu Shan, Ruobing Chen 0002, Xueyan Tang, Wentong Cai 0001, Shanjiang Tang, Xiaoguang Liu 0001, Gang Wang 0001, Xiaoli Gong, Ying Zhang 0015
HPDC3