Guoyao Xu

dblp:141/5593 · DBLP profile ↗
← Back
16ranked-venue papers
2as first author
12since 2021 · last 2025
0000-0002-1136-2678ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 14 · 2 first-author · 11 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021Artificial intelligence and machine learning · 1Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 EXIST: Enabling Extremely Efficient Intra-Service Tracing Observability in Datacenters
abstract
The complexity of online applications is rapidly increasing, bringing more sophisticated performance anomalies in today's cloud datacenter. To fully understand application behaviors, we should obtain both inter-service communication data via RPC-level tracing and intra-service execution traces via application-level tracing to precisely reason about event causality. However, the average time overhead of existing intra-service tracing schemes on the traced applications is generally about 5-10%, possibly reaching 18% in the worst case. To realize practical intra-service tracing in shared and stressed datacenters, one must achieve extreme tracing efficiency with an overhead at the per-mille level.
Xinkai Wang 0003, Xiaofeng Hou, Chao Li 0009, Yuancheng Li 0001, Du Liu, Guoyao Xu, Liping Zhang 0013, Yuemin Wu, Xiaopeng Yuan, Quan Chen 0002, Minyi Guo
ASPLOS (2)6
2025 WDP: Mitigating Interference in CPU Sharing Through Wake-up Delay Driven Preemption for QoS-aware Co-location
abstract
As Latency-critical (LC) tasks often experience diurnal load patterns, co-locating them with best-effort (BE) tasks improves resource utilization. Prior work allocates entire CPU cores between co-located tasks, due to the incapability of handling the interference with CPU sharing. We observed that the root cause of the interference on the same core is the inherent wake-up delay in the operating system scheduler, the wait time that a process can obtain the CPU cycles after it is woken up. Based on the finding, we propose WDP, a scheme that efficiently improves the throughput of BE tasks while ensuring QoS, leveraging CPU sharing. WDP comprises a wake-up delay-driven preemption mechanism and a preemption-based CPU manager. The preemption mechanism enables controlled preemption to reduce the wake-up delay of LC tasks with adjustable preemption capacity. Adopting the novel preemption mechanism, the CPU manager allocates CPU resources in a fine-grained manner among co-located tasks. Compared with the representative prior method, WDP improves the throughput of BE tasks by 31.2% on average while ensuring the QoS of co-located LC tasks.
Yaoxuan Li, Pu Pang, Yecheng Yang, Quan Chen 0002, Zhengxuan Yan, Guoyao Xu, Liping Zhang 0013, Minyi Guo
SoCC6
2025 Reducing the End-to-End Latency of DNN-Based Recommendation Systems in GPU Pools
abstract
While intelligent applications (e.g., recommendation systems) prefer different CPU-GPU ratios, GPU pooling technique that decouples the GPU and CPU resources yields substantial flexibility when serving diverse applications. With such architecture, DNN-based recommendation services often offload the compute-intensive neural network layers to the remote GPU pool for high resource utilization. However, such a paradigm results in the long end-to-end latency due to two causes: 1) the intermediate data is copied for multiple times during the entire process in current GPU pooling practices, incurring heavy overheads; 2) the content transferred to the GPU pool involves multiple small tensors, suffering from poor bandwidth efficiency. To solve these problems, we design Zero, a runtime system that incorporates a zero-copy transmission mechanism as well as a dynamic tensor merging policy. The zero-copy transmission mechanism unifies memory management across the inference framework and the RPC framework, accompanied by an elaborated serialization protocol to fully eliminate redundant data copying. Meanwhile, the tensor merging policy deliberately organizes small tensors into larger data blocks, so as to transfer them with higher efficiency. Experimental results show that, compared with prior work, Zero reduces the latency of typical recommendation models by up to 15.1% (10.1% on average).
Guangqiang Luan, Pu Pang, Quan Chen 0002, Chen Chen 0067, Guoyao Xu, Chi Zhang 0005, Yanyi Zi, Yinghao Yu, Liping Zhang 0013, Minyi Guo
IPDPS5
2024 DeployFix: Dynamic Repair of Software Deployment Failures via Constraint Solving
abstract
Software deployment misconfiguration often happens and has been one of the major causes of deployment failures that give rise to service interruptions. However, there is currently no existing approach to automatically repairing deployment failures. We propose DeployFix, which automatically repairs software deployment failures via constraint solving in the dynamic-changing deployment environments. DeployFix first defines DeployIR as a unified intermediate representation to achieve the translation of heterogeneous specifications from different schedulers with different syntaxes. By reducing the root-cause analysis of deployment failures to the conflict resolution in propositional logic, DeployFix uses off-the-shelf constraint solvers to achieve automatic localization and diagnosis of conflicting constraints, which are the root causes of deployment failures. DeployFix finally resolves the conflicting constraints and generates repaired deployment configurations in terms of practical requirements. We evaluate DeployFix in both simulation and production environments with tens of thousands of nodes at Alibaba, on which tens of thousands of applications are running guided by hundreds of thousands of deployment constraints. Experimental results demonstrate that DeployFix outperforms the state of the art and it correctly repairs the deployment failures in minutes, even in a large production data center.
Haoyu Liao, Jianmei Guo, Bo Huang 0002, Yujie Han, Dingyu Yang, Kai Shi 0006, Jonathan Ding, Guoyao Xu, Liping Zhang 0013
ASE8
2024 Optimizing Resource Management for Shared Microservices: A Scalable System Design
abstract
A common approach to improving resource utilization in data centers is to adaptively provision resources based on the actual workload. One fundamental challenge of doing this in microservice management frameworks, however, is that different components of a service can exhibit significant differences in their impact on end-to-end performance. To make resource management more challenging, a single microservice can be shared by multiple online services that have diverse workload patterns and SLA requirements. We present an efficient resource management system, namely Erms, for guaranteeing SLAs with high probability in shared microservice environments. Erms profiles microservice latency as a piece-wise linear function of the workload, resource usage, and interference. Based on this profiling, Erms builds resource scaling models to optimally determine latency targets for microservices with complex dependencies. Erms also designs new scheduling policies at shared microservices to further enhance resource efficiency. Experiments across microservice benchmarks as well as trace-driven simulations demonstrate that Erms can reduce SLA violation probability by 5× and more importantly, lead to a reduction in resource usage by 1.6×, compared to state-of-the-art approaches.
Shutian Luo, Chenyu Lin, Kejiang Ye, Guoyao Xu, Liping Zhang 0013, Huanle Xu, Cheng-Zhong Xu 0001
ACM Trans. Comput. Syst.4
2023 Erms: Efficient Resource Management for Shared Microservices with SLA Guarantees
abstract
A common approach to improving resource utilization in data centers is to adaptively provision resources based on the actual workload. One fundamental challenge of doing this in microservice management frameworks, however, is that different components of a service can exhibit significant differences in their impact on end-to-end performance. To make resource management more challenging, a single microservice can be shared by multiple online services that have diverse workload patterns and SLA requirements.
Shutian Luo, Huanle Xu, Kejiang Ye, Guoyao Xu, Liping Zhang 0013, Jian He 0004, Cheng-Zhong Xu 0001
ASPLOS (1)4
2023 Understanding and Optimizing Workloads for Unified Resource Management in Large Cloud Platforms
abstract
To fully utilize computing resources, cloud providers such as Google and Alibaba choose to co-locate online services with batch processing applications in their data centers. By implementing unified resource management policies, different types of complex computing jobs request resources in a consistent way, which can help data centers achieve global optimal scheduling and provide computing power with higher quality. To understand this new scheduling paradigm, in this paper, we first present an in-depth study of Alibaba's unified scheduling workloads. Our study focuses on the characterization of resource utilization, the application running performance, and scheduling scalability. We observe that although computing resources are significantly over-committed under unified scheduling, the resource utilization in Alibaba data centers is still low. In addition, existing resource usage predictors tend to make severe overestimations. At the same time, tasks within the same application behave fairly consistently, and the running performance of tasks can be well-profiled with respect to resource contention on the corresponding physical host.
Chengzhi Lu, Huanle Xu, Kejiang Ye, Guoyao Xu, Liping Zhang 0013, Cheng-Zhong Xu 0001
EuroSys4
2022 The power of prediction: microservice auto scaling via workload learning
abstract
When deploying microservices in production clusters, it is critical to automatically scale containers to improve cluster utilization and ensure service level agreements (SLA). Although reactive scaling approaches work well for monolithic architectures, they are not necessarily suitable for microservice frameworks due to the long delay caused by complex microservice call chains. In contrast, existing proactive approaches leverage end-to-end performance prediction for scaling, but cannot effectively handle microservice multiplexing and dynamic microservice dependencies.
Shutian Luo, Huanle Xu, Kejiang Ye, Guoyao Xu, Liping Zhang 0013, Cheng-Zhong Xu 0001
SoCC4
2022 Characterizing Job Microarchitectural Profiles at Scale: Dataset and Analysis
abstract
Understanding the microarchitectural resource characteristics of datacenter jobs has become increasingly critical to guarantee the performance of jobs while improving resource utilization. Prior work studied the resource characteristics of datacenter jobs at the OS level, little reveals the deep and detailed characteristics at the microarchitecture level due to the lack of related open traces. In this paper, we provide a new open trace, AMTrace (Alibaba Microarchitecture Trace) 1, which is profiled from 8,577 high-end physical hosts from Alibaba’s datacenter by a hardware/software co-design monitoring method. AMTrace provides the microarchitectural metrics of 9.8 × 105 Linux containers with ”Per-Container-Per-Logic CPU” granularity. Different from existing open traces, AMTrace provides a new perspective to analyze the microarchitectural resource characteristics of datacenter jobs. Based on AMTrace, we first reveal the uneven resource usage of jobs among multiple logic CPUs. Then, we analyze the impact of resource contention of CPU and memory bandwidth on job performance. Finally, we analyze the job performance under different CPU provisioning modes from microarchitecture perspective. These analyses lead to constructive insights for datacenter resource management and optimization. Furthermore, we discuss possible research opportunities on AMTrace and we believe that AMTrace will inspire more exciting research on microarchitecture and resource management.
Kangjin Wang, Ying Li 0012, Kingsum Chow, Yaoyong Dou, Guoyao Xu, Chuanjia Hou, Liping Zhang 0013
ICPP8
2022 Characterizing Co-Located Workloads in Alibaba Cloud Datacenters
abstract
Workload characteristics are vital for both data center operation and job scheduling in co-located data centers, where online services and batch jobs are deployed on the same production cluster. In this article, a comprehensive analysis is conducted on Alibaba's cluster-trace-v2018 of a production cluster of 4034 machines. The findings and insights are the following: (1) The workload on the production cluster poses a daily cyclical fluctuation, in terms of CPU and disk I/O utilization, and the memory system has become the performance bottleneck of a co-located cluster. (2) Batch jobs including their tasks and derived instances can be approximated as Zipf distribution. However, for all batch jobs with directed acyclic graph dependency, they suffer from co-location with online services since the online services are highly prioritized. (3) The resource usages of containers have similar cyclical fluctuation consistent with the whole cluster, while their memory usages remain approximately constant. (4) The number of batch jobs co-located with online services is dependent on the mispredictions per kilo instructions of online services. In order to guarantee the QoS of online services, when the MPKI of online services rises, the number of batch jobs to be co-located on the same machine should decrease.
Congfeng Jiang, Yitao Qiu, Weisong Shi, Zhefeng Ge, Shenglei Chen, Christophe Cérin, Zujie Ren, Guoyao Xu, Jiangbin Lin
IEEE Trans. Cloud Comput.9
2022 An In-Depth Study of Microservice Call Graph and Runtime Performance
abstract
Loosely-coupled and light-weight microservices running in containers are replacing monolithic applications gradually. Understanding the characteristics of microservices is critical to make good use of microservice architectures. However, there is no comprehensive study about microservice and its related systems in production environments so far. In this paper, we present a solid analysis of large-scale deployments of microservices at Alibaba clusters. Our study focuses on the characterization of microservice dependency as well as its runtime performance. We conduct an in-depth anatomy of microservice call graphs to quantify the difference between them and traditional DAGs of data-parallel jobs. In particular, we observe that microservice call graphs are heavy-tail distributed and their topology is similar to a tree and moreover, many microservices are hot-spots. We also discover that the structure of call graphs for long-term developed applications is much simpler so as to provide better performance. Our investigation on microservice runtime performance indicates most microservices are much more sensitive to CPU interference than memory interference. Moreover, we design resource management policies to efficiently tune memory resources.
Shutian Luo, Huanle Xu, Chengzhi Lu, Kejiang Ye, Guoyao Xu, Liping Zhang 0013, Jian He 0004, Cheng-Zhong Xu 0001
IEEE Trans. Parallel Distributed Syst.5
2021 Characterizing Microservice Dependency and Performance: Alibaba Trace Analysis
abstract
Loosely-coupled and light-weight microservices running in containers are replacing monolithic applications gradually. Understanding the characteristics of microservices is critical to make good use of microservice architectures. However, there is no comprehensive study about microservice and its related systems in production environments so far. In this paper, we present a solid analysis of large-scale deployments of microservices at Alibaba clusters. Our study focuses on the characterization of microservice dependency as well as its runtime performance. We conduct an in-depth anatomy of microservice call graphs to quantify the difference between them and traditional DAGs of data-parallel jobs. In particular, we observe that microservice call graphs are heavy-tail distributed and their topology is similar to a tree and moreover, many microservices are hot-spots. We reveal three types of meaningful call dependency that can be utilized to optimize microservice designs. Our investigation on microservice runtime performance indicates most microservices are much more sensitive to CPU interference than memory interference. To synthesize more representative microservice traces, we build a mathematical model to simulate call graphs. Experimental results demonstrate our model can well preserve those graph properties observed from Alibaba traces.
Shutian Luo, Huanle Xu, Chengzhi Lu, Kejiang Ye, Guoyao Xu, Liping Zhang 0013, Jian He 0004, Cheng-Zhong Xu 0001
SoCC5
2019 MEER: Online Estimation of Optimal Memory Reservations for Long Lived Containers in In-Memory Cluster Computing
abstract
Modern in-memory data-intensive computing systems like Spark create long-lived containers to execute diverse types of applications. They rely on a cluster manager like YARN or Mesos to perform resource allocation to the containers. The cluster manager or scheduler requires users of the containers to reserve resources beforehand. It is a challenge to estimate just right amounts of memory to run the applications before execution, so as to avoid over-or under-provisioning of memory space. We discover a general property of memory reservation elasticity, which allows applications to run with a reservation limit smaller than they would ideally need while only paying a moderate performance penalty. Based on the property, we designed a system, namely MEER, which performs online estimation of minimum necessary amount of memory limit that achieves nearly optimal performance. We referred to it as optimal reservation, which divides memory over-provisioning from under-provisioning. It is non-trivial to efficiently estimate optimal reservations on line through one step without runtime history. MEER uses a two-step approach to dealing with the challenge: 1) Do robust profiling and probability density analysis of applications' memory footprints in two pilot runs. By using confidence levels for the prediction, we reduce the negative effects of container footprints' randomness and achieve a highly accurate online initial estimation (over 80% accuracy) of optimal reservation. 2) By exploiting a self-decay property of the analytical results, MEER adaptively performs iterative search based on a feed-back control mechanism over subsequent recurring executions. We implemented MEER atop of YARN and evaluated the prototype by running 15 benchmark workloads on a 16-node local cluster. Evaluation results show that it achieves an average accuracy of more than 95%. By deploying MEER on schedulers and allocating memory according to the optimal reservations, one could improve cluster memory utilization by about 40%. It reduces individual application execution time by 2 to 6 times on average compared to the state-of-the-art approaches. A 90 times peak speedup for PageRank in comparison with the default Spark/Yarn is observed.
Guoyao Xu, Cheng-Zhong Xu 0001
ICDCS1
2018 How Does the Workload Look Like in Production Cloud? Analysis and Clustering of Workloads on Alibaba Cluster Trace
abstract
Cloud computing technology is widely used in today's datacenters due to the benefits such as high scalability, on-demand services and low cost. An in-depth understanding of the characteristics of workloads running in production cloud environments is very important for improving the resource management efficiency. In this paper, we make a detailed analysis with visualization techniques and clustering methods on the trace dataset released by Alibaba which contains 11089 online services and 12951 batch jobs running on 1313 machines. Our methodology for clustering workloads contains: i) Select effective feature vectors as the dimensions of clustering; ii) Identify the cluster boundaries of each dimension using K-Means algorithm; iii) Classify jobs by combining the feature vectors which uses the results from previous step; iv) Analyze the characteristics of workload groups at runtime. Our analysis reveals several insights which previous work has not found on Alibaba cluster trace. For batch jobs: a) Average CPU cores of all batch jobs show bimodal-distribution obviously. b) At a random sampling time, more than 50 % machines only run one group of jobs with a short duration, medium CPU cores and small memory utilization, the remaining machines run mixed groups of jobs. For online instances: a) The resource usage (CPU, Memory, and Disk) of most online instances is low; b) There are up to six groups running on the same machine according to our clustering method at a random sampling time.
Wenyan Chen 0001, Kejiang Ye, Yang Wang 0006, Guoyao Xu, Cheng-Zhong Xu 0001
ICPADS4
2017 Imbalance in the cloud: An analysis on Alibaba cluster trace
abstract
To improve resource efficiency and design intelligent scheduler for clouds, it is necessary to understand the workload characteristics and machine utilization in large-scale cloud data centers. In this paper, we perform a deep analysis on a newly released trace dataset by Alibaba in September 2017, consists of detail statistics of 11089 online service jobs and 12951 batch jobs co-locating on 1300 machines over 12 hours. To the best of our knowledge, this is one of the first work to analyze the Alibaba public trace. Our analysis reveals several important insights about different types of imbalance in the Alibaba cloud. Such imbalances exacerbate the complexity and challenge of cloud resource management, which might incur severe wastes of resources and low cluster utilization. 1) Spatial Imbalance: heterogeneous resource utilization across machines and workloads. 2) Temporal Imbalance: greatly time-varying resource usages per workload and machine. 3) Imbalanced proportion of multi-dimensional resources (CPU and memory) utilization per workload. 4) Imbalanced resource demands and runtime statistics (duration and task number) between online service and offline batch jobs. We argue accommodating such imbalances during resource allocation is critical to improve cluster efficiency, and will motivate the emergence of new resource managers and schedulers.
Chengzhi Lu, Kejiang Ye, Guoyao Xu, Cheng-Zhong Xu 0001, Tongxin Bai
IEEE BigData3
2017 Prometheus: online estimation of optimal memory demands for workers in in-memory distributed computation
abstract
Modern in-memory distributed computation frameworks like Spark adequately leverage memory resources to cache intermediate data across multi-stage tasks in pre-allocated worker processes, so as to speedup executions. They rely on a cluster resource manager like Yarn or Mesos to pre-reserve specific amount of CPU and memory for workers ahead of task scheduling. Since a worker is executed for an entire application and runs multiple batches of DAG tasks from multi-stages, its memory demands change over time [3].
Guoyao Xu, Cheng-Zhong Xu 0001
SoCC1