VLDB 2026 Research / reviewers in the wild / expert
Chengzhi Lu
dblp:213/1453
· DBLP profile ↗
14ranked-venue papers
4as first author
10since 2021 · last 2026
0000-0002-6421-3828ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 3 first-author · 8 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorArtificial intelligence and machine learning · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | High Throughput and Low Latency LLM Serving via Adaptive KV CachingabstractThe substantial memory demands of model weights and key-value (KV) caches often lead to severe memory bottlenecks in LLM serving. Existing systems address this by offloading KV caches to host memory and rapidly restoring them on demand before decoding. However, these approaches are too coarse-grained and fail to fully exploit the combined computational and storage capabilities of GPUs. Wenyan Chen 0001, Chengzhi Lu, Huanle Xu, Kejiang Ye, Cheng-Zhong Xu 0001 |
EuroSys | 2 |
| 2026 | FlexPipe: Adapting Dynamic LLM Serving Through Inflight Pipeline Refactoring in Fragmented Serverless ClustersabstractServing Large Language Models (LLMs) in production faces significant challenges from highly variable request patterns and severe resource fragmentation in serverless clusters. Current systems rely on static pipeline configurations that struggle to adapt to dynamic workload conditions, leading to substantial inefficiencies. Yanying Lin, Chengzhi Lu, Cheng-Zhong Xu 0001, Kejiang Ye |
EuroSys | 3 |
| 2025 | Multiplexing Dynamic Deep Learning Workloads with SLO-awareness in GPU ClustersabstractDeep learning (DL) inference services are widely recognized as crucial workloads in large-scale cloud clusters. However, due to the stringent latency requirements, cloud providers often over-provision GPU resources, resulting in underutilization of the available GPU potential. Although co-locating tasks on the same device can enhance utilization, ensuring Service Level Objectives (SLOs) guarantees for multiplexing highly dynamic inference services becomes extremely challenging due to significant resource interference. Wenyan Chen 0001, Chengzhi Lu, Huanle Xu, Kejiang Ye, Cheng-Zhong Xu 0001 |
EuroSys | 2 |
| 2025 | Serving LLM in Distributed GPU Cluster With Fine-Grain Pipeline ConstraintsabstractAs Large Language Models (LLMs) continue to advance, their parameter sizes are growing exponentially-far outpacing hardware capabilities. This widening gap necessitates distributed computing through pipeline parallelism for efficient inference. However, the uneven distribution of requests across pipeline stages creates significant performance bottlenecks in real-world deployments. To address this challenge, we presentPlanck, a performance optimization framework specifically designed for distributed LLM inference.Planck implements fine-grained control through two key mechanisms: a progressive SLO allocation strategy that dynamically adjusts time constraints based on workload patterns, and stage-specific performance controllers that prevent bottlenecks before they cascade through the system. By intelligently balancing resources across pipeline stages,Planck effectively eliminates queue buildup-essentially preventing traffic congestion before it forms. Evaluation using diverse workloads in real cloud environments demonstrates thatPlanck reduces P99 tail latency by up to 18% and decreases the longest queue lengths by as much as 47.8% across pipeline stages, significantly improving both system responsiveness and resource utilization. Yanying Lin, Shuaipeng Wu, Chengzhi Lu, Kejiang Ye, Cheng-Zhong Xu 0001 |
IEEE Trans. Serv. Comput. | 5 |
| 2024 | Planck: Optimizing LLM Inference Performance in Pipeline Parallelism with Fine-Grained SLO ConstraintabstractPipeline parallelism is an important strategy for improving inference performance in Large Language Models (LLMs). However, we find that different stages of LLM pipelines exhibit distinct performance and request characteristics, posing challenges to system performance in online inference scenarios. To address this issue, we propose Planck, a performance optimization framework tailored for LLM pipeline inference. By balancing request traffic, queue length, and execution time at each stage, Planck introduces a progressive SLO (Service Level Objective) allocation method and a stage instance performance controller. Planck fine-grainedly allocates SLOs to each pipeline stage and dynamically adjusts according to request distribution to control queue length. Through optimizing queue lengths across different stages of the model pipeline, Planck effectively reduces waiting time and tail latency. Evaluations conducted on a real cloud cluster using diverse workloads demonstrate that Planck effectively reduces P99 latency and queue length for each pipeline stage. Yanying Lin, Shuaipeng Wu, Chengzhi Lu, Cheng-Zhong Xu 0001, Kejiang Ye |
ICWS | 5 |
| 2024 | SMIless: Serving DAG-based Inference with Dynamic Invocations under Serverless ComputingabstractThe deployment of ML serving applications, featuring multiple inference functions on serverless platforms, has gained substantial popularity, leading to numerous developments of new systems. However, these systems often focus on optimizing resource provisioning and cold start management separately, ultimately resulting in higher monetary costs. This paper introduces SMIless, a highly efficient serverless system tailored for serving DAG-based ML inference in heterogeneous environments. SMIless effectively co-optimizes resource configuration and cold-start management in the context of dynamic invocations. This is achieved by seamlessly integrating adaptive pre-warming windows, striking an effective balance between performance and cost. We have implemented SMIless on top of OpenFaaS and conducted extensive evaluations using real-world ML serving applications. The experimental results demonstrate that SMIless can achieve up to a $5.73 \times$ reduction in the overall costs while meeting the SLA requirements for all user requests, surpassing the performance of state-of-the-art solutions. Chengzhi Lu, Huanle Xu, Yudan Li, Wenyan Chen 0001, Kejiang Ye, Cheng-Zhong Xu 0001 |
SC | 1 |
| 2023 | Understanding and Optimizing Workloads for Unified Resource Management in Large Cloud PlatformsabstractTo fully utilize computing resources, cloud providers such as Google and Alibaba choose to co-locate online services with batch processing applications in their data centers. By implementing unified resource management policies, different types of complex computing jobs request resources in a consistent way, which can help data centers achieve global optimal scheduling and provide computing power with higher quality. To understand this new scheduling paradigm, in this paper, we first present an in-depth study of Alibaba's unified scheduling workloads. Our study focuses on the characterization of resource utilization, the application running performance, and scheduling scalability. We observe that although computing resources are significantly over-committed under unified scheduling, the resource utilization in Alibaba data centers is still low. In addition, existing resource usage predictors tend to make severe overestimations. At the same time, tasks within the same application behave fairly consistently, and the running performance of tasks can be well-profiled with respect to resource contention on the corresponding physical host. Chengzhi Lu, Huanle Xu, Kejiang Ye, Guoyao Xu, Liping Zhang 0013, Cheng-Zhong Xu 0001 |
EuroSys | 1 |
| 2022 | An In-Depth Study of Microservice Call Graph and Runtime PerformanceabstractLoosely-coupled and light-weight microservices running in containers are replacing monolithic applications gradually. Understanding the characteristics of microservices is critical to make good use of microservice architectures. However, there is no comprehensive study about microservice and its related systems in production environments so far. In this paper, we present a solid analysis of large-scale deployments of microservices at Alibaba clusters. Our study focuses on the characterization of microservice dependency as well as its runtime performance. We conduct an in-depth anatomy of microservice call graphs to quantify the difference between them and traditional DAGs of data-parallel jobs. In particular, we observe that microservice call graphs are heavy-tail distributed and their topology is similar to a tree and moreover, many microservices are hot-spots. We also discover that the structure of call graphs for long-term developed applications is much simpler so as to provide better performance. Our investigation on microservice runtime performance indicates most microservices are much more sensitive to CPU interference than memory interference. Moreover, we design resource management policies to efficiently tune memory resources. Shutian Luo, Huanle Xu, Chengzhi Lu, Kejiang Ye, Guoyao Xu, Liping Zhang 0013, Jian He 0004, Cheng-Zhong Xu 0001 |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2021 | Characterizing Microservice Dependency and Performance: Alibaba Trace AnalysisabstractLoosely-coupled and light-weight microservices running in containers are replacing monolithic applications gradually. Understanding the characteristics of microservices is critical to make good use of microservice architectures. However, there is no comprehensive study about microservice and its related systems in production environments so far. In this paper, we present a solid analysis of large-scale deployments of microservices at Alibaba clusters. Our study focuses on the characterization of microservice dependency as well as its runtime performance. We conduct an in-depth anatomy of microservice call graphs to quantify the difference between them and traditional DAGs of data-parallel jobs. In particular, we observe that microservice call graphs are heavy-tail distributed and their topology is similar to a tree and moreover, many microservices are hot-spots. We reveal three types of meaningful call dependency that can be utilized to optimize microservice designs. Our investigation on microservice runtime performance indicates most microservices are much more sensitive to CPU interference than memory interference. To synthesize more representative microservice traces, we build a mathematical model to simulate call graphs. Experimental results demonstrate our model can well preserve those graph properties observed from Alibaba traces. Shutian Luo, Huanle Xu, Chengzhi Lu, Kejiang Ye, Guoyao Xu, Liping Zhang 0013, Jian He 0004, Cheng-Zhong Xu 0001 |
SoCC | 3 |
| 2021 | RPTCN: Resource Prediction for High-dynamic Workloads in Clouds based on Deep LearningabstractResource management is challenging in clouds due to the dynamics and sharing characteristics. The crucial problem is how to allocate resources accurately and satisfy demands of workloads timely. The traditional solution is to use historical data to predict future resource usage. Although these resource prediction methods can predict the periodicity, they can not accurately predict mutation points due to the high dynamics and uncertainty of resource usage. To tackle this issue, in this paper we propose a resource usage prediction method - RPTCN, which is based on a deep learning method - temporal convolutional networks (TCNs) in cloud systems. We add a fully connected layer and attention mechanism to TCNs to improve the prediction accuracy. In order to explore the relationship between the usage of different resources in the temporal dimension, we use correlation analysis to screen performance indicators as multidimensional feature input for prediction. Finally, we evaluate the performance of this method on Alibaba trace v2018. Evaluations show that RPTCN improves the overall MAE and MSE by 6.50%~89.03% and 0.41%~68.82% respectively compared to baselines in dynamic and long-term prediction of resource usage. Moreover, the convergence and generalization of RPTCN are also better than the baselines. Wenyan Chen 0001, Chengzhi Lu, Kejiang Ye, Yang Wang 0006, Cheng-Zhong Xu 0001 |
CLUSTER | 2 |
| 2020 | Interference Analysis of Co-Located Container Workloads: A Perspective from Hardware Performance Counters
Wenyan Chen 0001, Kejiang Ye, Chengzhi Lu, Dongdai Zhou, Cheng-Zhong Xu 0001 |
J. Comput. Sci. Technol. | 3 |
| 2019 | ADGS: Anomaly Detection and Localization Based on Graph Similarity in Container-Based CloudsabstractDocker container is experiencing rapid development with the support from the industry like Google and Alibaba and is being widely used in large scale production cloud environment. For example, Alibaba has deployed millions of containers for its internal business, and most of the online services are already migrated to the containers. Those services are usually very complex, spanning multiple containers with complex interaction and dependency relationship. Detecting potential anomalies in such a large container-based cloud platform is very challenging. Traditional detection models usually use system resource metrics like CPU and memory usage, but rarely consider the relationship among components, causing high false positive rate. In this paper, we present a novel Anomaly Detection and root cause localization method based on Graph Similarity (ADGS) in the container-based cloud environment. We first monitor the response time and resource usage of each component in the application to determine whether the system status is normal or not. Then, we propose a new mechanism to locate the root cause of the anomalies based on graph similarity, investigating the anomaly propagation rules among cluster components. We implement and evaluate our method in a container-based environment. The results show that the proposed method can detect and determine the root cause of anomalies efficiently and accurately. Chengzhi Lu, Kejiang Ye, Wenyan Chen 0001, Cheng-Zhong Xu 0001 |
ICPADS | 1 |
| 2018 | Modeling Application Performance in Docker Containers Using Machine Learning TechniquesabstractDocker container is experiencing a rapid development with the support from industry like Google and is being widely used in large scale production cloud environments. However the performance of applications running in Docker containers is still not clear due to the complex relationship between container resource allocation and application performance. In this paper, we first study the impact of key parameters in container resource allocation that affect the performance of containerized applications. Then, we present modeling techniques over CPU, memory and I/O resources to characterize the performance of applications running in containers. To address this multi-dimensional modeling problem, we propose three machine learning techniques, i.e. Linear Regression (LR), Support Vector Machine (SVM) and Artificial Neural Network (ANN). We implement and evaluate the modeling techniques for four complex benchmark workloads from Spark. Experimental results demonstrate the proposed models can achieve as low as 2.27% prediction error, with an average of 10.13% for most applications. Furthermore, the prediction accuracy of SVM and ANN models are substantially better than LR based approaches, with 48.13% and 29.30% improvement. Kejiang Ye, Yanmin Kou, Chengzhi Lu, Yang Wang 0006, Cheng-Zhong Xu 0001 |
ICPADS | 3 |
| 2017 | Imbalance in the cloud: An analysis on Alibaba cluster traceabstractTo improve resource efficiency and design intelligent scheduler for clouds, it is necessary to understand the workload characteristics and machine utilization in large-scale cloud data centers. In this paper, we perform a deep analysis on a newly released trace dataset by Alibaba in September 2017, consists of detail statistics of 11089 online service jobs and 12951 batch jobs co-locating on 1300 machines over 12 hours. To the best of our knowledge, this is one of the first work to analyze the Alibaba public trace. Our analysis reveals several important insights about different types of imbalance in the Alibaba cloud. Such imbalances exacerbate the complexity and challenge of cloud resource management, which might incur severe wastes of resources and low cluster utilization. 1) Spatial Imbalance: heterogeneous resource utilization across machines and workloads. 2) Temporal Imbalance: greatly time-varying resource usages per workload and machine. 3) Imbalanced proportion of multi-dimensional resources (CPU and memory) utilization per workload. 4) Imbalanced resource demands and runtime statistics (duration and task number) between online service and offline batch jobs. We argue accommodating such imbalances during resource allocation is critical to improve cluster efficiency, and will motivate the emergence of new resource managers and schedulers. Chengzhi Lu, Kejiang Ye, Guoyao Xu, Cheng-Zhong Xu 0001, Tongxin Bai |
IEEE BigData | 1 |