VLDB 2026 Research / reviewers in the wild / expert
Shutian Luo
dblp:304/9892
· DBLP profile ↗
12ranked-venue papers
6as first author
12since 2021 · last 2026
0000-0002-3064-5841ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 12 · 6 first-author · 12 since 2021Software engineering, systems software and programming languages · 3 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Cremes: Cost-Efficient and Reliable Microservice Execution on Spot InstancesabstractWhile spot instances offer a cost-effective alternative to on-demand cloud resources, they introduce reliability challenges for latency-sensitive microservices due to preemption risks and unpredictable provisioning delays. Conventional resource management systems, which often rely on assumptions of immediate instance availability, fail to account for these operational realities—resulting in increased risk of SLO violations when deployed in spot-based environments. Liao Chen 0001, Chenyu Lin, Junlin Chen, Shutian Luo, Huanle Xu, Cheng-Zhong Xu 0001 |
HPDC | 4 |
| 2026 | Workload-Adapted Resource Allocation for LLM Distributed Serving in Serverless ClustersabstractLarge language models increasingly rely on pipeline parallelism for distributed inference, but existing systems face critical challenges in serverless environments: heterogeneous request distributions across pipeline stages and unpredictable workload patterns requiring rapid elasticity. Traditional static resource allocation fails to address pipeline-specific bottlenecks and cold start delays inherent in serverless architectures. We propose QUART, a workload-adapted resource allocation system for LLM distributed serving in serverless clusters. QUART introduces pipeline-aware resource management through: (1) latency-aware critical stage identification using coefficient of variation (CV)-based burst propagation analysis, (2) dynamic replica allocation with proportional-integral-derivative (PID) control for congested stages, and (3) hierarchical parameter caching with copy-on-write mechanisms enabling sub-second serverless scaling. The system addresses serverless-specific challenges through cache-aware scheduling that maintains model parameters in memory, eliminating disk I/O overhead during rapid scaling events. Evaluation with real-world workloads shows QUART reduces average response latency by up to 87.1% compared to existing serverless inference systems while achieving 2.37x improvement in goodput. Yanying Lin, Shutian Luo, Haiying Shen, Kejiang Ye, Cheng-Zhong Xu 0001 |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2025 | Embracing Imbalance: Dynamic Load Shifting among Microservice Containers in Shared ClustersabstractIn a unified resource scheduling architecture, containers within the same microservice often encounter temporal and spatial performance imbalance when deployed in large-scale shared clusters. As a result, the commonly employed load-balancing approach often leads to substantial resource wastage as applications are frequently over-provisioned to meet service level agreements (SLAs). Shutian Luo, Jianxiong Liao, Chenyu Lin, Huanle Xu, Zhi Zhou 0006, Cheng-Zhong Xu 0001 |
ASPLOS (2) | 1 |
| 2025 | Understanding Diffusion Model Serving in Production: A Top-Down Analysis of Workload, Scheduling, and Resource EfficiencyabstractThis paper presents a comprehensive analysis of diffusion model serving challenges in production cloud environments. We examine the unique computational patterns and resource requirements that distinguish diffusion model serving from traditional ML workloads, revealing fundamental systemlevel challenges from their multi-stage pipeline architectures. Our analysis is based on a dataset collected from a commercial image generation service processing 3.5 million requests across 300+ GPUs of production operation. Yanying Lin, Shuaipeng Wu, Shutian Luo, Hong Xu 0001, Haiying Shen, Chong Ma 0005, Cheng-Zhong Xu 0001, Lin Qu, Kejiang Ye |
SoCC | 3 |
| 2025 | Grad: Intelligent Microservice Scaling by Harnessing Resource FungibilityabstractMicroservice applications are commonly deployed alongside other services to enhance resource utilization. However, this practice also leads to notable resource contention. While existing studies primarily focus on scaling critical microservices responsible for performance degradation to mitigate violations of SLAs regarding end-to-end latency in highly interfered environments, they often overlook the potential advantages of scaling non-critical microservices for optimized resource efficiency. In this paper, we introduce Grad, an intelligent microservice scaling framework by harnessing resource fungibility between critical and non-critical microservices. Addressing the challenges posed by the dynamic nature of resource fungibility during scaling, Grad incorporates three key components. First, Grad employs a modular learning approach to profile individual microservice latency in relation to environmental conditions. Utilizing gradient extracts from this profile, Grad designs a scalable optimization module to dynamically select the optimal set of microservices for scaling. To rapidly mitigate SLA violations, Grad also deploys an accurate end-to-end latency predictor, serving as an simulator to obtain real-time feedback. We evaluate Grad in our cluster using real microservice benchmarks and production traces, demonstrating its ability to reduce resource usage by $\mathbf{4 9. 1 \%}$ and lower the probability of SLA violations by $3.7 \times$ when compared to state-of-the-art solutions. Liao Chen 0001, Chenyu Lin, Shutian Luo, Huanle Xu, Cheng-Zhong Xu 0001 |
HPCA | 3 |
| 2024 | QUART: Latency-Aware FaaS System for Pipelining Large Model InferenceabstractPipeline parallelism is a key mechanism to ensure the performance of large model serving systems. These systems need to deal with unpredictable online workloads with low latency and high good put. However, due to the specific characteristics of large models and resource constraints in pipeline parallelism, existing systems struggle to balance resource allocation across pipeline stages. The primary challenge resides in the differential distribution of requests across various stages of the pipeline. We propose QUART, a large model serving system that focuses on optimizing the performance of key stages in pipeline parallelism. QUART dynamically identifies the key stages of the pipeline and introduces an innovative two-level model parameter caching system based on forks to achieve rapid scaling of key stages within seconds. In evaluations with real-world request workloads, QUART reduces average response latency by up to 87.1%) and increases good put by 2.37x compared to the baseline. The experiments demonstrate that QUART effectively reduces tail latency and the average queue length of the pipeline. Yanying Lin, Yingfei Tang, Shutian Luo, Haiying Shen, Cheng-Zhong Xu 0001, Kejiang Ye |
ICDCS | 5 |
| 2024 | Derm: SLA-aware Resource Management for Highly Dynamic MicroservicesabstractEnsuring efficient resource allocation while providing service level agreement (SLA) guarantees for end-to-end (E2E) latency is crucial for microservice applications. Although existing studies have made significant contributions towards achieving this objective, they primarily concentrate on static graphs. However, microservice graphs are inherently dynamic during runtime in production environments, necessitating more effective and scalable resource management solutions.In this paper, we present Derm, a new resource management system designed for microservice applications with highly dynamic graphs. Our principal finding is that prioritizing different microservice graphs can lead to a substantial reduction in resource allocation. To take advantage of this opportunity, we develop three main components. The first is a performance model that describes uncertainties of microservice latency through a conditional exponential distribution. The second is a probabilistic quantification of the dynamics of microservice graphs. The third is an optimization method for adjusting the resource allocation of microservices to minimize resource usage. We evaluate Derm in our cluster using real microservice benchmarks and production traces. The results highlight that Derm reduces the resource usage by $68.4 \%$ and lowers SLA violation probability by $6.7 \times$, compared to existing approaches. Liao Chen 0001, Shutian Luo, Chenyu Lin, Zizhao Mo, Huanle Xu, Kejiang Ye, Cheng-Zhong Xu 0001 |
ISCA | 2 |
| 2024 | Optimizing Resource Management for Shared Microservices: A Scalable System DesignabstractA common approach to improving resource utilization in data centers is to adaptively provision resources based on the actual workload. One fundamental challenge of doing this in microservice management frameworks, however, is that different components of a service can exhibit significant differences in their impact on end-to-end performance. To make resource management more challenging, a single microservice can be shared by multiple online services that have diverse workload patterns and SLA requirements. We present an efficient resource management system, namely Erms, for guaranteeing SLAs with high probability in shared microservice environments. Erms profiles microservice latency as a piece-wise linear function of the workload, resource usage, and interference. Based on this profiling, Erms builds resource scaling models to optimally determine latency targets for microservices with complex dependencies. Erms also designs new scheduling policies at shared microservices to further enhance resource efficiency. Experiments across microservice benchmarks as well as trace-driven simulations demonstrate that Erms can reduce SLA violation probability by 5× and more importantly, lead to a reduction in resource usage by 1.6×, compared to state-of-the-art approaches. Shutian Luo, Chenyu Lin, Kejiang Ye, Guoyao Xu, Liping Zhang 0013, Huanle Xu, Cheng-Zhong Xu 0001 |
ACM Trans. Comput. Syst. | 1 |
| 2023 | Erms: Efficient Resource Management for Shared Microservices with SLA GuaranteesabstractA common approach to improving resource utilization in data centers is to adaptively provision resources based on the actual workload. One fundamental challenge of doing this in microservice management frameworks, however, is that different components of a service can exhibit significant differences in their impact on end-to-end performance. To make resource management more challenging, a single microservice can be shared by multiple online services that have diverse workload patterns and SLA requirements. Shutian Luo, Huanle Xu, Kejiang Ye, Guoyao Xu, Liping Zhang 0013, Jian He 0004, Cheng-Zhong Xu 0001 |
ASPLOS (1) | 1 |
| 2022 | The power of prediction: microservice auto scaling via workload learningabstractWhen deploying microservices in production clusters, it is critical to automatically scale containers to improve cluster utilization and ensure service level agreements (SLA). Although reactive scaling approaches work well for monolithic architectures, they are not necessarily suitable for microservice frameworks due to the long delay caused by complex microservice call chains. In contrast, existing proactive approaches leverage end-to-end performance prediction for scaling, but cannot effectively handle microservice multiplexing and dynamic microservice dependencies. Shutian Luo, Huanle Xu, Kejiang Ye, Guoyao Xu, Liping Zhang 0013, Cheng-Zhong Xu 0001 |
SoCC | 1 |
| 2022 | An In-Depth Study of Microservice Call Graph and Runtime PerformanceabstractLoosely-coupled and light-weight microservices running in containers are replacing monolithic applications gradually. Understanding the characteristics of microservices is critical to make good use of microservice architectures. However, there is no comprehensive study about microservice and its related systems in production environments so far. In this paper, we present a solid analysis of large-scale deployments of microservices at Alibaba clusters. Our study focuses on the characterization of microservice dependency as well as its runtime performance. We conduct an in-depth anatomy of microservice call graphs to quantify the difference between them and traditional DAGs of data-parallel jobs. In particular, we observe that microservice call graphs are heavy-tail distributed and their topology is similar to a tree and moreover, many microservices are hot-spots. We also discover that the structure of call graphs for long-term developed applications is much simpler so as to provide better performance. Our investigation on microservice runtime performance indicates most microservices are much more sensitive to CPU interference than memory interference. Moreover, we design resource management policies to efficiently tune memory resources. Shutian Luo, Huanle Xu, Chengzhi Lu, Kejiang Ye, Guoyao Xu, Liping Zhang 0013, Jian He 0004, Cheng-Zhong Xu 0001 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2021 | Characterizing Microservice Dependency and Performance: Alibaba Trace AnalysisabstractLoosely-coupled and light-weight microservices running in containers are replacing monolithic applications gradually. Understanding the characteristics of microservices is critical to make good use of microservice architectures. However, there is no comprehensive study about microservice and its related systems in production environments so far. In this paper, we present a solid analysis of large-scale deployments of microservices at Alibaba clusters. Our study focuses on the characterization of microservice dependency as well as its runtime performance. We conduct an in-depth anatomy of microservice call graphs to quantify the difference between them and traditional DAGs of data-parallel jobs. In particular, we observe that microservice call graphs are heavy-tail distributed and their topology is similar to a tree and moreover, many microservices are hot-spots. We reveal three types of meaningful call dependency that can be utilized to optimize microservice designs. Our investigation on microservice runtime performance indicates most microservices are much more sensitive to CPU interference than memory interference. To synthesize more representative microservice traces, we build a mathematical model to simulate call graphs. Experimental results demonstrate our model can well preserve those graph properties observed from Alibaba traces. Shutian Luo, Huanle Xu, Chengzhi Lu, Kejiang Ye, Guoyao Xu, Liping Zhang 0013, Jian He 0004, Cheng-Zhong Xu 0001 |
SoCC | 1 |