VLDB 2026 Research / reviewers in the wild / expert
Liping Zhang 0013
dblp:48/6735-13
· DBLP profile ↗
24ranked-venue papers
0as first author
23since 2021 · last 2026
0000-0003-2334-3471ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 19 · 19 since 2021Software engineering, systems software and programming languages · 6 · 5 since 2021Computer networks · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GFS: A Preemption-aware Scheduling Framework for GPU Clusters with Predictive Spot Instance ManagementabstractThe surge in large language models (LLMs) has fundamentally reshaped the landscape of GPU usage patterns, creating an urgent need for more efficient management strategies. While cloud providers employ spot instances to reduce costs for low-priority (LP) tasks, existing schedulers still grapple with high eviction rates and lengthy queuing times. To address these limitations, we present GFS, a novel preemptive scheduling framework that enhances service-level objective (SLO) compliance for high-priority (HP) tasks while minimizing preemptions to LP tasks. Firstly, GFS utilizes a lightweight forecasting model that predicts GPU demand among different tenants, enabling proactive resource management. Secondly, GFS employs a dynamic allocation mechanism to adjust the spot quota for LP tasks with guaranteed durations. Lastly, GFS incorporates a preemptive scheduling policy that prioritizes HP tasks while minimizing the impact on LP tasks. We demonstrate the effectiveness of GFS through both real-world implementation and simulations. The results show that GFS reduces eviction rates by 33.0%, and cuts queuing delays by 44.1% for LP tasks. Furthermore, GFS enhances the GPU allocation rate by up to 22.8% in real production clusters. In a production cluster of more than 10,000 GPUs, GFS yields roughly $459,715 in monthly benefits. Jiaang Duan, Shenglin Xu, Shiyou Qian, Dingyu Yang, Kangjin Wang, Chenzhi Liao, Yinghao Yu, Qin Hua, Hanwen Hu, Dongqing Bao, Tianyu Lu, Jian Cao 0001, Guangtao Xue, Liping Zhang 0013, Gang Chen 0001 |
ASPLOS (1) | 17 |
| 2026 | FlashPS: Efficient Generative Image Editing with Mask-aware Caching and SchedulingabstractGenerative image editing using diffusion models has become a prevalent application in today's AI cloud services. In production environments, image editing typically involves a mask that specifies the regions of an image template to be edited. The use of mask provides direct control over the editing process and introduces sparsity in the model inference. In this paper, we present FlashPS, a system that efficiently serves image editing requests. The key insight behind FlashPS is that image editing only modifies the masked regions of image templates, while preserving the original content in the unmasked areas. Driven by this insight, FlashPS judiciously skips redundant computations associated with the unmask areas by reusing cached intermediate activations from previous inferences. To mitigate the high cache loading overhead, FlashPS employs a bubble-free pipeline scheme that overlaps computation with cache loading. Additionally, to reduce queuing latency in online serving while improving the GPU utilization, FlashPS proposes a novel continuous batching strategy for diffusion model serving, allowing newly arrived requests to join the running batch in just one step of denoising computation, without waiting for the entire batch to complete. As heterogenous masks induce imbalanced load, FlashPS also develops a load balancing strategy that takes into account the loads of both computation and cache loading. Collectively, FlashPS outperforms state-of-the-art diffusion serving systems for image editing, achieving up to 3× higher throughput and reducing average request latency by up to 14.7× while ensuring image quality. Xiaoxiao Jiang, Suyi Li 0002, Lingyun Yang, Tianyu Feng, Zhipeng Di, Weiyi Lu, Guoxuan Zhu, Xiu Lin, Yinghao Yu, Tao Lan, Lin Qu, Liping Zhang 0013, Wei Wang 0030 |
EuroSys | 14 |
| 2026 | AUM: Unleashing the Efficiency Potential of Shared Processors with Accelerator Units for LLM ServingabstractGenerative AI, especially LLM, is driving a fundamental shift in software paradigms, prompting cloud providers to build more efficient serving infrastructures. To meet the computational demands of emerging software, modern CPU processors are integrating Accelerator Units (AU) in the pipeline to accelerate key operations, such as Intel AMX for matrix multiplication. Current practices that dedicate AU-enabled CPU exclusively to LLM serving lead to significant resource waste and inferior efficiency. To this end, sharing AU-enabled CPU with general workloads is necessary to harvest redundant resources and improve platform performance-per-watt. However, perfectly sharing AU can be challenging since they introduce three-dimensional variations: variable usage patterns, compulsory frequency interferences, and dissimilar resource bounds. Existing resource managers are oblivious to complex Accelerator Unit Variations (AUV), resulting in performance and efficiency degradations of up to 50 % in shared environments. Therefore, this paper introduces AUM, a novel AU-aware resource manager designed to handle AUV and maximize the efficiency of shared processors. AUM has two cooperative components with three stages for three-dimensional AUV. The background profiler characterizes the usage, frequency, and resource information into a discrete model, guiding the runtime controller to analyze usage-aware requirements, select frequency-aware divisions, and make bound-aware resource decisions. Through extensive evaluations on production AU-enabled CPUs, we show that AUM improves CPU efficiency by$4.7-8.8 \%$while maintaining high-performance AU applications by reducing SLO violations by$\mathbf{7 - 1 1 \%}$compared with state-of-the-art resource managers. Xinkai Wang 0003, Chao Li 0009, Yiming Zhuansun, Jinyang Guo 0001, Xiaofeng Hou, Jing Wang 0055, Weigao Chen, Liping Zhang 0013, Minyi Guo |
HPCA | 11 |
| 2026 | PipeWeave: Synergizing Analytical and Learning Models for Unified GPU Performance Prediction
Yunfan Cui, Shuhao Zhang 0001, Chutong Ding, Shiyou Qian, Jian Cao 0001, Guangtao Xue, Liping Zhang 0013 |
ISCA | 11 |
| 2026 | Attack of the Bubbles: Straggler-Resilient Pipeline Parallelism for Large Model Training
Tianyuan Wu, Lunxi Cao, Hanfeng Lu, Xiaoxiao Jiang, Yinghao Yu, Siran Yang, Jiamang Wang, Lin Qu, Liping Zhang 0013, Wei Wang 0030 |
NSDI | 10 |
| 2025 | EXIST: Enabling Extremely Efficient Intra-Service Tracing Observability in DatacentersabstractThe complexity of online applications is rapidly increasing, bringing more sophisticated performance anomalies in today's cloud datacenter. To fully understand application behaviors, we should obtain both inter-service communication data via RPC-level tracing and intra-service execution traces via application-level tracing to precisely reason about event causality. However, the average time overhead of existing intra-service tracing schemes on the traced applications is generally about 5-10%, possibly reaching 18% in the worst case. To realize practical intra-service tracing in shared and stressed datacenters, one must achieve extreme tracing efficiency with an overhead at the per-mille level. Xinkai Wang 0003, Xiaofeng Hou, Chao Li 0009, Yuancheng Li 0001, Du Liu, Guoyao Xu, Liping Zhang 0013, Yuemin Wu, Xiaopeng Yuan, Quan Chen 0002, Minyi Guo |
ASPLOS (2) | 8 |
| 2025 | WDP: Mitigating Interference in CPU Sharing Through Wake-up Delay Driven Preemption for QoS-aware Co-locationabstractAs Latency-critical (LC) tasks often experience diurnal load patterns, co-locating them with best-effort (BE) tasks improves resource utilization. Prior work allocates entire CPU cores between co-located tasks, due to the incapability of handling the interference with CPU sharing. We observed that the root cause of the interference on the same core is the inherent wake-up delay in the operating system scheduler, the wait time that a process can obtain the CPU cycles after it is woken up. Based on the finding, we propose WDP, a scheme that efficiently improves the throughput of BE tasks while ensuring QoS, leveraging CPU sharing. WDP comprises a wake-up delay-driven preemption mechanism and a preemption-based CPU manager. The preemption mechanism enables controlled preemption to reduce the wake-up delay of LC tasks with adjustable preemption capacity. Adopting the novel preemption mechanism, the CPU manager allocates CPU resources in a fine-grained manner among co-located tasks. Compared with the representative prior method, WDP improves the throughput of BE tasks by 31.2% on average while ensuring the QoS of co-located LC tasks. Yaoxuan Li, Pu Pang, Yecheng Yang, Quan Chen 0002, Zhengxuan Yan, Guoyao Xu, Liping Zhang 0013, Minyi Guo |
SoCC | 8 |
| 2025 | Reducing the End-to-End Latency of DNN-Based Recommendation Systems in GPU PoolsabstractWhile intelligent applications (e.g., recommendation systems) prefer different CPU-GPU ratios, GPU pooling technique that decouples the GPU and CPU resources yields substantial flexibility when serving diverse applications. With such architecture, DNN-based recommendation services often offload the compute-intensive neural network layers to the remote GPU pool for high resource utilization. However, such a paradigm results in the long end-to-end latency due to two causes: 1) the intermediate data is copied for multiple times during the entire process in current GPU pooling practices, incurring heavy overheads; 2) the content transferred to the GPU pool involves multiple small tensors, suffering from poor bandwidth efficiency. To solve these problems, we design Zero, a runtime system that incorporates a zero-copy transmission mechanism as well as a dynamic tensor merging policy. The zero-copy transmission mechanism unifies memory management across the inference framework and the RPC framework, accompanied by an elaborated serialization protocol to fully eliminate redundant data copying. Meanwhile, the tensor merging policy deliberately organizes small tensors into larger data blocks, so as to transfer them with higher efficiency. Experimental results show that, compared with prior work, Zero reduces the latency of typical recommendation models by up to 15.1% (10.1% on average). Guangqiang Luan, Pu Pang, Quan Chen 0002, Chen Chen 0067, Guoyao Xu, Chi Zhang 0005, Yanyi Zi, Yinghao Yu, Liping Zhang 0013, Minyi Guo |
IPDPS | 10 |
| 2025 | GPU-Disaggregated Serving for Deep Learning Recommendation Models at Scale
Lingyun Yang, Yongchen Wang, Yinghao Yu, Qizhen Weng 0001, Jianbo Dong, Chi Zhang 0005, Yanyi Zi, Zechao Zhang, Menglei Zheng, Lanlan Xi, Binzhang Fu, Tao Lan, Liping Zhang 0013, Lin Qu, Wei Wang 0030 |
NSDI | 20 |
| 2025 | Katz: Efficient Workflow Serving for Diffusion Models with Many Adapters
Suyi Li 0002, Lingyun Yang, Xiaoxiao Jiang, Hanfeng Lu, Dakai An, Zhipeng Di, Weiyi Lu, Yinghao Yu, Tao Lan, Lin Qu, Liping Zhang 0013, Wei Wang 0030 |
USENIX ATC | 14 |
| 2025 | GREYHOUND: Hunting Fail-Slows in Hybrid-Parallel Training at Scale
Tianyuan Wu, Wei Wang 0030, Yinghao Yu, Siran Yang, Qinkai Duan, Jiamang Wang, Lin Qu, Liping Zhang 0013 |
USENIX ATC | 10 |
| 2024 | DeployFix: Dynamic Repair of Software Deployment Failures via Constraint SolvingabstractSoftware deployment misconfiguration often happens and has been one of the major causes of deployment failures that give rise to service interruptions. However, there is currently no existing approach to automatically repairing deployment failures. We propose DeployFix, which automatically repairs software deployment failures via constraint solving in the dynamic-changing deployment environments. DeployFix first defines DeployIR as a unified intermediate representation to achieve the translation of heterogeneous specifications from different schedulers with different syntaxes. By reducing the root-cause analysis of deployment failures to the conflict resolution in propositional logic, DeployFix uses off-the-shelf constraint solvers to achieve automatic localization and diagnosis of conflicting constraints, which are the root causes of deployment failures. DeployFix finally resolves the conflicting constraints and generates repaired deployment configurations in terms of practical requirements. We evaluate DeployFix in both simulation and production environments with tens of thousands of nodes at Alibaba, on which tens of thousands of applications are running guided by hundreds of thousands of deployment constraints. Experimental results demonstrate that DeployFix outperforms the state of the art and it correctly repairs the deployment failures in minutes, even in a large production data center. Haoyu Liao, Jianmei Guo, Bo Huang 0002, Yujie Han, Dingyu Yang, Kai Shi 0006, Jonathan Ding, Guoyao Xu, Liping Zhang 0013 |
ASE | 10 |
| 2024 | Optimizing Resource Management for Shared Microservices: A Scalable System DesignabstractA common approach to improving resource utilization in data centers is to adaptively provision resources based on the actual workload. One fundamental challenge of doing this in microservice management frameworks, however, is that different components of a service can exhibit significant differences in their impact on end-to-end performance. To make resource management more challenging, a single microservice can be shared by multiple online services that have diverse workload patterns and SLA requirements. We present an efficient resource management system, namely Erms, for guaranteeing SLAs with high probability in shared microservice environments. Erms profiles microservice latency as a piece-wise linear function of the workload, resource usage, and interference. Based on this profiling, Erms builds resource scaling models to optimally determine latency targets for microservices with complex dependencies. Erms also designs new scheduling policies at shared microservices to further enhance resource efficiency. Experiments across microservice benchmarks as well as trace-driven simulations demonstrate that Erms can reduce SLA violation probability by 5× and more importantly, lead to a reduction in resource usage by 1.6×, compared to state-of-the-art approaches. Shutian Luo, Chenyu Lin, Kejiang Ye, Guoyao Xu, Liping Zhang 0013, Huanle Xu, Cheng-Zhong Xu 0001 |
ACM Trans. Comput. Syst. | 5 |
| 2023 | Erms: Efficient Resource Management for Shared Microservices with SLA GuaranteesabstractA common approach to improving resource utilization in data centers is to adaptively provision resources based on the actual workload. One fundamental challenge of doing this in microservice management frameworks, however, is that different components of a service can exhibit significant differences in their impact on end-to-end performance. To make resource management more challenging, a single microservice can be shared by multiple online services that have diverse workload patterns and SLA requirements. Shutian Luo, Huanle Xu, Kejiang Ye, Guoyao Xu, Liping Zhang 0013, Jian He 0004, Cheng-Zhong Xu 0001 |
ASPLOS (1) | 5 |
| 2023 | Understanding and Optimizing Workloads for Unified Resource Management in Large Cloud PlatformsabstractTo fully utilize computing resources, cloud providers such as Google and Alibaba choose to co-locate online services with batch processing applications in their data centers. By implementing unified resource management policies, different types of complex computing jobs request resources in a consistent way, which can help data centers achieve global optimal scheduling and provide computing power with higher quality. To understand this new scheduling paradigm, in this paper, we first present an in-depth study of Alibaba's unified scheduling workloads. Our study focuses on the characterization of resource utilization, the application running performance, and scheduling scalability. We observe that although computing resources are significantly over-committed under unified scheduling, the resource utilization in Alibaba data centers is still low. In addition, existing resource usage predictors tend to make severe overestimations. At the same time, tasks within the same application behave fairly consistently, and the running performance of tasks can be well-profiled with respect to resource contention on the corresponding physical host. Chengzhi Lu, Huanle Xu, Kejiang Ye, Guoyao Xu, Liping Zhang 0013, Cheng-Zhong Xu 0001 |
EuroSys | 5 |
| 2023 | Beware of Fragmentation: Scheduling GPU-Sharing Workloads with Fragmentation Gradient Descent
Qizhen Weng 0001, Lingyun Yang, Yinghao Yu, Wei Wang 0030, Xiaochuan Tang, Liping Zhang 0013 |
USENIX ATC | 7 |
| 2022 | The power of prediction: microservice auto scaling via workload learningabstractWhen deploying microservices in production clusters, it is critical to automatically scale containers to improve cluster utilization and ensure service level agreements (SLA). Although reactive scaling approaches work well for monolithic architectures, they are not necessarily suitable for microservice frameworks due to the long delay caused by complex microservice call chains. In contrast, existing proactive approaches leverage end-to-end performance prediction for scaling, but cannot effectively handle microservice multiplexing and dynamic microservice dependencies. Shutian Luo, Huanle Xu, Kejiang Ye, Guoyao Xu, Liping Zhang 0013, Cheng-Zhong Xu 0001 |
SoCC | 5 |
| 2022 | Workload consolidation in alibaba clusters: the good, the bad, and the uglyabstractWeb companies typically run latency-critical long-running services and resource-intensive, throughput-hungry batch jobs in a shared cluster for improved utilization and reduced cost. Despite many recent studies on workload consolidation, the production practice remains largely unknown. This paper describes our efforts to efficiently consolidate the two types of workloads in Alibaba clusters to support the company's e-commerce businesses. Yongkang Zhang 0003, Yinghao Yu, Wei Wang 0030, Qiukai Chen, Tianchen Ding, Qizhen Weng 0001, Lingyun Yang, Jian He 0004, Liping Zhang 0013 |
SoCC | 14 |
| 2022 | Characterizing Job Microarchitectural Profiles at Scale: Dataset and AnalysisabstractUnderstanding the microarchitectural resource characteristics of datacenter jobs has become increasingly critical to guarantee the performance of jobs while improving resource utilization. Prior work studied the resource characteristics of datacenter jobs at the OS level, little reveals the deep and detailed characteristics at the microarchitecture level due to the lack of related open traces. In this paper, we provide a new open trace, AMTrace (Alibaba Microarchitecture Trace) 1, which is profiled from 8,577 high-end physical hosts from Alibaba’s datacenter by a hardware/software co-design monitoring method. AMTrace provides the microarchitectural metrics of 9.8 × 105 Linux containers with ”Per-Container-Per-Logic CPU” granularity. Different from existing open traces, AMTrace provides a new perspective to analyze the microarchitectural resource characteristics of datacenter jobs. Based on AMTrace, we first reveal the uneven resource usage of jobs among multiple logic CPUs. Then, we analyze the impact of resource contention of CPU and memory bandwidth on job performance. Finally, we analyze the job performance under different CPU provisioning modes from microarchitecture perspective. These analyses lead to constructive insights for datacenter resource management and optimization. Furthermore, we discuss possible research opportunities on AMTrace and we believe that AMTrace will inspire more exciting research on microarchitecture and resource management. Kangjin Wang, Ying Li 0012, Kingsum Chow, Yaoyong Dou, Guoyao Xu, Chuanjia Hou, Liping Zhang 0013 |
ICPP | 11 |
| 2022 | MLaaS in the Wild: Workload Analysis and Scheduling in Large-Scale Heterogeneous GPU Clusters
Qizhen Weng 0001, Wencong Xiao, Yinghao Yu, Wei Wang 0030, Jian He 0004, Yong Li 0045, Liping Zhang 0013, Wei Lin 0016 |
NSDI | 8 |
| 2022 | An In-Depth Study of Microservice Call Graph and Runtime PerformanceabstractLoosely-coupled and light-weight microservices running in containers are replacing monolithic applications gradually. Understanding the characteristics of microservices is critical to make good use of microservice architectures. However, there is no comprehensive study about microservice and its related systems in production environments so far. In this paper, we present a solid analysis of large-scale deployments of microservices at Alibaba clusters. Our study focuses on the characterization of microservice dependency as well as its runtime performance. We conduct an in-depth anatomy of microservice call graphs to quantify the difference between them and traditional DAGs of data-parallel jobs. In particular, we observe that microservice call graphs are heavy-tail distributed and their topology is similar to a tree and moreover, many microservices are hot-spots. We also discover that the structure of call graphs for long-term developed applications is much simpler so as to provide better performance. Our investigation on microservice runtime performance indicates most microservices are much more sensitive to CPU interference than memory interference. Moreover, we design resource management policies to efficiently tune memory resources. Shutian Luo, Huanle Xu, Chengzhi Lu, Kejiang Ye, Guoyao Xu, Liping Zhang 0013, Jian He 0004, Cheng-Zhong Xu 0001 |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2021 | Characterizing Microservice Dependency and Performance: Alibaba Trace AnalysisabstractLoosely-coupled and light-weight microservices running in containers are replacing monolithic applications gradually. Understanding the characteristics of microservices is critical to make good use of microservice architectures. However, there is no comprehensive study about microservice and its related systems in production environments so far. In this paper, we present a solid analysis of large-scale deployments of microservices at Alibaba clusters. Our study focuses on the characterization of microservice dependency as well as its runtime performance. We conduct an in-depth anatomy of microservice call graphs to quantify the difference between them and traditional DAGs of data-parallel jobs. In particular, we observe that microservice call graphs are heavy-tail distributed and their topology is similar to a tree and moreover, many microservices are hot-spots. We reveal three types of meaningful call dependency that can be utilized to optimize microservice designs. Our investigation on microservice runtime performance indicates most microservices are much more sensitive to CPU interference than memory interference. To synthesize more representative microservice traces, we build a mathematical model to simulate call graphs. Experimental results demonstrate our model can well preserve those graph properties observed from Alibaba traces. Shutian Luo, Huanle Xu, Chengzhi Lu, Kejiang Ye, Guoyao Xu, Liping Zhang 0013, Jian He 0004, Cheng-Zhong Xu 0001 |
SoCC | 6 |
| 2021 | Morphling: Fast, Near-Optimal Auto-Configuration for Cloud-Native Model ServingabstractMachine learning models are widely deployed in production cloud to provide online inference services. Efficiently deploying inference services requires careful tuning of hardware and runtime configurations (e.g., GPU type, GPU memory, batch size), which can significantly improve the model serving performance and reduce cost. However, existing autoconfiguration approaches for general workloads, such as Bayesian optimization and white-box prediction, are inefficient in navigating the high-dimensional configuration space of model serving, incurring high sampling cost. Lingyun Yang, Yinghao Yu, Wei Wang 0030, Bo Li 0001, Xianchao Sun, Jian He 0004, Liping Zhang 0013 |
SoCC | 8 |
| 2018 | All-Spark: Using Simulation Tests Directly in Production Environments to Detect System Bottlenecks in Large-Scale SystemsabstractWith the rapid growth in e-commerce, large-scale promotional activities have become a popular concept. However, when the existing system cannot be adjusted efficiently to adapt to the tremendous traffic in the promotion period, which is hundreds of times more than the volume on normal days, it be-comes a bottleneck that restricts the continuous growth of the online business. Traditional capacity prediction methods have been proven to be incapable of making accurate predictions for such special scenarios, because of a variety of unpredictable system bottlenecks. Simulation testing in a completely new test environment for such a large scale has a number of defects and limitations, such as the high cost of setting up the environment and the difficulty of testing the entire environment. Moreover, bottlenecks found in the test server may be different from those in the production server. We investigated online simulations in the production environment and built a complete simulation test system called All-Sparks. This solution solved a long-standing problem of simulation testing with large traffic in the production environment without causing any data pollution. The simulation test revealed hundreds of bottlenecks under a high workload pressure every year to eliminate the hidden problems caused by new applications. The final capacity evaluation result was deviated by less than 5% from the actual capacity, and the error rate was small (<2%); both of these are significant improvements over the traditional prediction results. This solution also provided a framework with good expansibility to multiple scenarios other than stress testing. Jialiang Lin 0002, Liping Zhang 0013, Yin Han |
Middleware | 4 |