Wenyan Chen 0001

dblp:48/7090-1 · also Wen-Yan Chen 0001 · DBLP profile ↗
← Back
11ranked-venue papers
6as first author
8since 2021 · last 2026
0000-0001-8949-0816ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 5 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 High Throughput and Low Latency LLM Serving via Adaptive KV Caching
abstract
The substantial memory demands of model weights and key-value (KV) caches often lead to severe memory bottlenecks in LLM serving. Existing systems address this by offloading KV caches to host memory and rapidly restoring them on demand before decoding. However, these approaches are too coarse-grained and fail to fully exploit the combined computational and storage capabilities of GPUs.
Wenyan Chen 0001, Chengzhi Lu, Huanle Xu, Kejiang Ye, Cheng-Zhong Xu 0001
EuroSys1
2025 FedDance: Efficient Participant Selection for Federated Learning in Highly Dynamic Environments
abstract
Federated Learning (FL) is a rising distributed learning paradigm that facilitates multiple devices to jointly train a shared model. Given the presence of heterogeneous devices with distinct data distributions, it is critical to select an optimal subset of devices for engagement in the collaborative training process. However, the dynamic nature of FL, encompassing aspects like dynamic device availability and inherent training dynamics, significantly complicates participant selection, and current systems routinely fall short in adapting effectively to such dynamic environments.
Yuanhang Chen, Xiaosong Chen, Wenyan Chen 0001, Huanle Xu
SoCC3
2025 Rock: Serving Multimodal Models in Cloud with Heterogeneous-Aware Resource Orchestration for Thousands of LoRA Adapters
abstract
In this paper, we present ROCK, a novel system for efficiently serving thousands of LoRA adapters for multimodal models in cloud environments. Through extensive analysis of production workloads, we identify key challenges in current cloud-based image generation services: extreme request burstiness (up to$90 \times$normal rates), heterogeneous task characteristics, and inefficient adapter management that wastes 40 % of GPU memory and increases delays by$3 x$during peak times. ROCK addresses these challenges through a three-layer architecture that decouples hardware, adapters, and requests. Our system features dynamic heterogeneous queues that match tasks to appropriate resources based on multidimensional feature vectors, and a multilevel orchestration framework that intelligently manages adapter placement across heterogeneous storage. Experiments on a 64-GPU testbed demonstrate that ROCK reduces average response latency by$16-26 \%$, and achieves an 84.1 % cache hit rate for LoRA adapters-outperforming traditional approaches while reducing adapter update frequency by up to 77 %.
Shuaipeng Wu, Yanying Lin, Wenyan Chen 0001, Chong Ma 0005, Cheng-Zhong Xu 0001, Kejiang Ye
CLUSTER4
2025 Multiplexing Dynamic Deep Learning Workloads with SLO-awareness in GPU Clusters
abstract
Deep learning (DL) inference services are widely recognized as crucial workloads in large-scale cloud clusters. However, due to the stringent latency requirements, cloud providers often over-provision GPU resources, resulting in underutilization of the available GPU potential. Although co-locating tasks on the same device can enhance utilization, ensuring Service Level Objectives (SLOs) guarantees for multiplexing highly dynamic inference services becomes extremely challenging due to significant resource interference.
Wenyan Chen 0001, Chengzhi Lu, Huanle Xu, Kejiang Ye, Cheng-Zhong Xu 0001
EuroSys1
2024 EINS: Edge-Cloud Deep Model Inference with Network-Efficiency Schedule in Serverless
abstract
Model inference in edge is often regarded as an effective method to alleviate high latency and enhance data privacy in edge-cloud collaborative computing environment. In this paper, we demonstrate that optimizing network communication in edge-cloud environment with limited bandwidth can enhance model inference performance. We first analyze network bottlenecks and the characteristics in model inference, then design a serverless inference system - EINS, to support collaborative optimization of network transmission and inference performance in edge-cloud environment. This system identifies concurrent network communication bottlenecks in multi-model deployment, dynamically scales capacity, and optimizes placement strategies and model transfer sequences. Real-world workload evaluation reveal that EINS can achieve a 5.7x throughput improvement and an average reduction of 62% latency in model instance startup.
Yanying Lin, Wenyan Chen 0001, Yingfei Tang, Xu Duan, Kejiang Ye
CSCWD3
2024 SMIless: Serving DAG-based Inference with Dynamic Invocations under Serverless Computing
abstract
The deployment of ML serving applications, featuring multiple inference functions on serverless platforms, has gained substantial popularity, leading to numerous developments of new systems. However, these systems often focus on optimizing resource provisioning and cold start management separately, ultimately resulting in higher monetary costs. This paper introduces SMIless, a highly efficient serverless system tailored for serving DAG-based ML inference in heterogeneous environments. SMIless effectively co-optimizes resource configuration and cold-start management in the context of dynamic invocations. This is achieved by seamlessly integrating adaptive pre-warming windows, striking an effective balance between performance and cost. We have implemented SMIless on top of OpenFaaS and conducted extensive evaluations using real-world ML serving applications. The experimental results demonstrate that SMIless can achieve up to a $5.73 \times$ reduction in the overall costs while meeting the SLA requirements for all user requests, surpassing the performance of state-of-the-art solutions.
Chengzhi Lu, Huanle Xu, Yudan Li, Wenyan Chen 0001, Kejiang Ye, Cheng-Zhong Xu 0001
SC4
2023 Interference-aware Multiplexing for Deep Learning in GPU Clusters: A Middleware Approach
abstract
A common strategy for improving efficiency in training deep learning entails multiplexing tasks on a single GPU. To mitigate the interference caused by multiplexing, existing approaches primarily employ kernel-level solutions to regulate GPU kernel execution, or harness hardware-level techniques to explicitly restrict GPU streaming multiprocessors and memory. Nevertheless, none of them perform satisfactorily in optimizing the completion time of tasks.
Wenyan Chen 0001, Zizhao Mo, Huanle Xu, Kejiang Ye, Cheng-Zhong Xu 0001
SC1
2021 RPTCN: Resource Prediction for High-dynamic Workloads in Clouds based on Deep Learning
abstract
Resource management is challenging in clouds due to the dynamics and sharing characteristics. The crucial problem is how to allocate resources accurately and satisfy demands of workloads timely. The traditional solution is to use historical data to predict future resource usage. Although these resource prediction methods can predict the periodicity, they can not accurately predict mutation points due to the high dynamics and uncertainty of resource usage. To tackle this issue, in this paper we propose a resource usage prediction method - RPTCN, which is based on a deep learning method - temporal convolutional networks (TCNs) in cloud systems. We add a fully connected layer and attention mechanism to TCNs to improve the prediction accuracy. In order to explore the relationship between the usage of different resources in the temporal dimension, we use correlation analysis to screen performance indicators as multidimensional feature input for prediction. Finally, we evaluate the performance of this method on Alibaba trace v2018. Evaluations show that RPTCN improves the overall MAE and MSE by 6.50%~89.03% and 0.41%~68.82% respectively compared to baselines in dynamic and long-term prediction of resource usage. Moreover, the convergence and generalization of RPTCN are also better than the baselines.
Wenyan Chen 0001, Chengzhi Lu, Kejiang Ye, Yang Wang 0006, Cheng-Zhong Xu 0001
CLUSTER1
2020 Interference Analysis of Co-Located Container Workloads: A Perspective from Hardware Performance Counters
Wenyan Chen 0001, Kejiang Ye, Chengzhi Lu, Dongdai Zhou, Cheng-Zhong Xu 0001
J. Comput. Sci. Technol.1
2019 ADGS: Anomaly Detection and Localization Based on Graph Similarity in Container-Based Clouds
abstract
Docker container is experiencing rapid development with the support from the industry like Google and Alibaba and is being widely used in large scale production cloud environment. For example, Alibaba has deployed millions of containers for its internal business, and most of the online services are already migrated to the containers. Those services are usually very complex, spanning multiple containers with complex interaction and dependency relationship. Detecting potential anomalies in such a large container-based cloud platform is very challenging. Traditional detection models usually use system resource metrics like CPU and memory usage, but rarely consider the relationship among components, causing high false positive rate. In this paper, we present a novel Anomaly Detection and root cause localization method based on Graph Similarity (ADGS) in the container-based cloud environment. We first monitor the response time and resource usage of each component in the application to determine whether the system status is normal or not. Then, we propose a new mechanism to locate the root cause of the anomalies based on graph similarity, investigating the anomaly propagation rules among cluster components. We implement and evaluate our method in a container-based environment. The results show that the proposed method can detect and determine the root cause of anomalies efficiently and accurately.
Chengzhi Lu, Kejiang Ye, Wenyan Chen 0001, Cheng-Zhong Xu 0001
ICPADS3
2018 How Does the Workload Look Like in Production Cloud? Analysis and Clustering of Workloads on Alibaba Cluster Trace
abstract
Cloud computing technology is widely used in today's datacenters due to the benefits such as high scalability, on-demand services and low cost. An in-depth understanding of the characteristics of workloads running in production cloud environments is very important for improving the resource management efficiency. In this paper, we make a detailed analysis with visualization techniques and clustering methods on the trace dataset released by Alibaba which contains 11089 online services and 12951 batch jobs running on 1313 machines. Our methodology for clustering workloads contains: i) Select effective feature vectors as the dimensions of clustering; ii) Identify the cluster boundaries of each dimension using K-Means algorithm; iii) Classify jobs by combining the feature vectors which uses the results from previous step; iv) Analyze the characteristics of workload groups at runtime. Our analysis reveals several insights which previous work has not found on Alibaba cluster trace. For batch jobs: a) Average CPU cores of all batch jobs show bimodal-distribution obviously. b) At a random sampling time, more than 50 % machines only run one group of jobs with a short duration, medium CPU cores and small memory utilization, the remaining machines run mixed groups of jobs. For online instances: a) The resource usage (CPU, Memory, and Disk) of most online instances is low; b) There are up to six groups running on the same machine according to our clustering method at a random sampling time.
Wenyan Chen 0001, Kejiang Ye, Yang Wang 0006, Guoyao Xu, Cheng-Zhong Xu 0001
ICPADS1