VLDB 2026 Research / reviewers in the wild / expert
Yilun Wang 0001
dblp:80/7412-1
· DBLP profile ↗
9ranked-venue papers
1as first author
7since 2021 · last 2025
0000-0001-9262-8652ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 4 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 3 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Mint: Cost-Efficient Tracing with All Requests Collection via Commonality and Variability AnalysisabstractDistributed traces contain valuable information but are often massive in volume, posing a core challenge in tracing framework design: balancing the tradeoff between preserving essential trace information and reducing trace volume. To address this tradeoff, previous approaches typically used a '1 or 0' sampling strategy: retaining sampled traces while completely discarding unsampled ones. However, based on an empirical study on real-world production traces, we discover that the '1 or 0' strategy actually fails to effectively balance this tradeoff. Haiyu Huang 0002, Cheng Chen 0056, Kunyi Chen, Pengfei Chen 0002, Guangba Yu, Yilun Wang 0001, Huxing Zhang, Qi Zhou 0001 |
ASPLOS (1) | 7 |
| 2025 | Defragmentation Scheduling with Deep Reinforcement Learning in Shared GPU ClustersabstractModern GPU clusters in computing centers face significant challenges in resource utilization due to fragmentation caused by GPU-sharing mechanisms, job diversity and asynchronous job lifecycles. Existing methods fail to address GPU fragmentation in dynamic scheduling scenarios under GPU sharing. To tackle this issue, this paper proposes DRR, a defragmentation scheduler with deep reinforcement learning (DRL) and rescheduling to mitigate GPU fragmentation. DRR employs a DRL agent trained via imitation learning from heuristic algorithms to overcome cold-start issues, and further enhanced by multi-scale policy optimization for balanced exploration and exploitation to reduce GPU fragmentation. Additionally, the rescheduling strategy in DRR further optimizes GPU utilization by relocating the running jobs. Evaluations conducted on the physical Kubernetes-based testbed and large-scale simulated clusters demonstrate that DRR reduces the average GPU fragmentation rate by 50% compared to state-of-the-art methods, while maintaining Quality of Service (QoS) and ensuring fairness for users. Qingfu Wu, Pengfei Chen 0002, Yilun Wang 0001 |
SoCC | 3 |
| 2025 | LLMConf: Knowledge-Enhanced Configuration Optimization for Large Language Model InferenceabstractAs large language models (LLMs) are widely applied across various domains, improving the quality of LLM inference services is essential. In this paper, we find that optimizing configuration parameters of LLM inference engines can significantly improve LLM inference performance in terms of latency and throughput. Therefore, we propose LLMConf, an automated performance tuning system that optimizes multiple LLM inference performance metrics by searching for the optimal configuration parameters of the LLM inference engine. We first introduce a knowledge-enhanced approach to identify the set of configuration parameters (LLMConfigs) that most significantly impact LLM performance from the adjustable parameters provided by the LLM inference engine. We then perform automated data collection to build functional relationships between LLMConfigs and each performance metric. Additionally, LLMConf employs a multi-objective optimization module to obtain optimal LLMConfigs for simultaneously optimizing multiple performance metrics. The experimental results show that LLMConf significantly outperforms existing methods. Compared to the default configuration parameters of the LLM inference engine, LLMConf achieves an average improvement of 20.1% across 7 key performance metrics. Moreover, experiments demonstrate that LLMConf has strong transferability across diverse datasets, varying concurrency levels and different LLM base models. Jingkai He, Pengfei Chen 0002, Yilun Wang 0001, Haiyu Huang 0002, Chuanfu Zhang, Haojia Huang, Danwen Chen |
IWQoS | 3 |
| 2024 | FaaSRCA: Full Lifecycle Root Cause Analysis for Serverless ApplicationsabstractServerless becomes popular as a novel computing paradigms for cloud native services. However, the complexity and dynamic nature of serverless applications present significant challenges to ensure system availability and performance. There are many root cause analysis (RCA) methods for microservice systems, but they are not suitable for precise modeling serverless applications. This is because: (1) Compared to microservice, serverless applications exhibit a highly dynamic nature. They have short lifecycle and only generate instantaneous pulse-like data, lacking long-term continuous information. (2) Existing methods solely focus on analyzing the running stage and overlook other stages, failing to encompass the entire lifecycle of serverless applications. To address these limitations, we propose FaaSRCA, a full lifecycle root cause analysis method for serverless applications. It integrates multi-modal observability data generated from platform and application side by using Global Call Graph. We train a Graph Attention Network (GAT) based graph autoencoder to compute reconstruction scores for the nodes in global call graph. Based on the scores, we determine the root cause at the granularity of the lifecycle stage of serverless functions. We conduct experimental evaluations on two serverless benchmarks, the results show that FaaSRCA outperforms other baseline methods with a top-k precision improvement ranging from 21.25% to 81.63%. Pengfei Chen 0002, Guangba Yu, Yilun Wang 0001, Haiyu Huang 0002 |
ISSRE | 4 |
| 2024 | FaaSConf: QoS-aware Hybrid Resources Configuration for Serverless WorkflowsabstractServerless computing, also known as Function-as-a-Service (FaaS), is a significant development trend in modern software system architecture. The workflow composition of multiple short-lived functions has emerged as a prominent pattern in FaaS, exposing a considerable resources configuration challenge compared to individual independent serverless functions. This challenge unfolds in two ways. Firstly, workflows frequently encounter dynamic and concurrent user workloads, increasing the risk of QoS violations. Secondly, the performance of a function can be affected by the resource reprovision of other functions within the workflow. Yilun Wang 0001, Pengfei Chen 0002, Yiwen Zhang 0001, Guangba Yu, Haiyu Huang 0002 |
ASE | 1 |
| 2024 | DeepCAT+: A Low-Cost and Transferrable Online Configuration Auto-Tuning Approach for Big Data FrameworksabstractBig data frameworks usually provide a large number of performance-related parameters. Online auto-tuning these parameters based on deep reinforcement learning (DRL) to achieve a better performance has shown their advantages over search-based and machine learning-based approaches. Unfortunately, the time cost during the online tuning phase of conventional DRL-based methods is still heavy, especially for Big Data applications. Therefore, in this paper, we propose DeepCAT$^+$, a low-cost and transferrable deep reinforcement learning-based approach to achieve online configuration auto-tuning for Big Data frameworks. To reduce the total online tuning cost and increase the adaptability: 1) DeepCAT$^+$utilizes the TD3 algorithm instead of DDPG to alleviate value overestimation; 2) DeepCAT$^+$modifies the conventional experience replay to fully utilize the rare but valuable transitions via a novel reward-driven prioritized experience replay mechanism; 3) DeepCAT$^+$designs a Twin-Q Optimizer to estimate the execution time of each action without the costly configuration evaluation and optimize the sub-optimal ones to achieve a low-cost exploration-exploitation tradeoff; 4) Furthermore, DeepCAT$^+$also implements an Online Continual Learner module based on Progressive Neural Networks to transfer knowledge from historical tuning experiences. Experimental results based on a lab Spark cluster with HiBench benchmark applications show that DeepCAT$^+$is able to speed up the best execution time by a factor of 1.49×, 1.63× and 1.65× on average respectively over the baselines, while consuming up to 50.08%, 53.39% and 70.79% less total tuning time. In addition, DeepCAT$^+$also has a strong adaptability to the time-varying environment of Big Data frameworks. Yilun Wang 0001, Yiwen Zhang 0001, Pengfei Chen 0002, Zibin Zheng |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2022 | DeepCAT: A Cost-Efficient Online Configuration Auto-Tuning Approach for Big Data FrameworksabstractTo support different application scenarios, big data frameworks usually provide a large number of performance-related configuration parameters. Online auto-tuning these parameters based on deep reinforcement learning to achieve a better performance has shown their advantages over search-based and machine learning-based approaches. Unfortunately, the time consumption during the online tuning phase of conventional DRL-based methods is still heavy, especially for big data applications. Therefore, in this paper, we propose DeepCAT, a cost-efficient deep reinforcement learning-based approach to achieve online configuration auto-tuning for big data frameworks. To reduce the total online tuning cost: 1) DeepCAT utilizes the TD3 algorithm instead of DDPG to alleviate value overestimation; 2) DeepCAT modifies the conventional experience replay to fully utilize the rare but valuable transitions via a novel reward-driven prioritized experience replay mechanism; 3) DeepCAT designs a Twin-Q Optimizer to estimate the execution time of each action without the costly configuration evaluation and optimize the sub-optimal ones to achieve a low-cost exploration-exploitation trade off. Experimental results based on a local 3-node Spark cluster and HiBench benchmark applications show that DeepCAT is able to speed up the best execution time by a factor of 1.45 × and 1.65 × on average respectively over CDBTune and OtterTune, while consuming up to 50.08% and 53.39% less total tuning time. Yilun Wang 0001, Yiwen Zhang 0001, Pengfei Chen 0002 |
ICPP | 2 |
| 2014 | Data Augmented Maximum Margin Matrix Factorization for Flickr Group Recommendation
Liang Chen 0001, Yilun Wang 0001, Tingting Liang, Lichuan Ji, Jian Wu 0001 |
PAKDD (1) | 2 |
| 2013 | WT-LDA: User Tagging Augmented LDA for Web Service Clustering
Liang Chen 0001, Yilun Wang 0001, Qi Yu 0001, Zibin Zheng, Jian Wu 0001 |
ICSOC | 2 |