Yuewen Wu

dblp:238/8878 · DBLP profile ↗
← Back
14ranked-venue papers
5as first author
10since 2021 · last 2025
0000-0002-1323-2455ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 2 first-author · 7 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Scheduling based on Block Features for Concurrent Inference with Unseen DNN Models on GPU
abstract
Efficiently scheduling concurrent deep neural network (DNN) inference on the same GPU can significantly optimize resource utilization. Such scheduling requires accurate prediction of concurrent inference time. Existing approaches primarily rely on model-level features for prediction and scheduling, which necessitates retraining and resampling when encountering unseen models to ensure prediction accuracy. However, in MLOps pipelines, the rapid iteration of models introduces numerous unseen models, making accurate predictions highly challenging and increasing the risk of SLA violations. To address these challenges posed by unseen models, we present SKADI, a scheduling framework based on block-level feature extraction and two-stage greedy scheduling. First, SKADI introduces block-level feature extraction, decomposing DNN models into homogeneous blocks (contiguous operator sequences) to enable zero-shot inference time prediction for unseen models. Second, it proposes a round-based and two-stage greedy scheduling strategy that rapidly selects optimal model pairs and overlaps their critical operators. Experimental results show that for unseen models, SKADI reduces the MAPE of concurrent inference time prediction by 59.25% and 57.55% compared to DeInfer and Abacus. Additionally, SKADI reduces SLA violation rates by 70.6% while increasing throughput by 7.3%.
Diaohan Luo, Heran Gao, Yuewen Wu, Heng Wu 0001, Wenbo Zhang 0006
ICPP4
2024 ETS: Deep Learning Training Iteration Time Prediction based on Execution Trace Sliding Window
abstract
Deep learning (DL) has become essential across various computer science domains. Accurately predicting iteration time for DL models in diverse cloud data center environments is critical for making high-quality scheduling decisions. Existing approaches neglect the sequential features inherent in the runtime execution, leading to issues such as overlooking DL framework overhead and struggling to handle diverse sizes of DL models, resulting in either low accuracy or slow convergence of the prediction model. This paper introduces ETS, a novel iteration time prediction method utilizing execution trace sliding windows. Our observation reveals that DL models exhibit a highly sequential runtime execution nature. Building upon this insight, we leverage sliding windows to extract a novel type of sequential features from the runtime execution trace. These features comprehensively capture DL framework overhead and address the diversity challenge in DL model sizes. By combining a best-practice method to train a prediction model, we achieve high accuracy and rapid convergence simultaneously. Experimental validation on over 14,000 DL model configurations demonstrates ETS's effectiveness in predicting the iteration time of DL models, achieving a mere 5.9% prediction error with a training time at the 10-minute level, and improving scheduling outcomes by reducing job completion time by 17%.
Heng Wu 0001, Yuewen Wu, Hua Zhong 0007, Wenbo Zhang 0006, Yan Liu 0102
HPDC4
2023 SPLIT: QoS-Aware DNN Inference on Shared GPU via Evenly-Sized Model Splitting
abstract
Improving QoS by simultaneously reducing the latency violation rate and jitter in the presence of multiple deep learning inference (DLI) tasks sharing a single edge computing processor remains a challenge. However, existing DLI systems at the edge, designed to maximize throughput, face performance challenges when confronted with requests with varying QoS.
Diaohan Luo, Yuewen Wu, Heng Wu 0001, Tao Wang 0030, Wenbo Zhang 0006
ICPP3
2023 Topology-Aware Self-Adaptive Resource Provisioning for Microservices
abstract
Microservice architecture is a popular technology for deploying services in cloud computing, with benefits like loose coupling, high fault tolerance, and scalability. The heterogeneous resource requirements and complex interaction relations have increased the difficulty in provisioning resources for microservices with intricacy topology. Existing approaches allocate resources for different microservices separately, and thus cannot achieve optimal global performance. Moreover, these approaches extract features from specific microservice topologies. We propose a topology-aware self-adaptive resource provisioning approach for microservices. Firstly, we propose a microservice state graph to characterize the status of each microservice in an application. Then, we use graph neural networks and attention to extract the resource requirements and correlation features of microservices. Thirdly, we use a reinforcement learning-based approach to allocate resources for microservices uniformly. Finally, we evaluate our approach by conducting a series of experiments on three typical microservice applications deployed in a heterogeneous cluster. The results show that our approach is efficient in extracting resource and correlation features of microservices, and can guarantee QoS with efficient resource utilization. Our approach can reduce the End-to-End latency by 22%, and can improve resource utilization by 18% with guaranteed latency.
Tao Wang 0030, Yuewen Wu, Heng Wu 0001, Wenbo Zhang 0006
ICWS4
2023 Hydra: Deadline-Aware and Efficiency-Oriented Scheduling for Deep Learning Jobs on Heterogeneous GPUs
abstract
With the rapid proliferation of deep learning (DL) jobs running on heterogeneous GPUs, scheduling DL jobs to meet various scheduling requirements, such as meeting deadlines and reducing job completion time (JCT), is critical. Unfortunately, existing efficiency-oriented and deadline-aware efforts are still rudimentary. They lack the capability of scheduling jobs to meet deadline requirements while reducing total JCT, especially when the jobs have various execution times on heterogeneous GPUs. Therefore, we present Hydra, a novel quantitative cost comparison approach, to address this scheduling issue. Here, the cost represents the total JCT plus a dynamic penalty calculated from the total tardiness (i.e., the delay time of exceeding the deadline) of all jobs. Hydra adopts a sampling approach that exploits the inherent iterative periodicity of DL jobs to estimate job execution times accurately on heterogeneous GPUs. Then, Hydra considers various combinations of job sequences and GPUs to obtain the minimized cost by leveraging an efficient branch-and-bound algorithm. Finally, the results of evaluation experiments on Alibaba traces show that Hydra can reduce total tardiness by 85.8% while reducing total JCT as much as possible, compared with state-of-the-art efforts.
Heng Wu 0001, Yuanjia Xu, Yuewen Wu, Hua Zhong 0001, Wenbo Zhang 0006
IEEE Trans. Computers4
2022 Serving unseen deep learning models with near-optimal configurations: a fast adaptive search approach
abstract
Public clouds provide a bewildering choice of configurations for Deep Learning (DL) models, and the choice of configuration will significantly impact the performance and budget. However, it is an obvious challenge to recommend a near-optimal configuration for a particular DL model from a wide range of candidates. The huge search overhead of finding such a configuration is the notorious cold start problem in state-of-the-art efforts, and this problem becomes more severe when they are faced with unseen DL models.
Yuewen Wu, Heng Wu 0001, Diaohan Luo, Yuanjia Xu, Wenbo Zhang 0006, Hua Zhong 0007
SoCC1
2021 Configurable In-Database Similarity Search of Electronic Medical Records
Yuewen Wu
WISA1
2021 Talos: A Weighted Speedup-Aware Device Placement of Deep Learning Models
abstract
Efficient device placement of deep learning (DL) models, which consist of many operations, is a big challenge when heterogeneous devices (e.g., CPU, GPU) are considered. Existing average speedup and transient speedup approaches do not make full use of operation-level speedups, and the Total Operation Completion Time (TOCT) cannot be optimized efficiently.To address this challenge, we present Talos, a weighted speedup-awareness approach to optimize device placement of multiple DL models. Talos reveals operations within or across DL models have diverse speedups (from 10−1to 102) on heterogeneous devices. In addition, the execution time of operations are widely ranged (from 0.1ms to 100ms). Talos considers the two features simultaneously as weighted speedups, and treats them as costs in an incremental minimum-cost flow. Compared with state-of-the-art efforts, experiment results show that Talos can reduce TOCT by up to 50%.
Yuanjia Xu, Heng Wu 0001, Wenbo Zhang 0006, Yuewen Wu, Heran Gao, Tao Wang 0030
ASAP5
2021 Best VM Selection for Big Data Applications across Multiple Frameworks by Transfer Learning
abstract
Cloud providers are presented with a bewildering choice of VM types for a range of contemporary data processing frameworks today. However, existing performance modeling and machine learning efforts cannot pick optimal VM types for multiple frameworks simultaneously, since they are difficult to balance model accuracy and model training cost.
Yuewen Wu, Heng Wu 0001, Yuanjia Xu, Wenbo Zhang 0006, Hua Zhong 0007, Tao Huang 0001
ICPP1
2021 Apollo: Rapidly Picking the Optimal Cloud Configurations for Big Data Analytics Using a Data-Driven Approach
Yuewen Wu, Yuanjia Xu, Heng Wu 0001, Lin-Gang Su, Wenbo Zhang 0006, Hua Zhong 0007
J. Comput. Sci. Technol.1
2020 Nuka: A Generic Engine with Millisecond Initialization for Serverless Computing
abstract
Serverless computing is becoming one of the mainstream trends in cloud computing due to its advantages of simplified programming and cost saving. However, existing serverless platforms still adopt Docker container as its execution engine, which has the cold start problem and causes high common-case invocation latency. In this work, we analyze the lifecycle of common-case serverless invocation on existing serverless platforms and find that current container startup and pulling remote images are the two main reasons causing cold start so slow. Based on the study, we implement Nuka, a generic engine with millisecond initialization for serverless computing. Nuka is fully compatible with Docker interface and can smoothly re-place Docker as the execution engine of existing serverless plat-forms. Through the isolation pool that reuses Linux's isolation configurations, Nuka avoids the high cost of container startup's scalability bottleneck, which reduces container's startup time with high concurrency scale. Nuka also avoids pulling remote images through dynamically resolving and importing required software packages from local package caching. A self-adaptive container reuse strategy dynamically controls container's pause time and replica numbers, which effectively reduces the frequency of cold start. Compared with Docker, Nuka can get a millisecond initialization with high concurrency and significantly reduces average time cost of cold startup by 6× on existing serverless platforms.
Shijun Qin, Heng Wu 0001, Yuewen Wu, Yuanjia Xu, Wenbo Zhang 0006
JCC3
2020 A framework to support multi-cloud collaboration
abstract
With the rapid development of cloud computing, major cloud providers have launched various cloud services with different functions to meet customer's needs. Therefore, flexibility is extremely important when developers use these cloud services. However, APIs of cloud services change dozens of times annually without backward compatibility. It means developers have to adapt these clouds with manual efforts. Such efforts make the multi-cloud collaboration extremely complex and cannot meet the demand of flexibility. This paper describes a configuration-based multi-cloud collaboration framework, which can support new clouds with comprehensible configurations. Meanwhile, if cloud APIs are updated without backward compatibility, it can restore services during runtime with minimized configuration. The main technologies used in this article include automatic discovery, unified abstraction, dynamic mapping and incremental update. We tested the virtual machine and container services of seven well-known cloud providers. The system can support heterogeneous clouds well. When the APIs are updated, the system can restore services in less than 200 milliseconds. At the same time, the extra cost of our framework is acceptable to cloud users.
Ting Tang, Heng Wu 0001, Yuewen Wu, Yuanjia Xu, Wenbo Zhang 0006
SERVICES4
2019 Dockerfile TF Smell Detection Based on Dynamic and Static Analysis Methods
abstract
Dockerfile is used to build docker image. In the image building process, temporary files are frequently used to import applications and data. A careless use of Dockerfile may cause temporary file left in the image, which can increase the image size, thus effects the elastic scale ability and QoS. This problem is identified as temporary file smell. The feature of UnionFS that docker image used is different from traditional filesystems. If users are not paying attention, they are too apt to make such mistakes. To address this problem, we propose two different methods to detect temporary file smell with dynamic analysis and static analysis respectively. We use the really-world cases to evaluate our methods. Experimental results show that our methods can effectively detect the temporary file smell.
Yuewen Wu, Zhigang Lu 0004, Tao Wang 0030
COMPSAC (1)2
2018 HW3C: A Heuristic based Workload Classification and Cloud Configuration Approach for Big Data Analytics
abstract
It is a big challenge to pick up the best cloud configuration for recurring big data analytics jobs running in clouds. Prior efforts may get in a sub-optimal configuration due to a broad spectrum of cloud configurations with a few test runs, such as CherryPick. We present HW3C which is a heuristic based workload classification and cloud configuration system for big data analytics jobs, our insight is classifying a job by comparing its resource preference and usage informantion with other jobs, and then using heuristic rules to distinguish bad samples from good ones in Bayesian Optimization algorithm. Our experiments on HiBench and SparkBench in Aliyun ECS show that the performance of job had been improved by 53% in average comparing with CherryPick, meanwhile the resource cost had been reduced by 40% in average.
Yuewen Wu, Heng Wu 0001, Wenbo Zhang 0006, Yuanjia Xu, Jun Wei 0001, Hua Zhong 0001
Internetware1