EDBT 2026 Demo / reviewers in the wild / expert
Yuanjia Xu
dblp:238/8881
· DBLP profile ↗
10ranked-venue papers
2as first author
6since 2021 · last 2023
0000-0003-0939-2381ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 2 first-author · 5 since 2021Software engineering, systems software and programming languages · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Hydra: Deadline-Aware and Efficiency-Oriented Scheduling for Deep Learning Jobs on Heterogeneous GPUsabstractWith the rapid proliferation of deep learning (DL) jobs running on heterogeneous GPUs, scheduling DL jobs to meet various scheduling requirements, such as meeting deadlines and reducing job completion time (JCT), is critical. Unfortunately, existing efficiency-oriented and deadline-aware efforts are still rudimentary. They lack the capability of scheduling jobs to meet deadline requirements while reducing total JCT, especially when the jobs have various execution times on heterogeneous GPUs. Therefore, we present Hydra, a novel quantitative cost comparison approach, to address this scheduling issue. Here, the cost represents the total JCT plus a dynamic penalty calculated from the total tardiness (i.e., the delay time of exceeding the deadline) of all jobs. Hydra adopts a sampling approach that exploits the inherent iterative periodicity of DL jobs to estimate job execution times accurately on heterogeneous GPUs. Then, Hydra considers various combinations of job sequences and GPUs to obtain the minimized cost by leveraging an efficient branch-and-bound algorithm. Finally, the results of evaluation experiments on Alibaba traces show that Hydra can reduce total tardiness by 85.8% while reducing total JCT as much as possible, compared with state-of-the-art efforts. Heng Wu 0001, Yuanjia Xu, Yuewen Wu, Hua Zhong 0001, Wenbo Zhang 0006 |
IEEE Trans. Computers | 3 |
| 2022 | Serving unseen deep learning models with near-optimal configurations: a fast adaptive search approachabstractPublic clouds provide a bewildering choice of configurations for Deep Learning (DL) models, and the choice of configuration will significantly impact the performance and budget. However, it is an obvious challenge to recommend a near-optimal configuration for a particular DL model from a wide range of candidates. The huge search overhead of finding such a configuration is the notorious cold start problem in state-of-the-art efforts, and this problem becomes more severe when they are faced with unseen DL models. Yuewen Wu, Heng Wu 0001, Diaohan Luo, Yuanjia Xu, Wenbo Zhang 0006, Hua Zhong 0007 |
SoCC | 4 |
| 2022 | EOP: efficient operator partition for deep learning inference over edge serversabstractRecently, Deep Learning (DL) models have demonstrated great success for its attractive ability of high accuracy used in artificial intelligence Internet of Things applications. A common deployment solution is to run such DL inference tasks on edge servers. In a DL inference, each operator takes tensors as input and run in a tensor virtual machine, which isolates resource usage among operators. Nevertheless, existing edge-based DL inference approaches can not efficiently use heterogeneous resources (e.g., CPU and low-end GPU) on edge servers and result in sub-optimal DL inference performance, since they can only partition operators in a DL inference with equal or fixed ratios. It is still a big challenge to support partition optimizations over edge servers for a wide range of DL models, such as Convolution Neural Network (CNN), Recurrent Neural Network (RNN) and Transformers. Yuanjia Xu, Heng Wu 0001, Wenbo Zhang 0006 |
VEE | 1 |
| 2021 | Talos: A Weighted Speedup-Aware Device Placement of Deep Learning ModelsabstractEfficient device placement of deep learning (DL) models, which consist of many operations, is a big challenge when heterogeneous devices (e.g., CPU, GPU) are considered. Existing average speedup and transient speedup approaches do not make full use of operation-level speedups, and the Total Operation Completion Time (TOCT) cannot be optimized efficiently.To address this challenge, we present Talos, a weighted speedup-awareness approach to optimize device placement of multiple DL models. Talos reveals operations within or across DL models have diverse speedups (from 10−1to 102) on heterogeneous devices. In addition, the execution time of operations are widely ranged (from 0.1ms to 100ms). Talos considers the two features simultaneously as weighted speedups, and treats them as costs in an incremental minimum-cost flow. Compared with state-of-the-art efforts, experiment results show that Talos can reduce TOCT by up to 50%. Yuanjia Xu, Heng Wu 0001, Wenbo Zhang 0006, Yuewen Wu, Heran Gao, Tao Wang 0030 |
ASAP | 1 |
| 2021 | Best VM Selection for Big Data Applications across Multiple Frameworks by Transfer LearningabstractCloud providers are presented with a bewildering choice of VM types for a range of contemporary data processing frameworks today. However, existing performance modeling and machine learning efforts cannot pick optimal VM types for multiple frameworks simultaneously, since they are difficult to balance model accuracy and model training cost. Yuewen Wu, Heng Wu 0001, Yuanjia Xu, Wenbo Zhang 0006, Hua Zhong 0007, Tao Huang 0001 |
ICPP | 3 |
| 2021 | Apollo: Rapidly Picking the Optimal Cloud Configurations for Big Data Analytics Using a Data-Driven Approach
Yuewen Wu, Yuanjia Xu, Heng Wu 0001, Lin-Gang Su, Wenbo Zhang 0006, Hua Zhong 0007 |
J. Comput. Sci. Technol. | 2 |
| 2020 | Nuka: A Generic Engine with Millisecond Initialization for Serverless ComputingabstractServerless computing is becoming one of the mainstream trends in cloud computing due to its advantages of simplified programming and cost saving. However, existing serverless platforms still adopt Docker container as its execution engine, which has the cold start problem and causes high common-case invocation latency. In this work, we analyze the lifecycle of common-case serverless invocation on existing serverless platforms and find that current container startup and pulling remote images are the two main reasons causing cold start so slow. Based on the study, we implement Nuka, a generic engine with millisecond initialization for serverless computing. Nuka is fully compatible with Docker interface and can smoothly re-place Docker as the execution engine of existing serverless plat-forms. Through the isolation pool that reuses Linux's isolation configurations, Nuka avoids the high cost of container startup's scalability bottleneck, which reduces container's startup time with high concurrency scale. Nuka also avoids pulling remote images through dynamically resolving and importing required software packages from local package caching. A self-adaptive container reuse strategy dynamically controls container's pause time and replica numbers, which effectively reduces the frequency of cold start. Compared with Docker, Nuka can get a millisecond initialization with high concurrency and significantly reduces average time cost of cold startup by 6× on existing serverless platforms. Shijun Qin, Heng Wu 0001, Yuewen Wu, Yuanjia Xu, Wenbo Zhang 0006 |
JCC | 5 |
| 2020 | A framework to support multi-cloud collaborationabstractWith the rapid development of cloud computing, major cloud providers have launched various cloud services with different functions to meet customer's needs. Therefore, flexibility is extremely important when developers use these cloud services. However, APIs of cloud services change dozens of times annually without backward compatibility. It means developers have to adapt these clouds with manual efforts. Such efforts make the multi-cloud collaboration extremely complex and cannot meet the demand of flexibility. This paper describes a configuration-based multi-cloud collaboration framework, which can support new clouds with comprehensible configurations. Meanwhile, if cloud APIs are updated without backward compatibility, it can restore services during runtime with minimized configuration. The main technologies used in this article include automatic discovery, unified abstraction, dynamic mapping and incremental update. We tested the virtual machine and container services of seven well-known cloud providers. The system can support heterogeneous clouds well. When the APIs are updated, the system can restore services in less than 200 milliseconds. At the same time, the extra cost of our framework is acceptable to cloud users. Ting Tang, Heng Wu 0001, Yuewen Wu, Yuanjia Xu, Wenbo Zhang 0006 |
SERVICES | 6 |
| 2019 | Aladdin: Optimized Maximum Flow Management for Shared Production ClustersabstractThe rise in popularity of long-lived applications (LLAs), such as deep learning and latency-sensitive online Web services, has brought new challenges for cluster schedulers in shared production environments. Scheduling LLAs needs to support complex placement constraints (e.g., to run multiple containers of an application on different machines) and larger degrees of parallelism to provide global optimization. But existing schedulers usually suffer severe constraint violations, high latency and low resource efficiency. This paper describes Aladdin, a novel cluster scheduler that can maximize resource efficiency while avoiding constraint violations: (i) it proposes a multidimensional and nonlinear capacity function to support constraint expressions; (ii) it applies an optimized maximum flow algorithm to improve resource efficiency. Experiments with an Alibaba workload trace from a 10,000-machine cluster show that Aladdin can reduce violated constraints by as mush as 20%. Meanwhile, it improves resource efficiency by 50% compared with state-of-the-art schedulers. Heng Wu 0001, Wenbo Zhang 0006, Yuanjia Xu, Tao Huang 0001, Haiyang Ding |
IPDPS | 3 |
| 2018 | HW3C: A Heuristic based Workload Classification and Cloud Configuration Approach for Big Data AnalyticsabstractIt is a big challenge to pick up the best cloud configuration for recurring big data analytics jobs running in clouds. Prior efforts may get in a sub-optimal configuration due to a broad spectrum of cloud configurations with a few test runs, such as CherryPick. We present HW3C which is a heuristic based workload classification and cloud configuration system for big data analytics jobs, our insight is classifying a job by comparing its resource preference and usage informantion with other jobs, and then using heuristic rules to distinguish bad samples from good ones in Bayesian Optimization algorithm. Our experiments on HiBench and SparkBench in Aliyun ECS show that the performance of job had been improved by 53% in average comparing with CherryPick, meanwhile the resource cost had been reduced by 40% in average. Yuewen Wu, Heng Wu 0001, Wenbo Zhang 0006, Yuanjia Xu, Jun Wei 0001, Hua Zhong 0001 |
Internetware | 4 |