Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Wujie Shao

dblp:221/4689 · DBLP profile ↗
← Back
2ranked-venue papers
1as first author
0since 2021 · last 2019
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Cloud and datacenter computing · 61% Performance modeling and evaluation · 30% Distributed systems · 9%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Cloud and datacenter computing
big data analytics
0.412019
Cost-Effective Cloud Server Provisioning for Predictable Performance of Big Data Analytics · IEEE Trans. Parallel Distributed Syst. 2019
Performance modeling and evaluation
performance prediction
0.412019
Cost-Effective Cloud Server Provisioning for Predictable Performance of Big Data Analytics · IEEE Trans. Parallel Distributed Syst. 2019
Cloud and datacenter computing › resource provisioning
server provisioning
0.412019
Cost-Effective Cloud Server Provisioning for Predictable Performance of Big Data Analytics · IEEE Trans. Parallel Distributed Syst. 2019
Distributed systems › fault tolerance
checkpointing
0.112019
Cost-Effective Cloud Server Provisioning for Predictable Performance of Big Data Analytics · IEEE Trans. Parallel Distributed Syst. 2019

Methods — techniques the papers use, named apart from their topics

price prediction · 0.4performance modeling · 0.4LSTM · 0.4
YearPublicationVenuePosition
2019 Stage Delay Scheduling: Speeding up DAG-style Data Analytics Jobs with Resource Interleaving
abstract
To increase the resource utilization of datacenters, big data analytics jobs are commonly running stages in parallel which are organized into and scheduled according to the Directed Acyclic Graph (DAG). Through an in-depth analysis of the latest Alibaba cluster trace and our motivation experiments on Amazon EC2, however, we show that the CPU and network resources are still under-utilized due to the unwise stage scheduling, thereby prolonging the completion time of a DAG-style job (e.g., Spark). While existing works on reducing the job completion time focus on either task scheduling or job scheduling, stage scheduling has received comparably little attention. In this paper, we design and implement DelayStage, a simple yet effective stage delay scheduling strategy to interleave the cluster resources across the parallel stages, so as to increase the cluster resource utilization and speed up the job performance. With the aim of minimizing the makespan of parallel stages, DelayStage judiciously arranges the execution of stages in a pipelined manner to maximize the performance benefits of resource interleaving. Extensive prototype experiments on 30 Amazon EC2 instances and complementary trace-driven simulations show that DelayStage can improve the cluster resource utilization by up to 81.8% and reduce the job completion time by up to 41.3%, in comparison to the stock Spark and the state-of-the-art stage scheduling strategies, yet with acceptable runtime overhead.
Wujie Shao, Fei Xu 0009, Li Chen 0019, Haoyue Zheng, Fangming Liu
ICPP1
2019 Cost-Effective Cloud Server Provisioning for Predictable Performance of Big Data Analytics
abstract
Cloud datacenters are underutilized due to server over-provisioning. To increase datacenter utilization, cloud providers offer users an option to run workloads such as big data analytics on the underutilized resources, in the form of cheap yet revocable transient servers (e.g., EC2 spot instances, GCE preemptible instances). Though at highly reduced prices, deploying big data analytics on the unstable cloud transient servers can severely degrade the job performance due to instance revocations. To tackle this issue, this paper proposes iSpot, a cost-effective transient server provisioning framework for achieving predictable performance in the cloud, by focusing on Spark as a representative Directed Acyclic Graph (DAG)-style big data analytics workload. It first identifies the stable cloud transient servers during the job execution by devising an accurate Long Short-Term Memory (LSTM)-based price prediction method. Leveraging automatic job profiling and the acquired DAG information of stages, we further build an analytical performance model and present a lightweight critical data checkpointing mechanism for Spark, to enable our design of iSpot provisioning strategy for guaranteeing the job performance on stable transient servers. Extensive prototype experiments on both EC2 spot instances and GCE preemptible instances demonstrate that, iSpot is able to guarantee the performance of big data analytics running on cloud transient servers while reducing the job budget by up to 83.8 percent in comparison to the state-of-the-art server provisioning strategies, yet with acceptable runtime overhead.
Fei Xu 0009, Haoyue Zheng, Wujie Shao, Haikun Liu, Zhi Zhou 0006
IEEE Trans. Parallel Distributed Syst.4