Zhiyao Hu

dblp:167/1434 · DBLP profile ↗
← Back
11ranked-venue papers
5as first author
4since 2021 · last 2026
0000-0001-9863-2919ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 6 · 3 first-authorArtificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
Cloud and datacenter computing · 72% Parallel and multicore computing · 24% Distributed systems · 4%
Computer networks
1 paper
Datacenter networks · 100%

Topics — the 7 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Cloud and datacenter computing › cluster resource management and scheduling
cluster resource management
0.922021
Optimizing Resource Allocation for Data-Parallel Jobs Via GCN-Based Prediction · IEEE Trans. Parallel Distributed Syst. 2021
ReLoca: Optimize Resource Allocation for Data-parallel Jobs using Deep Learning · INFOCOM 2020
Cloud and datacenter computing
cluster resource management and scheduling
0.512021
Coordinative Scheduling of Computation and Communication in Data-Parallel Systems · IEEE Trans. Computers 2021
Cloud and datacenter computing › cluster resource management and scheduling
coflow scheduling
0.512021
Coordinative Scheduling of Computation and Communication in Data-Parallel Systems · IEEE Trans. Computers 2021
Parallel and multicore computing › data parallelism
data-parallel systems
0.512021
Coordinative Scheduling of Computation and Communication in Data-Parallel Systems · IEEE Trans. Computers 2021
Cloud and datacenter computing
job scheduling
0.512021
Coordinative Scheduling of Computation and Communication in Data-Parallel Systems · IEEE Trans. Computers 2021
Parallel and multicore computing › task scheduling › communication-aware scheduling
joint communication and computation scheduling
0.512021
Coordinative Scheduling of Computation and Communication in Data-Parallel Systems · IEEE Trans. Computers 2021
Cloud and datacenter computing
resource allocation
0.512021
Optimizing Resource Allocation for Data-Parallel Jobs Via GCN-Based Prediction · IEEE Trans. Parallel Distributed Syst. 2021

Methods — techniques the papers use, named apart from their topics

urgency-based scheduling · 1.0simulation · 1.0adaptive sampling · 0.9transfer learning · 0.5graph convolutional network · 0.5fully connected network · 0.5deep neural network · 0.4
YearPublicationVenuePosition
2026 FedBal: a self-adaptive and semi-asynchronous federated learning framework against the imbalance of client updates
Zhiyao Hu
Neural Comput. Appl.1
2023 Outranking-based failure mode and effects analysis considering interactions between risk factors and its application to food cold chain management
Huchang Liao, Zhiyao Hu, Zhiying Zhang 0001, Ming Tang 0003, Audrius Banaitis
Eng. Appl. Artif. Intell.2
2021 Coordinative Scheduling of Computation and Communication in Data-Parallel Systems
abstract
For many data-parallel computing systems like Spark, a job usually consists of multiple computation stages and inter-stage communication (i.e., coflows). Many efforts have been done to schedule coflows and jobs independently. The simple combination of coflow scheduling and job scheduling, however, would prolong the average job completion time (JCT) due to the conflict. For this reason, we propose a new abstraction of scheduling unit, named coBranch, which takes the dependency between computation stages and coflows into consideration, to schedule coflows and jobs jointly. Besides, mainstream coflow schedulers are order-preserving, i.e., all coflows of a high-priority job are prioritized than those of a low-priority job. We observe that the order-preserving constraint incurs low inter-job parallelism. To overcome the problem, we employ an urgency-based mechanism to schedule coBranches, which aims to decrease the average JCT by enhancing the inter-job parallelism. We implement the urgency-based coBranch Scheduling (BS) method on Apache Spark, conduct prototype-based experiments, and evaluate the performance of our method against the shortest-job-first critical-path method and the FIFO method. Results show that our method achieves around 10 and 15 percent reduction in the average JCT, respectively. Large-scale simulations based on the Google trace show that our method performs better and reduces JCT by 23 and 35 percent, respectively.
Dongsheng Li 0001, Zhiyao Hu, Zhiquan Lai, Yiming Zhang 0003, Kai Lu 0001
IEEE Trans. Computers2
2021 Optimizing Resource Allocation for Data-Parallel Jobs Via GCN-Based Prediction
abstract
Under-allocating or over-allocating computation resources (e.g., CPU cores) can prolong the completion time of data-parallel jobs in a distributed system. We present a predictor, ReLocag, to find the near-optimal number of CPU cores to minimize job completion time (JCT). ReLocag includes a graph convolutional network (GCN) and a fully-connected network (FCNN). The GCN learns the dependency between operations from the workflow of a job, and then the FCNN takes the workflow dependency together with other features (e.g., the input size, the number of CPU cores, the amount of memory, and the number of computation tasks) as input for JCT prediction. The prediction result can guide the user to determine the near-optimal number of CPU cores. Besides, we propose two effective strategies to overcome the time-consuming issue of training sample collection in big data applications. First, we develop an adaptive sampling method to collect essential samples judiciously. Second, we further design a cross-application transfer learning model to exploit the training samples collected from other applications. We conduct extensive experiments in a Spark cluster for 7 types of exemplary Spark applications. Results show that ReLocag improves the JCT prediction accuracy by 4-14 percent. Moreover, the CPU core consumption decreases by 58.2 percent.
Zhiyao Hu, Dongsheng Li 0001, Dongxiang Zhang, Yiming Zhang 0003, Baoyun Peng
IEEE Trans. Parallel Distributed Syst.1
2020 ReLoca: Optimize Resource Allocation for Data-parallel Jobs using Deep Learning
abstract
Since under-allocating computation resource (e.g., CPU cores) causes suboptimal JCTs of data-parallel jobs, users are inclined to request excessive computation resource to decrease JCTs. However, over-allocating computation resource for data-parallel jobs incurs considerable system overheads (e.g., network communication and disk I/O overhead), which prolong the job completion time. In this paper, we propose ReLoca towards the optimal allocation of computation resource with the objective of minimizing the job completion time. ReLoca employs a deep neural network to guide the allocation of computation resource, by learning the impact of the operations in data-parallel jobs on the system overhead and computation time. Since training samples are time-consuming to collect, we develop an adaptive sampling method to preferably collect high-quality samples and thus overcome the issue of data scarcity. We apply ReLoca to improve Spark and conduct real experiments with five typical applications in big data analytics. Results show that ReLoca significantly reduces the average job completion time. Compared with the state-of-the-art method, ReLoca has higher prediction accuracy, needs fewer training samples and decreases the sampling overhead. With the prediction by ReLoca, the JCT decreases by 29.85%.
Zhiyao Hu, Dongsheng Li 0001, Dongxiang Zhang, Yixin Chen 0004
INFOCOM1
2020 Uncertain multicast under dynamic behaviors
Yudong Qin, Deke Guo, Zhiyao Hu, Bangbang Ren
Frontiers Comput. Sci.3
2019 Branch scheduling: dag-aware scheduling for speeding up data-parallel jobs
abstract
A data-parallel job is characterized as a directed acyclic graph (DAG) which usually consists of multiple computation stages and across-stage data transfers. However, classical DAG scheduling strategies like the critical path method ignore other stages off the main path and do not give specific consideration of data locality, transfer costs, etc. In practice, complicated DAGs include multiple paths which overlap with each other. The intersection of different paths in a DAG job forms an important synchronization. The synchronization significantly impacts the job completion time and however, no known scheduling methods are designed to speed up the completion time of the synchronization especially when complicated DAGs involve nested and hierarchical synchronizations. To the end, we propose a new abstraction, named branch, which is referred to a disjoint path in DAGs, and design a branch scheduling method to decrease the average job completion time of multiple data-parallel jobs. The branch scheduling method leverages the urgency of branches to speed up the synchronization of multiple parallel branches. We have implemented the BS method on Apache Spark and conducted prototype-based experiments. Compared with Spark FIFO and the shortest-job-first with the critical path methods, results show that the branch scheduling method achieves around 10-15% reduction in the average job completion time.
Zhiyao Hu, Dongsheng Li 0001, Yiming Zhang 0003, Deke Guo, Ziyang Li 0003
IWQoS1
2017 Reliable multicast routing with uncertain sources
abstract
Multicast can jointly utilize the network resources when delivering the same content to a set of destinations; hence, it can effectively reduce the consumption of network resources more than individual unicast. The source of a multicast, however, does not need to be in a specific location as long as certain constraints are satisfied. This means the multicast can have uncertain sources, which could reduce the network bandwidth consumption more than a traditional multicast. Meanwhile, a reliable multicast becomes crucial in providing reliable services for many important applications. However, a prior minimal cost forest (MCF) for such a new multicast is not designed to support reliable transmissions. In this paper, we propose a novel reliable multicast routing with uncertain sources named ReMUS. To the best of our knowledge, we are the first to study the reliable multicast under uncertain sources. The goal is to minimize the sum of the transfer cost and the recovery cost, although finding such a ReMUS is very challenging. Thus, we design a sourcebased multicast method to solve this problem by exploiting the flexibility of uncertain sources when no recovery nodes exist in the network. Furthermore, we design a general multicast method to jointly exploit the benefits of uncertain sources and recovery nodes to minimize the total cost of ReMUS. We conduct extensive evaluations based on the real topology of Internet2. The results indicate that our methods can efficiently realize the reliable and bandwidth-efficient multicast with uncertain sources, irrespective of the settings of networks and multicasts.
Deke Guo, Zhiyao Hu, Jie Wu 0001, Tao Chen 0013, Honghui Chen
IWQoS3
2017 Source selection problem in multi-source multi-destination multicasting
Deke Guo, Xiaoqiang Teng, Zhiyao Hu, Bangbang Ren
Comput. Networks3
2016 Multicast routing with uncertain sources in software-defined network
abstract
Multicast is designed to jointly deliver content from a single source to a set of destinations. It can efficiently save the bandwidth consumption and reduce the load on the source. The appearance of SDN provides opportunities to deploy flexible protocols, including multicast and its variants. However, in many important applications, it is not necessary that the source of a multicast transfer has to be in specific location as long as certain constraints are satisfied. Such facts bring a novel multicast with uncertain sources, abbreviated as uncertain multicast. It brings new opportunities and challenges to reduce the bandwidth consumption. In this paper, we focus on the uncertain multicast and construct a forest with the minimum cost (MCF), to enable that each destination reaches to one and only one source. Prior approaches, relying on traditional multicast, remain inapplicable to the MCF problem. Therefore, we propose two (2+ε)-approximation methods, named P-MCF and E-MCF, which can be deployed in SDN controllers. We conduct experiments on our SDN testbed together with large-scale simulations under the random SDN network. All manifest that our MCF approach always occupies less network links and incurs less network cost for an uncertain multicast than the traditional Steiner minimum tree (SMT) of any related multicast, irrespective of the used network topology and the setting of multicast transmissions.
Zhiyao Hu, Deke Guo, Bangbang Ren
IWQoS1
2015 Control plane of software defined networks: A survey
Deke Guo, Zhiyao Hu, Ting Qu 0003
Comput. Commun.3