VLDB 2026 Research / reviewers in the wild / expert
Guangshun Yao
dblp:62/105
· DBLP profile ↗
13ranked-venue papers
5as first author
7since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 2 first-author · 2 since 2021Systems, architecture and hardware · 3 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 3 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Graph-Based Batch Job Load Balancing Scheduling for Multi-Dimensional Resources in Heterogeneous GPU ClustersabstractGPU clusters serve as a cornerstone of high-performance computing and support a wide range of batch jobs with complex resource demands. However, the diverse requirements of batch jobs and resource heterogeneity present significant challenges to efficient scheduling. Existing approaches either rely on static rules or overlook the interdependencies among virtual machines introduced by resource heterogeneity, making it difficult to address the diverse resource demands and dynamic load balancing. In this paper, we propose a novel scheduling model based on Graph Neural Networks (GNNs) and Double Deep Q-Networks (DDQNs), termed GNN-DDQN, for batch job load balancing and scheduling of multi-dimensional resources (e.g. GPU, CPU and memory) in heterogeneous clusters. We propose a system model that integrates batch job resource requests, multi-dimensional resource configurations, and a multi-objective optimization framework. The scheduling problem is formulated as a Markov Decision Process based on this model. A GNN is employed to effectively capture the interdependencies among virtual machines, while a DDQN optimizes scheduling decisions using a dynamic target network update mechanism. Extensive experiments are conducted using two real-world Alibaba cluster traces. Results demonstrate the effectiveness and generalization capabilities of the proposed scheduling model. Compared to baseline methods, the results confirm its superiority in load balancing, job latency, fairness, and time efficiency. Yumei Shi, Shiping Chen 0002, Guangshun Yao, Shengxiang Wang |
IEEE Trans. Computers | 4 |
| 2026 | A Hybrid Fault-Tolerant Workflow Scheduling With Performance Fluctuated Cloud ResourcesabstractWith the increasing complexity of cloud systems, resource performance fluctuation and failure have become two significant factors that affect task execution in the cloud, particularly for workflow tasks with precedence constraints. The former often leads to uncertain task execution time, while the latter can result in task abortion. Although many cloud workflow scheduling algorithms have been proposed for fault tolerance or uncertain task execution time separately, these two issues simultaneously exist in practice and have not yet been thoroughly explored. This paper proposes a hybrid fault-tolerant cloud workflow scheduling algorithm designed to study the failures of cloud resources and uncertain task execution time simultaneously. The proposed scheduling process consists of four phases: preprocessing, initial scheduling, online scheduling, and online adjustment. Two fundamental and widely recognized strategies for fault tolerance, replication and resubmission, are integrated into the algorithm. An elastic resource provisioning mechanism is also designed to adjust active resources dynamically. Performance evaluations on both randomly generated and real-world workflows demonstrate that the algorithm effectively schedules deadline-constrained workflows under uncertain task execution time while guaranteeing fault tolerance. Qian Ren, Guangshun Yao |
IEEE Trans. Serv. Comput. | 2 |
| 2025 | AdaGap: An adaptive gap-aware resource allocation strategy for GPU sharing in heterogeneous clusters
Shiping Chen 0004, Yumei Shi, Guangshun Yao |
Future Gener. Comput. Syst. | 4 |
| 2023 | Input-to-state Stabilization of Delayed Semi-Markovian Jump Neural Networks Via Sampled-Data Control
Wenhuang Wu, Guangshun Yao, Jianping Zhou 0003 |
Neural Process. Lett. | 3 |
| 2023 | Failure-Aware Elastic Cloud Workflow SchedulingabstractWith an increasing complexity and functionality in cloud data centers, fault tolerance becomes an essential requirement for tasks executed in clouds, especially for workflows with task precedences. Hosts and network devices are the main physical components in a cloud data center. The PB (Primary-Backup) model is a desirable approach to fault tolerance. Many PB-based workflow scheduling algorithms have been proposed for host faults. However, only a few studies focus on cloud workflow scheduling considering network device faults. This paper analyzes the fault-tolerant properties for scheduling dependent tasks and migrating VMs based on the PB model, considering both host and network device faults in a cloud data center. A failure-aware elastic cloud workflow scheduling algorithm is designed for both host and network device fault tolerance. Additionally, an elastic resource provisioning mechanism is proposed and incorporated into the proposed algorithm to improve resource utilization. Performance evaluations on both randomly generated and real-world workflows show that the proposal effectively improves resource utilization while guaranteeing fault tolerance. Guangshun Yao, Xiaoping Li 0001, Qian Ren, Rubén Ruiz |
IEEE Trans. Serv. Comput. | 1 |
| 2022 | A Hybrid Fault-Tolerant Scheduling for Deadline-Constrained Tasks in Cloud SystemsabstractAmong multiple fault-tolerant strategies, resubmission, and replication are fundamental and widely recognized in distributed computing systems. In recent years, many algorithms based on replication or resubmission have been proposed. However, few of them consider these two techniques together, especially in Cloud systems. In this article, we propose a Hybrid Fault-Tolerant Scheduling Algorithm (HFTSA) for independent tasks with deadlines by integrating the above techniques in virtualized Cloud systems. During the task scheduling process, HFTSA selects fault-tolerant strategies from resubmission and replication for each accepted task based on the characteristics of both task and Cloud resources and then reserves suitable resources. During the task execution process, HFTSA adopts an online adjustment scheme for fault-tolerant strategies of some tasks if necessary while providing an online scheduling scheme for faults. Moreover, an elastic resource provisioning mechanism is designed and incorporated into HFTSA to dynamically adjust the provided resources to improve resource utilization. Experiments on a real cloud platform and a simulated platform are conducted to verify the effectiveness of the proposed HFTSA. The results demonstrate that HFTSA can provide an efficient fault-tolerant scheduling strategy for deadline-constrained tasks with high resource utilization and performs better than corresponding competitors. Guangshun Yao, Qian Ren, Xiaoping Li 0001, Rubén Ruiz |
IEEE Trans. Serv. Comput. | 1 |
| 2021 | Towards an Efficient Framework for Data Extraction from Chart Images
Weihong Ma, Hesuo Zhang, Shuang Yan, Guangshun Yao, Yichao Huang, Yaqiang Wu |
ICDAR (1) | 4 |
| 2017 | Fault-tolerant elastic scheduling algorithm for workflow in Cloud systems
Yongsheng Ding, Guangshun Yao, Kuangrong Hao |
Inf. Sci. | 2 |
| 2017 | Endocrine-based coevolutionary multi-swarm for multi-objective workflow scheduling in a cloud system
Guangshun Yao, Yongsheng Ding, Yaochu Jin, Kuangrong Hao |
Soft Comput. | 1 |
| 2017 | An improved immune system-inspired routing recovery scheme for energy harvesting wireless sensor networks
Xiangfei Zhang, Guangshun Yao, Yongsheng Ding, Kuangrong Hao |
Soft Comput. | 2 |
| 2017 | Using Imbalance Characteristic for Fault-Tolerant Workflow Scheduling in Cloud SystemsabstractResubmission and replication are two fundamental and widely recognized techniques in distributed computing systems for fault tolerance. The resubmission based strategy has an advantage in resource utilization, while the replication based strategy can reduce the task completed time in the context of fault. However, few researches take these two techniques together for fault-tolerant workflow scheduling, especially in Cloud systems. In this paper, we present a novel fault-tolerant workflow scheduling (ICFWS) algorithm for Cloud systems by combining the aforementioned two strategies together to play their respective advantages for fault tolerance while trying to meet the soft deadline of workflow. First, it divides the soft deadline of workflow into multiple sub-deadlines for all tasks. Then, it selects a reasonable fault-tolerant strategy and reserves suitable resource for each task by taking the imbalance sub-deadlines among tasks and on-demand resource provisioning of Cloud systems into consideration. Finally, an online scheduling and reservation adjustment scheme is designed to select a suitable resource for the task with resubmission strategy and adjust the sub-deadlines as well as fault-tolerant strategies of some unexecuted tasks during the task execution process, respectively. The proposed algorithm is evaluated on both real-world and randomly generated workflows. The results demonstrate that the ICFWS outperforms some well-known approaches on corresponding metrics. Guangshun Yao, Yongsheng Ding, Kuangrong Hao |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2016 | An adaptive clustering routing algorithm for energy harvesting-wireless sensor networksabstractIn order to address the influence caused by unstable and uneven harvested energy among sensor nodes for clustering routing in energy harvesting wireless sensor networks (EH-WSNs), a novel clustering routing algorithm (CREW) is proposed in this work. The CREW is composed by cluster building phase and data transmission phase. In cluster building phase, the CREW uses two new concepts, Network Gradient and Waiting Time for cluster head Competition, to divide the network into unequal clusters and select the cluster heads based on residual energy and energy gain of nodes, respectively. In data transmission phase, it adopts an adaptive inter-cluster communication mechanism, to sufficiently store and utilize the harvesting energy. Finally, to verify the effectiveness of the proposed CREW, a series of experiments are conducted and compared with other related clustering routings for the EH-WSNs. Simulation results manifest that the CREW can provide effective clustering routing for the EH-WSNs and highlight the better performance of the proposed approach than that of similar techniques. Xiangfei Zhang, Yongsheng Ding, Guangshun Yao, Kuangrong Hao |
CEC | 3 |
| 2016 | An immune system-inspired rescheduling algorithm for workflow in Cloud systems
Guangshun Yao, Yongsheng Ding, Lihong Ren, Kuangrong Hao, Lei Chen 0064 |
Knowl. Based Syst. | 1 |