Ji Wang 0002

dblp:64/856-2 · DBLP profile ↗
← Back
50ranked-venue papers
8as first author
33since 2021 · last 2026
0000-0002-4199-2793ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 2 first-author · 13 since 2021Systems, architecture and hardware · 14 · 5 first-author · 4 since 2021Databases, data management, data science and information retrieval · 9 · 1 first-author · 6 since 2021Computer networks · 6 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 PurMM: Attention-Guided Test-Time Backdoor Purification in Multimodal Large Language Models
abstract
Downstream fine-tuning of Multimodal Large Language Models (MLLMs) is advancing rapidly, allowing general models to achieve superior performance on domain-specific tasks. Yet most prior research focuses on performance gains and overlooks the vulnerability of the fine-tuning pipeline: attackers can easily poison the dataset to implant backdoors into MLLMs. We conduct an in-depth investigation of backdoor attacks on MLLMs and reveal the phenomenon of Attention Hijacking and its Hierarchical Mechanism. Guided by this insight, we propose PurMM, a test-time backdoor purification framework that removes visual tokens exhibiting anomalous attention, thereby avoiding targeted outputs while restoring correct answers. PurMM contains three stages: (1) locating tokens with abnormal attention, (2) filtering them using deep-layer cues, and (3) zeroing out their corresponding components in the visual embeddings. Unlike existing defences, PurMM dispenses with retraining and training-process modifications, operating at test-time to restore model performance while eliminating the backdoor. Extensive experiments across multiple MLLMs and datasets show that PurMM maintains normal performance, sharply reduces attack success rates, and consistently converts backdoor outputs to benign ones, offering a new perspective for safeguarding MLLMs.
Wenzheng Jiang, Ke Liang 0006, Xuankun Rong, Jingxuan Zhou, Zhengyi Zhong, Guancheng Wan, Ji Wang 0002
AAAI7
2026 Unveiling and Mitigating Untargeted Poisoning Attacks on Federated Knowledge Graph Embedding
Wenzheng Jiang, Ke Liang 0006, Wenke Huang 0003, Xiongtao Zhang, Guancheng Wan, Cheston Tan, Flint Xiaofeng Fan, Ji Wang 0002
WWW9
2026 Dynamic demand-aware UAV scheduling for IoT data collection using deep reinforcement learning approach
Xiaoqing Li 0006, Weidong Bao 0001, Qingbao Liu, Ji Wang 0002, Xiaomin Zhu 0001
Future Gener. Comput. Syst.6
2025 Yes is Harder than No: A Behavioral Study of Framing Effects in Large Language Models Across Downstream Tasks
abstract
Framing effect is a well-known cognitive bias in which individuals' responses to the same underlying question vary depending on how the question is phrased. Recent studies suggest that large language models (LLMs) also exhibit framing effects, but existing work has primarily replicated psychological experiments using hand-crafted prompts, leaving their impact on practical downstream tasks underexplored. To fill in the gap, in this paper, we conduct a systematic empirical investigation into framing effects in LLMs across multiple real-world downstream tasks. We construct semantically equivalent prompts with positive and negative framings and evaluate a wide range of LLMs under these conditions. We uncover several behavioral regularities of framing effects in LLMs, among which the most notable one is a consistent response asymmetry: LLMs find answering ''yes'' harder than ''no''. That is, LLMs tend to issue affirmative responses (i.e., ''yes'') only when they are highly confident, while they incline to answer negatively (i.e., ''no'') under uncertainty. We interpret this asymmetry through the lens of Error Management Theory (EMT), which posits that rational agents adopt risk-averse strategies to minimize the more costly error. We empirically show that this behavior is partially attributable to a statistical imbalance in the frequency of positive versus negative framing cues in pretraining corpora. Furthermore, we demonstrate that the framing-induced bias in LLMs can inform prompt engineering and active in-context learning, i.e., using framing-sensitive samples as demonstrations can improve model performance. Finally, we offer a preliminary strategy to mitigate the framing effect, i.e., injecting debiasing instructions, which shows promise. In all, our work uncovers a fundamental behavioral bias in LLMs and offers practical guidance for their reliable deployment across downstream tasks.
Weixin Zeng, Jiuyang Tang, Ji Wang 0002, Xiang Zhao 0002
CIKM4
2025 Unlearning through Knowledge Overwriting: Reversible Federated Unlearning via Selective Sparse Adapter
abstract
Federated Learning is a promising paradigm for privacy-preserving collaborative model training. In practice, it is essential not only to continuously train the model to acquire new knowledge but also to guarantee old knowledge the right to be forgotten (i.e., federated unlearning), especially for privacy-sensitive information or harmful knowledge. However, current federated unlearning methods face several challenges, including indiscriminate unlearning of cross-client knowledge, irreversibility of unlearning, and significant unlearning costs. To this end, we propose a method named FUSED, which first identifies critical layers by analyzing each layer’s sensitivity to knowledge and constructs sparse unlearning adapters for sensitive ones. Then, the adapters are trained without altering the original parameters, overwriting the unlearning knowledge with the remaining knowledge. This knowledge overwriting process enables FUSED to mitigate the effects of indiscriminate unlearning. Moreover, the introduction of independent adapters makes unlearning reversible and significantly reduces the unlearning costs. Finally, extensive experiments on three datasets across various unlearning scenarios demonstrate that FUSED’s effectiveness is comparable to Retraining, surpassing all other baselines while greatly reducing unlearning costs.
Zhengyi Zhong, Weidong Bao 0001, Ji Wang 0002, Shuai Zhang 0004, Jingxuan Zhou, Lingjuan Lyu, Wei Yang Bryan Lim
CVPR3
2025 CUT: Pruning Pre-trained Multi-task Models into Compact Models for Edge Devices
Jingxuan Zhou, Weidong Bao 0001, Ji Wang 0002, Zhengyi Zhong
ICIC (17)3
2025 FedHPD: Heterogeneous Federated Reinforcement Learning via Policy Distillation
Wenzheng Jiang, Ji Wang 0002, Xiongtao Zhang, Weidong Bao 0001, Cheston Tan, Flint Xiaofeng Fan
AAMAS2
2025 Gains: Fine-grained Federated Domain Adaptation in Open Set
abstract
Conventional federated learning (FL) assumes a closed world with a fixed total number of clients. In contrast, new clients continuously join the FL process in real-world scenarios, introducing new knowledge. This raises two critical demands: detecting new knowledge, i.e., knowledge discovery, and integrating it into the global model, i.e., knowledge adaptation. Existing research focuses on coarse-grained knowledge discovery, and often sacrifices source domain performance and adaptation efficiency. To this end, we propose a fine-grained federated domain adaptation approach in open set (Gains). Gains splits the model into an encoder and a classifier, empirically revealing features extracted by the encoder are sensitive to domain shifts while classifier parameters are sensitive to class increments. Based on this, we develop fine-grained knowledge discovery and contribution-driven aggregation techniques to identify and incorporate new knowledge. Additionally, an anti-forgetting mechanism is designed to preserve source domain performance, ensuring balanced adaptation. Experimental results on multi-domain datasets across three typical data-shift scenarios demonstrate that Gains significantly outperforms other baselines in performance for both source-domain and target-domain clients. Code is available at: https://github.com/Zhong-Zhengyi/Gains.
Zhengyi Zhong, Wenzheng Jiang, Weidong Bao 0001, Ji Wang 0002, Cheems Wang, Guanbo Wang, Yongheng Deng, Ju Ren 0001
NeurIPS4
2025 Efficient Multi-Task Modeling through Automated Fusion of Trained Models
abstract
Although multi-task learning is widely applied in intelligent services, traditional multi-task modeling methods often require customized designs based on specific task combinations, resulting in a cumbersome modeling process. Inspired by the rapid development and excellent performance of single-task models, this paper proposes an efficient multi-task modeling method that can automatically fuse trained single-task models with different structures and tasks to form a multi-task model. As a general framework, this method allows modelers to simply prepare trained models for the required tasks, simplifying the modeling process while fully utilizing the knowledge contained in the trained models. This eliminates the need for excessive focus on task relationships and model structure design. To achieve this goal, we consider the structural differences among various trained models and employ model decomposition techniques to hierarchically decompose them into multiple operable model components. Furthermore, we design an Adaptive Knowledge Fusion (AKF) module based on Transformer, which adaptively integrates intra-task and inter-task knowledge based on model components. Through the proposed method, we achieve efficient and automated construction of multi-task models, and its effectiveness is verified through extensive experiments on three datasets. Our code and related baseline methods can be found at: https://github.com/zxccvdql/EMM.
Jingxuan Zhou, Weidong Bao 0001, Ji Wang 0002, Dayu Zhang, Zhengyi Zhong
SMC3
2025 Revisiting the Byzantine Resilience of Federated Reinforcement Learning: A Distillation Perspective
abstract
Federated reinforcement learning (FRL) enhances sample efficiency while preserving data privacy. However, standard FRL frameworks rely on aggregating model parameters or gradients, making them vulnerable to Byzantine attacks. Current Byzantine-resilient approaches primarily focus on server-side robust aggregations, leaving the fundamental vulnerability of transmitting parameters unaddressed. In this paper, we revisit Byzantine resilience in FRL from the knowledge distillation (KD) perspective. KD-based FRL uploads policy representations instead of policy parameters. This framework-level shift fundamentally constrains the attack surface. We theoretically prove traditional FRL suffers unbounded corruption from Byzantine agents, whereas KD-based FRL converges to an ${\mathcal{O}}(\alpha )$-stationary point under α-fraction adversaries, formalizing the accuracy-robustness trade-off. Empirical validation confirms the Byzantine resilience of KD-based FRL: it maintains near-optimal performance across diverse attacks and even withstands Byzantine fractions up to 0.9. Our theoretical guarantees and experiments demonstrate distillation endows FRL with fundamentally stronger resilience.
Wenzheng Jiang, Ji Wang 0002, Zhengyi Zhong, Jiangzhou Liao, Xiaomin Zhu 0001, Flint Xiaofeng Fan
TrustCom2
2025 Enhancing long-term memory in federated class continual learning with lightweight adapters
Ji Wang 0002, Zhengyi Zhong, Weidong Bao 0001, Yaohong Zhang, Jianguo Chen 0001
Neurocomputing2
2025 All on board: Efficient reinforcement learning with milestone aggregation in asynchronous distributed training for RTS games
Dayu Zhang, Weidong Bao 0001, Ji Wang 0002, Xiongtao Zhang, Jingxuan Zhou, Yaohong Zhang
Neurocomputing3
2025 Improving Generalization and Personalization in Model-Heterogeneous Federated Learning
abstract
Conventional federated learning (FL) assumes the homogeneity of models, necessitating clients to expose their model parameters to enhance the performance of the server model. However, this assumption cannot reflect real-world scenarios. Sharing models and parameters raises security concerns for users, and solely focusing on the server-side model neglects clients' personalization requirements, potentially impeding expected performance improvements of users. On the other hand, prioritizing personalization may compromise the generalization of the server model, thereby hindering extensive knowledge migration. To address these challenges, we put forth an important problem: How can FL ensure both generalization and personalization when clients' models are heterogeneous? In this work, we introduce FedTED, which leverages a twin-branch structure and data-free knowledge distillation (DFKD) to address the challenges posed by model heterogeneity and diverse objectives in FL. The employed techniques in FedTED yield significant improvements in both personalization and generalization, while effectively coordinating the updating process of clients' heterogeneous models and successfully reconstructing a satisfactory global model. Our empirical evaluation demonstrates that FedTED outperforms many representative algorithms, particularly in scenarios where clients' models are heterogeneous, achieving a remarkable 19.37% enhancement in generalization performance and up to 9.76% improvement in personalization performance.
Xiongtao Zhang, Ji Wang 0002, Weidong Bao 0001, Yaohong Zhang, Xiaomin Zhu 0001, Hao Peng 0001, Xiang Zhao 0002
IEEE Trans. Neural Networks Learn. Syst.2
2025 SacFL: Self-Adaptive Federated Continual Learning for Resource-Constrained End Devices
abstract
The proliferation of end devices has led to a distributed computing paradigm, wherein on-device machine learning models continuously process diverse data generated by these devices. The dynamic nature of this data, characterized by continuous changes or data drift, poses significant challenges for on-device models. To address this issue, continual learning (CL) is proposed, enabling machine learning models to incrementally update their knowledge and mitigate catastrophic forgetting. However, the traditional centralized approach to CL is unsuitable for end devices due to privacy and data volume concerns. In this context, federated CL (FCL) emerges as a promising solution, preserving user data locally while enhancing models through collaborative updates. Aiming at the challenges of limited storage resources for CL, poor autonomy in task shift detection, and difficulty in coping with new adversarial tasks in the FCL scenario, we propose a novel FCL framework named self-adaptive federated CL (SacFL). $\rm {SacFL}$ employs an encoder-decoder architecture to separate task-robust and task-sensitive components, significantly reducing storage demands by retaining lightweight task-sensitive components for resource-constrained end devices. Moreover, $\rm {SacFL}$ leverages contrastive learning to introduce an autonomous data shift detection mechanism, enabling it to discern whether a new task has emerged and whether it is a benign task. This capability ultimately allows the device to autonomously trigger CL or attack defense strategy without additional information, which is more practical for end devices. Comprehensive experiments conducted on multiple text and image datasets, such as Cifar100 and THUCNews, have validated the effectiveness of $\rm {SacFL}$ in both class-incremental and domain-incremental scenarios. Furthermore, a demo system has been developed to verify its practicality.
Zhengyi Zhong, Weidong Bao 0001, Ji Wang 0002, Jianguo Chen 0001, Lingjuan Lyu, Wei Yang Bryan Lim
IEEE Trans. Neural Networks Learn. Syst.3
2024 Self-adaptive asynchronous federated optimizer with adversarial sharpness-aware minimization
Xiongtao Zhang, Ji Wang 0002, Weidong Bao 0001, Wenhua Xiao, Yaohong Zhang, Lihua Liu 0002
Future Gener. Comput. Syst.2
2024 Fault-Tolerant Scheduling of Heterogeneous UAVs for Data Collection of IoT Applications
abstract
UAV-enabled data collection is considered a promising paradigm of emergency data transmission for IoT applications when the communication infrastructure is damaged. UAV scheduling for data collection as critical technology has attracted widespread attention. Most studies default to the absolute reliability of UAVs for data collection, yet it is inevitable for UAVs to fail in flight. It is unacceptable if some critical data is lost due to UAV faults. Therefore, we research fault-tolerant scheduling for data collection enabled by heterogeneous UAVs. Firstly, a three-layer data collection motivation scenario is proposed, where the fault tolerance issue is involved for the first time. Then, we propose a utility-based fault tolerance model-UBFT to balance the reliability and efficiency of data collection. The UAV fault-tolerant scheduling is modeled as a multi-objective optimization problem to concurrently optimize the data throughput and load balancing of data collection. Combining the characteristics of optimization objectives, an alternating coordinate optimization method-ACTOR is presented to solve this problem efficiently. Numerous simulation experiments and real-machine experiments demonstrate that ACTOR-UBFT achieves excellent performance in fault tolerance, data throughput, adaptation, algorithm complexity, etc.
Weidong Bao 0001, Xiaoqing Li 0006, Xiaomin Zhu 0001, Yaohong Zhang, Ji Wang 0002, Ling Liu 0001
IEEE Internet Things J.6
2024 Structural graph federated learning: Exploiting high-dimensional information of statistical heterogeneity
Xiongtao Zhang, Ji Wang 0002, Weidong Bao 0001, Hao Peng 0001, Yaohong Zhang, Xiaomin Zhu 0001
Knowl. Based Syst.2
2024 A Hybrid Heuristic-Exact Optimization for Large-Scale Home Health Care Problem
abstract
During the COVID-19 pandemic, numerous people experiencing illness or senescence choose to receive home health care (HHC) services. However, a rapid increase in patients makes it a challenge to reasonably allocate nurses to provide HHC services under the condition of a paucity of nurse resources and patient time window constraints. To solve the large-scale HHC problem, a hybrid heuristic-exact optimization algorithm is proposed with three novel contributions. First, a framework of hybrid heuristic-exact optimization is designed to solve the large-scale problem where a reasonable solution is difficult to obtain under constraints. Second, a multi-objective mixed-integer linear programming modelization is formulated to get a more diverse nurse assignment. Finally, an improved branch and bound algorithm is proposed to speed up computation for the large-scale problem. Computational results on different HHC instances from 25 to 1000 patients demonstrate that the proposed algorithm can optimize the HHC problem with more than 100 patients and can provide various assignments for different numbers of nurses, which the common algorithm cannot optimize.
Xiaomin Zhu 0001, Mingyin Zou, Daqian Liu, Ji Wang 0002, Jun Tang 0001, Weidong Bao 0001
IEEE Trans. Comput. Biol. Bioinform.4
2023 Personalized Federated Relation Classification over Heterogeneous Texts
abstract
Relation classification detects the semantic relation between two annotated entities from a piece of text, which is a useful tool for structurization of knowledge. Recently, federated learning has been introduced to train relation classification models in decentralized settings. Current methods strive for a strong server model by decoupling the model training at server from direct access to texts at clients while taking advantage of them. Nevertheless, they overlook the fact that clients have heterogeneous texts (i.e., texts with diversely skewed distribution of relations), which renders existing methods less practical. In this paper, we propose to investigate personalized federated relation classification, in which strong client models adapted to their own data are desired. To further meet the challenges brought by heterogeneous texts, we present a novel framework, namely pf-RC, with several optimized designs. It features a knowledge aggregation method that exploits a relation-wise weighting mechanism, and a feature augmentation method that leverages prototypes to adaptively enhance the representations of instances of long-tail relations. We experimentally validate the superiority of pf-RC against competing baselines in various settings, and the results suggest that the tailored techniques mitigate the challenges.
Ning Pang, Xiang Zhao 0002, Weixin Zeng, Ji Wang 0002, Weidong Xiao 0003
SIGIR4
2023 Dyna-PPO reinforcement learning with Gaussian process for the continuous action decision-making in autonomous driving
Guanlin Wu, Wenqi Fang, Ji Wang 0002, Pin Ge, Jiang Cao, Yang Ping, Peng Gou
Appl. Intell.3
2023 Data Offloading Enabled by Heterogeneous UAVs for IoT Applications Under Uncertain Environments
abstract
With the continuous expansion of the Internet of Things (IoT) application scope, there are growing IoT scenarios that lack the coverage of wireless communication networks have the demand for data offloading. Efficient data transmission has been the main concern of these applications. Thus, data offloading within the limited communication environment has become a hotspot in both industry and academia. Since unmanned aerial vehicles (UAVs) can move across regions to make up for the communication gap caused by the loss of wireless communication networks, a lot of in-depth studies on UAV-enabled data offloading have been conducted. Nevertheless, few studies to date consider the uncertain user status and the heterogeneous UAV capabilities, which is however more practical and needs more attention. In this article, we propose an innovative framework to dynamically estimate user status information and determine the UAV scheduling strategy. On this basis, the heterogeneous UAV-enabled data offloading is modeled as a constrained multiobjective optimization problem, whose purpose is to lower the user data queue length while extending the working time of the UAV. Moreover, a differential evolution-based dynamic objective approximation method—RUDDER is proposed to solve the constrained multiobjective optimization problem. Through rigorous mathematical proof, we prove that RUDDER can consistently guide the population to approach the optimization solution with polynomial-level time complexity. To verify the effectiveness of the proposed RUDDER, extensive experiments are conducted to compare it with five comparison algorithms. The experimental results demonstrate the superiority of the RUDDER in terms of energy saving, time efficiency, and adaptability.
Weidong Bao 0001, Xiaomin Zhu 0001, Ji Wang 0002, Ling Liu 0001
IEEE Internet Things J.4
2023 PRETTY: A parallel transgenerational learning-assisted evolutionary algorithm for computationally expensive multi-objective optimization
Mingyin Zou, Xiaomin Zhu 0001, Ye Tian 0009, Ji Wang 0002, Huangke Chen
Inf. Sci.4
2023 Adversarial Attack and Defense on Graph Data: A Survey
abstract
Deep neural networks (DNNs) have been widely applied to various applications, including image classification, text generation, audio recognition, and graph data analysis. However, recent studies have shown that DNNs are vulnerable to adversarial attacks. Though there are several works about adversarial attack and defense strategies on domains such as images and natural language processing, it is still difficult to directly transfer the learned knowledge to graph data due to its representation structure. Given the importance of graph analysis, an increasing number of studies over the past few years have attempted to analyze the robustness of machine learning models on graph data. Nevertheless, existing research considering adversarial behaviors on graph data often focuses on specific types of attacks with certain assumptions. In addition, each work proposes its own mathematical formulation, which makes the comparison among different methods difficult. Therefore, this review is intended to provide an overall landscape of more than 100 papers on adversarial attack and defense strategies for graph data, and establish a unified formulation encompassing most graph adversarial learning models. Moreover, we also compare different graph attacks and defenses along with their contributions and limitations, as well as summarize the evaluation metrics, datasets and future trends. We hope this survey can help fill the gap in the literature and facilitate further development of this promising new field We also have created an online resource to keep track of relevant research on the basis of this survey athttps://github.com/safe-graph/graph-adversarial-learning-literature.
Lichao Sun 0001, Yingtong Dou, Carl Yang 0001, Kai Zhang 0039, Ji Wang 0002, Philip S. Yu, Lifang He 0001, Bo Li 0026
IEEE Trans. Knowl. Data Eng.5
2023 UNION: Fault-tolerant Cooperative Computing in Opportunistic Mobile Edge Cloud
abstract
Opportunistic Mobile Edge Cloud in which opportunistically connected mobile devices run in a cooperative way to augment the capability of a single device has become a timely and essential topic due to its widespread prospect under resource-constrained scenarios (e.g., disaster rescue). Because of the mobility of devices and the uncertainty of environments, it is inevitable that failures occur among the mobile nodes. Being different from existing studies that mainly focus on either data offloading or computing offloading among mobile devices in an ideal environment, we concentrate on how to guarantee the reliability of the task execution with the consideration of both data offloading and computing offloading under opportunistically connected mobile edge cloud. To this end, an optimization of mobile task offloading when considering reliability is formulated. Then, we propose a probabilistic model for task offloading and a reliability model for task execution, which estimates the probability of successful execution for a specific opportunistic path and describes the dynamic reliability of the task execution. Based on these models, a heuristic algorithm UNION (Fa u lt-Tolera n t Cooperat i ve C o mputi n g) is proposed to solve this NP-hard problem. Theoretical analysis shows that the complexity of UNION is 𝒪(|ℐ| 2 +|𝒩|) with guaranteeing the reliability of 0.99. Also, extensive experiments on real-world traces validate the superiority of the proposed algorithm UNION over existing typical strategies.
Wenhua Xiao, Xudong Fang, Bixin Liu, Ji Wang 0002, Xiaomin Zhu 0001
ACM Trans. Internet Techn.4
2023 YISHAN: Managing Large-scale Cloud Database Instances via Machine Learning
abstract
Efficiently managing database instances over cloud-scale clusters is significant for increasing service quality and reducing operational cost, especially confronting the growing cluster size and heterogeneous application services. Alibaba Cloud provides a large-scale Relational Database Service (RDS) for millions of users including enterprises from start-ups to large international corporations. To manage tremendous amount of RDS instances in the Cloud with the goal of reducing cost while guaranteeing service level agreement(SLA), YISHAN, an intelligent database instance management system, is designed to dynamically manage the placement of instances using machine learning techniques. YISHAN collects historical performance data to analyze patterns of the resource utilization of instances and hosts. By learning “good packings” in which instances colocate harmoniously, YISHAN is able to optimize the instances placement to provide better quality of service and improve the efficiency of CPU, memory, and disk resources. We deploy and run YISHAN in Alibaba Cloud RDS. The running logs show that YISHAN successfully saves 17% of the resources in hosts and efficiently reduces the burdens and crash risks of RDS instances.
Wenhua Xiao, Ji Wang 0002, Xiaomin Zhu 0001, Weidong Bao 0001, Xiaojie Feng, Wei Cao 0006, Feng Yu 0022, Ling Liu 0001
IEEE Trans. Serv. Comput.3
2022 Qauxi: Cooperative multi-agent reinforcement learning with knowledge transferred from auxiliary task
Wenqian Liang, Ji Wang 0002, Weidong Bao 0001, Xiaomin Zhu 0001, Guanlin Wu, Dayu Zhang, Liyuan Niu
Neurocomputing2
2022 FLEE: A Hierarchical Federated Learning Framework for Distributed Deep Neural Network over Cloud, Edge, and End Device
abstract
With the development of smart devices, the computing capabilities of portable end devices such as mobile phones have been greatly enhanced. Meanwhile, traditional cloud computing faces great challenges caused by privacy-leakage and time-delay problems, there is a trend to push models down to edges and end devices. However, due to the limitation of computing resource, it is difficult for end devices to complete complex computing tasks alone. Therefore, this article divides the model into two parts and deploys them on multiple end devices and edges, respectively. Meanwhile, an early exit is set to reduce computing resource overhead, forming a hierarchical distributed architecture. In order to enable the distributed model to continuously evolve by using new data generated by end devices, we comprehensively consider various data distributions on end devices and edges, proposing a hierarchical federated learning framework FLEE , which can realize dynamical updates of models without redeploying them. Through image and sentence classification experiments, we verify that it can improve model performances under all kinds of data distributions, and prove that compared with other frameworks, the models trained by FLEE consume less global computing resource in the inference stage.
Zhengyi Zhong, Weidong Bao 0001, Ji Wang 0002, Xiaomin Zhu 0001, Xiongtao Zhang
ACM Trans. Intell. Syst. Technol.3
2022 DANCE: Distributed Generative Adversarial Networks with Communication Compression
abstract
Generative adversarial networks (GANs) have shown great success in deep representations learning, data generation, and security enhancement. With the development of the Internet of Things, 5th generation wireless systems (5G), and other technologies, the large volume of data collected at the edge of networks provides a new way to improve the capabilities of GANs. Due to privacy, bandwidth, and legal constraints, it is not appropriate to upload all the data to the cloud or servers for processing. Therefore, this article focuses on deploying and training GANs at the edge rather than converging edge data to the central node. To address this problem, we designed a novel distributed learning architecture for GANs, called DANCE. DANCE can adaptively perform communication compression based on the available bandwidth, while supporting both data and model parallelism training of GANs. In addition, inspired by the gossip mechanism and Stackelberg game, a compatible algorithm, AC-GAN is proposed. The theoretical analysis guarantees the convergence of the model and the existence of approximate equilibrium in AC-GAN. Both simulation and prototype system experiments show that AC-GAN can achieve better training effectiveness with less communication overhead than the SOTA algorithms, i.e., FL-GAN and MD-GAN.
Xiongtao Zhang, Xiaomin Zhu 0001, Ji Wang 0002, Weidong Bao 0001, Laurence T. Yang
ACM Trans. Internet Techn.3
2021 ADAPT: Adaptive distributed optimization approach for uploading data with redundancy in cooperative mobile cloud
abstract
Summary With the development of information technology and the ubiquity of mobile devices, increasing amounts of data are generated, processed, and transmitted by mobile devices. To alleviate the tension between the energy poverty of mobile devices and the increasing demand for transmitting data, the energy‐efficient data transmission problem attracts considerable interests. Nonetheless, how to upload data with redundancy efficiently lacks a thorough study despite the wide existence of this problem in many situations like data storage among mobile devices and mobile crowd sensing. Since uploading redundant data brings little value while still consuming precious energy, it is important to design an efficient approach for mobile devices to upload data with redundancy cooperatively. In this work, we formulate the uploading data with redundancy in cooperative mobile cloud as an energy‐constrained utility maximization problem. To solve this problem, we propose an adaptive distributed optimization approach consisting of the correlated upload decision and the online distributed scheduling algorithm. By the correlated upload decision, each mobile device can make adaptive decisions on how much data to upload and which data to upload according to its own observations independently. The online distributed scheduling algorithm enables mobile devices to optimally upload data. A series of simulation experiments are conducted to demonstrate the effectiveness of our approach. Finally, we test our approach on a real demo system to verify its practicability in reality.
Ji Wang 0002, Weidong Bao 0001, Xiaomin Zhu 0001
Concurr. Comput. Pract. Exp.1
2021 EASE: Energy-efficient task scheduling for edge computing under uncertain runtime and unstable communication conditions
abstract
Summary Continuously growing network traffic has become a major technical bottleneck of the cloud service to develop the Internet of Things (IoTs) and mobile applications. Edge computing as a promising computing pattern deployed close to service users is expected to improve the quality of service (QoS). To fully utilize the capabilities of edge devices, a Device‐to‐Device (D2D)–based computing resource sharing and aggregation framework is proposed. Under this framework, this paper exploits the Beta distributions to characterize the uncertain communication rate and processing capability of the edge environment. The reliability and energy consumption of local computing and shared computing under uncertain conditions are, respectively, studied. We model the task scheduling as an Integer Programming problem, whose objective is to minimize the energy consumption while ensuring the reliability. Based on that, a heuristic task scheduling algorithm named EASE is proposed. Through a lot of simulation experiments, the performance of EASE is effectively evaluated under the static and dynamic environments. Compared with three comparison algorithms, EASE shows many advantages in terms of reliability, adaptability, and energy saving.
Xiaomin Zhu 0001, Dayu Zhang, Ji Wang 0002, Huangke Chen, Weidong Bao 0001
Concurr. Comput. Pract. Exp.5
2021 Distributed Learning on Mobile Devices: A New Approach to Data Mining in the Internet of Things
abstract
It is well known that deep learning is one of the most important methods for data mining. With the development of the fifth-generation mobile networks (5G) and the Internet of Things (IoT), the large volume of data collected in IoTs provides a new way to improve the capability of deep learning. Due to privacy, bandwidth, and legal concerns, it is impractical to send the data to a server or the cloud. The computing power of mobile devices makes it possible to process the data. Therefore, this article focuses on training these models in mobile devices. To solve the challenges, including unreliable networks, constrained resources, and slow convergence, we let multiple mobile devices learn a shared model collaboratively. We propose a novel architecture, GREAT, where each node chooses partners to share local model parameters according to link reliability. To balance the constrained resources and learning effectiveness, an optimization problem is developed by taking the reliability threshold as the variable of controlling the resources’ overhead. To implement this architecture, a dynamic control algorithm called Alpha-GossipSGD has been proposed. Its performance is evaluated by extensive experiments, which show that Alpha-GossipSGD can realize stable learning effectiveness over unreliable networks with constrained resources.
Xiongtao Zhang, Xiaomin Zhu 0001, Weidong Bao 0001, Laurence T. Yang, Ji Wang 0002, Huangke Chen
IEEE Internet Things J.5
2021 Knowledge graph embedding by relational and entity rotation
abstract
Knowledge graphs are typical large-scale multi-relational structures and useful for many artificial intelligence tasks. However, knowledge graphs often have missing facts, which limits the development of downstream tasks. To refine the knowledge graphs, knowledge graph embedding models have been developed. Knowledge graph embedding models aim to learn distributed representations for entities and relations and predict unknown triplets by scoring candidate triplets. Nevertheless, state-of-the-art works either aim to capture different relation patterns, or to model the multi-fold relations, and yet fail to consider these two aspects simultaneously. To fill this gap, in this paper, we propose a novel knowledge graph embedding model, MRotatE. It exploits triplet features from the perspective of relational and entity rotations, which can model and infer various relation patterns and handle with multi-fold relations at the same time. The experimental results demonstrate that MRotatE outperforms existing approaches and attains the state-of-the-art performance.
Xuqian Huang, Jiuyang Tang, Weixin Zeng, Ji Wang 0002, Xiang Zhao 0002
Knowl. Based Syst.5
2021 Kollector: Detecting Fraudulent Activities on Mobile Devices Using Deep Learning
abstract
With the rapid growth in smartphone usage, preventing leakage of personal information and privacy has become a challenging task. One major consequence of such leakage is impersonation. This type of illegal usage is nearly impossible to prevent as existing preventive mechanisms (e.g., passcode and fingerprinting), are not capable of continuously monitoring usage and determining whether the user is authorized. Once unauthorized users can defeat the initial protection mechanisms, they would have full access to the devices including using stored passwords to access high-value websites. We present Kollector, a new framework to detect impersonation based on a multi-view bagging deep learning approach to capture sequential tapping information on the smart-phone's keyboard. We construct a sequential-tapping biometrics model to continuously authenticate the user while typing. We empirically evaluated our system using real-world phone usage sessions from 26 users over eight weeks. We then compared our model against commonly used shallow machine techniques and find that our system performs better than other approaches and can achieve an 8.42 percent equal error rate, a 94.24 percent accuracy and a 94.41 percent H-mean using only the accelerometer and only five keyboard taps. We also experiment with using only three keyboard taps and find that the system still yields high accuracy while giving additional opportunities to make more decisions that can result in more accurate final decisions.
Lichao Sun 0001, Bokai Cao, Ji Wang 0002, Witawas Srisa-an, Philip S. Yu, Alex D. Leow, Stephen Checkoway
IEEE Trans. Mob. Comput.3
2020 Benign: An Automatic Optimization Framework for the Logic of Swarm Behaviors
abstract
In the field of swarm intelligence, it is usually complicated to express the logic of swarm behaviors. Behavior tree has drawn a lot of attention to be a practical approach to solving this problem in recent years. However, how to automatically design the logic of swarm behaviors according to the target of a task is the focus of swarm intelligence. Hence, we propose an automatic optimizing framework named Benign which is capable of using gene expression programming (GEP) to optimize the logic of swarm behaviors. In Benign, the basic swarm behaviors and the relationships among those behaviors are mapped to nodes of behavior tree by the method named Matt firstly. With these nodes, we design an artificial behavior tree. After that, the artificial behavior tree is transformed into an expression tree in GEP according to the method named Meet. Finally, GEP is used for optimization to generate the expected logic of swarm behaviors. We conduct simulation experiments to validate the efficiency of Benign. The experimental results show the superiority of Benign. Compared with the logic of the artificial behavior tree before optimization, the conduction of the optimized logic of swarm behaviors increases efficiency by more than 50%.
Jingjing Tao, Xiaomin Zhu 0001, Weidong Bao 0001, Ji Wang 0002
SMC6
2020 Federated learning with adaptive communication compression under dynamic bandwidth and unreliable networks
Xiongtao Zhang, Xiaomin Zhu 0001, Ji Wang 0002, Huangke Chen, Weidong Bao 0001
Inf. Sci.3
2019 Private Model Compression via Knowledge Distillation
abstract
The soaring demand for intelligent mobile applications calls for deploying powerful deep neural networks (DNNs) on mobile devices. However, the outstanding performance of DNNs notoriously relies on increasingly complex models, which in turn is associated with an increase in computational expense far surpassing mobile devices’ capacity. What is worse, app service providers need to collect and utilize a large volume of users’ data, which contain sensitive information, to build the sophisticated DNN models. Directly deploying these models on public mobile devices presents prohibitive privacy risk. To benefit from the on-device deep learning without the capacity and privacy concerns, we design a private model compression framework RONA. Following the knowledge distillation paradigm, we jointly use hint learning, distillation learning, and self learning to train a compact and fast neural network. The knowledge distilled from the cumbersome model is adaptively bounded and carefully perturbed to enforce differential privacy. We further propose an elegant query sample selection method to reduce the number of queries and control the privacy loss. A series of empirical evaluations as well as the implementation on an Android mobile device show that RONA can not only compress cumbersome models efficiently but also provide a strong privacy guarantee. For example, on SVHN, when a meaningful (9.83,10−6)-differential privacy is guaranteed, the compact model trained by RONA can obtain 20× compression ratio and 19× speed-up with merely 0.97% accuracy loss.
Ji Wang 0002, Weidong Bao 0001, Lichao Sun 0001, Xiaomin Zhu 0001, Bokai Cao, Philip S. Yu
AAAI1
2019 An Attention-augmented Deep Architecture for Hard Drive Status Monitoring in Large-scale Storage Systems
abstract
Data centers equipped with large-scale storage systems are critical infrastructures in the era of big data. The enormous amount of hard drives in storage systems magnify the failure probability, which may cause tremendous loss for both data service users and providers. Despite a set of reactive fault-tolerant measures such as RAID, it is still a tough issue to enhance the reliability of large-scale storage systems. Proactive prediction is an effective method to avoid possible hard-drive failures in advance. A series of models based on the SMART statistics have been proposed to predict impending hard-drive failures. Nonetheless, there remain some serious yet unsolved challenges like the lack of explainability of prediction results. To address these issues, we carefully analyze a dataset collected from a real-world large-scale storage system and then design an attention-augmented deep architecture for hard-drive health status assessment and failure prediction. The deep architecture, composed of a feature integration layer, a temporal dependency extraction layer, an attention layer, and a classification layer, cannot only monitor the status of hard drives but also assist in failure cause diagnoses. The experiments based on real-world datasets show that the proposed deep architecture is able to assess the hard-drive status and predict the impending failures accurately. In addition, the experimental results demonstrate that the attention-augmented deep architecture can reveal the degradation progression of hard drives automatically and assist administrators in tracing the cause of hard drive failures.
Ji Wang 0002, Weidong Bao 0001, Lei Zheng 0001, Xiaomin Zhu 0001, Philip S. Yu
ACM Trans. Storage1
2018 A Parallel Fast Fourier Transform Algorithm for Large-Scale Signal Data Using Apache Spark in Cloud
Weidong Bao 0001, Xiaomin Zhu 0001, Ji Wang 0002, Wenhua Xiao
ICA3PP (3)4
2018 Deep Learning towards Mobile Applications
abstract
Recent years have witnessed an explosive growth of mobile devices. Mobile devices are permeating every aspect of our daily lives. With the increasing usage of mobile devices and intelligent applications, there is a soaring demand for mobile applications with machine learning services. Inspired by the tremendous success achieved by deep learning in many machine learning tasks, it becomes a natural trend to push deep learning towards mobile applications. However, there exist many challenges to realize deep learning in mobile applications, including the contradiction between the miniature nature of mobile devices and the resource requirement of deep neural networks, the privacy and security concerns about individuals' data, and so on. To resolve these challenges, during the past few years, great leaps have been made in this area. In this paper, we provide an overview of the current challenges and representative achievements about pushing deep learning on mobile devices from three aspects: training with mobile data, efficient inference on mobile devices, and applications of mobile deep learning. The former two aspects cover the primary tasks of deep learning. Then, we go through our two recent applications that apply the data collected by mobile devices to inferring mood disturbance and user identification. Finally, we conclude this paper with the discussion of the future of this area.
Ji Wang 0002, Bokai Cao, Philip S. Yu, Lichao Sun 0001, Weidong Bao 0001, Xiaomin Zhu 0001
ICDCS1
2018 Layerwise Perturbation-Based Adversarial Training for Hard Drive Health Degree Prediction
abstract
With the development of cloud computing and big data, the reliability of data storage systems becomes increasingly important. Previous researchers have shown that machine learning algorithms based on SMART attributes are effective methods to predict hard drive failures. In this paper, we use SMART attributes to predict hard drive health degrees which are helpful for taking different fault tolerant actions in advance. Given the highly imbalanced SMART datasets, it is a nontrivial work to predict the health degree precisely. The proposed model would encounter overfitting and biased fitting problems if it is trained by the traditional methods. In order to resolve this problem, we propose two strategies to better utilize imbalanced data and improve performance. Firstly, we design a layerwise perturbation-based adversarial training method which can add perturbations to any layers of a neural network to improve the generalization of the network. Secondly, we extend the training method to the semi-supervised settings. Then, it is possible to utilize unlabeled data that have a potential of failure to further improve the performance of the model. Our extensive experiments on two real-world hard drive datasets demonstrate the superiority of the proposed schemes for both supervised and semi-supervised classification. The model trained by the proposed method can correctly predict the hard drive health status 5 and 15 days in advance.
Jianguo Zhang 0005, Ji Wang 0002, Lifang He 0001, Zhao Li 0007, Philip S. Yu
ICDM2
2018 Not Just Privacy: Improving Performance of Private Deep Learning in Mobile Cloud
abstract
The increasing demand for on-device deep learning services calls for a highly efficient manner to deploy deep neural networks (DNNs) on mobile devices with limited capacity. The cloud-based solution is a promising approach to enabling deep learning applications on mobile devices where the large portions of a DNN are offloaded to the cloud. However, revealing data to the cloud leads to potential privacy risk. To benefit from the cloud data center without the privacy risk, we design, evaluate, and implement a cloud-based framework ARDEN which partitions the DNN across mobile devices and cloud data centers. A simple data transformation is performed on the mobile device, while the resource-hungry training and the complex inference rely on the cloud data center. To protect the sensitive information, a lightweight privacy-preserving mechanism consisting of arbitrary data nullification and random noise addition is introduced, which provides strong privacy guarantee. A rigorous privacy budget analysis is given. Nonetheless, the private perturbation to the original data inevitably has a negative impact on the performance of further inference on the cloud side. To mitigate this influence, we propose a noisy training method to enhance the cloud-side network robustness to perturbed data. Through the sophisticated design, ARDEN can not only preserve privacy but also improve the inference performance. To validate the proposed ARDEN, a series of experiments based on three image datasets and a real mobile application are conducted. The experimental results demonstrate the effectiveness of ARDEN. Finally, we implement ARDEN on a demo system to verify its practicality.
Ji Wang 0002, Jianguo Zhang 0005, Weidong Bao 0001, Xiaomin Zhu 0001, Bokai Cao, Philip S. Yu
KDD1
2018 SP-Partitioner: A novel partition method to handle intermediate data skew in spark streaming
Guipeng Liu, Xiaomin Zhu 0001, Ji Wang 0002, Deke Guo, Weidong Bao 0001, Hui Guo 0001
Future Gener. Comput. Syst.3
2017 A Lightweight Recommendation Framework for Mobile User's Link Selection in Dense Network
abstract
With the proliferation of mobile devices and the development of communication technology, mobile devices have permeated every aspect of our daily lives. However, in dense network where large crowd of mobile devices try to access to the network simultaneously, the severe interference between mobile devices may incur a remarkable deterioration of the wireless communication quality. How to improve individual's experience in such scenario is a critical yet open problem. Inspired by the mobile device users' usage pattern as well as the characteristic of most wireless communication systems, we propose a framework offering uplink/downlink selection recommendation to different mobile device users to enhance their utility in this paper. The design of the framework starts with formulating the problem as a link selection game. Analysis shows that the game can be categorized as a generalized ordinal potential game whose Nash Equilibrium is guaranteed. We then devise a distributed link selection algorithm to generate a Nash Equilibrium of the game. To accommodate to the characteristic of dense network and the capacity limitation of mobile device, the design of the algorithm shows a light-weight property and does not require each mobile device user to know others' current selection. The probability of incomplete information gathering is also considered. Extensive experiments are conducted to demonstrate the effectiveness and superiority of the proposed framework. Experimental results show that the global average utility increase rate reaches above 20%, and about 70% mobile device users can benefit from using our framework.
Ji Wang 0002, Xiaomin Zhu 0001, Weidong Bao 0001, Guanlin Wu
ICDCS1
2017 Towards collaborative storage scheduling using alternating direction method of multipliers for mobile edge cloud
Guanlin Wu, Junjie Chen 0007, Weidong Bao 0001, Xiaomin Zhu 0001, Wenhua Xiao, Ji Wang 0002
J. Syst. Softw.6
2016 A Utility-Aware Approach to Redundant Data Upload in Cooperative Mobile Cloud
abstract
With the proliferation of mobile devices and the improvement of wireless communication technology, an increasing number of mobile devices are utilized for emergency management and healthcare monitoring. Redundant data upload to the cloud datacenters is gaining growing interest and attraction. One of the main challenges for redundant data upload in the cooperative mobile cloud is the optimization problem of how to provide high utility and high energy efficiency for data upload in the presence of intermittent connectivity and unpredictable bandwidth of wireless and mobile network. In this paper, we formulate the problem of redundant data upload in the cooperative mobile cloud as an energy-constrained utility maximization problem that aims at maximizing the amount of effective data uploaded under the energy consumption constraints. We propose an online distributed approach to enabling mobile devices to optimally make upload decisions without depending on the current state information of other devices and the prior knowledge of its own future context. We provide a rigorous theoretical analysis and an extensive suite of simulation experiments to demonstrate the effectiveness and superiority of our approach.
Ji Wang 0002, Xiaomin Zhu 0001, Weidong Bao 0001, Ling Liu 0001
CLOUD1
2016 Fault-Tolerant Scheduling for Real-Time Scientific Workflows with Elastic Resource Provisioning in Virtualized Clouds
abstract
Clouds are becoming an important platform for scientific workflow applications. However, with many nodes being deployed in clouds, managing reliability of resources becomes a critical issue, especially for the real-time scientific workflow execution where deadlines should be satisfied. Therefore, fault tolerance in clouds is extremely essential. The PB (primary backup) based scheduling is a popular technique for fault tolerance and has effectively been used in the cluster and grid computing. However, applying this technique for real-time workflows in a virtualized cloud is much more complicated and has rarely been studied. In this paper, we address this problem. We first establish a real-time workflow fault-tolerant model that extends the traditional PB model by incorporating the cloud characteristics. Based on this model, we develop approaches for task allocation and message transmission to ensure faults can be tolerated during the workflow execution. Finally, we propose a dynamic fault-tolerant scheduling algorithm, FASTER, for realtime workflows in the virtualized cloud. FASTER has three key features: 1) it employs a backward shifting method to make full use of the idle resources and incorporates task overlapping and VM migration for high resource utilization, 2) it applies the vertical/horizontal scaling-up technique to quickly provision resources for a burst of workflows, and 3) it uses the vertical scaling-down scheme to avoid unnecessary and ineffective resource changes due to fluctuated workflow requests. We evaluate our FASTER algorithm with synthetic workflows and workflows collected from the real scientific and business applications and compare it with six baseline algorithms. The experimental results demonstrate that FASTER can effectively improve the resource utilization and schedulability even in the presence of node failures in virtualized clouds.
Xiaomin Zhu 0001, Ji Wang 0002, Hui Guo 0001, Dakai Zhu 0001, Laurence T. Yang, Ling Liu 0001
IEEE Trans. Parallel Distributed Syst.2
2015 FESTAL: Fault-Tolerant Elastic Scheduling Algorithm for Real-Time Tasks in Virtualized Clouds
abstract
As clouds have been deployed widely in various fields, the reliability and availability of clouds become the major concern of cloud service providers and users. Thereby, fault tolerance in clouds receives a great deal of attention in both industry and academia, especially for real-time applications due to their safety critical nature. Large amounts of researches have been conducted to realize fault tolerance in distributed systems, among which fault-tolerant scheduling plays a significant role. However, few researches on the fault-tolerant scheduling study the virtualization and the elasticity, two key features of clouds, sufficiently. To address this issue, this paper presents a fault-tolerant mechanism which extends the primary-backup model to incorporate the features of clouds. Meanwhile, for the first time, we propose an elastic resource provisioning mechanism in the fault-tolerant context to improve the resource utilization. On the basis of the fault-tolerant mechanism and the elastic resource provisioning mechanism, we design novel fault-tolerant elastic scheduling algorithms for real-time tasks in clouds named FESTAL, aiming at achieving both fault tolerance and high resource utilization in clouds. Extensive experiments injecting with random synthetic workloads as well as the workload from the latest version of the Google cloud tracelogs are conducted by CloudSim to compare FESTAL with three baseline algorithms, i.e., Non-M igration-FESTAL (NMFESTAL), Non-Overlapping-FESTAL (NOFESTAL), and Elastic First Fit (EFF). The experimental results demonstrate that FESTAL is able to effectively enhance the performance of virtualized clouds.
Ji Wang 0002, Weidong Bao 0001, Xiaomin Zhu 0001, Laurence T. Yang, Yang Xiang 0001
IEEE Trans. Computers1
2015 Fault-Tolerant Scheduling for Real-Time Tasks on Multiple Earth-Observation Satellites
abstract
Fault-tolerance plays an important role in improving the reliability of multiple earth-observing satellites, especially in emergent scenarios such as obtaining photographs on battlefields or earthquake areas. Fault tolerance can be implemented through scheduling approaches. Unfortunately, little attention has been paid to fault-tolerant scheduling on satellites. To address this issue, we propose a novel dynamic fault-tolerant scheduling model for real-time tasks running on multiple observation satellites. In this model, the primary-backup policy is employed to tolerate one satellite's permanent failure at one time instant. In the light of the fault-tolerant model, we develop a novel fault-tolerant satellite scheduling algorithm named FTSS. To improve the resource utilization, we apply the overlapping technology that includes primary-backup copy overlapping (i.e., PB overlapping) and backup-backup copy overlapping (i.e., BB overlapping). According to the satellites characterized with time windows for observations, we extensively analyze the overlapping mechanism on satellites. We integrate the overlapping mechanism with FTSS, which employs the task merging strategies including primary-backup copy merging (i.e., PB merging), backup-backup copy merging (i.e., BB merging) and primary-primary copy merging (i.e., PP merging). These merging strategies are used to decrease the number of tasks required to be executed, thereby enhancing system schedulability. To demonstrate the superiority of our FTSS, we conduct extensive experiments using the real-world satellite parameters supplied from the satellite tool kit or STK; we compare FTSS with the three baseline algorithms, namely, NMFTSS, NOFTSS, and NMNOFTSS. The experimental results indicate that FTSS efficiently improves the scheduling quality of others and is suitable for fault-tolerant satellite scheduling.
Xiaomin Zhu 0001, Jianjiang Wang, Xiao Qin 0001, Ji Wang 0002, Zhong Liu 0002, Erik Demeulemeester
IEEE Trans. Parallel Distributed Syst.4
2014 Analysis and Design of Fault-Tolerant Scheduling for Real-Time Tasks on Earth-Observation Satellites
abstract
Fault-tolerant scheduling is an efficient approach to improving the reliability of multiple earth-observing satellites especially in some emergent scenarios such as obtaining photographs on battlefields or earthquake areas. Unfortunately, little work has been done to deal with the fault-tolerant scheduling on satellites. To address this issue, this paper presents a novel dynamic fault-tolerant scheduling model using primary-backup policy to tolerate one satellite's permanent failure at one time instant. On this basis, we propose a novel fault-tolerant satellite scheduling algorithm named FTSS, in which an overlapping technology is adopted to improve the resource utilization. Besides, the FTSS employs the task merging strategies to further enhance the schedulability. To demonstrate the superiority of our FTSS, we conduct extensive experiments by simulations using real-world satellite parameters from STK to compare FTSS with other baseline algorithms. The experimental results indicate that FTSS efficiently improves the scheduling quality of others and is suitable for fault-tolerant satellite scheduling.
Xiaomin Zhu 0001, Jianjiang Wang, Ji Wang 0002, Xiao Qin 0001
ICPP3
2014 Real-Time Tasks Oriented Energy-Aware Scheduling in Virtualized Clouds
abstract
Energy conservation is a major concern in cloud computing systems because it can bring several important benefits such as reducing operating costs, increasing system reliability, and prompting environmental protection. Meanwhile, power-aware scheduling approach is a promising way to achieve that goal. At the same time, many real-time applications, e.g., signal processing, scientific computing have been deployed in clouds. Unfortunately, existing energy-aware scheduling algorithms developed for clouds are not real-time task oriented, thus lacking the ability of guaranteeing system schedulability. To address this issue, we first propose in this paper a novel rolling-horizon scheduling architecture for real-time task scheduling in virtualized clouds. Then a task-oriented energy consumption model is given and analyzed. Based on our scheduling architecture, we develop a novel energy-aware scheduling algorithm named EARH for real-time, aperiodic, independent tasks. The EARH employs a rolling-horizon optimization policy and can also be extended to integrate other energy-aware scheduling algorithms. Furthermore, we propose two strategies in terms of resource scaling up and scaling down to make a good trade-off between task’s schedulability and energy conservation. Extensive simulation experiments injecting random synthetic tasks as well as tasks following the last version of the Google cloud tracelogs are conducted to validate the superiority of our EARH by comparing it with some baselines. The experimental results show that EARH significantly improves the scheduling quality of others and it is suitable for real-time task scheduling in virtualized clouds.
Xiaomin Zhu 0001, Laurence T. Yang, Huangke Chen, Ji Wang 0002, Shu Yin 0001, Xiaocheng Liu
IEEE Trans. Cloud Comput.4