Xiaomin Zhu 0001

dblp:09/4144-1 · DBLP profile ↗
← Back
87ranked-venue papers
14as first author
33since 2021 · last 2026
0000-0003-1301-7840ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 34 · 7 first-author · 4 since 2021Software engineering, systems software and programming languages · 14 · 3 first-author · 4 since 2021Computer networks · 12 · 1 first-author · 9 since 2021Databases, data management, data science and information retrieval · 11 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 6 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 2 since 2021Security and privacy · 3 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 since 2021Theory of computation · 3 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2026 Dynamic demand-aware UAV scheduling for IoT data collection using deep reinforcement learning approach
Xiaoqing Li 0006, Weidong Bao 0001, Qingbao Liu, Ji Wang 0002, Xiaomin Zhu 0001
Future Gener. Comput. Syst.7
2026 CROSS: Feedback-Oriented Multi-Modal Dynamic Alignment in Recommendation Systems
abstract
Aligning the multi-modal content and ID embeddings is crucial in multi-modal recommendation systems. Existing solutions typically adopt a bidirectional alignment paradigm. Our prior work, FETTLE , challenges this paradigm by proposing a one-way directional alignment at the item level, thus reducing the negative impact of low-quality modalities. However, FETTLE leaves two open questions: (1) when is one-way directional alignment optimal, and (2) how to incorporate collaborative signals to enhance alignment? We present CROSS (feedba C k-o R iented multi-m O dal alignment in recommendation S y S tem), a plug-and-play framework that extends FETTLE by introducing three major advancements. First, we introduce Dynamic Item-Level Alignment , which dynamically calibrates the “strength” of each modality via a variance-based compensation mechanism, mitigating the risk of overshadowing weaker modalities in the early stages of training. Second, we develop Multi-grained Collaborative Alignment , which introduces a medium-granularity alignment strategy based on neighboring items that share similar user feedback profiles. This neighbor-level alignment effectively balances noisy user interactions and excessive smoothing across items. Third, we conduct extensive experiments on more real-world datasets and show that CROSS significantly boosts the performance of both collaborative filtering (CF) models and multi-modal recommendation (MRS) approaches, achieving 21.52%–70.78% average improvement on CF backbones and 8.70%–20.73% on MRS backbones. Compared with FETTLE , CROSS achieves additional improvements of 3.82%–5.24%.
Yang Li 0213, Junpeng Du, Chenzhan Wang, Zunlong Liu, Xiaomin Zhu 0001, Chen Lin 0001
Trans. Recomm. Syst.5
2025 Revisiting the Byzantine Resilience of Federated Reinforcement Learning: A Distillation Perspective
abstract
Federated reinforcement learning (FRL) enhances sample efficiency while preserving data privacy. However, standard FRL frameworks rely on aggregating model parameters or gradients, making them vulnerable to Byzantine attacks. Current Byzantine-resilient approaches primarily focus on server-side robust aggregations, leaving the fundamental vulnerability of transmitting parameters unaddressed. In this paper, we revisit Byzantine resilience in FRL from the knowledge distillation (KD) perspective. KD-based FRL uploads policy representations instead of policy parameters. This framework-level shift fundamentally constrains the attack surface. We theoretically prove traditional FRL suffers unbounded corruption from Byzantine agents, whereas KD-based FRL converges to an ${\mathcal{O}}(\alpha )$-stationary point under α-fraction adversaries, formalizing the accuracy-robustness trade-off. Empirical validation confirms the Byzantine resilience of KD-based FRL: it maintains near-optimal performance across diverse attacks and even withstands Byzantine fractions up to 0.9. Our theoretical guarantees and experiments demonstrate distillation endows FRL with fundamentally stronger resilience.
Wenzheng Jiang, Ji Wang 0002, Zhengyi Zhong, Jiangzhou Liao, Xiaomin Zhu 0001, Flint Xiaofeng Fan
TrustCom5
2025 Task-driven multi-UAV path planning via three-stage optimization strategy for urban region surveillance
Bowen Fei, Daqian Liu, Weidong Bao 0001, Xiaomin Zhu 0001, Xiaoqing Li 0006
Adv. Eng. Informatics4
2025 Improving Generalization and Personalization in Model-Heterogeneous Federated Learning
abstract
Conventional federated learning (FL) assumes the homogeneity of models, necessitating clients to expose their model parameters to enhance the performance of the server model. However, this assumption cannot reflect real-world scenarios. Sharing models and parameters raises security concerns for users, and solely focusing on the server-side model neglects clients' personalization requirements, potentially impeding expected performance improvements of users. On the other hand, prioritizing personalization may compromise the generalization of the server model, thereby hindering extensive knowledge migration. To address these challenges, we put forth an important problem: How can FL ensure both generalization and personalization when clients' models are heterogeneous? In this work, we introduce FedTED, which leverages a twin-branch structure and data-free knowledge distillation (DFKD) to address the challenges posed by model heterogeneity and diverse objectives in FL. The employed techniques in FedTED yield significant improvements in both personalization and generalization, while effectively coordinating the updating process of clients' heterogeneous models and successfully reconstructing a satisfactory global model. Our empirical evaluation demonstrates that FedTED outperforms many representative algorithms, particularly in scenarios where clients' models are heterogeneous, achieving a remarkable 19.37% enhancement in generalization performance and up to 9.76% improvement in personalization performance.
Xiongtao Zhang, Ji Wang 0002, Weidong Bao 0001, Yaohong Zhang, Xiaomin Zhu 0001, Hao Peng 0001, Xiang Zhao 0002
IEEE Trans. Neural Networks Learn. Syst.5
2024 GENET: Unleashing the Power of Side Information for Recommendation via Hypergraph Pre-training
Yang Li 0213, Qi'ao Zhao, Chen Lin 0001, Xiaomin Zhu 0001, Jinsong Su
DASFAA (3)5
2024 DAWN: Dynamic Task Planning of Multi-UAV With Two-Layer Optimization Mechanism in Uncertain Environments
abstract
UAV cooperative formation provides rescue and material delivery for the industrial Internet of Things (IIoT). To solve issues, such as low material distribution efficiency and poor mobility during disaster rescue, we propose a two-layer optimization mechanism-based multiple UAV dynamic task planning method (DAWN), which can cope with the problem of the global communication link unreachable caused by disasters. Specifically, we consider the global task allocation as a dynamic vehicle routing problem (VRP) and use deep reinforcement learning (DRL) to solve it so as to minimize the global flight path and energy consumption. Second, based on the current communication structure, we establish a local path planning approach based on the trust network that maximizes the regional coverage rate while minimizing the flight paths. On the basis of these two layers, an UAV formation dynamic task planning approach is realized. Experimental results prove that the proposed DAWN can obtain the optimal flight paths and achieve higher energy efficiency while providing reasonable region coverage to discover more potential tasks.
Daqian Liu, Bowen Fei, Weidong Bao 0001, Xiaomin Zhu 0001, Xiaoqing Li 0006
IEEE Internet Things J.4
2024 Fault-Tolerant Scheduling of Heterogeneous UAVs for Data Collection of IoT Applications
abstract
UAV-enabled data collection is considered a promising paradigm of emergency data transmission for IoT applications when the communication infrastructure is damaged. UAV scheduling for data collection as critical technology has attracted widespread attention. Most studies default to the absolute reliability of UAVs for data collection, yet it is inevitable for UAVs to fail in flight. It is unacceptable if some critical data is lost due to UAV faults. Therefore, we research fault-tolerant scheduling for data collection enabled by heterogeneous UAVs. Firstly, a three-layer data collection motivation scenario is proposed, where the fault tolerance issue is involved for the first time. Then, we propose a utility-based fault tolerance model-UBFT to balance the reliability and efficiency of data collection. The UAV fault-tolerant scheduling is modeled as a multi-objective optimization problem to concurrently optimize the data throughput and load balancing of data collection. Combining the characteristics of optimization objectives, an alternating coordinate optimization method-ACTOR is presented to solve this problem efficiently. Numerous simulation experiments and real-machine experiments demonstrate that ACTOR-UBFT achieves excellent performance in fault tolerance, data throughput, adaptation, algorithm complexity, etc.
Weidong Bao 0001, Xiaoqing Li 0006, Xiaomin Zhu 0001, Yaohong Zhang, Ji Wang 0002, Ling Liu 0001
IEEE Internet Things J.4
2024 Structural graph federated learning: Exploiting high-dimensional information of statistical heterogeneity
Xiongtao Zhang, Ji Wang 0002, Weidong Bao 0001, Hao Peng 0001, Yaohong Zhang, Xiaomin Zhu 0001
Knowl. Based Syst.6
2024 A Hybrid Heuristic-Exact Optimization for Large-Scale Home Health Care Problem
abstract
During the COVID-19 pandemic, numerous people experiencing illness or senescence choose to receive home health care (HHC) services. However, a rapid increase in patients makes it a challenge to reasonably allocate nurses to provide HHC services under the condition of a paucity of nurse resources and patient time window constraints. To solve the large-scale HHC problem, a hybrid heuristic-exact optimization algorithm is proposed with three novel contributions. First, a framework of hybrid heuristic-exact optimization is designed to solve the large-scale problem where a reasonable solution is difficult to obtain under constraints. Second, a multi-objective mixed-integer linear programming modelization is formulated to get a more diverse nurse assignment. Finally, an improved branch and bound algorithm is proposed to speed up computation for the large-scale problem. Computational results on different HHC instances from 25 to 1000 patients demonstrate that the proposed algorithm can optimize the HHC problem with more than 100 patients and can provide various assignments for different numbers of nurses, which the common algorithm cannot optimize.
Xiaomin Zhu 0001, Mingyin Zou, Daqian Liu, Ji Wang 0002, Jun Tang 0001, Weidong Bao 0001
IEEE Trans. Comput. Biol. Bioinform.1
2023 TCCM: Time and Content-Aware Causal Model for Unbiased News Recommendation
abstract
Popularity bias significantly impacts news recommendation systems, as popular news articles receive more exposure and are often delivered to irrelevant users, resulting in unsatisfactory performance. Existing methods have not adequately addressed the issue of popularity bias in news recommendations, largely due to the neglect of the time factor and the impact of news content on popularity. In this paper, we propose a novel approach called Time and Content-aware Causal Model, namely TCCM. It models the effects of three factors on user interaction behavior, i.e., the time factor, the news popularity, and the matching between news content and user interest. TCCM also estimates news popularity more accurately by incorporating the news content, i.e., the popularity of entity and words. Causal intervention techniques are applied to obtain debiased recommendations. Extensive experiments on well-known benchmark datasets demonstrate that the proposed approach outperforms a range of state-of-the-art techniques.
Yewang Chen, Weiyao Ye, Guipeng Xv, Chen Lin 0001, Xiaomin Zhu 0001
CIKM5
2023 Data Offloading Enabled by Heterogeneous UAVs for IoT Applications Under Uncertain Environments
abstract
With the continuous expansion of the Internet of Things (IoT) application scope, there are growing IoT scenarios that lack the coverage of wireless communication networks have the demand for data offloading. Efficient data transmission has been the main concern of these applications. Thus, data offloading within the limited communication environment has become a hotspot in both industry and academia. Since unmanned aerial vehicles (UAVs) can move across regions to make up for the communication gap caused by the loss of wireless communication networks, a lot of in-depth studies on UAV-enabled data offloading have been conducted. Nevertheless, few studies to date consider the uncertain user status and the heterogeneous UAV capabilities, which is however more practical and needs more attention. In this article, we propose an innovative framework to dynamically estimate user status information and determine the UAV scheduling strategy. On this basis, the heterogeneous UAV-enabled data offloading is modeled as a constrained multiobjective optimization problem, whose purpose is to lower the user data queue length while extending the working time of the UAV. Moreover, a differential evolution-based dynamic objective approximation method—RUDDER is proposed to solve the constrained multiobjective optimization problem. Through rigorous mathematical proof, we prove that RUDDER can consistently guide the population to approach the optimization solution with polynomial-level time complexity. To verify the effectiveness of the proposed RUDDER, extensive experiments are conducted to compare it with five comparison algorithms. The experimental results demonstrate the superiority of the RUDDER in terms of energy saving, time efficiency, and adaptability.
Weidong Bao 0001, Xiaomin Zhu 0001, Ji Wang 0002, Ling Liu 0001
IEEE Internet Things J.3
2023 PRETTY: A parallel transgenerational learning-assisted evolutionary algorithm for computationally expensive multi-objective optimization
Mingyin Zou, Xiaomin Zhu 0001, Ye Tian 0009, Ji Wang 0002, Huangke Chen
Inf. Sci.2
2023 UNION: Fault-tolerant Cooperative Computing in Opportunistic Mobile Edge Cloud
abstract
Opportunistic Mobile Edge Cloud in which opportunistically connected mobile devices run in a cooperative way to augment the capability of a single device has become a timely and essential topic due to its widespread prospect under resource-constrained scenarios (e.g., disaster rescue). Because of the mobility of devices and the uncertainty of environments, it is inevitable that failures occur among the mobile nodes. Being different from existing studies that mainly focus on either data offloading or computing offloading among mobile devices in an ideal environment, we concentrate on how to guarantee the reliability of the task execution with the consideration of both data offloading and computing offloading under opportunistically connected mobile edge cloud. To this end, an optimization of mobile task offloading when considering reliability is formulated. Then, we propose a probabilistic model for task offloading and a reliability model for task execution, which estimates the probability of successful execution for a specific opportunistic path and describes the dynamic reliability of the task execution. Based on these models, a heuristic algorithm UNION (Fa u lt-Tolera n t Cooperat i ve C o mputi n g) is proposed to solve this NP-hard problem. Theoretical analysis shows that the complexity of UNION is 𝒪(|ℐ| 2 +|𝒩|) with guaranteeing the reliability of 0.99. Also, extensive experiments on real-world traces validate the superiority of the proposed algorithm UNION over existing typical strategies.
Wenhua Xiao, Xudong Fang, Bixin Liu, Ji Wang 0002, Xiaomin Zhu 0001
ACM Trans. Internet Techn.5
2023 YISHAN: Managing Large-scale Cloud Database Instances via Machine Learning
abstract
Efficiently managing database instances over cloud-scale clusters is significant for increasing service quality and reducing operational cost, especially confronting the growing cluster size and heterogeneous application services. Alibaba Cloud provides a large-scale Relational Database Service (RDS) for millions of users including enterprises from start-ups to large international corporations. To manage tremendous amount of RDS instances in the Cloud with the goal of reducing cost while guaranteeing service level agreement(SLA), YISHAN, an intelligent database instance management system, is designed to dynamically manage the placement of instances using machine learning techniques. YISHAN collects historical performance data to analyze patterns of the resource utilization of instances and hosts. By learning “good packings” in which instances colocate harmoniously, YISHAN is able to optimize the instances placement to provide better quality of service and improve the efficiency of CPU, memory, and disk resources. We deploy and run YISHAN in Alibaba Cloud RDS. The running logs show that YISHAN successfully saves 17% of the resources in hosts and efficiently reduces the burdens and crash risks of RDS instances.
Wenhua Xiao, Ji Wang 0002, Xiaomin Zhu 0001, Weidong Bao 0001, Xiaojie Feng, Wei Cao 0006, Feng Yu 0022, Ling Liu 0001
IEEE Trans. Serv. Comput.4
2022 User-level parallel file system: Case studies and performance optimizations
abstract
Abstract User‐level file systems are usually adopted to bridge the gap between efficacy and efficiency of file system developments for new applications' I/O demands. And the widely known user‐space file system framework, FUSE, is commonly utilized to deployed user‐level file systems. This article first uses a popular stack‐able file system as a case study to exam how FUSE affects I/O performance. Based on the testing and analytical results, this article then presents SHC, an implementation method to implement a user‐level file system without FUSE intervention. Experimental results indicate that SHC improves write bandwidth by up to 5.6x compared with that of FUSE and present leading superiority on read cases.
Yanliang Zou, Chen Chen 0124, Tongliang Deng, Jian Zhang 0070, Xiaomin Zhu 0001, Si Chen 0009, Shu Yin 0001
Concurr. Comput. Pract. Exp.5
2022 Qauxi: Cooperative multi-agent reinforcement learning with knowledge transferred from auxiliary task
Wenqian Liang, Ji Wang 0002, Weidong Bao 0001, Xiaomin Zhu 0001, Guanlin Wu, Dayu Zhang, Liyuan Niu
Neurocomputing4
2022 Autonomous Cooperative Search Model for Multi-UAV With Limited Communication Network
abstract
With the rapid development of artificial intelligence technology, the multi-UAV cooperative search has wide applications in the field of Internet of Things, such as resource exploration, emergency rescue, intelligent transportation, etc. However, the communication network in an unknown environment may be inaccessible, and the real-time information sharing among UAVs cannot be guaranteed, resulting in the failure of cooperative search. Aiming at this issue, this article is devoted to the design of the multi-UAV flight strategy to improve the cooperative search capability in an uncertain communication environment. Specifically, a new cooperative architecture oriented to a local communication network is devised to control the observation locations of multiple UAVs in the search process, and some local communication networks are established based on the distance among UAVs to meet the requirements of the search task. On this foundation, we develop a multi-UAV cooperative search model (MCSM) with communication cost and formation benefit as an optimization function to ensure the effectiveness of multi-UAV search. Moreover, in the process of model solving, an improved sparrow search algorithm (ISSA) is presented with some different search strategies to enhance the optimization capability. To verify the superiority of the proposed method, we designed several groups of simulation experiments to analyze the performance of MCSM. Experimental results illustrate that our method can not only maintain high cooperative search accuracy but also has high stability and convergence speed.
Bowen Fei, Weidong Bao 0001, Xiaomin Zhu 0001, Daqian Liu, Tong Men, Zhenliang Xiao
IEEE Internet Things J.3
2022 Cooperative Path Optimization for Multiple UAVs Surveillance in Uncertain Environment
abstract
Research on multiple unmanned aerial vehicles (UAVs) cooperative surveillance systems serving Internet of Things (IoT) applications, such as smart cities, precision logistics, etc., has become a hot topic. However, the target movement is unpredictable in an uncertain environment, and multiple UAVs are affected by obstacles or inaccessible regions, resulting in the decreased surveillance performance and even the loss of the target. This article is dedicated to determine the current surveillance environment through the 2-D laser scanner. At the cost of the energy consumption and the transmission unreliability, a multi-UAV cooperative path optimization (MCPO) model is designed to adjust the surveillance location of each UAV, which improves the target surveillance performance. Specifically, for different types of obstacles or inaccessible regions, we present a novel obstacle-avoidance selection strategy with two mechanisms in mind: 1) when some of UAVs encounter obstacles, but others can accurately monitor the target, a strict constraint mechanism is established to promptly adjust the surveillance location of each UAV, which ensures the accuracy of formation surveillance and 2) when all UAVs have to avoid obstacles, a fuzzy constraint mechanism is presented and combined with Lucas–Kanade (LK) method to expand the search range of the multi-UAV and enhance the flexible adjustment capability of the formation. To verify the superiority of the proposed optimization method, we develop a 3-D simulation experiment environment based on the UE4 platform and design several groups of experiments to analyze the effectiveness of MCPO. The experimental results demonstrate that MCPO can not only maintain the flight stability of multiple UAVs but also has satisfactory formation flexibility and surveillance accuracy.
Daqian Liu, Weidong Bao 0001, Xiaomin Zhu 0001, Bowen Fei, Tong Men, Zhenliang Xiao
IEEE Internet Things J.3
2022 FLEE: A Hierarchical Federated Learning Framework for Distributed Deep Neural Network over Cloud, Edge, and End Device
abstract
With the development of smart devices, the computing capabilities of portable end devices such as mobile phones have been greatly enhanced. Meanwhile, traditional cloud computing faces great challenges caused by privacy-leakage and time-delay problems, there is a trend to push models down to edges and end devices. However, due to the limitation of computing resource, it is difficult for end devices to complete complex computing tasks alone. Therefore, this article divides the model into two parts and deploys them on multiple end devices and edges, respectively. Meanwhile, an early exit is set to reduce computing resource overhead, forming a hierarchical distributed architecture. In order to enable the distributed model to continuously evolve by using new data generated by end devices, we comprehensively consider various data distributions on end devices and edges, proposing a hierarchical federated learning framework FLEE , which can realize dynamical updates of models without redeploying them. Through image and sentence classification experiments, we verify that it can improve model performances under all kinds of data distributions, and prove that compared with other frameworks, the models trained by FLEE consume less global computing resource in the inference stage.
Zhengyi Zhong, Weidong Bao 0001, Ji Wang 0002, Xiaomin Zhu 0001, Xiongtao Zhang
ACM Trans. Intell. Syst. Technol.4
2022 SMART: Vision-Based Method of Cooperative Surveillance and Tracking by Multiple UAVs in the Urban Environment
abstract
UAV surveillance and tracking have attracted great enthusiasm in intelligent transportation, and various approaches have been reported up to now. However, these approaches often ignored the uncertainties in the urban environment, such as occlusion, view change, and background clutter. Ignoring these uncertain factors often leads to a reduction in surveillance performance and tracking quality. This study devotes to improving the cooperative surveillance capability of multi-UAV formation by designing different cooperative strategies in the urban environment. To be specific, a novel cooperative architecture is designed to control the observation locations of multiple UAVs throughout the formation process. For different types of interference, we introduce a novel target recognition rate of each UAV as the decision factor and design corresponding cooperative strategies to guarantee the accuracy of cooperative surveillance. Based on this architecture, we develop a vision-based method of cooperative surveillance and tracking by multiple UAVs (SMART) whose objective function is the motion cost and flight reliability of UAVs to ensure that each UAV can be in the optimal surveillance location for the target. The proposed SMART skillfully integrates the strict, elastic, and flight constraint strategies. During the execution of the multi-UAV formation, the inherent safety constraints of multiple UAVs and the designed strategies are used to solve the quadratic optimization model to adjust the locations of these UAVs. To demonstrate the superiority of our method, we conduct a 3D simulation urban environment and devise several experiments to analyze the performance of SMART on it. The experimental results demonstrate that SMART can not only maintain the high cooperative flight capability, but also provide high flexibility and fault tolerance.
Daqian Liu, Xiaomin Zhu 0001, Weidong Bao 0001, Bowen Fei, Jianhong Wu
IEEE Trans. Intell. Transp. Syst.2
2022 Introduction to the Special Issue on Multiagent Systems and Services in the Internet of Things
abstract
research-article Share on Introduction to the Special Issue on Multiagent Systems and Services in the Internet of Things Authors: Andrei Ciortea University of St. Gallen, Switzerland University of St. Gallen, Switzerland 0000-0003-0721-4135Search about this author , Xiaomin Zhu National University of Defense Technology, China National University of Defense Technology, China 0000-0003-1301-7840Search about this author , Calton Pu Georgia Institute of Technology, USA Georgia Institute of Technology, USA 0000-0002-6616-8987Search about this author , Munindar P. Singh North Carolina State University, USA North Carolina State University, USA 0000-0003-3599-3893Search about this author Authors Info & Claims ACM Transactions on Internet TechnologyVolume 22Issue 4November 2022 Article No.: 99pp 1–3https://doi.org/10.1145/3584744Published:03 March 2023Publication History 0citation0DownloadsMetricsTotal Citations0Total Downloads0Last 12 Months0Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Andrei Ciortea, Xiaomin Zhu 0001, Calton Pu, Munindar P. Singh
ACM Trans. Internet Techn.2
2022 DANCE: Distributed Generative Adversarial Networks with Communication Compression
abstract
Generative adversarial networks (GANs) have shown great success in deep representations learning, data generation, and security enhancement. With the development of the Internet of Things, 5th generation wireless systems (5G), and other technologies, the large volume of data collected at the edge of networks provides a new way to improve the capabilities of GANs. Due to privacy, bandwidth, and legal constraints, it is not appropriate to upload all the data to the cloud or servers for processing. Therefore, this article focuses on deploying and training GANs at the edge rather than converging edge data to the central node. To address this problem, we designed a novel distributed learning architecture for GANs, called DANCE. DANCE can adaptively perform communication compression based on the available bandwidth, while supporting both data and model parallelism training of GANs. In addition, inspired by the gossip mechanism and Stackelberg game, a compatible algorithm, AC-GAN is proposed. The theoretical analysis guarantees the convergence of the model and the existence of approximate equilibrium in AC-GAN. Both simulation and prototype system experiments show that AC-GAN can achieve better training effectiveness with less communication overhead than the SOTA algorithms, i.e., FL-GAN and MD-GAN.
Xiongtao Zhang, Xiaomin Zhu 0001, Ji Wang 0002, Weidong Bao 0001, Laurence T. Yang
ACM Trans. Internet Techn.2
2022 Elastic Resource Provisioning Using Data Clustering in Cloud Service Platform
abstract
Currently, cloud computing has received great attention in commerce and scientific research due to its flexibility and strong data processing capability. However, in view of the fact that the types of tasks display an upward trend as the growth of service demands, and the different types of tasks arrive at the system without regularity. Moreover, the resources deployed in cloud are insufficiency to be flexibly provisioned in the face of obvious workload fluctuations. In this article, we present a method of elastic resource provisioning using date clustering in cloud service platform. The framework of proposed method consists of three core components: tasks clustering, the amount of tasks prediction in cluster, dynamic resource provisioning and scheduling. In workload classification, we propose a clustering ensemble method, which utilizes a novel distance decision-making method to obtain the final results. Our method can effectively partition the arriving tasks into several clusters based on similarity among tasks. For each cluster, we forecast the amount of tasks arriving at next moment by prediction model based on time-series to provide reference for the follow-up resource provisioning. Afterwards, an energy-saving resource provisioning method is designed to dynamically provide resources for tasks in each cluster to meet their performance requirements. We implement the experiments in Google cloud traces dataset and the results show that our method achieves 92.3, 91.2 percent, and 3679.2 kW$ \cdot$·h respectively in terms of guarantee ratio, resource utilization and total energy consumption, which demonstrates the effectiveness of proposed method for dynamic resource provisioning.
Bowen Fei, Xiaomin Zhu 0001, Daqian Liu, Junjie Chen 0007, Weidong Bao 0001, Ling Liu 0001
IEEE Trans. Serv. Comput.2
2022 An Edge Storage Acceleration Service for Collaborative Mobile Devices
abstract
Fueled by the advances in the Internet of Things, and the growing capacity of smart mobile devices at the edge of the Internet, we have witnessed a growing trend in research and development for edge computing and edge storage, which extends the abilities of single mobile device on the edge through on-demand collaboration among multiple geographically distributed mobile devices. In this article, we address several technical challenges that are unique for collaborative storage at the edge due to the unique characteristics of mobile devices. First, we formalize the collaborative storage problem as an optimization problem. Second, we design an Acceleration Algorithm for Collaborative Storage, called A2CS, based on the architecture of Alternating Direction Method of Multipliers (ADMM). Specifically, we use the Nesterov’s Acceleration strategy and the step size rules in the process of updating variables and determining the optimal speed of convergence. We develop a novel collaborative storage policy in order to guide the whole lifecycle of collaborative storage. Finally, we conduct a series of experiments for acceleration performance analysis and validation. We show that A2CS delivers a better convergence performance with different step size rules, compared with two existing approaches: the ADMM baseline and the ADMM-OR (ADMM with Over-Relaxation), achieving the acceleration percentage by at least 25.33 percent and at most 64.01 percent. In addition, by conducting the utility performance comparison analysis with the existing Average Distribution Strategy (ADS) and the existing Distance Preferred Distribution Strategy (DPDS), we show the advantage of A2CS over both ADS and DPDS with respect to the total utility and energy consumption.
Xiong Gao, Weidong Bao 0001, Xiaomin Zhu 0001, Guanlin Wu, Ling Liu 0001
IEEE Trans. Serv. Comput.3
2022 Noisy Optimization by Evolution Strategies With Online Population Size Learning
abstract
Optimization modeling of real-world application problems usually involves noise from various sources. Noisy optimization imposes challenges to optimization methods since the objective values can be different for multiple evaluations. In this article, we propose a novel online population size learning (OPL) technique of evolution strategies for handling noisy optimization problems. By re-evaluating a fraction of the candidates, we measure the strength of noise level of the re-evaluated candidate solutions and adapt the population size according to the noise level. The proposed OPL combines the advantages of both explicit averaging by re-evaluations and the implicit averaging by large population size and overcomes their limitations. We incorporate it with the covariance matrix adaptation evolution strategy (CMA-ES) and obtain OPL-CMA-ES. Compared with the existing noise handling technique, the proposed OPL is much simpler in both concepts and computation. We conduct comprehensive experiments to evaluate the algorithm’s performance on standard problems with Gaussian noise. We further evaluate the performance of OPL-CMA-ES on the black-box optimization benchmarks (BBOBs) noisy testbed, which is a standard platform for comparing black-box optimization algorithms, compared with the state-of-the-art noise-handling algorithms. The experimental results show that OPL-CMA-ES achieves remarkable performance and outperforms the compared variants.
Zhenhua Li 0005, Xinye Cai, Qingfu Zhang 0001, Xiaomin Zhu 0001, Zhun Fan, Xiuyi Jia
IEEE Trans. Syst. Man Cybern. Syst.5
2021 Adaptive Clustering Ensemble Method Based on Uncertain Entropy Decision-Making
abstract
As an unsupervised data mining method, clustering can extract valuable information in complex and redundan-t network data analysis. However, the existing methods are sensitive to the selection of initial cluster centers, and cannot automatically determine the number of clusters, which fails to adapt to various types of network data. To solve these issues, this paper proposes a method of adaptive clustering ensemble based on uncertain entropy decision-making. Firstly, K-means is used as the base clustering algorithm of clustering ensemble, and several base clustering members are randomly generated according to different the number of clusters, and the members with high stability and quality are selected as clustering ensemble inputs. Furthermore, the uncertainty of clusters in the base clusterings are calculated based on the information entropy criterion, and then the co-association matrix is established. The obtained co-association matrix is transformed into a distance matrix by Bhattacharyya distance among data samples. Finally, we use the distance matrix as the input of the density peaks (DP) algorithm, and further calculate the final clustering result. The experimental results on real-world datasets illustrate that the proposed method has better performance than other clustering methods.
Xiaomin Zhu 0001, Bowen Fei, Daqian Liu, Weidong Bao 0001
TrustCom1
2021 Multi-UAV Cooperative Obstacle Avoidance and Surveillance in Intelligent Transportation
abstract
In intelligent transportation system, UAV surveillance plays an important role, and it has wide applications in traffic detection and order management, etc. However, the interference of extensive buildings and inaccessible regions in the urban environment directly lead to the failure of the surveillance task. Aiming at this issue, this paper proposes a method of multi-UAV UAV Cooperative Obstacle Avoidance and Surveillance (COAS). The ellipse tangent method is used to avoid obstacles for the interference of urban obstacles. Furthermore, taking into account the cooperation of multi-UAV formation, the cooperative model based on moving cost and formation stability is established. Due to the timeliness requirement of multi-UAV cooperative surveillance task, we use a sparrow search algorithm with fast convergence speed and strong optimization capability to solve the cooperative model. Finally, the simulation experimental results in an urban environment with obstacle information demonstrate the effectiveness of the proposed method in tackling the issues of cooperative obstacle avoidance and target surveillance.
Daqian Liu, Weidong Bao 0001, Bowen Fei, Xiaomin Zhu 0001, Zhenliang Xiao, Tong Men
TrustCom4
2021 ADAPT: Adaptive distributed optimization approach for uploading data with redundancy in cooperative mobile cloud
abstract
Summary With the development of information technology and the ubiquity of mobile devices, increasing amounts of data are generated, processed, and transmitted by mobile devices. To alleviate the tension between the energy poverty of mobile devices and the increasing demand for transmitting data, the energy‐efficient data transmission problem attracts considerable interests. Nonetheless, how to upload data with redundancy efficiently lacks a thorough study despite the wide existence of this problem in many situations like data storage among mobile devices and mobile crowd sensing. Since uploading redundant data brings little value while still consuming precious energy, it is important to design an efficient approach for mobile devices to upload data with redundancy cooperatively. In this work, we formulate the uploading data with redundancy in cooperative mobile cloud as an energy‐constrained utility maximization problem. To solve this problem, we propose an adaptive distributed optimization approach consisting of the correlated upload decision and the online distributed scheduling algorithm. By the correlated upload decision, each mobile device can make adaptive decisions on how much data to upload and which data to upload according to its own observations independently. The online distributed scheduling algorithm enables mobile devices to optimally upload data. A series of simulation experiments are conducted to demonstrate the effectiveness of our approach. Finally, we test our approach on a real demo system to verify its practicability in reality.
Ji Wang 0002, Weidong Bao 0001, Xiaomin Zhu 0001
Concurr. Comput. Pract. Exp.3
2021 EASE: Energy-efficient task scheduling for edge computing under uncertain runtime and unstable communication conditions
abstract
Summary Continuously growing network traffic has become a major technical bottleneck of the cloud service to develop the Internet of Things (IoTs) and mobile applications. Edge computing as a promising computing pattern deployed close to service users is expected to improve the quality of service (QoS). To fully utilize the capabilities of edge devices, a Device‐to‐Device (D2D)–based computing resource sharing and aggregation framework is proposed. Under this framework, this paper exploits the Beta distributions to characterize the uncertain communication rate and processing capability of the edge environment. The reliability and energy consumption of local computing and shared computing under uncertain conditions are, respectively, studied. We model the task scheduling as an Integer Programming problem, whose objective is to minimize the energy consumption while ensuring the reliability. Based on that, a heuristic task scheduling algorithm named EASE is proposed. Through a lot of simulation experiments, the performance of EASE is effectively evaluated under the static and dynamic environments. Compared with three comparison algorithms, EASE shows many advantages in terms of reliability, adaptability, and energy saving.
Xiaomin Zhu 0001, Dayu Zhang, Ji Wang 0002, Huangke Chen, Weidong Bao 0001
Concurr. Comput. Pract. Exp.3
2021 Distributed Learning on Mobile Devices: A New Approach to Data Mining in the Internet of Things
abstract
It is well known that deep learning is one of the most important methods for data mining. With the development of the fifth-generation mobile networks (5G) and the Internet of Things (IoT), the large volume of data collected in IoTs provides a new way to improve the capability of deep learning. Due to privacy, bandwidth, and legal concerns, it is impractical to send the data to a server or the cloud. The computing power of mobile devices makes it possible to process the data. Therefore, this article focuses on training these models in mobile devices. To solve the challenges, including unreliable networks, constrained resources, and slow convergence, we let multiple mobile devices learn a shared model collaboratively. We propose a novel architecture, GREAT, where each node chooses partners to share local model parameters according to link reliability. To balance the constrained resources and learning effectiveness, an optimization problem is developed by taking the reliability threshold as the variable of controlling the resources’ overhead. To implement this architecture, a dynamic control algorithm called Alpha-GossipSGD has been proposed. Its performance is evaluated by extensive experiments, which show that Alpha-GossipSGD can realize stable learning effectiveness over unreliable networks with constrained resources.
Xiongtao Zhang, Xiaomin Zhu 0001, Weidong Bao 0001, Laurence T. Yang, Ji Wang 0002, Huangke Chen
IEEE Internet Things J.2
2021 Uncertainty-Aware Online Scheduling for Real-Time Workflows in Cloud Service Environment
abstract
Scheduling workflows in cloud service environment has attracted great enthusiasm, and various approaches have been reported up to now. However, these approaches often ignored the uncertainties in the scheduling environment, such as the uncertain task start/execution/finish time, the uncertain data transfer time among tasks, the sudden arrival of new workflows. Ignoring these uncertain factors often leads to the violation of workflow deadlines and increases service renting costs of executing workflows. This study devotes to improving the performance for cloud service platforms by minimizing uncertainty propagation in scheduling workflow applications that have both uncertain task execution time and data transfer time. To be specific, a novel scheduling architecture is designed to control the count of workflow tasks directly waiting on each service instance (e.g., virtual machine and container). Once a task is completed, its start/execution/finish time are available, which means its uncertainties disappearing, and will not affect the subsequent waiting tasks on the same service instance. Thus, controlling the count of waiting tasks on service instances can prohibit the propagation of uncertainties. Based on this architecture, we develop an unceRtainty-aware Online Scheduling Algorithm (ROSA) to schedule dynamic and multiple workflows with deadlines. The proposed ROSA skillfully integrates both the proactive and reactive strategies. During the execution of the generated baseline schedules, the reactive strategy in ROSA will be dynamically called to produce new proactive baseline schedules for dealing with uncertainties. Then, on the basis of real-world workflow traces, five groups of simulation experiments are carried out to compare ROSA with five typical algorithms. The comparison results reveal that ROSA performs better than the five compared algorithms with respect to costs (up to 56 percent), deviation (up to 70 percent), resource utilization (up to 37 percent), and fairness (up to 37 percent).
Huangke Chen, Xiaomin Zhu 0001, Guipeng Liu, Witold Pedrycz
IEEE Trans. Serv. Comput.2
2021 An Adaptive Resource Allocation Strategy for Objective Space Partition-Based Multiobjective Optimization
abstract
In evolutionary computation, balancing the diversity and convergence of the population for multiobjective evolutionary algorithms (MOEAs) is one of the most challenging topics. Decomposition-based MOEAs are efficient for population diversity, especially when the branch partitions the objective space of multiobjective optimization problem (MOP) into a series of subspaces, and each subspace retains a set of solutions. However, a persisting challenge is how to strengthen the population convergence while maintaining diversity for decomposition-based MOEAs. To address this issue, we first define a novel metric to measure the contributions of subspaces to the population convergence. Then, we develop an adaptive strategy that allocates computational resources to each subspace according to their contributions to the population. Based on the above two strategies, we design an objective space partition-based adaptive MOEA, called OPE-MOEA, to improve population convergence, while maintaining population diversity. Finally, 41 widely used MOP benchmarks are used to compare the performance of the proposed OPE-MOEA with other five representative algorithms. For the 41 MOP benchmarks, the OPE-MOEA significantly outperforms the five algorithms on 28 MOP benchmarks in terms of the metric hypervolume.
Huangke Chen, Guohua Wu 0001, Witold Pedrycz, Ponnuthurai N. Suganthan, Lining Xing 0001, Xiaomin Zhu 0001
IEEE Trans. Syst. Man Cybern. Syst.6
2020 FILT: Optimizing KV-Embedded File Systems through Flat Indexing
abstract
The effectiveness of applying key-value store mechanisms to manage metadata of file systems has been demonstrated recently. However, traditional indirect metadata indexing schemes are not in concert with modern key-value data structures, which could degrade the performance of a KV-embedded file system due to the overhead of hierarchical path queries. In this paper, we propose FILT, a proof-of-concept file system middleware that can solve this problem by employing flat indexing. FILT exploits the benefits of both flat indexing and LSM-tree structure to eliminate redundant path lookups. Our extensive performance evaluation studies show that FILT can offer up to 5.8x performance gain compared with sophisticated local file systems.
Chen Chen 0124, Tongliang Deng, Jian Zhang 0070, Yanliang Zou, Xiaomin Zhu 0001, Shu Yin 0001
ICDCS5
2020 Benign: An Automatic Optimization Framework for the Logic of Swarm Behaviors
abstract
In the field of swarm intelligence, it is usually complicated to express the logic of swarm behaviors. Behavior tree has drawn a lot of attention to be a practical approach to solving this problem in recent years. However, how to automatically design the logic of swarm behaviors according to the target of a task is the focus of swarm intelligence. Hence, we propose an automatic optimizing framework named Benign which is capable of using gene expression programming (GEP) to optimize the logic of swarm behaviors. In Benign, the basic swarm behaviors and the relationships among those behaviors are mapped to nodes of behavior tree by the method named Matt firstly. With these nodes, we design an artificial behavior tree. After that, the artificial behavior tree is transformed into an expression tree in GEP according to the method named Meet. Finally, GEP is used for optimization to generate the expected logic of swarm behaviors. We conduct simulation experiments to validate the efficiency of Benign. The experimental results show the superiority of Benign. Compared with the logic of the artificial behavior tree before optimization, the conduction of the optimized logic of swarm behaviors increases efficiency by more than 50%.
Jingjing Tao, Xiaomin Zhu 0001, Weidong Bao 0001, Ji Wang 0002
SMC2
2020 Federated learning with adaptive communication compression under dynamic bandwidth and unreliable networks
Xiongtao Zhang, Xiaomin Zhu 0001, Ji Wang 0002, Huangke Chen, Weidong Bao 0001
Inf. Sci.2
2020 Minimal Fault-Tolerant Coverage of Controllers in IaaS Datacenters
abstract
Large-scale datacenters are the key infrastructures of cloud computing. Inside a datacenter, a large number of servers are interconnected using a specific datacenter network to deliver the infrastructure as a service (IaaS) for tenants. To realize novel cloud applications like the network virtualization and network isolation among tenants, the principle of software-defined network (SDN) has been applied to datacenters. In the setting, multiple distributed controllers are deployed to offer a control plane over the entire datacenter to efficiently manage the network usage. Despite such efforts, cloud datacenters, however, still lack a scalable and resilient control plane. Consequently, this paper systematically studies the coverage problem of controllers, which means to cover all network devices using the least number of controllers. More precisely, we tackle this essential problem from three aspects, including the minimal coverage, the minimal fault-tolerant coverage, and the minimal communication overhead among controllers. After modelling and analyzing such three problems, we design efficient approaches to approximate the optimal solution, respectively. Extensive evaluation results indicate that our approaches can significantly save the number of required controllers, improve the fault-tolerant capability of the control plane and reduce the communication overhead of state synchronization among controllers. The design methodologies proposed in this paper can be applied to cloud datacenters with other networking structures after minimal modifications.
Deke Guo, Xiaomin Zhu 0001, Bangbang Ren, Honghui Chen
IEEE Trans. Serv. Comput.3
2019 Private Model Compression via Knowledge Distillation
abstract
The soaring demand for intelligent mobile applications calls for deploying powerful deep neural networks (DNNs) on mobile devices. However, the outstanding performance of DNNs notoriously relies on increasingly complex models, which in turn is associated with an increase in computational expense far surpassing mobile devices’ capacity. What is worse, app service providers need to collect and utilize a large volume of users’ data, which contain sensitive information, to build the sophisticated DNN models. Directly deploying these models on public mobile devices presents prohibitive privacy risk. To benefit from the on-device deep learning without the capacity and privacy concerns, we design a private model compression framework RONA. Following the knowledge distillation paradigm, we jointly use hint learning, distillation learning, and self learning to train a compact and fast neural network. The knowledge distilled from the cumbersome model is adaptively bounded and carefully perturbed to enforce differential privacy. We further propose an elegant query sample selection method to reduce the number of queries and control the privacy loss. A series of empirical evaluations as well as the implementation on an Android mobile device show that RONA can not only compress cumbersome models efficiently but also provide a strong privacy guarantee. For example, on SVHN, when a meaningful (9.83,10−6)-differential privacy is guaranteed, the compact model trained by RONA can obtain 20× compression ratio and 19× speed-up with merely 0.97% accuracy loss.
Ji Wang 0002, Weidong Bao 0001, Lichao Sun 0001, Xiaomin Zhu 0001, Bokai Cao, Philip S. Yu
AAAI4
2019 DEFT: Dynamic Fault-Tolerant Elastic scheduling for tasks with uncertain runtime in cloud
Xiaomin Zhu 0001, Huangke Chen, Hui Guo 0001, Wen Zhou 0013, Weidong Bao 0001
Inf. Sci.2
2019 Cooperative Data Sharing for Mobile Cloudlets Under Heterogeneous Environments
abstract
As accessing remote cloud via cellular network is costly due to the lower bandwidth, higher wide area network (WAN) latency, and higher energy consumption, mobile cloudlet that formed by several edge mobile devices has become an emerging computing paradigm and attracted increasing attention recently. Being different from the existing studies that mainly focus on the issue of computation offloading among the peers, this paper investigates the problem of cooperative data sharing among peers to overcome the data dissymmetry, especially with the presence of dynamic network context. First, a publish/subscribe-based data sharing model is designed to cope with the unpredictable communication condition. Then, the data transmission scheduling within cooperative mobile devices is formulated as a utility maximization optimization considering the limited channel capacity, heterogeneous quality of experience (QoE) requirements, and incentive mechanism for participation. To encourage cooperation among mobile devices, a data downloading/uploading queuing mechanism is elegantly designed. Furthermore, an online algorithm without predicting the future information on request arrivals and network changes is developed to simultaneously optimize data transmission and communication interface selection in the long run. Theoretical analysis shows that the proposed algorithm is able to obtain a utility arbitrarily close to the offline optimum and guarantee the delay bound. Simulations demonstrate the effectiveness and the superiority of the proposed algorithm over some existing typical strategies.
Wenhua Xiao, Xiaomin Zhu 0001, Weidong Bao 0001, Ling Liu 0001
IEEE Trans. Netw. Serv. Manag.2
2019 An Attention-augmented Deep Architecture for Hard Drive Status Monitoring in Large-scale Storage Systems
abstract
Data centers equipped with large-scale storage systems are critical infrastructures in the era of big data. The enormous amount of hard drives in storage systems magnify the failure probability, which may cause tremendous loss for both data service users and providers. Despite a set of reactive fault-tolerant measures such as RAID, it is still a tough issue to enhance the reliability of large-scale storage systems. Proactive prediction is an effective method to avoid possible hard-drive failures in advance. A series of models based on the SMART statistics have been proposed to predict impending hard-drive failures. Nonetheless, there remain some serious yet unsolved challenges like the lack of explainability of prediction results. To address these issues, we carefully analyze a dataset collected from a real-world large-scale storage system and then design an attention-augmented deep architecture for hard-drive health status assessment and failure prediction. The deep architecture, composed of a feature integration layer, a temporal dependency extraction layer, an attention layer, and a classification layer, cannot only monitor the status of hard drives but also assist in failure cause diagnoses. The experiments based on real-world datasets show that the proposed deep architecture is able to assess the hard-drive status and predict the impending failures accurately. In addition, the experimental results demonstrate that the attention-augmented deep architecture can reveal the degradation progression of hard drives automatically and assist administrators in tracing the cause of hard drive failures.
Ji Wang 0002, Weidong Bao 0001, Lei Zheng 0001, Xiaomin Zhu 0001, Philip S. Yu
ACM Trans. Storage4
2019 Minimizing Traffic Migration During Network Update in IaaS Datacenters
abstract
The cloud datacenter network is consistently undergoing changing, due to a variety of topology and traffic updates, such as the VM migrations. Given an update event, prior methods focus on finding a sequence of lossless transitions from an initial network state to an end network state. They, however, suffer frequent and global search of the feasible end network states. This incurs non-trivial computation overhead and decision-making delay, especially in large-scale networks. Moreover, in each round of transition, prior methods usually cause the cascaded migrations of existing flows; hence, significantly disrupt production services in IaaS data centers. To tackle such severe issues, we present a simple update mechanism to minimize the amount of flow migrations during the congestion-free network update. The basic idea is to replace performing the sequence of global transitions of network states with local reschedule of involved flows, caused by an update event. We first model all involved flows due to an update event as a set of new flows, and then propose a heuristic method Lupdate. It motivates to locally schedule each new flow into the shortest path, at the cost of causing the extra migration of at most one existing flow if needed. To minimize the amount of migrated traffic, the migrated flow should be as small as possible. To further improve the success rate, we propose an enhanced method Lupdate-S. It shares the similar design of Lupdate, but permits to migrate multiple necessary flows on the shortest path allocated to each new flow. We conduct large-scale trace-driven evaluations under widely used Fat-Tree and ER data centers. The experimental results indicate that our methods can realize congestion-free network with as less amount of traffic migration as possible even when the link utilization of a majority of links is very high. The amount of traffic migration caused by our Ludpate method is 1.2 times and 1.12 times of the optimal result in the Fat-Tree and ER random networks, respectively.
Ting Qu 0003, Deke Guo, Yulong Shen 0001, Xiaomin Zhu 0001, Lailong Luo, Zhong Liu 0002
IEEE Trans. Serv. Comput.4
2018 WITCAT: A Workload Spike Targeted Cloud Management Solution
abstract
The cloud computing technology offers consistent access to large-scale computing capabilities, thereby bringing convenience to life. However, the virtualized cloud systems are still too vulnerable to maintain performance scalability and service agility once a task burst surges in without any warning. A mounting account of research has been conducted on proper strategies for accurate workload prediction as well as effective resource reservation and arrangement, but commonly cloud providers seek help to strategies that deploy excessive resources, adding overhead cost and sacrificing the cloud's advantage of scalability, or otherwise fail to reconfigure timely and properly, causing dissatisfaction and even financial loss, which are not expected by both cloud providers and clients.
Junjie Chen 0007, Xiaomin Zhu 0001, Weidong Bao 0001, Zhong Liu 0002, Ling Liu 0001
SoCC2
2018 A Parallel Fast Fourier Transform Algorithm for Large-Scale Signal Data Using Apache Spark in Cloud
Weidong Bao 0001, Xiaomin Zhu 0001, Ji Wang 0002, Wenhua Xiao
ICA3PP (3)3
2018 PEA: Parallel Evolutionary Algorithm by Separating Convergence and Diversity for Large-Scale Multi-Objective Optimization
abstract
Running evolutionary algorithms in parallel is an intuitive way to speed up the process of solving large-scale multi-objective optimization problems, which have hundreds or thousands of decision variables. However, the framework of the existing multi-objective evolutionary algorithms seriously limits their parallelization. During each iteration, the environmental selection operators present in the existing framework need to collect and compare all the candidate solutions to balance the convergence and diversity, thus dividing the whole evolutionary process into a series of dependent sub-processes and resulting in frequent data transmission. To address this issue, we propose a novel parallel framework that separates the environmental selection operator from the entire evolutionary process, evidently removing the dependencies among sub-processes and reducing the data transmission. On the basis of the parallel framework, a new parallel evolutionary algorithm, namely PEA, is designed. In PEA, the convergence is achieved by a series of independent sub-populations, and the diversity is merely emphasized at the converged solutions from each subpopulation, which is helpful for avoiding that the environmental selection operator limits the parallelization of the algorithm. Moreover, a new environmental selection strategy is proposed to improve the diversity without considering the convergence. To assess the performance of the proposed PEA, we compare it with five representative multi-objective evolutionary algorithms in terms of both the convergence and diversity. The performance of the parallel framework is also analyzed by comparing with two existing parallel models. The experimental results demonstrate the superiority of the proposed parallel algorithms in terms of the convergence, diversity, and speedup.
Huangke Chen, Xiaomin Zhu 0001, Witold Pedrycz, Shu Yin 0001, Guohua Wu 0001
ICDCS2
2018 Deep Learning towards Mobile Applications
abstract
Recent years have witnessed an explosive growth of mobile devices. Mobile devices are permeating every aspect of our daily lives. With the increasing usage of mobile devices and intelligent applications, there is a soaring demand for mobile applications with machine learning services. Inspired by the tremendous success achieved by deep learning in many machine learning tasks, it becomes a natural trend to push deep learning towards mobile applications. However, there exist many challenges to realize deep learning in mobile applications, including the contradiction between the miniature nature of mobile devices and the resource requirement of deep neural networks, the privacy and security concerns about individuals' data, and so on. To resolve these challenges, during the past few years, great leaps have been made in this area. In this paper, we provide an overview of the current challenges and representative achievements about pushing deep learning on mobile devices from three aspects: training with mobile data, efficient inference on mobile devices, and applications of mobile deep learning. The former two aspects cover the primary tasks of deep learning. Then, we go through our two recent applications that apply the data collected by mobile devices to inferring mood disturbance and user identification. Finally, we conclude this paper with the discussion of the future of this area.
Ji Wang 0002, Bokai Cao, Philip S. Yu, Lichao Sun 0001, Weidong Bao 0001, Xiaomin Zhu 0001
ICDCS6
2018 Not Just Privacy: Improving Performance of Private Deep Learning in Mobile Cloud
abstract
The increasing demand for on-device deep learning services calls for a highly efficient manner to deploy deep neural networks (DNNs) on mobile devices with limited capacity. The cloud-based solution is a promising approach to enabling deep learning applications on mobile devices where the large portions of a DNN are offloaded to the cloud. However, revealing data to the cloud leads to potential privacy risk. To benefit from the cloud data center without the privacy risk, we design, evaluate, and implement a cloud-based framework ARDEN which partitions the DNN across mobile devices and cloud data centers. A simple data transformation is performed on the mobile device, while the resource-hungry training and the complex inference rely on the cloud data center. To protect the sensitive information, a lightweight privacy-preserving mechanism consisting of arbitrary data nullification and random noise addition is introduced, which provides strong privacy guarantee. A rigorous privacy budget analysis is given. Nonetheless, the private perturbation to the original data inevitably has a negative impact on the performance of further inference on the cloud side. To mitigate this influence, we propose a noisy training method to enhance the cloud-side network robustness to perturbed data. Through the sophisticated design, ARDEN can not only preserve privacy but also improve the inference performance. To validate the proposed ARDEN, a series of experiments based on three image datasets and a real mobile application are conducted. The experimental results demonstrate the effectiveness of ARDEN. Finally, we implement ARDEN on a demo system to verify its practicality.
Ji Wang 0002, Jianguo Zhang 0005, Weidong Bao 0001, Xiaomin Zhu 0001, Bokai Cao, Philip S. Yu
KDD4
2018 DuoFS: A Hybrid Storage System Balancing Energy-Efficiency, Reliability, and Performance
abstract
As the Energy Wall and the Reliability Wall become unavoidable, it is a demanding and challenging task to reduce energy consumption in large-scale storage systems in modern data centers while retaining acceptable systems reliability. We propose a reliable energy-efficient storage system called DuoFS, which aims at balancing the energy efficiency, the reliability and the performance of parallel storage systems by seamlessly integrating one HDD-based file system and one SSD-based file system. At the heart of the DuoFS is a transformative middleware layer that dispatches files to the one of the two independent parallel file systems based on the files' I/O access popularity. By replicating popular files to the SSD-based file system and pushing the HDD-based file system into the low-power mode under light workload conditions, DuoFS can reduce significant energy consumption, avoid major factors that harm the storage systems reliability, and extract SSDs good I/O performance. Experimental results show that the DuoFS system saves up to 40% of energy, achieves up to 50% better I/O performance while only sacrificing less than 15% of the system's reliability.
Shu Yin 0001, Bing Jiao, Xiaomin Zhu 0001, Xiaojun Ruan, Si Chen 0009, Zhuo Tang
PDP3
2018 A server consolidation method with integrated deep learning predictor in local storage based clouds
abstract
Summary Server consolidation is one of the critical techniques for energy‐efficiency in cloud data centers. As it is often assumed that cloud service instances (eg, Amazon EC2 instances) utilize the shared storage only. In recent years, however, cloud service providers have been providing local storage for cloud users, since local storage can offer a better performance with identified price. However, these cloud instances usually contain much more data than shared storage cloud instances. Thus, in such local storage based cloud center, the migration cost can be really high and is in dire need of an efficient resource pre‐allocation. If we can predict the resource demand in advance, the migration oscillation will be reduced to minify the migration cost. We have found that there are some related work about server consolidation based on forecasting. Unfortunately, their latest work did not consider the background of “local storage” as we mentioned above. At the same time, some research about local storage did not involve the prediction strategy, which plays a significant part in server consolidation. To address this issue, this paper proposes Losari, a consolidation method, which takes numeric forecasting and local storage architecture into consideration. Losari consolidates servers on the basis of the resource demand predicted value using a statistical learning method. We model the workload from real cloud production environment as a time series. Taking deep learning as a frame of reference, multiple deep belief networks integrated with ARIMA model was trained to study the feature of historical workload. The experimental results have showed that its average predicted error is only 10.7% in the short term, which is much lower than the most common model based on threshold (19.8%) on the same dataset. What is more, the results show that Losari not only simulates the true sequences in high accuracy but also scales the compute resource well, which demonstrated the validity of this integrated deep learning model.
Weidong Bao 0001, Xiaomin Zhu 0001, Huining Yan
Concurr. Comput. Pract. Exp.3
2018 SP-Partitioner: A novel partition method to handle intermediate data skew in spark streaming
Guipeng Liu, Xiaomin Zhu 0001, Ji Wang 0002, Deke Guo, Weidong Bao 0001, Hui Guo 0001
Future Gener. Comput. Syst.2
2017 DuoFS: An Attempt at Energy-Saving and Retaining Reliability of Storage Systems
abstract
As issues of the Energy Wall and the Reliability Wall become unavoidable, it is a demanding and challenging task to reduce energy consumption in large-scale storage systems in modern data centres while retaining acceptable systems reliability. Most energy conservation techniques inevitably have adverse impacts on the parallel disk systems. To address the reliability issues of energy-efficient parallel storage systems, we propose a reliable energy-efficient storage system called DuoFS, which aims at improving both energy efficiency and reliability of parallel storage systems by seamlessly integrating HDDs and SSDs. With the help of the middleware layer, DuoFS can distribute popular data to SSD-based nodes and put HDD-based nodes into the low-power mode under light workload conditions without modification of the parallel systems.
Bing Jiao, Xiaomin Zhu 0001, Xiaojun Ruan, Xiao Qin 0001, Shu Yin 0001
ICDCS2
2017 An Event-Level Abstraction for Achieving Efficiency and Fairness in Network Update
abstract
Changes of network state are a common source of instability in networks. An update event typically involves multiple flows that compete for network resources at the cost of rescheduling and migrating some existing flows. Previous network updating schemes tackle such flows independently, rather than as the entity of an update event. They only optimize the flow-level metrics for the flows involved in an update event. In this paper, we present an event-level abstraction of network update which groups flows of an update event and schedules them together to minimize the event completion time (ECT). We then study the scheduling problem of multiple update events for achieving high scheduling efficiency and preserving fairness. The designed least migration traffic first (LMTF) method schedules all update events in the FIFO order, but avoids head-of-line blocking by randomly fine-tuning the queue order of some events. It can considerably reduce the update cost, the average, and tail ECTs of all update events. In addition, we design a general parallel-LMTF (P-LMTF) method to guarantee fairness and further improve scheduling efficiency among update events. It improves the LMTF method by opportunistically updating multiple events simultaneously. The comprehensive evaluation results indicate that the average ECT of our approach is up to 10× faster than the flow-level scheduling method for network update events, and its tail ECT is up to 6x faster. Our P-LMTF method incurs 75% reduction in the average ECT compared with FIFO when the network utilization exceeds 70%, and it achieves a 42% reduction in tail ECT.
Ting Qu 0003, Deke Guo, Xiaomin Zhu 0001, Jie Wu 0001, Xiaolei Zhou 0001, Zhong Liu 0002
ICDCS3
2017 A Lightweight Recommendation Framework for Mobile User's Link Selection in Dense Network
abstract
With the proliferation of mobile devices and the development of communication technology, mobile devices have permeated every aspect of our daily lives. However, in dense network where large crowd of mobile devices try to access to the network simultaneously, the severe interference between mobile devices may incur a remarkable deterioration of the wireless communication quality. How to improve individual's experience in such scenario is a critical yet open problem. Inspired by the mobile device users' usage pattern as well as the characteristic of most wireless communication systems, we propose a framework offering uplink/downlink selection recommendation to different mobile device users to enhance their utility in this paper. The design of the framework starts with formulating the problem as a link selection game. Analysis shows that the game can be categorized as a generalized ordinal potential game whose Nash Equilibrium is guaranteed. We then devise a distributed link selection algorithm to generate a Nash Equilibrium of the game. To accommodate to the characteristic of dense network and the capacity limitation of mobile device, the design of the algorithm shows a light-weight property and does not require each mobile device user to know others' current selection. The probability of incomplete information gathering is also considered. Extensive experiments are conducted to demonstrate the effectiveness and superiority of the proposed framework. Experimental results show that the global average utility increase rate reaches above 20%, and about 70% mobile device users can benefit from using our framework.
Ji Wang 0002, Xiaomin Zhu 0001, Weidong Bao 0001, Guanlin Wu
ICDCS2
2017 MidHDC: Advanced topics on middleware services for heterogeneous distributed computing. Part 2
Florin Pop, Xiaomin Zhu 0001, Laurence T. Yang
Future Gener. Comput. Syst.2
2017 Chord: Checkpoint-based scheduling using hybrid waiting list in shared clusters
Yiyang Shao, Weidong Bao 0001, Xiaomin Zhu 0001, Wenhua Xiao, Jian Wang 0105
J. Syst. Softw.3
2017 Towards collaborative storage scheduling using alternating direction method of multipliers for mobile edge cloud
Guanlin Wu, Junjie Chen 0007, Weidong Bao 0001, Xiaomin Zhu 0001, Wenhua Xiao, Ji Wang 0002
J. Syst. Softw.4
2017 Scheduling for Workflows with Security-Sensitive Intermediate Data by Selective Tasks Duplication in Clouds
abstract
With the wide deployment of cloud computing in many business enterprises as well as science and engineering domains, high quality security services are increasingly critical for processing workflow applications with sensitive intermediate data. Unfortunately, most existing worklfow scheduling approaches disregard the security requirements of the intermediate data produced by workflows, and overlook the performance impact of encryption time of intermediate data on the start of subsequent workflow tasks. Furthermore, the idle time slots on resources, resulting from data dependencies among workflow tasks, have not been adequately exploited to mitigate the impact of data encryption time on workflows' makespans and monetary cost. To address these issues, this paper presents a novel task-scheduling framework for security sensitive workflows with three novel features. First, we provide comprehensive theoretical analyses on how selectively duplicating a task's predecessor tasks is helpful for preventing both the data transmission time and encryption time from delaying task's start time. Then, we define workflow tasks' latest finish time, and prove that tasks can be completed before tasks' latest finish time by using cheapest resources to reduce monetary cost without delaying tasks' successors' start time and workflows' makespans. Based on these analyses, we devise a novel scheduling approach with selective tasks duplication, named SOLID, incorporating two important phases: 1) task scheduling with selectively duplicating predecessor tasks to idle time slots on resources; and 2) intermediate data encrypting by effectively exploiting tasks' laxity time. We evaluate our solution approach through rigorous performance evaluation study using both randomly generated workflows and some real-world workflow traces. Our results show that the proposed SOLID approach prevails over existing algorithms in terms of makespan, monetary costs and resource efficiency.
Huangke Chen, Xiaomin Zhu 0001, Dishan Qiu, Ling Liu 0001, Zhihui Du
IEEE Trans. Parallel Distributed Syst.2
2017 Cost-Aware Big Data Processing Across Geo-Distributed Datacenters
abstract
With the globalization of service, organizations continuously produce large volumes of data that need to be analysed over geo-dispersed locations. Traditionally central approach that moving all data to a single cluster is inefficient or infeasible due to the limitations such as the scarcity of wide-area bandwidth and the low latency requirement of data processing. Processing big data across geo-distributed datacenters continues to gain popularity in recent years. However, managing distributed MapReduce computations across geo-distributed datacenters poses a number of technical challenges: how to allocate data among a selection of geo-distributed datacenters to reduce the communication cost, how to determine the Virtual Machine (VM) provisioning strategy that offers high performance and low cost, and what criteria should be used to select a datacenter as the final reducer for big data analytics jobs. In this paper, these challenges is addressed by balancing bandwidth cost, storage cost, computing cost, migration cost, and latency cost, between the two MapReduce phases across datacenters. We formulate this complex cost optimization problem for data movement, resource provisioning and reducer selection into a joint stochastic integer nonlinear optimization problem by minimizing the five cost factors simultaneously. The Lyapunov framework is integrated into our study and an efficient online algorithm that is able to minimize the long-term time-averaged operation cost is further designed. Theoretical analysis shows that our online algorithm can provide a near optimum solution with a provable gap and can guarantee that the data processing can be completed within pre-defined bounded delays. Experiments on WorldCup98 web site trace validate the theoretical analysis results and demonstrate that our approach is close to the offline-optimum performance and superior to some representative approaches.
Wenhua Xiao, Weidong Bao 0001, Xiaomin Zhu 0001, Ling Liu 0001
IEEE Trans. Parallel Distributed Syst.3
2016 Uncertainty-Aware Real-Time Workflow Scheduling in the Cloud
abstract
Scheduling real-time workflows running in the Cloud often need to deal with uncertain task execution time and minimize uncertainty propagation during the workflow runtime. Efficient scheduling approaches can minimize the operational cost of Cloud providers and provide higher guarantee of the quality of services (QoS) for Cloud consumers. However, most of the existing workflow scheduling approaches are designed for the individual workflow runtime environments that are deterministic. Such static workflow schedulers are inadequate for multiple and dynamic workflows, each with possibly uncertain task execution time. In this paper, we address the problem of minimizing uncertainty propagation in real-time workflow scheduling. We first introduce an uncertainty-aware scheduling architecture to mitigate the impact of uncertainty factors on the quality of workflow schedules. Then we present a dynamic workflow scheduling algorithm (PRS) that can dynamically exploit proactive and reactive scheduling methods. Finally, we conduct extensive experiments using real-world workflow traces and our experimental results show that PRS outperforms two representative scheduling algorithms in terms of costs (up to 60%), resource utilization (up to 40%) and deviation (up to 70%).
Huangke Chen, Xiaomin Zhu 0001, Dishan Qiu, Ling Liu 0001
CLOUD2
2016 A Utility-Aware Approach to Redundant Data Upload in Cooperative Mobile Cloud
abstract
With the proliferation of mobile devices and the improvement of wireless communication technology, an increasing number of mobile devices are utilized for emergency management and healthcare monitoring. Redundant data upload to the cloud datacenters is gaining growing interest and attraction. One of the main challenges for redundant data upload in the cooperative mobile cloud is the optimization problem of how to provide high utility and high energy efficiency for data upload in the presence of intermittent connectivity and unpredictable bandwidth of wireless and mobile network. In this paper, we formulate the problem of redundant data upload in the cooperative mobile cloud as an energy-constrained utility maximization problem that aims at maximizing the amount of effective data uploaded under the energy consumption constraints. We propose an online distributed approach to enabling mobile devices to optimally make upload decisions without depending on the current state information of other devices and the prior knowledge of its own future context. We provide a rigorous theoretical analysis and an extensive suite of simulation experiments to demonstrate the effectiveness and superiority of our approach.
Ji Wang 0002, Xiaomin Zhu 0001, Weidong Bao 0001, Ling Liu 0001
CLOUD2
2016 General Framework for Task Scheduling and Resource Provisioning in Cloud Computing Systems
abstract
Clouds have become an important platform to deliver services for various applications. Task scheduling and resource provisioning are key components to improve system performance under provisioned resources and satisfy users' demands for quality of service (QoS). To address the diversity of cloud services and applications, much of recent research and development efforts have been engaged in designing and implementing scheduling strategies and algorithms for specific tasks, such as dependent or independent tasks, fault-tolerant tasks with real-time deadlines or energy-efficient tasks. However, these task scheduling and resource provisioning schemes, though optimized with specific objectives, suffer from several inherent problems in cloud execution environments. In this paper, we propose a general framework for task scheduling and resource provisioning in cloud computing systems with dynamic customizability. By utilizing software engineering framework as the design guideline, we incorporate multiple scheduling objectives and multiple types of tasks to be processed under varied resource constraints to enable cloud applications to dynamically select and assemble scheduling strategies and algorithms according to different runtime QoS requirements. We illustrate the flexibility and customizability of our framework through two example scheduling algorithms: EASU and RAS. We validate the ffectiveness of our proposed framework through experimental evaluation of the effectiveness of our proposed algorithms using both simulation and in real cloud platforms.
Xiaomin Zhu 0001, Yabing Zha, Ling Liu 0001, Peng Jiao
COMPSAC1
2016 CHIME: A Checkpoint-Based Approach to Improving the Performance of Shared Clusters
abstract
Due to the limitation of resources, preemption frequently occurs in almost all the commercial cloud platforms, such as Google cluster and Amazon cluster. Since preemption can ensure that once the system is in heavy workload, high-priority tasks will be executed primarily and at the same time, some low-priority tasks will be killed immediately. Then when more resources are available, the killed tasks will restart to execute. Especially, during the peak time, some low-priority tasks could possibly be preempted and restarted repeatedly resulting in much more consuming precious resources including CPU cores, RAM and hard drives. Thanks to the checkpoint technology, it provides an efficient solution to addressing the preemption issue. But checkpoint technology has limitations, e.g., making checkpoint frequently will add redundant overhead to the cluster and cause I/O congestion. In this paper, by leveraging checkpoint technology, we designed a novel approach to improving the performance of shared clusters. Specifically, by checking the occupancy of resources periodically, making decisions to checkpoint or not and checkpointing for certain tasks, our method can reduce unnecessary checkpoints and exalt the performance of the whole cloud, especially tasks with low-priority. Extensive simulation experiments injecting tasks following the Google cloud trace logs were conducted to validate the superiority of our approach by comparing it with some baselines.
Yiyang Shao, Xiaomin Zhu 0001, Weidong Bao 0001, Wen Zhou 0013, Wenhua Xiao
ICPADS2
2016 Improving the Performance of Data Sharing in Dynamic Peer-to-Peer Mobile Cloud
abstract
Mobile cloud computing has become an emerging computing paradigm to extend the capability of the mobile devices and it has gained increasing popularity in recent years. Existing studies mainly focus on how to leverage the computing capability of the individual device by employing the capability from remote cloud datacenters or local mobile cloud formed by nearby devices. Different from these studies, we investigate how to improve the performance of data sharing in the peer-to-peer mobile cloud, with the limited bandwidth and the presence of dynamic and unpredictable wireless channel state. Specifically, we first formulate the data transmission among devices as a utility maximization problem with the consideration of limited bandwidth, incentive participation and the QoE (Quality of Experience) heterogeneity, based on incorporating publish/subscribe component into the base station. Then, a dynamic online algorithm, which does not need the future context (e.g., channel state) of the mobile cloud, is developed to simultaneously make the decision of data transmission and communication interface selection. Rigorously theoretical analysis shows the optimality and the effectiveness of the proposed algorithm. Extensive experiments are conducted to verify the analysis results and the superiority of the proposed algorithm over existing strategies.
Wenhua Xiao, Weidong Bao 0001, Xiaomin Zhu 0001, Wen Zhou 0013, Peizhong Lu
ICPADS3
2016 RESS: A Reliable Energy-Efficient Storage System
abstract
Extracting high I/O performance from parallel file systems is no longer the only goal in modern data centres. As issues of the Energy Wall and the Reliability Wall become unavoidable, it is a demanding and challenging task to reduce energy consumption in large-scale storage systems in modern data centres while retaining acceptable systems reliability. Most energy conservation techniques inevitably have adverse impacts on the parallel disk systems. To address the reliability issues of energy-efficient parallel storage systems, we propose a reliable energy-efficient storage system called RESS, which aims at improving both energy efficiency and reliability of parallel storage systems by seamlessly integrating HDDs and SSDs. At the heart of the RESS is a transformative middleware layer, which reorganizes the I/O workload for the underlying parallel file systems. With the help of the middleware layer, RESS can distribute popular data to SSDs and put HDDs into the low-power mode under light workload conditions without modification of the parallel systems.
Shu Yin 0001, Zhaoyu Xiao, Kenli Li 0001, Jianzhong Huang 0001, Xiaojun Ruan, Xiaomin Zhu 0001, Xiao Qin 0001
ICPADS6
2016 MidHDC: Advanced Topics on Middleware Services for Heterogeneous Distributed Computing. Part 1
Florin Pop, Xiaomin Zhu 0001, Laurence T. Yang
Future Gener. Comput. Syst.2
2016 Dynamic Request Redirection and Resource Provisioning for Cloud-Based Video Services under Heterogeneous Environment
abstract
Cloud computing provides a new opportunity for Video Service Providers (VSP) to running compute-intensive video applications in a cost effective manner. Under this paradigm, a VSP may rent virtual machines (VMs) from multiple geo-distributed datacenters that are close to video requestors to run their services. As user demands are difficult to predict and the prices of the VMs vary in different time and region, optimizing the number of VMs of each type rented from datacenters located in different regions in a given time frame becomes essential to achieve cost effectiveness for VSPs. Meanwhile, it is equally important to guarantee users' Quality of Experience (QoE) with rented VMs. In this paper, we give a systematic method called Dynamical Request Redirection and Resource Provisioning (DYRECEIVE) to address this problem. We formulate the problem as a stochastic optimization problem and design a Lyapunov optimization framework based online algorithm to solve it. Our method is able to minimize the long-term time average cost of renting cloud resources while maintaining the user QoE. Theoretical analysis shows that our online algorithm can produce a solution within an upper bound to the optimal solution achieved through offline computing. Extensive experiments shows that our method is adaptive to request pattern changes along time and outperforms existing algorithms.
Wenhua Xiao, Weidong Bao 0001, Xiaomin Zhu 0001, Chen Wang 0008, Lidong Chen, Laurence T. Yang
IEEE Trans. Parallel Distributed Syst.3
2016 Fault-Tolerant Scheduling for Real-Time Scientific Workflows with Elastic Resource Provisioning in Virtualized Clouds
abstract
Clouds are becoming an important platform for scientific workflow applications. However, with many nodes being deployed in clouds, managing reliability of resources becomes a critical issue, especially for the real-time scientific workflow execution where deadlines should be satisfied. Therefore, fault tolerance in clouds is extremely essential. The PB (primary backup) based scheduling is a popular technique for fault tolerance and has effectively been used in the cluster and grid computing. However, applying this technique for real-time workflows in a virtualized cloud is much more complicated and has rarely been studied. In this paper, we address this problem. We first establish a real-time workflow fault-tolerant model that extends the traditional PB model by incorporating the cloud characteristics. Based on this model, we develop approaches for task allocation and message transmission to ensure faults can be tolerated during the workflow execution. Finally, we propose a dynamic fault-tolerant scheduling algorithm, FASTER, for realtime workflows in the virtualized cloud. FASTER has three key features: 1) it employs a backward shifting method to make full use of the idle resources and incorporates task overlapping and VM migration for high resource utilization, 2) it applies the vertical/horizontal scaling-up technique to quickly provision resources for a burst of workflows, and 3) it uses the vertical scaling-down scheme to avoid unnecessary and ineffective resource changes due to fluctuated workflow requests. We evaluate our FASTER algorithm with synthetic workflows and workflows collected from the real scientific and business applications and compare it with six baseline algorithms. The experimental results demonstrate that FASTER can effectively improve the resource utilization and schedulability even in the presence of node failures in virtualized clouds.
Xiaomin Zhu 0001, Ji Wang 0002, Hui Guo 0001, Dakai Zhu 0001, Laurence T. Yang, Ling Liu 0001
IEEE Trans. Parallel Distributed Syst.1
2015 REED: A Reliable Energy-Efficient RAID
abstract
Recent studies indicate that the energy cost and carbon footprint of data centers have become exorbitant. It is a demanding and challenging task to reduce energy consumption in large-scale storage systems in modern data centers. Most energy conservation techniques inevitably have adverse impacts on parallel disk systems. To address the reliability issues of energy-efficient parallel disks, we propose a reliable energy-efficient RAID system called REED, which aims at improving both energy efficiency and reliability of RAID systems by seamlessly integrating HDDs and SSDs. At the heart of REED is a high-performance cache mechanism powered by SSDs, which are serving popular data. Under light workload conditions, REED spins down HDDs into the low-power mode, thereby offering energy conservation. Importantly, during an I/O access turbulence (i.e., I/O load is dynamically and frequently changing), REED is conducive to reducing the number of disk power-state transitions by keeping HDDs in the low-power mode while serving requests with SSDs. We build a model to quantitatively show that REED is capable of improving the reliability of energy-efficient RAIDs. We implement the REED prototype in a real-world RAID-0 system. Our experimental results demonstrate that REED improves the energy-efficiency of conventional RAID-0 by up to 73% while maintaining good reliability.
Shu Yin 0001, Xuewu Li, Kenli Li 0001, Jianzhong Huang 0001, Xiaojun Ruan, Xiaomin Zhu 0001, Wei Cao 0006, Xiao Qin 0001
ICPP6
2015 Towards dynamic real-time scheduling for multiple earth observation satellites
Jianjiang Wang, Xiaomin Zhu 0001, Laurence T. Yang, Jianghan Zhu, Manhao Ma
J. Comput. Syst. Sci.2
2015 Towards energy-efficient scheduling for real-time tasks under uncertain cloud computing environment
Huangke Chen, Xiaomin Zhu 0001, Hui Guo 0001, Jianghan Zhu, Xiao Qin 0001, Jianhong Wu
J. Syst. Softw.2
2015 SGEESS: Smart green energy-efficient scheduling strategy with dynamic electricity price for data center
Hongtao Lei, Tao Zhang 0033, Yabing Zha, Xiaomin Zhu 0001
J. Syst. Softw.5
2015 FESTAL: Fault-Tolerant Elastic Scheduling Algorithm for Real-Time Tasks in Virtualized Clouds
abstract
As clouds have been deployed widely in various fields, the reliability and availability of clouds become the major concern of cloud service providers and users. Thereby, fault tolerance in clouds receives a great deal of attention in both industry and academia, especially for real-time applications due to their safety critical nature. Large amounts of researches have been conducted to realize fault tolerance in distributed systems, among which fault-tolerant scheduling plays a significant role. However, few researches on the fault-tolerant scheduling study the virtualization and the elasticity, two key features of clouds, sufficiently. To address this issue, this paper presents a fault-tolerant mechanism which extends the primary-backup model to incorporate the features of clouds. Meanwhile, for the first time, we propose an elastic resource provisioning mechanism in the fault-tolerant context to improve the resource utilization. On the basis of the fault-tolerant mechanism and the elastic resource provisioning mechanism, we design novel fault-tolerant elastic scheduling algorithms for real-time tasks in clouds named FESTAL, aiming at achieving both fault tolerance and high resource utilization in clouds. Extensive experiments injecting with random synthetic workloads as well as the workload from the latest version of the Google cloud tracelogs are conducted by CloudSim to compare FESTAL with three baseline algorithms, i.e., Non-M igration-FESTAL (NMFESTAL), Non-Overlapping-FESTAL (NOFESTAL), and Elastic First Fit (EFF). The experimental results demonstrate that FESTAL is able to effectively enhance the performance of virtualized clouds.
Ji Wang 0002, Weidong Bao 0001, Xiaomin Zhu 0001, Laurence T. Yang, Yang Xiang 0001
IEEE Trans. Computers3
2015 ANGEL: Agent-Based Scheduling for Real-Time Tasks in Virtualized Clouds
abstract
The success of cloud computing makes an increasing number of real-time applications such as signal processing and weather forecasting run in the cloud. Meanwhile, scheduling for real-time tasks is playing an essential role for a cloud provider to maintain its quality of service and enhance the system's performance. In this paper, we devise a novel agent-based scheduling mechanism in cloud computing environment to allocate real-time tasks and dynamically provision resources. In contrast to traditional contract net protocols, we employ a bidirectional announcement-bidding mechanism and the collaborative process consists of three phases, i.e., basic matching phase, forward announcement-bidding phase and backward announcement-bidding phase. Moreover, the elasticity is sufficiently considered while scheduling by dynamically adding virtual machines to improve schedulability. Furthermore, we design calculation rules of the bidding values in both forward and backward announcement-bidding phases and two heuristics for selecting contractors. On the basis of the bidirectional announcement-bidding mechanism, we propose an agent-based dynamic scheduling algorithm named ANGEL for real-time, independent and aperiodic tasks in clouds. Extensive experiments are conducted on CloudSim platform by injecting random synthetic workloads and the workloads from the last version of the Google cloud tracelogs to evaluate the performance of our ANGEL. The experimental results indicate that ANGEL can efficiently solve the real-time task scheduling problem in virtualized clouds.
Xiaomin Zhu 0001, Chao Chen 0027, Laurence T. Yang, Yang Xiang 0001
IEEE Trans. Computers1
2015 Exploiting Efficient and Scalable Shuffle Transfers in Future Data Center Networks
abstract
Distributed computing systems like MapReduce in data centers transfer massive amount of data across successive processing stages. Such shuffle transfers contribute most of the network traffic and make the network bandwidth become a bottleneck. In many commonly used workloads, data flows in such a transfer are highly correlated and aggregated at the receiver side. To lower down the network traffic and efficiently use the available network bandwidth, we propose to push the aggregation computation into the network and parallelize the shuffle and reduce phases. In this paper, we first examine the gain and feasibility of the in-network aggregation with BCube, a novel server-centric networking structure for future data centers. To exploit such a gain, we model the in-network aggregation problem that is NP-hard in BCube. We propose two approximate methods for building the efficient IRS-based incast aggregation tree and SRS-based shuffle aggregation subgraph, solely based on the labels of their members and the data center topology. We further design scalable forwarding schemes based on Bloom filters to implement in-network aggregation over massive concurrent shuffle transfers. Based on a prototype and large-scale simulations, we demonstrate that our approaches can significantly decrease the amount of network traffic and save the data center resources. Our approaches for BCube can be adapted to other servercentric network structures for future data centers after minimal modifications.
Deke Guo, Xiaolei Zhou 0001, Xiaomin Zhu 0001, Wei Wei 0006, Xueshan Luo
IEEE Trans. Parallel Distributed Syst.4
2015 Fault-Tolerant Scheduling for Real-Time Tasks on Multiple Earth-Observation Satellites
abstract
Fault-tolerance plays an important role in improving the reliability of multiple earth-observing satellites, especially in emergent scenarios such as obtaining photographs on battlefields or earthquake areas. Fault tolerance can be implemented through scheduling approaches. Unfortunately, little attention has been paid to fault-tolerant scheduling on satellites. To address this issue, we propose a novel dynamic fault-tolerant scheduling model for real-time tasks running on multiple observation satellites. In this model, the primary-backup policy is employed to tolerate one satellite's permanent failure at one time instant. In the light of the fault-tolerant model, we develop a novel fault-tolerant satellite scheduling algorithm named FTSS. To improve the resource utilization, we apply the overlapping technology that includes primary-backup copy overlapping (i.e., PB overlapping) and backup-backup copy overlapping (i.e., BB overlapping). According to the satellites characterized with time windows for observations, we extensively analyze the overlapping mechanism on satellites. We integrate the overlapping mechanism with FTSS, which employs the task merging strategies including primary-backup copy merging (i.e., PB merging), backup-backup copy merging (i.e., BB merging) and primary-primary copy merging (i.e., PP merging). These merging strategies are used to decrease the number of tasks required to be executed, thereby enhancing system schedulability. To demonstrate the superiority of our FTSS, we conduct extensive experiments using the real-world satellite parameters supplied from the satellite tool kit or STK; we compare FTSS with the three baseline algorithms, namely, NMFTSS, NOFTSS, and NMNOFTSS. The experimental results indicate that FTSS efficiently improves the scheduling quality of others and is suitable for fault-tolerant satellite scheduling.
Xiaomin Zhu 0001, Jianjiang Wang, Xiao Qin 0001, Ji Wang 0002, Zhong Liu 0002, Erik Demeulemeester
IEEE Trans. Parallel Distributed Syst.1
2014 Analysis and Design of Fault-Tolerant Scheduling for Real-Time Tasks on Earth-Observation Satellites
abstract
Fault-tolerant scheduling is an efficient approach to improving the reliability of multiple earth-observing satellites especially in some emergent scenarios such as obtaining photographs on battlefields or earthquake areas. Unfortunately, little work has been done to deal with the fault-tolerant scheduling on satellites. To address this issue, this paper presents a novel dynamic fault-tolerant scheduling model using primary-backup policy to tolerate one satellite's permanent failure at one time instant. On this basis, we propose a novel fault-tolerant satellite scheduling algorithm named FTSS, in which an overlapping technology is adopted to improve the resource utilization. Besides, the FTSS employs the task merging strategies to further enhance the schedulability. To demonstrate the superiority of our FTSS, we conduct extensive experiments by simulations using real-world satellite parameters from STK to compare FTSS with other baseline algorithms. The experimental results indicate that FTSS efficiently improves the scheduling quality of others and is suitable for fault-tolerant satellite scheduling.
Xiaomin Zhu 0001, Jianjiang Wang, Ji Wang 0002, Xiao Qin 0001
ICPP1
2014 Real-Time Tasks Oriented Energy-Aware Scheduling in Virtualized Clouds
abstract
Energy conservation is a major concern in cloud computing systems because it can bring several important benefits such as reducing operating costs, increasing system reliability, and prompting environmental protection. Meanwhile, power-aware scheduling approach is a promising way to achieve that goal. At the same time, many real-time applications, e.g., signal processing, scientific computing have been deployed in clouds. Unfortunately, existing energy-aware scheduling algorithms developed for clouds are not real-time task oriented, thus lacking the ability of guaranteeing system schedulability. To address this issue, we first propose in this paper a novel rolling-horizon scheduling architecture for real-time task scheduling in virtualized clouds. Then a task-oriented energy consumption model is given and analyzed. Based on our scheduling architecture, we develop a novel energy-aware scheduling algorithm named EARH for real-time, aperiodic, independent tasks. The EARH employs a rolling-horizon optimization policy and can also be extended to integrate other energy-aware scheduling algorithms. Furthermore, we propose two strategies in terms of resource scaling up and scaling down to make a good trade-off between task’s schedulability and energy conservation. Extensive simulation experiments injecting random synthetic tasks as well as tasks following the last version of the Google cloud tracelogs are conducted to validate the superiority of our EARH by comparing it with some baselines. The experimental results show that EARH significantly improves the scheduling quality of others and it is suitable for real-time task scheduling in virtualized clouds.
Xiaomin Zhu 0001, Laurence T. Yang, Huangke Chen, Ji Wang 0002, Shu Yin 0001, Xiaocheng Liu
IEEE Trans. Cloud Comput.1
2014 Dynamic Scheduling for Emergency Tasks on Distributed Imaging Satellites with Task Merging
abstract
Scheduling plays a significant role in improving observation effectiveness of distributed imaging satellites. Although extensive satellite scheduling algorithms have been proposed, none of them focuses on dynamic scheduling for emergency tasks. In this paper, a novel multi-objective dynamic scheduling model for emergency tasks on distributed imaging satellites is established for the first time. To improve user's satisfaction ratio and resource utilization, we propose the task merging strategy: establishing a task merging graph (TMG) model and proposing a task merging algorithm-CP-TM based on clique partition. In addition, a rehabilitation technique is suggested to overcome the disadvantage that task merging will make tasks have less imaging opportunities. To further enhance the schedulability, the task backward shift in the waiting sequence is considered in our study. Furthermore, a novel dynamic scheduling algorithm called TMBSR-DES is presented, which comprehensively considers task merging, backward shift, and rehabilitation. To demonstrate the superiority of our TMBSR-DES, we conduct extensive experiments by simulations to compare TMBSR-DES with three existing algorithm-RBHA, RTSSA, and LSA, as well as three baseline algorithms-BS-DES, TMR-DES, and TMBS-DES. The experimental results indicate that TMBSR-DES outperforms the others and is suitable for emergency task scheduling.
Jianjiang Wang, Xiaomin Zhu 0001, Dishan Qiu, Laurence T. Yang
IEEE Trans. Parallel Distributed Syst.2
2013 Novel clock synchronization algorithm of parametric difference for parallel and distributed simulations
Linjun Fan, Yunxiang Ling, Xiaomin Zhu 0001, Xiaoyong Tang
Comput. Networks4
2013 3E: Energy-efficient elastic scheduling for independent tasks in heterogeneous computing systems
Xiaomin Zhu 0001, Rong Ge 0002, Jinguang Sun
J. Syst. Softw.1
2012 An improved security-aware packet scheduling algorithm in real-time wireless networks
Xiaomin Zhu 0001, Shaoshuai Liang
Inf. Process. Lett.1
2012 Adaptive energy-efficient scheduling for real-time tasks on DVS-enabled heterogeneous clusters
Xiaomin Zhu 0001, Kenli Li 0001, Xiao Qin 0001
J. Parallel Distributed Comput.1
2012 Rolling-horizon scheduling for energy constrained distributed real-time embedded systems
Xiaomin Zhu 0001, Hui Guo 0001, Dishan Qiu, Jianqing Jiang
J. Syst. Softw.2
2011 Energy-efficient elastic scheduling in heterogeneous computing systems
abstract
Reducing energy consumption has become a major goal in designing modern heterogeneous computing systems. Conventional energy-efficient scheduling strategies developed on this kind of systems mainly focus on energy saving, lacking of the consideration of users' expectations such as expected finish time. Again, the system elasticity is not sufficiently taken into account while making scheduling decisions. In this paper, we develop a novel scheduling strategy named energy-efficient elastic (3E) scheduling for aperiodic, and independent tasks on DVS-enabled heterogeneous computing systems. The 3E strategy concentrates on adjusting tasks' supply voltages according to the system workload, thereby making the best trade-offs between energy conservation and users' expected finish times. Experimental results demonstrate that our 3E significantly improves the scheduling quality of the others, and is able to effectively enhance the system elasticity.
Xiaomin Zhu 0001, Jianjiang Wang
IPCCC1
2011 Improving adaptivity and fairness of processing real-time tasks with QoS requirements on clusters through dynamic scheduling
Jianghan Zhu, Xiaomin Zhu 0001, Jianqing Jiang
Inf. Process. Lett.2
2011 Boosting adaptivity of fault-tolerant scheduling for real-time tasks with service requirements on clusters
Xiaomin Zhu 0001, Rong Ge 0002, Peizhong Lu
J. Syst. Softw.1
2011 QoS-Aware Fault-Tolerant Scheduling for Real-Time Tasks on Heterogeneous Clusters
abstract
Fault-tolerant scheduling plays a significant role in improving system reliability of clusters. Although extensive fault-tolerant scheduling algorithms have been proposed for real-time tasks in parallel and distributed systems, quality of service (QoS) requirements of tasks have not been taken into account. This paper presents a fault-tolerant scheduling algorithm called QAFT that can tolerate one node's permanent failures at one time instant for real-time tasks with QoS needs on heterogeneous clusters. In order to improve system flexibility, reliability, schedulability, and resource utilization, QAFT strives to either advance the start time of primary copies and delay the start time of backup copies in order to help backup copies adopt the passive execution scheme, or to decrease the simultaneous execution time of the primary and backup copies of a task as much as possible to improve resource utilization. QAFT is capable of adaptively adjusting the QoS levels of tasks and the execution schemes of backup copies to attain high system flexibility. Furthermore, we employ the overlapping technology of backup copies. The latest start time of backup copies and their constraints are analyzed and discussed. We conduct extensive experiments to compare our QAFT with two existing schemes-NOQAFT and DYFARS. Experimental results show that QAFT significantly improves the scheduling quality of NOQAFT and DYFARS.
Xiaomin Zhu 0001, Xiao Qin 0001, Meikang Qiu
IEEE Trans. Computers1