Hailiang Zhao

dblp:19/538 · DBLP profile ↗
← Back
31ranked-venue papers
13as first author
21since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 10 · 5 first-author · 6 since 2021Artificial intelligence and machine learning · 6 · 3 first-author · 3 since 2021Computer networks · 5 · 1 first-author · 4 since 2021Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
YearPublicationVenuePosition
2026 PeerSync: Accelerating Containerized Model Inference at the Network Edge
abstract
Efficient container image distribution is crucial for enabling machine learning inference at the network edge, where resource limitations and dynamic network conditions create significant challenges. In this paper, we presentPeerSync, a decentralized P2P-based system designed to optimize image distribution in edge environments.PeerSyncemploys a popularity- and network-aware download engine that dynamically adapts to content popularity and real-time network conditions.PeerSyncfurther integrates automated tracker election for rapid peer discovery and dynamic cache management for efficient storage utilization. We implementPeerSyncwith 8000+ lines of Rust code and test its performance extensively on both large-scale Docker-based emulations and physical edge devices. Experimental results show thatPeerSyncdelivers a remarkable speed increase of 2.72×, 1.79×, and 1.28× compared to the Baseline solution, Dragonfly, and Kraken, respectively, while significantly reducing cross-network traffic by 90.72% under congested and varying network conditions.
Yinuo Deng, Hailiang Zhao, Dongjing Wang, Peng Chen 0051, Wenzhuo Qian, Jianwei Yin, Schahram Dustdar, Shuiguang Deng
IEEE Trans. Serv. Comput.2
2025 CADRef: Robust Out-of-Distribution Detection via Class-Aware Decoupled Relative Feature Leveraging
abstract
Deep neural networks (DNNs) have been widely criticized for their overconfidence when dealing with out-of-distribution (OOD) samples, highlighting the critical need for effective OOD detection to ensure the safe deployment of DNNs in real-world settings. Existing post-hoc OOD detection methods primarily enhance the discriminative power of logit-based approaches by reshaping sample features, yet they often neglect critical information inherent in the features themselves. In this paper, we propose the Class-Aware Relative Feature-based method (CARef), which utilizes the error between a sample’s feature and its class-aware average feature as a discriminative criterion. To further refine this approach, we introduce the Class-Aware Decoupled Relative Feature-based method (CADRef), which decouples sample features based on the alignment of signs between the relative feature and corresponding model weights, enhancing the discriminative capabilities of CARef. Extensive experimental results across multiple datasets and models demonstrate that both proposed methods exhibit effectiveness and robustness in OOD detection compared to state-of-the-art methods. Specifically, our two methods outperform the best baseline by 2.82% and 3.27% in AUROC, with improvements of 4.03% and 6.32% in FPR95, respectively.
Zhiwei Ling, Yachen Chang, Hailiang Zhao, Xinkui Zhao, Kingsum Chow, Shuiguang Deng
CVPR3
2025 SeMi: When Imbalanced Semi-Supervised Learning Meets Mining Hard Examples
abstract
Semi-Supervised Learning (SSL) can leverage abundant unlabeled data to boost model performance. However, the class-imbalanced data distribution in real-world scenarios poses great challenges to SSL, resulting in performance degradation. Existing class-imbalanced semi-supervised learning (CISSL) methods mainly focus on rebalancing datasets but ignore the potential of using hard examples to enhance performance, making it difficult to fully harness the power of unlabeled data even with sophisticated algorithms. To address this issue, we propose a method that enhances the performance of Imbalanced Semi-Supervised Learning by Mining Hard Examples (SeMi). This method distinguishes the entropy differences among logits of hard and easy examples, thereby identifying hard examples and increasing the utility of unlabeled data, better addressing the imbalance problem in CISSL. In addition, we maintain a class-balanced memory bank with confidence decay for storing high-confidence embeddings to enhance the pseudo-labels' reliability. Although our method is simple, it is effective and seamlessly integrates with existing approaches. We perform comprehensive experiments on standard CISSL benchmarks and experimentally demonstrate that our proposed SeMi outperforms existing state-of-the-art methods on multiple benchmarks, especially in reversed scenarios, where our best result shows approximately a 54.8% improvement over the baseline methods. Our code is available at https://github.com/pywin/SeMi.
Yin Wang 0004, Hao Lu 0009, Zhen Qin 0004, Hailiang Zhao, Guanjie Cheng, Xin Du 0002, Ge Su, Li Kuang, MengChu Zhou, Shuiguang Deng
ACM Multimedia5
2025 Robustifying Learning-Augmented Caching Efficiently without Compromising 1-Consistency
abstract
The online caching problem aims to minimize cache misses when serving a sequence of requests under a limited cache size. While naive learning-augmented caching algorithms achieve ideal $1$-consistency, they lack robustness guarantees. Existing robustification methods either sacrifice $1$-consistency or introduce excessive computational overhead. In this paper, we introduce Guard, a lightweight robustification framework that enhances the robustness of a broad class of learning-augmented caching algorithms to $2H_{k-1} + 2$, while preserving their $1$-consistency. Guard achieves the current best-known trade-off between consistency and robustness, with only $\mathcal{O}(1)$ additional per-request overhead, thereby maintaining the original time complexity of the base algorithm. Extensive experiments across multiple real-world datasets and prediction models validate the effectiveness of Guard in practice.
Peng Chen 0051, Hailiang Zhao, Xueyan Tang, Shuiguang Deng
NeurIPS2
2025 Ontology-Based Semantic Integration of Multi-Source Heterogeneous Industrial Devices
abstract
The 3C industry uses industrial internet platforms to tackle data fragmentation, poor collaboration, and semantic mismatches, driving digital transformation. Although the Object Linking and Embedding for Process Control Unified Architecture (OPC UA) standard supports industrial automation, it faces several limitations, including poor compatibility, ambiguous semantics, and incomplete models, which hinder automated reasoning and system interoperability. To address these issues, this work proposes an integrated framework that combines offline classification of heterogeneous devices using a Character-level Text Convolutional Neural Network (CTCNN) with online automated reasoning over OPC UA information models. CTCNN leverages character-level embeddings and convolutional layers to achieve fine-grained and accurate device type recognition. Furthermore, a deterministic Markov decision process is formulated, and an intelligent agent is trained to perform automatic reasoning based on the OPC UA model structure. This hybrid framework enables effective OPC UA device integration and interconnection. Experimental results demonstrate that CTCNN achieves up to an 18% improvement in device type identification precision, while the combined offline recognition and online reasoning approach enhances the accuracy of automatic OPC UA information model inference by approximately 12% compared to state-of-the-art methods.
Rina Wu, Hailiang Zhao
SMC5
2025 Decentralized Proactive Model Offloading and Resource Allocation for Split and Federated Learning
abstract
In the resource-constrained Internet of Things (IoT)-edge computing environment, split federated (SplitFed) learning is implemented to enhance training efficiency. This method involves each terminal device dividing its full deep neural network (DNN) model at a designated layer into a device-side model and a server-side model, then offloading the latter to the edge server. However, existing research overlooks four critical issues as follows: 1) the heterogeneity of end devices’ resource capacities and the sizes of their local data samples impact training efficiency; 2) the influence of the edge server’s computation and network resource allocation on training efficiency; 3) the data leakage risk associated with the offloaded server-side submodel; and 4) the privacy drawbacks of current centralized algorithms. Consequently, proactively identifying the optimal cut layer and server resource requirements for each end device to minimize training latency while adhering to data leakage risk rate constraint remains a challenging issue. To address these problems, this article first formulates the latency and data leakage risk of training DNN models using SplitFed learning. Next, we frame the SplitFed learning problem as a mixed-integer nonlinear programming challenge. To tackle this, we propose a decentralized proactive model offloading and resource allocation (DP-MORA) scheme, empowering each end device to determine its cut layer and resource requirements based on its local multidimensional training configuration, without knowledge of other devices’ configurations. Extensive experiments on two real-world datasets demonstrate that the DP-MORA scheme effectively reduces DNN model training latency, enhances training efficiency, and complies with data leakage risk constraints compared to several baseline algorithms across various experimental settings.
Binbin Huang 0006, Hailiang Zhao, Lingbin Wang, Wenzhuo Qian, Yuyu Yin, Shuiguang Deng
IEEE Internet Things J.2
2025 Tail-Learning: Adaptive Learning Method for Mitigating Tail Latency in Autonomous Edge Systems
abstract
In the field of edge computing, the increasing demand for high Quality of Service (QoS), particularly in dynamic multimedia streaming applications (e.g., Augmented Reality/Virtual Reality and online gaming), has prompted the need for effective solutions. Nevertheless, adopting an edge paradigm grounded in distributed computing has exacerbated the issue of tail latency. Given a limited variety of multimedia services supported by edge servers and the dynamic nature of user requests, employing traditional queuing methods to model tail latency in distributed edge computing is challenging, substantially exacerbating head-of-line (HoL) blocking. In response to this challenge, we have developed a learning-based scheduling method to mitigate the overall tail latency, which adaptively selects appropriate edge servers for execution as incoming distributed tasks vary with unknown size. To optimize the utilization of the edge computing paradigm, we leverage the Laplace transform to theoretically derive an upper bound for the response time of edge servers. Subsequently, we integrate this upper bound into reinforcement learning to facilitate tail-learning and enable informed decisions for autonomous distributed scheduling. The experiment results demonstrate the efficiency in reducing tail latency compared to existing methods.
Cheng Zhang 0010, Yinuo Deng, Hailiang Zhao, Tianlv Chen, Shuiguang Deng
ACM Trans. Auton. Adapt. Syst.3
2025 CATScaler: A Convolution-Augmented Transformer Scaling Framework for Cloud-Native Applications
abstract
Efficient container scaling is crucial for enhancing the availability and scalability of cloud-native applications through adaptive resource management. In cloud computing, the default autoscaling feature of Kubernetes scales pods only when the cluster or application exceeds a predefined threshold. However, this reactive approach often leads to significant resource waste during demand fluctuations because it cannot predict future workload changes and adjust resources in advance. This paper presents CATScaler, a novel Convolution-Augmented Transformer Scaler designed to proactively optimize resource allocation in serverless environments. CATScaler is a proactive approach composed of two modules: workload prediction and elastic auto-scaling. In the prediction module, we develop a convolution-augmented transformer to accurately predict workload changes at both local and global levels. Additionally, we incorporate reversible instance normalization to mitigate the shift caused by the difference between workload data and training data. In the auto-scaling module, we implement an instance-counting method to handle the nonlinear relationships between variables. Experiments using two real datasets from Alibaba Cloud and Huawei Cloud demonstrate the effectiveness of CATScaler. The tests conducted on a cluster of 4 servers demonstrated that CATScaler reduced response time latency by 1.1× compared to Kubernetes' default scaler and decreased service violation rates by 3.2×.
Fan'an Meng, Hongjun Dai, Guoqing Cong, Hailiang Zhao
IEEE Trans. Serv. Comput.5
2025 Data-Locality-Aware Task Assignment and Scheduling for Distributed Job Executions
abstract
This paper addresses the data-locality-aware task assignment and scheduling problem for distributed job executions. Our goal is to minimize job completion times without prior knowledge of future job arrivals. We propose an Optimal Balanced Task Assignment algorithm (OBTA), which achieves minimal job completion times while significantly reducing computational overhead through efficient narrowing of the solution search space. To balance performance and efficiency, we extend the approximate Water-Filling (WF) algorithm, providing a rigorous proof that its approximation factor equals the number of task groups in a job. We also introduce a novel heuristic, Replica-Deletion (RD), which outperforms WF by leveraging global optimization techniques. To further enhance scheduling efficiency, we incorporate job ordering strategies based on a shortest-estimated-time-first policy, reducing average job completion times across workloads. Extensive trace-driven evaluations validate the effectiveness and scalability of the proposed algorithms.
Hailiang Zhao, Xueyan Tang, Peng Chen 0051, Jianwei Yin, Shuiguang Deng
IEEE Trans. Serv. Comput.1
2025 Online Workload Scheduling for Social Welfare Maximization in the Computing Continuum
abstract
Computing ecosystems are shifting toward a computing continuum paradigm designed to handle the diverse and dynamic nature of computing resources spread across various locations. It demonstrates significant potential in providing high-bandwidth and low-latency services for users. However, as a large number of users request services from distributed computing continuum systems, it is critical to schedule numerous delay-sensitive, fractional workloads and maximum parallelism-bound jobs to appropriate backend resources,e.g., cloud container instances. In addition, the scheduling strategy also needs to maximize the social welfare that incorporates the utilities of jobs and the revenue of service providers. However, current workload scheduling algorithms are based on simple heuristics and lack performance guarantees. Due to the unpredictability of online requests, the distribution of requests should not be assumed. Therefore, designing an online workload scheduling strategy without assumptions on request distributions is essential for balancing the online workload. This work first establishes a spatiotemporal integrated resource pool to reflect the computational resources provided by distributed computing continuum systems. Then, several pseudo-social welfare functions and marginal cost functions are constructed, where the latter is used to estimate the marginal cost of provisioning services to each newly arrived job based on the current resource surplus. We propose an online workload scheduling strategy namedOnSocMaxto solve the above problems. It operates by following the solutions to several convex pseudo-social welfare maximization problems and is proven to be$\alpha$-competitive for some$\alpha$with a value of at least 2. The evaluation results demonstrate thatOnSocMaxoutperforms several benchmark strategies in maximizing social welfare.
Hailiang Zhao, Ziqi Wang 0011, Guanjie Cheng, Wenzhuo Qian, Peng Chen 0051, Jianwei Yin, Schahram Dustdar, Shuiguang Deng
IEEE Trans. Serv. Comput.1
2024 Learning-Augmented Algorithms for the Bahncard Problem
abstract
In this paper, we study learning-augmented algorithms for the Bahncard problem. The Bahncard problem is a generalization of the ski-rental problem, where a traveler needs to irrevocably and repeatedly decide between a cheap short-term solution and an expensive long-term one with an unknown future. Even though the problem is canonical, only a primal-dual-based learning-augmented algorithm was explicitly designed for it. We develop a new learning-augmented algorithm, named PFSUM, that incorporates both history and short-term future to improve online decision making. We derive the competitive ratio of PFSUM as a function of the prediction error and conduct extensive experiments to show that PFSUM outperforms the primal-dual-based algorithm.
Hailiang Zhao, Xueyan Tang, Peng Chen 0051, Shuiguang Deng
NeurIPS1
2024 A Game-Theoretic Approach-Based Task Offloading and Resource Pricing Method for Idle Vehicle Devices Assisted VEC
abstract
Vehicle Edge Computing (VEC), as an emerging computing paradigm, aims to achieve the high efficiencies and quality of service by distributing computation tasks to vehicles and cloud-edge servers. The resource pricing problem focuses on how to reasonably price the resources of VEC to encourage their allocation and utilization. However, VEC server overloading may lead to performance degradation, especially in urban congested areas. Meanwhile, idle resources near VEC roads, such as parked vehicles and RSUs, are underutilized and can provide additional computation and communication resources to the system. Inspired by this, this paper introduces a model to assist vehicle edge computing by attracting Idle Vehicles (IVs) to share resources. We use a two-stage Stackelberg game model to address the resource pricing and task offloading problem, analyzing the interaction between requesting vehicles and cloud-edge servers. Through a backward induction method, we transform the problem into a convex optimization problem and theoretically prove the existence of a unique Nash equilibrium. In the first stage, optimal offloading ratio strategy is solved using convex optimization theory. In the second stage, the original problem is decomposed into 2N sub-problems and solved using the Lagrangian dual method and Karush-Kuhn-Tucker (KKT) conditions for optimal resource pricing. Additionally, a price incentive mechanism and a task-vehicle stable matching game model are employed to recruit idle vehicles around the roads to spontaneously participate in the task offloading process. Finally, simulation results reveal our solution effectively reduces offloading costs, latency, energy use, and enhances task completion compared to others.
Yishan Chen 0001, Junxiao Han, Hailiang Zhao, Shuiguang Deng
IEEE Internet Things J.4
2024 Cloud-Native Computing: A Survey From the Perspective of Services
abstract
The development of cloud computing delivery models inspires the emergence of cloud-native computing. Cloud-native computing, as the most influential development principle for web applications, has already attracted increasingly more attention in both industry and academia. Despite the momentum in the cloud-native industrial community, a clear research roadmap on this topic is still missing. As a contribution to this knowledge, this article surveys key issues during the life cycle of cloud-native applications, from the perspective of services. Specifically, we elaborate on the research domains by decoupling the life cycle of cloud-native applications into four states: building, orchestration, operation, and maintenance. We also discuss the fundamental necessities and summarize the key performance metrics that play critical roles during the development and management of cloud-native applications. We highlight the key implications and limitations of existing works in each state. The challenges, future directions, and research opportunities are also discussed.
Shuiguang Deng, Hailiang Zhao, Binbin Huang 0006, Cheng Zhang 0010, Feiyi Chen, Yinuo Deng, Jianwei Yin, Schahram Dustdar, Albert Y. Zomaya
Proc. IEEE2
2024 Scheduling Multi-Server Jobs With Sublinear Regrets via Online Learning
abstract
Multi-server jobs that request multiple computing resources and hold onto them during their execution dominate modern computing clusters. When allocating the multi-type resources to several co-located multi-server jobs simultaneously in online settings, it is difficult to make the tradeoff between the parallel computation gain and the internal communication overhead, apart from the resource contention between jobs. To study the computation-communication tradeoff, we model the computation gain as the speedup on the job completion time when it is executed in parallelism on multiple computing instances, and fit it with utilities of different concavities. Meanwhile, we take the dominant communication overhead as the penalty to be subtracted. To achieve a better gain-overhead tradeoff, we formulate an cumulative reward maximization program and design an online algorithm, namedOgaSched, to schedule multi-server jobs.OgaSchedallocates the multi-type resources to each arrived job in the ascending direction of the reward gradients. It has several parallel sub-procedures to accelerate its computation, which greatly reduces the complexity. We proved that it has a sublinear regret with general concave rewards. We also conduct extensive trace-driven simulations to validate the performance ofOgaSched. The results demonstrate thatOgaSchedoutperforms widely used heuristics by 11.33%, 7.75%, 13.89%, and 13.44%, respectively.
Hailiang Zhao, Shuiguang Deng, Zhengzhe Xiang, Xueqiang Yan, Jianwei Yin, Schahram Dustdar, Albert Y. Zomaya
IEEE Trans. Serv. Comput.1
2023 Learning to Schedule Multi-Server Jobs With Fluctuated Processing Speeds
abstract
Multi-server jobs are imperative in modern cloud computing systems. A noteworthy feature of multi-server jobs is that, they usually request multiple computing devices simultaneously for their execution. How to schedule multi-server jobs online with a high system efficiency is a topic of great concern. First, the scheduling decisions have to satisfy the service locality constraints. Second, the scheduling decisions needs to be made online without the knowledge of future job arrivals. Third, and most importantly, the actual service rate experienced by a job is usually in fluctuation because of the dynamic voltage and frequency scaling (DVFS) and power oversubscription techniques when multiple types of jobs co-locate. A majority of online algorithms with theoretical performance guarantees are proposed. However, most of them require the processing speeds to be knowable, thereby the job completion times can be exactly calculated. To present a theoretically guaranteed online scheduling algorithm for multi-server jobs without knowing actual processing speeds apriori, in this article, we proposeEsdp(Efficient Sampling-based Dynamic Programming), which learns the distribution of the fluctuated processing speeds over time and simultaneously seeks to maximize the cumulative overall utility. The cumulative overall utility is formulated as the sum of the utilities of successfully serving each multi-server job minus the penalty on the operating, maintaining, and energy cost.Esdpis proved to have a polynomial complexity and a logarithmic regret, which is a State-of-the-Art result. We also validate it with extensive simulations and the results show that the proposed algorithm outperforms several benchmark policies with improvements by up to 73%, 36%, and 28%, respectively.
Hailiang Zhao, Shuiguang Deng, Feiyi Chen, Jianwei Yin, Schahram Dustdar, Albert Y. Zomaya
IEEE Trans. Parallel Distributed Syst.1
2022 DPoS: Decentralized, Privacy-Preserving, and Low-Complexity Online Slicing for Multi-Tenant Networks
abstract
Network slicing is the key to enable virtualized resource sharing among vertical industries in the era of 5G communication. Efficient resource allocation is of vital importance to realize network slicing in real-world business scenarios. To deal with the high algorithm complexity, privacy leakage, and unrealistic offline setting of current network slicing algorithms, in this paper we propose a fully decentralized and low-complexity online algorithm, DPoS, for multi-resource slicing. We first formulate the problem as a global social welfare maximization problem. Next, we design the online algorithm DPoS based on the primal-dual approach and posted price mechanism. In DPoS, each tenant is incentivized to make its own decision based on its true preferences without disclosing any private information to the mobile virtual network operator and other tenants. We provide a rigorous theoretical analysis to show that DPoS has the optimal competitive ratio when the cost function of each resource is linear. Extensive simulation experiments are conducted to evaluate the performance of DPoS. The results show that DPoS can not only achieve close-to-offline-optimal performance, but also have low algorithmic overheads.
Hailiang Zhao, Shuiguang Deng, Zhengzhe Xiang, Jianwei Yin, Schahram Dustdar, Albert Y. Zomaya
IEEE Trans. Mob. Comput.1
2022 Mobility-Aware Offloading and Resource Allocation for Distributed Services Collaboration
abstract
In mobile edge computing (MEC) systems, mobile users (MUs) are capable of allocating local resources (CPU frequency and transmission power) and offloading tasks to edge servers in the vicinity in order to enhance their computation capabilities and reduce back-and-forth transmission over backhaul link. Nevertheless, mobile environment makes it hard to draw offloading and resource allocation decisions under dynamical wireless channel state and users’ locations. In real life, social relationship is also provably a significant factor affecting integral performance in collaborative work, which results in MUs decisions strongly coupled and renders this problem further intractable. Most of previous works ignore the impact of inter-user dependency (or data dependency among IoT devices). To bridge this gap, we study the service collaboration with master-slave dependency among service chains of MUs and formulate this combinational optimization problem as a mixed integer non-linear programming (MINLP) problem. To this end, we derive the closed-form expression of resource allocation solution by convex optimization and transform it to integer linear programming (ILP) problem. Subsequently, we propose a distributed algorithm based on Markov approximation which has polynomial computation complexity. Experimental result on real-world dataset substantiates the usefulness and superiority of our scheme, in terms of reducing latency and energy consumption.
Shuiguang Deng, Hongze Zhu, Hailiang Zhao, Schahram Dustdar, Albert Y. Zomaya
IEEE Trans. Parallel Distributed Syst.4
2022 Dependent Function Embedding for Distributed Serverless Edge Computing
abstract
Edge computing is booming as a promising paradigm to extend service provisioning from the centralized cloud to the network edge. Benefit from the development of serverless computing, an edge server can be configured as a carrier of limited serverless functions, in the way of deploying Docker runtime and Kubernetes engine. Meanwhile, an application generally takes the form of directed acyclic graphs (DAGs), where vertices represent dependent functions and edges represent data traffic. The status quo of minimizing the completion time (a.k.a. makespan) of the application motivates the study on optimal function placement. However, current approaches lose sight of proactively splitting and mapping the traffic to the logical data paths between the heterogeneous edge servers, which could affect the makespan significantly. To remedy that, we propose an algorithm, termed as Dependent Function Embedding (DPE), to get the optimal edge server for each function to execute and the moment it starts executing. DPE finds the best segmentation of each data traffic by exquisitely solving several infinity norm minimization problems. DPE is theoretically verified to achieve the global optimality. Extensive experiments on Alibaba cluster trace show that DPE significantly outperforms two baseline algorithms in makespan by 43.19% and 40.71%, respectively.
Shuiguang Deng, Hailiang Zhao, Zhengzhe Xiang, Cheng Zhang 0010, Ying Li 0001, Jianwei Yin, Schahram Dustdar, Albert Y. Zomaya
IEEE Trans. Parallel Distributed Syst.2
2022 Distributed Redundant Placement for Microservice-based Applications at the Edge
abstract
Multi-access edge computing (MEC) is booming as a promising paradigm to push the computation and communication resources from cloud to the network edge to provide services and to perform computations. With container technologies, mobile devices with small memory footprint can run composite microservice-based applications without time-consuming backbone. Service placement at the edge is of importance to put MEC from theory into practice. However, current state-of-the-art research does not sufficiently take the composite property of services into consideration. Besides, although Kubernetes has certain abilities to heal container failures, high availability cannot be ensured due to heterogeneity and variability of edge sites. To deal with these problems, we propose a distributed redundant placement framework SAA-RP and a GA-based Server Selection (GASS) algorithm for microservice-based applications with sequential combinatorial structure. We formulate a stochastic optimization problem with the uncertainty of microservice request considered, and then decide for each microservice, how it should be deployed and with how many instances as well as on which edge sites to place them. Benchmark policies are implemented in two scenarios, where redundancy is allowed and not, respectively. Numerical results based on a real-world dataset verify that GASS significantly outperforms all the benchmark policies.
Hailiang Zhao, Shuiguang Deng, Jianwei Yin, Schahram Dustdar
IEEE Trans. Serv. Comput.1
2021 Distributed Redundancy Scheduling for Microservice-based Applications at the Edge
abstract
Multi-access Edge Computing is booming as a promising paradigm to push the computation and communication resources from cloud to the edge to provision services and to perform computations. With container technologies, mobile devices with small memory footprint can run composite microservice-based applications without the time-consuming backbone transmission. Service placement at the edge is of importance to put MEC from theory into practice. However, current state-of-the-art research does not sufficiently take the composite property of services into consideration but study the to-be-placed services in an atomic way. Besides, although Kubernetes has certain abilities to heal container failures, high availability cannot be ensured due to heterogeneity and variability of edge sites.
Hailiang Zhao, Shuiguang Deng, Jianwei Yin, Schahram Dustdar
SERVICES1
2021 Incentive-Driven Computation Offloading in Blockchain-Enabled E-Commerce
abstract
Blockchain is regarded as one of the most promising technologies to upgrade e-commerce. This article analyzes the challenges that current e-commerce is facing and introduces a new scenario of e-commerce enabled by blockchain. A framework is proposed for mining tasks in this scenario offloaded onto edge servers based on mobile edge computing. Then, the offloading issue is modeled as a multi-constrained optimization problem, and evolutionary algorithms are utilized and re-designed as solvers. The experimental results validate the efficiency of the framework and algorithms and also show that the lower bound of computation resources exists to obtain the maximum overall revenue.
Shuiguang Deng, Guanjie Cheng, Hailiang Zhao, Honghao Gao, Jianwei Yin
ACM Trans. Internet Techn.3
2020 Edge Intelligence: The Confluence of Edge Computing and Artificial Intelligence
abstract
Along with the rapid developments in communication technologies and the surge in the use of mobile devices, a brand-new computation paradigm, edge computing, is surging in popularity. Meanwhile, the artificial intelligence (AI) applications are thriving with the breakthroughs in deep learning and the many improvements in hardware architectures. Billions of data bytes, generated at the network edge, put massive demands on data processing and structural optimization. Thus, there exists a strong demand to integrate edge computing and AI, which gives birth to edge intelligence. In this article, we divide edge intelligence into AI for edge (intelligence-enabled edge computing) and AI on edge (artificial intelligence on edge). The former focuses on providing more optimal solutions to key problems in edge computing with the help of popular and effective AI technologies while the latter studies how to carry out the entire process of building AI models, i.e., model training and inference, on the edge. This article provides insights into this new interdisciplinary field from a broader perspective. It discusses the core concepts and the research roadmap, which should provide the necessary background for potential future research initiatives in edge intelligence.
Shuiguang Deng, Hailiang Zhao, Weijia Fang, Jianwei Yin, Schahram Dustdar, Albert Y. Zomaya
IEEE Internet Things J.2
2019 Data-Intensive Application Deployment at Edge: A Deep Reinforcement Learning Approach
abstract
Mobile Edge Computing (MEC) has already developed into a key component of the future mobile broadband network due to its low latency. In MEC, mobile devices can access data-intensive applications deployed at edge, which are facilitated by service and computing resources available on edge servers. However, it is difficult to handle such issues while data transmission, user mobility and load balancing conditions change constantly among mobile devices, edge servers and the cloud. In this paper, we propose an approach for formulating Data-intensive Application Edge Deployment Policy (DAEDP) that maximizes the latency reduction for mobile devices while minimizing the monetary cost for Application Service Providers (ASPs). The deployment problem is modelled as a Markov decision process, and a deep reinforcement learning strategy is proposed to formulate the optimal policy with maximization of the long-term discount reward. Extensive experiments are conducted to evaluate DAEDP. The results show that DAEDP outperforms four baseline approaches.
Yishan Chen 0001, Shuiguang Deng, Hailiang Zhao, Qiang He 0001, Ying Li 0001, Honghao Gao
ICWS3
2019 Multiple Energy Harvesting Devices Enabled Joint Computation Offloading and Dynamic Resource Allocation for Mobile-Edge Computing Systems
abstract
A mobile-edge computing (MEC) system integrating energy harvesting (EH) techniques is a promising paradigm for supporting computation-intensive and delay-sensitive mobile applications. While computation offloading reduces users' perceived latency, EH techniques mitigate the limitation of mobile devices' battery capacity. However, when considering a scenario with multiple EH devices, the advantages of MEC systems with EH devices may be compromised due to the competition among multiple devices for available computational resources and wireless bandwidth. In this paper, the joint computation offloading and dynamic resource allocation (JCODRA) that minimizes the long-term average execution cost is formulated as a stochastic optimization problem. In particular, both the long-term average execution delay and the penalty delay are included in the optimization objective. The former intends to handle the competition and random and uncontrollable EH processes properly and the latter aims to reduce the ratio of dropped tasks. An online algorithm based on Lyapunov optimization is proposed to transform the original problem into a per-time slot deterministic problem. The results of experiments demonstrate that our algorithm significantly outperforms three representative baseline approaches.
Wei Du 0001, Qiwang Lei, Qiang He 0001, Wei Liu 0011, Feifei Chen 0001, Lei Pan 0002, Hailiang Zhao
ICWS8
2019 Service Capacity Enhanced Task Offloading and Resource Allocation in Multi-Server Edge Computing Environment
abstract
An edge computing environment features multiple edge servers and multiple service clients. In this environment, mobile service providers can offload client-side computation tasks from service clients' devices onto edge servers to reduce service latency and power consumption experienced by the clients. A critical issue that has yet to be properly addressed is how to allocate edge computing resources to achieve two optimization objectives: 1) minimize the service cost measured by the service latency and the power consumption experienced by service clients; and 2) maximize the service capacity measured by the number of service clients that can offload their computation tasks in the long term. This paper formulates this long-term problem as a stochastic optimization problem and solves it with an online algorithm based on Lyapunov optimization. This NPhard problem is decomposed into three sub-problems, which are then solved with a suite of techniques. The experimental results show that our approach significantly outperforms two baseline approaches.
Wei Du 0001, Qiang He 0001, Wei Liu 0011, Qiwang Lei, Hailiang Zhao, Wei Wang 0033
ICWS6
2019 A Mobility-Aware Cross-Edge Computation Offloading Framework for Partitionable Applications
abstract
Mobile Edge Computing has already become a new paradigm to reduce the latency in data transmission for resource-limited mobile devices by offloading computation tasks onto edge servers. However, for mobility-aware computation-intensive services, existing offloading strategies cannot handle the offloading procedure properly because of the lack of collaboration among edge servers. A data stream application is partitionable if it can be presented by a directed acyclic dataflow graph, which makes cross-edge collaboration possible. In this paper, we propose a cross-edge computation offloading (CCO) framework for partitionable applications. The transmission, execution and coordination cost, as well as the penalty for task failure, are considered. An online algorithm based on Lyapunov optimization is proposed to jointly determine edge site-selection and energy harvesting without priori knowledge. By stabilizing the battery energy level of each mobile device around a positive constant, the proposed algorithm can obtain asymptotic optimality. The-oretical analysis about the complexity and the effectiveness of the proposed framework is provided. Experimental results based on a real-life dataset corroborate that CCO can achieve superior performance compared with benchmarks where crossedge collaboration is not allowed.
Hailiang Zhao, Shuiguang Deng, Cheng Zhang 0010, Wei Du 0001, Qiang He 0001, Jianwei Yin
ICWS1
2018 A Multi-objective Optimization Model for Determining the Optimal Standard Feasible Neighborhood of Intelligent Vehicles
Hailiang Zhao
PRICAI (1)3
2004 Rule chain and dominant rule control algorithm for unknown nonlinear systems
abstract
This paper is focused on the control method based upon rule chains of a rule-base for unknown nonlinear systems. At first, the fuzzy point is employed to simplify the expression of multi-input multi-output fuzzy rules. Some relations between the fuzzy point and fuzzy number are obtained. Second, the input dominant region and output dominant region of a rule are proposed. Third, the dominant rule control algorithm is constructed by means of the extension rule, which contains not only the fuzzy couple of system state and system input, but also the information about its action time and the prospective system state after its action time. Following that, the rule chain and /spl epsi/-level stability are presented. Furthermore, a sufficient and necessary condition for the /spl epsi/-level stability of the rule-base control systems is obtained. A method to analyze the /spl epsi/-level stability of the rule-base control systems is presented. An upper bound about the settling time and the maximum steady-error for the system are estimated. Finally, an application result is employed to show the effectivity of the method.
Hailiang Zhao
FUZZ-IEEE1
2003 Research on multiobjective optimization control for nonlinear unknown systems
abstract
This paper is focused on the optimal control problem of the nonlinear unknown systems with multiobjectives. The concept of system output response function based upon system input sequence over time is proposed, and the problem is formulated accordingly. Considering the fact that a finite response curve set is easy to be obtained, we employ the Pareto rule-base and the approximate Pareto control algorithm for multiobjective control optimization based on the finite response curve set. In this way, an easy method to find a Pareto rule-base for the complicated multiobjective optimal control problem that converts the problem to the one that can be resolved only in a finite set consisting of input-output date and curves over time is presented. It can guarantee that every rule's input and output base point is optimally matched in Pareto sense within the known set of input and output of the system. Moreover, some sufficient conditions are obtained for the conventional fuzzy control algorithm to be a Pareto one. It is shown that if the rule-base is composed of Pareto rules, then for any inputs between two rule base-point, the corresponding output of the algorithm is also bounded by the two corresponding out base-points of the two rules. From the view of approximation, the Pareto algorithm can guarantee the system response is of Pareto performance relative to the objectives. As an illustration, the theory is applied to Monotone Inertial System. Simulation results agree with theories presented in this paper, and show that the fuzzy controller based on the Pareto rule-base presents very good behaviors in adaptivity, robustness and tracking with time-varying setpoint.
Hailiang Zhao, Tsu-Tian Lee
FUZZ-IEEE1
2003 Monotone inertial system model for unknown nonlinear systems
abstract
In this paper, a qualitative mode and a quantitative control model for unknown system are proposed. Based on the system output process trend according to control input sequence, a qualitative model- monotone inertial system (MIS, for short) is proposed, which can be used to describe most MIMO plants with no exactly mathematical model or unknown systems in practical engineering. In a MIS, the relation between input and output is monotonic and the output process is of time inertial. In a quantitative way, we suggest a new concept- the effective rule-base. After getting some qualitative properties of the MIS, a conclusion about the control stability is obtained by a way, which differs from Lyapunov method. Furthermore, a method to find an effective rule-base is presented for an unknown MIS based on finite knowledge. Simulation results on a challenging nonlinear magnetic bearing system support the conclusions of this paper.
Hailiang Zhao, Tsu-Tian Lee
SMC1
2000 Monotone fuzzy control method and its control performance
abstract
For most process controls the relation between the input and the output is of monotonicity. So a question immediately arises as follows: For two precise inputs, say x/spl les/y, can the fuzzy control method guarantee the corresponding crisp outputs preserving the same inequality? Or whether or not the fuzzy control method is of monotone? In this paper, we will focus on the problems mentioned above. Therefore, some new concepts about this topic are proposed, such as the monotone and the rough monotone control algorithm, the monotone rule base and the monotone smooth transition system model, etc. Analysis shows that as long as the rule base is monotone, the single-input single-output conventional Mamdani fuzzy control algorithm (CMFCA for short) is monotone and the two-input single-output CMFCA is rough monotone. Finally, a sufficient condition for monotonicity of the CMFCA is obtained, i.e. if the rule base is monotonic and the universes is partitioned by isosceles triangle fuzzy numbers with base point single resonating, then the two-input single-output conventional Mamdani algorithm is of monotone.
Hailiang Zhao, Changqian Zhu
SMC1