EDBT 2026 Demo / reviewers in the wild / expert
Ao Zhou 0001
dblp:123/6949-1
· DBLP profile ↗
96ranked-venue papers
13as first author
61since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 42 · 2 first-author · 32 since 2021Software engineering, systems software and programming languages · 21 · 4 first-author · 15 since 2021Systems, architecture and hardware · 13 · 4 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 5 · 5 since 2021Security and privacy · 3Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-authorArtificial intelligence and machine learning · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FedCurrMM: A Federated Map Matching Framework with Curriculum-Aware Client Selection
Minxiao Chen, Haitao Yuan 0002, Zhihan Zheng, Ao Zhou 0001, Shangguang Wang |
ICDE | 6 |
| 2026 | Temperature- and Energy-Aware Dynamic Task Scheduling and Computing Resource Allocation for Satellite ComputingabstractSatellite computing, as an emerging edge computing paradigm, extends computing and networking services into space. Due to the internal design constraints of low-Earth orbit (LEO) satellites and the challenges posed by the external environment, satellite computing faces inherent limitations, including severely constrained resources, non-rechargeable batteries, poor heat dissipation, and highly dynamic operating conditions, leading to unreliable and unsustainable quality of service. To address the above challenges and fully realize the potential of satellite computing, this paper investigates temperature- and energy-aware dynamic task scheduling and computing resource allocation, aiming to optimize service latency, reduce onboard energy consumption, and enhance operational profit. Solving this problem requires coordinating task scheduling and resource allocation, balancing communication and computation latency, and addressing the challenge of a vast search space. To solve the above challenges, we first formulate this problem as a repeated Stackelberg game by developing temperature and energy models. Through theoretical analysis, we show that this game leads to a convex optimization framework that exhibits exponential complexity. To accelerate the search for the Stackelberg equilibrium solution, we propose a dynamic task scheduling algorithm based on the interior point method, which reduces the computational complexity to polynomial order. Trace-driven simulations demonstrate that the proposed algorithm reduces task scheduling latency by 28.4% and improves utility by 13% on average. Chao Wang 0093, Xiao Ma 0009, Chuanxiu Chi, Ao Zhou 0001, Ruolin Xing, Shangguang Wang |
IEEE Trans. Mob. Comput. | 4 |
| 2026 | Prototyping and Analyzing Mobile SoC Clusters as Modern Edge ServersabstractThe rapidly growing edge computing platforms, coupled with the still imperfect edge infrastructure, present an excellent opportunity for the emergence of new edge hardware. However, it remains unclear whether alternative architectures built from energy-efficient mobile System-on-Chips (SoCs) can meet the stringent performance, cost, and energy demands of modern edge workloads. In this paper, we propose a new type of edge server composed of 60 Qualcomm Snapdragon 865 mobile SoCs in a 2U rack, referred to as SoC Cluster. We demonstrate its successful deployment on existing edge cloud platforms and its ability to natively serve mobile cloud gaming services. Despite the emergence of new hardware on edge platforms and its successful operation in serving mobile cloud gaming, our trace analysis revealed low hardware utilization and significant dynamic fluctuations in usage. To assess its broader applicability, we conducted the first measurement study of SoC Cluster to reveal its ability to run two popular and modern edge applications: deep learning inference and video transcoding. We developed a cross-platform benchmark suite to evaluate throughput, latency, power consumption, and application-specific metrics like video quality. We then directly compare SoC Cluster with a traditional edge server equipped with Intel CPUs and NVIDIA GPUs in terms of energy efficiency, space efficiency, and monetary cost. Results show that SoC Cluster exhibits up to 6.5? higher energy efficiency and 7.7? higher space efficiency. We also disclose its limitations in serving computation-intensive workloads such as large deep learning models. The outcomes provide insightful implications and offer practical direction for refining SoC Cluster toward broader deployment in edge scenarios. Li Zhang 0133, Boqing Shi, Xiang Li 0067, Ao Zhou 0001, Xiao Ma 0009, Shangguang Wang, Mengwei Xu 0001 |
IEEE Trans. Mob. Comput. | 4 |
| 2026 | Exploring Image Similarity to Optimize Resource Provisioning in Container-Enabled Edge ComputingabstractContainerization offers great flexibility and agility for resource provisioning in edge clouds. However, this benefit is not freely available, as substantial network traffic incurred by container image pulling heavily burdens the back-haul networks. Our in-depth measurements on 516.3 GB images from Docker Hub reveal that many similar images share identical layers, with only 11.21% having no shared layers. Building on this, we investigate the resource provisioning problem by leveraging image similarity to avoid repeated layer transmissions, aiming to reduce network traffic and improve overall performance. We formulate resource provisioning as a mixed-integer non-linear programming problem, which is challenging due to the coupling of four issues and their conflicting effects on overall performance, including offloading decisions, container instance deployment, image pulling, and resource allocation. To tackle these complexities, we propose a novel Similarity-Aware Resource Provisioning approach, which decomposes the problem into independent sub-problems using counterfactual multi-agent deep reinforcement learning and then solves sub-problems individually with convex optimization and fractional programming techniques. We conduct extensive evaluations with images from Docker Hub. The results show that our approach enables up to 32.6% traffic reduction and 19.9% utility improvement, outperforming the state-of-the-art solutions. Ao Zhou 0001, Xiao Ma 0009, Jinfeng Wen, Shangguang Wang |
IEEE Trans. Serv. Comput. | 2 |
| 2025 | Reliability-Aware Resource Allocation for Vehicular Services in Mobile Edge ComputingabstractMobile edge computing is envisioned as a promising paradigm to overcome the computational limitations of vehicles. However, current resource allocation optimizations for vehicular services in mobile edge computing are typically performed under quasi-static scenarios. Due to the time-varying wireless communication environments between vehicles and the base station, the communication uncertainty becomes a key factor to hinder the service reliability. To tackle this problem, the reliability-aware mobile edge resource allocation for vehicular services is explored in this paper. Firstly, the model of computation, communication, and reliability is reviewed integratively, and the problem is then formulated as a fractional programming model. Secondly, conditional value-at-risk (CVaR) is adopted to transform the original problem into a solvable semidefinite programming model. Finally, we propose an efficient two-layer reliability-aware resource allocation approach that decomposes the original problem into two sub-problems with lower complexity. Lagrange multiplier and Dinkelbach strategy are then applied for these sub-problems, which minimize average energy consumption while ensuring reliability for vehicles. Simulation results examine the performance of our proposed approach and demonstrate its superiority over other approaches. Lipei Yang, Ao Zhou 0001, Xiao Ma 0009, Shangguang Wang |
ICWS | 3 |
| 2025 | SateLight: A Satellite Application Update Framework for Satellite ComputingabstractSatellite computing is an emerging paradigm that empowers satellites to perform onboard processing tasks (i.e., satellite applications), thereby reducing reliance on ground-based systems and improving responsiveness. However, enabling application software updates in this context remains a fundamental challenge due to application heterogeneity, limited ground-to-satellite bandwidth, and harsh space conditions. Existing software update approaches, designed primarily for terrestrial systems, fail to address these constraints, as they assume abundant computational capacity and stable connectivity.To address this gap, we propose SateLight, a practical and effective satellite application update framework tailored for satellite computing. SateLight leverages containerization to encapsulate heterogeneous applications, enabling efficient deployment and maintenance. SateLight further integrates three capabilities: (1) a content-aware differential strategy that minimizes communication data volume, (2) a fine-grained onboard update design that reconstructs target applications, and (3) a layer-based fault-tolerant recovery mechanism to ensure reliability under failure-prone space conditions. Experimental results on a satellite simulation environment with 10 representative satellite applications demonstrate that SateLight reduces transmission latency by up to 91.18% (average 56.54%) compared to the best currently available baseline. It also consistently ensures 100% update correctness across all evaluated applications. Furthermore, a case study on a real-world in-orbit satellite demonstrates the practicality of our approach. Jinfeng Wen, Jianshu Zhao, Zixi Zhu, Ao Zhou 0001, Shangguang Wang |
ASE | 6 |
| 2025 | An RDMA Congestion Control Enhancement Framework for Edge DatacentersabstractWith the rapid expansion of Internet of Things (IoT) devices and the growing demand for real-time data processing, Remote Direct Memory Access (RDMA) has become increasingly vital in edge datacenters due to its high performance and low CPU utilization. To fully harness RDMA’s potential, a lossless underlying network is essential, typically achieved through hop-by-hop Priority Flow Control (PFC). However, existing RDMA congestion control mechanisms in edge datacenters struggle to achieve rapid rate convergence and may exacerbate PFC side effects such as head-of-line blocking, unfairness, and even deadlocks. In this paper, we propose a reinforcement learning-based RDMA congestion control enhancement framework called RDI. By leveraging receiver-side information and network congestion levels, RDI provides precise rate guidance for congested flows, alleviating congestion and mitigating PFC issues. Furthermore, RDI combines both online and offline learning to assist receiver-side information, achieving finer rate adjustments to dynamically adapt to changing network workloads and optimize congestion control performance. RDI is transparent and compatible with existing RDMA network architectures, requiring no modifications to network devices. Extensive simulations under realistic traffic patterns show that the congestion control scheme enhanced by RDI significantly outperforms the original mechanisms in terms of throughput and flow completion time (FCT), while also reducing PFC side effects. RDI-enhanced congestion control reduces the 99th-percentile tail FCT by up to 92% and the average FCT by up to 60%. Ao Zhou 0001, Shangguang Wang |
IEEE Internet Things J. | 3 |
| 2025 | RLOMM: An Efficient and Robust Online Map Matching Framework with Reinforcement LearningabstractOnline map matching is a fundamental problem in location-based services, aiming to incrementally match trajectory data step-by-step onto a road network. However, existing methods fail to meet the needs for efficiency, robustness, and accuracy required by large-scale online applications, making this task still challenging. This paper introduces a novel framework that achieves high accuracy and efficient matching while ensuring robustness in handling diverse scenarios. To improve efficiency, we begin by modeling the online map matching problem as an Online Markov Decision Process (OMDP) based on its inherent characteristics. This approach helps efficiently merge historical and real-time data, reducing unnecessary calculations. Next, to enhance robustness, we design a reinforcement learning method, enabling robust handling of real-time data from dynamically changing environments. In particular, we propose a novel model learning process and a comprehensive reward function, allowing the model to make reasonable current matches from a future-oriented perspective, and to continuously update and optimize during the decision-making process based on feedback. Lastly, to address the heterogeneity between trajectories and roads, we design distinct graph structures, facilitating efficient representation learning through graph and recurrent neural networks. To further align trajectory and road data, we introduce contrastive learning to decrease their distance in the latent space, thereby promoting effective integration of the two. Extensive evaluations on three real-world datasets confirm that our method significantly outperforms existing state-of-the-art solutions in terms of accuracy, efficiency and robustness. Minxiao Chen, Haitao Yuan 0002, Zhihan Zheng, Sai Wu, Ao Zhou 0001, Shangguang Wang |
Proc. ACM Manag. Data | 6 |
| 2025 | S-MGHSTN: Towards An Effective Streaming Traffic Accident Risk Prediction FrameworkabstractTraffic accidents pose a significant risk to human health and property safety. To address this issue, predicting their risks has garnered growing interest. We argue that a desired prediction solution should demonstrate resilience to the complexity of traffic accidents. In particular, it should adequately consider the streaming nature of data and key related aspects, such as regional background, accurately capture both proximity and similarity while bridging the disparities, and effectively address the sparsity. However, these factors are often overlooked or difficult to incorporate. In this paper, we propose a novel streaming multi-granularity hierarchical spatio-temporal network. Initially, we innovate by incorporating remote sensing data, facilitating the creation of hierarchical multi-granularity structure and the comprehension of regional background. We construct multiple high-level risk prediction tasks to enhance model's ability to cope with sparsity. Subsequently, to capture and bridge spatial proximity and semantic similarity, region features and multi-view graph undergo encoding processes to distill effective representations, followed by a graph-enhanced representation alignment module that reconciles their disparities. At last, an alternating experience replay with a dual-memory buffer is employed to accommodate streaming data scenarios. Extensive experiments on two real datasets verify the superiority of our model against the state-of-the-art methods. Minxiao Chen, Haitao Yuan 0002, Zhihan Zheng, Zhifeng Bao, Ao Zhou 0001, Shangguang Wang |
IEEE Trans. Knowl. Data Eng. | 6 |
| 2025 | Rethinking Cost-Efficient VM Scheduling on Public Edge Platforms: A Service Provider's PerspectiveabstractPrior studies on traditional centralized clouds independently optimizing static resource utilization and dynamic bandwidth cost are not applicable to edge scenarios, where edge sites are interconnected by wide area networks (WAN) rather than local area networks (LAN) as within clouds. Due to the lack of knowledge about the actual status of public edge platforms and real-world edge datasets, existing influential literature on edge scenarios demonstrates significant disparities in optimization objectives and perspectives. To bridge this gap, we collaborate with a public edge platform and perform a comprehensive measurement, which reveals limitations of the status quo VM scheduling schemes and potential opportunities for improvement. However, resolving VM scheduling considering static resource utilization, dynamic bandwidth cost, and end users’ QoE in a cost-efficient manner faces several challenges, including coupled objectives, exponentially increased complexity, and spatiotemporal dynamics. To address the above challenges, in this work, we propose a holistic online framework that integrates combinatorial bandit-based VM migration and seasonality-aware VM request allocation at two distinct time granularities. Large-scale experiments based on a real-world dataset confirm that our online framework achieves near-offline bandwidth cost and resource utilization while significantly lowering time consumption. Xiao Ma 0009, Zhe Fu 0005, Ao Zhou 0001, Mengwei Xu 0001, Shangguang Wang |
IEEE Trans. Mob. Comput. | 4 |
| 2025 | Delay- and Resource-Aware Satellite UPF Service OptimizationabstractExecuting 5G core network functions on satellites has become crucial to enhance satellite network management and service capabilities. The User Plane Function (UPF) is responsible for efficient data traffic forwarding and is envisioned as a key and pioneering core network function that will be deployed on satellites. However, managing and providing services with satellite UPFs face dual challenges. Limited satellite resources constrain the user scale that a satellite UPF can service, resulting in an unguaranteed service delay. Moreover, the extremely rapid mobility of satellites renders it difficult for satellite UPFs to provide seamless services. To address the above challenges, this paper presents the first-of-its-kind service optimization scheme for satellite UPFs in terms of switch control, state migration, and traffic routing. To provide guaranteed service delay, we provide a theoretical analysis based on the M/G/1 queue model, demonstrating the service delay-resource consumption trade-off. A satellite UPF switch control scheme is integrated into the service optimization process, which can decrease satellite UPF service delay while saving satellite resources by adjusting the switch control parameters. To provide seamless services, we propose a satellite UPF-oriented state-aware service migration and traffic routing (UPF service optimization) algorithm. A policy network-based reinforcement learning approach is employed to dynamically perceive the satellite network’s state as well as the satellite UPF switch state. Building upon the optimization of service delay through satellite UPF switch control, the processes of state-aware state migration and traffic routing are further employed to reduce delay, ensuring seamless service effectively. Experiments reveal that the proposed algorithm outperforms other benchmark algorithms under different metrics. The service delay is reduced by an average of 23.2% compared with other algorithms. Chao Wang 0093, Xiao Ma 0009, Ruolin Xing, Ao Zhou 0001, Shangguang Wang |
IEEE Trans. Mob. Comput. | 5 |
| 2025 | EdgeMoE: Empowering Sparse Large Language Models on Mobile DevicesabstractLarge language models (LLMs) such as GPTs and Mixtral-8x7B have revolutionized machine intelligence due to their exceptional abilities in generic ML tasks. Transiting LLMs from datacenters to edge devices brings benefits like better privacy and availability, but is challenged by their massive parameter size and thus unbearable runtime costs. To this end, we presentEdgeMoE, an on-device inference engine for mixture-of-expert (MoE) LLMs – a popular form of sparse LLM that scales its parameter size with almost constant computing complexity.EdgeMoEachieves both memory- and compute-efficiency by partitioning the model into the storage hierarchy: non-expert weights are held in device memory; while expert weights are held on external storage and fetched to memory only when activated. This design is motivated by a key observation that expert weights are bulky but infrequently used due to sparse activation. To further reduce the expert I/O swapping overhead,EdgeMoEincorporates two novel techniques: (1) expert-wise bitwidth adaptation that reduces the expert sizes with tolerable accuracy loss; (2) expert preloading that predicts the activated experts ahead of time and preloads it with the compute-I/O pipeline. On popular MoE LLMs and edge devices,EdgeMoEshowcase significant memory savings and speedup over competitive baselines. Rongjie Yi, Shiyun Wei, Ao Zhou 0001, Shangguang Wang, Mengwei Xu 0001 |
IEEE Trans. Mob. Comput. | 4 |
| 2025 | SLICE: Energy-Efficient Satellite-Ground Co-Inference via Layer-Wise Scheduling OptimizationabstractRecent advancements in Low Earth Orbit (LEO) satellites are facilitating the provision of Deep Neural Networks (DNNs)-inherent services to achieve ubiquitous coverage via satellite computing. However, the computational demands and energy consumption of DNN models present significant challenges for satellite computing with limited power and computation resources. Based on the layered characteristics of DNN models, a satellite-ground co-inference strategy has been introduced, which executes certain layers on satellites and the remaining layers on ground servers. Determining the optimal layers for in-orbit processing, however, is non-trivial due to the under-explored energy consumption of satellite computing across different models and restricted yet varying communication conditions of satellite-ground links. In this paper, we first conduct a comprehensive measurement to uncover energy consumption of satellite computing across different layers and models. By summarizing the key observations, we develop a layer-specific energy consumption model tailored to diverse DNN architectures and kernels. We then investigate the energy-efficient satellite-ground co-inference problem and formulate it as an integer-nonlinear programming problem, which presents high computational complexity. To tackle these difficulties, we propose a satellite-ground co-inference algorithm that employs a branch-and-bound strategy, combined with the Sobol sequence and Lagrange multiplier, to reduce complexity and ensure stability across diverse DNN architectures. To evaluate the proposed algorithm, we conduct experiments based on real-world satellite parameters. The results demonstrate that our proposed algorithm can achieve an average energy savings of 96% under various data volumes compared to the existing benchmarks. Qiyang Zhang 0001, Ruolin Xing, Yuanzhe Li 0001, Xiao Ma 0009, Ao Zhou 0001, Shangguang Wang |
IEEE Trans. Serv. Comput. | 7 |
| 2025 | A Resource-Efficient Multiple Recognition Services Framework for IoT DevicesabstractDeploying the convolutional neural network (CNN) model on Internet of Things (IoT) devices to provide diverse recognition services has received increasing attention. Due to the limited storage, computing, and other resources of IoT devices, it has become mainstream to first train the CNN model on the edge/cloud server and then send the trained CNN to the IoT device. However, most existing related methods suffer from two limitations, (i) low performance due to service interference or insufficient mutual assistance, and (ii) large memory resources and switching resource overhead. To this end, this article proposes a resource-efficient multiple recognition services framework for IoT devices. The proposed framework is based on the edge server-assisted IoT device training of the CNN model, and the framework includes a deeper weight adaptation (DeepWAdapt) algorithm to mitigate service interference. The DeepWAdapt algorithm consists of a set of learnable masks, and by inserting these masks into the appropriate layers of the CNN model, it mitigates mutual interference between services caused by training a single CNN model for multiple services. Each service has a specific set of masks. These learnable masks work like keys for each service, selecting appropriate and specific features for each service from a shared feature set. Experimental results demonstrate that the DeepWAdapt outperforms other state-of-the-art methods on image-level classification services and pixel-level dense prediction services. Specifically, when executing 40 services based on ResNet18, the proposed DeepWAdapt achieves 66.82% F1-score on the CelebA dataset, which is +2.61% F1-score than the previous state-of-the-art result. In addition, compared with the routing method, our proposed DeepWAdapt also reduces network transmission traffic by approximately 35%. Chuntao Ding, Ao Zhou 0001, Yidong Li, Shangguang Wang |
IEEE Trans. Serv. Comput. | 3 |
| 2025 | ReFrame: A Resource-Friendly Cloud-Assisted On-Device Deep Learning Framework for Vision ServicesabstractCloud-assisted Internet of Things (IoT) device deployment of deep neural networks (DNNs) promotes On-device deep learning to provide users with ubiquitous high-quality services by solving the contradiction between insufficient IoT device resources and intensive demand for high-performance DNN resources. However, most existing methods optimize DNNs by considering one or two terms of transmission, computation, and storage resources, but do not consider all three terms at the same time in cloud-assisted IoT device deployment and updating DNNs. To this end, we propose a non-learnable module-based ResNet and a cloud-assisted on-device deep learning framework, ReFrame, based on the consideration of three indicators: model transmission parameters, computation resources, and storage resources. In the proposed method, we first specify that some parameters in DNNs are non-learnable and randomly initialized, so that, these parameters can be saved and reproduced with a few random seeds. By doing so, the cloud only transmits random seeds and learnable parameters to reduce the number of parameter transmissions. Second, we reduce the computation resource consumption of the model by introducing computation-friendly operators, such as pooling, to replace vanilla convolutions. Finally, since random seeds are used to save non-learnable model parameters, on IoT devices we only need to store random seeds and learnable parameters to reproduce the well-trained model. Compared with saving the complete model, our method greatly reduces IoT device storage resource consumption. Experimental results on image classification, object detection, and semantic segmentation tasks demonstrate the effectiveness of the proposed method. Specifically, on the CIFAR-10, our proposed method reduces approximately 89% of FLOPs and 90% of transmitted data in the prototype system compared to ResNet-18. Jianhang Xie, Chuntao Ding, Qingji Guan, Ao Zhou 0001, Yidong Li |
IEEE Trans. Serv. Comput. | 4 |
| 2025 | Flexible Shadow: Resource-Efficient Reliability Enhancement for Edge Services Through Dynamic Shadow CoordinationabstractEdge computing plays a pivotal role in supporting services necessitating sub-second latency, notably in domains like Industry 4.0 and autonomous driving. However, unpredictable failure occurring at edge servers can result in prolonged response time and decreased service reliability, posing significant risks to both safety and property. Traditional reliability mechanisms, namely task re-execution and task replication, are often inadequate for edge environments. The former struggles to meet the stringent end-to-end service latency requirements, while the latter imposes a high resource consumption burden on resource-limited edge clouds. To address this issue, this paper introduces a novel Flexible Shadow mechanism, where the backup instance, referred to as the Flexible Shadow, is allocated fewer computation resources compared to its primary instance to conserve computation resources, and temporally preempts a portion of resources from neighboring shadows to accelerate when necessary. To support the implementation of this mechanism, we propose the Flexible Shadow Backup Framework, a resource-efficient reliability enhancement framework for edge services through dynamic shadow coordination. This framework integrates three key components: a deployment algorithm for resource allocation, an adjustment algorithm for migration cost-latency tradeoffs, and a reconfiguration algorithm for adaptation optimization. Comprehensive experiments conducted on a Docker-based prototype demonstrate the effectiveness of the Flexible Shadow mechanism, achieving nearly 60% reduction in computing resource consumption compared to traditional approaches while maintaining sub-second latency. Lipei Yang, Ao Zhou 0001, Xiao Ma 0009, Qing Li 0028, Yuanzhe Li 0001, Shangguang Wang |
IEEE Trans. Serv. Comput. | 2 |
| 2025 | FedCLR+: Tackling Onboard Label Constraints for Accurate Federated Satellite ComputingabstractThe rapid growth of Low Earth Orbit (LEO) satellites, particularly with the increasing deployment of intelligent computing capabilities using commercial off-the-shelf (COTS) hardware, presents significant opportunities to enhance the quality of in-orbit services. However, the current onboard conditions remain insufficient to enhance model accuracy by increasing model size, and inadequate accuracy hampers the effectiveness of in-orbit services. The satellite-ground federated learning (FL) paradigm, leveraging collaborative fine-tuning, offers a promising solution to continuously improve onboard model performance. Prior studies have focused on optimizing fine-tuning under constraints like limited bandwidth and computational resources, they often overlook two critical challenges: the scarcity and skewness of labeled onboard data and the long revisit cycles of satellites. To address these challenges and better support in-orbit services, this paper designs a realistic simulation methodology for the onboard fine-tuning process and conducts a comprehensive measurement study. Based on insights from the measurement results, we propose an efficient satellite-ground federated fine-tuning system,FedCLR+. In this system, we design a FedCLR algorithm to enhance system accuracy through representation optimization. Additionally, we propose a hybrid bias-compensated strategy to further mitigate accuracy loss by enriching the diversity of aggregation information. Experimental results show thatFedCLR+significantly enhances accuracy by up to 21.61×, reduces transmission volume by an average of 7.29%, and maintaining acceptable additional overhead compared to baselines. Chen Yang 0043, Qiyang Zhang 0001, Qibo Sun, Shufeng Ouyang, Ao Zhou 0001, Shangguang Wang, Mengwei Xu 0001 |
IEEE Trans. Serv. Comput. | 5 |
| 2024 | Profit-Aware Task Allocation in Satellite ComputingabstractThe rapid evolution of satellite networks promises to expand global internet service. However, optimizing task allocation for the efficient and sustainable operation of satellite computing presents complex challenges. Existing approaches usually prioritize energy considerations while neglecting economic aspects, which restricts satellite networks from achieving their full economic potential. In this paper, we address this gap by investigating task allocation in satellite computing. Our approach encourages satellites to consistently provide resources and optimizes battery usage, enabling the completion of more tasks and ultimately maximizing profit. The task allocation approach involves two key components: task pricing and task scheduling. Firstly, we introduce a unique task pricing algorithm that adheres to economic properties, establishing a direct link between satellite utilization and financial income, ensuring economically viable satellite operations. Moreover, we develop two distinct task scheduling algorithms tailored for offline and online scenarios, exploiting dynamic programming and reinforcement learning respectively. Extensive simulations demonstrate that our proposed algorithms effectively enhance task completion rates and optimize total satellite profit. Jie Huang 0021, Ruolin Xing, Xiao Ma 0009, Ao Zhou 0001, Shangguang Wang |
ICWS | 4 |
| 2024 | Flexible Shadow: Enhancing Service Reliability in Resource-Constrained Edge ComputingabstractEdge computing plays a pivotal role in supporting services necessitating sub-second latency, notably in domains like Industry 4.0 and autonomous driving. However, unpredictable failure occurring at edge servers can result in prolonged response time and decreased service reliability, posing significant risks to both safety and property. Traditional reliability mechanisms, namely re-execution and replication, are often inadequate for edge environments, with the former often struggle to meet latency requirements and the latter imposing a high resource consumption burden on resource-limited edge clouds. To address this issue, this paper introduces a novel "Flexible Shadow" mechanism, where the backup instance, referred to as the "Flexible Shadow", is allocated fewer computation resources compared to its primary instance to conserve computation resources, and temporally preempts a portion of resources from neighboring shadows to accelerate when necessary. To tackle the implementation challenges of this framework arising from diverse service requirements, dynamic nature of edge environments and potential deadlock in the reconfiguration process, we devised the Flexible Shadow Deployment Algorithm for accurate shadow deployment and the Flexible Shadow Reconfiguration Algorithm for dynamic strategy adjustment. We have implemented our Flexible Shadow framework on Docker and evaluated it via comprehensive experiments. The experiment results demonstrate a nearly 60% reduction in computing resource consumption while ensuring sub-second latency. Lipei Yang, Ao Zhou 0001, Xiao Ma 0009, Yuanzhe Li 0001, Shangguang Wang |
ICWS | 2 |
| 2024 | Energy-Aware Satellite-Ground Co-Inference via Layer-Wise Processing Schedule OptimizationabstractRecent advancements in Low Earth Orbit (LEO) satellites are facilitating the provision of Deep Neural Networks (DNNs)-inherent services to achieve ubiquitous coverage via satellite computing. However, the computational demands and energy consumption of DNN models pose significant challenges for satellite computing with limited power and computation resources. Based on the hierarchical characteristics of DNN models, we propose a satellite-ground co-inference strategy that executing certain layers on satellites and the remaining layers on ground servers. However, identifying the optimal layers for in-orbit processing with latency constraints is challenging due to the uncertain energy consumption across diverse models. To explore the correlation between energy consumption and layer types, we conduct comprehensive measurements on a hardware device commonly found in commercial LEO satellites and develop a layer-based energy consumption prediction model. Then, we formulate an optimization problem of minimizing the energy consumption on the satellite within the latency constraint as an integer nonlinear programming problem. Solving this problem is difficult due to combinatorial explosion in the discrete solution space. To address this, we propose an improved algorithm based on genetic algorithms. Using configurations from a real satellite, we conduct simulation experiments, concluding that our algorithm significantly improves energy savings by an average of 27 ×. Qiyang Zhang 0001, Ruolin Xing, Yuanzhe Li 0001, Xiao Ma 0009, Chaoxin Yu, Ao Zhou 0001, Shangguang Wang |
Internetware | 8 |
| 2024 | Deciphering the Enigma of Satellite Computing with COTS Devices: Measurement and AnalysisabstractIn the wake of the rapid deployment of large-scale low-Earth orbit satellite constellations, exploiting the full computing potential of Commercial Off-The-Shelf (COTS) devices in these environments has become a pressing issue. However, understanding this problem is far from straightforward due to the inherent differences between the terrestrial infrastructure and the satellite platform in space. In this paper, we take an important step towards closing this knowledge gap by presenting the first measurement study on the thermal control, power management, and performance of COTS computing devices on satellites. Our measurements reveal that the satellite platform and COTS computing devices significantly interplay in terms of the temperature and energy, forming the main constraints on satellite computing. Further, we analyze the critical factors that shape the characteristics of onboard COTS computing devices. We provide guidelines for future research on optimizing the use of such devices for computing purposes. Finally, we have released the datasets to facilitate further study in satellite computing. Ruolin Xing, Mengwei Xu 0001, Ao Zhou 0001, Qing Li 0028, Feng Qian 0001, Shangguang Wang |
MobiCom | 3 |
| 2024 | Poster: Service Orchestration for Satellite ComputingabstractSatellite computing is emerging as a promising domain for delivering mobile services that meet stringent Quality of Service (QoS) requirements, such as low latency, to users. However, the inherent mobility of satellites as computing nodes can precipitate QoS degradation, a challenge not encountered in terrestrial cloud systems. This discrepancy poses significant adaptation challenges for cloud service orchestration systems, such as Kubernetes, due to the rapid movement of satellites. This poster introduces a service orchestration system and a corresponding service placement strategy tailored for satellite computing environments. Our proposed architecture and strategy surpass traditional fixed instance deployment by not only achieving lower average latency but also maintaining an optimal balance between benefits and costs. Ruolin Xing, Qibo Sun, Ao Zhou 0001, Xiao Ma 0009 |
MobiSys | 4 |
| 2024 | An Enhancement Framework for RDMA Congestion Control in Multi-tenant DatacentersabstractRecently, Remote Direct Memory Access (RDMA) is gradually gaining popularity in multi-tenant datacenters due to its high performance and low CPU utilization. RDMA requires a lossless underlying network to fully realize its potential. Thus, hop-by-hop Priority Flow Control (PFC) is deployed to ensure losslessness, and congestion control is also needed to allocate per-flow bandwidth. However, we find that existing RDMA congestion control schemes in multi-tenant datacenters fail to achieve rapid rate convergence due to the heuristic nature and may exacerbate PFC side effects including head-of-line blocking, unfairness, and even deadlocks. In this paper, we propose an enhancement framework for RDMA congestion control called RDI. RDI has several noteworthy properties. First, RDI utilizes valuable receiver-side information and network congestion level to provide a precise guide rate for congested flows, thus alleviating congestion and corresponding PFC issues; Second, RDI is transparent to tenants and compatible with existing RDMA network architectures without the need of modifying in-network devices. We evaluate RDI under realistic traffic traces. The results show that the congestion control scheme enhanced by RDI significantly outperforms the original in terms of throughput and flow completion time, while also reducing the side effects of PFC. For instance, RDI-enhanced congestion control shortens by up to 44% and 60% for the 99th-percentile tail and average FCT, respectively. Ao Zhou 0001, Shangguang Wang |
NOMS | 3 |
| 2024 | Exploring Real-Time Satellite Computing: From Energy and Thermal PerspectivesabstractSmall satellites (SmallSats) are now widely used in various fields, such as real-time communication and earth observation. These increasingly complex space applications face limited support from conventional radiation-hardened processors onboard. Hence, many SmallSats are designed to utilize high performance commercial off-the-shelf (COTS) computing devices to address this problem but it remains unclear how the unique energy and thermal characteristics of SmallSats impact computing efficiency onboard. This work conducts a systematic and quantitative measurement study of COTS devices’ computing efficiency on two real orbiting SmallSats. The key findings are: 1) inadequate energy management may lead to electricity wastage in sunlit zones and shortages in eclipse zones, impacting onboard computing availability and 2) the weak heat dissipation onboard may compromise COTS computing efficiency by incurring thermal throttling. To address such challenges, we design ProScale, a lightweight application-aware power management and thermal control system to improve computing efficiency under both electrical and thermal energy constraints. Evaluation shows that ProScale can improve the average task completion latency by $2.1 \times$ for computation-intensive applications compared with baselines. Qing Li 0028, Shangguang Wang, Chenren Xu, Xiao Ma 0009, Mengwei Xu 0001, Ao Zhou 0001, Ruolin Xing, Zuo Zhu, Ying Zhang 0012, Xuanzhe Liu |
RTSS | 6 |
| 2024 | More is Different: Prototyping and Analyzing a New Form of Edge Server with Massive Mobile SoCs
Li Zhang 0133, Zhe Fu 0005, Boqing Shi, Xiang Li 0067, Rujin Lai, Chenyang Yang 0004, Ao Zhou 0001, Xiao Ma 0009, Shangguang Wang, Mengwei Xu 0001 |
USENIX ATC | 7 |
| 2024 | Reliability-Aware Task Replication for Mobile Edge ComputingabstractAs infrastructure deployment continues to expand worldwide, the development of the Internet of Vehicles has become increasingly feasible. With widespread cellular connectivity and powerful roadside computing capabilities, advanced driving assistance systems can now rely on roadside decision models in addition to vehicle-side ones, overcoming the limitations of a single vehicle’s perception range. This shift has led to improvements in manufacturing efficiency, cruising range, and battery life of intelligent vehicles. However, maintaining the ultra-low latency and high-reliability requirements of on-vehicle services is still a challenge due to air interface fluctuations and edge server computing loads, which could potentially jeopardize driving safety. To tackle this issue, we conducted real-world measurements of edge server access delay in LTE and 5G cellular networks. Our analysis identified key factors affecting delay distribution, leading to the development of an approximate fitting function for the delay probability density function. We also proposed a reliability-aware task replication algorithm that leverages delay samples and edge server status information to make real-time task replication and offloading decisions, minimizing replication while ensuring service reliability. Simulations based on real-world datasets indicate our approach reduces task completion delay by up to 42.11% and limits the maximum task replication redundancy peak value to 63.37%, effectively ensuring the reliability of on-vehicle services during driving. Lipei Yang, Ao Zhou 0001, Xiao Ma 0009, Shangguang Wang |
IEEE Internet Things J. | 2 |
| 2024 | A Reinforcement Learning-Based Incentive Mechanism for Task Allocation Under Spatiotemporal CrowdsensingabstractWith the development of the Industrial Internet of Things (IoT), the work of large-scale data collection makes spatiotemporal crowdsensing (SC) play an important role. Mobile devices equipped with sensors could act as workers to collect and process data for uploading. In the task allocation process, a fully static allocation fails to meet the needs of realistic conditions, while a completely dynamic allocation fails to achieve the desired results. Therefore, we assume a task-scheduled execution scenario that combines the above two conditions. In the pre-allocation process, an original time location constraints (ORTA) allocation algorithm is first proposed. Then it is optimized (OPTA) to fully utilize the remaining time of the workers and increase the matched number. In addition, the design of the incentive mechanism is an effective means to improve the task completion rate of the platform. To efficiently utilize the limited platform budget in the long run, a Q-learning-based algorithm is proposed to identify target inspire tasks and subsequently increase their reward to attract workers’ participation. Finally, comparison experiments are conducted on real datasets to verify the effectiveness of our algorithm. Furthermore, the experiments on a Raspberry Pi local terminal are conducted under a satellite-based environment. Kaige Jiang, Yingjie Wang 0002, Zhaowei Liu 0001, Qilong Han, Ao Zhou 0001, Chaocan Xiang, Zhipeng Cai 0001 |
IEEE Trans. Comput. Soc. Syst. | 6 |
| 2024 | Seamless Cross-Edge Service Migration for Real-Time Rendering ApplicationsabstractSeamless cross-edge migration for real-time rendering applications is challenging. The strong interactive nature of real-time rendering applications demands a downtime lower than 15ms to achieve an imperceptible migration. Existing methods based on virtual machine migration and container migration suffer from unpleasant downtime brought by dirty page retransmission-induced repeated memory data copy and the shared storage failure-induced extensive disk data copy. In this paper, we propose Cloud-assisted Service Migration (CSM) which leverages cloud-edge collaboration to achieve seamless service migration for real-time rendering applications. CSM improves service migration user experience in three folds: First, it introduces a dual rendering mechanism to bypass the peer-to-peer data copy and compresses the freezing stage. Second, a user equipment-centric session switch mechanism is proposed to save time by well coordinating application session switches and 5G user plane session switches. Third, a smooth switching mechanism is leveraged to prevent unpleasant frame flickers during session switching. We implement CSM in edge-rendering multiplayer games and deploy it on a 5G test bed with a full-stack user plane protocol stack. The evaluation results show that CSM can reduce downtime to < 14ms and the service migration process is user imperceptible. Yuanzhe Li 0001, Shangguang Wang, Yuanchun Li 0003, Ao Zhou 0001, Mengwei Xu 0001, Xiao Ma 0009, Yunxin Liu 0001 |
IEEE Trans. Mob. Comput. | 4 |
| 2024 | Communication-Efficient Satellite-Ground Federated Learning Through Progressive Weight QuantizationabstractLarge constellations of Low Earth Orbit (LEO) satellites have been launched for Earth observation and satellite-ground communication, which collect massive imagery and sensor data. These data can enhance the AI capabilities of satellites to address global challenges such as real-time disaster navigation and mitigation. Prior studies proposed leveraging federated learning (FL) across satellite-ground to collaboratively train a share machine learning (ML) model in a privacy-preserving mechanism. However, they mostly focus on single unique challenges such as limited ground-to-satellite bandwidth, short connection window, and long connection cycle, while ignoring the completeness of these challenges in deploying efficient FL frameworks in space. In this paper, we propose an efficient satellite-ground FL framework, SatelliteFL, to address these three challenges collectively. Its key idea is to ensure that each satellite must complete per-round training within each connection window. Moreover, we design a progressive block-wise quantization algorithm that determines a unique bitwidth for each block of the ML model to maximize the model utility while not exceeding the connection window. We evaluate SatelliteFL by plugging an implemented FL platform into real-world satellite networks and satellite images. The results show that SatelliteFL highly accelerates the convergence by up to 2.8× and improves the bandwidth utilization ratio by up to 9.3× compared to the state-of-the-art methods. Chen Yang 0043, Jinliang Yuan, Yaozong Wu, Qibo Sun, Ao Zhou 0001, Shangguang Wang, Mengwei Xu 0001 |
IEEE Trans. Mob. Comput. | 5 |
| 2024 | Battery-Aware Energy Optimization for Satellite Edge ComputingabstractSatellite edge computing can incur dramatically increased energy demand onboard, which is met by satellite batteries during eclipses. Excessive energy usage during regular operations accelerates battery wear. Therefore, it is important and timely to optimize the energy consumption onboard to extend satellite batteries life. This paper investigates battery-aware energy optimization for satellite edge computing under energy harvesting dynamics and wireless environment uncertainty. Inspired by the periodical energy harvesting and satellite-ground connection, we develop a pattern-aware online energy scheduling algorithm within an online convex optimization framework. This learning algorithm achieves theoretical guarantees of no regret and gradually zeroing constraint violations. We further exploit inter-satellites collaboration to extend the average battery life in a whole constellation where satellites have different battery capacity degradation. Trace-driven simulations show that our algorithm can significantly extend the battery life by 1.32× and effectively adapt to the energy harvesting dynamics and wireless environment uncertainty. Qing Li 0028, Shangguang Wang, Xiao Ma 0009, Ao Zhou 0001, Yue Wang 0072, Gang Huang 0001, Xuanzhe Liu |
IEEE Trans. Serv. Comput. | 4 |
| 2023 | Cooperative Content Caching and Distribution for Satellite CDNsabstractContent distribution networks play a crucial role in the services provided by streaming applications, but current terrestrial CDNs have uneven coverage and poor performance in some remote and rural areas. Satellite communications have emerged as an effective way to address these issues. However, the satellite environment is highly dynamic, and how to reasonably deploy CDNs in the satellite constellation to realize efficient content caching and distribution is an urgent problem to be solved. To this end, this paper proposes a content caching and content distribution approach for satellite CDNs. The approach uses Lyapunov optimization, Gibbs sampling, and matching theory to control the traffic cost of satellite-cached content copies while reducing the content distribution latency. A cost-effective content cache deployment and distribution strategy is proposed for large low-orbiting satellite constellations operating at high speeds. Extensive simulation results show that the algorithm can quickly converge to favorable values compared to the reference algorithm, effectively reducing latency for end users while maintaining low deployment traffic costs. Xiao Ma 0009, Ao Zhou 0001, Shangguang Wang |
ICNP | 3 |
| 2023 | Energy and Time-Aware Inference Offloading for DNN-based Applications in LEO SatellitesabstractIn recent years, Low Earth Orbit (LEO) satellites have witnessed rapid development, with inference based on Deep Neural Network (DNN) models emerging as the prevailing technology for remote sensing satellite image recognition. However, the substantial computation capability and energy demands of DNN models, coupled with the instability of the satellite-ground link, pose significant challenges, burdening satellites with limited power intake and hindering the timely completion of tasks. Existing approaches, such as transmitting all images to the ground for processing or executing DNN models on the satellite, is unable to effectively address this issue. By exploiting the internal hierarchical structure of DNNs and treating each layer as an independent subtask, we propose a satellite-ground collaborative computation partial offloading approach to address this challenge. We formulate the problem of minimizing the inference task execution time and onboard energy consumption through offloading as an integer linear programming (ILP) model. The complexity in solving the problem arises from the combinatorial explosion in the discrete solution space. To address this, we have designed an improved optimization algorithm based on branch and bound. Simulation results illustrate that, compared to the existing approaches, our algorithm improve the performance by 10%-18%. Qiyang Zhang 0001, Xiao Ma 0009, Ao Zhou 0001 |
ICNP | 5 |
| 2023 | Freshness-aware Content Update for Earth Observation: Trading Off AoI and AccuracyabstractSpace networks composed of Low Earth Orbit (LEO) satellites play a significant role in Earth observation systems. LEO satellite constellations, which provide wide-ranging coverage, continuously observe the Earth and transmit the observed data to satellites in higher orbits or ground stations for further analysis, benefiting from their more powerful computation capacity. In such an observation and analysis network, the timeliness of data update and the quality of analysis are crucial factors, both constrained by limited energy resources. In this paper, we jointly optimize the timeliness and analysis quality by constructing an update scheduling problem with a trade-off between Age of Information (AoI) and accuracy. Considering the unpredictable and changeable environment dynamics, we formulate this problem into a Markov decision process and propose an algorithm based on deep reinforcement learning to obtain an optimal policy that effectively tackles the tradeoff problem. Simulation results demonstrate that our algorithm can jointly reduce AoI and increase accuracy, outperforming benchmark algorithms in different cases. Yuran Guo, Xiao Ma 0009, Qibo Sun, Ao Zhou 0001 |
ICNP | 4 |
| 2023 | Optimizing Space-Borne Computation: A Reliability Enhancement Framework for LEO ConstellationabstractEdge computing is extending the frontier of computation beyond terrestrial boundaries, with Low Earth Orbit (LEO) constellations emerging as the cutting-edge paradigm, aspiring to deliver ubiquitous computing capabilities globally. However, the unresolved service reliability issues within LEO constellations continue to barricade the realization of the vision for space computing. Terrestrial-native classical methods struggle to meet the stringent environmental conditions of space computing and often fail to satisfy the rigorous requirements of LEO constellations for failure response times and computational resource overhead. To address these challenges, we introduce the LEO Service Reliability Enhancement Framework (LSREF). LSREF employs the Variable Speed Replication technology to balance failure response times and computational resource overhead and adopts a decentralized design, more congruent with the characteristics of LEO constellations. Specifically, to counter the immense scale and high dynamism of LEO constellations, LSREF proposes orbital plane-based autonomous domains and leverages the Satellite Autonomous Backup Deployment Algorithm in conjunction with the Registry Polling Mechanism to enable autonomous decision-making for service backup strategies on each satellite. Our simulation experiments demonstrate that, compared to classical methods, LSREF reduces average failure response times by 10.76% and diminishes computational resource consumption by nearly 15.53%. Lipei Yang, Ao Zhou 0001, Xiao Ma 0009, Shangguang Wang |
ICPADS | 2 |
| 2023 | Container Image Similarity-Aware Resource Provisioning for Serverless Edge ComputingabstractContainer-enabled serverless computing has become a widely adopted approach for resource provisioning in the edge cloud. However, traffic incurred by container image pulling heavily burdens the already congested back-haul network. To relieve the problem, we do an analysis on Docker Hub, and find that instance deployment strategy has a significant impact on the back-haul traffic due to the varying similarity levels of different images. We incorporate this feature into task offloading decision and resource provisioning, and formulate the problem with a mixed integer non-linear programming (MINLP) problem. To address the challenges arising from the coupling and contradiction of instance deployment, image pulling, offloading decision, and resource allocation, we employ multi-agent deep reinforcement learning to decompose the problem into several simpler sub-problems, and design an algorithm for each sub-problem individually by exploiting convex optimization and fractional programming techniques. Simulations are conducted to validate the effectiveness of the proposed algorithm. The experiment results illustrate that our algorithm outperforms current notable solutions and improves the global utility by 13%–74%. Ao Zhou 0001, Xiao Ma 0009, Shangguang Wang |
ICWS | 1 |
| 2023 | Privacy as a Resource in Differentially Private Federated LearningabstractDifferential privacy (DP) enables model training with a guaranteed bound on privacy leakage, therefore is widely adopted in federated learning (FL) to protect the model update. However, each DP-enhanced FL job accumulates privacy leakage, which necessitates a unified platform to enforce a global privacy budget for each dataset owned by users. In this work, we present a novel DP-enhanced FL platform that treats privacy as a resource and schedules multiple FL jobs across sensitive data. It first introduces a novel notion of device-time blocks for distributed data streams. Such data abstraction enables fine-grained privacy consumption composition across multiple FL jobs. Regarding the non-replenishable nature of the privacy resource (that differs it from traditional hardware resources like CPU and memory), it further employs an allocation-then-recycle scheduling algorithm. Its key idea is to first allocate an estimated upper-bound privacy budget for each arrived FL job, and then progressively recycle the unused budget as training goes on to serve further FL jobs. Extensive experiments show that our platform is able to deliver up to 2.1× as many completed jobs while reducing the violation rate by up to 55.2% under limited privacy budget constraint. Jinliang Yuan, Shangguang Wang, Shihe Wang, Yuanchun Li 0003, Xiao Ma 0009, Ao Zhou 0001, Mengwei Xu 0001 |
INFOCOM | 6 |
| 2023 | Boosting DNN Cold Inference on Edge DevicesabstractDNNs are ubiquitous on edge devices nowadays. With its increasing importance and use cases, it's not likely to pack all DNNs into device memory and expect that each inference has been warmed up. Therefore, cold inference, the process to read, initialize, and execute a DNN model, is becoming commonplace and its performance is urgently demanded to be optimized. To this end, we present NNV12, the first on-device inference engine optimizing cold inference. NNV12 is built atop three novel optimization knobs: selecting a proper kernel (i.e., operator implementation) for each DNN operator, bypassing the weights transformation process by caching the post-transformed weights on disk, and pipelined execution of many kernels on asymmetric processors. To tackle with the huge search space, NNV12 employs a heuristic-based scheme to obtain a near-optimal kernel scheduling plan. We fully implement a prototype of NNV12 and evaluate its performance across extensive experiments. It shows that NNV12 achieves up to 15.2× speedup compared to the state-of-the-art DNN engines on edge CPUs and 401.5× speedup on edge GPUs, respectively. Rongjie Yi, Ting Cao 0003, Ao Zhou 0001, Xiao Ma 0009, Shangguang Wang, Mengwei Xu 0001 |
MobiSys | 3 |
| 2023 | Planning-based mobile crowdsourcing bidirectional multi-stage online task assignment
Yingjie Wang 0002, Bingyi Xie, Lingkang Meng, Zhaowei Liu 0001, Xiangrong Tong, Ao Zhou 0001, Zhipeng Cai 0001 |
Comput. Networks | 7 |
| 2023 | Service Routing in Multi-Tier Edge Computing: A Matching Game ApproachabstractAlthough microservice envisioned as a promising approach for edge applications which can improves the development efficiency and the deployment productivity, it also leads to the operational complexity during service runtime. This leads to the emergence of service mesh, a decided infrastructure layer over microservices for reliable service-to-service communication. However, when migrating the service mesh to the multi-tier edge computing, the corresponding extensions on the top of service mesh is needed to overcome the challenges brought by the multi-tier edge servers, shared microservices and diverse QoS requirements. Therefore, in this paper, we investigate the problem of service routing to achieve the efficient microservice-based service provision in multi-tier edge computing. The objective is to route user requests to the optimal microservice instances with low service delay and resource cost. To this end, we first formulate the service routing problem as an integer non-linear optimization problem, which is NP-hard. Then we map it to a many-to-one matching problem by employing the matching game theory and propose a dependency-aware deferred acceptance algorithm with dynamic quota. The experimental results based on a real-world dataset demonstrate that our proposed algorithm can significantly outperform existing representative algorithms in terms of service delay and resource cost. Shangguang Wang, Yan Guo 0004, Xiao Liu 0004, Ao Zhou 0001 |
IEEE J. Sel. Areas Commun. | 4 |
| 2023 | Towards Diversified IoT Image Recognition Services in Mobile Edge ComputingabstractWith the rapid development of the Internet of Things (IoT) and emerging Mobile Edge Computing (MEC) technologies, various IoT image recognition services are revolutionizing our lives by providing diverse cognitive assistance. However, most existing related approaches are difficult to meet the diversified needs of users because they believe that the MEC platform is a single layer. In addition, due to the mutual interference between the data, it is not easy for them to extract the discriminative features (DFs) necessary to analyze the input data. To this end, this article proposes an IoT image recognition services framework for different needs in the MEC environment, which consists of Hierarchical Discriminative Feature Extraction (HDFE) and Sub-extractor Deployment (Sub-ED) algorithms. We first propose HDFE, which can avoid mutual interference between data by separately optimizing the data structure, thereby generating an extractor that extracts effective DFs. Then there is Sub-ED, which divides the extractor into a series of sub-extractors and deploys them on appropriate MEC platforms. By doing so, the IoT device can connect to the corresponding MEC platform according to its service types, and use the sub-extract to extract DFs. Then, the MEC platform uploads the extracted feature data to the cloud server for further processing, e.g., feature matching. Finally, the cloud server sends the processed result back to the IoT device. Experimental results show that compared with the state-of-the-art approaches, the proposed framework improves recognition accuracy by about 6% and reduces network traffic by up to 94%. Chuntao Ding, Ao Zhou 0001, Xiao Ma 0009, Ning Zhang 0007, Ching-Hsien Hsu, Shangguang Wang |
IEEE Trans. Cloud Comput. | 2 |
| 2023 | Online Service Request Duplicating for Vehicular ApplicationsabstractVehicles on roads have increasingly powerful computing capabilities and edge nodes are being widely deployed. They can work together to provide computing services for onboard driving systems, passengers, and pedestrians. Typical applications in vehicular systems have service requirements such as low latency and high reliability. Most studies in vehicular networks concerning latency and reliability focus on vehicular communication at the network level. Based on these fundamental works, an increasing proportion of vehicles boast complex applications that require service-level end-to-end performance guarantees. Several works guarantee service-level latency or reliability while new and innovative applications are demanding a joint optimization of the above two metrics. To address the critical challenges induced by the joint modeling of latency and reliability, system uncertainty, and performance and cost trade-off, we employ service request duplication to ensure both latency and reliability performance at the service level. We propose an online learning-based service request duplication algorithm based on a multi-armed bandit framework and Lyapunov optimization theory. The proposed algorithm achieves an upper-bounded regret compared to the oracle algorithm. Simulations are based on real-world datasets and the results demonstrate that the proposed algorithm outperforms the benchmarks. Qing Li 0028, Xiao Ma 0009, Ao Zhou 0001, Changhee Joo, Shangguang Wang |
IEEE Trans. Mob. Comput. | 3 |
| 2023 | User-Oriented Edge Node Grouping in Mobile Edge ComputingabstractIn mobile edge computing networks, densely deployed access points are empowered with computation and storage capacities. This brings benefits of enlarged edge capacity, ultra-low latency, and reduced backhaul congestion. This paper concerns edge node grouping in mobile edge computing, where multiple edge nodes serve one end user cooperatively to enhance user experience. Most existing studies focus on centralized schemes that have to collect global information and thus induce high overhead. Although some recent studies propose efficient decentralized schemes, most of them did not consider the system uncertainty from both the wireless environment and other users. To tackle the aforementioned problems, we first formulate the edge node grouping problem as a game that is proved to be an exact potential game with a unique Nash equilibrium. Then, we propose a novel decentralized learning-based edge node grouping algorithm, which guides users to make decisions by learning from historical feedback. Furthermore, we investigate two extended scenarios by generalizing our computation model and communication model, respectively. We further prove that our algorithms converge to the Nash equilibrium with upper-bounded learning loss. Simulation results show that our mechanisms can achieve up to 96.99% of the oracle benchmark. Qing Li 0028, Xiao Ma 0009, Ao Zhou 0001, Xiapu Luo, Fangchun Yang, Shangguang Wang |
IEEE Trans. Mob. Comput. | 3 |
| 2023 | Dynamic Task Scheduling in Cloud-Assisted Mobile Edge ComputingabstractThe cloud-assisted mobile edge computing system is a critical architecture to process computation-intensive and delay-sensitive mobile applications in close proximity to mobile users with high resource efficiency. Due to the heterogenous dynamics of task arrivals at edge nodes and the distributed nature of the system, the workloads of edge nodes are prone to be unbalanced, which can cause high task response time and resource cost. This paper solves the dynamic task scheduling problem in cloud-assisted mobile edge computing (including both peer task scheduling among edge nodes and cross-layer task scheduling from edge nodes to the cloud), aiming at minimizing average task response time within resource budget limit. To overcome the challenges of task arrival dynamics, edge node heterogeneity, and computation-communication delay tradeoff, we propose aWater-filling BasedDynamic TaskScheduling (WiDaS) algorithm. WiDaS dynamically tunes the usage of cloud resources based on the Lyapunov optimization method and efficiently schedules mobile tasks among edge nodes (and the cloud) by exploiting the idea of water filling. Extensive simulations are conducted to evaluate WiDaS under a trace-driven traffic pattern and two mathematic traffic patterns. The results demonstrate that WiDaS shows two-fold benefits of efficiency and effectiveness. In terms of efficiency, WiDaS can achieve the approximate results with the KKT-based algorithm while reducing the computation complexity from exponential order to polynomial order. In terms of effectiveness, WiDaS can reduce the average task response time by up to 64.4% and 47.2% over the Fair-ratio and the Edge-first algorithm. Xiao Ma 0009, Ao Zhou 0001, Shan Zhang 0001, Qing Li 0028, Alex X. Liu, Shangguang Wang |
IEEE Trans. Mob. Comput. | 2 |
| 2022 | Collaborative Mobile Edge Computing Through UPF Selection
Yuanzhe Li 0001, Ao Zhou 0001, Xiao Ma 0009, Shangguang Wang |
CollaborateCom (2) | 2 |
| 2022 | Commutativity-guaranteed Docker Image Reconstruction towards Effective Layer SharingabstractOwing to the benefit of light weight, containers have become a promising enabler for cloud native computing. Container images composed of applications and dependencies support flexible service deployment and migration. Rapid adoption and integration of containers generate millions of images to be stored. Additionally, non-local images have to be frequently downloaded from the registry, resulting in huge amounts of traffic. Content Addressable Storage (CAS) has been adopted for saving storage and networking by enabling identical layers sharing across images. However, according to our measurements, the implication of CAS is significantly limited as layers are rarely fully identical in practice. In this paper, we propose to reconstruct the docker images to raise the number of identical layers and thereby reduce storage and network consumption. We explore the layered structure of images and define the commutativity of files to assure image validity. The image reconstruction is formulated as an integer nonlinear programming problem. Inspired by the observed similarity of layers, we design a similarity-aware online image reconstruction algorithm. Extensive evaluations are conducted to verify the performance of the proposed approach. Ao Zhou 0001, Xiao Ma 0009, Mengwei Xu 0001, Shangguang Wang |
WWW | 2 |
| 2022 | A Comprehensive Benchmark of Deep Learning Libraries on Mobile DevicesabstractDeploying deep learning (DL) on mobile devices has been a notable trend in recent years. To support fast inference of on-device DL, DL libraries play a critical role as algorithms and hardware do. Unfortunately, no prior work ever dives deep into the ecosystem of modern DL libs and provides quantitative results on their performance. In this paper, we first build a comprehensive benchmark that includes 6 representative DL libs and 15 diversified DL models. We then perform extensive experiments on 10 mobile devices, which help reveal a complete landscape of the current mobile DL libs ecosystem. For example, we find that the best-performing DL lib is severely fragmented across different models and hardware, and the gap between those DL libs can be rather huge. In fact, the impacts of DL libs can overwhelm the optimizations from algorithms or hardware, e.g., model quantization and GPU/DSP-based heterogeneous computing. Finally, atop the observations, we summarize practical implications to different roles in the DL lib ecosystem. Qiyang Zhang 0001, Xiang Li 0067, Xiangying Che, Xiao Ma 0009, Ao Zhou 0001, Mengwei Xu 0001, Shangguang Wang, Yun Ma 0002, Xuanzhe Liu |
WWW | 5 |
| 2022 | Price-Aware Service Deployment in Hierarchical Mobile-Edge ComputingabstractMobile-edge computing (MEC) is considered as a promising solution to release pressure on the core network and reduce service response time. Edge nodes with storage and computation resources are able to cache various services and process tasks rather than offloading to remote clouds. However, it is difficult to make service deployment decisions appropriately since resources are limited in edge nodes and requirements of services are diverse. The hierarchical MEC structure in the 5G network causes extra complication, and the cost of service deployment aggravates the hardness, especially for service providers. In this article, we focus on the service deployment problem considering service caching, resource allocation, and task scheduling in the hierarchical MEC network, aiming at minimizing monetary cost. To address the heterogeneous limitations and requirements, we formulate the problem as a mixed-integer nonlinear programming problem and develop an iterative service deployment algorithm by exploiting Gibbs sampling. We make the service caching strategies of edge nodes iteratively. Furthermore, we transform the resource allocation and task scheduling optimization into the linear programming problem and employ a typical optimization function. Simulation results show that our algorithm always obtains minimal monetary cost for various number of services and task arrival rates, compared with benchmarks. Jie Huang 0021, Ao Zhou 0001, Shangguang Wang |
IEEE Internet Things J. | 2 |
| 2022 | Profit-Aware Edge Server PlacementabstractIn a 5G network, mobile-edge computing (MEC) plays a key role in providing low access delay services. The placement of edge servers not only determines the quality of services on the user side but also affects the profit of running a MEC system. In this article, we study how to properly place edge servers so as to guarantee the access delay and maximize the profit of edge providers. We first propose a profit model which involves both access delay and energy consumption. In this model, we take the 5G user plane function (UPF) into consideration to calculate access delay for the first time. Then, we devise a particle swarm optimization-based algorithm to optimize the profit. In the algorithm, we introduce a weight value$q$to guarantee the access delay and assign base stations properly. Moreover, a service-level agreement is adopted to balance the tradeoff between access delay and energy consumption. We take advantage of our 5G network emulator called mini5Gedge and data set from Shanghai Telecom to conduct massive experiments. The results show that our algorithm stands out in terms of achieving the highest profit. Yuanzhe Li 0001, Ao Zhou 0001, Xiao Ma 0009, Shangguang Wang |
IEEE Internet Things J. | 2 |
| 2022 | Reliability-Enhanced Task Offloading in Mobile Edge Computing EnvironmentsabstractInternet of Things (IoT) devices have become an integral part of our lives and are increasingly used in almost every field. Subsequently, there are a large number of latency-sensitive IoT applications (e.g., face recognition and autonomous driving) targeted for mobile edge computing environments. These IoT applications are often split into multiple collaborative tasks and offloaded onto containers or virtual machines (VMs) with certain failure rates and recovery rates. If these containers or VMs are not deployed in the same edge servers, the bandwidth resources of edge clouds must be consumed to transfer data. These factors increase the completion time of IoT applications to different degrees, and then affect their reliability level. Therefore, there exists equilibrium between the reliability level and bandwidth consumption. In this article, we investigate the equilibrium of minimizing the bandwidth consumption of IoT applications while maximizing the reliability level of these IoT applications during task offloading. We propose a multiobjective optimization problem, and transform it to a single-objective optimization problem. Furthermore, we introduce two efficient approaches to acquire two near-optimal solutions. The results of simulation experiments demonstrate that our proposed approaches can observably enhance the reliability level and reduce the bandwidth consumption of IoT applications compared with other related approaches. Meanwhile, we also make a comparative analysis of our proposed approaches. Jialei Liu, Ao Zhou 0001, Chunhong Liu, Tongguang Zhang, Lianyong Qi, Shangguang Wang, Rajkumar Buyya |
IEEE Internet Things J. | 2 |
| 2022 | Service-Oriented Resource Allocation for Blockchain-Empowered Mobile Edge ComputingabstractIntegrating dense small cell (DSC) networks with mobile edge computing is employed by 5G to tackle the contradiction between the computation limitations of user equipment (UE) and the stringent latency requirement of services. This paper investigates the service-oriented edge resource allocation problem in DSC networks, determining where to deploy the service entity, how many service entities should be deployed at each edge cloud, and how to assign the UEs to service entities. The problem is challenging for the following three aspects: 1) Service entity deployment and UE assignment are highly coupled. 2) Due to the overlap of coverage regions of densely deployed small cells, the allocation mechanism of different base stations has mutual effects on the overall service performance. 3) Considering the limited resources of edge clouds, it is a thorny problem to encourage edge clouds to cache and share service startup images. We devote the following efforts to tackle the problem under these challenges. First, we explore blockchain’s decentralized, traceable, and secure characteristics, and propose a scheme to encourage image sharing in mobile edge computing. Second, we formulate the service-oriented edge resource allocation as mixed integer non-linear programming. Third, towards the target of reducing the computational complexity, we decouple UE assignment from service entity deployment and solve it through Gibbs sampling. Moreover, the power of Lyapunov optimization and convex optimization is incorporated to reduce the long-term power consumption and budget. Experiment results demonstrate the superiority of our approach over current notable solutions. Ao Zhou 0001, Xiao Ma 0009, Shangguang Wang |
IEEE J. Sel. Areas Commun. | 1 |
| 2022 | A Cloud-Edge Collaboration Framework for Cognitive ServiceabstractMobile applications can leverage high-quality deep learning models such as convolutional neural networks and deep neural networks to provide high-performance cognitive services. Prior work on deep learning models-based mobile applications in a cloud-edge computing environment focuses on performing lightweight data pre-processing tasks on edge servers for cloud-hosted cognitive servers. These approaches have two major limitations. First, it is uneasy for the mobile applications to assure satisfactory user experience in terms of network communication delay, because the intermediary edge servers are used only to pre-process data (e.g., images and videos) and the cloud servers are used to complete the tasks. Second, these approaches assume the pre-trained deep learning models deployed on cloud servers are static, and will not attempt to automatically upgrade in a context-aware manner. In this article, we propose a cloud-edge collaboration framework that facilitates delivering cognitive services with long-lasting, fast response, and high accuracy properties. We fist deploy a shallow model (i.e., EdgeCNN) on the edge server and a deep model (i.e., CloudCNN) on the cloud server. EdgeCNN can provide durable and rapid response cognitive services, because edge servers not only provide computing resources for mobile applications, but also close to users. Then, we enable CloudCNN to assist in training EdgeCNN to improve the performance of the latter. Thus, EdgeCNN also provides high-accuracy cognitive services. Furthermore, because users may continue to upload data to edge servers in real-world scenarios, we propose to use the ongoing assistance of CloudCNN to further improve the accuracy of the shallow model. Experimental results show that EdgeCNN can reduce the average response time of cognitive services by up to 55.08 percent and improve accuracy by up to 26.70 percent. Chuntao Ding, Ao Zhou 0001, Yunxin Liu 0001, Rong Chang 0001, Ching-Hsien Hsu, Shangguang Wang |
IEEE Trans. Cloud Comput. | 2 |
| 2022 | Resource-Aware Feature Extraction in Mobile Edge ComputingabstractMobile image recognition services, which provide people with image recognition services through the cameras of mobile devices, are revolutionizing our lives. However, most existing cloud/edge-based approaches suffer from two major limitations, (i) Low recognition accuracy and high network bandwidth pressure, and (ii) Not easy to extract features based on currently available resources of mobile devices. In this paper, we propose a resource-aware feature extraction framework for mobile image recognition services. The proposed framework consists of discriminative feature extraction (DFE) and NestDFE algorithms. The DFE algorithm can generate an extractor${{\mathbf E}}$to extract discriminative features from the image data set on the edge server and images on mobile devices. Thus, the proposed framework can achieve higher recognition accuracy and require mobile devices to upload less feature data to the edge server. The NestDFE algorithm generates a single multi-capacity extractor that acts as a series of sub-extractors and enables mobile devices to dynamically select sub-extractors. Experimental results show that the proposed framework improves recognition accuracy by about 23 percent and reduces network traffic by about 76 percent compared with existing approaches. Chuntao Ding, Ao Zhou 0001, Xiulong Liu 0001, Xiao Ma 0009, Shangguang Wang |
IEEE Trans. Mob. Comput. | 2 |
| 2022 | QoS Driven Task Offloading With Statistical Guarantee in Mobile Edge ComputingabstractIn mobile edge computing, popular mobile applications, such as augmented reality, usually offload their tasks to resource-rich edge servers. The user experience can be considerably affected when many mobile users compete for the limited communication and computation resources. The key technical challenge in task offloading is to guarantee the Quality of Service (QoS) for such applications. Existing work on task offloading focus on deterministic QoS (delay) guarantee, which means that tasks have to complete before the given deadline with 100 percent. However, it is impractical to impose a deterministic QoS guarantee for tasks due to the high dynamics of the wireless environment when offloading to edge servers. In this paper, we focus on task offloading with statistical QoS guarantee (tasks are allowed to complete before a given deadline with a probability above the given threshold), which can further save more energy by loosing the QoS requirement. Specially, we first propose a statistical computation model and a statistical transmission model to quantify the correlation between the statistical QoS guarantee and task offloading strategy. Then, we formulate the task offloading problem as a mixed integer non-Linear programming problem with the statistical delay constraint. We transform the statistical delay constraint into the constraints on CPU cycle numbers and the delay exponent respectively. We propose an algorithm to provide the statistical QoS guarantee for tasks using convex optimization theory and Gibbs sampling method. Experiment results show that the proposed algorithm outperforms the three baselines. Qing Li 0028, Shangguang Wang, Ao Zhou 0001, Xiao Ma 0009, Fangchun Yang, Alex X. Liu |
IEEE Trans. Mob. Comput. | 3 |
| 2022 | Providing Reliable Service for Parked-vehicle-assisted Mobile Edge ComputingabstractNowadays, a growing number of computation-intensive applications appear in our daily life. Those applications make the loads of both the core network and the mobile devices, in terms of energy and bandwidth, hugely increase. Offloading computation-intensive tasks to edge cloud is proposed to address this issue. Since edge clouds have limited computation resources compared with the remote cloud, they would get over-loaded because of the heavy computation burden. Parked-vehicle-assisted mobile edge computing becomes one of the promising solutions for this problem. However, several critical issues in parked-vehicle-assisted mobile edge computing would result in low reliable edge service. The open environment would bring about uncertainty, and the data privacy is hard to ensure. In addition, different from edge cloud, each parked vehicle only has limited parking duration and can leave unexpectedly for personal reasons. Moreover, edge cloud and vehicle adopt different execution models of computation and communication. The heterogeneous environment may result in negative effect on cooperativeness. Ignoring those issues can result in substantial performance degradation. To tackle this challenge and explore the benefits of parked-vehicle-assisted offloading, we study the task offloading and resource-allocation problem by fully considering the above issues. First, we propose a resource-management scheme to address the privacy issue. Second, we review the execution model of computation and communication in parked-vehicle-assisted computation offloading. Then, we formulate the problem into a mixed-integer nonlinear programming. The problem is hard to tackle due to its non-convex nature, which means that the time complexity of finding global optimal solution is unaffordable. Finally, we decompose the original problem into two sub-problems with lower complexity, and related algorithms are given to deal with the sub-problems. Simulation results demonstrate the effectiveness of the proposed solution. Ao Zhou 0001, Xiao Ma 0009, Siyi Gao, Shangguang Wang |
ACM Trans. Internet Techn. | 1 |
| 2021 | Towards Green Service Composition Approach in the CloudabstractWith the popularization of the cloud, an increasing number of users and developers need to combine multiple different services to satisfy their requirements in the cloud. Consequently, the question of how to combine these services that are published by different service providers in the cloud has become an active focus of research in service-oriented cloud systems. Service composition is the core technology of service-oriented cloud systems. Shangguang Wang, Ao Zhou 0001, Ruo Bao, Chou Wu |
SERVICES | 2 |
| 2021 | Joint Placement of UPF and Edge Server for 6G NetworkabstractThe emerging 6G network will make it possible for cybertwin, which relies deeply on the low latency and powerful computation provided by the edge network. To this end, the convergence of computing and network has been attached great importance. Most existing work study either placing edge servers or deploying user plane functions (UPFs), seldom considers the two processes jointly. In this article, we study how to minimize the latency with cost limitation by means of jointly deploying edge servers and UPFs in 6G scenario. We have shown that the problem is NP-hard. Then, we simplify the problem by analyzing the placement relationship between edge servers and UPFs and prune the solution space of the problem. To solve the problem effectively, a UPF and edge server placement algorithm is proposed. Massive experiments are conducted based on real-world data set and an edge core network emulator. The evaluation results show that our algorithm outperforms the benchmark algorithms. Yuanzhe Li 0001, Xiao Ma 0009, Mengwei Xu 0001, Ao Zhou 0001, Qibo Sun, Ning Zhang 0007, Shangguang Wang |
IEEE Internet Things J. | 4 |
| 2021 | Freshness-Aware Information Update and Computation Offloading in Mobile-Edge ComputingabstractMobile-edge computing is a promising computing paradigm with the advantages of reduced delay and relieved outsourcing traffic to the core network. In mobile-edge computing, reducing the computation offloading cost of mobile users and maintaining fresh information at edge nodes are two critical while conflicted objectives, as both consume the limited wireless bandwidth of edge nodes. Although extensive efforts have been devoted to optimizing computation offloading decisions and some works have investigated freshness-aware channel allocation issues recently, no prior works have considered the above conflict. This article is the first work to jointly optimize the channel allocation and computation offloading decisions, aiming at reducing the computation offloading cost within freshness requirements of sensors. We analyze the recursiveness of Age of Information (AoI) in analogy to the evolvement of a queue and formulate the problem as a nonlinear integer dynamic optimization problem. To overcome the challenges of AoI-computation cost tradeoff, AoI time dependency and high complexity caused by the heterogeneity of users, we propose an algorithm to solve the problem with reduced computation complexity. Specifically, we first transform the original problem into a static optimization problem in each time slot (which is NP-hard) based on Lyapunov optimization techniques. To reduce the computation complexity, we exploit the finite improvement property of potential games and further enforce centralized control to reduce the number of improvement iterations. Simulations have been conducted and the results demonstrate that the proposed algorithm shows good effectiveness and scalability. Xiao Ma 0009, Ao Zhou 0001, Qibo Sun, Shangguang Wang |
IEEE Internet Things J. | 2 |
| 2021 | BCEdge: Blockchain-based resource management in D2D-assisted mobile edge computingabstractSummary In recent decades, newly emerging mobile applications and services are becoming increasingly resource hungry and computation intensive. Portable size mobile devices fall short in providing such services. Cloud computing has been a core computation technology to provide visualized resources in a scalable way for mobile services. However, the unpredictable network transmission latency makes cloud computing not efficient enough for time‐sensitive mobile services, which requires major changes in underlying computing platform. Mobile edge computing is widely known as one of the novel technologies has emerged in recent years to address this issue. For being impractical to construct huge edge cloud in network edge, edge clouds can still be overloaded in rush time. Offloading to near‐user facilities using device‐to‐device (D2D) or other technologies becomes an augmentation approach. However, how to manage the facilities/resources effectively become a new issue. In this paper, we proposed a resource management scheme named BCEdge based on blockchain in D2D‐assisted mobile edge computing. BCEdge is a reliable scheme that operates in a distributed way to relieve the load of edge clouds. We illustrate the advantages and technical details of BCEdge using flow charts and interaction charts. The experiment results validate the effectiveness of our scheme. Finally, we discuss the possible future extensions. Ao Zhou 0001, Qibo Sun |
Softw. Pract. Exp. | 1 |
| 2021 | A Cloud-Guided Feature Extraction Approach for Image Retrieval in Mobile Edge ComputingabstractMobile Edge Computing (MEC) can facilitate various important image retrieval applications for mobile users by offloading partial computation tasks from resource-limited mobile devices to edge servers. However, existing related works suffer from two major limitations. (i) High network bandwidth cost: they need to extract numerous features from the image and upload these feature data to the cloud server. (ii) Lowretrieval accuracy: they separate the feature extraction processes from the image data set in the cloud server, thus unable to provide effective features for accurate image retrieval. In this paper, we propose a cloud-guided feature extraction approach for mobile image retrieval. In the proposed approach, the cloud server first leverages the relationships among labeled images in the data set to learn a projection matrix P. Then, it uses the matrix P to extract discriminative features from the image data set and form a low-dimensional feature data set. Following that, the cloud server sends the matrix P to the edge server and uses it to multiply the image χ. The result PTχ, i.e., image features, is uploaded to the cloud server to find the label of the image with the most similar multiplying result. The label is regarded as the retrieval result and returned to the mobile user. In the cloud-guided feature extraction approach, the matrix P can extract a small number of effective image features, which not only reduces network traffic but also improves retrieval accuracy. We have implemented a prototype system to validate the proposed approach and evaluate its performance by conducting extensive experiments using a real MEC environment and data set. The experimental results show that the proposed approach reduces the network traffic by nearly 93 percent and improves the retrieval accuracy by nearly 6.9 percent compared with the state-of-the-art image retrieval approaches in MEC. Shangguang Wang, Chuntao Ding, Ning Zhang 0007, Xiulong Liu 0001, Ao Zhou 0001, Jiannong Cao 0001, Xuemin Shen |
IEEE Trans. Mob. Comput. | 5 |
| 2021 | Delay-Aware Microservice Coordination in Mobile Edge Computing: A Reinforcement Learning ApproachabstractAs an emerging service architecture, microservice enables decomposition of a monolithic web service into a set of independent lightweight services which can be executed independently. With mobile edge computing, microservices can be further deployed in edge clouds dynamically, launched quickly, and migrated across edge clouds easily, providing better services for users in proximity. However, the user mobility can result in frequent switch of nearby edge clouds, which increases the service delay when users move away from their serving edge clouds. To address this issue, this article investigates microservice coordination among edge clouds to enable seamless and real-time responses to service requests from mobile users. The objective of this work is to devise the optimal microservice coordination scheme which can reduce the overall service delay with low costs. To this end, we first propose a dynamic programming-based offline microservice coordination algorithm, that can achieve the globally optimal performance. However, the offline algorithm heavily relies on the availability of the prior information such as computation request arrivals, time-varying channel conditions and edge cloud's computation capabilities required, which is hard to be obtained. Therefore, we reformulate the microservice coordination problem using Markov decision process framework and then propose a reinforcement learning-based online microservice coordination algorithm to learn the optimal strategy. Theoretical analysis proves that the offline algorithm can find the optimal solution while the online algorithm can achieve near-optimal performance. Furthermore, based on two real-world datasets, i.e., the Telecom's base station dataset and Taxi Track dataset from Shanghai, experiments are conducted. The experimental results demonstrate that the proposed online algorithm outperforms existing algorithms in terms of service delay and migration costs, and the achieved performance is close to the optimal performance obtained by the offline algorithm. Shangguang Wang, Yan Guo 0004, Ning Zhang 0007, Peng Yang 0004, Ao Zhou 0001, Xuemin Shen |
IEEE Trans. Mob. Comput. | 5 |
| 2021 | Towards Green Service Composition Approach in the CloudabstractWith the increasing popularity of cloud computing, many notable quality of service (QoS)-aware service composition approaches have been incorporated in service-oriented cloud computing systems. However, these approaches are implemented without considering the energy and network resource consumption of the composite services. The increases in energy and network resource consumption resulting from these compositions can incur a high cost in data centers. In this paper, the trade-off among QoS performance, energy consumption, and network resource consumption in a service composition process is first analyzed. Then, a green service composition approach is proposed. It gives priority to those composite services that are hosted on the same virtual machine, physical server, or edge switch with end-to-end QoS guarantee. It fulfills the green service composition optimization by minimizing the energy and network resource consumption on physical servers and switches in cloud data centers. Experimental results indicate that, with comparisons to other approaches, our approach saves 20-50 percent of energy consumption and 10-50 percent of network resource consumption. Shangguang Wang, Ao Zhou 0001, Ruo Bao, Wu Chou, Stephen S. Yau |
IEEE Trans. Serv. Comput. | 2 |
| 2020 | Scheduling of Time Constrained Workflows in Mobile Edge Computing
Xican Chen, Siyi Gao, Qibo Sun, Ao Zhou 0001 |
BlockSys | 4 |
| 2020 | Adaptive Edge Resource Allocation for Maximizing the Number of Tasks Completed on Time: A Deep Q-Learning Approach
Qibo Sun, Ao Zhou 0001, Shangguang Wang, Tao Lei 0006 |
BlockSys | 3 |
| 2020 | Cognitive Service in Mobile Edge ComputingabstractCognitive services have revolutionized the way we live, work and interact with the world. In recent years, deep neural networks have become the mainstream approach in cognitive service, and mobile edge computing facilitates a variety of cognitive services for users by offloading computation tasks from resource-limited mobile devices to relatively wealthy edge servers. Combining the two to provide users with a higher quality of cognitive service is an issue worth researching. However, many related studies are not easy to provide fast responses because in these systems, edge servers are only used to pre-process data, and the cloud server is used to perform tasks. In this paper, we aim to study deploying deep neural network models on edge servers to provide fast services. However, a single edge server collects only a small amount of data, which results in low inference accuracy. To address this problem, we propose a cloud and edge collaboration framework. The key idea of the proposed framework is to use a cloud model to assist in training an edge model to improve the latter's inference accuracy and enable the latter to provide fast response and high-performance cognitive service. Experimental results demonstrate the effectiveness of our proposed framework. Chuntao Ding, Ao Zhou 0001, Xiao Ma 0009, Shangguang Wang |
ICWS | 2 |
| 2020 | Cooperative Service Caching and Workload Scheduling in Mobile Edge ComputingabstractMobile edge computing is beneficial for reducing service response time and core network traffic by pushing cloud functionalities to network edge. Equipped with storage and computation capacities, edge nodes can cache services of resource-intensive and delay-sensitive mobile applications and process the corresponding computation tasks without outsourcing to central clouds. However, the heterogeneity of edge resource capacities and mismatch of edge storage and computation capacities make it difficult to fully utilize both the storage and computation capacities in the absence of edge cooperation. To address this issue, we consider cooperation among edge nodes and investigate cooperative service caching and workload scheduling in mobile edge computing. This problem can be formulated as a mixed integer nonlinear programming problem, which has non-polynomial computation complexity. Addressing this problem faces challenges of sub-problem coupling, computation-communication tradeoff, and edge node heterogeneity. We develop an iterative algorithm named ICE to solve this problem. It is designed based on Gibbs sampling, which has provably near-optimal performance, and the idea of water filling, which has polynomial computation complexity. Simulation results demonstrate that our algorithm can jointly reduce the service response time and the outsourcing traffic, compared with the benchmark algorithms. Xiao Ma 0009, Ao Zhou 0001, Shan Zhang 0001, Shangguang Wang |
INFOCOM | 2 |
| 2020 | SLA-driven container consolidation with usage prediction for green cloud computing
Jialei Liu, Shangguang Wang, Ao Zhou 0001, Jinliang Xu, Fangchun Yang |
Frontiers Comput. Sci. | 3 |
| 2020 | Dependency-Aware Task Scheduling in Vehicular Edge ComputingabstractVehicular edge computing (VEC) offers a new paradigm to improve vehicular services and augment the capabilities of vehicles. In this article, we study the problem of task scheduling in VEC, where multiple computation-intensive vehicular applications can be offloaded to roadside units (RSUs) and each application can be further divided into multiple tasks with task dependency. The tasks can be scheduled to different mobile-edge computing servers on RSUs for execution to minimize the average completion time of multiple applications. Considering the completion time constraint of each application and the processing dependency of multiple tasks belonging to the same application, we formulate the multiple tasks scheduling problem as an optimization problem that is NP-hard. To solve the optimization problem, we develop an efficient task scheduling algorithm. The basic idea is to prioritize multiple applications and prioritize multiple tasks so as to guarantee the completion time constraints of applications and the processing dependency requirements of tasks. The numerical results demonstrate that our proposed algorithm can significantly reduce the average completion time of multiple applications compared with benchmark algorithms. Yujiong Liu, Shangguang Wang, Qinglin Zhao, Ao Zhou 0001, Xiao Ma 0009, Fangchun Yang |
IEEE Internet Things J. | 5 |
| 2020 | Path Selection for Seamless Service Migration in Vehicular Edge ComputingabstractMobile-edge computing provisions computing and storage resources by deploying edge servers (ESs) at the edge of the network to support ultralow delay and high bandwidth services. To ensure QoS of latency-sensitive services in vehicular networks, service migration is required to migrate data of the ongoing services to the closest ES seamlessly when users move across different ESs. To achieve seamless service migration, path selection is proposed to obtain one or more paths (consisting of several switches and ESs) to transfer service data. We focus on the following problems about path selection: 1) where to implement path selection? 2) how to coordinate interests of mobile users (i.e., vehicles) and network providers since they have conflicting interests during path selection? and 3) how to ensure seamless service migration during the migration of vehicles? To address the above problems, this article investigates path selection for seamless service migration. We propose a path-selection algorithm to jointly optimize both interests of the network plane (i.e., the cost for network providers) and service plane (i.e., QoE of users). We first formulate it as a multiobjective optimization problem and further prove theoretically that the proposed algorithm can give aweakly Pareto-optimal solution. Moreover, to improve the scalability of the proposed algorithm, a distance-based filter strategy is designed to eliminate undesired switches in advance. We conduct experiments on two synthesized data sets and the results validate the effectiveness of the proposed algorithm. Jinliang Xu, Xiao Ma 0009, Ao Zhou 0001, Qiang Duan 0002, Shangguang Wang |
IEEE Internet Things J. | 3 |
| 2020 | LMM: latency-aware micro-service mashup in mobile edge computing environment
Ao Zhou 0001, Shangguang Wang, Shaohua Wan 0001, Lianyong Qi |
Neural Comput. Appl. | 1 |
| 2020 | User allocation-aware edge cloud placement in mobile edge computingabstractSummary Mobile edge computing is emerging as a novel ubiquitous computing platform to overcome the limit resources of mobile devices and bandwidth bottleneck of the core network in mobile cloud computing. In mobile edge computing, it is a significant issue for cost reduction and QoS improvement to place edge clouds at the edge network as a small data center to serve users. In this paper, we study the edge cloud placement problem, which is to place the edge clouds at the candidate locations and allocate the mobile users to the edge clouds. Specifically, we formulate it as a multiobjective optimization problem with objective to balance the workload between edge clouds and minimize the service communication delay of mobile users. To this end, we propose an approximate approach that adopted the K‐means and mixed‐integer quadratic programming. Furthermore, we conduct experiments based on Shanghai Telecom's base station data set and compare our approach with other representative approaches. The results show that our approach performs better to some extent in terms of workload balance and communication delay and validate the proposed approach. Yan Guo 0004, Shangguang Wang, Ao Zhou 0001, Jinliang Xu, Ching-Hsien Hsu |
Softw. Pract. Exp. | 3 |
| 2020 | Towards Network-Aware Service Composition in the CloudabstractComposing several API-defined services into one composite service per user requirements has become an important service creation approach in the cloud-enabled API economy. Various service selection approaches in support of service composition on demand have been proposed. They usually assume that networking resources are over-provisioned and their usage needs not be considered when making quality-aware service composition decisions. In practice, these approaches often lead to wasteful network resource consumption and impractical end-to-end QoS optimality for cloud-based services. This paper proposes a network-aware cloud service composition approach, named NetMIP, with comparative experimental evaluations for the clouds that adopt the widely deployed fat-tree network topology. By formalizing the service composition goal as a multi-objective constraint optimization problem, we have validated the proposed approach can be used to effectively reduce network resource consumption and deliver QoS optimality while satisfying the end-to-end QoS constraints for the candidate composite services in the cloud. The comparative experimental evaluations are done via a credible cloud infrastructure simulation system, named WebCloudSim. Extensive evaluation results show that NetMIP outperforms several representative cloud service composition approaches in terms of network resource consumption, QoS optimality, and computation time under various service selection workloads and fat-tree network topology settings. Shangguang Wang, Ao Zhou 0001, Fangchun Yang, Rong Chang 0001 |
IEEE Trans. Cloud Comput. | 2 |
| 2020 | Availability-Aware Virtual Cluster Allocation in Bandwidth-Constrained DatacentersabstractAs greater numbers of data-intensive applications are required to process big data in bandwidth-constrained datacenters with heterogeneous physical machines (PMs) and virtual machines (VMs), network core traffic is experiencing rapid growth. The VMs of a virtual cluster (VC) must be allocated as compactly as possible to avoid bandwidth-related bottlenecks. Since each PM/switch has a certain failure probability, a VC may not be executed when it meets with any PM/switch fault. Although the VMs of a VC can be spread out across different fault domains to minimize the risk of violating the availability requirement of the VC, this increases the network core traffic. Therefore, avoiding the decrease in availability caused by the heterogeneous PM/switch failure probabilities and bandwidth-related bottlenecks has been a constant challenge. In this paper, we first introduce a joint optimization function to measure the overall risk cost and overall bandwidth usage in the network core to allocate the same set of data-intensive applications. We then introduce an approach to maximize the value of the joint optimization function. Finally, we performed a side-by-side comparison with prior algorithms, and the experimental results show that our approach outperforms the other existing algorithms. Jialei Liu, Shangguang Wang, Ao Zhou 0001, Rajkumar Buyya, Fangchun Yang |
IEEE Trans. Serv. Comput. | 3 |
| 2020 | Towards Service Composition Aware Virtual Machine Migration Approach in the CloudabstractThere is a growing trend for service providers to migrate their services from local clusters to the cloud data center. When there is no single service can satisfy the functionality requirement of the end user, existing services are combined together to fulfill the requirements. The data communication between component service hosting servers imposes a heavy burden on the data center network. In this article, we seek to reduce the data center network resource consumption by designing a novel service composition aware virtual machine migration approach. First, we formulate the problem as a multi-object integer non-linear(INLP) programming problem. The problem, which can be reduced into a well-known multi-object quadratic assignment problem, is proved to be NP-hard. Second, we simplify the multiple-objects INLP formulation into an equivalent, but much simplified single object ILP formulation. Then, we prove that the simplified formulation can also lead to the optimal solutions. Finally, optimization problem solvers, such as LPSolver, are employed to solve the problem. Experimental results in a large scale cloud data center demonstrate that our method significantly reduce the network resource consumption than other approaches. Ao Zhou 0001, Shangguang Wang, Xiao Ma 0009, Stephen S. Yau |
IEEE Trans. Serv. Comput. | 1 |
| 2019 | Redundant Virtual Machine Placement in Mobile Edge Computing
Siyi Gao, Ao Zhou 0001, Xican Chen, Qibo Sun |
BlockSys | 2 |
| 2019 | Predicting Fine-Grained Traffic Conditions via Spatio-Temporal LSTMabstractPredicting traffic conditions for road segments is the prelude of working on intelligent transportation. Many existing methods can be used for short-term or long-term traffic prediction, but they focus more on regions than on road segments. The lack of fine-grained traffic predicting approach hinders the development of ITS. Therefore, MapLSTM, a spatio-temporal long short-term memory network preluded by map-matching, is proposed in this paper to predict fine-grained traffic conditions. MapLSTM first obtains the historical and real-time traffic conditions of road segments via map-matching. Then LSTM is used to predict the conditions of the corresponding road segments in the future. Breaking the single-index forecasting, MapLSTM can predict the vehicle speed, traffic volume, and the travel time in different directions of road segments simultaneously. Experiments confirmed MapLSTM can not only achieve prediction for road segments based a large scale of GPS trajectories effectively but also have higher predicting accuracy than GPR and ConvLSTM. Moreover, we demonstrate that MapLSTM can serve various applications in a lightweight way, such as cognizing driving preferences, learning navigation, and inferring traffic emissions. Xiaojuan Wei, Quan Yuan 0004, Kaihui Chen, Ao Zhou 0001, Fangchun Yang |
Wirel. Commun. Mob. Comput. | 5 |
| 2018 | The Performance Evaluation of Virtual Machine Placement Algorithm Based on WebCloudSimabstractEffective virtual machine placement algorithms can improve network resource utilization in cloud data centers. In order to design a more effective VM placement algorithm, we usually rely on a large-scale experimental platform to evaluate and verify its performance. However, large-scale experiment may have a huge impact on data center network, which makes it very unlikely to happen. Therefore, researchers need an experiment platform that can support large-scale environment to meet the above requirements. In order to evaluate the performance of virtual machine deployment algorithm, this paper first designs and implements a cloud data center network experiment system, WebCloudSim, which can support the joint experimental verification of real environment and simulation environment. Then, based on WebCloudSim, the paper implements three classic algorithms for virtual machine deployment, and verifies and analyzes the results. The experimental results show that WebCloudSim can effectively support the performance evaluation of virtual machine placement algorithm. Songtai Dai, Ao Zhou 0001, Shangguang Wang |
IEEE CLOUD | 2 |
| 2018 | FMSR: A Fairness-Aware Mobile Service Recommendation MethodabstractWith the development of mobile Internet, mobile service is emerging one after another, and the problem of information overload is becoming ever more serious. As an important tool to alleviate information overload, mobile service recommendation has attracted more and more attention. However, traditional recommendation algorithms always recommend popular services to users, which result into a rich-get-richer problem and become a barrier for the unpopular services to startup and growth. In order to promote the healthy development of the service ecosystem, it is necessary to guarantee the fairness of unpopular services. To address this problem, this paper proposes a fairness-aware mobile service recommendation method (FMSR), which gives a relatively fair recommendation opportunity for unpopular services. FMSR makes a tradeoff between recommendation accuracy and fairness, and can recommend popular services and unpopular services respectively. For unpopular services, we design a fair efficiency function and use combinatorial optimization techniques to achieve recommendations. For popular services, bias matrix factorization is utilized to implement recommendations. Experimental results based on real-world demonstrate that FMSR significantly improve the fairness of mobile service recommendation in the evolving mobile service ecosystem. Qiliang Zhu, Ao Zhou 0001, Qibo Sun, Shangguang Wang, Fangchun Yang |
ICWS | 2 |
| 2018 | Towards Bandwidth Guaranteed Virtual Cluster Reallocation in the CloudabstractCloud data center traffic is experiencing a rapid growth as more and more data-intensive applications are required to process big data in a cloud data center. Although fat-tree networks own rich path multiplicity, and have been widely adopted as network topologies in cloud data center networks to transmit vast bisection bandwidth, they lead to bandwidth-related bottlenecks. To address this issue, in this paper, we propose a traffic-aware virtual cluster reallocation approach via biogeography-based optimization to allocate some reallocated virtual machines (VMs) as compact as possible with those VMs in the same virtual clusters. In order to validate our approach, we build a system model to perform a thorough evaluation of its performance. Experimental results show that our proposed approach outperforms six existed approaches in term of total transmission cost, total processing time, and total network resource consumption. Jialei Liu, Shangguang Wang, Ao Zhou 0001, Sathish A. P. Kumar, Fangchun Yang |
Comput. J. | 3 |
| 2018 | Using Proactive Fault-Tolerance Approach to Enhance Cloud Service ReliabilityabstractThe large-scale utilization of cloud computing services for hosting industrial/enterprise applications has led to the emergence of cloud service reliability as an important issue for both cloud service providers and users. To enhance cloud service reliability, two types of fault tolerance schemes, reactive and proactive, have been proposed. Existing schemes rarely consider the problem of coordination among multiple virtual machines (VMs) that jointly complete a parallel application. Without VM coordination, the parallel application execution results will be incorrect. To overcome this problem, we first propose an initial virtual cluster allocation algorithm according to the VM characteristics to reduce the total network resource consumption and total energy consumption in the data center. Then, we model CPU temperature to anticipate a deteriorating physical machine (PM). We migrate VMs from a detected deteriorating PM to some optimal PMs. Finally, the selection of the optimal target PMs is modeled as an optimization problem that is solved using an improved particle swarm optimization algorithm. We evaluate our approach against five related approaches in terms of the overall transmission overhead, overall network resource consumption, and total execution time while executing a set of parallel applications. Experimental results demonstrate the efficiency and effectiveness of our approach. Jialei Liu, Shangguang Wang, Ao Zhou 0001, Sathish A. P. Kumar, Fangchun Yang, Rajkumar Buyya |
IEEE Trans. Cloud Comput. | 3 |
| 2018 | Service Migration in Mobile Edge Computing
Shangguang Wang, Wu Chou, Kok-Seng Wong, Ao Zhou 0001, Victor C. M. Leung |
Wirel. Commun. Mob. Comput. | 4 |
| 2018 | AIMING: Resource Allocation with Latency Awareness for Federated-Cloud ApplicationsabstractFederated‐cloud has been widely deployed due to the growing popularity of real‐time applications, and hence allocating resources among clouds becomes nontrivial to meet the stringent service requirements. The challenges lie in achieving minimized latency constrained by virtual machines rental overhead and resource requirement. This becomes further complicated by the issues of datacenter selection. To this end, we propose AIMING, a novel resource allocation approach which aims to minimize the latency constrained by monetary overhead in the context of federated‐cloud. Specifically, the network resources are deployed and selected according to k‐means clustering. Meanwhile, the total latency among datacenters is optimized based on binary quadratic programming. The evaluation is conducted with real data traces. The results show that AIMING can reduce total datacenter latency effectively compared with other approaches. Ao Zhou 0001, Fangchun Yang |
Wirel. Commun. Mob. Comput. | 2 |
| 2018 | Overview on Fault Tolerance Strategies of Composite Service in Service ComputingabstractIn order to build highly reliable composite service via Service Oriented Architecture (SOA) in the Mobile Fog Computing environment, various fault tolerance strategies have been widely studied and got notable achievements. In this paper, we provide a comprehensive overview of key fault tolerance strategies. Firstly, fault tolerance strategies are categorized into static and dynamic fault tolerance according to the phase of their adoption. Secondly, we review various static fault tolerance strategies. Then, dynamic fault tolerance implementation mechanisms are analyzed. Finally, main challenges confronted by fault tolerance for composite service are reviewed. Junna Zhang, Ao Zhou 0001, Qibo Sun, Shangguang Wang, Fangchun Yang |
Wirel. Commun. Mob. Comput. | 2 |
| 2017 | CMIP: Data Transmission Latency Optimization for Cooperative Group in Multi-cloud by Adaptive RoutingabstractThe proliferation of real-time applications, such as online gaming, video chatting, brings unprecedented pressure for multi-cloud application providers to secure the good quality of experience. These applications are designed for cooperative group users/members/clients and are expected to have low data transmission latency due to frequent interactions. We use total data transmission latency (TDTL) to represent the group latency in this paper. However, existing approaches perform poorly since they often ignore the feature of the cooperative group scenario and cannot provide a flexible rental cost model for application providers. Minimizing TDTL is challenging due to the difficulties in 1) Finding an effective method to transmit data through multi-cloud datacenters, 2) Taking the cooperative group scenario into consideration, 3) Providing a flexible transmission rental cost model. To this end, we propose CMIP, an innovative approach that aims to minimize the TDTL by renting virtual machines of well-connected datacenters as the proxy for the cooperative group scenario in multi-cloud providers. First, a weighted graph is constructed to illustrate the multi-cloud network. Second, we define a path latency model and a rental cost model for the cooperative group. After that, Yen's algorithm and mixed integer programming are adopted to optimize the TDTL with rental cost constraint. To study the performance of CMIP, we conduct extensive evaluations using multiple real data traces and compare it with related approaches. The comprehensive evaluation analysis shows that the CMIP achieves much lower total latency and is more flexible than other approaches. Shangguang Wang, Ao Zhou 0001, Fangchun Yang |
ICPADS | 3 |
| 2017 | Network failure-aware redundant virtual machine placement in a cloud data centerabstractSummary Cloud has become a very popular infrastructure for many smart city applications. A growing number of smart city applications from all over the world are deployed on the clouds. However, node failure events from the cloud data center have negative impact on the performance of smart city applications. Survivable virtual machine placement has been proposed by the researchers to enhance the service reliability. Because of the ignorance of switch failure, current survivable virtual machine placement approaches cannot achieve the best effect. In this paper, we study to enhance the service reliability by designing a novel network failure–aware redundant virtual machine placement approach in a cloud data center. Firstly, we formulate the network failure–aware redundant virtual machine placement problem as an integer nonlinear programming problem and prove that the problem is NP‐hard. Secondly, we propose a heuristic algorithm to solve the problem. Finally, extensive simulation results show the effectiveness of our algorithm. Ao Zhou 0001, Shangguang Wang, Ching-Hsien Hsu, Kok-Seng Wong |
Concurr. Comput. Pract. Exp. | 1 |
| 2017 | Cloud Service Reliability Enhancement via Virtual Machine Placement OptimizationabstractWith rapid adoption of the cloud computing model, many enterprises have begun deploying cloud-based services. Failures of virtual machines (VMs) in clouds have caused serious quality assurance issues for those services. VM replication is a commonly used technique for enhancing the reliability of cloud services. However, when determining the VM redundancy strategy for a specific service, many state-of-the-art methods ignore the huge network resource consumption issue that could be experienced when the service is in failure recovery mode. This paper proposes a redundant VM placement optimization approach to enhancing the reliability of cloud services. The approach employs three algorithms. The first algorithm selects an appropriate set of VM-hosting servers from a potentially large set of candidate host servers based upon the network topology. The second algorithm determines an optimal strategy to place the primary and backup VMs on the selected host servers with k-fault-tolerance assurance. Lastly, a heuristic is used to address the task-to-VM reassignment optimization problem, which is formulated as finding a maximum weight matching in bipartite graphs. The evaluation results show that the proposed approach outperforms four other representative methods in network resource consumption in the service recovery stage. Ao Zhou 0001, Shangguang Wang, Bo Cheng 0001, Zibin Zheng, Fangchun Yang, Rong Chang 0001, Michael R. Lyu, Rajkumar Buyya |
IEEE Trans. Serv. Comput. | 1 |
| 2016 | Machine Status Prediction for Dynamic and Heterogenous Cloud EnvironmentabstractThe widespread utilization of cloud computing services has brought in the emergence of cloud service reliability as an important issue for both cloud providers and users. To enhance cloud service reliability and reduce the subsequent losses, the future status of virtual machines should be monitored in real time and predicted before they crash. However, most existing methods ignore the following two characteristics of actual cloud environment, and will result in bad performance of status prediction: 1. cloud environment is dynamically changing, 2. cloud environment consists of many heterogeneous physical and virtual machines. In this paper, we investigate the predictive power of collected data from cloud environment, and propose a simple yet general machine learning model StaP to predict multiple machine status. We introduce the motivation, the model development and optimization of the proposed StaP. The experimental results validated the effectiveness of the proposed StaP. Jinliang Xu, Ao Zhou 0001, Shangguang Wang, Qibo Sun, Fangchun Yang |
CLUSTER | 2 |
| 2016 | A Novel Trust Update Mechanism Based on Sliding Window for Trust Management System
Qibo Sun, Ao Zhou 0001 |
ICCSA (1) | 3 |
| 2016 | Task rescheduling optimization to minimize network resource consumption
Ao Zhou 0001, Shangguang Wang, Ching-Hsien Hsu, Qibo Sun, Fangchun Yang |
Multim. Tools Appl. | 1 |
| 2016 | On Cloud Service Reliability Enhancement with Optimal Resource UsageabstractAn increasing number of companies are beginning to deploy services/applications in the cloud computing environment. Enhancing the reliability of cloud service has become a critical and challenging research problem. In the cloud computing environment, all resources are commercialized. Therefore, a reliability enhancement approach should not consume too much resource. However, existing approaches cannot achieve the optimal effect because of checkpoint image-sharing neglect, and checkpoint image inaccessibility caused by node crashing. To address this problem, we propose a cloud service reliability enhancement approach for minimizing network and storage resource usage in a cloud data center. In our proposed approach, the identical parts of all virtual machines that provide the same service are checkpointed once as the service checkpoint image, which can be shared by those virtual machines to reduce the storage resource consumption. Then, the remaining checkpoint images only save the modified page. To persistently store the checkpoint image, the checkpoint image storage problem is modeled as an optimization problem. Finally, we present an efficient heuristic algorithm to solve the problem. The algorithm exploits the data center network architecture characteristics and the node failure predicator to minimize network resource usage. To verify the effectiveness of the proposed approach, we extend the renowned cloud simulator Cloudsim and conduct experiments on it. Experimental results based on the extended Cloudsim show that the proposed approach not only guarantees cloud service reliability, but also consumes fewer network and storage resources than other approaches. Ao Zhou 0001, Shangguang Wang, Zibin Zheng, Ching-Hsien Hsu, Michael R. Lyu, Fangchun Yang |
IEEE Trans. Cloud Comput. | 1 |
| 2016 | Optimal mobile device selection for mobile cloud service providing
Ao Zhou 0001, Shangguang Wang, Qibo Sun, Fangchun Yang |
J. Supercomput. | 1 |
| 2016 | Enhanced User Context-Aware Reputation Measurement of Multimedia ServiceabstractReputation plays an important role for users in choosing or paying for multimedia applications or services. Some efficient multimedia reputation-measurement approaches have been proposed to achieve accurate reputation measurement based on feedback ratings that users give to a multimedia service after invoking. However, the implementation of these approaches suffers from the problems of wide abuse and low utilization of user context. In this article, we study the relationship between user context and feedback ratings according to which one user often gives different feedback ratings to the same multimedia service in different user contexts. We further propose an enhanced user context-aware reputation-measurement approach for multimedia services that is accurate in two senses: (1) Each multimedia service has three reputation values with three different user context levels when its feedback ratings are sufficient and (2) the reputation of a multimedia service with different user context levels is found using user context sensitivity and user similarity when its feedback ratings are limited or not available. Experimental results based on a real-world dataset show that our approach outperforms other approaches in terms of accuracy. Shangguang Wang, Ao Zhou 0001, Wei Lei, Zhiwen Yu 0001, Ching-Hsien Hsu, Fangchun Yang |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2015 | Minimizing Data Transmission Latency by Bipartite Graph in MapReduceabstractMany factors affect the time cost of Cloud computing tasks. One of the most serious factors is data transmission latency, which reduces the efficiency of Cloud computing. Existing notable schemes ignore the communication cost among virtual machines (VMs) in the MapReduce environment. In this paper, we propose a VM placement approach to reduce data transmission latency with the communication cost among VMs. We first construct bipartite graph and classify VMs as two groups according to their transmission latency with data nodes. Then we propose two VM placement optimization algorithms to minimize the total data transmission latency (TDTL) and the maximum data transmission latency (MDTL) in the MapReduce environment. Finally, we place VMs for Reduce phase. The evaluation results show that our approach reduces the average data transmission latency by 26.3% compared with other approaches. Shangguang Wang, Ao Zhou 0001, Qibo Sun, Ruisheng Shi, Fangchun Yang |
CLUSTER | 4 |
| 2015 | An online cloud data center simulation systemabstractCurrently, many researchers play more attention on cloud data center network and hope to verify their achievements. However, existing cloud data center simulation tools are difficult to provide an online cloud data center simulation. In this paper, we design an online cloud data center simulation system based on CloudSim and some extensions. This system support Browser/Server mode and some simple configuration can simplify the simulation process. Ao Zhou 0001, Shangguang Wang |
IWQoS | 3 |
| 2015 | PFT-CCKP: A proactive fault tolerance mechanism for data center networkabstractExiting schemes rarely take into account coordinated problem among multiple virtual machines (VMs) which collectively complete a task, since if these VMs are uncoordinated, the execution results of a task are not correct. To solve the problem, we first exploit a proactive prediction scheme to predict the VM's status. Then, if VM's status is deteriorating, coordinated checkpoint is adopted to suspend the current task to search an optimal target host. Finally, we introduce an efficient heuristic algorithm to solve the optimal target host selection problem. The experimental results demonstrate the efficiency and effectiveness of our proposed approach. Jialei Liu, Shangguang Wang, Ao Zhou 0001, Fangchun Yang |
IWQoS | 3 |
| 2013 | A Dynamic Virtual Resource Renting Method for Maximizing the Profit of Cloud Service Provider under SLA ConstraintabstractTo maximize the profit of cloud service provider, we present a dynamic virtual resource renting method under SLA (service level agreement) constraints. According to the price distribution and current task emergency, the method attempts to adjust acceptable price of each virtual resource type at different price interval. If there is price acceptable resource, we choose to rent the most profitable one. Otherwise, if SLA permits, we even suspend the task and restart it when price falls. Partial experimental result in simulation environment is also presented. Ao Zhou 0001, Shangguang Wang, Qibo Sun, Hua Zou 0001, Fangchun Yang |
IEEE CLOUD | 1 |
| 2013 | Dynamic Virtual Resource Renting Method for Maximizing the Profits of a Cloud Service Provider in a Dynamic Pricing ModelabstractWith an increasing number of cloud service providers (CSP) delivering services to customers from the cloud, maximizing the profits of CSPs becomes a critical problem. Existing methods are difficult to solve the problem because they do not make full use of temporal price differences. This paper introduces a dynamic virtual resource renting method that attempts to dynamically adjust the virtual resource rental strategy according to price distribution and task urgency. We first pretreat the historical price series and adopt the outlier detection technique to filter the extreme price. Then, considering task urgency and price distribution, we design a weak equilibrium operator to calculate the acceptable price for each type of virtual resource. All types of virtual resources that are at an acceptable price are inserted into a set. Finally, we design a novel rental decision-making algorithm to select the most profitable resource from the set. We provide an extensive evaluation of our method using Amazon EC2 spot price dataset and normally distributed price dataset. The results demonstrate the effectiveness of our method. Ao Zhou 0001, Shangguang Wang, Qibo Sun, Hua Zou 0001, Fangchun Yang |
ICPADS | 1 |