VLDB 2026 Research / reviewers in the wild / expert
Xiao Ma 0009
dblp:35/573-9
· DBLP profile ↗
71ranked-venue papers
7as first author
55since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 43 · 6 first-author · 34 since 2021Software engineering, systems software and programming languages · 13 · 10 since 2021Systems, architecture and hardware · 6 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 6 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 4 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Temperature- and Energy-Aware Dynamic Task Scheduling and Computing Resource Allocation for Satellite ComputingabstractSatellite computing, as an emerging edge computing paradigm, extends computing and networking services into space. Due to the internal design constraints of low-Earth orbit (LEO) satellites and the challenges posed by the external environment, satellite computing faces inherent limitations, including severely constrained resources, non-rechargeable batteries, poor heat dissipation, and highly dynamic operating conditions, leading to unreliable and unsustainable quality of service. To address the above challenges and fully realize the potential of satellite computing, this paper investigates temperature- and energy-aware dynamic task scheduling and computing resource allocation, aiming to optimize service latency, reduce onboard energy consumption, and enhance operational profit. Solving this problem requires coordinating task scheduling and resource allocation, balancing communication and computation latency, and addressing the challenge of a vast search space. To solve the above challenges, we first formulate this problem as a repeated Stackelberg game by developing temperature and energy models. Through theoretical analysis, we show that this game leads to a convex optimization framework that exhibits exponential complexity. To accelerate the search for the Stackelberg equilibrium solution, we propose a dynamic task scheduling algorithm based on the interior point method, which reduces the computational complexity to polynomial order. Trace-driven simulations demonstrate that the proposed algorithm reduces task scheduling latency by 28.4% and improves utility by 13% on average. Chao Wang 0093, Xiao Ma 0009, Chuanxiu Chi, Ao Zhou 0001, Ruolin Xing, Shangguang Wang |
IEEE Trans. Mob. Comput. | 2 |
| 2026 | Prototyping and Analyzing Mobile SoC Clusters as Modern Edge ServersabstractThe rapidly growing edge computing platforms, coupled with the still imperfect edge infrastructure, present an excellent opportunity for the emergence of new edge hardware. However, it remains unclear whether alternative architectures built from energy-efficient mobile System-on-Chips (SoCs) can meet the stringent performance, cost, and energy demands of modern edge workloads. In this paper, we propose a new type of edge server composed of 60 Qualcomm Snapdragon 865 mobile SoCs in a 2U rack, referred to as SoC Cluster. We demonstrate its successful deployment on existing edge cloud platforms and its ability to natively serve mobile cloud gaming services. Despite the emergence of new hardware on edge platforms and its successful operation in serving mobile cloud gaming, our trace analysis revealed low hardware utilization and significant dynamic fluctuations in usage. To assess its broader applicability, we conducted the first measurement study of SoC Cluster to reveal its ability to run two popular and modern edge applications: deep learning inference and video transcoding. We developed a cross-platform benchmark suite to evaluate throughput, latency, power consumption, and application-specific metrics like video quality. We then directly compare SoC Cluster with a traditional edge server equipped with Intel CPUs and NVIDIA GPUs in terms of energy efficiency, space efficiency, and monetary cost. Results show that SoC Cluster exhibits up to 6.5? higher energy efficiency and 7.7? higher space efficiency. We also disclose its limitations in serving computation-intensive workloads such as large deep learning models. The outcomes provide insightful implications and offer practical direction for refining SoC Cluster toward broader deployment in edge scenarios. Li Zhang 0133, Boqing Shi, Xiang Li 0067, Ao Zhou 0001, Xiao Ma 0009, Shangguang Wang, Mengwei Xu 0001 |
IEEE Trans. Mob. Comput. | 5 |
| 2026 | Exploring Image Similarity to Optimize Resource Provisioning in Container-Enabled Edge ComputingabstractContainerization offers great flexibility and agility for resource provisioning in edge clouds. However, this benefit is not freely available, as substantial network traffic incurred by container image pulling heavily burdens the back-haul networks. Our in-depth measurements on 516.3 GB images from Docker Hub reveal that many similar images share identical layers, with only 11.21% having no shared layers. Building on this, we investigate the resource provisioning problem by leveraging image similarity to avoid repeated layer transmissions, aiming to reduce network traffic and improve overall performance. We formulate resource provisioning as a mixed-integer non-linear programming problem, which is challenging due to the coupling of four issues and their conflicting effects on overall performance, including offloading decisions, container instance deployment, image pulling, and resource allocation. To tackle these complexities, we propose a novel Similarity-Aware Resource Provisioning approach, which decomposes the problem into independent sub-problems using counterfactual multi-agent deep reinforcement learning and then solves sub-problems individually with convex optimization and fractional programming techniques. We conduct extensive evaluations with images from Docker Hub. The results show that our approach enables up to 32.6% traffic reduction and 19.9% utility improvement, outperforming the state-of-the-art solutions. Ao Zhou 0001, Xiao Ma 0009, Jinfeng Wen, Shangguang Wang |
IEEE Trans. Serv. Comput. | 3 |
| 2025 | Towards Robust Trajectory Embedding for Similarity Computation: When Triangle Inequality Violations in Distance Metrics MatterabstractTrajectory similarity is a cornerstone of trajectory data management and analysis. Traditional similarity functions often suffer from high computational complexity and a reliance on specific distance metrics, prompting a shift towards deep representation learning in Euclidean space. However, existing Euclidean-based trajectory embeddings often face challenges due to the triangle inequality constraints that do not universally hold for trajectory data. To address this issue, this paper introduces a novel approach by incorporating non-Euclidean geometry, specifically hyperbolic space, into trajectory representation learning. We present the first-ever integration of hyperbolic space to resolve the inherent limitations of the triangle inequality in Euclidean embeddings. In particular, we achieve it by designing a Lorentz distance measure, which is proven to overcome triangle inequality constraints. Additionally, we design a model-agnostic framework LH-plugin to seamlessly integrate hyperbolic embeddings into existing representation learning pipelines. This includes a novel projection method optimized with the Cosh function to prevent the diminishment of distances, supported by a theoretical foundation. Furthermore, we propose a dynamic fusion distance that intelligently adapts to variations in triangle inequality constraints across different trajectory pairs, blending Lorentzian and Euclidean distances for more robust similarity calculations. Comprehensive experimental evaluations demonstrate that our approach effectively enhances the accuracy of trajectory similarity measures in state-of-the-art models across multiple real-world datasets. The LH-plugin not only addresses the triangle inequality issues but also significantly refines the precision of trajectory similarity computations, marking a substantial advancement in the field of trajectory representation learning. Jianing Si, Haitao Yuan 0002, Minxiao Chen, Xiao Ma 0009, Shangguang Wang |
ICDE | 5 |
| 2025 | Having It Both Ways: Single Trajectory Embedding for Similarity Computation with Pairwise LearningabstractTrajectory similarity measure is a fundamental component in trajectory databases, supporting many down-stream trajectory tasks. Existing similarity functions often exhibit unacceptable time complexities, hampering their efficiency for real-world scenarios. To address this limitation, learning-based approximation techniques utilizing trajectory embeddings have been proposed. However, creating a robust embedding model presents challenges, including the lack of direct involvement in the computational similarity process, adherence to non-metric similarity spaces, and the integration of precise similarity computation alignments. To address these challenges, we introduce DTisT, a novel embedding framework that enhances trajectory embeddings by pairwise learning from dual-trajectory input models. DTisT not only captures the dynamics of trajectory similarity computation through a dual-trajectory learning model but also integrates a learnable virtual trajectory to align the embedding space with non-metric similarity spaces effectively. Additionally, we incorporate aligned information from actual similarity computations into our embedding process using an attention mask mechanism. To ensure effective learning, we adopt a pre-train and fine-tune strategy, utilizing contrastive learning during the pre-training stage. Extensive experiments conducted on two real datasets demonstrate that DTisT surpasses state-of-the-art methods, showcasing its effectiveness in trajectory similarity embedding. Jianing Si, Haitao Yuan 0002, Xiang Li 0067, Xiao Ma 0009, Guoliang Li 0001, Shangguang Wang |
ICDE | 5 |
| 2025 | Reliability-Aware Resource Allocation for Vehicular Services in Mobile Edge ComputingabstractMobile edge computing is envisioned as a promising paradigm to overcome the computational limitations of vehicles. However, current resource allocation optimizations for vehicular services in mobile edge computing are typically performed under quasi-static scenarios. Due to the time-varying wireless communication environments between vehicles and the base station, the communication uncertainty becomes a key factor to hinder the service reliability. To tackle this problem, the reliability-aware mobile edge resource allocation for vehicular services is explored in this paper. Firstly, the model of computation, communication, and reliability is reviewed integratively, and the problem is then formulated as a fractional programming model. Secondly, conditional value-at-risk (CVaR) is adopted to transform the original problem into a solvable semidefinite programming model. Finally, we propose an efficient two-layer reliability-aware resource allocation approach that decomposes the original problem into two sub-problems with lower complexity. Lagrange multiplier and Dinkelbach strategy are then applied for these sub-problems, which minimize average energy consumption while ensuring reliability for vehicles. Simulation results examine the performance of our proposed approach and demonstrate its superiority over other approaches. Lipei Yang, Ao Zhou 0001, Xiao Ma 0009, Shangguang Wang |
ICWS | 4 |
| 2025 | Rethinking Cost-Efficient VM Scheduling on Public Edge Platforms: A Service Provider's PerspectiveabstractPrior studies on traditional centralized clouds independently optimizing static resource utilization and dynamic bandwidth cost are not applicable to edge scenarios, where edge sites are interconnected by wide area networks (WAN) rather than local area networks (LAN) as within clouds. Due to the lack of knowledge about the actual status of public edge platforms and real-world edge datasets, existing influential literature on edge scenarios demonstrates significant disparities in optimization objectives and perspectives. To bridge this gap, we collaborate with a public edge platform and perform a comprehensive measurement, which reveals limitations of the status quo VM scheduling schemes and potential opportunities for improvement. However, resolving VM scheduling considering static resource utilization, dynamic bandwidth cost, and end users’ QoE in a cost-efficient manner faces several challenges, including coupled objectives, exponentially increased complexity, and spatiotemporal dynamics. To address the above challenges, in this work, we propose a holistic online framework that integrates combinatorial bandit-based VM migration and seasonality-aware VM request allocation at two distinct time granularities. Large-scale experiments based on a real-world dataset confirm that our online framework achieves near-offline bandwidth cost and resource utilization while significantly lowering time consumption. Xiao Ma 0009, Zhe Fu 0005, Ao Zhou 0001, Mengwei Xu 0001, Shangguang Wang |
IEEE Trans. Mob. Comput. | 2 |
| 2025 | A Collaborative Cloud-Edge Approach for Robust Edge Workload ForecastingabstractWith the rapid development of edge computing in the post-COVID19 pandemic period, precise workload forecasting is considered the basis for making full use of the edge-limited resources, and both edge service providers (ESPs) and edge service consumers (ESCs) can benefit significantly from it. Existing paradigms of workload forecasting (i.e., edge-only or cloud-only) are improper, due to failing to consider the inter-site correlations and might suffer from significant data transmission delays. With the increasing adoption of edge platforms by web services, it is critical to balance both accuracy and efficiency in workload forecasting. In this paper, we propose XELASTIC, which offers three key improvements over the conference version. First, we redesigned the aggregation and disaggregation layers using GCNs to capture more complex relationships among workload series. Second, we introduced a supervised contrastive loss to enhance robustness against outliers, particularly for handling missing or abnormal data in real-world scenarios. Finally, we expanded the evaluation with additional baselines and larger datasets. Extensive experiments on realistic edge workload datasets collected from China’s largest edge service provider (Alibaba ENS) show that XELASTIC outperforms state-of-the-art methods, decreases time consumption, and reduces communication costs. Penghong Zhao, Xiao Ma 0009, Haitao Yuan 0002, Zhe Fu 0005, Mengwei Xu 0001, Shangguang Wang |
IEEE Trans. Mob. Comput. | 3 |
| 2025 | Delay- and Resource-Aware Satellite UPF Service OptimizationabstractExecuting 5G core network functions on satellites has become crucial to enhance satellite network management and service capabilities. The User Plane Function (UPF) is responsible for efficient data traffic forwarding and is envisioned as a key and pioneering core network function that will be deployed on satellites. However, managing and providing services with satellite UPFs face dual challenges. Limited satellite resources constrain the user scale that a satellite UPF can service, resulting in an unguaranteed service delay. Moreover, the extremely rapid mobility of satellites renders it difficult for satellite UPFs to provide seamless services. To address the above challenges, this paper presents the first-of-its-kind service optimization scheme for satellite UPFs in terms of switch control, state migration, and traffic routing. To provide guaranteed service delay, we provide a theoretical analysis based on the M/G/1 queue model, demonstrating the service delay-resource consumption trade-off. A satellite UPF switch control scheme is integrated into the service optimization process, which can decrease satellite UPF service delay while saving satellite resources by adjusting the switch control parameters. To provide seamless services, we propose a satellite UPF-oriented state-aware service migration and traffic routing (UPF service optimization) algorithm. A policy network-based reinforcement learning approach is employed to dynamically perceive the satellite network’s state as well as the satellite UPF switch state. Building upon the optimization of service delay through satellite UPF switch control, the processes of state-aware state migration and traffic routing are further employed to reduce delay, ensuring seamless service effectively. Experiments reveal that the proposed algorithm outperforms other benchmark algorithms under different metrics. The service delay is reduced by an average of 23.2% compared with other algorithms. Chao Wang 0093, Xiao Ma 0009, Ruolin Xing, Ao Zhou 0001, Shangguang Wang |
IEEE Trans. Mob. Comput. | 2 |
| 2025 | SatCooper: Enhancing Cooperative Inference Analytics for Satellite Service via Multi-Exit DNNsabstractAs a key technology of intelligent satellite-enabled services in B5G or 6G networks, deploying Deep Neural Networks (DNN) models on satellites has been a notable trend, catering to the daily demand for extensive computing-intensive and latency-sensitive tasks. The computing resources are strategically deployed on satellites where sensor data is generated or collected, facilitating the fine-grained computational inference of DNN-based tasks. However, no prior study has comprehensively explored the crucial inference challenges – e.g., the trade-off between the number of tasks completed and accuracy and partitioning models in multi-exit models – in the resource-constrained space environment. Effective scheduling frameworks cater to various streams of inference tasks are scarce because inference performance may deviate from the ideal situation due to changes in task system status, such as task profiles and network state. To this end, we first formulate a gain-aware in-orbit computing inference problem to strike a proper trade-off between inference latency and the number of tasks completed by dynamically selecting optimal early exit points and model partitioning points. We propose an offline dynamic programming-based algorithm that provides an effective solution when comprehensive system details are to be predicted. We have developed an online learning-based method to schedule inference tasks with uncertain and dynamic system statuses in real-world situations. Our evaluation shows that, compared to baseline methods, the online learning-based algorithm can improve task gain by an average of 87.3% across various tasks. Qiyang Zhang 0001, Shangguang Wang, Jinglong Guan, Praveen Kumar Donta, Xiao Ma 0009, R. Venkatesha Prasad, Schahram Dustdar, Xuanzhe Liu |
IEEE Trans. Mob. Comput. | 5 |
| 2025 | SLICE: Energy-Efficient Satellite-Ground Co-Inference via Layer-Wise Scheduling OptimizationabstractRecent advancements in Low Earth Orbit (LEO) satellites are facilitating the provision of Deep Neural Networks (DNNs)-inherent services to achieve ubiquitous coverage via satellite computing. However, the computational demands and energy consumption of DNN models present significant challenges for satellite computing with limited power and computation resources. Based on the layered characteristics of DNN models, a satellite-ground co-inference strategy has been introduced, which executes certain layers on satellites and the remaining layers on ground servers. Determining the optimal layers for in-orbit processing, however, is non-trivial due to the under-explored energy consumption of satellite computing across different models and restricted yet varying communication conditions of satellite-ground links. In this paper, we first conduct a comprehensive measurement to uncover energy consumption of satellite computing across different layers and models. By summarizing the key observations, we develop a layer-specific energy consumption model tailored to diverse DNN architectures and kernels. We then investigate the energy-efficient satellite-ground co-inference problem and formulate it as an integer-nonlinear programming problem, which presents high computational complexity. To tackle these difficulties, we propose a satellite-ground co-inference algorithm that employs a branch-and-bound strategy, combined with the Sobol sequence and Lagrange multiplier, to reduce complexity and ensure stability across diverse DNN architectures. To evaluate the proposed algorithm, we conduct experiments based on real-world satellite parameters. The results demonstrate that our proposed algorithm can achieve an average energy savings of 96% under various data volumes compared to the existing benchmarks. Qiyang Zhang 0001, Ruolin Xing, Yuanzhe Li 0001, Xiao Ma 0009, Ao Zhou 0001, Shangguang Wang |
IEEE Trans. Serv. Comput. | 5 |
| 2025 | Flexible Shadow: Resource-Efficient Reliability Enhancement for Edge Services Through Dynamic Shadow CoordinationabstractEdge computing plays a pivotal role in supporting services necessitating sub-second latency, notably in domains like Industry 4.0 and autonomous driving. However, unpredictable failure occurring at edge servers can result in prolonged response time and decreased service reliability, posing significant risks to both safety and property. Traditional reliability mechanisms, namely task re-execution and task replication, are often inadequate for edge environments. The former struggles to meet the stringent end-to-end service latency requirements, while the latter imposes a high resource consumption burden on resource-limited edge clouds. To address this issue, this paper introduces a novel Flexible Shadow mechanism, where the backup instance, referred to as the Flexible Shadow, is allocated fewer computation resources compared to its primary instance to conserve computation resources, and temporally preempts a portion of resources from neighboring shadows to accelerate when necessary. To support the implementation of this mechanism, we propose the Flexible Shadow Backup Framework, a resource-efficient reliability enhancement framework for edge services through dynamic shadow coordination. This framework integrates three key components: a deployment algorithm for resource allocation, an adjustment algorithm for migration cost-latency tradeoffs, and a reconfiguration algorithm for adaptation optimization. Comprehensive experiments conducted on a Docker-based prototype demonstrate the effectiveness of the Flexible Shadow mechanism, achieving nearly 60% reduction in computing resource consumption compared to traditional approaches while maintaining sub-second latency. Lipei Yang, Ao Zhou 0001, Xiao Ma 0009, Qing Li 0028, Yuanzhe Li 0001, Shangguang Wang |
IEEE Trans. Serv. Comput. | 3 |
| 2024 | Bidirectional Buffer-Constrained Bitrate Adaptation for Layered 360-degree Video StreamingabstractBy cutting the 360-degree video into temporal chunks and spacial tiles of different quality levels, layered 360-degree video streaming enables fine-grained bitrate adaptation. A critical challenge is whether to transmit the expected viewport at a higher quality level or pre-fetch the low-rate base layer while avoiding buffer overflow or rebuffering, given the fluctuated and constrained fronthaul bandwidth. In this work, we investigate this interplay to provision better quality of experience (QoE) while guaranteeing the smoothness of 360-degree video streaming. Specifically, an optimization problem is formulated to maximize the expected video quality under a bidirectional constraint on the long-term average buffer. To address this problem, we propose the Lyapunov Stable Point Optimization(LSPO), which is a Lyapunov optimization variant. LSPO achieves a near-optimal time-averaged QoE within an O(1/V )margin, while also upholding bidirectional buffer constraints. A buffer-based algorithm is then proposed and evaluated based on real-trace data. The results demonstrate superior performance in video quality and reduced rebuffering probability compared to existing algorithms. Binbin Hu, Shan Zhang 0001, Chang Feng, Zhiyuan Wang 0004, Hongbin Luo, Xiao Ma 0009 |
GLOBECOM | 6 |
| 2024 | Profit-Aware Task Allocation in Satellite ComputingabstractThe rapid evolution of satellite networks promises to expand global internet service. However, optimizing task allocation for the efficient and sustainable operation of satellite computing presents complex challenges. Existing approaches usually prioritize energy considerations while neglecting economic aspects, which restricts satellite networks from achieving their full economic potential. In this paper, we address this gap by investigating task allocation in satellite computing. Our approach encourages satellites to consistently provide resources and optimizes battery usage, enabling the completion of more tasks and ultimately maximizing profit. The task allocation approach involves two key components: task pricing and task scheduling. Firstly, we introduce a unique task pricing algorithm that adheres to economic properties, establishing a direct link between satellite utilization and financial income, ensuring economically viable satellite operations. Moreover, we develop two distinct task scheduling algorithms tailored for offline and online scenarios, exploiting dynamic programming and reinforcement learning respectively. Extensive simulations demonstrate that our proposed algorithms effectively enhance task completion rates and optimize total satellite profit. Jie Huang 0021, Ruolin Xing, Xiao Ma 0009, Ao Zhou 0001, Shangguang Wang |
ICWS | 3 |
| 2024 | Flexible Shadow: Enhancing Service Reliability in Resource-Constrained Edge ComputingabstractEdge computing plays a pivotal role in supporting services necessitating sub-second latency, notably in domains like Industry 4.0 and autonomous driving. However, unpredictable failure occurring at edge servers can result in prolonged response time and decreased service reliability, posing significant risks to both safety and property. Traditional reliability mechanisms, namely re-execution and replication, are often inadequate for edge environments, with the former often struggle to meet latency requirements and the latter imposing a high resource consumption burden on resource-limited edge clouds. To address this issue, this paper introduces a novel "Flexible Shadow" mechanism, where the backup instance, referred to as the "Flexible Shadow", is allocated fewer computation resources compared to its primary instance to conserve computation resources, and temporally preempts a portion of resources from neighboring shadows to accelerate when necessary. To tackle the implementation challenges of this framework arising from diverse service requirements, dynamic nature of edge environments and potential deadlock in the reconfiguration process, we devised the Flexible Shadow Deployment Algorithm for accurate shadow deployment and the Flexible Shadow Reconfiguration Algorithm for dynamic strategy adjustment. We have implemented our Flexible Shadow framework on Docker and evaluated it via comprehensive experiments. The experiment results demonstrate a nearly 60% reduction in computing resource consumption while ensuring sub-second latency. Lipei Yang, Ao Zhou 0001, Xiao Ma 0009, Yuanzhe Li 0001, Shangguang Wang |
ICWS | 3 |
| 2024 | Resource-efficient In-orbit Detection of Earth ObjectsabstractWith the rapid proliferation of large Low Earth Orbit (LEO) satellite constellations, a huge amount of in-orbit data is generated and needs to be transmitted to the ground for processing. However, traditional LEO satellite constellations, which downlink raw data to the ground, are significantly restricted in transmission capability. Orbital edge computing (OEC), which exploits the computation capacities of LEO satellites and processes the raw data in orbit, is envisioned as a promising solution to relieve the downlink burden. Yet, with OEC, the bottleneck is shifted to the inelastic computation capacities. The computational bottleneck arises from two primary challenges that existing satellite systems have not adequately addressed: the inability to process all captured images and the limited energy supply available for satellite operations. In this work, we seek to fully exploit the scarce satellite computation and communication resources to achieve satellite-ground collaboration and present a satellite-ground collaborative system named TargetFuse for onboard object detection. TargetFuse incorporates a combination of techniques to minimize detection errors under energy and bandwidth constraints. Extensive experiments show that TargetFuse can reduce detection errors by 3.4× on average, compared to onboard computing. TargetFuse achieves a 9.6× improvement in bandwidth efficiency compared to the vanilla baseline under the limited bandwidth budget constraint. Qiyang Zhang 0001, Ruolin Xing, Zimu Zheng, Xiao Ma 0009, Mengwei Xu 0001, Schahram Dustdar, Shangguang Wang |
INFOCOM | 6 |
| 2024 | Energy-Aware Satellite-Ground Co-Inference via Layer-Wise Processing Schedule OptimizationabstractRecent advancements in Low Earth Orbit (LEO) satellites are facilitating the provision of Deep Neural Networks (DNNs)-inherent services to achieve ubiquitous coverage via satellite computing. However, the computational demands and energy consumption of DNN models pose significant challenges for satellite computing with limited power and computation resources. Based on the hierarchical characteristics of DNN models, we propose a satellite-ground co-inference strategy that executing certain layers on satellites and the remaining layers on ground servers. However, identifying the optimal layers for in-orbit processing with latency constraints is challenging due to the uncertain energy consumption across diverse models. To explore the correlation between energy consumption and layer types, we conduct comprehensive measurements on a hardware device commonly found in commercial LEO satellites and develop a layer-based energy consumption prediction model. Then, we formulate an optimization problem of minimizing the energy consumption on the satellite within the latency constraint as an integer nonlinear programming problem. Solving this problem is difficult due to combinatorial explosion in the discrete solution space. To address this, we propose an improved algorithm based on genetic algorithms. Using configurations from a real satellite, we conduct simulation experiments, concluding that our algorithm significantly improves energy savings by an average of 27 ×. Qiyang Zhang 0001, Ruolin Xing, Yuanzhe Li 0001, Xiao Ma 0009, Chaoxin Yu, Ao Zhou 0001, Shangguang Wang |
Internetware | 5 |
| 2024 | Poster: Service Orchestration for Satellite ComputingabstractSatellite computing is emerging as a promising domain for delivering mobile services that meet stringent Quality of Service (QoS) requirements, such as low latency, to users. However, the inherent mobility of satellites as computing nodes can precipitate QoS degradation, a challenge not encountered in terrestrial cloud systems. This discrepancy poses significant adaptation challenges for cloud service orchestration systems, such as Kubernetes, due to the rapid movement of satellites. This poster introduces a service orchestration system and a corresponding service placement strategy tailored for satellite computing environments. Our proposed architecture and strategy surpass traditional fixed instance deployment by not only achieving lower average latency but also maintaining an optimal balance between benefits and costs. Ruolin Xing, Qibo Sun, Ao Zhou 0001, Xiao Ma 0009 |
MobiSys | 5 |
| 2024 | Exploring Real-Time Satellite Computing: From Energy and Thermal PerspectivesabstractSmall satellites (SmallSats) are now widely used in various fields, such as real-time communication and earth observation. These increasingly complex space applications face limited support from conventional radiation-hardened processors onboard. Hence, many SmallSats are designed to utilize high performance commercial off-the-shelf (COTS) computing devices to address this problem but it remains unclear how the unique energy and thermal characteristics of SmallSats impact computing efficiency onboard. This work conducts a systematic and quantitative measurement study of COTS devices’ computing efficiency on two real orbiting SmallSats. The key findings are: 1) inadequate energy management may lead to electricity wastage in sunlit zones and shortages in eclipse zones, impacting onboard computing availability and 2) the weak heat dissipation onboard may compromise COTS computing efficiency by incurring thermal throttling. To address such challenges, we design ProScale, a lightweight application-aware power management and thermal control system to improve computing efficiency under both electrical and thermal energy constraints. Evaluation shows that ProScale can improve the average task completion latency by $2.1 \times$ for computation-intensive applications compared with baselines. Qing Li 0028, Shangguang Wang, Chenren Xu, Xiao Ma 0009, Mengwei Xu 0001, Ao Zhou 0001, Ruolin Xing, Zuo Zhu, Ying Zhang 0012, Xuanzhe Liu |
RTSS | 4 |
| 2024 | More is Different: Prototyping and Analyzing a New Form of Edge Server with Massive Mobile SoCs
Li Zhang 0133, Zhe Fu 0005, Boqing Shi, Xiang Li 0067, Rujin Lai, Chenyang Yang 0004, Ao Zhou 0001, Xiao Ma 0009, Shangguang Wang, Mengwei Xu 0001 |
USENIX ATC | 8 |
| 2024 | Multiparticipant Double Auction for Resource Allocation and Pricing in Edge ComputingabstractEdge computing serves as a critical solution for latency-sensitive services on mobile and IoT devices. However, the high cost and limited edge resources present significant challenges for service and infrastructure providers in establishing efficient collaborations, particularly with conflicting profit objectives. Inspired by the pseudo elbow formation of octopuses, we propose a multi-participant double auction for resource allocation and pricing between service and infrastructure providers. We introduce a neutral third-party auctioneer to eliminate direct bargaining among participants, leading to an improved amount of allocated resources and matching efficiency. The presence of heterogeneous participants, many-to-many mapping and an advisable payment strategy that satisfies economic properties exacerbate the difficulty. To address these challenges, we propose a Matching and Pricing Resource Allocation algorithm for a long-term steady market, and a Truthful Resource Allocation algorithm for a short-term market. Simulation results demonstrate that the proposed algorithms exhibit superior performance not only in maximizing social welfare and utility of both service and infrastructure providers, but also in improving resource utilization. Jie Huang 0021, Lipei Yang, Jianing Si, Xiao Ma 0009, Shangguang Wang |
IEEE Internet Things J. | 5 |
| 2024 | Online Request Replication for Obtaining Fresh Information Under Pull ModelabstractAge of Information (AoI) has gained widespread usage and emerged as a pivotal metric for assessing timeliness performance in information-update systems. Such systems often entail service requirements for rapidly obtaining requested data in real-time. For instance, in the financial market, users rely on up-to-date and low-latency information to make appropriate trading decisions to maximize profits within their financial budge. Much of the existing research on real-time services focuses on ensuring AoI or service-level latency, but there is a growing demand for joint optimization of these two metrics to accommodate a broader range of potential applications. Therefore, this article investigates the problem of minimizing AoI within the context of statistical latency guarantees. To tackle the critical challenges posed by the joint modeling of AoI and statistical latency, system uncertainty, as well as tradeoff between performance and user’s budget, we employ a replication scheme to ensure both AoI and statistical service-level latency. To address the critical challenges posed by the unknown distribution in data updating processes and response times across providers, we formulate the AoI minimization problem with statistical latency constraints as a combinatorial multiarmed bandit problem utilizing the Lyapunov optimization theory. Subsequently, we propose an online learning-based request replication algorithm to address this problem. Our proposed algorithm achieves a cumulative regret of$O(T\sqrt {\log (T)})$compared to the genie-aided algorithm. Simulation results demonstrate the superior performance of the proposed algorithm against benchmarks. Qibo Sun, Qing Li 0028, Xiao Ma 0009, Shangguang Wang |
IEEE Internet Things J. | 4 |
| 2024 | Reliability-Aware Task Replication for Mobile Edge ComputingabstractAs infrastructure deployment continues to expand worldwide, the development of the Internet of Vehicles has become increasingly feasible. With widespread cellular connectivity and powerful roadside computing capabilities, advanced driving assistance systems can now rely on roadside decision models in addition to vehicle-side ones, overcoming the limitations of a single vehicle’s perception range. This shift has led to improvements in manufacturing efficiency, cruising range, and battery life of intelligent vehicles. However, maintaining the ultra-low latency and high-reliability requirements of on-vehicle services is still a challenge due to air interface fluctuations and edge server computing loads, which could potentially jeopardize driving safety. To tackle this issue, we conducted real-world measurements of edge server access delay in LTE and 5G cellular networks. Our analysis identified key factors affecting delay distribution, leading to the development of an approximate fitting function for the delay probability density function. We also proposed a reliability-aware task replication algorithm that leverages delay samples and edge server status information to make real-time task replication and offloading decisions, minimizing replication while ensuring service reliability. Simulations based on real-world datasets indicate our approach reduces task completion delay by up to 42.11% and limits the maximum task replication redundancy peak value to 63.37%, effectively ensuring the reliability of on-vehicle services during driving. Lipei Yang, Ao Zhou 0001, Xiao Ma 0009, Shangguang Wang |
IEEE Internet Things J. | 3 |
| 2024 | Seamless Cross-Edge Service Migration for Real-Time Rendering ApplicationsabstractSeamless cross-edge migration for real-time rendering applications is challenging. The strong interactive nature of real-time rendering applications demands a downtime lower than 15ms to achieve an imperceptible migration. Existing methods based on virtual machine migration and container migration suffer from unpleasant downtime brought by dirty page retransmission-induced repeated memory data copy and the shared storage failure-induced extensive disk data copy. In this paper, we propose Cloud-assisted Service Migration (CSM) which leverages cloud-edge collaboration to achieve seamless service migration for real-time rendering applications. CSM improves service migration user experience in three folds: First, it introduces a dual rendering mechanism to bypass the peer-to-peer data copy and compresses the freezing stage. Second, a user equipment-centric session switch mechanism is proposed to save time by well coordinating application session switches and 5G user plane session switches. Third, a smooth switching mechanism is leveraged to prevent unpleasant frame flickers during session switching. We implement CSM in edge-rendering multiplayer games and deploy it on a 5G test bed with a full-stack user plane protocol stack. The evaluation results show that CSM can reduce downtime to < 14ms and the service migration process is user imperceptible. Yuanzhe Li 0001, Shangguang Wang, Yuanchun Li 0003, Ao Zhou 0001, Mengwei Xu 0001, Xiao Ma 0009, Yunxin Liu 0001 |
IEEE Trans. Mob. Comput. | 6 |
| 2024 | Towards Timely Video Analytics Services at the Network EdgeabstractReal-time video analytics services aim to provide users with accurate recognition results timely. However, existing studies usually fall into the dilemma between reducing delay and improving accuracy. The edge computing scenario imposes strict transmission and computation resource constraints, making balancing these conflicting metrics under dynamic network conditions difficult. In this regard, we introduce the age of processed information (AoPI) concept, which quantifies the time elapsed since the generation of the latest accurately recognized frame. AoPI depicts the integrated impact of recognition accuracy, transmission, and computation efficiency. We derive closed-form expressions for AoPI under preemptive and non-preemptive computation scheduling policies w.r.t. the transmission/computation rate and recognition accuracy of video frames. We then investigate the joint problem of edge server selection, video configuration adaptation, and bandwidth/computation resource allocation to minimize the long-term average AoPI over all cameras. We propose an online method, i.e., Lyapunov-based block coordinate descent (LBCD), to solve the problem, which decouples the original problem into two subproblems to optimize the video configuration/resource allocation and edge server selection strategy separately. We prove that LBCD achieves asymptotically optimal performance. According to the testbed experiments and simulation results, LBCD reduces the average AoPI by up to 10.94X compared to state-of-the-art baselines. Xishuo Li, Shan Zhang 0001, Yuejiao Huang, Xiao Ma 0009, Zhiyuan Wang 0004, Hongbin Luo |
IEEE Trans. Mob. Comput. | 4 |
| 2024 | A Comprehensive Deep Learning Library Benchmark and Optimal Library SelectionabstractDeploying deep learning (DL) on mobile devices has been a notable trend in recent years. To support fast inference of on-device DL, DL libraries play a critical role as algorithms and hardware do. Unfortunately, no prior work ever dives deep into the ecosystem of modern DL libraries and provides quantitative results on their performance. In this paper, we first build a comprehensive benchmark that includes 6 representative DL libraries and 15 diversified DL models. Then we perform extensive experiments on 10 mobile devices, and the results reveal the current landscape of mobile DL libraries. For example, we find that the best-performing DL library is severely fragmented across different models and hardware, and the gap between DL libraries can be rather huge. In fact, the impacts of DL libraries can overwhelm the optimizations from algorithms or hardware, e.g., model quantization and GPU/DSP-based heterogeneous computing. Motivated by the fragmented performance of DL libraries across models and hardware, we propose an effective DL Library selection framework to obtain the optimal library on a new dataset that has been created. We evaluate the DL Library selection algorithm, and the results show that the framework at it can improve the prediction accuracy by about 10% than benchmark approaches on average. Qiyang Zhang 0001, Xiangying Che, Xiao Ma 0009, Mengwei Xu 0001, Schahram Dustdar, Xuanzhe Liu, Shangguang Wang |
IEEE Trans. Mob. Comput. | 4 |
| 2024 | Battery-Aware Energy Optimization for Satellite Edge ComputingabstractSatellite edge computing can incur dramatically increased energy demand onboard, which is met by satellite batteries during eclipses. Excessive energy usage during regular operations accelerates battery wear. Therefore, it is important and timely to optimize the energy consumption onboard to extend satellite batteries life. This paper investigates battery-aware energy optimization for satellite edge computing under energy harvesting dynamics and wireless environment uncertainty. Inspired by the periodical energy harvesting and satellite-ground connection, we develop a pattern-aware online energy scheduling algorithm within an online convex optimization framework. This learning algorithm achieves theoretical guarantees of no regret and gradually zeroing constraint violations. We further exploit inter-satellites collaboration to extend the average battery life in a whole constellation where satellites have different battery capacity degradation. Trace-driven simulations show that our algorithm can significantly extend the battery life by 1.32× and effectively adapt to the energy harvesting dynamics and wireless environment uncertainty. Qing Li 0028, Shangguang Wang, Xiao Ma 0009, Ao Zhou 0001, Yue Wang 0072, Gang Huang 0001, Xuanzhe Liu |
IEEE Trans. Serv. Comput. | 3 |
| 2023 | Cooperative Content Caching and Distribution for Satellite CDNsabstractContent distribution networks play a crucial role in the services provided by streaming applications, but current terrestrial CDNs have uneven coverage and poor performance in some remote and rural areas. Satellite communications have emerged as an effective way to address these issues. However, the satellite environment is highly dynamic, and how to reasonably deploy CDNs in the satellite constellation to realize efficient content caching and distribution is an urgent problem to be solved. To this end, this paper proposes a content caching and content distribution approach for satellite CDNs. The approach uses Lyapunov optimization, Gibbs sampling, and matching theory to control the traffic cost of satellite-cached content copies while reducing the content distribution latency. A cost-effective content cache deployment and distribution strategy is proposed for large low-orbiting satellite constellations operating at high speeds. Extensive simulation results show that the algorithm can quickly converge to favorable values compared to the reference algorithm, effectively reducing latency for end users while maintaining low deployment traffic costs. Xiao Ma 0009, Ao Zhou 0001, Shangguang Wang |
ICNP | 2 |
| 2023 | Energy and Time-Aware Inference Offloading for DNN-based Applications in LEO SatellitesabstractIn recent years, Low Earth Orbit (LEO) satellites have witnessed rapid development, with inference based on Deep Neural Network (DNN) models emerging as the prevailing technology for remote sensing satellite image recognition. However, the substantial computation capability and energy demands of DNN models, coupled with the instability of the satellite-ground link, pose significant challenges, burdening satellites with limited power intake and hindering the timely completion of tasks. Existing approaches, such as transmitting all images to the ground for processing or executing DNN models on the satellite, is unable to effectively address this issue. By exploiting the internal hierarchical structure of DNNs and treating each layer as an independent subtask, we propose a satellite-ground collaborative computation partial offloading approach to address this challenge. We formulate the problem of minimizing the inference task execution time and onboard energy consumption through offloading as an integer linear programming (ILP) model. The complexity in solving the problem arises from the combinatorial explosion in the discrete solution space. To address this, we have designed an improved optimization algorithm based on branch and bound. Simulation results illustrate that, compared to the existing approaches, our algorithm improve the performance by 10%-18%. Qiyang Zhang 0001, Xiao Ma 0009, Ao Zhou 0001 |
ICNP | 4 |
| 2023 | Freshness-aware Content Update for Earth Observation: Trading Off AoI and AccuracyabstractSpace networks composed of Low Earth Orbit (LEO) satellites play a significant role in Earth observation systems. LEO satellite constellations, which provide wide-ranging coverage, continuously observe the Earth and transmit the observed data to satellites in higher orbits or ground stations for further analysis, benefiting from their more powerful computation capacity. In such an observation and analysis network, the timeliness of data update and the quality of analysis are crucial factors, both constrained by limited energy resources. In this paper, we jointly optimize the timeliness and analysis quality by constructing an update scheduling problem with a trade-off between Age of Information (AoI) and accuracy. Considering the unpredictable and changeable environment dynamics, we formulate this problem into a Markov decision process and propose an algorithm based on deep reinforcement learning to obtain an optimal policy that effectively tackles the tradeoff problem. Simulation results demonstrate that our algorithm can jointly reduce AoI and increase accuracy, outperforming benchmark algorithms in different cases. Yuran Guo, Xiao Ma 0009, Qibo Sun, Ao Zhou 0001 |
ICNP | 2 |
| 2023 | Optimizing Space-Borne Computation: A Reliability Enhancement Framework for LEO ConstellationabstractEdge computing is extending the frontier of computation beyond terrestrial boundaries, with Low Earth Orbit (LEO) constellations emerging as the cutting-edge paradigm, aspiring to deliver ubiquitous computing capabilities globally. However, the unresolved service reliability issues within LEO constellations continue to barricade the realization of the vision for space computing. Terrestrial-native classical methods struggle to meet the stringent environmental conditions of space computing and often fail to satisfy the rigorous requirements of LEO constellations for failure response times and computational resource overhead. To address these challenges, we introduce the LEO Service Reliability Enhancement Framework (LSREF). LSREF employs the Variable Speed Replication technology to balance failure response times and computational resource overhead and adopts a decentralized design, more congruent with the characteristics of LEO constellations. Specifically, to counter the immense scale and high dynamism of LEO constellations, LSREF proposes orbital plane-based autonomous domains and leverages the Satellite Autonomous Backup Deployment Algorithm in conjunction with the Registry Polling Mechanism to enable autonomous decision-making for service backup strategies on each satellite. Our simulation experiments demonstrate that, compared to classical methods, LSREF reduces average failure response times by 10.76% and diminishes computational resource consumption by nearly 15.53%. Lipei Yang, Ao Zhou 0001, Xiao Ma 0009, Shangguang Wang |
ICPADS | 3 |
| 2023 | Towards Timely Edge-assisted Video Analytics ServicesabstractReal-time video analytics services are expected to deliver accurate recognition results to users timely. However, existing studies usually fail in the dilemma between reducing delay and improving accuracy. Balancing such conflicting metrics is extremely hard under dynamic network conditions. In this regard, we introduce the age of processed information (AoPI) concept, i.e., the time elapsed since the generation of the latest accurately recognized frame. As a systematic metric, AoPI depicts the integrated impact of recognition accuracy, transmission, and computation efficiency. We derive the closed-form expressions of AoPI under preemptive and non-preemptive computation scheduling policies w.r.t. the transmission/computation rate and recognition accuracy of video frames, which are validated using prototype experiments. Based on the derived results, we study the joint video configuration selection, bandwidth, and computation resource allocation problem to minimize the long-term average AoPI among all cameras. An efficient method is proposed to solve the problem, which utilizes Lyapunov optimization and block coordinate descent to make decisions online without requiring future information about network variations and video content. We prove that our method achieves asymptotically optimal performance. Extensive simulations show that our method reduces the AoPI by up to 4.06X compared with the state-of-the-art baselines. Xishuo Li, Shan Zhang 0001, Yuejiao Huang, Xiao Ma 0009, Zhiyuan Wang 0004, Hongbin Luo |
ICWS | 4 |
| 2023 | Container Image Similarity-Aware Resource Provisioning for Serverless Edge ComputingabstractContainer-enabled serverless computing has become a widely adopted approach for resource provisioning in the edge cloud. However, traffic incurred by container image pulling heavily burdens the already congested back-haul network. To relieve the problem, we do an analysis on Docker Hub, and find that instance deployment strategy has a significant impact on the back-haul traffic due to the varying similarity levels of different images. We incorporate this feature into task offloading decision and resource provisioning, and formulate the problem with a mixed integer non-linear programming (MINLP) problem. To address the challenges arising from the coupling and contradiction of instance deployment, image pulling, offloading decision, and resource allocation, we employ multi-agent deep reinforcement learning to decompose the problem into several simpler sub-problems, and design an algorithm for each sub-problem individually by exploiting convex optimization and fractional programming techniques. Simulations are conducted to validate the effectiveness of the proposed algorithm. The experiment results illustrate that our algorithm outperforms current notable solutions and improves the global utility by 13%–74%. Ao Zhou 0001, Xiao Ma 0009, Shangguang Wang |
ICWS | 3 |
| 2023 | Privacy as a Resource in Differentially Private Federated LearningabstractDifferential privacy (DP) enables model training with a guaranteed bound on privacy leakage, therefore is widely adopted in federated learning (FL) to protect the model update. However, each DP-enhanced FL job accumulates privacy leakage, which necessitates a unified platform to enforce a global privacy budget for each dataset owned by users. In this work, we present a novel DP-enhanced FL platform that treats privacy as a resource and schedules multiple FL jobs across sensitive data. It first introduces a novel notion of device-time blocks for distributed data streams. Such data abstraction enables fine-grained privacy consumption composition across multiple FL jobs. Regarding the non-replenishable nature of the privacy resource (that differs it from traditional hardware resources like CPU and memory), it further employs an allocation-then-recycle scheduling algorithm. Its key idea is to first allocate an estimated upper-bound privacy budget for each arrived FL job, and then progressively recycle the unused budget as training goes on to serve further FL jobs. Extensive experiments show that our platform is able to deliver up to 2.1× as many completed jobs while reducing the violation rate by up to 55.2% under limited privacy budget constraint. Jinliang Yuan, Shangguang Wang, Shihe Wang, Yuanchun Li 0003, Xiao Ma 0009, Ao Zhou 0001, Mengwei Xu 0001 |
INFOCOM | 5 |
| 2023 | Boosting DNN Cold Inference on Edge DevicesabstractDNNs are ubiquitous on edge devices nowadays. With its increasing importance and use cases, it's not likely to pack all DNNs into device memory and expect that each inference has been warmed up. Therefore, cold inference, the process to read, initialize, and execute a DNN model, is becoming commonplace and its performance is urgently demanded to be optimized. To this end, we present NNV12, the first on-device inference engine optimizing cold inference. NNV12 is built atop three novel optimization knobs: selecting a proper kernel (i.e., operator implementation) for each DNN operator, bypassing the weights transformation process by caching the post-transformed weights on disk, and pipelined execution of many kernels on asymmetric processors. To tackle with the huge search space, NNV12 employs a heuristic-based scheme to obtain a near-optimal kernel scheduling plan. We fully implement a prototype of NNV12 and evaluate its performance across extensive experiments. It shows that NNV12 achieves up to 15.2× speedup compared to the state-of-the-art DNN engines on edge CPUs and 401.5× speedup on edge GPUs, respectively. Rongjie Yi, Ting Cao 0003, Ao Zhou 0001, Xiao Ma 0009, Shangguang Wang, Mengwei Xu 0001 |
MobiSys | 4 |
| 2023 | ELASTIC: Edge Workload Forecasting based on Collaborative Cloud-Edge Deep LearningabstractWith the rapid development of edge computing in the post-COVID19 pandemic period, precise workload forecasting is considered the basis for making full use of the edge limited resources, and both edge service providers (ESPs) and edge service consumers (ESCs) can benefit significantly from it. Existing paradigms of workload forecasting (i.e., edge-only or cloud-only) are improper, due to failing to consider the inter-site correlations and might suffer from significant data transmission delays. With the increasing adoption of edge platforms by web services, it is critical to balance both accuracy and efficiency in workload forecasting. In this paper, we propose ELASTIC, which is the first study that leverages a cloud-edge collaborative paradigm for edge workload forecasting with multi-view graphs. Specifically, at the global stage, we design a learnable aggregation layer on each edge site to reduce the time consumption while capturing the inter-site correlation. Additionally, at the local stage, we design a disaggregation layer combining both the intra-site correlation and inter-site correlation to improve the prediction accuracy. Extensive experiments on realistic edge workload datasets collected from China’s largest edge service provider show that ELASTIC outperforms state-of-the-art methods, decreases time consumption, and reduces communication cost. Haitao Yuan 0002, Zhe Fu 0005, Xiao Ma 0009, Mengwei Xu 0001, Shangguang Wang |
WWW | 4 |
| 2023 | Placing Timely Refreshing Services at the Network EdgeabstractAccommodating services at the network edge is favorable for time-sensitive applications. However, maintaining service usability is resource consuming in terms of pulling service images to the edge, synchronizing databases of service containers, and hot updates of service modules. Accordingly, it is critical to determine which service to place based on the received user requests and service refreshing (maintaining) cost, which is usually neglected in existing studies. In this work, we study how to cooperatively place timely refreshing services and offload user requests among edge servers to minimize the backhaul transmission costs. We formulate an integer nonlinear programming problem and prove its NP-hardness. This problem is highly nontractable due to the complex spatial-and-temporal coupling effect among service placement, offloading, and refreshing costs. We first decouple the problem in the temporal domain by transforming it into a Markov shortest path problem. We then propose a lightweighted discounted value approximation (DVA) method, which further decouples the problem in the spatial domain by estimating the offloading costs among edge servers. The worst performance of DVA is proved to be bounded. 5G service placement testbed experiments and real-trace simulations show that DVA reduces the total transmission cost by up to 59.1% compared with the state-of-the-art baselines. Xishuo Li, Shan Zhang 0001, Hongbin Luo, Xiao Ma 0009 |
IEEE Internet Things J. | 4 |
| 2023 | Towards Diversified IoT Image Recognition Services in Mobile Edge ComputingabstractWith the rapid development of the Internet of Things (IoT) and emerging Mobile Edge Computing (MEC) technologies, various IoT image recognition services are revolutionizing our lives by providing diverse cognitive assistance. However, most existing related approaches are difficult to meet the diversified needs of users because they believe that the MEC platform is a single layer. In addition, due to the mutual interference between the data, it is not easy for them to extract the discriminative features (DFs) necessary to analyze the input data. To this end, this article proposes an IoT image recognition services framework for different needs in the MEC environment, which consists of Hierarchical Discriminative Feature Extraction (HDFE) and Sub-extractor Deployment (Sub-ED) algorithms. We first propose HDFE, which can avoid mutual interference between data by separately optimizing the data structure, thereby generating an extractor that extracts effective DFs. Then there is Sub-ED, which divides the extractor into a series of sub-extractors and deploys them on appropriate MEC platforms. By doing so, the IoT device can connect to the corresponding MEC platform according to its service types, and use the sub-extract to extract DFs. Then, the MEC platform uploads the extracted feature data to the cloud server for further processing, e.g., feature matching. Finally, the cloud server sends the processed result back to the IoT device. Experimental results show that compared with the state-of-the-art approaches, the proposed framework improves recognition accuracy by about 6% and reduces network traffic by up to 94%. Chuntao Ding, Ao Zhou 0001, Xiao Ma 0009, Ning Zhang 0007, Ching-Hsien Hsu, Shangguang Wang |
IEEE Trans. Cloud Comput. | 3 |
| 2023 | Online Service Request Duplicating for Vehicular ApplicationsabstractVehicles on roads have increasingly powerful computing capabilities and edge nodes are being widely deployed. They can work together to provide computing services for onboard driving systems, passengers, and pedestrians. Typical applications in vehicular systems have service requirements such as low latency and high reliability. Most studies in vehicular networks concerning latency and reliability focus on vehicular communication at the network level. Based on these fundamental works, an increasing proportion of vehicles boast complex applications that require service-level end-to-end performance guarantees. Several works guarantee service-level latency or reliability while new and innovative applications are demanding a joint optimization of the above two metrics. To address the critical challenges induced by the joint modeling of latency and reliability, system uncertainty, and performance and cost trade-off, we employ service request duplication to ensure both latency and reliability performance at the service level. We propose an online learning-based service request duplication algorithm based on a multi-armed bandit framework and Lyapunov optimization theory. The proposed algorithm achieves an upper-bounded regret compared to the oracle algorithm. Simulations are based on real-world datasets and the results demonstrate that the proposed algorithm outperforms the benchmarks. Qing Li 0028, Xiao Ma 0009, Ao Zhou 0001, Changhee Joo, Shangguang Wang |
IEEE Trans. Mob. Comput. | 2 |
| 2023 | User-Oriented Edge Node Grouping in Mobile Edge ComputingabstractIn mobile edge computing networks, densely deployed access points are empowered with computation and storage capacities. This brings benefits of enlarged edge capacity, ultra-low latency, and reduced backhaul congestion. This paper concerns edge node grouping in mobile edge computing, where multiple edge nodes serve one end user cooperatively to enhance user experience. Most existing studies focus on centralized schemes that have to collect global information and thus induce high overhead. Although some recent studies propose efficient decentralized schemes, most of them did not consider the system uncertainty from both the wireless environment and other users. To tackle the aforementioned problems, we first formulate the edge node grouping problem as a game that is proved to be an exact potential game with a unique Nash equilibrium. Then, we propose a novel decentralized learning-based edge node grouping algorithm, which guides users to make decisions by learning from historical feedback. Furthermore, we investigate two extended scenarios by generalizing our computation model and communication model, respectively. We further prove that our algorithms converge to the Nash equilibrium with upper-bounded learning loss. Simulation results show that our mechanisms can achieve up to 96.99% of the oracle benchmark. Qing Li 0028, Xiao Ma 0009, Ao Zhou 0001, Xiapu Luo, Fangchun Yang, Shangguang Wang |
IEEE Trans. Mob. Comput. | 2 |
| 2023 | Dynamic Task Scheduling in Cloud-Assisted Mobile Edge ComputingabstractThe cloud-assisted mobile edge computing system is a critical architecture to process computation-intensive and delay-sensitive mobile applications in close proximity to mobile users with high resource efficiency. Due to the heterogenous dynamics of task arrivals at edge nodes and the distributed nature of the system, the workloads of edge nodes are prone to be unbalanced, which can cause high task response time and resource cost. This paper solves the dynamic task scheduling problem in cloud-assisted mobile edge computing (including both peer task scheduling among edge nodes and cross-layer task scheduling from edge nodes to the cloud), aiming at minimizing average task response time within resource budget limit. To overcome the challenges of task arrival dynamics, edge node heterogeneity, and computation-communication delay tradeoff, we propose aWater-filling BasedDynamic TaskScheduling (WiDaS) algorithm. WiDaS dynamically tunes the usage of cloud resources based on the Lyapunov optimization method and efficiently schedules mobile tasks among edge nodes (and the cloud) by exploiting the idea of water filling. Extensive simulations are conducted to evaluate WiDaS under a trace-driven traffic pattern and two mathematic traffic patterns. The results demonstrate that WiDaS shows two-fold benefits of efficiency and effectiveness. In terms of efficiency, WiDaS can achieve the approximate results with the KKT-based algorithm while reducing the computation complexity from exponential order to polynomial order. In terms of effectiveness, WiDaS can reduce the average task response time by up to 64.4% and 47.2% over the Fair-ratio and the Edge-first algorithm. Xiao Ma 0009, Ao Zhou 0001, Shan Zhang 0001, Qing Li 0028, Alex X. Liu, Shangguang Wang |
IEEE Trans. Mob. Comput. | 1 |
| 2022 | Collaborative Mobile Edge Computing Through UPF Selection
Yuanzhe Li 0001, Ao Zhou 0001, Xiao Ma 0009, Shangguang Wang |
CollaborateCom (2) | 3 |
| 2022 | Commutativity-guaranteed Docker Image Reconstruction towards Effective Layer SharingabstractOwing to the benefit of light weight, containers have become a promising enabler for cloud native computing. Container images composed of applications and dependencies support flexible service deployment and migration. Rapid adoption and integration of containers generate millions of images to be stored. Additionally, non-local images have to be frequently downloaded from the registry, resulting in huge amounts of traffic. Content Addressable Storage (CAS) has been adopted for saving storage and networking by enabling identical layers sharing across images. However, according to our measurements, the implication of CAS is significantly limited as layers are rarely fully identical in practice. In this paper, we propose to reconstruct the docker images to raise the number of identical layers and thereby reduce storage and network consumption. We explore the layered structure of images and define the commutativity of files to assure image validity. The image reconstruction is formulated as an integer nonlinear programming problem. Inspired by the observed similarity of layers, we design a similarity-aware online image reconstruction algorithm. Extensive evaluations are conducted to verify the performance of the proposed approach. Ao Zhou 0001, Xiao Ma 0009, Mengwei Xu 0001, Shangguang Wang |
WWW | 3 |
| 2022 | A Comprehensive Benchmark of Deep Learning Libraries on Mobile DevicesabstractDeploying deep learning (DL) on mobile devices has been a notable trend in recent years. To support fast inference of on-device DL, DL libraries play a critical role as algorithms and hardware do. Unfortunately, no prior work ever dives deep into the ecosystem of modern DL libs and provides quantitative results on their performance. In this paper, we first build a comprehensive benchmark that includes 6 representative DL libs and 15 diversified DL models. We then perform extensive experiments on 10 mobile devices, which help reveal a complete landscape of the current mobile DL libs ecosystem. For example, we find that the best-performing DL lib is severely fragmented across different models and hardware, and the gap between those DL libs can be rather huge. In fact, the impacts of DL libs can overwhelm the optimizations from algorithms or hardware, e.g., model quantization and GPU/DSP-based heterogeneous computing. Finally, atop the observations, we summarize practical implications to different roles in the DL lib ecosystem. Qiyang Zhang 0001, Xiang Li 0067, Xiangying Che, Xiao Ma 0009, Ao Zhou 0001, Mengwei Xu 0001, Shangguang Wang, Yun Ma 0002, Xuanzhe Liu |
WWW | 4 |
| 2022 | Service Coverage for Satellite Edge ComputingabstractRecently, increasing investments in satellite-related technologies make the low earth orbit (LEO) satellite constellation a strong complement to terrestrial networks. To mitigate the limitations of the traditional satellite constellation “bent-pipe” architecture, satellite edge computing (SEC) has been proposed by placing computing resources at the LEO satellite constellation. Most existing works focus on space-air-ground integrated network architecture and SEC computing framework. Beyond these works, we are the first to investigate how to efficiently deploy services on the SEC nodes to realize robustness aware service coverage with constrained resources. Facing the challenges of spatial-temporal system dynamics and service coverage-robustness conflict, we propose a novel online service placement algorithm with a theoretical performance guarantee by leveraging Lyapunov optimization and Gibbs sampling. Extensive simulation results show that our algorithm can improve the service coverage by$4.3\times $compared with the baseline. Qing Li 0028, Shangguang Wang, Xiao Ma 0009, Qibo Sun, Houpeng Wang, Suzhi Cao, Fangchun Yang |
IEEE Internet Things J. | 3 |
| 2022 | Profit-Aware Edge Server PlacementabstractIn a 5G network, mobile-edge computing (MEC) plays a key role in providing low access delay services. The placement of edge servers not only determines the quality of services on the user side but also affects the profit of running a MEC system. In this article, we study how to properly place edge servers so as to guarantee the access delay and maximize the profit of edge providers. We first propose a profit model which involves both access delay and energy consumption. In this model, we take the 5G user plane function (UPF) into consideration to calculate access delay for the first time. Then, we devise a particle swarm optimization-based algorithm to optimize the profit. In the algorithm, we introduce a weight value$q$to guarantee the access delay and assign base stations properly. Moreover, a service-level agreement is adopted to balance the tradeoff between access delay and energy consumption. We take advantage of our 5G network emulator called mini5Gedge and data set from Shanghai Telecom to conduct massive experiments. The results show that our algorithm stands out in terms of achieving the highest profit. Yuanzhe Li 0001, Ao Zhou 0001, Xiao Ma 0009, Shangguang Wang |
IEEE Internet Things J. | 3 |
| 2022 | Service-Oriented Resource Allocation for Blockchain-Empowered Mobile Edge ComputingabstractIntegrating dense small cell (DSC) networks with mobile edge computing is employed by 5G to tackle the contradiction between the computation limitations of user equipment (UE) and the stringent latency requirement of services. This paper investigates the service-oriented edge resource allocation problem in DSC networks, determining where to deploy the service entity, how many service entities should be deployed at each edge cloud, and how to assign the UEs to service entities. The problem is challenging for the following three aspects: 1) Service entity deployment and UE assignment are highly coupled. 2) Due to the overlap of coverage regions of densely deployed small cells, the allocation mechanism of different base stations has mutual effects on the overall service performance. 3) Considering the limited resources of edge clouds, it is a thorny problem to encourage edge clouds to cache and share service startup images. We devote the following efforts to tackle the problem under these challenges. First, we explore blockchain’s decentralized, traceable, and secure characteristics, and propose a scheme to encourage image sharing in mobile edge computing. Second, we formulate the service-oriented edge resource allocation as mixed integer non-linear programming. Third, towards the target of reducing the computational complexity, we decouple UE assignment from service entity deployment and solve it through Gibbs sampling. Moreover, the power of Lyapunov optimization and convex optimization is incorporated to reduce the long-term power consumption and budget. Experiment results demonstrate the superiority of our approach over current notable solutions. Ao Zhou 0001, Xiao Ma 0009, Shangguang Wang |
IEEE J. Sel. Areas Commun. | 3 |
| 2022 | Resource-Aware Feature Extraction in Mobile Edge ComputingabstractMobile image recognition services, which provide people with image recognition services through the cameras of mobile devices, are revolutionizing our lives. However, most existing cloud/edge-based approaches suffer from two major limitations, (i) Low recognition accuracy and high network bandwidth pressure, and (ii) Not easy to extract features based on currently available resources of mobile devices. In this paper, we propose a resource-aware feature extraction framework for mobile image recognition services. The proposed framework consists of discriminative feature extraction (DFE) and NestDFE algorithms. The DFE algorithm can generate an extractor${{\mathbf E}}$to extract discriminative features from the image data set on the edge server and images on mobile devices. Thus, the proposed framework can achieve higher recognition accuracy and require mobile devices to upload less feature data to the edge server. The NestDFE algorithm generates a single multi-capacity extractor that acts as a series of sub-extractors and enables mobile devices to dynamically select sub-extractors. Experimental results show that the proposed framework improves recognition accuracy by about 23 percent and reduces network traffic by about 76 percent compared with existing approaches. Chuntao Ding, Ao Zhou 0001, Xiulong Liu 0001, Xiao Ma 0009, Shangguang Wang |
IEEE Trans. Mob. Comput. | 4 |
| 2022 | QoS Driven Task Offloading With Statistical Guarantee in Mobile Edge ComputingabstractIn mobile edge computing, popular mobile applications, such as augmented reality, usually offload their tasks to resource-rich edge servers. The user experience can be considerably affected when many mobile users compete for the limited communication and computation resources. The key technical challenge in task offloading is to guarantee the Quality of Service (QoS) for such applications. Existing work on task offloading focus on deterministic QoS (delay) guarantee, which means that tasks have to complete before the given deadline with 100 percent. However, it is impractical to impose a deterministic QoS guarantee for tasks due to the high dynamics of the wireless environment when offloading to edge servers. In this paper, we focus on task offloading with statistical QoS guarantee (tasks are allowed to complete before a given deadline with a probability above the given threshold), which can further save more energy by loosing the QoS requirement. Specially, we first propose a statistical computation model and a statistical transmission model to quantify the correlation between the statistical QoS guarantee and task offloading strategy. Then, we formulate the task offloading problem as a mixed integer non-Linear programming problem with the statistical delay constraint. We transform the statistical delay constraint into the constraints on CPU cycle numbers and the delay exponent respectively. We propose an algorithm to provide the statistical QoS guarantee for tasks using convex optimization theory and Gibbs sampling method. Experiment results show that the proposed algorithm outperforms the three baselines. Qing Li 0028, Shangguang Wang, Ao Zhou 0001, Xiao Ma 0009, Fangchun Yang, Alex X. Liu |
IEEE Trans. Mob. Comput. | 4 |
| 2022 | Providing Reliable Service for Parked-vehicle-assisted Mobile Edge ComputingabstractNowadays, a growing number of computation-intensive applications appear in our daily life. Those applications make the loads of both the core network and the mobile devices, in terms of energy and bandwidth, hugely increase. Offloading computation-intensive tasks to edge cloud is proposed to address this issue. Since edge clouds have limited computation resources compared with the remote cloud, they would get over-loaded because of the heavy computation burden. Parked-vehicle-assisted mobile edge computing becomes one of the promising solutions for this problem. However, several critical issues in parked-vehicle-assisted mobile edge computing would result in low reliable edge service. The open environment would bring about uncertainty, and the data privacy is hard to ensure. In addition, different from edge cloud, each parked vehicle only has limited parking duration and can leave unexpectedly for personal reasons. Moreover, edge cloud and vehicle adopt different execution models of computation and communication. The heterogeneous environment may result in negative effect on cooperativeness. Ignoring those issues can result in substantial performance degradation. To tackle this challenge and explore the benefits of parked-vehicle-assisted offloading, we study the task offloading and resource-allocation problem by fully considering the above issues. First, we propose a resource-management scheme to address the privacy issue. Second, we review the execution model of computation and communication in parked-vehicle-assisted computation offloading. Then, we formulate the problem into a mixed-integer nonlinear programming. The problem is hard to tackle due to its non-convex nature, which means that the time complexity of finding global optimal solution is unaffordable. Finally, we decompose the original problem into two sub-problems with lower complexity, and related algorithms are given to deal with the sub-problems. Simulation results demonstrate the effectiveness of the proposed solution. Ao Zhou 0001, Xiao Ma 0009, Siyi Gao, Shangguang Wang |
ACM Trans. Internet Techn. | 2 |
| 2021 | From cloud to edge: a first look at public edge platformsabstractPublic edge platforms have drawn increasing attention from both academia and industry. In this study, we perform a first-of-its-kind measurement study on a leading public edge platform that has been densely deployed in China. Based on this measurement, we quantitatively answer two critical yet unexplored questions. First, from end users' perspective, what is the performance of commodity edge platforms compared to cloud, in terms of the end-to-end network delay, throughput, and the application QoE. Second, from the edge service provider's perspective, how are the edge workloads different from cloud, in terms of their VM subscription, monetary cost, and resource usage. Our study quantitatively reveals the status quo of today's public edge platforms, and provides crucial insights towards developing and operating future edge services. Mengwei Xu 0001, Zhe Fu 0005, Xiao Ma 0009, Li Zhang 0133, Feng Qian 0001, Shangguang Wang, Xuanzhe Liu |
Internet Measurement Conference | 3 |
| 2021 | Joint Placement of UPF and Edge Server for 6G NetworkabstractThe emerging 6G network will make it possible for cybertwin, which relies deeply on the low latency and powerful computation provided by the edge network. To this end, the convergence of computing and network has been attached great importance. Most existing work study either placing edge servers or deploying user plane functions (UPFs), seldom considers the two processes jointly. In this article, we study how to minimize the latency with cost limitation by means of jointly deploying edge servers and UPFs in 6G scenario. We have shown that the problem is NP-hard. Then, we simplify the problem by analyzing the placement relationship between edge servers and UPFs and prune the solution space of the problem. To solve the problem effectively, a UPF and edge server placement algorithm is proposed. Massive experiments are conducted based on real-world data set and an edge core network emulator. The evaluation results show that our algorithm outperforms the benchmark algorithms. Yuanzhe Li 0001, Xiao Ma 0009, Mengwei Xu 0001, Ao Zhou 0001, Qibo Sun, Ning Zhang 0007, Shangguang Wang |
IEEE Internet Things J. | 2 |
| 2021 | Freshness-Aware Information Update and Computation Offloading in Mobile-Edge ComputingabstractMobile-edge computing is a promising computing paradigm with the advantages of reduced delay and relieved outsourcing traffic to the core network. In mobile-edge computing, reducing the computation offloading cost of mobile users and maintaining fresh information at edge nodes are two critical while conflicted objectives, as both consume the limited wireless bandwidth of edge nodes. Although extensive efforts have been devoted to optimizing computation offloading decisions and some works have investigated freshness-aware channel allocation issues recently, no prior works have considered the above conflict. This article is the first work to jointly optimize the channel allocation and computation offloading decisions, aiming at reducing the computation offloading cost within freshness requirements of sensors. We analyze the recursiveness of Age of Information (AoI) in analogy to the evolvement of a queue and formulate the problem as a nonlinear integer dynamic optimization problem. To overcome the challenges of AoI-computation cost tradeoff, AoI time dependency and high complexity caused by the heterogeneity of users, we propose an algorithm to solve the problem with reduced computation complexity. Specifically, we first transform the original problem into a static optimization problem in each time slot (which is NP-hard) based on Lyapunov optimization techniques. To reduce the computation complexity, we exploit the finite improvement property of potential games and further enforce centralized control to reduce the number of improvement iterations. Simulations have been conducted and the results demonstrate that the proposed algorithm shows good effectiveness and scalability. Xiao Ma 0009, Ao Zhou 0001, Qibo Sun, Shangguang Wang |
IEEE Internet Things J. | 1 |
| 2021 | Cost-Efficient Resource Provisioning for Dynamic Requests in Cloud Assisted Mobile Edge ComputingabstractMobile edge computing is emerging as a new computing paradigm that provides enhanced experience to mobile users via low latency connections and augmented computation capacity. As the amount of user requests is time-varying, while the computation capacity of edge hosts is limited, Cloud Assisted Mobile Edge (CAME) computing framework is introduced to improve the scalability of the edge platform. By outsourcing mobile requests to clouds with various types of instances, the CAME framework can accommodate dynamic mobile requests with diverse quality of service requirements. In order to provide guaranteed services at minimal system cost, the edge resource provisioning and cloud outsourcing of the CAME framework should be carefully designed in a cost-efficient manner. Specifically, two fundamental issues should be answered: (1) what is the optimal edge computation capacity configuration? and (2) what types of cloud instances should be tenanted and what is the amount of each type? To solve these issues, we formulate the resource provisioning in CAME framework as an optimization problem. By exploiting the piecewise convex property of this problem, the Optimal Resource Provisioning (ORP) algorithms with different instances are proposed, so as to optimize the computation capacity of edge hosts and meanwhile dynamically adjust the cloud tenancy strategy. The proposed algorithms are proved to be with polynomial computational complexity. To evaluate the performance of the ORP algorithms, extensive simulations and experiments are conducted based on both the widely-used traffic models and the Google cluster usage tracelogs, respectively. It is shown that the proposed ORP algorithms outperform the local-first and cloud-first benchmark algorithms in system flexibility and cost-efficiency. Xiao Ma 0009, Shangguang Wang, Shan Zhang 0001, Peng Yang 0004, Chuang Lin 0002, Xuemin Shen |
IEEE Trans. Cloud Comput. | 1 |
| 2021 | AoI-Delay Tradeoff in Mobile Edge Caching With Freshness-Aware Content RefreshingabstractMobile edge caching can effectively reduce service delay but may introduce information staleness, calling for timely content refreshing. However, content refreshing consumes additional transmission resources and may degrade the delay performance of mobile systems. In this work, we propose a freshness-aware refreshing scheme to balance the service delay and content freshness measured by Age of Information (AoI). Specifically, the cached content items will be refreshed to the up-to-date version upon user requests if the AoI exceeds a certain threshold (named as refreshing window). The average AoI and service delay are derived in closed forms approximately, which reveals an AoI-delay tradeoff relationship with respect to the refreshing window. In addition, the refreshing window is optimized to minimize the average delay while meeting the AoI requirements, and the results indicate to set a smaller refreshing window for the popular content items. Extensive simulations are conducted on the OMNeT++ platform to validate the analytical results. The results indicate that the proposed scheme can restrain frequent refreshing as the request arrival rate increases, whereby the average delay can be reduced by around 80% while maintaining the AoI below one second in heavily-loaded scenarios. Shan Zhang 0001, Liudi Wang, Hongbin Luo, Xiao Ma 0009, Sheng Zhou 0001 |
IEEE Trans. Wirel. Commun. | 4 |
| 2020 | Cognitive Service in Mobile Edge ComputingabstractCognitive services have revolutionized the way we live, work and interact with the world. In recent years, deep neural networks have become the mainstream approach in cognitive service, and mobile edge computing facilitates a variety of cognitive services for users by offloading computation tasks from resource-limited mobile devices to relatively wealthy edge servers. Combining the two to provide users with a higher quality of cognitive service is an issue worth researching. However, many related studies are not easy to provide fast responses because in these systems, edge servers are only used to pre-process data, and the cloud server is used to perform tasks. In this paper, we aim to study deploying deep neural network models on edge servers to provide fast services. However, a single edge server collects only a small amount of data, which results in low inference accuracy. To address this problem, we propose a cloud and edge collaboration framework. The key idea of the proposed framework is to use a cloud model to assist in training an edge model to improve the latter's inference accuracy and enable the latter to provide fast response and high-performance cognitive service. Experimental results demonstrate the effectiveness of our proposed framework. Chuntao Ding, Ao Zhou 0001, Xiao Ma 0009, Shangguang Wang |
ICWS | 3 |
| 2020 | Cooperative Service Caching and Workload Scheduling in Mobile Edge ComputingabstractMobile edge computing is beneficial for reducing service response time and core network traffic by pushing cloud functionalities to network edge. Equipped with storage and computation capacities, edge nodes can cache services of resource-intensive and delay-sensitive mobile applications and process the corresponding computation tasks without outsourcing to central clouds. However, the heterogeneity of edge resource capacities and mismatch of edge storage and computation capacities make it difficult to fully utilize both the storage and computation capacities in the absence of edge cooperation. To address this issue, we consider cooperation among edge nodes and investigate cooperative service caching and workload scheduling in mobile edge computing. This problem can be formulated as a mixed integer nonlinear programming problem, which has non-polynomial computation complexity. Addressing this problem faces challenges of sub-problem coupling, computation-communication tradeoff, and edge node heterogeneity. We develop an iterative algorithm named ICE to solve this problem. It is designed based on Gibbs sampling, which has provably near-optimal performance, and the idea of water filling, which has polynomial computation complexity. Simulation results demonstrate that our algorithm can jointly reduce the service response time and the outsourcing traffic, compared with the benchmark algorithms. Xiao Ma 0009, Ao Zhou 0001, Shan Zhang 0001, Shangguang Wang |
INFOCOM | 1 |
| 2020 | Dependency-Aware Task Scheduling in Vehicular Edge ComputingabstractVehicular edge computing (VEC) offers a new paradigm to improve vehicular services and augment the capabilities of vehicles. In this article, we study the problem of task scheduling in VEC, where multiple computation-intensive vehicular applications can be offloaded to roadside units (RSUs) and each application can be further divided into multiple tasks with task dependency. The tasks can be scheduled to different mobile-edge computing servers on RSUs for execution to minimize the average completion time of multiple applications. Considering the completion time constraint of each application and the processing dependency of multiple tasks belonging to the same application, we formulate the multiple tasks scheduling problem as an optimization problem that is NP-hard. To solve the optimization problem, we develop an efficient task scheduling algorithm. The basic idea is to prioritize multiple applications and prioritize multiple tasks so as to guarantee the completion time constraints of applications and the processing dependency requirements of tasks. The numerical results demonstrate that our proposed algorithm can significantly reduce the average completion time of multiple applications compared with benchmark algorithms. Yujiong Liu, Shangguang Wang, Qinglin Zhao, Ao Zhou 0001, Xiao Ma 0009, Fangchun Yang |
IEEE Internet Things J. | 6 |
| 2020 | Path Selection for Seamless Service Migration in Vehicular Edge ComputingabstractMobile-edge computing provisions computing and storage resources by deploying edge servers (ESs) at the edge of the network to support ultralow delay and high bandwidth services. To ensure QoS of latency-sensitive services in vehicular networks, service migration is required to migrate data of the ongoing services to the closest ES seamlessly when users move across different ESs. To achieve seamless service migration, path selection is proposed to obtain one or more paths (consisting of several switches and ESs) to transfer service data. We focus on the following problems about path selection: 1) where to implement path selection? 2) how to coordinate interests of mobile users (i.e., vehicles) and network providers since they have conflicting interests during path selection? and 3) how to ensure seamless service migration during the migration of vehicles? To address the above problems, this article investigates path selection for seamless service migration. We propose a path-selection algorithm to jointly optimize both interests of the network plane (i.e., the cost for network providers) and service plane (i.e., QoE of users). We first formulate it as a multiobjective optimization problem and further prove theoretically that the proposed algorithm can give aweakly Pareto-optimal solution. Moreover, to improve the scalability of the proposed algorithm, a distance-based filter strategy is designed to eliminate undesired switches in advance. We conduct experiments on two synthesized data sets and the results validate the effectiveness of the proposed algorithm. Jinliang Xu, Xiao Ma 0009, Ao Zhou 0001, Qiang Duan 0002, Shangguang Wang |
IEEE Internet Things J. | 2 |
| 2020 | Towards Service Composition Aware Virtual Machine Migration Approach in the CloudabstractThere is a growing trend for service providers to migrate their services from local clusters to the cloud data center. When there is no single service can satisfy the functionality requirement of the end user, existing services are combined together to fulfill the requirements. The data communication between component service hosting servers imposes a heavy burden on the data center network. In this article, we seek to reduce the data center network resource consumption by designing a novel service composition aware virtual machine migration approach. First, we formulate the problem as a multi-object integer non-linear(INLP) programming problem. The problem, which can be reduced into a well-known multi-object quadratic assignment problem, is proved to be NP-hard. Second, we simplify the multiple-objects INLP formulation into an equivalent, but much simplified single object ILP formulation. Then, we prove that the simplified formulation can also lead to the optimal solutions. Finally, optimization problem solvers, such as LPSolver, are employed to solve the problem. Experimental results in a large scale cloud data center demonstrate that our method significantly reduce the network resource consumption than other approaches. Ao Zhou 0001, Shangguang Wang, Xiao Ma 0009, Stephen S. Yau |
IEEE Trans. Serv. Comput. | 3 |
| 2019 | Age of Information and Delay Tradeoff with Freshness-Aware Mobile Edge Cache UpdateabstractMobile edge caching is an effective way to reduce the service delay of content delivery, where the popular contents can be pro-actively stored in proximity to users. In practice, the cached contents should be updated timely to avoid information staleness, in case that the information of a content changes with time and environment. However, cache update consumes additional transmission resources, which can degrade the delay performance. This work studies the fundamental tradeoff relationship between the content freshness (depicted by the age of information (AoI)) and service delay in mobile edge caching networks, and proposes a freshness-aware cache update scheme to achieve the AoI-delay balance. In specific, the base station will fetch the latest version of a content before delivery, if the AoI is larger than a certain threshold (i.e., update window size). The average AoI and service delay are derived in closed forms through approximated analysis of queueing systems, revealing a tradeoff relationship with respect to the update window size. Extensive simulations are conducted on the OMNeT++ platform, which validates the analytical results. Both the analytical and simulation results show that the proposed scheme can flexibly balance the average AoI and delay on demand, by tuning the update window size. Furthermore, the proposed scheme can also avoid frequent update in case of heavy content requests, whereby the AoI and delay are regulated by setting the appropriate update window size. Shan Zhang 0001, Liudi Wang, Hongbin Luo, Xiao Ma 0009, Sheng Zhou 0001 |
GLOBECOM | 4 |
| 2019 | Two-level task scheduling with multi-objectives in geo-distributed and large-scale SaaS cloud
Puheng Zhang, Xiao Ma 0009, Yanping Xiao, Chuang Lin 0002 |
World Wide Web | 2 |
| 2017 | Cost-Efficient Resource Provisioning in Cloud Assisted Mobile Edge ComputingabstractMobile edge computing (MEC) is emerging as an effective computing paradigm which alleviates the conflict between computation-intensive mobile applications and resource-constrained mobile devices. In this paper, a Cloud Assisted Mobile Edge computing (CAME) framework is adopted to enhance the adaptability of MEC to time-varying mobile requests. The resource provisioning problem is investigated to provide guaranteed quality of service (QoS) with minimum system cost. By exploiting the piecewise convexity of the problem, the Optimal Resource Provisioning (ORP) algorithm is developed, which determines the computation capacity at mobile edge and dynamically tunes the usage of cloud resources. Extensive simulations demonstrate that the ORP algorithm yields the minimum system cost, compared with the local-first and the cloud-first algorithms. In addition, the ORP algorithm can flexibly adapt to the time-varying mobile requests. Xiao Ma 0009, Shan Zhang 0001, Peng Yang 0004, Ning Zhang 0007, Chuang Lin 0002, Xuemin Shen |
GLOBECOM | 1 |
| 2017 | Cost-efficient workload scheduling in Cloud Assisted Mobile Edge ComputingabstractMobile edge computing is envisioned as a promising computing paradigm with the advantage of low latency. However, compared with conventional mobile cloud computing, mobile edge computing is constrained in computing capacity, especially under the scenario of dense population. In this paper, we propose a Cloud Assisted Mobile Edge computing (CAME) framework, in which cloud resources are leased to enhance the system computing capacity. To balance the tradeoff between system delay and cost, mobile workload scheduling and cloud outsourcing are further devised. Specifically, the system delay is analyzed by modeling the CAME system as a queuing network. In addition, an optimization problem is formulated to minimize the system delay and cost. The problem is proved to be convex, which can be solved by using the Karush-Kuhn-Tucker (KKT) conditions. Instead of directly solving the KKT conditions, which incurs exponential complexity, an algorithm with linear complexity is proposed by exploiting the linear property of constraints. Extensive simulations are conducted to evaluate the proposed algorithm. Compared with the fair ratio algorithm and the greedy algorithm, the proposed algorithm can reduce the system delay by up to 33% and 46%, respectively, at the same outsourcing cost. Furthermore, the simulation results demonstrate that the proposed algorithm can effectively deal with the challenge of heterogeneous mobile users and balance the tradeoff between computation delay and transmission overhead. Xiao Ma 0009, Shan Zhang 0001, Puheng Zhang, Chuang Lin 0002, Xuemin Shen |
IWQoS | 1 |
| 2017 | Long-Term Multi-objective Task Scheduling with Diff-Serv in Hybrid Clouds
Puheng Zhang, Chuang Lin 0002, Xiao Ma 0009 |
WISE (1) | 4 |
| 2016 | Monitoring-Based Task Scheduling in Large-Scale SaaS Cloud
Puheng Zhang, Chuang Lin 0002, Xiao Ma 0009, Fengyuan Ren |
ICSOC | 3 |
| 2015 | Joint Media Streaming Optimization of Energy and Rebuffering Time in Cellular NetworksabstractStreaming services are gaining popularity and have contributed a tremendous fraction of today's cellular network traffic. Both playback fluency and battery endurance are significant performance metrics for mobile streaming services. However, because of the unpredictable network condition and the loose coupling between upper layer streaming protocols and underlying network configurations, jointly optimizing rebuffering time and energy consumption for mobile streaming services remains a significant challenge. In this paper, we propose a novel framework that effectively addresses the above limitations and optimizes video transmission in cellular networks. We design two complementary algorithms, Rebuffering Time Minimization Algorithm (RTMA) and Energy Minimization Algorithm (EMA) in this framework, to achieve smoothed playback and energy-efficiency on demand over multi-user scenarios. Our algorithms integrate cross-layer parameters to schedule video delivery. Specifically, RTMA aims at achieving the minimum rebuffering time with limited energy and EMA tries to obtain the minimum energy consumption while meeting the rebuffering time constraint. Extensive simulation demonstrates that RTMA is able to reduce at least 68% rebuffering time and EMA can achieve more than 27% energy reduction compared with other state-of-the-art solutions. Zeqi Lai, Yong Cui 0001, Yayun Bao, Jiangchuan Liu, Yingchao Zhao 0001, Xiao Ma 0009 |
ICPP | 6 |
| 2015 | Game-theoretic Analysis of Computation Offloading for Cloudlet-based Mobile Cloud ComputingabstractMobile cloud computing (MC2) is emerging as a promising computing paradigm which helps alleviate the conflict between resource-constrained mobile devices and resource-consuming mobile applications through computation offloading. In this paper, we analyze the computation offloading problem in cloudlet-based mobile cloud computing. Different from most of the previous works which are either from the perspective of a single user or under the setting of a single wireless access point (AP), we research the computation offloading strategy of multiple users via multiple wireless APs. With the widespread deployment of WLAN, offloading via multiple wireless APs will obtain extensive application. Taking energy consumption and delay (including computing and transmission delay) into account, we present a game-theoretic analysis of the computation offloading problem while mimicking the selfish nature of the individuals. In the case of homogeneous mobile users, conditions of Nash equilibrium are analyzed, and an algorithm that admits a Nash equilibrium is proposed. For heterogeneous users, we prove the existence of Nash equilibrium by introducing the definition of exact potential game and design a distributed computation offloading algorithm to help mobile users choose proper offloading strategies. Numerical extensive simulations have been conducted and results demonstrate that the proposed algorithm can achieve desired system performance. Xiao Ma 0009, Chuang Lin 0002, Xudong Xiang, Congjie Chen |
MSWiM | 1 |
| 2015 | Cooperative Coverage Extension for Relay-Union NetworksabstractMulti-hop coverage extension can be utilized as a feasible approach to facilitating uncovered users to get Internet service in public area WLANs. In this paper we introduce a relay-union network (RUN), which refers to a public area WLAN in which users often wander in the same area and have the ability to provide data forwarding services for others. We develop a RUN framework to model the cost of providing forwarding services and the utility obtained by gaining services. The objective of the RUN is to maximize the total Quality of Cooperation (QoC) of users in the RUN. Two optimal bandwidth allocation schemes are proposed for both free and dynamic bandwidth demand models. To make our scheme more pragmatic, we then consider a more practical scenario in which the bandwidth capacity of the relays and the minimum demand of the clients are bounded. We prove that the problems under both the single relay and the multi-relay scenario are NP-hard. Three heuristic algorithms are proposed to deal with bandwidth allocation and relay-client association. We also propose a distributed signaling protocol and divide the centralized MRMC algorithm into three distributed ones to better adapt for real network environment. Finally, extensive simulations demonstrate that our RUN framework can significantly improve the efficiency of cooperation in the long term. Yong Cui 0001, Xiao Ma 0009, Xiuzhen Cheng, Minming Li, Jiangchuan Liu, Tianze Ma, Yihua Guo, Biao Chen 0002 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2014 | Policy-based flow control for multi-homed mobile terminals with IEEE 802.11u standard
Yong Cui 0001, Xiao Ma 0009, Jiangchuan Liu, Yuri Ismailov |
Comput. Commun. | 2 |
| 2013 | A Survey of Energy Efficient Wireless Transmission and Modeling in Mobile Cloud Computing
Yong Cui 0001, Xiao Ma 0009, Hongyi Wang 0004, Ivan Stojmenovic, Jiangchuan Liu |
Mob. Networks Appl. | 2 |