Zhi Zhou 0006

dblp:04/2090-6 · DBLP profile ↗
← Back
118ranked-venue papers
11as first author
82since 2021 · last 2026
0000-0002-0987-9344ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 70 · 7 first-author · 49 since 2021Systems, architecture and hardware · 31 · 2 first-author · 22 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 1 first-author · 5 since 2021Software engineering, systems software and programming languages · 5 · 5 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Security and privacy · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Cappuccino: Cost-Efficient Heterogeneous Lora Fine-Tuning Via Replica-Level Orchestration
Lingxuan Weng, Shiqi Cheng, Zhi Zhou 0006
ICDCS7
2026 Laser: Unlocking Layer-Level Scheduling for Efficient Multi-SLO LLM Serving
abstract
Engaging applications with diverse SLO requirements has become indispensable for production-scale LLM serving systems. However, existing systems rely on iteration-level scheduling, which enforces inflexible, unified execution across multi-SLO workloads, significantly constraining the serving efficiency.
Jianxiong Liao, Quanxing Dong, Yunkai Liang, Zhi Zhou 0006, Xu Chen 0004
PPoPP4
2026 Energy-Efficient and Dequantization-Free Quantization of LLMs: A Spiking Neural Network Approach to Salient Value Mitigation
Chenyu Wang 0004, Zhanglu Yan, Zhi Zhou 0006, Xu Chen 0004, Weng-Fai Wong
WWW3
2026 Adaptive load balance scheme for the distributed control plane in SDN
Yuwen Zhou, Bangbang Ren, Zhi Zhou 0006, Xu Chen 0004, Zhiguang Chen 0001, Deke Guo
Frontiers Comput. Sci.4
2026 Duba: Cost-Efficient Serverless Cloud-Edge Collaborative Machine Learning Serving with Dual-Batching
Jianxiong Liao, Zhi Zhou 0006, Fei Xu 0009
J. Comput. Sci. Technol.3
2026 AIGC-Enhanced Federated Learning: Addressing Data Scarcity in Preference-Based Scenarios
Chenyu Wang 0004, Zhi Zhou 0006, Zixin Xu, Shaoquan Wang, Xu Chen 0004
IEEE Trans. Mob. Comput.2
2026 MicroEdge: An Online Optimization Framework for Cost-Efficient Microservice Orchestration in Edge Native Applications
abstract
The rapid proliferation of edge computing infrastructure has significantly accelerated the adoption of edge-native applications, ranging from autonomous vehicles to augmented reality and real-time analytics. Microservice, renowned for its lightweight, loosely coupled, and modular architecture, has emerged as the de-facto standard for developing edge native applications. However, the resource scarcity and heterogeneity, coupled with request dynamics in edge environments, pose substantial challenges for effective microservice orchestration. To address these challenges, we propose MicroEdge, an online optimization framework designed for cost-efficient microservice orchestration in edge-native environments. MicroEdge employs a multi-level optimization approach by strategically coordinating four key dimensions: microservice placement, layer placement, layer pulling, and user request scheduling. The framework confronts two fundamental challenges in solving this joint optimization problem: (1) the time-coupled nature of long-term holistic cost minimization, and (2) the NP-hardness of the underlying problem. MicroEdge tackles these dual challenges by integrating a regularization method for online algorithm design and a dependent rounding technique for approximation algorithm design. Both rigorous theoretical analysis and extensive simulations driven by realistic Alibaba microservice workload traces validate the efficacy of MicroEdge.
Weihan Zeng, Kongyange Zhao, Jianxiong Liao, Zhi Zhou 0006, Deke Guo, Xu Chen 0004
IEEE Trans. Mob. Comput.4
2026 Online Location Planning for AI-Defined Vehicles: Optimizing Joint Tasks of Order Serving and Spatio-Temporal Heterogeneous Model Fine-Tuning
abstract
Advances in artificial intelligence (AI) including foundation models (FMs), are increasingly transforming human society, with smart city driving the evolution of urban living. Meanwhile, vehicle crowdsensing (VCS) has emerged as a key enabler, leveraging vehicles' mobility and sensor-equipped capabilities. In particular, ride-hailing vehicles can effectively facilitate flexible data collection and contribute towards urban intelligence, despite resource limitations. Therefore, this work explores a promising scenario, where edge-assisted vehicles perform joint tasks of order serving and the emerging foundation model finetuning using various urban data. However, integrating the VCS AI task with the conventional order serving task is challenging, due to their inconsistent spatio-temporal characteristics: (i) The distributions of ride orders and data point-of-interests (PoIs) may not coincide in geography, both following a priori unknown patterns; (ii) they have distinct forms of temporal effects, i.e., prolonged waiting makes orders become instantly invalid while data with increased staleness gradually reduces its utility for model fine-tuning. To overcome these obstacles, we propose an online framework based on multi-agent reinforcement learning (MARL) with careful augmentation. A new quality-of-service (QoS) metric is designed to characterize and balance the utility of the two joint tasks, under the effects of varying data volumes and staleness. We also integrate graph neural networks (GNNs) with MARL to enhance state representations, capturing graph-structured, time-varying dependencies among vehicles and across locations. Extensive experiments on our testbed simulator, utilizing various real-world foundation model fine-tuning tasks and the New York City Taxi ride order dataset, demonstrate the advantage of our proposed method.
Bokeng Zheng, Bo Rao, Tianxiang Zhu, Chee-Wei Tan 0001, Jingpu Duan, Zhi Zhou 0006, Xu Chen 0004, Xiaoxi Zhang 0001
IEEE Trans. Mob. Comput.6
2026 OSGS: A Framework for Online Scheduling of Satellite-Ground Collaborative Inference With Space Edge Computing
Kongyange Zhao, Yuanming Wang, Zhi Zhou 0006, Ruiting Zhou, Xiaoxi Zhang 0001, Xu Chen 0004, Dechao Ran, Fei Zhang 0005, Lu Cao 0001
IEEE Trans. Serv. Comput.3
2025 Embracing Imbalance: Dynamic Load Shifting among Microservice Containers in Shared Clusters
abstract
In a unified resource scheduling architecture, containers within the same microservice often encounter temporal and spatial performance imbalance when deployed in large-scale shared clusters. As a result, the commonly employed load-balancing approach often leads to substantial resource wastage as applications are frequently over-provisioned to meet service level agreements (SLAs).
Shutian Luo, Jianxiong Liao, Chenyu Lin, Huanle Xu, Zhi Zhou 0006, Cheng-Zhong Xu 0001
ASPLOS (2)5
2025 Online Context Caching for Distributed Large Language Models Serving
Bin Gao 0013, Zhuomin He, Yizhen Yao, Zhanzhi Lou Lou, Zhi Zhou 0006, Weng-Fai Wong
INFOCOM5
2025 AdaRAG: Adaptive Optimization for Retrieval Augmented Generation with Multilevel Retrievers at the Edge
Tao Ouyang, Guihang Hong, Kongyange Zhao, Zhi Zhou 0006, Weigang Wu, Zhaobiao Lv, Xu Chen 0004
INFOCOM4
2025 Towards Federated Inference: An Online Model Ensemble Framework for Cooperative Edge AI
Zhi Zhou 0006, Mengke Huang, Tao Ouyang, Fangming Liu, Xu Chen 0004
INFOCOM1
2025 Espresso: Cost-Efficient Large Model Training by Exploiting GPU Heterogeneity in the Cloud
Qiannan Zhou, Fei Xu 0009, Lingxuan Weng, Ruixing Li, Li Chen 0019, Zhi Zhou 0006, Fangming Liu
INFOCOM7
2025 Poster: A Unified Framework for Simultaneous Video Analytics and Streaming on UAVs
abstract
With the increasing adoption of unmanned aerial vehicles (UAVs) in critical applications such as infrastructure inspection and emergency response, efficient on-site recognition via live video analytics and streaming has become essential. However, the inherent resource limitation poses significant challenges for performing simultaneous and real-time video analytics and streaming on UAVs. To address this issue, we propose a unified framework that orchestrate the Neural Processing Unit (NPU) and Graph Processing Unit (GPU) of the Systems-on-Chip (SoC) processor to accelerate and carefully schedule the pipeline of video analytics and streaming on UAVs. Additionally, our system incorporates frame interpolation to enable real-time streaming of video analytics results, providing immediate visual feedback to on-site operators. Empirical results on a commercial UAV equipped with Snapdragon 865 SoC platform show that our system reduces per-frame inference latency from 163ms (GPU) to 63ms (NPU), achieving a 2.6× speedup. Combined with optimized pre-processing and frame interpolation, our system increases effective streaming throughput from 2 to 30 FPS, enabling smooth and simultaneous real-time video analytics and streaming.
Zhi Zhou 0006, Rouyi Wang, Xu Chen 0004
MobiCom2
2025 Demo: WasmSD-Edge: A Lightweight Edge Stable Diffusion Image Generation Framework Based on WebAssembly
abstract
The growing demand for deploying Artificial Intelligence Generated Content (AIGC) models like Stable Diffusion on resource-constrained edge devices challenges balancing quality, lightweight implementation, and portability. The emergence of WebAssembly (WASM) offers a compactcross-platform, and isolated runtime environment, making it a promising solution for efficient edge AIGC inference. However, current WASM based AI inference solutions are restricted to text interactions, offering limited support for image generation. To solve the challenges, we propose WebAssembly-Rust based WasmSD-Edge, a lightweight, edge-oriented AI image generation framework for high performance on-device Stable Diffusion inference on various edge devices. WasmSD-Edge employs a plugin-based architecture by integrating stable-diffusion.cpp as a WASM backend plugin for the WasmEdge runtime. It exposes a set of WebAssembly System Interfaces (WASI) to support text-to-image, image-to-image, and model convertion. Additionally, a Rust Crate SDK further enables developers to parametrically control inference process and output generation. To evaluate usability and portability of WasmSD-Edge on heterogeneous devices, we deployed it on heterogeneous devices. It achieves high inference speed and image quality with low resource consumption, offering a practical and efficient solution for deploying edge AIGC workflow. The implementation has been merged into WasmEdge — one of the largest WASM community, and source code are available at: https://github.com/WasmEdge/wasmedge-stable-diffusion.
Rouyi Wang, Zhi Zhou 0006, Xu Chen 0004
MobiCom2
2025 CoEdge-RAG: Optimizing Hierarchical Scheduling for Retrieval-Augmented LLMs in Collaborative Edge Computing
abstract
Motivated by the imperative for real-time responsiveness and data privacy preservation, large language models (LLMs) are increasingly deployed on resource-constrained edge devices to enable localized inference. To improve output quality, retrieval-augmented generation (RAG) is an efficient technique that seamlessly integrates local data into LLMs. However, existing edge computing paradigms primarily focus on single-node optimization, neglecting opportunities to holistically exploit distributed data and heterogeneous resources through cross-node collaboration. To bridge this gap, we propose CoEdge-RAG, a hierarchical scheduling framework for retrieval-augmented LLMs in collaborative edge computing. In general, privacy constraints preclude accurate a priori acquisition of heterogeneous data distributions across edge nodes, directly impeding RAG performance optimization. Thus, we first design an online query identification mechanism using proximal policy optimization (PPO), which autonomously infers query semantics and establishes cross-domain knowledge associations in an online manner. Second, we devise a dynamic inter-node scheduling strategy that balances workloads across heterogeneous edge nodes by synergizing historical performance analytics with real-time resource thresholds. Third, we develop an intra-node scheduler based on online convex optimization, adaptively allocating query processing ratios and memory resources to optimize the latency-quality trade-off under fluctuating assigned loads. Comprehensive evaluations across diverse QA benchmarks demonstrate that our proposed method significantly boosts the performance of collaborative retrieval-augmented LLMs, achieving performance gains of 4.23 % to 91.39% over baseline methods across all tasks.
Guihang Hong, Tao Ouyang, Kongyange Zhao, Zhi Zhou 0006, Xu Chen 0004
RTSS4
2025 Hetis: Serving LLMs in Heterogeneous GPU Clusters with Fine-grained and Dynamic Parallelism
abstract
The significant resource demands in LLM serving prompts production clusters to fully utilize heterogeneous hardware by partitioning LLM models across a mix of high-end and low-end GPUs. However, existing parallelization approaches often struggle to scale efficiently in heterogeneous environments due to their coarse-grained and static parallelization strategies.
Zizhao Mo, Jianxiong Liao, Huanle Xu, Zhi Zhou 0006, Cheng-Zhong Xu 0001
SC4
2025 Opara: Exploiting Operator Parallelism for Expediting DNN Inference on GPUs
abstract
GPUs have become thedefactohardware devices for accelerating Deep Neural Network (DNN) inference workloads. However, the conventionalsequential execution mode of DNN operatorsin mainstream deep learning frameworks cannot fully utilize GPU resources, even with the operator fusion enabled, due to the increasing complexity of model structures and a greater diversity of operators. Moreover, theinadequate operator launch orderin parallelized execution scenarios can lead to GPU resource wastage and unexpected performance interference among operators. In this paper, we proposeOpara, a resource- and interference-aware DNNOperatorparallel scheduling framework to accelerate DNN inference on GPUs. Specifically,Oparafirst employsCUDA StreamsandCUDA Graphtoparallelizethe execution of multiple operators automatically. To further expedite DNN inference,Oparaleverages the resource demands of operators to judiciously adjust the operator launch order on GPUs, overlapping the execution of compute-intensive and memory-intensive operators. We implement and open source a prototype ofOparabased on PyTorch in anon-intrusivemanner. Extensive prototype experiments with representative DNN and Transformer-based models demonstrate thatOparaoutperforms the default sequentialCUDA Graphin PyTorch and the state-of-the-art operator parallelism systems by up to$1.68\boldsymbol{\times}$and$1.29\boldsymbol{\times}$, respectively, yet with acceptable runtime overhead.
Aodong Chen, Fei Xu 0009, Li Han 0001, Li Chen 0019, Zhi Zhou 0006, Fangming Liu
IEEE Trans. Computers6
2025 Delay-Sensitive Task Offloading With Edge Caching Through Martingale-Based Deep Reinforcement Learning
abstract
In the forthcoming era of 6G networks, delay-sensitive applications for Internet of Things (IoT) are poised to become the prevailing services with ultra-reliable and low-latency (URLLC) requirements. Unlike traditional video caching, IoT-based edge caching faces unique challenges due to diverse data types, update frequencies, and computational needs, requiring integrated storage and computational resource management. To support the more stringent requirements for these innovative applications, mobile edge computing (MEC) is introduced to enhance the service reliability of delay-sensitive applications in the 6G era. However, task offloading, as an indispensable procedure in MEC, would encounter many challenges, such as network jitter and resource insufficiency, possibly leading to unpredictable queuing delays and other negative issues. To ensure reliable services in a dynamical MEC environment, the caching-enabled MEC network has emerged as a novel architecture, placing computing and storage resources in the edge network. In this paper, we investigate the caching-enabled MEC to support reliable task offloading for delay-sensitive applications, with a focus on IoT scenarios. In our system model, we formulate the task process as a two-hop tandem queuing system with limited capacity, including task transmission and computation queues. The Martingale theory is leveraged to analyze the delay violation probability in this system, demonstrating how the offloading and caching decisions affect the end-to-end (E2E) delay. Besides, task offloading and resource allocation policies are integrated to reduce high system costs, including energy consumption and cache resource rental costs. Based on the delay analysis of martingale theory, we propose an advanced deep reinforcement learning (DRL) algorithm called Dynamic Request Aware Soft Actor-Critic (DRA-SAC) algorithm to achieve minimal system costs by obtaining the optimal task offloading and resource allocation policies, including caching and computation resources. We conduct some illustrative studies to evaluate the proposed scheme. The algorithm we have put forward outperforms benchmark algorithms regarding both cache hit ratio and system cost.
Chongwu Dong, Zhi Zhou 0006, Xu Chen 0004, Zhihong Tian 0001, Wushao Wen
IEEE Trans. Mob. Comput.3
2025 Efficient Coordination of Federated Learning and Inference Offloading at the Edge: A Proactive Optimization Paradigm
abstract
Benefiting from hardware upgrades and deep learning techniques, more and more end devices can independently support a variety of intelligent applications. Further powered by edge computing technologies, the end-edge collaboration paradigm becomes one mainstream approach for achieving advanced edge intelligence (EI). To fully exploit the system resources, it is desirable to coordinate diverse EI services efficiently. Thus, we present a novel framework to jointly optimize the cost-performance trade-off for two distinct but typical EI services, where end devices simultaneously perform federated learning (FL) model training and conduct model inference with the assistance of edge offloading. However, balancing the long-term cost-performance trade-off is highly non-trivial, especially in the absence of knowledge of future system dynamics. Moreover, the capacity heterogeneity further increases the difficulty of service coordination among resource-limited end devices. To overcome these challenges, we first analyze the optimality of inference offloading decisions with and without FL model training and quantify their mutual effects due to local resource contention. By incorporating the loss estimation of FL training model, we then propose a novel proactive policy with theoretical guarantees, which proactively controls the stopping of FL training procedure to balance well the trade-offs between FL model performance and resource costs while fulfilling the inference performance requirements. Extensive results show the efficiency and robustness of our proposed algorithm for EI service coordination in dynamic end-edge collaboration scenarios.
Ke Luo 0001, Kongyange Zhao, Tao Ouyang, Xiaoxi Zhang 0001, Zhi Zhou 0006, Xu Chen 0004
IEEE Trans. Mob. Comput.5
2025 Quality-of-Service Aware LLM Routing for Edge Computing With Multiple Experts
abstract
Large Language Models (LLMs) have demonstrated remarkable capabilities, leading to a significant increase in user demand for LLM services. However, cloud-based LLM services often suffer from high latency, unstable responsiveness, and privacy concerns. Therefore, multiple LLMs are usually deployed at the network edge to boost real-time responsiveness and protect data privacy, particularly for many emerging smart mobile and IoT applications. Given the varying response quality and latency of LLM services, a critical issue is how to route user requests from mobile and IoT devices to an appropriate LLM service (i.e., edge LLM expert) to ensure acceptable quality-of-service (QoS). Existing routing algorithms fail to simultaneously address the heterogeneity of LLM services, the interference among requests, and the dynamic workloads necessary for maintaining long-term stable QoS. To meet these challenges, in this paper we propose a novel deep reinforcement learning (DRL)-based QoS-aware LLM routing framework for sustained high-quality LLM services. Due to the dynamic nature of the global state, we propose a dynamic state abstraction technique to compactly represent global state features with a heterogeneous graph attention network (HAN). Additionally, we introduce an action impact estimator and a tailored reward function to guide the DRL agent in maximizing QoS and preventing latency violations. Extensive experiments on both Poisson and real-world workloads demonstrate that our proposed algorithm significantly improves average QoS and computing resource efficiency compared to existing baselines.
Qiong Wu 0009, Zhiying Feng, Zhi Zhou 0006, Deke Guo, Xu Chen 0004
IEEE Trans. Mob. Comput.4
2025 Dynamic Edge-Centric Resource Provisioning for Online and Offline Services Co-Location via Reactive and Predictive Approaches
abstract
Due to the penetration of edge computing, a wide variety of workloads are sunk down to the network edge to alleviate huge pressure of the cloud. With the presence of high input workload dynamics and intensive edge resource contention, it is highly non-trivial for an edge proxy to optimize the scheduling of heterogeneous services with diverse QoS requirements. In general, online services should be quickly completed in a quite stable running environment to meet their tight latency constraint, while offline services can be processed loosely for their elastic soft deadlines. To well coordinate such services at the resource-limited edge cluster, in this paper, we study an edge-centric resource provisioning optimization for dynamic online and offline services co-location, where the proxy seeks to maximize timely online service performances while maintaining satisfactory long-term offline service performances. However, intricate hybrid couplings for provisioning decisions arise due to heterogeneous constraints of the co-located services and their different time-scale performances. We hence first propose a reactive provisioning approach without requiring a prior knowledge of future system dynamics, which leverages a Lagrange relaxation for devising constraint-aware stochastic subgradient algorithm to deal with the challenge of hybrid couplings. To further boost the performance by integrating powerful machine learning techniques, we then advocate a predictive provisioning approach, where future request arrivals can be estimated accurately. To align with practical deployments, we incorporate a tunable prediction window mechanism, which well balances the potential improvement and degradation of online performance in imperfect prediction scenarios. With rigorous theoretical analysis and extensive trace-driven evaluations, we show the superior performance of our proposed algorithms for online and offline services co-location at the edge.
Tao Ouyang, Kongyange Zhao, Guihang Hong, Xiaoxi Zhang 0001, Zhi Zhou 0006, Xu Chen 0004
IEEE Trans. Netw.5
2025 MEC-Enabled Task Replication With Resource Allocation for Reliability-Sensitive Services in 5G mMTC Networks
abstract
The increasing demand for connectivity in 5G networks has led to a focus on massive machine-type communication (mMTC) in mobile edge computing (MEC) for IoTs. However, the proliferation of IoT devices has resulted in densely deployed networks and led to a high volume of task offloading to the same edge servers simultaneously. As a consequence, mMTC applications may experience service congestion, negatively impacting service reliability. To enhance the service reliability of latency-sensitive applications, task replication with resource allocation is proposed in MEC, in which a task can be sent simultaneously to multiple computing nodes. Task replication can reduce task latency and improve service reliability at the cost of consuming more computation resources. However, unconstrained task replication may result in too many uploading links, leading to severe costs in network operation. To handle the above challenge, we propose a constrained stochastic optimization problem by task replication with wireless resource block (RB) allocation and edge server queue management. To ensure queue stability while minimizing cost, we design one strategy based on the Lyapunov optimization framework. Accordingly, we further model RB allocation as a mean-field game (MFG) due to the intensive coupling of the RB pool for massive users. Tractable partial differential equations are used to analyze MFG equilibrium, and we derive the optimal edge server queue management based on a given task replication strategy and RB allocation scheme. Our theoretical analysis demonstrates that our algorithm closely approaches the optimal overall costs within a small gap, and simulation results show that our strategy generates a significantly lower cumulative cost than other alternative strategies.
Rui Huang 0016, Wushao Wen, Zhi Zhou 0006, Chongwu Dong, Xu Chen 0004
IEEE Trans. Serv. Comput.3
2024 SAFE: Intelligent Online Scheduling for Collaborative DNN Inference in Vehicular Network
abstract
Recent years have witnessed a widespread use of deep neural networks (DNNs) in providing various intelligent services, and vehicular networks are no exception. Given the limited computing capabilities of vehicles, collaborative vehicle-edge DNN inference has emerged as a viable alternative. This approach employs DNN partitioning, where a part of DNN is computed on vehicles, and the other part on the edge, e.g., roadside unit (RSU), aiming to enhance the inference accuracy and reduce the inference latency. In this setting, deriving an optimal DNN partitioning scheme becomes critical, yet challenging given the constant movement of vehicles and the highly dynamic wireless connections. Furthermore, vehicles may move out of the signal coverage of an RSU, making it difficult to receive the inference results. To this end, we propose a two-stage intelligent scheduling framework named Soft Actor-critic for discrete actions (SAC-D) based collaborative DNN inference FramEwork (SAFE). SAFE engages multiple RSUs to assist vehicles in completing inference tasks sequentially and ensuring reliable data transmission. It can learn the dynamic vehicular network and make scheduling decisions to minimize the overall latency of vehicle inference tasks. Extensive experimental results show that SAFE can reduce up to 80% of the overall latency with a lower failure rate, compared to four baselines.
Ruiting Zhou, Ziyi Han, Zhi Zhou 0006, Wei Wang 0030
CSCWD4
2024 Bridging the Data Gap in Federated Preference Learning with AIGC
abstract
Federated learning (FL), a decentralized machine learning approach, enables privacy-preserving and collaborative model training without centralizing sensitive data. It has been successfully applied in various domains, including e-commerce, healthcare, and finance. However, existing FL schemes often fail to address personalized task requirements, such as prior-itizing the accuracy of specific classes within a dataset. The recent surge in Artificial Intelligence Generated Content (AIGC) offers potential to meet these personalized requirements by augmenting the training data of specific classes with generative models. Nevertheless, integrating generative models with FL introduces challenges, such as non-compliant data, disorganized distributions, and limited computing power on edge devices. To address these challenges, we propose AIGC-augmented Federated Preference Learning (FPL), which focuses on training specific data classes, referred to as preference classes (PCs). To improve the quality of AI -generated data, we implement strategies such as pre-training and fine-tuning across various datasets. Additionally, we enhance FL efficiency through a client selection strategy that matches generated data tasks with suitable clients and an AIGC data distribution strategy that optimally allocates data where it is most needed. We validate the feasibility and effectiveness of AIGC-augmented FPL by conducting experiments on the MNIST and CIFAR-10 datasets from various perspectives.
Chenyu Wang 0004, Zhi Zhou 0006, Xiaoxi Zhang 0001, Xu Chen 0004
ICDCS2
2024 COUPLE: Orchestrating Video Analytics on Heterogeneous Mobile Processors
abstract
Video analytics is considered the killer application of edge computing and has been successfully deployed across diverse domains. Yet, executing video analytics on mobile devices presents notable challenges owing to the considerable computational demands and frame rate requirements of DNN models. Current mobile inference frameworks often concentrate on enhancing model inference performance on the CPU or GPU, overlooking the potential of the Digital Signal Processor (DSP) – an emerging heterogeneous processor increasingly integrated into modern mobile processors. In this paper, we introduce COUPLE, an orchestration framework for video analytics on heterogeneous mobile processors, with the goal of optimizing real-time video analysis through the collaboration of CPU, GPU and DSP. To tackle the accuracy loss of DSP inference, we introduce the Anchor Frame Calibration mechanism, utilizing high-precision GPU inference results and frame similarities to mitigate accuracy loss on the DSP. Additionally, we design a lightweight progressive scheduler to distribute video frames to GPU and DSP, maximizing inference Average Precision (AP) under performance (i.e., frame rate) and power constraints. COUPLE has been implemented on the Qualcomm's Snapdragon 888 mobile SoC, extensive evaluation results demonstrate its efficacy in imnroving the inference performance and accuracy.
Hao Bao, Zhi Zhou 0006, Fei Xu 0009, Xu Chen 0004
ICDE2
2024 IMI: In-memory Multi-job Inference Acceleration for Large Language Models
abstract
Large Language Models (LLMs) are increasingly used in various applications but are computationally complex and energy-consuming due to the high volume of off-chip memory accesses. Processing-in-Memory (PIM) has emerged as a potential solution for efficient inference. However, existing PIM accelerators designed for deep neural networks (DNNs) aren’t suitable for LLMs because of differences in operations, input sizes, and job completion times. This leads to performance issues like head-of-line blocking where an earlier job can monopolize resources at the expense of later jobs, and low resource utilization. To improve efficiency, a time-multiplex solution and job colocation accelerator could be beneficial. However, facilitating multi-job execution with in-memory acceleration is challenging due to limitations in memristor architecture, inefficiency of non-stationary weight programming, difficulty in dynamic partitioning of hardware resources, and complexity in dynamic job scheduling. This work proposes a PIM-based LLM accelerator to enable the concurrent LLM inference job execution. The experiment shows that IMI can significantly improve resource utilization as well as the rate of satisfying the service level requirements of jobs.
Bin Gao 0013, Zhehui Wang, Zhuomin He, Tao Luo 0014, Weng-Fai Wong, Zhi Zhou 0006
ICPP6
2024 Online Resource Allocation for Edge Intelligence with Colocated Model Retraining and Inference
abstract
With edge intelligence, AI models are increasingly push to the edge to serve ubiquitous users. However, due to the drift of model, data, and task, AI model deployed at the edge suffers from degraded accuracy in the inference serving phase. Model retraining handles such drifts by periodically retraining the model with newly arrived data. When colocating model retraining and model inference serving for the same model on resource-limited edge servers, a fundamental challenge arises in balancing the resource allocation for model retraining and inference, aiming to maximize long-term inference accuracy. This problem is particularly difficult due to the underlying mathematical formulation being time-coupled, non-convex, and NP-hard. To address these challenges, we introduce a lightweight and explainable online approximation algorithm, named ORRIC, designed to optimize resource allocation for adaptively balancing the accuracy of model training and inference. The competitive ratio of ORRIC outperforms that of the traditional Inference-Only paradigm, especially when data drift persists for a sufficiently lengthy time. This highlights the advantages and applicable scenarios of colocating model retraining and inference. Notably, ORRIC can be translated into several heuristic algorithms for different resource environments. Experiments conducted in real scenarios validate the effectiveness of ORRIC.
Huaiguang Cai, Zhi Zhou 0006, Qianyi Huang
INFOCOM2
2024 HarmonyBatch: Batching multi-SLO DNN Inference with Heterogeneous Serverless Functions
abstract
Deep Neural Network (DNN) inference on serverless functions is gaining prominence due to its potential for substantial budget savings. Existing works on serverless DNN inference solely optimize batching requests from one application with a single Service Level Objective (SLO) on CPU functions. However, production serverless DNN inference traces indicate that the request arrival rate of applications is surprisingly low, which inevitably causes a long batching time and SLO violations. Hence, there is an urgent need for batching multiple DNN inference requests with diverse SLOs (i.e., multi-SLO DNN inference) in serverless platforms. Moreover, the potential performance and cost benefits of deploying heterogeneous (i.e., CPU and GPU) functions for DNN inference have received scant attention.In this paper, we present HarmonyBatch, a cost-efficient resource provisioning framework designed to achieve predictable performance for multi-SLO DNN inference with heterogeneous serverless functions. Specifically, we construct an analytical performance and cost model of DNN inference on both CPU and GPU functions, by explicitly considering the GPU time-slicing scheduling mechanism and request arrival rate distribution. Based on such a model, we devise a two-stage merging strategy in HarmonyBatch to judiciously batch the multi-SLO DNN inference requests into application groups. It aims to minimize the budget of function provisioning for each application group while guaranteeing diverse performance SLOs of inference applications. We have implemented a prototype of HarmonyBatch on Alibaba Cloud Function Compute. Extensive prototype experiments with representative DNN inference workloads demonstrate that HarmonyBatch can provide predictable performance to serverless DNN inference workloads while reducing the monetary cost by up to 82.9% compared to the state-of-the-art methods.
Fei Xu 0009, Yikun Gu, Li Chen 0019, Fangming Liu, Zhi Zhou 0006
IWQoS6
2024 Can You Do Both? Balancing Order Serving and Crowdsensing for Ride-Hailing Vehicles
abstract
Given the high mobility and sensor-carrying capability, vehicle crowdsensing (VCS) has become a significant part of urban crowdsensing tasks in the development of smart cities. Ride-hailing vehicles, which are widely distributed in cities, can be a powerful tool for carrying out VCS. However, dispatching the vehicles to jointly benefit VCS and order serving is challenging, as the goals of these two tasks may not be consistent or even conflict. The distribution of ride orders and the distribution of point-of-interests (PoIs) may not coincide in time and geography. In addition, these orders and data PoIs have distinct forms of timeliness: prolonged waiting makes orders invalid and data with a larger age-of-information (AoI) has lower utility. We propose an online framework by extending multi-agent reinforcement learning (MARL) with careful augmentation to optimize the profit of order-serving and the data utility of crowdsensing. A new quality-of-service (QoS) metric is designed to characterize the utility of the two joint tasks, and formal mathematical modeling drives our MARL design. In particular, we integrated graph neural networks (GNN) to enhance state representations and capture the graph-structured dependencies among vehicles. We developed a simulator and conducted extensive experiments utilizing the New York City Taxi dataset. Experimental results demonstrate the advantage of our method in QoS improvement.
Bo Rao, Xiaoxi Zhang 0001, Tianxiang Zhu, Yufei You, Jingpu Duan, Zhi Zhou 0006, Xu Chen 0004
IWQoS7
2024 MEGA: Mesh-Aligned 3DGS Towards Geometry-Preserving Online Reconstruction
Ke Luo 0001, Shengyuan Ye, Tao Ouyang, Zhi Zhou 0006
NPC (1)4
2024 Computing Power Networking Meets Blockchain: A Reputation-Enhanced Trading Framework for Decentralized IoT Cloud Services
abstract
Computing Power Networking (CPN) represents a transformative paradigm in distributed computing, harnessing the collective capabilities of edge servers dispersed across diverse geographical locations. CPN’s core strengths lie in its ability to accelerate data processing, diminish latency, and scale efficiently, rendering it particularly apt for real-time applications and the Internet of Things. When coupled with blockchain technology, CPN extends its potential by facilitating secure and transparent allocation and trading of computing resources, bolstering data integrity and reliability. However, current research at the intersection of CPN and blockchain primarily focuses on framework development and technology integration, often overlooking the challenge of delivering dependable computing services, especially in the presence of potentially unreliable nodes. To tackle this issue, we introduce a reputation-enhanced resource trading framework, designed to ensure equitable and trustworthy computing power transactions. We establish a decentralized reputation model, capable of accurately assessing node behavior over extended periods. Additionally, we present three optimization mechanisms for reputation updates, accounting for transaction history, quality of service, and transaction amount. Furthermore, our work introduces a reputation-enhanced consensus mechanism within the trading system, strategically employing incentives to motivate participants to deliver high-quality services, thereby increasing their rewards. Simultaneously, it effectively mitigates wealth inequality among resource providers of varying sizes. To validate our approach, we develop a prototype system and conduct performance evaluations, which affirm the superiority of our system in enhancing reputation and delivering robust economic features.
Li Lin 0001, Jiapeng Wu, Zhi Zhou 0006, Jin Zhao 0003, Peng Li 0017, Jinbo Xiong
IEEE Internet Things J.3
2024 Learning With Side Information: Elastic Multi-Resource Control for the Open RAN
abstract
The open radio access network (O-RAN) architecture provides enhanced opportunities for integrating machine learning in 5G/6G resource management by decomposing RAN functionalities. Yet, generic learning mechanisms either do not fully exploit the disaggregated non-real-time and near-real-time RAN controllers or ignore the potential elasticity of application demands, another degree of freedom in managing RAN resources. We introduce a two-timescale framework aimed at optimizing users’ long-term total QoS. Rather than reactive resource allocation, our approach proactively modifies multi-resource user demands using congestion indicators, prior to enforcing any allocation rules. Addressing the issue of insufficient user feedback on individual resource utilities, we employ a bandit-feedback version of the combinatorial multi-armed bandit framework to deduce resource-specific signals. Also, to compensate for insufficient and infrequent feedback, we’ve developed an algorithm that gleans side information from live network traffic to refine predictions on user resource sensitivities. This streamlines the algorithm’s optimality convergence and leverages the two-tier O-RAN controller structure. We validate our algorithms’ efficacy through analysis and 5G usage experiments, revealing our proposed method improves application utility by 13-60%, throughput by 8-19%, and reduces latency by 10-18%.
Xiaoxi Zhang 0001, Jinhang Zuo, Zhe Huang 0001, Zhi Zhou 0006, Xu Chen 0004, Carlee Joe-Wong
IEEE J. Sel. Areas Commun.4
2024 Dynamic Task Offloading for Multi-UAVs in Vehicular Edge Computing With Delay Guarantees: A Consensus ADMM-Based Optimization
abstract
Within the paradigm of forthcoming 6G network infrastructures, unmanned aerial vehicles (UAVs), functioning as principal conveyances, are projected to emerge as pivotal enablers in the nascent domain of the low-altitude economy. UAVs are poised to embrace various innovative applications, including latency-sensitive and compute-intensive services. However, UAVs are constrained by their energy capacity and computational resources, rendering them insufficient for fulfilling the increasingly rigorous service demands in the future. To address these challenges, our investigation focuses on the innovative UAV-based Vehicular Edge Computing (UVEC) framework, incorporating Vehicular Edge Computing (VEC) in UAV systems to bolster service reliability. A UAV can enhance its mission duration by dynamically selecting suitable vehicles for computation offloading and adaptively adjusting the task offloading ratio between vehicles and the edge server. By integrating vehicle selection and task offloading scheduling in the UVEC framework, we investigate the optimization of energy efficiency while satisfying the statistical delay and the buffer constraints for UAVs. To deal with the proposed problem, a distributed algorithm is designed by jointly considering the vehicle selection for task offloading radio to vehicles and the edge server. The stochastic network calculus (SNC) is employed to derive performance bounds for the statistical delay and constraints, enabling robust analysis and optimization of network performance. After that, we leverage linear transformation techniques to reformulate the original problem into a linear framework, enabling the application of the Alternating Direction Method of Multipliers (ADMM) algorithm to efficiently solve the transformed problem. Theoretical analysis and simulation results show that our algorithm converges while effectively satisfying service reliability constraints within the desired targets, outperforming benchmark schemes in terms of efficiency while meeting task delay and error-rate bounded constraints.
Rui Huang 0016, Wushao Wen, Zhi Zhou 0006, Chongwu Dong, Cheng Qiao, Zhihong Tian 0001, Xu Chen 0004
IEEE Trans. Mob. Comput.3
2024 DYNAMITE: Dynamic Interplay of Mini-Batch Size and Aggregation Frequency for Federated Learning With Static and Streaming Datasets
abstract
Federated Learning (FL) is a distributed learning paradigm that can coordinate heterogeneous edge devices to perform model training without sharing private data. While prior works have focused on analyzing FL convergence with respect to hyperparameters like batch size and aggregation frequency, the joint effects of adjusting these parameters on model performance, training time, and resource consumption have been overlooked, especially when facing dynamic data streams and network characteristics. This paper introduces novel analytical models and optimization algorithms that leverage the interplay between batch size and aggregation frequency to navigate the trade-offs among convergence, cost, and completion time for dynamic FL training. We establish a new convergence bound for training error considering heterogeneous datasets across devices and derive closed-form solutions for co-optimized batch size and aggregation frequency that are consistent across all devices. Additionally, we design an efficient algorithm for assigning different batch configurations across devices, improving model accuracy and addressing the heterogeneity of both data and system characteristics. Further, we propose an adaptive control algorithm that dynamically estimates network states, efficiently samples appropriate data batches, and effectively adjusts batch sizes and aggregation frequency on the fly. Extensive experiments demonstrate the superiority of our offline optimal solutions and online adaptive algorithm.
Xiaoxi Zhang 0001, Jingpu Duan, Carlee Joe-Wong, Zhi Zhou 0006, Xu Chen 0004
IEEE Trans. Mob. Comput.5
2024 Taming Serverless Cold Start of Cloud Model Inference With Edge Computing
abstract
Serverless computing is envisioned as the de-facto standard for next-generation cloud computing. However, the cold start dilemma has impeded its adoption by delay-sensitive and burst applications. In this paper, we propose to tame serverless cold start in a cloud inference system with edge computing. Specifically, the proposed solution smooths the serverless cloud workload with user-owned edge computing, reducing the number of cold starts. Leveraging the configurability of requests and serverless functions, the proposed solution further reduces the transmission latency and serverless cost by adapting request configuration (e.g., image resolution) and function configuration (e.g., memory). To alleviate the potential inference accuracy degradation incurred by configuration adaption, we aim to strike a nice balance between inference latency, cost, and accuracy. However, achieving this goal is non-trivial since the underlying optimization is non-convex and involves future uncertain information. To simultaneously address dual challenges, the presented cold-start-aware online algorithms apply the regularization technique to decompose the problem into separate convex subproblems. Then, it applies lazy switching to smooth the number of provisioned functions and thus reduces the cold start. Through rigorous theoretical analysis, realistic prototype evaluations on AWS Lambda, and trace-driven simulations, we comprehensively validate the theoretical and empirical performance of our proposed solution.
Kongyange Zhao, Zhi Zhou 0006, Lei Jiao 0002, Shen Cai, Fei Xu 0009, Xu Chen 0004
IEEE Trans. Mob. Comput.2
2024 Online Optimization of DNN Inference Network Utility in Collaborative Edge Computing
abstract
Collaborative Edge Computing (CEC) is an emerging paradigm that collaborates heterogeneous edge devices as a resource pool to compute DNN inference tasks in proximity such as edge video analytics. Nevertheless, as the key knob to improve network utility in CEC, existing works mainly focus on the workload routing strategies among edge devices with the aim of minimizing the routing cost, remaining an open question for joint workload allocation and routing optimization problem from a system perspective. To this end, this paper presents a holistic, learned optimization for CEC towards maximizing the total network utility in an online manner, even though the utility functions of task input rates are unknown a priori. In particular, we characterize the CEC system in a flow model and formulate an online learning problem in a form of cross-layer optimization. We propose a nested-loop algorithm to solve workload allocation and distributed routing iteratively, using the tools of gradient sampling and online mirror descent. To improve the convergence rate over the nested-loop version, we further devise a single-loop algorithm. Rigorous analysis is provided to show its inherent convexity, efficient convergence, as well as algorithmic optimality. Finally, extensive numerical simulations demonstrate the superior performance of our solutions.
Rui Li 0062, Tao Ouyang, Liekang Zeng, Guocheng Liao, Zhi Zhou 0006, Xu Chen 0004
IEEE/ACM Trans. Netw.5
2024 Serving Graph Neural Networks With Distributed Fog Servers for Smart IoT Services
abstract
Graph Neural Networks (GNNs) have gained growing interest in miscellaneous applications owing to their outstanding ability in extracting latent representation on graph structures. To render GNN-based service for IoT-driven smart applications, traditional model serving paradigms usually resort to the cloud by fully uploading geo-distributed input data to remote datacenters. However, our empirical measurements reveal the significant communication overhead of such cloud-based serving and highlight the profound potential in applying the emerging fog computing. To maximize the architectural benefits brought by fog computing, in this paper, we present Fograph, a novel distributed real-time GNN inference framework that leverages diverse and dynamic resources of multiple fog nodes in proximity to IoT data sources. By introducing heterogeneity-aware execution planning and GNN-specific compression techniques, Fograph tailors its design to well accommodate the unique characteristics of GNN serving in fog environments. Prototype-based evaluation and case study demonstrate that Fograph significantly outperforms the state-of-the-art cloud serving and fog deployment by up to 5.39$\times$execution speedup and 6.84$\times$throughput improvement.
Liekang Zeng, Xu Chen 0004, Ke Luo 0001, Xiaoxi Zhang 0001, Zhi Zhou 0006
IEEE/ACM Trans. Netw.6
2024 Cost-Aware Dispersed Resource Probing and Offloading at the Edge: A User-Centric Online Layered Learning Approach
abstract
To meet the stringent requirement of edge intelligence applications, resource-constrained devices can offload their task to nearby resource-rich devices. Resource awareness, as a prime prerequisite for offloading decision-making, is critical for achieving efficient collaborative computation performance. Although major works have explored computation offloading in dynamic edge environments, the impact of fresh resource information perception has not been formally investigated. To bridge the gap, we design a cost-aware edge resource probing (CERP) framework for infrastructure-free edge computing, where a task device self-organizes its resource probing to enable informed computation offloading. We first formulate the joint optimization of device probing and offloading as a multi-stage optimal stopping problem and derive a multi-threshold-based optimal strategy with theoretical guarantees. Accordingly, we devise a data-driven layered learning mechanism to handle more complex real-world scenarios. The layered learning enables the task device to adaptively learn the optimal probing sequence and decision thresholds on the fly, aiming to strike a good balance between the gain of choosing the best edge device and the accumulated cost of deep resource probing. To further boost its learning efficiency, we replace the$\epsilon$-greedy method with a tailored UCB-based adaptive exploration scheme in layered learning, thus better navigating the exploration and exploitation trade-off during probing processes. Finally, we conduct a thorough performance evaluation of the proposed CERP schemes using both extensive numerical simulations and realistic system prototype implementation, which demonstrate the superior performance of CERP in diverse application scenarios.
Tao Ouyang, Xu Chen 0004, Liekang Zeng, Zhi Zhou 0006
IEEE Trans. Serv. Comput.4
2024 Tetris: Proactive Container Scheduling for Long-Term Load Balancing in Shared Clusters
abstract
Long-running containerized workloads (e.g., machine learning), which typically showtime-varyingpatterns, are increasingly prevailing in shared production clusters. To improve workload performance, current schedulers mainly focus on optimizingshort-termbenefits of cluster load balancing orinitial container placementon servers. However, this would inevitably bring manyinvalid migrations(i.e., containers are migrated back and forth among servers over a short time window), leading to significant service level objective (SLO) violations. This paper introducesTetris, amodel predictive control(MPC)-based container scheduling strategy to proactively migrate long-running workloads for cluster load balancing. Specifically, we first build a discrete-time dynamic model forlong-termoptimization of container scheduling. To solve such an optimization problem,Tetristhen employs two main components: (1) a container resource predictor, which leverages time-series analysis approaches to accurately predict the container resource consumption; (2) an MPC-based container scheduler that jointly optimizes the cluster load balancing and container migration costover a certain sliding time window. We implement and open source a prototype ofTetrisbased on K8s. Extensive prototype experiments and trace-driven simulations demonstrate thatTetriscan improve the cluster load balancing degree by up to 77.8% without incurring any SLO violations, compared to the state-of-the-art container scheduling strategies.
Fei Xu 0009, Xiyue Shen, Shuohao Lin, Li Chen 0019, Zhi Zhou 0006, Fen Xiao, Fangming Liu
IEEE Trans. Serv. Comput.5
2024 Joint Power Allocation and Task Offloading for Reliability-Aware Services in NOMA-Enabled MEC
abstract
With the proliferation of 5G networks, mobile edge computing (MEC) has emerged as a promising technology to fulfill the stringent requirements for reliability-aware services in the Internet of Things (IoT). However, in such networks, the wireless channel states and the task arrivals are stochastic and hard to predict well. Under this scenario, tasks generated from mobile devices would pile up in the transmission queue and edge computing queue when offloading to the edge cloud via a 5G network, resulting in quality degradation for reliability-aware services. To tackle the above challenges, we introduce non-orthogonal multiple access (NOMA) in MEC to meet the requirements of ultra-reliable and low-latency communications (URLLC), in which task queuing delay violation probability and transmission error probability are both considered. Furthermore, we explore the closed-form expression based on the effective capacity (EC) to derive the performance boundary of service reliability under a general model that multiple data sources are from different IoT devices and tasks are offloaded through two-stage transmission-computing tandem queues. Based on the above mathematical analysis for service reliability, we propose an efficient strategy combining power allocation and task offloading to reduce energy consumption for all devices in NOMA-enabled MEC. Extensive simulation studies are further conducted to validate the advantage of our strategy and show the significant performance gain of nearly up to 20% over other alternatives.
Chongwu Dong, Yirui Tian, Zhi Zhou 0006, Wushao Wen, Xu Chen 0004
IEEE Trans. Wirel. Commun.3
2023 EdgeOrcher: Predictive Function Orchestration for Serverless-Based Edge Native Applications
abstract
Serverless computing is becoming prevalent to develop resource-demanding and delay-sensitive edge native applications across the edge and cloud. The unique pricing mechanism of serverless computing brings new opportunities to reduce the cost of edge native applications, by orchestrating function fusion and placement across the edge and cloud. However, function fusion potentially increases the latency of the serverless workflow. To navigate this performance-cost tradeoff, we present an online predictive function orchestration framework which leverages predictions to dynamically optimize the function fusion and placement. Preliminary evaluation results verify the efficacy of the proposed framework.
Yunkai Liang, Zhi Zhou 0006, Xu Chen 0004
ICDCS2
2023 Behavior Tree-based Workflow Modeling and Scheduling for Serverless Edge Computing
abstract
Despite the popularity of Serverless computing, there are insufficient efforts dedicated to Serverless workflows (i.e., Serverless function orchestration), particularly for Serverless edge computing. In this paper, we first identify the challenges of deploying the state-of-the-art cloud-oriented Serverless workflow scheduling on resource-constrained edge devices, then propose to model Serverless workflows with behavior trees, and finally reveal our key observations and preliminary results for behavior tree-based Serverless workflow scheduling.
Ke Luo 0001, Tao Ouyang, Zhi Zhou 0006, Xu Chen 0004
ICDCS3
2023 Learning to Be Green: Carbon-Aware Online Control for Edge Intelligence with Colocated Learning and Inference
abstract
Edge intelligence is an emerging paradigm that leverages edge computing to pave the last mile delivery of artificial intelligence. While pilot efforts on edge intelligence have mostly focused on the performance and power issues, the sustainability dilemma along with the upcoming carbon peaking and neutrality era has largely been overlooked. To green edge intelligence, we propose a carbon-aware online control framework (CARE) in this paper. CARE colocates learning and inference tasks within an edge node and dynamically adapts their configurations based on the temporal variation of carbon intensity and renewable energy availability. With such a colocation setup, CARE aims to minimize the long-term inference accuracy loss under the long-term carbon emission cap. The underlying long-term optimization problem is nontrivial since it involves uncertain information (e.g., renewable energy availability) and is NP-hard. To address these dual challenges, CARE first designs an online learning module to make fractional decisions by learning from previous system dynamics and configuration adaptation results. Then, CARE further designs a randomized rounding module, which converts the fractional decision into integer without violating the long-term carbon emission cap. The effectiveness of CARE is verified by rigorous theoretical analysis and extensive trace-driven simulations.
Shuomiao Su, Zhi Zhou 0006, Tao Ouyang, Ruiting Zhou, Xu Chen 0004
ICDCS2
2023 Fair DNN Model Selection in Edge AI via A Cooperative Game Approach
abstract
Edge intelligence is an emerging paradigm that leverages edge computing to pave the last-mile delivery of artificial intelligence (AI). To adapt to the resource restriction, model selection which adaptively selects DNN model variants is widely applied to shape the resource demand of edge AI inference tasks. Unfortunately, in current edge AI serving systems, applications are suffering unfairness since the DNN model selection is performed in a best-effort manner to maximize the system-wide inference accuracy. To achieve a predictable inference accuracy for the applications, edge AI serving systems should guarantee the minimum inference accuracy in a fair fashion at the application level. At the same time, edge resources should be efficiently utilized to minimize operational costs. In this paper, we model the edge DNN model selection problem as a Nash Bargaining Game (NBG), and propose the model selection principles by guaranteeing a base accuracy for each application. Based on the rigorous cooperative game-theoretic approach, we design an approximate algorithm to achieve computationally-efficient and fair model selection, corresponding to the Nash Bargaining Solution (NBS). With extensive trace-driven simulations, we show that our strategy can meet two desirable requirements towards the predictable inference accuracy for applications as well as low operational costs for the system.
Zhi Zhou 0006, Tao Ouyang, Xiaoxi Zhang 0001, Xu Chen 0004
ICDCS2
2023 Cost-Efficient Cloud-Edge Video Analytics with Hybird IaaS and FaaS Resources
abstract
Video analytics services are extensively employed in various real-time applications, including crime monitoring, business intelligence, and traffic flow control. These applications typically depend on cloud data centers to aggregate video stream tasks using Infrastructure as a Service (IaaS). However, the conventional approach of employing limited-term leased virtual machines (VMs) often leads to leads to high costs and idle time. Serverless computing or Function as a Service (FaaS) offers a more adaptable, pay-as-you-go solution but at a higher price, introducing latency challenges. Edge servers can reduce latency and costs, but limited physical resources affect accuracy. Opting for high-quality analytic services with better frame rates and resolutions can maintain accuracy but may not control costs. Therefore, the key challenge in video analytics is choosing the right configuration and computing methods for high accuracy, low cost, and low latency in edge, IaaS and FaaS environments. To address the challenge, we consider a scenario where the video analytics task is offloaded by combining the above three placements (i.e., VM, serverless, and edge node) to optimize the cost and latency while keeping the accuracy. The optimization problem can be formulated as an Integer Programming (IP) problem of which the NP-hardness is proved. To further deal with it, we propose an efficient online algorithm, which can optimize cost and latency using the Lyapunov optimization analysis while ensuring accuracy over long periods. Finally, the multi-angle simulation experiment results show that OCPA can effectively reduce the cost and delay compared with the benchmarks. At the same time, the achieved accuracy is close to the set long-term time-averaged value.
Zhi Zhou 0006, Kongyange Zhao, Huirong Ma, Xu Chen 0004
ICPADS2
2023 DAG-Aware Optimization for Geo-Distributed Data Analytics
abstract
Geo-distributed data analytics has been proposed to analyze geographically distributed data. Existing studies have achieved significant reductions in execution time and data transfer cost ($) of data analytics jobs by optimizing task placement. Given a directed acyclic graph (DAG)-style job, however, they mainly optimize each stage independently, and they tend to distribute tasks and intermediate data across all locations, potentially inflating execution time and data transfer cost of descendent stages and the whole job.
Qingyuan Wang 0005, Bin Gao 0013, Zhi Zhou 0006, Fei Xu 0009, Chenghao Ouyang
ICPP3
2023 Dynamic Edge-centric Resource Provisioning for Online and Offline Services Co-location
abstract
Due to the penetration of edge computing, a wide variety of workloads are sunk down to the network edge to alleviate huge pressure of the cloud. With the presence of high input workload dynamics and intensive edge resource contention, it is highly non-trivial for an edge proxy to optimize the scheduling of heterogeneous services with diverse QoS requirements. In general, online services should be quickly completed in a quite stable running environment to meet their tight latency constraint, while offline services can be processed in a loose manner for their elastic soft deadlines. To well coordinate such services at the resource-limited edge cluster, in this paper, we study an edge-centric resource provisioning optimization for dynamic online and offline services co-location, where the proxy seeks to maximize timely online service performances while maintaining satisfactory long-term offline service performances. However, intricate hybrid couplings for provisioning decisions arise due to heterogeneous constraints of the co-located services and their different time-scale performances. We hence first propose a reactive provisioning approach without requiring a prior knowledge of future system dynamics, which leverages a Lagrange relaxation for devising constraint-aware stochastic subgradient algorithm to deal with the challenge of hybrid couplings. To further boost the performance by integrating the powerful machine learning techniques, we also advocate a predictive provisioning approach, where the future request arrivals can be estimated accurately. With rigorous theoretical analysis and extensive trace-driven evaluations, we show the superior performance of our proposed algorithms for online and offline services co-location at the edge.
Tao Ouyang, Kongyange Zhao, Xiaoxi Zhang 0001, Zhi Zhou 0006, Xu Chen 0004
INFOCOM4
2023 AdaCoOpt: Leverage the Interplay of Batch Size and Aggregation Frequency for Federated Learning
abstract
Federated Learning (FL) is a distributed learning paradigm that can coordinate heterogeneous edge devices to perform model training without sharing private raw data. Many prior works have analyzed the FL convergence with respect to important hyperparameters, including batch size and aggregation frequency. However, adjusting the batch size and the number of local updates can affect the model performance, training time, and the cost of consuming computation and communication resources, in different and perhaps complex forms. Their joint effects have been overlooked and should be exploited to achieve accurate models with controllable operational expenditure. This paper proposes novel analytical models and optimization algorithms that leverage the interplay of batch size and aggregation frequency to navigate the trade-offs among convergence, cost, and completion time for FL. We first obtain a new convergence bound of the training error under heterogeneous training datasets across devices. Based on this bound, we derive closed-form solutions of a co-optimized batch size and aggregation frequency, a single configuration for all the devices. We then design an efficient exact algorithm for assigning different batch configurations across devices that can further improve the model accuracy to address the heterogeneity of both data and system characteristics. Further, we propose an adaptive control algorithm to dynamically adjust the solutions with estimated network states. Extensive experiments demonstrate the superiority of our offline optimal solutions and online adaptive algorithm.
Xiaoxi Zhang 0001, Jingpu Duan, Carlee Joe-Wong, Zhi Zhou 0006, Xu Chen 0004
IWQoS5
2023 spotDNN: Provisioning Spot Instances for Predictable Distributed DNN Training in the Cloud
abstract
Distributed Deep Neural Network (DDNN) training on cloud spot instances is increasingly compelling as it can significantly save the user budget. To handle unexpected instance revocations, provisioning a heterogeneous cluster using the asynchronous parallel mechanism becomes the dominant method for DDNN training with spot instances. However, blindly provisioning a cluster of spot instances can easily result in unpre-dictable DDNN training performance, mainly because bottlenecks occur on the parameter server network bandwidth and PCIe bandwidth resources, as well as the inadequate cluster heterogeneity. To address the challenges above, we propose spotDNN, a heterogeneity-aware spot instance provisioning framework that provides predictable performance for DDNN training in the cloud. By explicitly considering the contention for bottle-neck resources, we first build an analytical performance model of DDNN training in heterogeneous clusters. It leverages the weighted average batch size and convergence coefficient to quantify the DDNN training loss in heterogeneous clusters. Through a lightweight workload profiling, we further design a cost-efficient instance provisioning strategy which incorporates the bounds calculation and sliding window techniques to effectively guarantee the training performance service level objectives (SLOs). We have implemented a prototype of spotDNN and conducted extensive experiments on Amazon EC2. Experiment results show that spotDNN can deliver predictable DDNN training performance while reducing the monetary cost by up to 68.1% compared to the existing solutions, yet with acceptable runtime overhead.
Ruitao Shang, Fei Xu 0009, Zhuoyan Bai, Li Chen 0019, Zhi Zhou 0006, Fangming Liu
IWQoS5
2023 A Budget-aware Incentive Mechanism for Vehicle-to-Grid via Reinforcement Learning
abstract
With the increasing penetration of renewable energy and electric vehicles (EVs), the behavior of EVs' charging and discharging has shown great impact on the Micro Grid power load, motivating the development of Vehicle-to-Grid (V2G) technologies. However, the V2G market is still in its infancy, due to insufficient understanding of EV users' willingness and concerns. While many studies consider direct EV control, it's more realistic to indirectly affect users' behavior through monetary incentives. For better implementation flexibility, we advocate to display at charging piles strategically chosen incentives that are combined with electricity prices. Technically, this is the first model-free learning algorithm that can optimize incentives under unknown EV user reactions, increase the load control effectiveness and users' quality-of-service (QoS) simultaneously under a long-term incentive budget, and provide theoretical performance guarantees. We first construct a bi-level optimization framework to model the time-dependencies across our solutions. We then integrate primal-dual theories and upper-confidence bounds into reinforcement learning to balance power control and incentive consumption. A dynamic programming based algorithm is also proposed to maximize the aggregate user QoS. Finally, we prove bounded sub-optimality of our learning algorithm through theoretical analysis and conduct trace-driven simulations to demonstrate the advantages of our bi-level framework.
Tianxiang Zhu, Xiaoxi Zhang 0001, Jingpu Duan, Zhi Zhou 0006, Xu Chen 0004
IWQoS4
2023 COUPLE: Accelerating Video Analytics on Heterogeneous Mobile Processors
abstract
Deep learning has achieved tremendous success in various fields, but its significant computational demands make inference on mobile devices extremely challenging. To address this issue, we propose the COUPLE system, which enables heterogeneous processors to collaborate on mobile devices for accelerating video analytics. Additionally, we design the Co-Optimize strategy which utilizes the inference results of GPU to mitigate the accuracy loss caused by DSP. Experimental results demonstrate that COUPLE can improve the inference Average Precision by up to 5% compared to existing solutions.
Hao Bao, Zhi Zhou 0006, Qianyi Huang, Fei Xu 0009, Xu Chen 0004
MobiCom2
2023 RLink: Accelerate On-Device Deep Reinforcement Learning with Inference Knowledge at the Edge
abstract
Deep reinforcement learning (DRL) has been a successful paradigm in machine learning that enables solving complex control problems at the human level. However, the sampling and training efficiency of state-of-the-art DRL frameworks can not satisfy the stringent latency and throughput requirements of today’s mobile environments. Existing distributed and offline reinforcement learning algorithms along with the libraries for training acceleration are inherently designed for DRL tasks performed in the cloud rather than on distributed mobile devices, on which the computing resources are highly constrained, heterogeneous, and possibly dynamically changing. With the rise of edge computing and intelligence services, this paper presents RLink, a novel distributed training library to accelerate on-device deep reinforcement learning with inference knowledge at the edge. We leverage knowledge distillation to realize lightweight interaction between our on-device training task and the remote models that can provide inference knowledge. In this way, RLink is designed to be event-driven and agnostic to heterogeneous deep reinforcement learning algorithms and libraries. To tackle the communication bottleneck, a novel asynchronous sampling algorithm is proposed to facilitate real-time training in RLink. Tuned for unstable-connected mobile devices, RLink is robust and efficient by using a semantic-aware communication pipeline for lossless data compression. Extensive experimental results show that, compared with state-of-the-art algorithms and libraries, RLink can accelerate deep reinforcement learning at the edge with up to decuple speedups in convergence and ideal computational performance.
Tianyu Zeng, Xiaoxi Zhang 0001, Daipeng Feng, Jingpu Duan, Zhi Zhou 0006, Xu Chen 0004
MSN5
2023 QoS-aware Resource Optimization for Hierarchical Cross-Edge Video Analytics
abstract
As the killer application of edge computing, video analytics typically involves multiple vision components in the pipeline, which together determine the quality of service (QoS) for users. By exploiting diverse resource demands of different components, a fine-grained cross-layer orchestration with QoS-aware configuration adaptation can further boost the system efficiency of heterogeneous resources. Thus, we study a video analytics pipeline system with vertical and horizontal resource collaboration across device-edge-cloud hierarchy to achieve QoS-aware cost optimization. To judiciously match the component diversity and the resource heterogeneity, we explore smooth configuration adaptation to model a mixed-integer nonlinear problem, which jointly optimizes long-term resource cost and QoS (including accuracy and latency). However, it is nontrivial to efficiently solve such a NP-hard problem in an online manner without the future information as a prior knowledge due to the time-coupling deployment cost caused by fluctuating input traffic. To address the above challenges, we decouple the intractable problem according to the traffic routing constraints. By leveraging the lazy-switching method, we derive the component orchestration decisions for the decoupled subproblems in each slot and further design a dependent rounding scheme to obtain an efficient feasible solution while guaranteeing the knapsack resource constraints. We rigorously analyze the performance guarantee of our online algorithms by a parameterized competitive ratio, and further verify the empirical performance of our approach through extensive trace-driven experiments.
Kongyange Zhao, Zhi Zhou 0006, Tao Ouyang, Mingliao Zhao, Xu Chen 0004
SECON2
2023 Online Scheduling of CPU-NPU Co-inference for Edge AI Tasks
abstract
Edge AI is an emerging paradigm that leverages edge computing to pave the last mile delivery of artificial intelligence. To satisfy the stringent timeliness and energy-efficiency requirements of emerging edge AI tasks, specialized AI accelerator of Neural Processing Units (NPU) have been widely equipped by edge nodes. Compared to the traditional centralized processing units (CPU), NPU has better performance and energy-efficiency. However, these benefits come at the cost of reduced inference accuracy. As a result, existing coarse-grained scheduling mechanisms that schedule a whole DNN task to either the CPU or NPU are unable to make the best use of NPU. To address this issue, we propose an online NPU-CPU co-inference scheduling mechanism to schedule the DNN task at the fine-grained layer level, and thus to fully utilize the performance, accuracy, and power diversities of the NPU and CPU. By applying Lyapunov optimization to schedule the network layers dynamically, our proposed online scheduling mechanism is able to ensure the real-time inference speed and cap the long-term time-averaged power consumption, while still approximately minimizes the long-term inference accuracy loss. Via rigorous theoretical analysis as well as realistic trace-driven simulations, we demonstrate the effectiveness of our proposed online scheduling mechanism.
Xiancheng Lin, Zhi Zhou 0006, Xu Chen 0004, Zhilan Huang
WCNC5
2023 Reliability-Aware Online Scheduling for DNN Inference Tasks in Mobile-Edge Computing
abstract
Mobile-edge computing (MEC) is widely envisioned as a promising technique for provisioning artificial intelligence (AI) capability for resource-limited Internet of Things (IoT) devices by leveraging edge servers (ESs) for executing deep neural network (DNN) inference tasks in proximity. However, scheduling DNN inference tasks at the network edge under unknown system dynamics (e.g., uncertain availability of ESs) may suffer from failures, making it difficult to guarantee reliable services for the IoT device. To overcome this challenge, we propose a reliability-aware online scheduling scheme for DNN inference tasks in MEC by leveraging both online feedback and offline data to learn the uncertain availability of ESs to maximize both the inference accuracy and service reliability of DNN inference tasks (i.e., the number of DNN inference tasks processed during the system span). We first formulate the reliability-aware DNN inference tasks scheduling problem as a novel constrained combinatorial multiarmed bandit (CMAB) problem. Then by integrating the Lyapunov optimization technique, bandit learning, approximated submodular maximization, and historical data organically, we design a reliability-aware task scheduling scheme with a bandit learning (RTBL) algorithm to solve this problem. Unfortunately, even with an accurate prediction of the system uncertainties, the task scheduling problem is still NP-hard. To deal with it, we, therefore, design an advanced approximation algorithm based on the submodularity of the scheduling problem which obtains a near-optimal solution and provides a satisfactory performance guarantee. Finally, we conduct rigorous theoretical analysis and race-driven simulations to show RTBL’s brilliant performance.
Huirong Ma, Rui Li 0062, Xiaoxi Zhang 0001, Zhi Zhou 0006, Xu Chen 0004
IEEE Internet Things J.4
2023 Toward Carbon-Neutral Edge Computing: Greening Edge AI by Harnessing Spot and Future Carbon Markets
abstract
Provisioning dynamic machine learning (ML) inference as a service for artificial intelligence (AI) applications of edge devices faces many challenges, including the trade-off among accuracy loss, carbon emission, and unknown future costs. Besides, many governments are launching carbon emission rights (CER) for operators to reduce carbon emissions further to reverse climate change. Facing these challenges, to achieve carbon-aware ML task offloading under limited carbon emission rights thus to achieve green edge AI, we establish a joint ML task offloading and CER purchasing problem, intending to minimize the accuracy loss under the long-term time-averaged cost budget of purchasing the required CER. However, considering the uncertainty of the resource prices, the CER purchasing prices, the carbon intensity of sites, and ML tasks’ arrivals, it is hard to decide the optimal policy online over a long-running period time. To overcome this difficulty, we leverage the two-timescale Lyapunov optimization technique, of which the T-slot drift-plus-penalty methodology inspires us to propose an online algorithm that purchases CER in multiple timescales (on-preserved in carbon future market and on-demanded in the carbon spot market) and makes decisions about where to offload ML tasks. Considering the NP-hardness of the T-slot problems, we further propose the resource-restricted randomized dependent rounding algorithm to help to gain the near-optimal solution with no help of any future information. Our theoretical analysis and extensive simulation results driven by the real carbon intensity trace show the superior performance of the proposed algorithms.
Huirong Ma, Zhi Zhou 0006, Xiaoxi Zhang 0001, Xu Chen 0004
IEEE Internet Things J.2
2023 BeeFlow: Behavior tree-based Serverless workflow modeling and scheduling for resource-constrained edge clusters
Ke Luo 0001, Tao Ouyang, Zhi Zhou 0006, Xu Chen 0004
J. Syst. Archit.3
2023 GNN at the Edge: Cost-Efficient Graph Neural Network Processing Over Distributed Edge Servers
abstract
Edge intelligence has arisen as a promising computing paradigm for supporting miscellaneous smart applications that rely on machine learning techniques. While the community has extensively investigated multi-tier edge deployment for traditional deep learning models (e.g. CNNs, RNNs), the emerging Graph Neural Networks (GNNs) are still under exploration, presenting a stark disparity to its broad edge adoptions such as traffic flow forecasting and location-based social recommendation. To bridge this gap, this paper formally studies the cost optimization for distributed GNN processing over a multi-tier heterogeneous edge network. We build a comprehensive modeling framework that can capture a variety of different cost factors, based on which we formulate a cost-efficient graph layout optimization problem that is proved to be NP-hard. Instead of trivially applying traditional data placement wisdom, we theoretically reveal the structural property of quadratic submodularity implicated in GNN’s unique computing pattern, which motivates our design of an efficient iterative solution exploiting graph cuts. Rigorous analysis shows that it provides parameterized constant approximation ratio, guaranteed convergence, and exact feasibility. To tackle potential graph topological evolution in GNN processing, we further devise an incremental update strategy and an adaptive scheduling algorithm for lightweight dynamic layout optimization. Evaluations with real-world datasets and various GNN benchmarks demonstrate that our approach achieves superior performance over de facto baselines with more than 95.8% cost reduction in a fast convergence speed.
Liekang Zeng, Chongyu Yang, Zhi Zhou 0006, Shuai Yu 0001, Xu Chen 0004
IEEE J. Sel. Areas Commun.4
2023 Adaptive User-Managed Service Placement for Mobile Edge Computing via Contextual Multi-Armed Bandit Learning
abstract
Mobile Edge Computing (MEC), envisioned as a cloud extension, pushes cloud resource from the network core to the network edge, thereby meeting the stringent service requirements of many emerging computation-intensive mobile applications. Many existing works have focused on studying the system-wide MEC service placement issues, personalized service performance optimization yet receives much less attention. As motivated, in this paper we propose a novel adaptive user-managed service placement mechanism, which jointly optimizes a users perceived-latency and service migration cost, weighted by user-specific preferences. We first formulate the user-managed dynamic service placement process with limited system information as a contextual multi-armed bandit learning problem. In particular, we investigate both cases without and with neighboring edge feedbacks, where the later considers edge information sharing for more informed decision making. For both cases, we design lightweight Thompson-sampling based online learning algorithms, which can efficiently assist the user to make adaptive service placement decisions. We further conduct a novel information-directed theoretical analysis on the regret bound of the proposed online learning algorithms and reveal the structural impact of edge information sharing. Extensive evaluations demonstrate the superior performance gain of the proposed adaptive user-managed service placement mechanism over existing learning schemes.
Tao Ouyang, Xu Chen 0004, Zhi Zhou 0006, Rui Li 0062
IEEE Trans. Mob. Comput.3
2023 Online Control of Service Function Chainings Across Geo-Distributed Datacenters
abstract
Network Function Virtualization (NFV) provides the possibility to implement complex network functions from dedicated hardware to software instances called Virtual Network Functions (VNF) by leveraging the virtualization technology. Service Function Chaining (SFC) is therefore defined as a chain-ordered set of placed VNFs that handles the traffic of the delivery and control of a specific application. Due to the advantages of flexibility, efficiency, scalability, and short deployment cycles, NFV has been widely recognized as the next-generation network service provisioning paradigm. In this paper, we study the problem of online SFC control across geo-distributed datacenters, which is to dynamically place required VNFs on datacenter nodes and find routing paths between each adjacent VNF pair for each NFV service flow that varies over time. To that end, we first formulate this problem as an offline optimization problem whose goal is to minimize the average delay such that each datacenter's average cost does not exceed a given expense value. Considering that the offline optimization requires complete offline network information which is difficult to obtain or predict in practice, we present an online SFC control framework without requiring any future information about the traffic demands. More specifically, we leverage the Lyapunov optimization technique to formulate the problem as a series of one-time slot offline optimization problems and then apply a primal-decomposition method to solve each one-time slot problem. Simulation results reveal that our proposed online SFC control framework can efficiently reduce long-term average delay while keeping datacenter's long-term average cost consumption low.
Song Yang 0002, Fan Li 0001, Zhi Zhou 0006, Xu Chen 0004, Yu Wang 0003, Xiaoming Fu 0001
IEEE Trans. Mob. Comput.3
2023 EdgeAdaptor: Online Configuration Adaption, Model Selection and Resource Provisioning for Edge DNN Inference Serving at Scale
abstract
The accelerating convergence of artificial intelligence and edge computing has sparked a recent wave of interest in edge intelligence. While pilot efforts focused on edge DNN inference serving for a single user or DNN application, scaling edge DNN inference serving to multiple users and applications is however nontrivial. In this paper, we propose an online optimization framework EdgeAdaptor for multi-user and multi-application edge DNN inference serving at scale, which aims to navigate the three-way trade-off between inference accuracy, latency, and resource cost via jointly optimizing the application configuration adaption, DNN model selection and edge resource provisioning on-the-fly. The underlying long-term optimization problem is difficult since it is NP-hard and involves future uncertain information. To address these dual challenges, we fuse the power of online optimization and approximate optimization into a joint optimization framework, via i) decomposing the long-term problem into a series of single-shot fractional problems with a regularization technique, and ii) rounding the fractional solution to a near-optimal integral solution with a randomized dependent scheme. Rigorous theoretical analysis derives a parameterized competition ratio of our online algorithms, and extensive trace-driven simulations verify that its empirical value is no larger than 1.4 in typical scenarios.
Kongyange Zhao, Zhi Zhou 0006, Xu Chen 0004, Ruiting Zhou, Xiaoxi Zhang 0001, Shuai Yu 0001, Di Wu 0001
IEEE Trans. Mob. Comput.2
2023 HiFlash: Communication-Efficient Hierarchical Federated Learning With Adaptive Staleness Control and Heterogeneity-Aware Client-Edge Association
abstract
Federated learning (FL) is a promising paradigm that enables collaboratively learning a shared model across massive clients while keeping the training data locally. However, for many existing FL systems, clients need to frequently exchange model parameters of large data size with the remote cloud server directly via wide-area networks (WAN), leading to significant communication overhead and long transmission time. To mitigate the communication bottleneck, we resort to the hierarchical federated learning paradigm of HiFL, which reaps the benefits of mobile edge computing and combines synchronous client-edge model aggregation and asynchronous edge-cloud model aggregation together to greatly reduce the traffic volumes of WAN transmissions. Specifically, we first analyze the convergence bound of HiFL theoretically and identify the key controllable factors for model performance improvement. We then advocate an enhanced design of HiFlash by innovatively integrating deep reinforcement learning based adaptive staleness control and heterogeneity-aware client-edge association strategy to boost the system efficiency and mitigate the staleness effect without compromising model accuracy. Extensive experiments corroborate the superior performance of HiFlash in model accuracy, communication reduction, and system efficiency.
Qiong Wu 0009, Xu Chen 0004, Tao Ouyang, Zhi Zhou 0006, Xiaoxi Zhang 0001, Shusen Yang, Junshan Zhang
IEEE Trans. Parallel Distributed Syst.4
2023 iGniter: Interference-Aware GPU Resource Provisioning for Predictable DNN Inference in the Cloud
abstract
GPUs are essential to accelerating the latency-sensitive deep neural network (DNN) inference workloads in cloud datacenters. To fully utilize GPU resources,spatial sharingof GPUs among co-located DNN inference workloads becomes increasingly compelling. However, GPU sharing inevitably bringssevere performance interferenceamong co-located inference workloads, as motivated by an empirical measurement study of DNN inference on EC2 GPU instances. While existing works on guaranteeing inference performance service level objectives (SLOs) focus on eithertemporal sharingof GPUs orreactiveGPU resource scaling and inference migration techniques, how toproactivelymitigate such severe performance interference has received comparatively little attention. In this paper, we proposeiGniter, aninterference-awareGPU resource provisioning framework for cost-efficiently achieving predictable DNN inference in the cloud.iGniteris comprised of two key components: (1) alightweightDNN inference performance model, which leverages the system and workload metrics that are practically accessible to capture the performance interference; (2) Acost-efficientGPU resource provisioning strategy thatjointlyoptimizes the GPU resource allocation and adaptive batching based on our inference performance model, with the aim of achieving predictable performance of DNN inference workloads. We implement a prototype ofiGniterbased on the NVIDIA Triton inference server hosted on EC2 GPU instances. Extensive prototype experiments on four representative DNN models and datasets demonstrate thatiGnitercan guarantee the performance SLOs of DNN inference workloads with practically acceptable runtime overhead, while saving the monetary cost by up to$25\%$in comparison to the state-of-the-art GPU resource provisioning strategies.
Fei Xu 0009, Jianian Xu, Li Chen 0019, Ruitao Shang, Zhi Zhou 0006, Fangming Liu
IEEE Trans. Parallel Distributed Syst.6
2022 Fograph: Enabling Real-Time Deep Graph Inference with Fog Computing
abstract
Graph Neural Networks (GNNs) have gained growing interest in miscellaneous applications owing to their outstanding ability in extracting latent representation on graph structures. To render GNN-based service for IoT-driven smart applications, the traditional model serving paradigm resorts to the cloud by fully uploading the geo-distributed input data to the remote datacenter. However, our empirical measurements reveal the significant communication overhead of such cloud-based serving and highlight the profound potential in applying the emerging fog computing. To maximize the architectural benefits brought by fog computing, in this paper, we present Fograph, a novel distributed real-time GNN inference framework that leverages diverse resources of multiple fog nodes in proximity to IoT data sources. By introducing heterogeneity-aware execution planning and GNN-specific compression techniques, Fograph tailors its design to well accommodate the unique characteristics of GNN serving in fog environment. Prototype-based evaluation and case study demonstrate that Fograph significantly outperforms the state-of-the-art cloud serving and vanilla fog deployment by up to 5.39 × execution speedup and 6.84 × throughput improvement.
Liekang Zeng, Ke Luo 0001, Xiaoxi Zhang 0001, Zhi Zhou 0006, Xu Chen 0004
WWW5
2022 Edge Robotics: Edge-Computing-Accelerated Multirobot Simultaneous Localization and Mapping
Liekang Zeng, Xu Chen 0004, Ke Luo 0001, Zhi Zhou 0006, Shuai Yu 0001
IEEE Internet Things J.5
2022 Cost-Efficient Continuous Edge Learning for Artificial Intelligence of Things
abstract
The accelerating convergence of artificial intelligence (AI) and Internet of Things (IoT) has sparked a recent wave of interest in Artificial Intelligence of Things (AIoT). By exploiting the novel paradigm of edge intelligence, emerging computational intensive and resource demanding AIoT applications can be efficiently supported at the network edge. However, due to the limited resource capacity and/or power budget of the edge node, AIoT applications typically deploy compressed AI models to achieve the goal of low-latency and energy-efficient model inference. However, compressed models inherently suffer from the curse of data drift, i.e., the inference data at the deployment stage diverges from the training data at the training stage, leading to reduced model inference accuracy. To handle this issue, continuous learning has been proposed to periodically retrain the AI models on new data in an incremental manner. In this article, we investigate how to coordinate the edge and the cloud resources to perform cost-efficient continuous learning, with the goal of simultaneously optimizing the model performance (in terms of accuracy and robustness) and resource cost. Leveraging the Lyapunov optimization theory, we design and analyze a cost-efficient optimization framework for making online decisions upon admission control, transmission scheduling, and resource provisioning, for the dynamically arrived new data samples of various AIoT applications. We examine the effectiveness of the proposed framework on navigating the performance–cost tradeoff theoretically and empirically through trace-driven simulations.
Zhi Zhou 0006, Fei Xu 0009, Hai Jin 0001
IEEE Internet Things J.2
2022 λDNN: Achieving Predictable Distributed DNN Training With Serverless Architectures
abstract
Serverless computing is becoming a promising paradigm for Distributed Deep Neural Network (DDNN) training in the cloud, as it allows users to decompose complex model training into a number offunctionswithout managing virtual machines or servers. Though provided with a simpler resource interface (i.e., function number and memory size), inadequate function resource provisioning (either under-provisioning or over-provisioning) easily leads tounpredictableDDNN training performance in serverless platforms. Our empirical studies on AWS Lambda indicate that, suchunpredictable performanceof serverless DDNN training is mainly caused by the resource bottleneck of Parameter Servers (PS) and small local batch size. In this article, we design and implement$\lambda$λDNN, a cost-efficient function resource provisioning framework to provide predictable performance for serverless DDNN training workloads, while saving the budget of provisioned functions. Leveraging the PS network bandwidth and function CPU utilization, we build alightweightanalytical DDNN training performance model to enable our design of$\lambda$λDNNresource provisioning strategy, so as to guarantee DDNN training performance with serverless functions. Extensive prototype experiments on AWS Lambda and complementary trace-driven simulations demonstrate that,$\lambda$λDNNcan deliver predictable DDNN training performance and save the monetary cost of function resources by up to 66.7 percent, compared with the state-of-the-art resource provisioning strategies, yet with an acceptable runtime overhead.
Fei Xu 0009, Yiling Qin, Li Chen 0019, Zhi Zhou 0006, Fangming Liu
IEEE Trans. Computers4
2022 Resource Price-Aware Offloading for Edge-Cloud Collaboration: A Two-Timescale Online Control Approach
abstract
Computation offloading is envisioned as a promising technique for prolonging the battery lives and enhancing the computation capability of mobile devices. In this paper, we study the task offloading and resource purchasing problems in an edge-cloud collaborative system. The purpose of this system is to minimize the cost of task offloading while ensuring that the tasks can be served before their maximum acceptable delays. Due to the uncertainty of both the task arrival rates and the prices of the computing resources, it is impossible to make an optimal decision online for a long-running time. Therefore, we propose a two-timescale Lyapunov optimization algorithm to overcome the uncertainty of the system’s future information and make the optimal decisions only based on the system’s current states. By purchasing computation resources in different timescales from the public cloud and making online decisions on where and how many requests should be offloaded, we can achieve an efficient outcome such that the system performance can approach the offline optimum without requiring a priori knowledge of system statistics. Rigorous theoretical analysis confirms the effectiveness of the proposed two-timescale Lyapunov optimization algorithm and extensive trace-driven experimental results show that the algorithm achieves outstanding performance gains over existing benchmarks.
Rui Li 0062, Zhi Zhou 0006, Xu Chen 0004, Qing Ling 0001
IEEE Trans. Cloud Comput.2
2022 An Online Framework for Joint Network Selection and Service Placement in Mobile Edge Computing
abstract
With the rapid development and deployment of 5G wireless technology, mobile edge computing (MEC) has emerged as a new computing paradigm to facilitate a large variety of infrastructures at the network edge to reduce user-perceived communication delay. One of the fundamental problems in this new paradigm is to preserve satisfactory quality-of-service (QoS) for mobile users in light of densely dispersed wireless communication environment and often capacity-constrained MEC nodes. Such user-perceived QoS, typically in terms of the end-to-end delay, is highly vulnerable to both access network bottleneck and communication delay. Previous works have primarily focused on optimizing the communication delay through dynamic service placement, while ignoring the critical effect of access network selection on the access delay. In this work, we study the problem of jointly optimizing the access network selection and service placement for MEC, with the objective of improving the QoS in a cost-efficient manner by judiciously balancing the access delay, communication delay, and service switching cost. Specifically, we propose an efficient online framework to decompose a long-term time-varying optimization problem into a series of one-shot subproblems. To address the NP-hardness of the one-shot problem, we design a computationally-efficient two-phase algorithm based on matching and game theory, which achieves a near-optimal solution. Both rigorous theoretical analysis on the optimality gap and extensive trace-driven simulations are conducted to validate the efficacy of our proposed solution.
Bin Gao 0013, Zhi Zhou 0006, Fangming Liu, Fei Xu 0009, Bo Li 0001
IEEE Trans. Mob. Comput.2
2022 Graph Attention Spatial-Temporal Network With Collaborative Global-Local Learning for Citywide Mobile Traffic Prediction
abstract
With the rapid development of mobile cellular technologies and the increasing popularity of mobile and Internet of Things (IoT) devices, timely mobile traffic forecasting with high accuracy becomes more and more critical for proactive network service provisioning and efficient network resource allocation in smart cities. Traditional traffic forecasting methods mostly rely on time series prediction techniques, which fail to capture the complicated dynamic nature and spatial relations of mobile traffic demand. In this paper, we propose a novel deep learning framework, graph attention spatial-temporal network (GASTN), for accurate citywide mobile traffic forecasting, which can capture not only local geographical dependency but also distant inter-region relationship when considering spatial factor. Specifically, GASTN considers spatial correlation through our constructed spatial relation graph and utilizes structural recurrent neural networks to model the global near-far spatial relationships as well as the temporal dependencies. In the framework of GASTN, two attention mechanisms are designed to integrate different effects in a holistic way. Besides, in order to further enhance the prediction performance, we propose a collaborative global-local learning strategy for the training of GASTN, which takes full advantage of the knowledge from both the global model and local models for individual regions and enhance the effectiveness of our model. Extensive experiments on a large-scale real-world mobile traffic dataset demonstrate that our GASTN model dramatically outperforms the state-of-the-art methods. And it reveals that a significant enhancement in the prediction performance of GASTN can be obtained by leveraging the collaborative global-local learning strategy.
Kaiwen He 0001, Xu Chen 0004, Qiong Wu 0009, Shuai Yu 0001, Zhi Zhou 0006
IEEE Trans. Mob. Comput.5
2022 FedHome: Cloud-Edge Based Personalized Federated Learning for In-Home Health Monitoring
abstract
In-home health monitoring has attracted great attention for the ageing population worldwide. With the abundant user health data accessed by Internet of Things (IoT) devices and recent development in machine learning, smart healthcare has seen many successful stories. However, existing approaches for in-home health monitoring do not pay sufficient attention to user data privacy and thus are far from being ready for large-scale practical deployment. In this paper, we propose FedHome, a novel cloud-edge based federated learning framework for in-home health monitoring, which learns a shared global model in the cloud from multiple homes at the network edges and achieves data privacy protection by keeping user data locally. To cope with the imbalanced and non-IID distribution inherent in user’s monitoring data, we design a generative convolutional autoencoder (GCAE), which aims to achieve accurate and personalized health monitoring by refining the model with a generated class-balanced dataset from user’s personal data. Besides, GCAE is lightweight to transfer between the cloud and edges, which is useful to reduce the communication cost of federated learning in FedHome. Extensive experiments based on realistic human activity recognition data traces corroborate that FedHome significantly outperforms existing widely-adopted methods.
Qiong Wu 0009, Xu Chen 0004, Zhi Zhou 0006, Junshan Zhang
IEEE Trans. Mob. Comput.3
2022 Online Task Offloading for 5G Small Cell Networks
abstract
Small cells are deployed in 5G networks to complement the macro cells for improving coverage and capacity. Small cells and edge computing are natural partners which can improve users’ experience. Small cell nodes (SCNs) equipped with edge servers can support emerging computing services, such as virtual reality which impose low-latency and precise contextual requirements. With the proliferation of wireless devices, there is an increasing demand for offloading tasks to SCNs. Given limited computation and communication resources, the fundamental problem for a small cell network is how to select computing tasks to maximize effective rewards in an uncertain and stochastic environment. To this end, we propose an online learning framework, LFSC, which has the performance guarantee to guide task offloading in a small cell network. LFSC balances between reward and constraint violations, and it consists of three subroutines: i) a randomized algorithm which calculates selection probability of each task based on task weights; ii) a greedy assignment algorithm which cooperatively allocates tasks among different SCNs based on the selection probability; iii) an update algorithm which exploits the multi-armed bandit (MAB) technique to update task weights according to the feedback. Our theoretical analysis shows that both the regret and violations metrics of LFSC have the sub-linear property. Extensive simulation studies based on real world data confirm that LFSC achieves a close-to-optimal reward with low violations, and outperforms many state-of-the-art algorithms.
Ruiting Zhou, Shixin Qin, John C. S. Lui, Zhi Zhou 0006, Hao Huang 0001, Zongpeng Li
IEEE Trans. Mob. Comput.5
2022 Joint Application Placement and Request Routing Optimization for Dynamic Edge Computing Service Management
abstract
As mobile edge computing (MEC) hosting applications at the network edge with limited capacities, service providers are facing the new challenge of how to make full use of the scarce edge resources to maximize the system performance. Accommodating this challenge requires careful application placement and request routing to coordinate diverse MEC nodes. However, frequent application re-placement would greatly increase the system reconfiguration cost, indicating a performance-cost trade-off. In response, in this paper, we study the problem of joint optimization on application placement and request routing to maximize the system performance, under a long-term budget of the application reconfiguration cost. Solving this problem is non-trivial since the long-term budget is coupled with the future system states (e.g., user request arrivals) that are typically unpredictable. To address this challenge, we first advocate an approximated dynamic optimization framework to decompose the long-term optimization problem into a series of one-shot problems which do not require the future system states. Moreover, since the decomposed problem is a mixed integer linear program (MILP) which is proven to be NP-hard, we then devise an efficient dependent rounding based approximation algorithm, which can achieve the near-optimal performance in a fast manner. Both rigorous theoretical analysis and extensive trace-driven evaluations demonstrate the proposed framework can achieve superior performance gain over existing schemes.
Rui Li 0062, Zhi Zhou 0006, Xiaoxi Zhang 0001, Xu Chen 0004
IEEE Trans. Parallel Distributed Syst.2
2021 Flying MEC: Online Task Offloading, Trajectory Planning and Charging Scheduling for UAV-Assisted MEC
Tao Ouyang, Zhi Zhou 0006, Xu Chen 0004
ICA3PP (1)3
2021 Rationing bandwidth resources for mitigating network resource contention in distributed DNN training clusters
Qiang Qi, Fei Xu 0009, Li Chen 0019, Zhi Zhou 0006
CCF Trans. High Perform. Comput.4
2021 Boosting Edge Intelligence With Collaborative Cross-Edge Analytics
abstract
Edge intelligence is emerging as a new interdiscipline to accelerate the convergence of artificial intelligence (AI) and Internet of Things (IoT). By applying edge computing, edge intelligence enables resource-limited IoT devices to offload computation-intensive AI applications to the network edge for execution. While existing efforts have focused on optimizing the model training and inference phases of edge intelligence, the initial phase of edge intelligence-collaborative cross-edge analytics that preprocess the unlabeled raw data dispersed at multiple edge sites to obtain labeled and trainable data samples-has greatly overlooked. To bridge this gap, in this article, we study how to jointly optimize the input data and task placement, with the goal of speeding up collaborative cross-edge analytics at low network traffic cost. This problem is by no means trivial since it is nonconvex and involving future uncertainty of the query characteristics. To accommodate these dual challenges, we blend the advantages of the convex relaxation method and a two-stage optimization. Specifically, based on a prediction of the query characteristics, the input data placement is first determined when it is generated, and then when the query job arrives, the actual value of the query characteristics is used to optimize the task placement. The problem becomes more complicated when we have multiple queries simultaneously. In response, we develop an efficient flow scheduling for the intermediate data transfers of competing queries by extending the classic shortest remaining processing time (SRPT) policy. Extensive trace-driven simulations verify the efficacy of our proposed solution.
Hai Jin 0001, Zhi Zhou 0006
IEEE Internet Things J.3
2021 Age of Processing: Age-Driven Status Sampling and Processing Offloading for Edge-Computing-Enabled Real-Time IoT Applications
abstract
The freshness of status information is of great importance for time-critical Internet-of-Things (IoT) applications. A metric measuring status freshness is the Age of Information (AoI), which captures the time elapsed from the status being generated at the source node (e.g., a sensor) to the latest status update. However, in intelligent IoT applications such as video surveillance, the status information is revealed after some computation-intensive and time-consuming data processing operations, which would affect the status freshness. In this article, we propose a novel metric, Age of Processing (AoP), to quantify such status freshness, which captures the time elapsed of the newest received processed status data since it is generated. Compared with AoI, AoP further takes the data processing time into account. Since an IoT device has limited computation and energy resources, the IoT device can choose to offload the data processing to the nearby edge server under constrained status sampling frequency. We aim to minimize theaverageAoP in a long-term process by jointly optimizing the status sampling frequency and processing offloading policy. We first formulate this online problem as an infinite-horizon constrained Markov decision process (CMDP) with an average reward criterion. We then transform the CMDP problem into an unconstrained Markov decision process (MDP) by leveraging a Lagrangian method, and accordingly propose a Lagrangian transformation framework for the original CMDP problem. Furthermore, we integrate the framework with a perturbation-based refinement mechanism for achieving the optimal policy of the CMDP problem. Our investigation shows that to minimize the average AoP: 1) for processing offloading: the policy exploits good channel state to offload processing to the edge server and 2) for status sampling: the waiting time presents a threshold structure. Extensive numerical evaluations show that the proposed algorithm outperforms the benchmarks, with an average AoP reduction up to 30%.
Rui Li 0062, Qian Ma 0002, Jie Gong 0003, Zhi Zhou 0006, Xu Chen 0004
IEEE Internet Things J.4
2021 When Deep Reinforcement Learning Meets Federated Learning: Intelligent Multitimescale Resource Management for Multiaccess Edge Computing in 5G Ultradense Network
abstract
Recently, smart cities, healthcare system, and smart vehicles have raised challenges on the capability and connectivity of state-of-the-art Internet-of-Things (IoT) devices, especially for the devices in hotspots area. Multiaccess edge computing (MEC) can enhance the ability of emerging resource-intensive IoT applications and has attracted much attention. However, due to the time-varying network environments, as well as the heterogeneous resources of network devices, it is hard to achieve stable, reliable, and real-time interactions between edge devices and their serving edge servers, especially in the 5G ultradense network (UDN) scenarios. Ultradense edge computing (UDEC) has the potential to fill this gap, especially in the 5G era, but it still faces challenges in its current solutions, such as the lack of: 1) efficient utilization of multiple 5G resources (e.g., computation, communication, storage, and service resources); 2) low overhead offloading decision making and resource allocation strategies; and 3) privacy and security protection schemes. Thus, we first propose an intelligent UDEC (I-UDEC) framework, which integrates blockchain and artificial intelligence (AI) into 5G UDEC networks. Then, in order to achieve real-time and low overhead computation offloading decisions and resource allocation strategies, we design a novel two-timescale deep reinforcement learning (2Ts-DRL) approach, consisting of a fast-timescale and a slow-timescale learning process, respectively. The primary objective is to minimize the total offloading delay and network resource usage by jointly optimizing computation offloading, resource allocation, and service caching placement. We also leverage federated learning (FL) to train the 2Ts-DRL model in a distributed manner, aiming to protect the edge devices' data privacy. Simulation results corroborate the effectiveness of both the 2Ts-DRL and FL in the I-UDEC framework and prove that our proposed algorithm can reduce task execution time up to 31.87%.
Shuai Yu 0001, Xu Chen 0004, Zhi Zhou 0006, Xiaowen Gong, Di Wu 0001
IEEE Internet Things J.3
2021 Deep Reinforcement Learning With Spatio-Temporal Traffic Forecasting for Data-Driven Base Station Sleep Control
abstract
To meet the ever increasing mobile traffic demand in 5G era, base stations (BSs) have been densely deployed in radio access networks (RANs) to increase the network coverage and capacity. However, as the high density of BSs is designed to accommodate peak traffic, it would consume an unnecessarily large amount of energy if BSs are on during off-peak time. To save the energy consumption of cellular networks, an effective way is to deactivate some idle base stations that do not serve any traffic demand. In this paper, we develop a traffic-aware dynamic BS sleep control framework, named DeepBSC, which presents a novel data-driven learning approach to determine the BS active/sleep modes while meeting lower energy consumption and satisfactory Quality of Service (QoS) requirements. Specifically, the traffic demands are predicted by the proposed GS-STN model, which leverages the geographical and semantic spatial-temporal correlations of mobile traffic. With accurate mobile traffic forecasting, the BS sleep control problem is cast as a Markov Decision Process that is solved by Actor-Critic reinforcement learning methods. To reduce the variance of cost estimation in the dynamic environment, we propose a benchmark transformation method that provides robust performance indicator for policy update. To expedite the training process, we adopt a Deep Deterministic Policy Gradient (DDPG) approach, together with an explorer network, which can strengthen the exploration further. Extensive experiments with a real-world dataset corroborate that our proposed framework significantly outperforms the existing methods.
Qiong Wu 0009, Xu Chen 0004, Zhi Zhou 0006, Liang Chen 0009, Junshan Zhang
IEEE/ACM Trans. Netw.3
2021 CoEdge: Cooperative DNN Inference With Adaptive Workload Partitioning Over Heterogeneous Edge Devices
abstract
Recent advances in artificial intelligence have driven increasing intelligent applications at the network edge, such as smart home, smart factory, and smart city. To deploy computationally intensive Deep Neural Networks (DNNs) on resource-constrained edge devices, traditional approaches have relied on either offloading workload to the remote cloud or optimizing computation at the end device locally. However, the cloud-assisted approaches suffer from the unreliable and delay-significant wide-area network, and the local computing approaches are limited by the constrained computing capability. Towards high-performance edge intelligence, the cooperative execution mechanism offers a new paradigm, which has attracted growing research interest recently. In this paper, we propose CoEdge, a distributed DNN computing system that orchestrates cooperative DNN inference over heterogeneous edge devices. CoEdge utilizes available computation and communication resources at the edge and dynamically partitions the DNN inference workload adaptive to devices' computing capabilities and network conditions. Experimental evaluations based on a realistic prototype show that CoEdge outperforms status-quo approaches in saving energy with close inference latency, achieving up to 25.5% ~ 66.9% energy reduction for four widely-adopted CNN models.
Liekang Zeng, Xu Chen 0004, Zhi Zhou 0006, Lei Yang 0001, Junshan Zhang
IEEE/ACM Trans. Netw.3
2020 An Online Learning-Based Task Offloading Framework for 5G Small Cell Networks
abstract
Small cells are deployed in 5G networks to complement the macro cells for improving coverage and capacity. Small cells and edge computing are natural partners which can improve users’ experience. Small cell nodes (SCNs) equipped with edge servers can support emerging computing services such as virtual reality which impose low-latency and precise contextual requirements. With the proliferation of wireless devices, there is an increasing demand for offloading tasks to SCNs. Given limited computation and communication resources, the fundamental problem for a small cell network is how to select computing tasks to maximize effective rewards in an uncertain and stochastic environment. To this end, we propose an online learning framework, LFSC, which has the performance guarantee to guide task offloading in a small cell network. LFSC balances between reward and constraint violations, and it consists of three subroutines: i) a randomized algorithm which calculates selection probability of each task based on task weights; ii) a greedy assignment algorithm which cooperatively allocates tasks among different SCNs based on the selection probability; iii) an update algorithm which exploits the multi-armed bandit (MAB) technique to update task weights according to the feedback. Our theoretical analysis shows that both the regret and violations metrics of LFSC have the sub-linear property. Extensive simulation studies based on real world data confirm that LFSC achieves a close-to-optimal reward with low violations, and outperforms many state-of-the-art algorithms.
Ruiting Zhou, Zhi Zhou 0006, John C. S. Lui, Zongpeng Li
ICPP3
2020 CEFL: Online Admission Control, Data Scheduling, and Accuracy Tuning for Cost-Efficient Federated Learning Across Edge Nodes
abstract
With the proliferation of Internet of Things (IoT), zillions of bytes of data are generated at the network edge, incurring an urgent need to push the frontiers of artificial intelligence (AI) to network edge so as to fully unleash the potential of the IoT big data. To materialize such a vision which is known as edge intelligence, federated learning is emerging as a promising solution to enable edge nodes to collaboratively learn a shared model in a privacy-preserving and communication-efficient manner, by keeping the data at the edge nodes. While pilot efforts on federated learning have mostly focused on reducing the communication overhead, the computation efficiency of those resource-constrained edge nodes has been largely overlooked. To bridge this gap, in this article, we investigate how to coordinate the edge and the cloud to optimize the system-wide cost efficiency of federated learning. Leveraging the Lyapunov optimization theory, we design and analyze a cost-efficient optimization framework CEFL to make online yet near-optimal control decisions on admission control, load balancing, data scheduling, and accuracy tuning for the dynamically arrived training data samples, reducing both computation and communication cost. In particular, our control framework CEFL can be flexibly extended to incorporate various design choices and practical requirements of federated learning, such as exploiting the cheaper cloud resource for model training with better cost efficiency yet still facilitating on-demand privacy preservation. Via both rigorous theoretical analysis and extensive trace-driven evaluations, we verify the cost efficiency of our proposed CEFL framework.
Zhi Zhou 0006, Song Yang 0002, Lingjun Pu, Shuai Yu 0001
IEEE Internet Things J.1
2020 CE-IoT: Cost-Effective Cloud-Edge Resource Provisioning for Heterogeneous IoT Applications
abstract
With the great advance in the Internet-of-Things (IoT) sector, the recent years have witnessed an unprecedented wave of the proliferation of heterogeneous IoT devices and applications. Among them, some have stringent hard deadlines which can only be satisfied by the emerging paradigm of mobile-edge computing (MEC), while the others may pose elastic soft deadlines which can be flexibly fulfilled by cloud computing. However, with the presence of both temporal and spatial diversities of the resource cost of MEC and cloud, it remains a practical challenge how to efficiently provision the MEC and cloud resource to minimize the long-term operational cost, while still guaranteeing both hard and soft deadlines for heterogeneous IoT applications. To navigate such an inherent performance-cost tradeoff, an efficient online cloud-edge resource provisioning framework is proposed, based on the delay-aware Lyapunov optimization technique. Without requiring a priori knowledge of the statistics of the cloud-edge system, the proposed framework allows to make online greedy decisions on how much MEC and cloud resources to be provisioned to heterogeneous IoT applications. Through rigorous theoretical analysis, we prove that without violating both the hard and soft deadlines of heterogeneous IoT applications, the long-term operational cost can be pushed arbitrarily close to the offline optimum. With extensive evaluations driven by realistic traffic and cost traces, we empirically demonstrate the cost efficiency of the proposed cloud-edge resource provisioning framework.
Zhi Zhou 0006, Shuai Yu 0001, Wuhui Chen, Xu Chen 0004
IEEE Internet Things J.1
2020 DeepCP: Deep Learning Driven Cascade Prediction-Based Autonomous Content Placement in Closed Social Network
abstract
Online social networks (OSNs) are emerging as the most popular mainstream platform for content cascade diffusion. In order to provide satisfactory quality of experience (QoE) for users in OSNs, much research dedicates to proactive content placement by using the propagation pattern, user's personal profiles and social relationships in open social network scenarios (e.g., Twitter and Weibo). In this paper, we take a new direction of popularity-aware content placement in a closed social network (e.g., WeChat Moment) where user's privacy is highly enhanced. We propose a novel data-driven holistic deep learning framework, namely DeepCP, for joint diffusion-aware cascade prediction and autonomous content placement without utilizing users' personal and social information. We first devise a time-window LSTM model for content popularity prediction and cascade geo-distribution estimation. Accordingly, we further propose a novel autonomous content placement mechanism CP-GAN which adopts the generative adversarial network (GAN) for agile placement decision making to reduce the content access latency and enhance users' QoE. We conduct extensive experiments using cascade diffusion traces in WeChat Moment (WM). Evaluation results corroborate that the proposed DeepCP framework can predict the content popularity with a high accuracy, generate efficient placement decision in a real-time manner, and achieve significant content access latency reduction over existing schemes.
Qiong Wu 0009, Muhong Wu, Xu Chen 0004, Zhi Zhou 0006, Kaiwen He 0001, Liang Chen 0009
IEEE J. Sel. Areas Commun.4
2020 Mobile App Usage Patterns Aware Smart Data Pricing
abstract
The explosive growth of traffic-consumption by mobile devices is leading to severe cellular network congestion, which is posing challenges for Internet Service Providers (ISPs) to provide good quality services with limited cellular capacity and impacting the user's experience. Data pricing has been proven to be an effective way to enhance both the service quality and ISP's profit. However, traditional data pricing schemes do not consider the real Mobile Application (App) Usage Patterns (MAUPs) among large scale cellular networks. In this paper, MAUPs aware smart data pricing scheme is proposed. In our work, we firstly extract and model the users' app usage behaviors of approximately 9,600 cellular towers as two-dimensional MAUPs (time, app category). Then 7 distinct derived MAUPs are considered to be incorporated into the user satisfaction model and ISP's profit model. The performance of our proposal is evaluated and verified by numerical experiments from the aspects of ISP's profit, consumption surplus, capacity utilization and traffic efficiency. The MAUPs based pricing scheme can be periodically updated according to the operational conditions and therefore significantly instructive for ISPs.
Jieli Yin, Yali Fan, Tong Xia, Yong Li 0008, Xiang Chen 0007, Zhi Zhou 0006, Xu Chen 0004
IEEE J. Sel. Areas Commun.6
2020 A Truthful and Efficient Incentive Mechanism for Demand Response in Green Datacenters
abstract
Datacenter demand response is envisioned as a promising tool for mitigating operational stability issues faced by smart grids. It enables significant potentials in peak load reduction and facilitates the incorporation of distributed generation. Monetary refund from the smart grid can also alleviate the cloud's burden in escalating electricity cost. However, the current demand response paradigm is inefficient towards incentivizing a cloud service provider (CSP) that operates geo-distributed datacenters. To incentivize CSP participation, this work presents an auction mechanism that enables smart grids to voluntarily submit bids to the CSP to procure diverse amounts of demand response with different payments. To maximize the social welfare of the auction, the CSP that acts as the auctioneer needs to solve the winner determination problem at large-scale. By applying the proximal Jacobian alternating direction method of multipliers, we propose a distributed algorithm for each datacenter to solve a small-scale problem in a parallel fashion. Desirable properties of the proposed auction, such as social welfare maximization and truthfulness are achieved through Vickrey-Clarke-Groves (VCG) payment. Through extensive evaluations based on real datacenter workload traces and IEEE 14-bus test systems, we demonstrate that our incentive mechanism constitutes a win-win mechanism for both the geo-distributed cloud and the smart grid.
Zhi Zhou 0006, Fangming Liu, Zongpeng Li
IEEE Trans. Parallel Distributed Syst.1
2020 Edge AI: On-Demand Accelerating Deep Neural Network Inference via Edge Computing
abstract
As a key technology of enabling Artificial Intelligence (AI) applications in 5G era, Deep Neural Networks (DNNs) have quickly attracted widespread attention. However, it is challenging to run computation-intensive DNN-based tasks on mobile devices due to the limited computation resources. What’s worse, traditional cloud-assisted DNN inference is heavily hindered by the significant wide-area network latency, leading to poor real-time performance as well as low quality of user experience. To address these challenges, in this paper, we proposeEdgent, a framework that leverages edge computing for DNN collaborative inference through device-edge synergy.Edgentexploits two design knobs: (1) DNN partitioning that adaptively partitions computation between device and edge for purpose of coordinating the powerful cloud resource and the proximal edge resource for real-time DNN inference; (2) DNN right-sizing that further reduces computing latency via early exiting inference at an appropriate intermediate DNN layer. In addition, considering the potential network fluctuation in real-world deployment,Edgentis properly design to specialize for both static and dynamic network environment. Specifically, in a static environment where the bandwidth changes slowly,Edgentderives the best configurations with the assist of regression-based prediction models, while in a dynamic environment where the bandwidth varies dramatically,Edgentgenerates the best execution plan through the online change point detection algorithm that maps the current bandwidth state to the optimal configuration. We implementEdgentprototype based on the Raspberry Pi and the desktop PC and the extensive experimental evaluations demonstrateEdgent’s effectiveness in enabling on-demand low-latency edge intelligence.
Liekang Zeng, Zhi Zhou 0006, Xu Chen 0004
IEEE Trans. Wirel. Commun.3
2020 HFEL: Joint Edge Association and Resource Allocation for Cost-Efficient Hierarchical Federated Edge Learning
abstract
Federated Learning (FL) has been proposed as an appealing approach to handle data privacy issue of mobile devices compared to conventional machine learning at the remote cloud with raw user data uploading. By leveraging edge servers as intermediaries to perform partial model aggregation in proximity and relieve core network transmission overhead, it enables great potentials in low-latency and energy-efficient FL. Hence we introduce a novel Hierarchical Federated Edge Learning (HFEL) framework in which model aggregation is partially migrated to edge servers from the cloud. We further formulate a joint computation and communication resource allocation and edge association problem for device users under HFEL framework to achieve global cost minimization. To solve the problem, we propose an efficient resource scheduling algorithm in the HFEL framework. It can be decomposed into two subproblems: resource allocation given a scheduled set of devices for each edge server and edge association of device users across all the edge servers. With the optimal policy of the convex resource allocation subproblem for a set of devices under a single edge server, an efficient edge association strategy can be achieved through iterative global cost reduction adjustment process, which is shown to converge to a stable system point. Extensive performance evaluations demonstrate that our HFEL framework outperforms the proposed benchmarks in global cost saving and achieves better training performance compared to conventional federated learning.
Xu Chen 0004, Qiong Wu 0009, Zhi Zhou 0006, Shuai Yu 0001
IEEE Trans. Wirel. Commun.4
2020 Incentive-Aware Micro Computing Cluster Formation for Cooperative Fog Computing
abstract
Fog computing is envisioned as a promising approach for supporting emerging computation-intensive applications on capacity and battery constrained mobile Internet of Things (IoT) devices. Technically speaking, a massive crowd of devices in close proximity can be harvested and collaborate for computation and communication resource sharing. Hence fog computing enables significant potentials in low-latency and energy-efficient mobile task execution. However, without an efficient incentive mechanism to stimulate resource sharing among devices, the benefits of fog computing cannot be fully realized. Leveraging coalitional game theory, this work presents an efficient incentive mechanism to incentivize mutually-beneficial resource cooperation among the devices for collaborative task execution. In particular, to efficiently achieve mutually beneficial task execution, the proposed mechanism groups the devices into multiple micro computing clusters (MCCs). Within each MCC, devices can exchange mutually beneficial actions by helping to compute or transmit tasks, making all of their performances no worse than local execution or execution in the fog server. The solution to the MCC formation is devised by both centralized and decentralized schemes and further proven to admit nice properties such as top coalition, core solution, individual rationality and computational efficiency. Extensive numerical studies demonstrate the superior performance of our MCC formation mechanisms.
Xu Chen 0004, Zhi Zhou 0006, Xiang Chen 0007, Weigang Wu
IEEE Trans. Wirel. Commun.3
2020 Leveraging the Power of Prediction: Predictive Service Placement for Latency-Sensitive Mobile Edge Computing
abstract
Mobile edge computing (MEC) is emerging to support delay-sensitive 5G applications at the edge of mobile networks. When a user moves erratically among multiple MEC nodes, the challenge of how to dynamically migrate its service to maintain service performance (i.e., user-perceived latency) arises. However, frequent service migration can significantly increase operational cost, incurring the conflict between improving performance and reducing cost. To address these mis-aligned objectives, this paper studies the performance optimization of mobile edge service placement under the constraint of long-term cost budget. It is challenging because the budget involves the future uncertain information (e.g., user mobility). To overcome this difficulty, we devote to leveraging the power of prediction and advocate predictive service placement with predicted near-future information. By using two-timescale Lyapunov optimization method, we propose a T-slot predictive service placement (PSP) algorithm to incorporate the prediction of user mobility based on a frame-based design. We characterize the performance bounds of PSP in terms of cost-delay trade-off theoretically. Furthermore, we propose a new weight adjustment scheme for the queue in each frame named PSP-WU to exploit the historical queue information, which greatly reduces the length of queue while improving the quality of user-perceived latency. Rigorous theoretical analysis and extensive evaluations using realistic data traces demonstrate the superior performance of the proposed predictive schemes.
Huirong Ma, Zhi Zhou 0006, Xu Chen 0004
IEEE Trans. Wirel. Commun.2
2019 Graph Attention Spatial-Temporal Network for Deep Learning Based Mobile Traffic Prediction
abstract
With the rapid development of mobile cellular technologies and the popularity of mobile devices, timely mobile traffic forecasting with high accuracy becomes more and more critical for proactive network service provisioning and efficient network resource allocation. Due to the complicated dynamic nature of mobile traffic demand, traditional time series methods cannot satisfy the requirements of prediction tasks well and often neglect the important spatial factors. In addition, while some recent approaches model mobile traffic prediction problem using temporal and spatial features, they only consider local geographical dependency and do not take influential distant regions into consideration. In this paper, we propose Graph Attention Spatial-Temporal Network (GASTN), a novel deep learning framework to tackle the mobile traffic forecasting problem. Specifically, GASTN considers spatial correlation through the geographical relation graph and utilizes structural recurrent neural networks to model the global near-far spatial relationships as well as capture the temporal dependencies between future demand for mobile traffic and historical traffic volume. Besides, two attention mechanisms are proposed to integrate different effects in a holistic way. Extensive experiments on a large-scale real-world mobile traffic dataset demonstrate that our model significantly outperforms the state-of-the-art methods.
Kaiwen He 0001, Yufen Huang, Xu Chen 0004, Zhi Zhou 0006, Shuai Yu 0001
GLOBECOM4
2019 F3C: Fog-enabled Joint Computation, Communication and Caching Resource Sharing for Energy-Efficient IoT Data Stream Processing
abstract
Fog/edge computing has been recently regarded as a promising approach for supporting emerging mission-critical Internet of Things (IoT) applications on capacity and battery constrained devices. By harvesting and collaborating a massive crowd of devices in close proximity for computation, communication and caching resource sharing (i.e., 3C resources), it enables great potentials in low-latency and energy-efficient IoT task execution. To efficiently exploit 3C resources of fog devices in proximity, we propose F3C, a fog-enabled 3C resource sharing framework for energy-efficient IoT data stream processing by solving an energy cost minimization problem under 3C constraints. Nevertheless, the minimization problem proves to be NP-hard via reduction to a Generalized Assignment Problem (GAP). To cope with such challenge, we propose an efficient F3C algorithm based on an iterative task team formation mechanism which regards each task's 3C resource sharing as a subproblem solved by the elaborated min cost flow transformation. Via utility improving iterations, the proposed F3C algorithm is shown to converge to a stable system point. Extensive performance evaluations demonstrate that our F3C algorithm can achieve superior performance in energy saving compared to various benchmarks.
Xu Chen 0004, Zhi Zhou 0006
ICDCS3
2019 GreenEdge: Greening Edge Datacenters with Energy-Harvesting IoT Devices
abstract
Mobile edge computing, with its promise to fulfill the urgent need for richer applications and better experience of resource-hungry IoT devices, is emerging as a new computing paradigm and has quickly ascended to the spotlight. It is readily acknowledged, however that edge infrastructures are less capable of improving power usage efficiency and integrating renewable energy. To address this challenge, we propose a new framework - GreenEdge, which leverages device-to-device (D2D) communication and energy-harvesting (EH) to realize sustainable and collaborative task execution. Specifically, we first introduce the motivations of combining D2D and EH to green edge infrastructure. We next validate the feasibility and economic-efficiency of combining D2D and EH, with the help of two emerging commercial-applicable IoT applications: smart street lighting and smart bike-sharing. We further present the basic architecture, model and optimization of GreenEdge. For research inspirations, practical challenges and directions towards GreenEdge are identified. Finally, we acknowledge that GreenEdge is not the only road towards sustainability, future alternatives that can work in conjunction with GreenEdge to comprehensively green edge computing are discussed.
Zhi Zhou 0006
ICNP1
2019 Cynthia: Cost-Efficient Cloud Resource Provisioning for Predictable Distributed Deep Neural Network Training
abstract
It becomes an increasingly popular trend for deep neural networks with large-scale datasets to be trained in a distributed manner in the cloud. However, widely known as resource-intensive and time-consuming, distributed deep neural network (DDNN) training suffers from unpredictable performance in the cloud, due to the intricate factors of resource bottleneck, heterogeneity and the imbalance of computation and communication which eventually cause severe resource under-utilization. In this paper, we propose Cynthia, a cost-efficient cloud resource provisioning framework to provide predictable DDNN training performance and reduce the training budget. To explicitly explore the resource bottleneck and heterogeneity, Cynthia predicts the DDNN training time by leveraging a lightweight analytical performance model based on the resource consumption of workers and parameter servers. With an accurate performance prediction, Cynthia is able to optimally provision the cost-efficient cloud instances to jointly guarantee the training performance and minimize the training budget. We implement Cynthia on top of Kubernetes by launching a 56-docker cluster to train four representative DNN models. Extensive prototype experiments on Amazon EC2 demonstrate that Cynthia can provide predictable training performance while reducing the monetary cost for DDNN workloads by up to 50.6%, in comparison to state-of-the-art resource provisioning strategies, yet with acceptable runtime overhead.
Haoyue Zheng, Fei Xu 0009, Li Chen 0019, Zhi Zhou 0006, Fangming Liu
ICPP4
2019 Winning at the Starting Line: Joint Network Selection and Service Placement for Mobile Edge Computing
abstract
Mobile Edge Computing (MEC) is an emerging computing paradigm in which computational capabilities are pushed from the central cloud to the network edges. However, preserving the satisfactory quality-of-service (QoS) for user applications is non-trivial among multiple densely dispersed yet capacity constrained MEC nodes. This is mainly because both the access network and edge nodes are vulnerable to network congestion. Previous works are mostly limited to optimizing the QoS through dynamic service placement, while ignoring the critical effects of access network selection on the network congestion. In this paper, we study the problem of jointly optimizing the access network selection and service placement for MEC, towards the goal of improving the QoS by balancing the access, switching and communication delay. Specifically, we first design an efficient online framework to decompose the long-term optimization problem into a series of one-shot problems. To address the NP-hardness of the one-shot problem, we further propose an iteration-based algorithm to derive a computation efficient solution. Both rigorous theoretical analysis on the optimality gap and extensive trace-driven simulations validate the efficacy of our proposed solution.
Bin Gao 0013, Zhi Zhou 0006, Fangming Liu, Fei Xu 0009
INFOCOM2
2019 Adaptive User-managed Service Placement for Mobile Edge Computing: An Online Learning Approach
abstract
Mobile Edge Computing (MEC), envisioned as a cloud extension, pushes cloud resource from the network core to the network edge, thereby meeting the stringent service requirements of many emerging computation-intensive mobile applications. Many existing works have focused on studying the system-wide MEC service placement issues, personalized service performance optimization yet receives much less attention. Thus, in this paper we propose a novel adaptive user-managed service placement mechanism, which jointly optimizes a user's perceived-latency and service migration cost, weighted by user preferences. To overcome the unavailability of future information and unknown system dynamics, we formulate the dynamic service placement problem as a contextual Multi-armed Bandit (MAB) problem, and then propose a Thompson-sampling based online learning algorithm to explore the dynamic MEC environment, which further assists the user to make adaptive service placement decisions. Rigorous theoretical analysis and extensive evaluations demonstrate the superior performance of the proposed adaptive user-managed service placement mechanism.
Tao Ouyang, Rui Li 0062, Xu Chen 0004, Zhi Zhou 0006
INFOCOM4
2019 Predictive Online Server Provisioning for Cost-Efficient IoT Data Streaming Across Collaborative Edges
abstract
Edge computing is envisioned to be the de-facto paradigm of hosting emerging low latency Internet-of-Things (IoT) data streaming services.For IoT data streaming in edge computing, cost management is of strategic significance, due to the low cost-efficiency of edge servers. While existing literature adopts a reactive approach to dynamically provisioning edge servers to reduce cost, the delay of server activation and instantiation has been mostly ignored. In this paper, we target a proactive approach to dynamic edge server provisioning for real-time IoT data streaming across edge nodes, which adjusts server provisioning ahead of time, based on prediction of the upcoming workload. To effectively predict upcoming workload, a learning-based method online gradient descent is applied. We further combine the online learning method with an online optimization algorithm for server provisioning in a joint online optimization framework, through (1) minimizing of the regret incurred by inaccurate workload prediction, and (2) minimizing the cost incurred by near-optimal online decisions. The resulting predictive online algorithm can well leverage the power of prediction and achieve a good performance guarantee, as verified by both rigorous theoretical analysis and extensive trace-driven evaluations.
Zhi Zhou 0006, Xu Chen 0004, Weigang Wu, Di Wu 0001, Junshan Zhang
MobiHoc1
2019 On-demand Privacy Preservation for Cost-Efficient Edge Intelligence Model Training
Zhi Zhou 0006, Xu Chen 0004
ProvSec1
2019 Cost-Aware Edge Resource Probing for Infrastructure-Free Edge Computing: From Optimal Stopping to Layered Learning
abstract
To meet the stringent requirement of artificial intelligence applications, such as face recognition and video streaming analytics, a resource-constrained device can offload its task to nearby resource-rich devices in edge computing. Resource awareness, as a prime prerequisite for offloading decision-making, is critical for achieving efficient collaborative computation performance. In this paper, we consider cost-aware edge resource probing (CERP) framework design for infrastructure-free edge computing wherein a task device self-organizes its resource probing for informed computation offloading. We first propose a multi-stage optimal stopping formulation for the problem, and derive the optimal probing strategy which reveals a nice multi-threshold structure. Accordingly, we then devise a data-driven layered learning mechanism for more practical and complicated application environments. Layered learning enables the task device to adaptively learn the optimal probing sequence and decision thresholds at runtime, aiming at deriving a good balance between the gain of choosing the best edge device and the accumulated cost of deep resource probing. We further conduct thorough performance evaluation of the proposed CERP schemes using both extensive numerical simulations and realistic system prototype implementation, which demonstrate the superior performance of CERP in the diverse application scenarios.
Tao Ouyang, Xu Chen 0004, Liekang Zeng, Zhi Zhou 0006
RTSS4
2019 ERP: Edge Resource Pooling for Data Stream Mobile Computing
abstract
Recently, the explosion of resource-hungry and delay-sensitive Internet-of-Things (IoT) applications as exemplified by wearable appliances, video surveillance, and connected vehicles have posed great challenges on the underlying IoT devices which typically have limited computation resource. In response, computation offloading is envisioned as a promising approach to augmenting capability of IoT devices. Toward real-time and efficient computation offloading, in this paper we propose a novel edge resource pooling framework, in which a massive crowd of devices at the network edge exploit device-to-device (D2D) collaboration for pooling and sharing computation resource with each other. Specifically, we first formulate the utility maximization problem under both computation and communication constraints as a mixed-integer linear programming problem, which is further proven to be NP-hard. To address this challenge, we propose a greedy heuristic based on the classical maximum network flow problem, and thus to schedule the task offloading in a cost-efficient manner. By considering the case that a centralized controller (e.g., a network operator) is not available, a decentralized task offloading scheme is further proposed, in which IoT devices communicate and determine D2D offloading strategy locally. Rigorous theoretical analysis and extensive evaluations demonstrate the effectiveness of the proposed algorithms.
Ke Luo 0001, Zhi Zhou 0006, Xu Chen 0004
IEEE Internet Things J.3
2019 Mobile Social Data Learning for User-Centric Location Prediction With Application in Mobile Edge Service Migration
abstract
Recently, location prediction has attracted considerable research effort because of the popularity of location-based services, such as mobile advertising and recommendations. With the unprecedented proliferation of mobile social networks, such as WeChat and Twitter, we are able to use location service to bridge the online and offline worlds, which is of great significance to many smart city applications. Different from existing studies, in this paper, we promote a user-centric location prediction approach by leveraging a user's local mobile social information without involving other users' location privacy. We propose a factor graph learning model that integrates not only user's social and network information but also the correlations between a user's locations into a unified framework. Furthermore, we use ReliefF algorithm to select user-specific significant features for location prediction and define the measure of location entropy to study the similarity between location, network status, and social behavior. To show the benefit of precise location prediction, we further apply it to personalized service migration in mobile edge computing (MEC) and accordingly propose prediction-based amortizing algorithm and lazy migration algorithm that can well balance the tradeoff between migration cost and non-migration latency in a cost-efficient manner. We conduct extensive experiments using a real-world data trace, which shows that our model performs much better in location prediction compared with several classic methods and the MEC service quality can be significantly enhanced by leveraging the location prediction.
Qiong Wu 0009, Xu Chen 0004, Zhi Zhou 0006, Liang Chen 0009
IEEE Internet Things J.3
2019 Online Orchestration of Cross-Edge Service Function Chaining for Cost-Efficient Edge Computing
abstract
Edge computing (EC) has quickly ascended to be the de-facto standard for hosting emerging low-latency applications, as exemplified by intelligent video surveillance, Internet of Vehicles, and augmented reality. For EC, service function chaining is envisioned as a promising approach to configure various services in an agile, flexible, and cost-efficient manner. When running on top of geographically dispersed edge clouds, fully unleashing the benefits of service function chaining is, however, by no means trivial. In this paper, we propose an online orchestration framework for cross-edge service function chaining, which aims to maximize the holistic cost efficiency, via jointly optimizing the resource provisioning and traffic routing on-the-fly. This long-term cost minimization problem is difficult since it is NP-hard and involves future uncertain information. To simultaneously address these dual challenges, we carefully combine an online optimization technique with an approximate optimization method in a joint optimization framework, through: 1) decomposing the long-term problem into a series of one-shot fractional problem with a regularization technique and 2) rounding the fractional solution to a near-optimal integral solution with a randomized dependent scheme that preserves the solution feasibility. The resulting online algorithm achieves an outstanding performance guarantee, as verified by both rigorous theoretical analysis and extensive trace-driven simulations.
Zhi Zhou 0006, Qiong Wu 0009, Xu Chen 0004
IEEE J. Sel. Areas Commun.1
2019 Edge Intelligence: Paving the Last Mile of Artificial Intelligence With Edge Computing
abstract
With the breakthroughs in deep learning, the recent years have witnessed a booming of artificial intelligence (AI) applications and services, spanning from personal assistant to recommendation systems to video/audio surveillance. More recently, with the proliferation of mobile computing and Internet of Things (IoT), billions of mobile and IoT devices are connected to the Internet, generating zillions bytes of data at the network edge. Driving by this trend, there is an urgent need to push the AI frontiers to the network edge so as to fully unleash the potential of the edge big data. To meet this demand, edge computing, an emerging paradigm that pushes computing tasks and services from the network core to the network edge, has been widely recognized as a promising solution. The resulted new interdiscipline, edge AI or edge intelligence (EI), is beginning to receive a tremendous amount of interest. However, research on EI is still in its infancy stage, and a dedicated venue for exchanging the recent advances of EI is highly desired by both the computer system and AI communities. To this end, we conduct a comprehensive survey of the recent research efforts on EI. Specifically, we first review the background and motivation for AI running at the network edge. We then provide an overview of the overarching architectures, frameworks, and emerging key technologies for deep learning model toward training/inference at the network edge. Finally, we discuss future research opportunities on EI. We believe that this survey will elicit escalating attentions, stimulate fruitful discussions, and inspire further research ideas on EI.
Zhi Zhou 0006, Xu Chen 0004, Liekang Zeng, Ke Luo 0001, Junshan Zhang
Proc. IEEE1
2019 Cost-Effective Cloud Server Provisioning for Predictable Performance of Big Data Analytics
abstract
Cloud datacenters are underutilized due to server over-provisioning. To increase datacenter utilization, cloud providers offer users an option to run workloads such as big data analytics on the underutilized resources, in the form of cheap yet revocable transient servers (e.g., EC2 spot instances, GCE preemptible instances). Though at highly reduced prices, deploying big data analytics on the unstable cloud transient servers can severely degrade the job performance due to instance revocations. To tackle this issue, this paper proposes iSpot, a cost-effective transient server provisioning framework for achieving predictable performance in the cloud, by focusing on Spark as a representative Directed Acyclic Graph (DAG)-style big data analytics workload. It first identifies the stable cloud transient servers during the job execution by devising an accurate Long Short-Term Memory (LSTM)-based price prediction method. Leveraging automatic job profiling and the acquired DAG information of stages, we further build an analytical performance model and present a lightweight critical data checkpointing mechanism for Spark, to enable our design of iSpot provisioning strategy for guaranteeing the job performance on stable transient servers. Extensive prototype experiments on both EC2 spot instances and GCE preemptible instances demonstrate that, iSpot is able to guarantee the performance of big data analytics running on cloud transient servers while reducing the job budget by up to 83.8 percent in comparison to the state-of-the-art server provisioning strategies, yet with acceptable runtime overhead.
Fei Xu 0009, Haoyue Zheng, Wujie Shao, Haikun Liu, Zhi Zhou 0006
IEEE Trans. Parallel Distributed Syst.6
2018 Automatic 3D Neuron Tracing Based on Terminations Detection
Chao Wang 0072, Weixun Chen, Min Liu 0008, Zhi Zhou 0006
BIBM4
2018 User-Centric Location Prediction in Mobile Social Networks: A Factor Graph Learning Approach
abstract
Recently, location prediction has attracted considerable research effort because of the popularity of location- based services, such as mobile advertising and recommendations. With the unprecedented proliferation of mobile social networks, we are able to use location service to bridge the online and offline worlds. Different from existing studies, in this paper we promote a user-centric location prediction approach by leveraging a user's local mobile social information without involving other users' location privacy. We propose a factor graph learning model that integrates not only user's social and network information, but also the correlations between user's locations into a unified framework. Furthermore, we use ReliefF algorithm to select user-specific significant features for location prediction and define the measure of location entropy to study the similarity between location, network status and social behavior. We conduct extensive experiments using a real-world dataset, which shows that our model performs much better in location prediction compared with several classic methods.
Qiong Wu 0009, Xu Chen 0004, Zhi Zhou 0006, Liang Chen 0009
GLOBECOM3
2018 eBrowser: Making Human-Mobile Web Interactions Energy Efficient with Event Rate Learning
abstract
Due to the limited screen size of mobile devices, finger movements on touchscreen, such as scrolling and pinching (i.e., zooming in or out), are frequently used on mobile Web browsers and WebView-based apps, consuming considerable energy on mobile devices. While existing works on mobile Web browsers focus on reducing the power consumption or optimizing the performance of webpage loading, the power consumption of mobile Web interactions, especially after webpage loading, has received comparatively little attention. Motivated by an empirical study of the power consumption and user experience survey of human-mobile interactions, we design and implement eBrowser, an energy-efficient mobile Web interaction framework. It leverages a cloud-based machine learning model to enable personalized interaction event rate for individual users according to the interaction speed of their finger movement and the content of rendered webpages. To adapt to user behavior changes, eBrowser continuously monitors the interaction experience on each mobile device and periodically updates the personalized event rate model with incremental learning in the cloud. We implement eBrowser in Chromium and deploy the event rate model in a remote Aliyun cloud instance. Experimental results show that eBrowser reduces the energy consumption of mobile Web interactions by up to 43.8% with negligible runtime overhead, while guaranteeing user satisfaction on both mobile browsers and WebView-based apps.
Fei Xu 0009, Zhi Zhou 0006, Jia Rao
ICDCS3
2018 Dewing in Fog: Incentive-Aware Micro Computing Cluster Formation for Fog Computing
abstract
Fog computing is envisioned as a promising approach for supporting emerging mission-critical applications on capacity and battery constrained mobile devices. By harvesting and collaborating a massive crowd of devices in close proximity for computation and communication resource sharing, it enables significant potentials in low-latency and energy-efficient mobile task execution. It is readily acknowledged, however, that without an efficient incentive mechanism that stimulates resources sharing among devices, the benefits of fog computing cannot be fully realized. Leveraging coalitional game theory, this work presents an efficient incentive mechanism to incentivize mutually-beneficial resource cooperation among the devices for collaborative task execution. Specially, to prevent the over-exploiting and free-riding behaviors that harm resource-rich device's willingness to collaborate, the proposed mechanism groups the devices into multiple micro computing clusters (MC-C). Within each MCC, devices can exchange mutually beneficial actions by helping to compute or transmit tasks, making all of them better off. The solution of the MCC formation is devised by a network-assisted mechanism, which is further proven to admit nice properties such as top coalition and core solution.
Zhi Zhou 0006, Xiang Chen 0007, Weigang Wu
ICPADS2
2018 Follow Me at the Edge: Mobility-Aware Dynamic Service Placement for Mobile Edge Computing
abstract
Mobile edge computing is a new computing paradigm in which cloud computing capabilities are pushed from the network core to the network edge to serve the end-user in proximity. However, with the sinking of computing capabilities, the new challenge incurred by user mobility arises: since end-users typically move erratically, the services should be dynamically migrated among multiple edges to maintain the service performance, i.e., user-perceived latency. Tackling this problem is non-trivial since frequent service migration would greatly increase the operational cost. To address this challenge in terms performance-cost trade-off, in this paper we study the mobile edge service performance optimization problem under long-term cost budget constraint. To address user mobility which is typically unpredictable, we first apply Lyapunov optimization to decompose the long-term optimization problem into a series of real-time optimization problems which do not require a priori knowledge such as user mobility. As the decomposed problem is NP-hard, we further propose an efficient heuristic based on the Markov approximation technique. Rigorous theoretical analysis and extensive evaluations demonstrate the efficacy of the proposed solution.
Tao Ouyang, Zhi Zhou 0006, Xu Chen 0004
IWQoS2
2018 A D2D offloading approach to efficient mobile edge resource pooling
abstract
The explosion of resource-hungry mobile applications has posed great challenges on the underlying mobile devices which typically have limited computation resource. In response, device-to-device (D2D) computation offloading is envisioned as a promising approach to the problem by gearing resource-rich devices and resource-poor devices. Towards real-time and efficient computation offloading, in this paper, we proposed a novel edge resource pooling framework called ERP, in which a massive crowd of devices at the network edge exploit D2D collaboration for pooling and sharing computation resource with each other. Specifically, we first formulate the utility maximization problem under both computation and communication constraints as a mixed-integer linear programming (MILP), which is further proven to be NP-hard. To address this challenge, we propose a centralized greedy heuristic based on the classical maximum network flow problem, which schedules the task offloading in a cost-efficient manner. Rigorous theoretical analysis and extensive evaluations demonstrate the effectiveness of the heuristic to some extent.
Ke Luo 0001, Zhi Zhou 0006, Xu Chen 0004
WiOpt3
2018 Follow Me at the Edge: Mobility-Aware Dynamic Service Placement for Mobile Edge Computing
abstract
Mobile edge computing is a new computing paradigm, which pushes cloud computing capabilities away from the centralized cloud to the network edge. However, with the sinking of computing capabilities, the new challenge incurred by user mobility arises: since end users typically move erratically, the services should be dynamically migrated among multiple edges to maintain the service performance, i.e., user-perceived latency. Tackling this problem is non-trivial since frequent service migration would greatly increase the operational cost. To address this challenge in terms of the performance-cost tradeoff, in this paper, we study the mobile edge service performance optimization problem under long-term cost budget constraint. To address user mobility which is typically unpredictable, we apply Lyapunov optimization to decompose the long-term optimization problem into a series of real-time optimization problems which do not require a priori knowledge such as user mobility. As the decomposed problem is NP-hard, we first design an approximation algorithm based on Markov approximation to seek a near-optimal solution. To make our solution scalable and amenable to future fifth-generation application scenario with large-scale user devices, we further propose a distributed approximation scheme with greatly reduced time complexity, based on the technique of the best response update. Rigorous theoretical analysis and extensive evaluations demonstrate the efficacy of the proposed centralized and distributed schemes.
Tao Ouyang, Zhi Zhou 0006, Xu Chen 0004
IEEE J. Sel. Areas Commun.2
2017 Predictive Resilience Analysis of Complex Systems Using Dynamic Bayesian Networks
abstract
Uncertain and potentially harsh operating environments are often known to alter the operational performance of a system. In order to maintain system performance while coping with varying operating environments and potential disruptions, the resilience of engineered systems is desirable. Engineering systems are often interconnected in a dimensional way inherently from basic components to subsystems to the system of systems, which poses a grand challenge for system designers to analyze the resilience of such a complex system. Moreover, further complications in the assessment of resilience in the engineering domain are attributed to time-varying system performances, random perturbation occurrences, and probable failures caused by adverse events. This paper presents a dynamic Bayesian network (DBN) approach for the modeling and predictive resilience analysis for dynamic engineered systems. With the inter-time-slice links and the conditional probability tables in a DBN, the system performance could be molded as changing in a discrete time slice, while capturing the temporal probabilistic dependencies between the variables. An industrial-based case study of an electricity distribution system is further studied to demonstrate the effectiveness of the DBN approach for resilience analysis. The approach presented in this paper hopes to aid in realizing resiliency in system designs and to pave the way toward enhancements in developing resilient engineered systems.
Nita Yodo, Pingfeng Wang, Zhi Zhou 0006
IEEE Trans. Reliab.3
2016 On-Demand and Reliable vSD-EON Provisioning with Correlated Data and Control Plane Embedding
abstract
Software-defined elastic optical networks (SD-EONs) provide operators more flexibility to customize their optical infrastructure dynamically and adaptively, and network virtualization, i.e., infrastructure-as-a-service (IaaS), enables multiple tenants to share the substrate infrastructure efficiently. In this paper, we study how to provision virtual SD-EONs (vSD-EONs) with the correlated data and control plane embedding (χ-VNE) that considers the quality-of-service (QoS) of virtual control plane (vCP), i.e., availability and control channel latency. We propose a χ-VNE algorithm to solve the problem, design the network system to realize it, and accomplish proof-of-concept experimental demonstrations in an OpenFlow-based network testbed. Numerical and experimental results indicate that the proposed algorithm and system function well and can realize on-demand and reliable vSD-EON provisioning.
Heqing Yin, Siqi Liu 0004, Zhi Zhou 0006, Xiaoliang Chen 0004, Zuqing Zhu
GLOBECOM4
2016 A coupled memcapacitor emulator based relaxation oscillator
abstract
Tremendous efforts have been put into dissecting the inherent characteristics and potential applications of Memcapacitor (MC), which possesses unique abilities of storing both information and energy [1]. Recently, coupling is disclosed as the third relation beyond series and parallel connections of memristive circuits in [2], of which the mechanical dynamic coupling of MCs is taken into account for illustration purpose. Coupled MCs could provide us more opportunities for developing new electronic devices with unique functions. However, very few works currently focus on the practical implementation of coupled MC emulators and its possible application in electronic circuits. In this letter, a practical emulator of coupled MC is newly proposed and then used for structuring Relaxation Oscillators (ROs), of which the period and duty cycle of output oscillating signal can be purposefully controlled in virtue of the coupling action.
Dongsheng Yu, Zhi Zhou 0006, Herbert H. C. Iu, Tyrone Fernando
ISCAS2
2016 Bilateral Electricity Trade Between Smart Grids and Green Datacenters: Pricing Models and Performance Evaluation
abstract
Datacenter demand response is a promising approach for mitigating operational instability faced by smart grids. It enables significant potentials in peak load shedding and facilitates the incorporation of distributed generation and intermittent energy sources. This paper considers two key aspects toward real-time electricity pricing for eliciting demand response: 1) two-way electricity flow between smart grids and large datacenters with hybrid green generation capabilities and 2) the geo-distributed nature of large cloud systems, and hence the potential competition among smart grids that serve different datacenters of the cloud. We propose a pricing scheme tailored for geo-distributed green datacenters, from a multi-leader (smart grids) single-follower (cloud) game point of view. At the cloud side, in quest for scalability, robustness, and performance, the energy cost minimization problem is solved in a distributed manner, based on the technique of alternating direction method of multipliers. At the smart grid side, a practical equilibrium of the multi-leader single-follower pricing game is desired. To this end, we employ the technique of equilibrium problem with equilibrium constraints and exact linearization, to accurately transform the multi-leader single-follower pricing game, which is non-convex into a mixed integer linear system that can be readily solved. The effectiveness of the proposed solutions is evaluated based on the real datacenter workload traces and the IEEE 14-bus test systems with real generation and demand data.
Zhi Zhou 0006, Fangming Liu, Zongpeng Li
IEEE J. Sel. Areas Commun.1
2016 Carbon-Aware Online Control of Geo-Distributed Cloud Services
abstract
Recently, datacenter carbon emission has become an emerging concern for the cloud service providers. Previous works are limited on cutting down the power consumption of datacenters to defuse such a concern. In this paper, we show how the spatial and temporal variabilities of the electricity carbon footprint can be fully exploited to further green the cloud running on top of geographically distributed datacenters. Specifically, we first verify that electricity cost minimization conflicts with carbon emission minimization, based on an empirical study of several representative geo-distributed cloud services. We then jointly consider the electricity cost, service level agreement (SLA) requirement, and emission reduction budget. To navigate such a three-way tradeoff, we take advantage of Lyapunov optimization techniques to design and analyze a carbon-aware control framework, which makes online decisions on geographical load balancing, capacity right-sizing, and server speed scaling. Results from rigorous mathematical analysis and real-world trace-driven evaluation demonstrate the effectiveness of our framework in reducing both electricity cost and carbon emission.
Zhi Zhou 0006, Fangming Liu, Ruolan Zou, Jiangchuan Liu, Hong Xu 0001, Hai Jin 0001
IEEE Trans. Parallel Distributed Syst.1