Liekang Zeng

dblp:241/6945 · DBLP profile ↗
← Back
38ranked-venue papers
6as first author
35since 2021 · last 2026
0000-0003-4800-8768ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 26 · 4 first-author · 25 since 2021Systems, architecture and hardware · 5 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Chimera: An Efficient Multimodal Embodied Inference Framework with Complexity-Aware Edge-Cloud Routing
Muen Xue, Liekang Zeng, Tao Ouyang, Shaoyong Guo, Xu Chen 0004
ICDCS3
2026 Venus: An Efficient Edge Memory-and-Retrieval System for VLM-based Online Video Understanding
Shengyuan Ye, Bei Ouyang, Tianyi Qian, Liekang Zeng, Mu Yuan, Xiaowen Chu 0001, Weijie Hong, Xu Chen 0004
INFOCOM4
2026 COACH: Adaptive Robust Human-Robot Collaboration for Efficient Smart Manufacturing
abstract
Modern smart manufacturing pipelines have pervasively collaborated human workers, mobile robots, and industrial Internet of Things (IIoT) in shared workspaces for versatile production tasks. Despite the promising capacity of individual entities, the performance of these IIoT systems largely relies on pipeline coordination, i.e., task dispatching between humans and robots, which is particularly challenging under heterogeneous physical constraints and complex environmental uncertainties. Nonetheless, existing works either rely on traditional operation frameworks that lack scalability for large-scale complex production, or propose customized solutions for fixed agent models, overlooking the evolving nature of IIoT environments. To address these limitations, this paper proposes COACH, a human-robot collaborative manufacturing system that enables robust constraint-aware coordination across humans, robots, and IIoT. Specifically, COACH designs a scalable contextual encoder to represent the evolving relationships among human and robot agents in dynamic heterogeneous graphs. With that, a novel experience-driven task dispatcher is developed, enabling both high-performance and computation-efficient policy generation concerning the status of IIoT. To accommodate changing human fatigue and pipeline scales, COACH further develops a curriculum-enhanced reinforcement learning module for efficient dispatcher adaptation. Extensive evaluations using both synthetic testbeds and real-world manufacturing datasets demonstrate that COACH improves the feasible ratio of manufacturing pipelines by up to 27.4% and achieves up to 13.9% improvement in time efficiency compared to competing baselines across diverse job scales and environmental settings.
Hui Wang 0011, Liekang Zeng, Zhiwen Yu 0001, Yao Zhang 0005, Di Duan, Mu Yuan, Bin Guo 0001, Guoliang Xing
SenSys2
2026 BOTH: Efficient Coordination of Mobile Agents With Graph-Enhanced Bayesian Online Learning
abstract
Collaborative agents, consisting of at least one human and one mobile robot agent working toward a common objective, are increasingly prevalent and effective in both social and industrial spheres, such as manufacturing. The inherent heterogeneity of these agents requires efficient and scalable Task Scheduling and Allocation (TSA) schemes that match individuals to tasks based on their abilities and meet specific temporal constraints, maximizing performance in less time. Existing works face challenges as exact methods rely on assumptions and deterministic models, which struggle to scale and infer time-varying, stochastic human task performance. While offline reinforcement learning shows promise, it is time-consuming and heavily dependent on training data that is often scarce in practical factory settings. To address these challenges, we formulate the TSA problem in mobile multi-agent teams as a temporal-constrained contextual decision-making process and propose the Bayesian Optimization-augmented Team coordination among Heterogeneous agents (BOTH), a novel scalable and training-free scheduling approach. The core idea is to use Gaussian Processes (GP) to iteratively infer agent dynamics in real-time, enabling the automatic derivation of a robust TSA solution that requires no prior data and adapts to varying problem sizes. We start by employing a heterogeneous graph-based encoder to extract representative context from the individual differences among team agents and tasks, considering strict temporal constraints. Following this, we propose a GP-driven Bayesian optimizer to intelligently explore and exploit optimal task assignments for each context, without making assumptions about the system. Experiments on synthetic and real datasets demonstrate that BOTH boosts accuracy and time efficiency compared to competing baselines, even within a few iterations.
Zhiwen Yu 0001, Yao Zhang 0005, Jiaqi Liu 0002, Liekang Zeng, Huan Zhou 0002, Bin Guo 0001, Guoliang Xing
IEEE Trans. Mob. Comput.5
2026 Resource-Efficient Personal Large Language Models Fine-Tuning With Collaborative Edge Computing
abstract
Large language models (LLMs) have unlocked a plethora of powerful applications at the network edge, such as intelligent personal assistants. Data privacy and security concerns have prompted a shift towards edge-based fine-tuning of personal LLMs, away from cloud reliance. However, this raises issues of computational intensity and resource scarcity, hindering training efficiency and feasibility. While current studies investigate parameter-efficient fine-tuning (PEFT) techniques to mitigate resource constraints, our analysis indicates that these techniques are not sufficiently resource-efficient for edge devices. Other studies focus on exploiting the potential of edge devices through resource management optimization, yet are ultimately bottlenecked by the resource wall of individual devices. To tackle these challenges, we proposePAC+, a resource efficient collaborative edge AI framework for in-situ personal LLMs fine-tuning.PAC+breaks the resource wall of personal LLMs fine-tuning with a sophisticated algorithm-system co-design. (1) Algorithmically,PAC+implements a personal LLMs fine-tuning technique that is efficient in terms of parameters, time, and memory. It utilizes Parallel Adapters to circumvent the need for a full backward pass through the LLM backbone. Additionally, an activation cache mechanism further streamlining the process by negating the necessity for repeated forward passes across multiple epochs. (2) Systematically,PAC+leverages edge devices in close proximity, pooling them as a collective resource for in-situ personal LLMs fine-tuning, utilizing a hybrid data and pipeline parallelism to orchestrate distributed training. The use of the activation cache eliminates the need for forward pass through the LLM backbone, enabling exclusive fine-tuning of the Parallel Adapters using data parallelism. Extensive evaluation of the prototype implementation demonstrates thatPAC+significantly outperforms existing collaborative edge training systems, achieving up to a$9.7\times$end-to-end speedup. Furthermore, compared to mainstream LLM fine-tuning algorithms,PAC+reduces memory footprint by up to$88.16\%$.
Shengyuan Ye, Bei Ouyang, Tianyi Qian, Liekang Zeng, Jiangsu Du, Xiaowen Chu 0001, Guoliang Xing, Xu Chen 0004
IEEE Trans. Parallel Distributed Syst.4
2025 Multi-Tier Multi-Node Scheduling of LLM for Collaborative AI Computing
Mulei Ma, Chenyu Gong, Liekang Zeng, Yang Yang 0001
INFOCOM3
2025 Jupiter: Fast and Resource-Efficient Collaborative Inference of Generative LLMs on Edge Devices
Shengyuan Ye, Bei Ouyang, Liekang Zeng, Tianyi Qian, Xiaowen Chu 0001, Jian Tang 0008, Xu Chen 0004
INFOCOM3
2025 Poster: Mobile Menstrual Health Advising with Multimodal Feature Engineering
Liekang Zeng, Zhenyu Yan 0002, Yunqi Guo, Hongkai Chen 0001, Guoliang Xing
MobiSys2
2025 ContextAgent: Context-Aware Proactive LLM Agents with Open-world Sensory Perceptions
abstract
Recent advances in Large Language Models (LLMs) have propelled intelligent agents from reactive responses to proactive support. While promising, existing proactive agents either rely exclusively on observations from enclosed environments (e.g., desktop UIs) with direct LLM inference or employ rule-based proactive notifications, leading to suboptimal user intent understanding and limited functionality for proactive service. In this paper, we introduce ContextAgent, the first context-aware proactive agent that incorporates extensive sensory contexts surrounding humans to enhance the proactivity of LLM agents. ContextAgent first extracts multi-dimensional contexts from massive sensory perceptions on wearables (e.g., video and audio) to understand user intentions. ContextAgent then leverages the sensory contexts and personas from historical data to predict the necessity for proactive services. When proactive assistance is needed, ContextAgent further automatically calls the necessary tools to assist users unobtrusively. To evaluate this new task, we curate ContextAgentBench, the first benchmark for evaluating context-aware proactive LLM agents, covering 1,000 samples across nine daily scenarios and twenty tools. Experiments on ContextAgentBench show that ContextAgent outperforms baselines by achieving up to 8.5% and 6.0% higher accuracy in proactive predictions and tool calling, respectively. We hope our research can inspire the development of more advanced, human-centric, proactive AI assistants. The code and dataset are publicly available at https://github.com/openaiotlab/ContextAgent.
Bufang Yang, Lilin Xu, Liekang Zeng, Kaiwei Liu 0001, Siyang Jiang, Wenrui Lu, Hongkai Chen 0001, Xiaofan Jiang 0001, Guoliang Xing, Zhenyu Yan 0002
NeurIPS3
2025 Grape: Efficient Spatiotemporal Prediction Services with Stale Sensing Streams
abstract
Emerging cyber-physical systems have embraced a large number of IoT devices spanning geo-distributed, which generate and consume massive volumes of data continuously. Accurate and timely spatiotemporal predictions (STP) over these streaming sensor data are critical and, in growing demand, ubiquitous across various edge scenarios such as traffic flow forecasting. Towards that, recent advanced systems have developed sophisticated optimizations among STP pipelines, aiming at optimal prediction performance. However, based on our empirical studies in real-world settings, we identify a previously overlooked bottleneck of end-to-end STP performance: data staleness. To mitigate this issue, in this work, we investigate a new task, namely stream interception, which deliberately terminates the acceptance of incoming sensor data and anticipates model execution with imputed missing features. We propose a novel dynamic interception strategy to determine the time slot to exit waiting and present Grape, an STP system that implements it with practical system designs. Extensive evaluations on real-world traces show that Grape can strike a superior tradeoff between prediction accuracy and serving latency, achieving 1.69-1.90× speedup against traditional all-waiting baselines across various STP services with high prediction accuracy on par with offline optimal cases.
Liekang Zeng, Shengyuan Ye, Mu Yuan, Di Duan, Xu Chen 0004, Guoliang Xing
RTSS1
2025 SCX: Stateless KV-Cache Encoding for Cloud-Scale Confidential Transformer Serving
abstract
Transformer models have revolutionized fields like natural language processing and computer vision but face privacy concerns in sensitive applications such as medical diagnostics. Existing confidential serving methods, including cryptography-based, memory isolation-based, and access control-based, offer trade-offs between privacy and efficiency but often struggle with high latency or hardware dependencies. This work proposes stateless KV-cache encoding (SCX), a novel framework that encodes the intermediate key-value cache during Transformer inference using user-controlled keys. SCX ensures that the cloud can neither recover the input nor independently complete the next token prediction, effectively preserving privacy. By introducing efficient encoding and decoding schemes, SCX addresses communication complexity and attack vulnerabilities while ensuring zero loss of inference quality. Experiments on large Transformer models demonstrate that SCX achieves lower latency (e.g., 36ms for LLaMA-7B), outperforming state-of-the-art cryptography and memory isolation methods by orders of magnitude. Moreover, SCX can complementarily work with advanced KV-cache management techniques to further enhance KV-cache communication efficiency by 85%, marking a significant step toward practical, privacy-preserving large Transformer serving.
Mu Yuan, Lan Zhang 0002, Liekang Zeng, Siyang Jiang, Bufang Yang, Di Duan, Guoliang Xing
SIGCOMM3
2025 Revisiting Location Privacy in MEC-Enabled Computation Offloading
abstract
Mobile Edge Computing (MEC) revolutionizes real-time applications by extending cloud capabilities to network edges, enabling efficient computation offloading from mobile devices. In recent years, the location privacy concern within MEC offloading has been recognized, prompting the proposal of various methodologies to mitigate this concern. However, this paper demonstrates that the prevailing privacy protection methods exhibit vulnerabilities. First, we analyze the shortcomings of current methodologies through both system modeling and evaluation metrics. Then, we introduce a Learning-based Trajectory Reconstruction Attack (LTRA) to expose the weaknesses, achieving up to 91.2% reconstruction accuracy against the state-of-the-art protection method. Further, based onw-event differential privacy, we propose an ℓ-trajectory differentially private mechanism, i.e., OffloadingBD. Compared to the existing works, OffloadingBD provides more flexible and enhanced protection with sound privacy theoretical guarantee. Lastly, we conduct extensive experiments to evaluate LTRA and OffloadingBD. The experiment results show that LTRA has good generalization ability and OffloadingBD showcases a superior balance between privacy and utility compared with baselines.
Wenzhong Ou, Bei Ouyang, Shengyuan Ye, Liekang Zeng, Lin Chen 0002, Xu Chen 0004
IEEE Trans. Inf. Forensics Secur.5
2025 Sequential Privacy Budget Recycling for Federated Vector Mean Estimation: A Game-Theoretic Approach
abstract
Privacy-preserving vector mean estimation is a crucial primitive in federated analytics. Existing practices usually resort to Local Differentiated Privacy (LDP) mechanisms that inject random noise into users’ vectors when communicating with users and the central server. Due to the privacy-utility trade-off, the privacy budget has been widely recognized as the bottleneck resource that requires well-provisioning. In this paper, we explore the possibility of privacy budget recycling and propose a novelChainDPframework enabling users to carry out data aggregation sequentially to recycle the privacy budget. We establish a sequential game to model the user interactions in our framework. We theoretically show the mathematical nature of the sequential game, solve its Nash Equilibrium, and design an incentive mechanism with provable economic properties. To alleviate potential privacy collusion attacks, we further derive a differentially privacy-guaranteed protocol to avoid holistic exposure. Our numerical simulation validates the effectiveness of ChainDP, showing that it can significantly save privacy budget as well as lower estimation error compared to the traditional LDP mechanism.
Guangjing Huang, Liekang Zeng, Lin Chen 0002, Xu Chen 0004
IEEE Trans. Mob. Comput.3
2025 Resource-Efficient Collaborative Edge Transformer Inference With Hybrid Model Parallelism
abstract
Transformer-based models have unlocked a plethora of powerful intelligent applications at the edge, such as voice assistant in smart home. Traditional deployment approaches offload the inference workloads to the remote cloud server, which would induce substantial pressure on the backbone network as well as raise users' privacy concerns. To address that, in-situ inference has been recently recognized for edge intelligence, but it still confronts significant challenges stemming from the conflict between intensive workloads and limited on-device computing resources. In this paper, we leverage our observation that many edge environments usually comprise a rich set of accompanying trusted edge devices with idle resources and proposeGalaxy+, a collaborative edge AI system that breaks the resource walls across heterogeneous edge devices for efficient Transformer inference acceleration.Galaxy+introduces a novel hybrid model parallelism to orchestrate collaborative inference, along with a heterogeneity and memory-aware parallelism planning for fully exploiting the resource potential. To mitigate the impact of tensor synchronizations on inference latency under bandwidth-constrained edge environments,Galaxy+devises a tile-based fine-grained overlapping of communication and computation. Furthermore, a fault-tolerant re-scheduling mechanism is developed to address device-level resource dynamics, ensuring stable and low-latency inference. Extensive evaluation based on prototype implementation demonstrates thatGalaxy+remarkably outperforms state-of-the-art approaches under various edge environment setups, achieving a$1.2\times$to$4.24\times$end-to-end latency reduction. Besides,Galaxy+can adapt to device-level resource dynamics, swiftly rescheduling and restoring inference in the presence of unexpected straggler devices.
Shengyuan Ye, Bei Ouyang, Jiangsu Du, Liekang Zeng, Tianyi Qian, Wenzhong Ou, Xiaowen Chu 0001, Deke Guo, Yutong Lu, Xu Chen 0004
IEEE Trans. Mob. Comput.4
2025 Mitigating Tail Latency for On-Device Inference With Load-Balanced Heterogeneous Models
abstract
Serving machine learning models on edge, mobile, and embedded devices places stringent requirements on inference latency. From operating a real enterprise service, we observed that even a fully optimized model could lead to severe violations of latency objectives when the load surges. A straightforward and mature approach is to auto-scale multiple models to balance the load. However, unlike cloud clusters, edge or mobile devices usually cannot afford to deploy multiple model replicas. Therefore, in this paper, we explore a new idea: in addition to the original model, we deploy one (or more) heterogeneous model(s) with much smaller resource overhead on the device, and perform load balancing among all models. We overcame the technical challenges posed by performance dynamics and developed InferRouter based on queuing theory. We implement and evaluate InferRouter on three real on-device inference systems, covering mobile sensing, video analytics, and natural language processing applications. Experimental results show that compared with strong baselines, InferRouter can decrease 85.2% P99 latency (5.8x faster) and improve 5.9% accuracy on the mobile workload. For a traffic video analytics task, InferRouter achieves 55.1% higher accuracy with zero deadline misses. InferRouter also shows its advantages in saving resources compared with auto-scaling and offloading approaches.
Mu Yuan, Lan Zhang 0002, Di Duan, Liekang Zeng, Miaohui Song, Zichong Li, Guoliang Xing, Xiang-Yang Li 0001
IEEE Trans. Mob. Comput.4
2024 MOGR: Multi-task Offloading via Graph Representation in Heterogeneous Computing Network
abstract
In the rapidly evolving field of heterogeneous computing networks, efficient task offloading plays a pivotal role in optimizing system throughput and resource utilization. However, existing task offloading methods often fall short of adequately modeling the dependency topology relationships between of-floaded tasks, which limits their effectiveness in capturing the complex interdependencies of task features. To address this limitation, we propose a framework named MOGR: Multi-task Offloading via Graph Representation. Our modeling approach takes into account factors such as task characteristics, network conditions, and available resources at the edge, and embeds these captured features into the graph structure. By utilizing Graph Convolutional Networks (GCN), our mechanism can capture and analyze the intricate relationships between task features, enabling a more comprehensive understanding of the underlying dependency topology. Through extensive evaluations in heteroge-neous networks, our proposed algorithm improves 15.1%-30.5% over greedy and approximate algorithms in optimizing system throughput and resource utilization. Our experiments showcase the advantage of considering the intricate interplay of task features using GCN-based modeling.
Mulei Ma, Chenyu Gong, Liekang Zeng, Yang Yang 0001
ICC3
2024 Pluto and Charon: A Time and Memory Efficient Collaborative Edge AI Framework for Personal LLMs Fine-tuning
abstract
Large language models (LLMs) have unlocked a plethora of powerful applications at the network edge, such as intelligent personal assistants. Data privacy and security concerns have prompted a shift towards edge-based fine-tuning of personal LLMs, away from cloud reliance. However, this raises issues of computational intensity and resource scarcity, hindering training efficiency and feasibility. While current studies investigate parameter-efficient fine-tuning (PEFT) techniques to mitigate resource constraints, our analysis indicates that these techniques are not sufficiently resource-efficient for edge devices. Other studies focus on exploiting the potential of edge devices through resource management optimization, yet are ultimately bottlenecked by the resource wall of individual devices.
Bei Ouyang, Shengyuan Ye, Liekang Zeng, Tianyi Qian, Xu Chen 0004
ICPP3
2024 Galaxy: A Resource-Efficient Collaborative Edge AI System for In-situ Transformer Inference
abstract
Transformer-based models have unlocked a plethora of powerful intelligent applications at the edge, such as voice assistant in smart home. Traditional deployment approaches offload the inference workloads to the remote cloud server, which would induce substantial pressure on the backbone network as well as raise users’ privacy concerns. To address that, in-situ inference has been recently recognized for edge intelligence, but it still confronts significant challenges stemming from the conflict between intensive workloads and limited on-device computing resources. In this paper, we leverage our observation that many edge environments usually comprise a rich set of accompanying trusted edge devices with idle resources and propose Galaxy, a collaborative edge AI system that breaks the resource walls across heterogeneous edge devices for efficient Transformer inference acceleration. Galaxy introduces a novel hybrid model parallelism to orchestrate collaborative inference, along with a heterogeneity-aware parallelism planning for fully exploiting the resource potential. Furthermore, Galaxy devises a tile-based fine-grained overlapping of communication and computation to mitigate the impact of tensor synchronizations on inference latency under bandwidth-constrained edge environments. Extensive evaluation based on prototype implementation demonstrates that Galaxy remarkably outperforms state-of-the-art approaches under various edge environment setups, achieving up to 2.5× end-to-end latency reduction.
Shengyuan Ye, Jiangsu Du, Liekang Zeng, Wenzhong Ou, Xiaowen Chu 0001, Yutong Lu, Xu Chen 0004
INFOCOM3
2024 SECO: Multi-Satellite Edge Computing Enabled Wide-Area and Real-Time Earth Observation Missions
abstract
Rapid advances in low Earth orbit (LEO) satellite technology and satellite edge computing (SEC) have facilitated a key role for LEO satellites in enhanced Earth observation missions (EOM). These missions (e.g., remote object detection) typically require multi-satellite cooperative observations of a large region of interest (RoI) area, as well as the observation image routing and computation processing, enabling accurate and real-time responsiveness. However, optimizing the resources of LEO satellite networks is nontrivial in the presence of its dynamic and heterogeneous properties. To this end, we propose SECO, a SEC-enabled framework that jointly optimizes multi-satellite observation scheduling, routing and computation node selection for enhanced EOM. Specifically, in the observation phase, we leverage the orbital motion and the rotatable onboard cameras of satellites, and propose a distributed game-based scheduling strategy to minimize the overall size of captured images while ensuring full (observation) coverage. In the sequent routing and computation phase, we first adopt image splitting technology to achieve parallel transmission and computation. Then, we propose an efficient iterative algorithm to jointly optimize image splitting, routing and computation node selection for each captured image. On this basis, we propose a theoretically guaranteed systemwide greedy-based strategy to reduce the total time cost (i.e., transmission, computation and queuing delay) over simultaneous processing for multiple images. Extensive experiments based on real-world datasets demonstrate that SECO can achieve up to a 60.7% reduction in overall time cost compared to baselines.
Zhiwei Zhai, Liekang Zeng, Tao Ouyang, Shuai Yu 0001, Qianyi Huang, Xu Chen 0004
INFOCOM2
2024 Asteroid: Resource-Efficient Hybrid Pipeline Parallelism for Collaborative DNN Training on Heterogeneous Edge Devices
abstract
On-device Deep Neural Network (DNN) training has been recognized as crucial for privacy-preserving machine learning at the edge. However, the intensive training workload and limited onboard computing resources pose significant challenges to the availability and efficiency of model training. While existing works address these challenges through native resource management optimization, we instead leverage our observation that edge environments usually comprise a rich set of accompanying trusted edge devices with idle resources beyond a single terminal. We propose Asteroid, a distributed edge training system that breaks the resource walls across heterogeneous edge devices for efficient model training acceleration. Asteroid adopts a hybrid pipeline parallelism to orchestrate distributed training, along with a judicious parallelism planning for maximizing throughput under certain resource constraints. Furthermore, a fault-tolerant yet lightweight pipeline replay mechanism is developed to tame the device-level dynamics for training robustness and performance stability. We implement Asteroid on heterogeneous edge devices with both vision and language models, demonstrating up to 12.2× faster training than conventional parallelism methods and 2.1× faster than state-of-the-art hybrid parallelism methods through evaluations. Furthermore, Asteroid can recover training pipeline 14× faster than baseline methods while preserving comparable throughput despite unexpected device exiting and failure.
Shengyuan Ye, Liekang Zeng, Xiaowen Chu 0001, Guoliang Xing, Xu Chen 0004
MobiCom2
2024 FlocOff: Data Heterogeneity Resilient Federated Learning With Communication-Efficient Edge Offloading
abstract
Federated Learning (FL) has emerged as a fundamental learning paradigm to harness massive data scattered at geo-distributed edge devices in a privacy-preserving way. Given the heterogeneous deployment of edge devices, however, their data are usually Non-IID, introducing significant challenges to FL including degraded training accuracy, intensive communication costs, and high computing complexity. Towards that, traditional approaches typically utilize adaptive mechanisms, which may suffer from scalability issues, increased computational overhead, and limited adaptability to diverse edge environments. To address that, this paper instead leverages the observation that the computation offloading involves inherent functionalities such as node matching and service correlation to achieve data reshaping and proposesFederatedlearning basedoncomputingOffloading (FlocOff) framework, to address data heterogeneity and resource-constrained challenges. Specifically, FlocOff formulates the FL process with Non-IID data in edge scenarios and derives rigorous analysis on the impact of imbalanced data distribution. Based on this, FlocOff decouples the optimization in two steps, namely: 1) Minimizes the Kullback-Leibler (KL) divergence via Computation Offloading scheduling (MKL-CO); 2) Minimizes the Communication Cost through Resource Allocation (MCC-RA). Extensive experimental results demonstrate that the proposed FlocOff effectively improves model convergence and accuracy by 14.3%-32.7% while reducing data heterogeneity under various data distributions.
Mulei Ma, Chenyu Gong, Liekang Zeng, Yang Yang 0001, Liantao Wu
IEEE J. Sel. Areas Commun.3
2024 Online Optimization of DNN Inference Network Utility in Collaborative Edge Computing
abstract
Collaborative Edge Computing (CEC) is an emerging paradigm that collaborates heterogeneous edge devices as a resource pool to compute DNN inference tasks in proximity such as edge video analytics. Nevertheless, as the key knob to improve network utility in CEC, existing works mainly focus on the workload routing strategies among edge devices with the aim of minimizing the routing cost, remaining an open question for joint workload allocation and routing optimization problem from a system perspective. To this end, this paper presents a holistic, learned optimization for CEC towards maximizing the total network utility in an online manner, even though the utility functions of task input rates are unknown a priori. In particular, we characterize the CEC system in a flow model and formulate an online learning problem in a form of cross-layer optimization. We propose a nested-loop algorithm to solve workload allocation and distributed routing iteratively, using the tools of gradient sampling and online mirror descent. To improve the convergence rate over the nested-loop version, we further devise a single-loop algorithm. Rigorous analysis is provided to show its inherent convexity, efficient convergence, as well as algorithmic optimality. Finally, extensive numerical simulations demonstrate the superior performance of our solutions.
Rui Li 0062, Tao Ouyang, Liekang Zeng, Guocheng Liao, Zhi Zhou 0006, Xu Chen 0004
IEEE/ACM Trans. Netw.3
2024 A3D: Adaptive, Accurate, and Autonomous Navigation for Edge-Assisted Drones
abstract
Accurate navigation is of paramount importance to ensure flight safety and efficiency for autonomous drones. Recent research starts to use Deep Neural Networks (DNN) to enhance drone navigation given their remarkable predictive capability for visual perception. However, existing solutions either run DNN inference tasks on drones in situ, impeded by the limited onboard resource, or offload the computation to external servers which may incur large network latency. Few works consider jointly optimizing the offloading decisions along with image transmission configurations and adapting them on the fly. In this paper, we propose A3D, an edge server assisted drone navigation framework that can dynamically adjust task execution location, input resolution, and image compression ratio in order to achieve low inference latency, high prediction accuracy, and long flight distances. Specifically, we first augment state-of-the-art convolutional neural networks for drone navigation and define a novel metric called Quality of Navigation as our optimization objective which can effectively capture the above goals. We then design a deep reinforcement learning (DRL) based neural scheduler at the drone side for which an information encoder is devised to reshape the state features and thus improve its learning ability. To further support simultaneous multi-drone serving, we extend the edge server design by developing a network-aware resource allocation algorithm, which allows provisioning containerized resources aligned with drones’ demand. We finally implement a proof-of-concept prototype with realistic devices and validate its performance in a real-world campus scene, as well as a simulation environment for thorough evaluation upon AirSim. Extensive experimental results show that A3D can reduce end-to-end latency by 28.06% and extend the flight distance by up to 27.28% compared with non-adaptive solutions.
Liekang Zeng, Daipeng Feng, Xiaoxi Zhang 0001, Xu Chen 0004
IEEE/ACM Trans. Netw.1
2024 Serving Graph Neural Networks With Distributed Fog Servers for Smart IoT Services
abstract
Graph Neural Networks (GNNs) have gained growing interest in miscellaneous applications owing to their outstanding ability in extracting latent representation on graph structures. To render GNN-based service for IoT-driven smart applications, traditional model serving paradigms usually resort to the cloud by fully uploading geo-distributed input data to remote datacenters. However, our empirical measurements reveal the significant communication overhead of such cloud-based serving and highlight the profound potential in applying the emerging fog computing. To maximize the architectural benefits brought by fog computing, in this paper, we present Fograph, a novel distributed real-time GNN inference framework that leverages diverse and dynamic resources of multiple fog nodes in proximity to IoT data sources. By introducing heterogeneity-aware execution planning and GNN-specific compression techniques, Fograph tailors its design to well accommodate the unique characteristics of GNN serving in fog environments. Prototype-based evaluation and case study demonstrate that Fograph significantly outperforms the state-of-the-art cloud serving and fog deployment by up to 5.39$\times$execution speedup and 6.84$\times$throughput improvement.
Liekang Zeng, Xu Chen 0004, Ke Luo 0001, Xiaoxi Zhang 0001, Zhi Zhou 0006
IEEE/ACM Trans. Netw.1
2024 Cost-Aware Dispersed Resource Probing and Offloading at the Edge: A User-Centric Online Layered Learning Approach
abstract
To meet the stringent requirement of edge intelligence applications, resource-constrained devices can offload their task to nearby resource-rich devices. Resource awareness, as a prime prerequisite for offloading decision-making, is critical for achieving efficient collaborative computation performance. Although major works have explored computation offloading in dynamic edge environments, the impact of fresh resource information perception has not been formally investigated. To bridge the gap, we design a cost-aware edge resource probing (CERP) framework for infrastructure-free edge computing, where a task device self-organizes its resource probing to enable informed computation offloading. We first formulate the joint optimization of device probing and offloading as a multi-stage optimal stopping problem and derive a multi-threshold-based optimal strategy with theoretical guarantees. Accordingly, we devise a data-driven layered learning mechanism to handle more complex real-world scenarios. The layered learning enables the task device to adaptively learn the optimal probing sequence and decision thresholds on the fly, aiming to strike a good balance between the gain of choosing the best edge device and the accumulated cost of deep resource probing. To further boost its learning efficiency, we replace the$\epsilon$-greedy method with a tailored UCB-based adaptive exploration scheme in layered learning, thus better navigating the exploration and exploitation trade-off during probing processes. Finally, we conduct a thorough performance evaluation of the proposed CERP schemes using both extensive numerical simulations and realistic system prototype implementation, which demonstrate the superior performance of CERP in diverse application scenarios.
Tao Ouyang, Xu Chen 0004, Liekang Zeng, Zhi Zhou 0006
IEEE Trans. Serv. Comput.3
2023 Real-Time High-Resolution Pedestrian Detection in Crowded Scenes via Parallel Edge Offloading
abstract
To identify dense and small-size pedestrians in surveillance systems, high-resolution cameras are widely deployed, where high-resolution images are captured and delivered to off-the-shelf pedestrian detection models. However, given the highly computation-intensive workload brought by the high resolution, the resource-constrained cameras fail to afford accurate inference in real time. To address that, we propose Hode, an offloaded video analytic framework that utilizes multiple edge nodes in proximity to expedite pedestrian detection with high-resolution inputs. Specifically, Hode can intelligently split high-resolution images into respective regions and then offload them to distributed edge nodes to perform pedestrian detection in parallel. A spatio-temporal flow filtering method is designed to enable context-aware region partitioning, as well as a DRL-based scheduling algorithm to allow accuracy-aware load balance among heterogeneous edge nodes. Extensive evaluation results using realistic prototypes show that Hode can achieve up to 2.01× speedup with very mild accuracy loss.
Hao Bao, Liekang Zeng, Ke Luo 0001, Xu Chen 0004
ICC3
2023 Chained-DP: Can We Recycle Privacy Budget?
abstract
Privacy-preserving vector mean estimation is a crucial primitive in federated analytics. Existing practices usually resort to Local Differentiated Privacy (LDP) mechanisms that inject random noise into users' vectors when communicating with users and the central server. Due to the privacy-utility trade-off, the privacy budget has been widely recognized as the bottleneck resource that requires well provisioning. In this paper, we explore the possibility of privacy budget recycling and propose a novel Chained-DP framework enabling users to carry out data aggregation sequentially to recycle the privacy budget. We establish a sequential game to model the user interactions in our framework. We theoretically show the mathematical nature of the sequential game, solve its Nash Equilibrium, and design an incentive mechanism with provable economic properties. Our numerical simulation validates the effectiveness of Chained-DP, showing that it can significantly save privacy budget as well as lower estimation error compared to the traditional LDP mechanism.
Guangjing Huang, Liekang Zeng, Lin Chen 0002, Xu Chen 0004
IWQoS3
2023 GNN at the Edge: Cost-Efficient Graph Neural Network Processing Over Distributed Edge Servers
abstract
Edge intelligence has arisen as a promising computing paradigm for supporting miscellaneous smart applications that rely on machine learning techniques. While the community has extensively investigated multi-tier edge deployment for traditional deep learning models (e.g. CNNs, RNNs), the emerging Graph Neural Networks (GNNs) are still under exploration, presenting a stark disparity to its broad edge adoptions such as traffic flow forecasting and location-based social recommendation. To bridge this gap, this paper formally studies the cost optimization for distributed GNN processing over a multi-tier heterogeneous edge network. We build a comprehensive modeling framework that can capture a variety of different cost factors, based on which we formulate a cost-efficient graph layout optimization problem that is proved to be NP-hard. Instead of trivially applying traditional data placement wisdom, we theoretically reveal the structural property of quadratic submodularity implicated in GNN’s unique computing pattern, which motivates our design of an efficient iterative solution exploiting graph cuts. Rigorous analysis shows that it provides parameterized constant approximation ratio, guaranteed convergence, and exact feasibility. To tackle potential graph topological evolution in GNN processing, we further devise an incremental update strategy and an adaptive scheduling algorithm for lightweight dynamic layout optimization. Evaluations with real-world datasets and various GNN benchmarks demonstrate that our approach achieves superior performance over de facto baselines with more than 95.8% cost reduction in a fast convergence speed.
Liekang Zeng, Chongyu Yang, Zhi Zhou 0006, Shuai Yu 0001, Xu Chen 0004
IEEE J. Sel. Areas Commun.1
2022 AdaDrone: Quality of Navigation Based Neural Adaptive Scheduling for Edge-Assisted Drones
abstract
Accurate navigation is of paramount importance to ensure flight safety and efficiency for autonomous drones. Recent research starts to use Deep Neural Networks (DNN) to enhance drone navigation given their remarkable predictive capability for visual perception. However, existing solutions either run DNN inference tasks on drones in-situ, impeded by the limited onboard resource, or offload the computation to external servers which may incur large network latency. Few works consider jointly optimizing the offloading decisions along with image transmission configurations and adapting them on the fly. In this paper, we propose AdaDrone, an edge computing assisted drone navigation framework that can dynamically adjust task execution location, input resolution, and image compression ratio in order to achieve low inference latency, high prediction accuracy, and long flight distances. Specifically, we first augment state-of-the-art convolutional neural networks for drone navigation and define a novel metric called Quality of Navigation as our optimization objective which can effectively capture the above goals. We then design a deep reinforcement learning (DRL) based neural scheduler for which an information encoder is devised to reshape the state features and thus improve its learning ability. We finally implement a prototype of our framework wherein a drone board for navigation and scheduling control interacts with edge servers for task offloading and a simulator for performance evaluation. Extensive experimental results show that AdaDrone can reduce end-to-end latency by 28.06% and extend the flight distance by up to 27.28% compared with non-adaptive solutions.
Liekang Zeng, Xiaoxi Zhang 0001, Xu Chen 0004
ICDCS2
2022 Eco-FL: Adaptive Federated Learning with Efficient Edge Collaborative Pipeline Training
abstract
Federated Learning (FL) has been a promising paradigm in distributed machine learning that enables in-situ model training and global model aggregation. While it can well preserve private data for end users, to apply it efficiently on IoT devices yet suffer from their inherent variants: their available computing resources are typically constrained, heterogeneous, and changing dynamically. Existing works deploy FL on IoT devices by pruning a sparse model or adopting a tiny counterpart, which alleviates the workload but may have negative impacts on model accuracy. To address these issues, we propose Eco-FL, a novel Edge Collaborative pipeline based Federated Learning framework. On the client side, each IoT device collaborates with trusted available devices in proximity to perform pipeline training, enabling local training acceleration with efficient augmented resource orchestration. On the server side, Eco-FL adopts a novel grouping-based hierarchical architecture that combines synchronous intra-group aggregation and asynchronous inter-group aggregation, where a heterogeneity-aware dynamic grouping strategy that jointly considers response latency and data distribution is developed. To tackle the resource fluctuation during the runtime, Eco-FL further applies an adaptive scheduling policy to judiciously adjust workload allocation and client grouping at different levels. Extensive experimental results using both prototype and simulation show that, compared to state-of-the-art methods, Eco-FL can upgrade the training accuracy by up to 26.3%, reduce the local training time by up to 61.5%, and improve the local training throughput by up to 2.6 ×.
Shengyuan Ye, Liekang Zeng, Qiong Wu 0009, Ke Luo 0001, Qingze Fang, Xu Chen 0004
ICPP2
2022 Adaptive Progressive Image Enhancement for Edge-Assisted Mobile Vision
abstract
Recent advances in deep learning models have pushed Super-Resolution (SR) techniques to an unprecedented altitude, enabling high-quality image rendering with variable scaling size and natural fidelity. To deploy them on resource-constrained mobile devices, however, confronts significant chal-lenges of excessively long latency and poor user experience. To this end, we propose Apie, an edge-assisted adaptive image rendering system that allows low-latency, progressive image enhancement for a smooth user experience. Apie adopts a data parallel strategy across the end device and the edge server, along with a residual learning mechanism to judiciously retrieve information for SR models. Besides, a novel progressive image reconstruction is developed by exploiting content-aware image blocking and incremental image rendering, towards improved quality of user experience. Furthermore, Apie can dynamically adjust the choice of employed SR models with respect to the networking conditions, striking a good balance upon the latency-quality trade-off. Extensive evaluations show that Apie performs 7.33x faster than on-device GPU execution and 1.42x faster compared to the partial offloading method, while achieves 2.84dB higher PSNR compared to the interpolation method using conventional JPEG image compression and 0.74dB higher PSNR compared to the partial offloading method.
Daipeng Feng, Liekang Zeng, Lingjun Pu, Xu Chen 0004
MSN2
2022 Fograph: Enabling Real-Time Deep Graph Inference with Fog Computing
abstract
Graph Neural Networks (GNNs) have gained growing interest in miscellaneous applications owing to their outstanding ability in extracting latent representation on graph structures. To render GNN-based service for IoT-driven smart applications, the traditional model serving paradigm resorts to the cloud by fully uploading the geo-distributed input data to the remote datacenter. However, our empirical measurements reveal the significant communication overhead of such cloud-based serving and highlight the profound potential in applying the emerging fog computing. To maximize the architectural benefits brought by fog computing, in this paper, we present Fograph, a novel distributed real-time GNN inference framework that leverages diverse resources of multiple fog nodes in proximity to IoT data sources. By introducing heterogeneity-aware execution planning and GNN-specific compression techniques, Fograph tailors its design to well accommodate the unique characteristics of GNN serving in fog environment. Prototype-based evaluation and case study demonstrate that Fograph significantly outperforms the state-of-the-art cloud serving and vanilla fog deployment by up to 5.39 × execution speedup and 6.84 × throughput improvement.
Liekang Zeng, Ke Luo 0001, Xiaoxi Zhang 0001, Zhi Zhou 0006, Xu Chen 0004
WWW1
2022 Edge Robotics: Edge-Computing-Accelerated Multirobot Simultaneous Localization and Mapping
Liekang Zeng, Xu Chen 0004, Ke Luo 0001, Zhi Zhou 0006, Shuai Yu 0001
IEEE Internet Things J.2
2021 Joint Multiuser DNN Partitioning and Computational Resource Allocation for Collaborative Edge Intelligence
abstract
Mobile-edge computing (MEC) has emerged as a promising supporting architecture providing a variety of resources to the network edge, thus acting as an enabler for edge intelligence services empowering massive mobile and Internet-of-Things (IoT) devices with artificial intelligence (AI) capability. With the assistance of edge servers, user equipments (UEs) are able to run deep neural network (DNN)-based AI applications, which are generally resource hungry and computation intensive such that an individual UE can hardly afford by itself in real time. However, the resources in each individual edge server are typically limited. Therefore, any resource optimization involving edge servers is by nature a resource-constrained optimization problem and needs to be tackled in such a realistic context. Motivated by this observation, we investigate the optimization problem of DNN partitioning (an emerging DNN offloading scheme) in a realistic multiuser resource-constrained condition that rarely considered in previous works. Despite the extremely large solution space, we reveal several properties of this specific optimization problem of joint multi-UE DNN partitioning and computational resource allocation. We propose an algorithm called iterative alternating optimization (IAO) that can achieve the optimal solution in polynomial time. In addition, we present a rigorous theoretic analysis of our algorithm in terms of time complexity and performance under realistic estimation error. Moreover, we build a prototype that implements our framework and conducts extensive experiments using realistic DNN models, whose results demonstrate its effectiveness and efficiency.
Xu Chen 0004, Liekang Zeng, Shuai Yu 0001, Lin Chen 0002
IEEE Internet Things J.3
2021 CoEdge: Cooperative DNN Inference With Adaptive Workload Partitioning Over Heterogeneous Edge Devices
abstract
Recent advances in artificial intelligence have driven increasing intelligent applications at the network edge, such as smart home, smart factory, and smart city. To deploy computationally intensive Deep Neural Networks (DNNs) on resource-constrained edge devices, traditional approaches have relied on either offloading workload to the remote cloud or optimizing computation at the end device locally. However, the cloud-assisted approaches suffer from the unreliable and delay-significant wide-area network, and the local computing approaches are limited by the constrained computing capability. Towards high-performance edge intelligence, the cooperative execution mechanism offers a new paradigm, which has attracted growing research interest recently. In this paper, we propose CoEdge, a distributed DNN computing system that orchestrates cooperative DNN inference over heterogeneous edge devices. CoEdge utilizes available computation and communication resources at the edge and dynamically partitions the DNN inference workload adaptive to devices' computing capabilities and network conditions. Experimental evaluations based on a realistic prototype show that CoEdge outperforms status-quo approaches in saving energy with close inference latency, achieving up to 25.5% ~ 66.9% energy reduction for four widely-adopted CNN models.
Liekang Zeng, Xu Chen 0004, Zhi Zhou 0006, Lei Yang 0001, Junshan Zhang
IEEE/ACM Trans. Netw.1
2020 Edge AI: On-Demand Accelerating Deep Neural Network Inference via Edge Computing
abstract
As a key technology of enabling Artificial Intelligence (AI) applications in 5G era, Deep Neural Networks (DNNs) have quickly attracted widespread attention. However, it is challenging to run computation-intensive DNN-based tasks on mobile devices due to the limited computation resources. What’s worse, traditional cloud-assisted DNN inference is heavily hindered by the significant wide-area network latency, leading to poor real-time performance as well as low quality of user experience. To address these challenges, in this paper, we proposeEdgent, a framework that leverages edge computing for DNN collaborative inference through device-edge synergy.Edgentexploits two design knobs: (1) DNN partitioning that adaptively partitions computation between device and edge for purpose of coordinating the powerful cloud resource and the proximal edge resource for real-time DNN inference; (2) DNN right-sizing that further reduces computing latency via early exiting inference at an appropriate intermediate DNN layer. In addition, considering the potential network fluctuation in real-world deployment,Edgentis properly design to specialize for both static and dynamic network environment. Specifically, in a static environment where the bandwidth changes slowly,Edgentderives the best configurations with the assist of regression-based prediction models, while in a dynamic environment where the bandwidth varies dramatically,Edgentgenerates the best execution plan through the online change point detection algorithm that maps the current bandwidth state to the optimal configuration. We implementEdgentprototype based on the Raspberry Pi and the desktop PC and the extensive experimental evaluations demonstrateEdgent’s effectiveness in enabling on-demand low-latency edge intelligence.
Liekang Zeng, Zhi Zhou 0006, Xu Chen 0004
IEEE Trans. Wirel. Commun.2
2019 Cost-Aware Edge Resource Probing for Infrastructure-Free Edge Computing: From Optimal Stopping to Layered Learning
abstract
To meet the stringent requirement of artificial intelligence applications, such as face recognition and video streaming analytics, a resource-constrained device can offload its task to nearby resource-rich devices in edge computing. Resource awareness, as a prime prerequisite for offloading decision-making, is critical for achieving efficient collaborative computation performance. In this paper, we consider cost-aware edge resource probing (CERP) framework design for infrastructure-free edge computing wherein a task device self-organizes its resource probing for informed computation offloading. We first propose a multi-stage optimal stopping formulation for the problem, and derive the optimal probing strategy which reveals a nice multi-threshold structure. Accordingly, we then devise a data-driven layered learning mechanism for more practical and complicated application environments. Layered learning enables the task device to adaptively learn the optimal probing sequence and decision thresholds at runtime, aiming at deriving a good balance between the gain of choosing the best edge device and the accumulated cost of deep resource probing. We further conduct thorough performance evaluation of the proposed CERP schemes using both extensive numerical simulations and realistic system prototype implementation, which demonstrate the superior performance of CERP in the diverse application scenarios.
Tao Ouyang, Xu Chen 0004, Liekang Zeng, Zhi Zhou 0006
RTSS3
2019 Edge Intelligence: Paving the Last Mile of Artificial Intelligence With Edge Computing
abstract
With the breakthroughs in deep learning, the recent years have witnessed a booming of artificial intelligence (AI) applications and services, spanning from personal assistant to recommendation systems to video/audio surveillance. More recently, with the proliferation of mobile computing and Internet of Things (IoT), billions of mobile and IoT devices are connected to the Internet, generating zillions bytes of data at the network edge. Driving by this trend, there is an urgent need to push the AI frontiers to the network edge so as to fully unleash the potential of the edge big data. To meet this demand, edge computing, an emerging paradigm that pushes computing tasks and services from the network core to the network edge, has been widely recognized as a promising solution. The resulted new interdiscipline, edge AI or edge intelligence (EI), is beginning to receive a tremendous amount of interest. However, research on EI is still in its infancy stage, and a dedicated venue for exchanging the recent advances of EI is highly desired by both the computer system and AI communities. To this end, we conduct a comprehensive survey of the recent research efforts on EI. Specifically, we first review the background and motivation for AI running at the network edge. We then provide an overview of the overarching architectures, frameworks, and emerging key technologies for deep learning model toward training/inference at the network edge. Finally, we discuss future research opportunities on EI. We believe that this survey will elicit escalating attentions, stimulate fruitful discussions, and inspire further research ideas on EI.
Zhi Zhou 0006, Xu Chen 0004, Liekang Zeng, Ke Luo 0001, Junshan Zhang
Proc. IEEE4