Tao Ouyang

dblp:227/3029 · DBLP profile ↗
← Back
41ranked-venue papers
9as first author
37since 2021 · last 2026
0000-0002-9380-7238ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 23 · 7 first-author · 20 since 2021Systems, architecture and hardware · 14 · 14 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Chimera: An Efficient Multimodal Embodied Inference Framework with Complexity-Aware Edge-Cloud Routing
Muen Xue, Liekang Zeng, Tao Ouyang, Shaoyong Guo, Xu Chen 0004
ICDCS4
2026 GALA: Generative Aligned Learning for Adaptive Multimodal Representation in the Taobao Shangou Recommender System
abstract
Modern recommender systems in food delivery increasingly leverage multimodal signals, including images, text, and user interaction histories, to enhance user experience, yet effective fusion of these heterogeneous modalities remains challenging, hindering both the joint modeling of multimodal signals and adaptation to evolving user intent. In mainstream two-stage approaches, the separation between content-semantic pretraining of image-text encoders and behavior-driven ranking models limits alignment between semantic understanding and user behavior patterns. To address these issues, we present GALA, a three-stage pipeline whose core innovation lies in an intermediate "generative RL alignment" stage that constructs multimodal pretraining data from user behavior and refines it via conversion-based rewards, effectively bridging the pretraining-fine-tuning gap to align with downstream objectives. GALA comprises three stages: first, behavior-aware triplet pretraining on query-image-text pairs from search logs to early capture user intent and content preferences; second, a novel intermediate stage that refines multimodal embeddings through reward-driven optimization (GRPO) to dynamically align them with user behavior and bridge the pretraining-fine-tuning gap; and finally, integration of multimodal and ID embeddings via adaptive gating with a hybrid loss, preserving multimodal contributions under long-term ID-dominant training. GALA has been deployed in the production environment at Taobao Shangou, serving over 200 million daily active users. Compared with state-of-the-art (SOTA) methods, it delivers consistent offline gains of +0.12/+0.20 AUC along with better PCOC metrics. Large-scale online A/B tests further report a 0.55 percent increase in order volume, confirming GALA's effectiveness at industrial scale and its robustness across diverse demand patterns.
Jiping Liu, Zhongmin Zhang, Zisen Sang, Zhijia Fang, Tao Ouyang, Ma Jiang, Shaopeng Liang, Zeyang Hou, Guodong Cao
ICDE5
2026 CoDrone: Autonomous Drone Navigation Assisted by Edge and Cloud Foundation Models
abstract
Autonomous navigation for Unmanned Aerial Vehicles (UAVs) presents significant challenges due to the limited onboard computational resources, which often restrict deployed deep neural networks to shallow architectures incapable of handling complex environments. Additionally, offloading tasks to remote edge servers introduces high latency, creating an inherent trade-off in system design. To address these limitations, we propose CoDrone—the first cloud-edge-end collaborative computing framework that integrates foundation models into autonomous UAV cruising scenarios—effectively leveraging foundation models to enhance the performance of resource-constrained unmanned aerial vehicle platforms. To reduce both onboard computation and data transmission overhead, CoDrone employs grayscale imagery for the navigation model. When enhanced environmental perception is required, CoDrone leverages the edge-assisted foundation model Depth Anything V2 for depth estimation and introduces a novel, one-dimensional occupancy grid–based navigation method—enabling fine-grained scene understanding while significantly advancing the efficiency and representational simplicity of autonomous navigation. A key component of CoDrone is a Deep Reinforcement Learning (DRL)-based neural scheduler that seamlessly integrates depth estimation with autonomous navigation decisions, enabling real-time adaptation to dynamic environments. Furthermore, the framework introduces a UAV-specific vision language interaction module, which incorporates domain-tailored low-level flight primitives to enable effective interaction between the cloud foundation model, the Vision Language model, and the UAV. The introduction of VLM enhances open-set reasoning capabilities in complex and previously unseen scenarios. We implement a prototype of CoDrone and conduct extensive evaluations in the AirSim simulation environment. Experimental results demonstrate that CoDrone significantly outperforms baseline methods under varying flight speeds and network conditions, achieving a 40% increase in average flight distance and a 5% improvement in average Quality of Navigation.
Tao Ouyang, Ke Luo 0001, Weijie Hong, Xu Chen 0004
IEEE Internet Things J.2
2026 Communication-Efficient Personalized Federated Learning With Incentive-Driven Adaptive Model Pruning and Neighbor Selection
abstract
Personalized Federated Learning (PFL) enables client-specific models to address data heterogeneity but suffers from high communication overhead and unstable participation in resource-constrained and self-interested environments. Existing mainstream approaches predominantly prioritize training process efficiency but lack explicit consideration of incentive mechanisms and rational client behaviors, which may lead to clients behaving conservatively, reducing participation and limiting the effectiveness and scalability of PFL in real-world deployments. In this paper, we proposeIncenPNS, the first incentive-driven adaptive framework that jointly optimizes model pruning, neighbor selection, and incentive mechanisms for communication-efficient PFL.IncenPNSformulates the joint design as a unified optimization objective that balances personalization performance, communication efficiency, and incentive utility under dynamic and heterogeneous environments. To solve the resulting coupled and high-dimensional decision problem, we develop a Multi-Agent Soft Actor-Critic (MASAC)-based learning algorithm that enables clients to adapt pruning rates and collaboration decisions through online interaction. Moreover, a budget-balanced incentive mechanism is incorporated to operate under partial observability, aligning individual rationality with system-level objectives and ensuring reliable participation of self-interested clients. Extensive experiments on representative benchmarks show thatIncenPNSachieves up to 34.8% reduction in communication cost, 68.8% faster convergence, and 83.3% improvement in personalization accuracy compared with state-of-the-art baselines.
Ting Li 0023, Huiting Mo, Tao Ouyang, Yinlong Liu, Kai Yang 0037
IEEE Internet Things J.3
2025 MFEL-HAM: Multimodal Federated Edge Learning with Heterogeneity-Aware Modality Balancing
abstract
The proliferation of Edge Intelligence (EI) and diverse user demands has led to the generation of vast amounts of heterogeneous multimodal data at the network edge. Multimodal Federated Learning (MFL) offers a promising solution for intelligent and personalized services by enabling collaborative training across distributed clients while preserving data privacy. However, existing MFL frameworks remain unsuitable for edge deployment, as they primarily assume homogeneous environments and fail to address heterogeneous client resources. To bridge this gap, we propose$M$ultimodal$F$ederated$E$dge$L$earning (MFEL), a novel paradigm that extends the conventional MFL framework to enable adaptive submodel deployment based on client capabilities. Building on MFEL, we propose MFEL-HAM, a heterogeneous-aware MFL approach that incorporates three core mechanisms: (1) Prototype Networks to align cross-client modality-specific representations, mitigating divergences caused by non-IID data and heterogeneous sensing environments; (2) Rebalanced Modality Gradient Modulation (R-MGM), which adaptively amplifies gradients of underrepresented modalities and suppresses those of dominant ones, alleviating intra-client modality imbalance; and (3) Momentum Knowledge Distillation (MKD), enabling efficient knowledge transfer without sharing raw data, effectively mitigating the impact of resource heterogeneity on collaborative training. Extensive experiments on heterogeneous multimodal datasets show that MFEL-HAM consistently outperforms baselines in accuracy, convergence speed, and training stability, while demonstrating strong generalization across diverse architectures and resource profiles.
Shihan Chen, Hui Jiang 0015, Tao Ouyang, Xu Chen 0004
ICPADS3
2025 AdaRAG: Adaptive Optimization for Retrieval Augmented Generation with Multilevel Retrievers at the Edge
Tao Ouyang, Guihang Hong, Kongyange Zhao, Zhi Zhou 0006, Weigang Wu, Zhaobiao Lv, Xu Chen 0004
INFOCOM1
2025 Towards Federated Inference: An Online Model Ensemble Framework for Cooperative Edge AI
Zhi Zhou 0006, Mengke Huang, Tao Ouyang, Fangming Liu, Xu Chen 0004
INFOCOM4
2025 CoEdge-RAG: Optimizing Hierarchical Scheduling for Retrieval-Augmented LLMs in Collaborative Edge Computing
abstract
Motivated by the imperative for real-time responsiveness and data privacy preservation, large language models (LLMs) are increasingly deployed on resource-constrained edge devices to enable localized inference. To improve output quality, retrieval-augmented generation (RAG) is an efficient technique that seamlessly integrates local data into LLMs. However, existing edge computing paradigms primarily focus on single-node optimization, neglecting opportunities to holistically exploit distributed data and heterogeneous resources through cross-node collaboration. To bridge this gap, we propose CoEdge-RAG, a hierarchical scheduling framework for retrieval-augmented LLMs in collaborative edge computing. In general, privacy constraints preclude accurate a priori acquisition of heterogeneous data distributions across edge nodes, directly impeding RAG performance optimization. Thus, we first design an online query identification mechanism using proximal policy optimization (PPO), which autonomously infers query semantics and establishes cross-domain knowledge associations in an online manner. Second, we devise a dynamic inter-node scheduling strategy that balances workloads across heterogeneous edge nodes by synergizing historical performance analytics with real-time resource thresholds. Third, we develop an intra-node scheduler based on online convex optimization, adaptively allocating query processing ratios and memory resources to optimize the latency-quality trade-off under fluctuating assigned loads. Comprehensive evaluations across diverse QA benchmarks demonstrate that our proposed method significantly boosts the performance of collaborative retrieval-augmented LLMs, achieving performance gains of 4.23 % to 91.39% over baseline methods across all tasks.
Guihang Hong, Tao Ouyang, Kongyange Zhao, Zhi Zhou 0006, Xu Chen 0004
RTSS2
2025 Efficient Multitask Asynchronous Federated Learning in Edge Computing: A Two-Layer Optimization Approach
abstract
Advances in hardware and AI have enabled edgebased IoT devices to leverage substantial computational and data resources, facilitating large-scale deployment of AI models, particularly through federated learning (FL). However, the high heterogeneity of devices and resource contention at the edge make collaborative optimization of resource scheduling for multiple FL tasks challenging. To tackle this, we propose a novel Multi-Task Asynchronous Federated Learning (MTAFL) architecture, which enhances resource utilization efficiency by enabling orthogonal multiplexing of computation and communication resources through adjusting local epochs on edge devices. Then, we formulate an optimization problem in the MTAFL framework to manage resources and local epochs, aiming to minimize energy consumption while achieving FL performance. However, intricate couplings between resource allocation and local control complicate the long-term FL process. To address this, we employ a two-step relaxation approach and develop an efficient optimization strategy based on the block coordinate descent algorithm. To enhance optimization granularity, we extend the MTAFL framework by incorporating device-level data characteristics. We propose a Gaussian Process-based client selection mechanism that dynamically characterizes and predicts training loss trajectories across clients. After selecting clients for each task, we optimize resource allocation and local control strategies in the system. Extensive numerical evaluations corroborate the superior performance of the proposed approaches over existing schemes.
Hui Jiang 0015, Tao Ouyang, Kongyange Zhao, Xu Chen 0004
IEEE Internet Things J.3
2025 Multi-Hop Task Offloading and Relay Selection for IoT Devices in Mobile Edge Computing
abstract
To bridge the gap of conventional single-hop task offloading schemes in infrastructure-free scenarios, multi-hop task offloading schemes for IoT devices in Mobile Edge Computing (MEC) are desired to jointly optimize task offloading decisions and routing paths. In this paper, we investigate a hierarchical multi-hop edge computing framework and propose a joint Task Offloading and Relay Selection (TORS) scheme. It considers real-time computation at each relay node and employs directional searches to facilitate the task execution and results reporting at the fastest speed. However, finding the optimal TORS solution is a formidable challenge due to the time-varying network environments, the strong interdependence of decision sets across different time slots, and the high computational complexity. To address these challenges, we first leverage Lyapunov optimization to transform the stochastic TORS problem into a deterministic per-slot block problem, avoiding the need for extensive system prior knowledge. Subsequently, we propose a Soft Actor-Critic (SAC)-based algorithm, SAC-TORS, to find a satisfactory TORS solution with minimal computational complexity in a distributed manner. Accordingly, each IoT device can independently make self-determined and directional decisions with observable network information. Through extensive experiments, we demonstrate that the SAC-TORS outperforms state-of-the-art solutions, achieving performance improvements of up to 66%.
Ting Li 0023, Yinlong Liu, Tao Ouyang, Hangsheng Zhang, Kai Yang 0037, Xu Zhang 0006
IEEE Trans. Mob. Comput.3
2025 Efficient Coordination of Federated Learning and Inference Offloading at the Edge: A Proactive Optimization Paradigm
abstract
Benefiting from hardware upgrades and deep learning techniques, more and more end devices can independently support a variety of intelligent applications. Further powered by edge computing technologies, the end-edge collaboration paradigm becomes one mainstream approach for achieving advanced edge intelligence (EI). To fully exploit the system resources, it is desirable to coordinate diverse EI services efficiently. Thus, we present a novel framework to jointly optimize the cost-performance trade-off for two distinct but typical EI services, where end devices simultaneously perform federated learning (FL) model training and conduct model inference with the assistance of edge offloading. However, balancing the long-term cost-performance trade-off is highly non-trivial, especially in the absence of knowledge of future system dynamics. Moreover, the capacity heterogeneity further increases the difficulty of service coordination among resource-limited end devices. To overcome these challenges, we first analyze the optimality of inference offloading decisions with and without FL model training and quantify their mutual effects due to local resource contention. By incorporating the loss estimation of FL training model, we then propose a novel proactive policy with theoretical guarantees, which proactively controls the stopping of FL training procedure to balance well the trade-offs between FL model performance and resource costs while fulfilling the inference performance requirements. Extensive results show the efficiency and robustness of our proposed algorithm for EI service coordination in dynamic end-edge collaboration scenarios.
Ke Luo 0001, Kongyange Zhao, Tao Ouyang, Xiaoxi Zhang 0001, Zhi Zhou 0006, Xu Chen 0004
IEEE Trans. Mob. Comput.3
2025 Adaptive Dynamic Scaling and Request Routing Optimization in the Multi-Edge Cluster Collaboration
abstract
With the rapid proliferation of mobile devices, a growing number of intelligent applications are being deployed at the network edge, placing immense strain on the processing capabilities of edge computing. Therefore, resourceconstrained edge servers frequently experience overload due to highly dynamic workloads. To address this, one approach involves forwarding user requests to the cloud or other edge servers, albeit at the cost of increased transmission latency. Alternatively, dynamic scaling of edge clusters can be employed to enhance processing capacity, thereby mitigating latency but at the expense of additional service configuration and hosting expenses. By integrating their complementary benefits, we study the joint optimization problem of dynamic scaling and request routing within a multi-edge cluster collaborative framework, which fully exploits cluster resources to manage the temporal and spatial varying edge workloads. This collaborative framework aims to minimize overall request latency while satisfying an acceptable time-averaged budget cost. However, the complex coupling between scaling and routing decisions, along with the uncertainty of future system information (e.g., user request workloads) impedes the derivation of an optimal offline policy over the long term. Thus, considering the different decision granularities, we employ the two-timescale Lyapunov optimization technique to decouple the original problem into a series of independent online optimization problems with the current system state. In particular, we make cluster scaling decisions in each large timescale and request routing decisions in each small timescale. Given that the decoupled large-timescale subproblems involve NP-hard mixed-integer linear programming, we design an edge resource-aware greedy rounding algorithm to efficiently produce approximate optimal solutions. Finally, both rigorous theoretical analysis and extensive trace-driven evaluations demonstrate the superiority of our proposed algorithm over its counterparts.
Tao Ouyang, Jie Gong 0003, Chao Hong, Xu Chen 0004
IEEE Trans. Mob. Comput.2
2025 Dynamic Edge-Centric Resource Provisioning for Online and Offline Services Co-Location via Reactive and Predictive Approaches
abstract
Due to the penetration of edge computing, a wide variety of workloads are sunk down to the network edge to alleviate huge pressure of the cloud. With the presence of high input workload dynamics and intensive edge resource contention, it is highly non-trivial for an edge proxy to optimize the scheduling of heterogeneous services with diverse QoS requirements. In general, online services should be quickly completed in a quite stable running environment to meet their tight latency constraint, while offline services can be processed loosely for their elastic soft deadlines. To well coordinate such services at the resource-limited edge cluster, in this paper, we study an edge-centric resource provisioning optimization for dynamic online and offline services co-location, where the proxy seeks to maximize timely online service performances while maintaining satisfactory long-term offline service performances. However, intricate hybrid couplings for provisioning decisions arise due to heterogeneous constraints of the co-located services and their different time-scale performances. We hence first propose a reactive provisioning approach without requiring a prior knowledge of future system dynamics, which leverages a Lagrange relaxation for devising constraint-aware stochastic subgradient algorithm to deal with the challenge of hybrid couplings. To further boost the performance by integrating powerful machine learning techniques, we then advocate a predictive provisioning approach, where future request arrivals can be estimated accurately. To align with practical deployments, we incorporate a tunable prediction window mechanism, which well balances the potential improvement and degradation of online performance in imperfect prediction scenarios. With rigorous theoretical analysis and extensive trace-driven evaluations, we show the superior performance of our proposed algorithms for online and offline services co-location at the edge.
Tao Ouyang, Kongyange Zhao, Guihang Hong, Xiaoxi Zhang 0001, Zhi Zhou 0006, Xu Chen 0004
IEEE Trans. Netw.1
2025 Evolving the Cloud Block Store with Performance, Elasticity, Availability, and Hardware Offloading
abstract
In this paper, we qualitatively and quantitatively discuss the design choices, production experience, and lessons in building the Elastic Block Storage ( EBS ) at Alibaba Cloud over the past decade. To cope with hardware advancement and users’ demands, we shift our focus from design simplicity in EBS1 to high performance and space efficiency in EBS2 , and finally reducing network traffic amplification in EBS3 . In addition to the architectural evolutions, we also summarize development lessons and experiences as four topics, including: (i) achieving high elasticity in latency, throughput, IOPS, and capacity; (ii) improving availability by minimizing the blast radius of individual, regional, and global failure events; (iii) identifying the motivations and key tradeoffs in various hardware offloading solutions; and (iv) identifying the pros/cons of alternative solutions and explaining why seemingly promising ideas would not work in practice.
Erci Xu, Weidong Zhang 0011, Qiuping Wang, Yuesheng Gu, Zhenwei Lu, Tao Ouyang, Guanqun Dong, Wenwen Peng, Yilei Peng, Tianyun Wang, Wenyuan Yan, Wenhui Yao, Zhongjie Wu, Lingjun Zhu, Yinhu Wang, Junping Wu, Jiaji Zhu, Jiesheng Wu
ACM Trans. Storage7
2024 What's the Story in EBS Glory: Evolutions and Lessons in Building Cloud Block Store
Weidong Zhang 0011, Erci Xu, Qiuping Wang, Yuesheng Gu, Zhenwei Lu, Tao Ouyang, Guanqun Dai, Wenwen Peng, Yilei Peng, Tianyun Wang, Wenyuan Yan, Wenhui Yao, Zhongjie Wu, Lingjun Zhu, Yinhu Wang, Junping Wu, Jiaji Zhu, Jiesheng Wu
FAST7
2024 SECO: Multi-Satellite Edge Computing Enabled Wide-Area and Real-Time Earth Observation Missions
abstract
Rapid advances in low Earth orbit (LEO) satellite technology and satellite edge computing (SEC) have facilitated a key role for LEO satellites in enhanced Earth observation missions (EOM). These missions (e.g., remote object detection) typically require multi-satellite cooperative observations of a large region of interest (RoI) area, as well as the observation image routing and computation processing, enabling accurate and real-time responsiveness. However, optimizing the resources of LEO satellite networks is nontrivial in the presence of its dynamic and heterogeneous properties. To this end, we propose SECO, a SEC-enabled framework that jointly optimizes multi-satellite observation scheduling, routing and computation node selection for enhanced EOM. Specifically, in the observation phase, we leverage the orbital motion and the rotatable onboard cameras of satellites, and propose a distributed game-based scheduling strategy to minimize the overall size of captured images while ensuring full (observation) coverage. In the sequent routing and computation phase, we first adopt image splitting technology to achieve parallel transmission and computation. Then, we propose an efficient iterative algorithm to jointly optimize image splitting, routing and computation node selection for each captured image. On this basis, we propose a theoretically guaranteed systemwide greedy-based strategy to reduce the total time cost (i.e., transmission, computation and queuing delay) over simultaneous processing for multiple images. Extensive experiments based on real-world datasets demonstrate that SECO can achieve up to a 60.7% reduction in overall time cost compared to baselines.
Zhiwei Zhai, Liekang Zeng, Tao Ouyang, Shuai Yu 0001, Qianyi Huang, Xu Chen 0004
INFOCOM3
2024 Efficient Multi-Task Asynchronous Federated Learning in Edge Computing
abstract
Driven by the continuous upgrading of hardware devices, a notable shift from traditional single-FL task to complicated multi-FL tasks is emerging in edge computing, supporting richer intelligent services. With the presence of high capacity heterogeneity and intensive resource contention at the edge, it is highly non-trivial to collaboratively optimize resource scheduling of multiple FL tasks with diverse QoS requirements. To well tackle the above challenge, we firstly propose a novel Multi-Task Asynchronous Federated Learning (MTAFL) architecture. This novel framework has the potential to enhance resource utilization efficiency by enabling the orthogonal multiplexing of computation and communication resources through adjusting the number of local epochs on edge clients. Then, we formulate an optimization problem within the MTAFL framework, aiming at managing resource allocation and client scheduling to minimize the system-wide energy consumption while achieving the target FL performance. However, intricate couplings for resource allocation and local training decisions arise during the long-term FL process. We hence employ a two-step relaxation approach to transform original non-convex problem into a multi-convex problem, and further devise an efficient optimization strategy based on the block coordinate descent algorithm. Extensive numerical evaluations corroborate the superior performance of the proposed MTAFL framework over existing schemes.
Tao Ouyang, Kongyange Zhao, Yousheng Li, Xu Chen 0004
IWQoS2
2024 FedCarbon: Carbon-Efficient Federated Learning with Double Flexible Controls for Green Edge AI
abstract
The deep integration of federated learning (FL) and edge computing holds great promise in delivering ubiquitous edge AI services. However, in light of the upcoming carbon peaking and neutrality era, existing research has largely overlooked the sustainability challenges of FL in future edge computing. Therefore, we first propose a novel carbon-efficient FL framework in this paper, which leverages client sampling and model pruning approaches to adjust carbon-aware local model training during the long-term FL procedure, adapt to heterogeneous and dynamic edge environments, such as time-varying renewable energy and edge workloads. We then conduct a theoretical analysis of its convergence bound, based on which we introduce an online control algorithm to efficiently balance the trade-off between training performance and carbon emission, i.e., maximizing carbon efficiency while ensuring satisfactory FL performance. The effectiveness of proposed algorithm is verified by extensive trace-driven simulations, reducing up-to 72% carbon emissions than other methods.
Yousheng Li, Tao Ouyang, Xu Chen 0004
IWQoS2
2024 MIX3D: A Mixed Representation for Communication-Efficient Distributed 3DGS Training
abstract
3D Gaussian Splatting (3DGS) has recently emerged as a prominent technique in novel view synthesis. The superior performance of 3DGS has catalyzed an increasing number of 3DGS- based applications in edge scenarios, where 3DGS is utilized for various purposes, such as scene representation, comprehension, and generation. Meanwhile, these edge applications also serve as primary sources of scene observations for producing 3DGS models. However, the intensive computation involved in 3DGS training and the massive number of 3D Gaussian primitives required for high-resolution scene repre-sentation hinder the effectiveness of in-situ 3DGS training on off-the-shelf edge devices, whether using standalone training or Data-Distributed-Parallel (DDP) training. To address this issue, this work proposes MIX3D, a novel mixed representation for communication-efficient distributed 3DGS training in edge scenarios. MIX3D features a global sparse sub-model and various local dense sub-models, where the sparse sub-model encodes coarse-grained appearance for the entire scene, and each dense sub-model targets fine-grained details for a specific region of the scene. Extensive evaluations on a four-device edge cluster demonstrate the effectiveness of our developed distributed 3DGS training workflow based on MIX3D, achieving reductions in training time up to 86.6% compared to vanilla DDP training and an average speedup of 3.767x over standalone training.
Ke Luo 0001, Kongyange Zhao, Shengyuan Ye, Tao Ouyang, Xu Chen 0004
MSN4
2024 MEGA: Mesh-Aligned 3DGS Towards Geometry-Preserving Online Reconstruction
Ke Luo 0001, Shengyuan Ye, Tao Ouyang, Zhi Zhou 0006
NPC (1)3
2024 Hydra: Hybrid-model federated learning for human activity recognition on heterogeneous devices
Tao Ouyang, Qiong Wu 0009, Qianyi Huang, Jie Gong 0003, Xu Chen 0004
J. Syst. Archit.2
2024 Online Optimization of DNN Inference Network Utility in Collaborative Edge Computing
abstract
Collaborative Edge Computing (CEC) is an emerging paradigm that collaborates heterogeneous edge devices as a resource pool to compute DNN inference tasks in proximity such as edge video analytics. Nevertheless, as the key knob to improve network utility in CEC, existing works mainly focus on the workload routing strategies among edge devices with the aim of minimizing the routing cost, remaining an open question for joint workload allocation and routing optimization problem from a system perspective. To this end, this paper presents a holistic, learned optimization for CEC towards maximizing the total network utility in an online manner, even though the utility functions of task input rates are unknown a priori. In particular, we characterize the CEC system in a flow model and formulate an online learning problem in a form of cross-layer optimization. We propose a nested-loop algorithm to solve workload allocation and distributed routing iteratively, using the tools of gradient sampling and online mirror descent. To improve the convergence rate over the nested-loop version, we further devise a single-loop algorithm. Rigorous analysis is provided to show its inherent convexity, efficient convergence, as well as algorithmic optimality. Finally, extensive numerical simulations demonstrate the superior performance of our solutions.
Rui Li 0062, Tao Ouyang, Liekang Zeng, Guocheng Liao, Zhi Zhou 0006, Xu Chen 0004
IEEE/ACM Trans. Netw.2
2024 Cost-Aware Dispersed Resource Probing and Offloading at the Edge: A User-Centric Online Layered Learning Approach
abstract
To meet the stringent requirement of edge intelligence applications, resource-constrained devices can offload their task to nearby resource-rich devices. Resource awareness, as a prime prerequisite for offloading decision-making, is critical for achieving efficient collaborative computation performance. Although major works have explored computation offloading in dynamic edge environments, the impact of fresh resource information perception has not been formally investigated. To bridge the gap, we design a cost-aware edge resource probing (CERP) framework for infrastructure-free edge computing, where a task device self-organizes its resource probing to enable informed computation offloading. We first formulate the joint optimization of device probing and offloading as a multi-stage optimal stopping problem and derive a multi-threshold-based optimal strategy with theoretical guarantees. Accordingly, we devise a data-driven layered learning mechanism to handle more complex real-world scenarios. The layered learning enables the task device to adaptively learn the optimal probing sequence and decision thresholds on the fly, aiming to strike a good balance between the gain of choosing the best edge device and the accumulated cost of deep resource probing. To further boost its learning efficiency, we replace the$\epsilon$-greedy method with a tailored UCB-based adaptive exploration scheme in layered learning, thus better navigating the exploration and exploitation trade-off during probing processes. Finally, we conduct a thorough performance evaluation of the proposed CERP schemes using both extensive numerical simulations and realistic system prototype implementation, which demonstrate the superior performance of CERP in diverse application scenarios.
Tao Ouyang, Xu Chen 0004, Liekang Zeng, Zhi Zhou 0006
IEEE Trans. Serv. Comput.1
2023 Behavior Tree-based Workflow Modeling and Scheduling for Serverless Edge Computing
abstract
Despite the popularity of Serverless computing, there are insufficient efforts dedicated to Serverless workflows (i.e., Serverless function orchestration), particularly for Serverless edge computing. In this paper, we first identify the challenges of deploying the state-of-the-art cloud-oriented Serverless workflow scheduling on resource-constrained edge devices, then propose to model Serverless workflows with behavior trees, and finally reveal our key observations and preliminary results for behavior tree-based Serverless workflow scheduling.
Ke Luo 0001, Tao Ouyang, Zhi Zhou 0006, Xu Chen 0004
ICDCS2
2023 Learning to Be Green: Carbon-Aware Online Control for Edge Intelligence with Colocated Learning and Inference
abstract
Edge intelligence is an emerging paradigm that leverages edge computing to pave the last mile delivery of artificial intelligence. While pilot efforts on edge intelligence have mostly focused on the performance and power issues, the sustainability dilemma along with the upcoming carbon peaking and neutrality era has largely been overlooked. To green edge intelligence, we propose a carbon-aware online control framework (CARE) in this paper. CARE colocates learning and inference tasks within an edge node and dynamically adapts their configurations based on the temporal variation of carbon intensity and renewable energy availability. With such a colocation setup, CARE aims to minimize the long-term inference accuracy loss under the long-term carbon emission cap. The underlying long-term optimization problem is nontrivial since it involves uncertain information (e.g., renewable energy availability) and is NP-hard. To address these dual challenges, CARE first designs an online learning module to make fractional decisions by learning from previous system dynamics and configuration adaptation results. Then, CARE further designs a randomized rounding module, which converts the fractional decision into integer without violating the long-term carbon emission cap. The effectiveness of CARE is verified by rigorous theoretical analysis and extensive trace-driven simulations.
Shuomiao Su, Zhi Zhou 0006, Tao Ouyang, Ruiting Zhou, Xu Chen 0004
ICDCS3
2023 Fair DNN Model Selection in Edge AI via A Cooperative Game Approach
abstract
Edge intelligence is an emerging paradigm that leverages edge computing to pave the last-mile delivery of artificial intelligence (AI). To adapt to the resource restriction, model selection which adaptively selects DNN model variants is widely applied to shape the resource demand of edge AI inference tasks. Unfortunately, in current edge AI serving systems, applications are suffering unfairness since the DNN model selection is performed in a best-effort manner to maximize the system-wide inference accuracy. To achieve a predictable inference accuracy for the applications, edge AI serving systems should guarantee the minimum inference accuracy in a fair fashion at the application level. At the same time, edge resources should be efficiently utilized to minimize operational costs. In this paper, we model the edge DNN model selection problem as a Nash Bargaining Game (NBG), and propose the model selection principles by guaranteeing a base accuracy for each application. Based on the rigorous cooperative game-theoretic approach, we design an approximate algorithm to achieve computationally-efficient and fair model selection, corresponding to the Nash Bargaining Solution (NBS). With extensive trace-driven simulations, we show that our strategy can meet two desirable requirements towards the predictable inference accuracy for applications as well as low operational costs for the system.
Zhi Zhou 0006, Tao Ouyang, Xiaoxi Zhang 0001, Xu Chen 0004
ICDCS3
2023 Dynamic Edge-centric Resource Provisioning for Online and Offline Services Co-location
abstract
Due to the penetration of edge computing, a wide variety of workloads are sunk down to the network edge to alleviate huge pressure of the cloud. With the presence of high input workload dynamics and intensive edge resource contention, it is highly non-trivial for an edge proxy to optimize the scheduling of heterogeneous services with diverse QoS requirements. In general, online services should be quickly completed in a quite stable running environment to meet their tight latency constraint, while offline services can be processed in a loose manner for their elastic soft deadlines. To well coordinate such services at the resource-limited edge cluster, in this paper, we study an edge-centric resource provisioning optimization for dynamic online and offline services co-location, where the proxy seeks to maximize timely online service performances while maintaining satisfactory long-term offline service performances. However, intricate hybrid couplings for provisioning decisions arise due to heterogeneous constraints of the co-located services and their different time-scale performances. We hence first propose a reactive provisioning approach without requiring a prior knowledge of future system dynamics, which leverages a Lagrange relaxation for devising constraint-aware stochastic subgradient algorithm to deal with the challenge of hybrid couplings. To further boost the performance by integrating the powerful machine learning techniques, we also advocate a predictive provisioning approach, where the future request arrivals can be estimated accurately. With rigorous theoretical analysis and extensive trace-driven evaluations, we show the superior performance of our proposed algorithms for online and offline services co-location at the edge.
Tao Ouyang, Kongyange Zhao, Xiaoxi Zhang 0001, Zhi Zhou 0006, Xu Chen 0004
INFOCOM1
2023 QoS-aware Resource Optimization for Hierarchical Cross-Edge Video Analytics
abstract
As the killer application of edge computing, video analytics typically involves multiple vision components in the pipeline, which together determine the quality of service (QoS) for users. By exploiting diverse resource demands of different components, a fine-grained cross-layer orchestration with QoS-aware configuration adaptation can further boost the system efficiency of heterogeneous resources. Thus, we study a video analytics pipeline system with vertical and horizontal resource collaboration across device-edge-cloud hierarchy to achieve QoS-aware cost optimization. To judiciously match the component diversity and the resource heterogeneity, we explore smooth configuration adaptation to model a mixed-integer nonlinear problem, which jointly optimizes long-term resource cost and QoS (including accuracy and latency). However, it is nontrivial to efficiently solve such a NP-hard problem in an online manner without the future information as a prior knowledge due to the time-coupling deployment cost caused by fluctuating input traffic. To address the above challenges, we decouple the intractable problem according to the traffic routing constraints. By leveraging the lazy-switching method, we derive the component orchestration decisions for the decoupled subproblems in each slot and further design a dependent rounding scheme to obtain an efficient feasible solution while guaranteeing the knapsack resource constraints. We rigorously analyze the performance guarantee of our online algorithms by a parameterized competitive ratio, and further verify the empirical performance of our approach through extensive trace-driven experiments.
Kongyange Zhao, Zhi Zhou 0006, Tao Ouyang, Mingliao Zhao, Xu Chen 0004
SECON3
2023 BeeFlow: Behavior tree-based Serverless workflow modeling and scheduling for resource-constrained edge clusters
Ke Luo 0001, Tao Ouyang, Zhi Zhou 0006, Xu Chen 0004
J. Syst. Archit.2
2023 Collaboration in Participant-Centric Federated Learning: A Game-Theoretical Perspective
abstract
Federated learning (FL) is a promising distributed framework for collaborative artificial intelligence model training while protecting user privacy. A bootstrapping component that has attracted significant research attention is the design of incentive mechanism to stimulate user collaboration in FL. The majority of works adopt a broker-centric approach to help the central operator to attract participants and further obtain a well-trained model. Few works consider forging participant-centric collaboration among participants to pursue an FL model for their common interests, which induces dramatic differences in incentive mechanism design from the broker-centric FL. To coordinate the selfish and heterogeneous participants, we propose a novel analytic framework for incentivizing effective and efficient collaborations for participant-centric FL. Specifically, we respectively propose two novel game models for contribution-oblivious FL (COFL) and contribution-aware FL (CAFL), where the latter one implements a minimum contribution threshold mechanism. We further analyze the uniqueness and existence for Nash equilibrium of both COFL and CAFL games and design efficient algorithms to achieve equilibrium solutions. Extensive performance evaluations show that there exists free-riding phenomenon in COFL, which can be greatly alleviated through the adoption of CAFL model with the optimized minimum threshold.
Guangjing Huang, Xu Chen 0004, Tao Ouyang, Qian Ma 0002, Lin Chen 0002, Junshan Zhang
IEEE Trans. Mob. Comput.3
2023 Adaptive User-Managed Service Placement for Mobile Edge Computing via Contextual Multi-Armed Bandit Learning
abstract
Mobile Edge Computing (MEC), envisioned as a cloud extension, pushes cloud resource from the network core to the network edge, thereby meeting the stringent service requirements of many emerging computation-intensive mobile applications. Many existing works have focused on studying the system-wide MEC service placement issues, personalized service performance optimization yet receives much less attention. As motivated, in this paper we propose a novel adaptive user-managed service placement mechanism, which jointly optimizes a users perceived-latency and service migration cost, weighted by user-specific preferences. We first formulate the user-managed dynamic service placement process with limited system information as a contextual multi-armed bandit learning problem. In particular, we investigate both cases without and with neighboring edge feedbacks, where the later considers edge information sharing for more informed decision making. For both cases, we design lightweight Thompson-sampling based online learning algorithms, which can efficiently assist the user to make adaptive service placement decisions. We further conduct a novel information-directed theoretical analysis on the regret bound of the proposed online learning algorithms and reveal the structural impact of edge information sharing. Extensive evaluations demonstrate the superior performance gain of the proposed adaptive user-managed service placement mechanism over existing learning schemes.
Tao Ouyang, Xu Chen 0004, Zhi Zhou 0006, Rui Li 0062
IEEE Trans. Mob. Comput.1
2023 HiFlash: Communication-Efficient Hierarchical Federated Learning With Adaptive Staleness Control and Heterogeneity-Aware Client-Edge Association
abstract
Federated learning (FL) is a promising paradigm that enables collaboratively learning a shared model across massive clients while keeping the training data locally. However, for many existing FL systems, clients need to frequently exchange model parameters of large data size with the remote cloud server directly via wide-area networks (WAN), leading to significant communication overhead and long transmission time. To mitigate the communication bottleneck, we resort to the hierarchical federated learning paradigm of HiFL, which reaps the benefits of mobile edge computing and combines synchronous client-edge model aggregation and asynchronous edge-cloud model aggregation together to greatly reduce the traffic volumes of WAN transmissions. Specifically, we first analyze the convergence bound of HiFL theoretically and identify the key controllable factors for model performance improvement. We then advocate an enhanced design of HiFlash by innovatively integrating deep reinforcement learning based adaptive staleness control and heterogeneity-aware client-edge association strategy to boost the system efficiency and mitigate the staleness effect without compromising model accuracy. Extensive experiments corroborate the superior performance of HiFlash in model accuracy, communication reduction, and system efficiency.
Qiong Wu 0009, Xu Chen 0004, Tao Ouyang, Zhi Zhou 0006, Xiaoxi Zhang 0001, Shusen Yang, Junshan Zhang
IEEE Trans. Parallel Distributed Syst.3
2022 Separating Data via Block Invalidation Time Inference for Write Amplification Reduction in Log-Structured Storage
Qiuping Wang, Patrick P. C. Lee, Tao Ouyang, Lilong Huang
FAST4
2022 Edge intelligence in motion: Mobility-aware dynamic DNN inference service migration with downtime in mobile edge computing
Tao Ouyang, Guocheng Liao, Jie Gong 0003, Shuai Yu 0001, Xu Chen 0004
J. Syst. Archit.2
2021 Flying MEC: Online Task Offloading, Trajectory Planning and Charging Scheduling for UAV-Assisted MEC
Tao Ouyang, Zhi Zhou 0006, Xu Chen 0004
ICA3PP (1)2
2021 Controller Placements for Optimizing Switch-to-Controller and Inter-controller Communication Latency in Software Defined Networks
Yuqi Fan 0001, Lunfei Wang, Tao Ouyang, Lei Shi 0011
WASA (1)4
2021 Multijob Associated Task Scheduling for Cloud Computing Based on Task Duplication and Insertion
abstract
With the emergence and development of various computer technologies, many jobs processed in cloud computing systems consist of multiple associated tasks which follow the constraint of execution order. The task of each job can be assigned to different nodes for execution, and the relevant data are transmitted between nodes to complete the job processing. The computing or communication capabilities of each node may be different due to processor heterogeneity, and hence, a task scheduling algorithm is of great significance for job processing performance. An efficient task scheduling algorithm can make full use of resources and improve the performance of job processing. The performance of existing research on associated task scheduling for multiple jobs needs to be improved. Therefore, this paper studies the problem of multijob associated task scheduling with the goal of minimizing the jobs’ makespan. This paper proposes a task Duplication and Insertion algorithm based on List Scheduling (DILS) which incorporates dynamic finish time prediction, task replication, and task insertion. The algorithm dynamically schedules tasks by predicting the completion time of tasks according to the scheduling of previously scheduled tasks, replicates tasks on different nodes, reduces transmission time, and inserts tasks into idle time slots to speed up task execution. Experimental results demonstrate that our algorithm can effectively reduce the jobs’ makespan.
Lei Shi 0011, Lunfei Wang, Zhifeng Jin, Tao Ouyang, Juan Xu 0002, Yuqi Fan 0001
Wirel. Commun. Mob. Comput.6
2019 Adaptive User-managed Service Placement for Mobile Edge Computing: An Online Learning Approach
abstract
Mobile Edge Computing (MEC), envisioned as a cloud extension, pushes cloud resource from the network core to the network edge, thereby meeting the stringent service requirements of many emerging computation-intensive mobile applications. Many existing works have focused on studying the system-wide MEC service placement issues, personalized service performance optimization yet receives much less attention. Thus, in this paper we propose a novel adaptive user-managed service placement mechanism, which jointly optimizes a user's perceived-latency and service migration cost, weighted by user preferences. To overcome the unavailability of future information and unknown system dynamics, we formulate the dynamic service placement problem as a contextual Multi-armed Bandit (MAB) problem, and then propose a Thompson-sampling based online learning algorithm to explore the dynamic MEC environment, which further assists the user to make adaptive service placement decisions. Rigorous theoretical analysis and extensive evaluations demonstrate the superior performance of the proposed adaptive user-managed service placement mechanism.
Tao Ouyang, Rui Li 0062, Xu Chen 0004, Zhi Zhou 0006
INFOCOM1
2019 Cost-Aware Edge Resource Probing for Infrastructure-Free Edge Computing: From Optimal Stopping to Layered Learning
abstract
To meet the stringent requirement of artificial intelligence applications, such as face recognition and video streaming analytics, a resource-constrained device can offload its task to nearby resource-rich devices in edge computing. Resource awareness, as a prime prerequisite for offloading decision-making, is critical for achieving efficient collaborative computation performance. In this paper, we consider cost-aware edge resource probing (CERP) framework design for infrastructure-free edge computing wherein a task device self-organizes its resource probing for informed computation offloading. We first propose a multi-stage optimal stopping formulation for the problem, and derive the optimal probing strategy which reveals a nice multi-threshold structure. Accordingly, we then devise a data-driven layered learning mechanism for more practical and complicated application environments. Layered learning enables the task device to adaptively learn the optimal probing sequence and decision thresholds at runtime, aiming at deriving a good balance between the gain of choosing the best edge device and the accumulated cost of deep resource probing. We further conduct thorough performance evaluation of the proposed CERP schemes using both extensive numerical simulations and realistic system prototype implementation, which demonstrate the superior performance of CERP in the diverse application scenarios.
Tao Ouyang, Xu Chen 0004, Liekang Zeng, Zhi Zhou 0006
RTSS1
2018 Follow Me at the Edge: Mobility-Aware Dynamic Service Placement for Mobile Edge Computing
abstract
Mobile edge computing is a new computing paradigm in which cloud computing capabilities are pushed from the network core to the network edge to serve the end-user in proximity. However, with the sinking of computing capabilities, the new challenge incurred by user mobility arises: since end-users typically move erratically, the services should be dynamically migrated among multiple edges to maintain the service performance, i.e., user-perceived latency. Tackling this problem is non-trivial since frequent service migration would greatly increase the operational cost. To address this challenge in terms performance-cost trade-off, in this paper we study the mobile edge service performance optimization problem under long-term cost budget constraint. To address user mobility which is typically unpredictable, we first apply Lyapunov optimization to decompose the long-term optimization problem into a series of real-time optimization problems which do not require a priori knowledge such as user mobility. As the decomposed problem is NP-hard, we further propose an efficient heuristic based on the Markov approximation technique. Rigorous theoretical analysis and extensive evaluations demonstrate the efficacy of the proposed solution.
Tao Ouyang, Zhi Zhou 0006, Xu Chen 0004
IWQoS1
2018 Follow Me at the Edge: Mobility-Aware Dynamic Service Placement for Mobile Edge Computing
abstract
Mobile edge computing is a new computing paradigm, which pushes cloud computing capabilities away from the centralized cloud to the network edge. However, with the sinking of computing capabilities, the new challenge incurred by user mobility arises: since end users typically move erratically, the services should be dynamically migrated among multiple edges to maintain the service performance, i.e., user-perceived latency. Tackling this problem is non-trivial since frequent service migration would greatly increase the operational cost. To address this challenge in terms of the performance-cost tradeoff, in this paper, we study the mobile edge service performance optimization problem under long-term cost budget constraint. To address user mobility which is typically unpredictable, we apply Lyapunov optimization to decompose the long-term optimization problem into a series of real-time optimization problems which do not require a priori knowledge such as user mobility. As the decomposed problem is NP-hard, we first design an approximation algorithm based on Markov approximation to seek a near-optimal solution. To make our solution scalable and amenable to future fifth-generation application scenario with large-scale user devices, we further propose a distributed approximation scheme with greatly reduced time complexity, based on the technique of the best response update. Rigorous theoretical analysis and extensive evaluations demonstrate the efficacy of the proposed centralized and distributed schemes.
Tao Ouyang, Zhi Zhou 0006, Xu Chen 0004
IEEE J. Sel. Areas Commun.1