EDBT 2026 Demo / reviewers in the wild / expert
Kongyange Zhao
dblp:353/2876
· DBLP profile ↗
16ranked-venue papers
4as first author
16since 2021 · last 2026
0000-0003-1927-782XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 12 · 3 first-author · 12 since 2021Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Edge-Assisted Real-Time Dynamic 3D Point Cloud Rendering for Multi-Party Mobile Virtual RealityabstractMulti-party Mobile Virtual Reality (MMVR) enables multiple mobile users to share virtual scenes for an immersive multimedia experience in scenarios such as gaming, social interaction, and industrial mission collaboration. Dynamic 3D Point Cloud (DPCL) is an emerging representation form of MMVR that can be consumed as a free-viewpoint video with 6 degrees of freedom. With limited on-device resources, it is a challenge to achieve a satisfying frame rate for DPCL rendering, which makes edge-assisted rendering a practical solution. However, repeated loading of DPCL scenes with a substantial amount of metadata introduces a significant redundancy overhead that cannot be overlooked when enabling multiple edge servers to support the rendering requirements of user groups. In this paper, we design PoClVR, an edge-assisted DPCL rendering system for MMVR applications, which introduces an object-level splitting mode to alleviate performance bottlenecks caused by redundant loading. In addition, PoClVR dynamically selects the splitting mode and scheduling decisions to adapt to varying task requirements and available computational resources, thereby improving overall system efficiency. To evaluate the performance of PoClVR, we implement and deploy a realistic prototype system and also conduct large-scale trace-driven simulations. The experimental results show that PoClVR can reduce resource usage by up to approximately 49.3% under different task requirements and resource conditions, while decreasing bottleneck performance degradation by up to 77.3%. Ximing Wu, Kongyange Zhao, Xu Chen 0004, Teng Liang, Weizhe Zhang |
IEEE Trans. Mob. Comput. | 2 |
| 2026 | MicroEdge: An Online Optimization Framework for Cost-Efficient Microservice Orchestration in Edge Native ApplicationsabstractThe rapid proliferation of edge computing infrastructure has significantly accelerated the adoption of edge-native applications, ranging from autonomous vehicles to augmented reality and real-time analytics. Microservice, renowned for its lightweight, loosely coupled, and modular architecture, has emerged as the de-facto standard for developing edge native applications. However, the resource scarcity and heterogeneity, coupled with request dynamics in edge environments, pose substantial challenges for effective microservice orchestration. To address these challenges, we propose MicroEdge, an online optimization framework designed for cost-efficient microservice orchestration in edge-native environments. MicroEdge employs a multi-level optimization approach by strategically coordinating four key dimensions: microservice placement, layer placement, layer pulling, and user request scheduling. The framework confronts two fundamental challenges in solving this joint optimization problem: (1) the time-coupled nature of long-term holistic cost minimization, and (2) the NP-hardness of the underlying problem. MicroEdge tackles these dual challenges by integrating a regularization method for online algorithm design and a dependent rounding technique for approximation algorithm design. Both rigorous theoretical analysis and extensive simulations driven by realistic Alibaba microservice workload traces validate the efficacy of MicroEdge. Weihan Zeng, Kongyange Zhao, Jianxiong Liao, Zhi Zhou 0006, Deke Guo, Xu Chen 0004 |
IEEE Trans. Mob. Comput. | 2 |
| 2026 | OSGS: A Framework for Online Scheduling of Satellite-Ground Collaborative Inference With Space Edge Computing
Kongyange Zhao, Yuanming Wang, Zhi Zhou 0006, Ruiting Zhou, Xiaoxi Zhang 0001, Xu Chen 0004, Dechao Ran, Fei Zhang 0005, Lu Cao 0001 |
IEEE Trans. Serv. Comput. | 1 |
| 2025 | AdaRAG: Adaptive Optimization for Retrieval Augmented Generation with Multilevel Retrievers at the Edge
Tao Ouyang, Guihang Hong, Kongyange Zhao, Zhi Zhou 0006, Weigang Wu, Zhaobiao Lv, Xu Chen 0004 |
INFOCOM | 3 |
| 2025 | CoEdge-RAG: Optimizing Hierarchical Scheduling for Retrieval-Augmented LLMs in Collaborative Edge ComputingabstractMotivated by the imperative for real-time responsiveness and data privacy preservation, large language models (LLMs) are increasingly deployed on resource-constrained edge devices to enable localized inference. To improve output quality, retrieval-augmented generation (RAG) is an efficient technique that seamlessly integrates local data into LLMs. However, existing edge computing paradigms primarily focus on single-node optimization, neglecting opportunities to holistically exploit distributed data and heterogeneous resources through cross-node collaboration. To bridge this gap, we propose CoEdge-RAG, a hierarchical scheduling framework for retrieval-augmented LLMs in collaborative edge computing. In general, privacy constraints preclude accurate a priori acquisition of heterogeneous data distributions across edge nodes, directly impeding RAG performance optimization. Thus, we first design an online query identification mechanism using proximal policy optimization (PPO), which autonomously infers query semantics and establishes cross-domain knowledge associations in an online manner. Second, we devise a dynamic inter-node scheduling strategy that balances workloads across heterogeneous edge nodes by synergizing historical performance analytics with real-time resource thresholds. Third, we develop an intra-node scheduler based on online convex optimization, adaptively allocating query processing ratios and memory resources to optimize the latency-quality trade-off under fluctuating assigned loads. Comprehensive evaluations across diverse QA benchmarks demonstrate that our proposed method significantly boosts the performance of collaborative retrieval-augmented LLMs, achieving performance gains of 4.23 % to 91.39% over baseline methods across all tasks. Guihang Hong, Tao Ouyang, Kongyange Zhao, Zhi Zhou 0006, Xu Chen 0004 |
RTSS | 3 |
| 2025 | Efficient Multitask Asynchronous Federated Learning in Edge Computing: A Two-Layer Optimization ApproachabstractAdvances in hardware and AI have enabled edgebased IoT devices to leverage substantial computational and data resources, facilitating large-scale deployment of AI models, particularly through federated learning (FL). However, the high heterogeneity of devices and resource contention at the edge make collaborative optimization of resource scheduling for multiple FL tasks challenging. To tackle this, we propose a novel Multi-Task Asynchronous Federated Learning (MTAFL) architecture, which enhances resource utilization efficiency by enabling orthogonal multiplexing of computation and communication resources through adjusting local epochs on edge devices. Then, we formulate an optimization problem in the MTAFL framework to manage resources and local epochs, aiming to minimize energy consumption while achieving FL performance. However, intricate couplings between resource allocation and local control complicate the long-term FL process. To address this, we employ a two-step relaxation approach and develop an efficient optimization strategy based on the block coordinate descent algorithm. To enhance optimization granularity, we extend the MTAFL framework by incorporating device-level data characteristics. We propose a Gaussian Process-based client selection mechanism that dynamically characterizes and predicts training loss trajectories across clients. After selecting clients for each task, we optimize resource allocation and local control strategies in the system. Extensive numerical evaluations corroborate the superior performance of the proposed approaches over existing schemes. Hui Jiang 0015, Tao Ouyang, Kongyange Zhao, Xu Chen 0004 |
IEEE Internet Things J. | 5 |
| 2025 | Efficient Coordination of Federated Learning and Inference Offloading at the Edge: A Proactive Optimization ParadigmabstractBenefiting from hardware upgrades and deep learning techniques, more and more end devices can independently support a variety of intelligent applications. Further powered by edge computing technologies, the end-edge collaboration paradigm becomes one mainstream approach for achieving advanced edge intelligence (EI). To fully exploit the system resources, it is desirable to coordinate diverse EI services efficiently. Thus, we present a novel framework to jointly optimize the cost-performance trade-off for two distinct but typical EI services, where end devices simultaneously perform federated learning (FL) model training and conduct model inference with the assistance of edge offloading. However, balancing the long-term cost-performance trade-off is highly non-trivial, especially in the absence of knowledge of future system dynamics. Moreover, the capacity heterogeneity further increases the difficulty of service coordination among resource-limited end devices. To overcome these challenges, we first analyze the optimality of inference offloading decisions with and without FL model training and quantify their mutual effects due to local resource contention. By incorporating the loss estimation of FL training model, we then propose a novel proactive policy with theoretical guarantees, which proactively controls the stopping of FL training procedure to balance well the trade-offs between FL model performance and resource costs while fulfilling the inference performance requirements. Extensive results show the efficiency and robustness of our proposed algorithm for EI service coordination in dynamic end-edge collaboration scenarios. Ke Luo 0001, Kongyange Zhao, Tao Ouyang, Xiaoxi Zhang 0001, Zhi Zhou 0006, Xu Chen 0004 |
IEEE Trans. Mob. Comput. | 2 |
| 2025 | Dynamic Edge-Centric Resource Provisioning for Online and Offline Services Co-Location via Reactive and Predictive ApproachesabstractDue to the penetration of edge computing, a wide variety of workloads are sunk down to the network edge to alleviate huge pressure of the cloud. With the presence of high input workload dynamics and intensive edge resource contention, it is highly non-trivial for an edge proxy to optimize the scheduling of heterogeneous services with diverse QoS requirements. In general, online services should be quickly completed in a quite stable running environment to meet their tight latency constraint, while offline services can be processed loosely for their elastic soft deadlines. To well coordinate such services at the resource-limited edge cluster, in this paper, we study an edge-centric resource provisioning optimization for dynamic online and offline services co-location, where the proxy seeks to maximize timely online service performances while maintaining satisfactory long-term offline service performances. However, intricate hybrid couplings for provisioning decisions arise due to heterogeneous constraints of the co-located services and their different time-scale performances. We hence first propose a reactive provisioning approach without requiring a prior knowledge of future system dynamics, which leverages a Lagrange relaxation for devising constraint-aware stochastic subgradient algorithm to deal with the challenge of hybrid couplings. To further boost the performance by integrating powerful machine learning techniques, we then advocate a predictive provisioning approach, where future request arrivals can be estimated accurately. To align with practical deployments, we incorporate a tunable prediction window mechanism, which well balances the potential improvement and degradation of online performance in imperfect prediction scenarios. With rigorous theoretical analysis and extensive trace-driven evaluations, we show the superior performance of our proposed algorithms for online and offline services co-location at the edge. Tao Ouyang, Kongyange Zhao, Guihang Hong, Xiaoxi Zhang 0001, Zhi Zhou 0006, Xu Chen 0004 |
IEEE Trans. Netw. | 2 |
| 2024 | Efficient Multi-Task Asynchronous Federated Learning in Edge ComputingabstractDriven by the continuous upgrading of hardware devices, a notable shift from traditional single-FL task to complicated multi-FL tasks is emerging in edge computing, supporting richer intelligent services. With the presence of high capacity heterogeneity and intensive resource contention at the edge, it is highly non-trivial to collaboratively optimize resource scheduling of multiple FL tasks with diverse QoS requirements. To well tackle the above challenge, we firstly propose a novel Multi-Task Asynchronous Federated Learning (MTAFL) architecture. This novel framework has the potential to enhance resource utilization efficiency by enabling the orthogonal multiplexing of computation and communication resources through adjusting the number of local epochs on edge clients. Then, we formulate an optimization problem within the MTAFL framework, aiming at managing resource allocation and client scheduling to minimize the system-wide energy consumption while achieving the target FL performance. However, intricate couplings for resource allocation and local training decisions arise during the long-term FL process. We hence employ a two-step relaxation approach to transform original non-convex problem into a multi-convex problem, and further devise an efficient optimization strategy based on the block coordinate descent algorithm. Extensive numerical evaluations corroborate the superior performance of the proposed MTAFL framework over existing schemes. Tao Ouyang, Kongyange Zhao, Yousheng Li, Xu Chen 0004 |
IWQoS | 3 |
| 2024 | Edge-assisted Real-time Dynamic 3D Point Cloud Rendering for Multi-party Mobile Virtual RealityabstractMulti-party Mobile Virtual Reality (MMVR) enables multiple mobile users to share virtual scenes for immersive multimedia experience in scenarios such as gaming, social interaction, and industrial mission collaboration. Dynamic 3D Point Cloud (DPCL) is an emerging representation form of MMVR that can be consumed as a free-viewpoint video with 6 degrees of freedom. Given that it is challenging to render DPCL at a satisfying frame rate with limited on-device resources, offloading rendering tasks to edge servers is recognized as a practical solution. However, repeated loading of DPCL scenes with a substantial amount of metadata introduces a significant redundancy overhead that cannot be overlooked when enabling multiple edge servers to support the rendering requirements of user groups. In this paper, we design PoClVR, an edge-assisted DPCL rendering system for MMVR applications, which breaks down the rendering process of the complete dynamic scene into multiple rendering tasks of dynamic objects. PoClVR significantly reduces the repetitive loading overhead of DPCL scenes on edge servers and periodically adjusts the rendering task allocation during the application running to accommodate rendering requirements. We deploy PoClVR based on a real-world implementation and the experimental evaluation results show that PoClVR can reduce GPU utilization by up to 15.1% and increase rendering frame rate by up to 34.6% compared to other baselines while ensuring that the image quality viewed by the user is virtually unchanged. Ximing Wu, Kongyange Zhao, Xu Chen 0004, Teng Liang |
ACM Multimedia | 2 |
| 2024 | MIX3D: A Mixed Representation for Communication-Efficient Distributed 3DGS Trainingabstract3D Gaussian Splatting (3DGS) has recently emerged as a prominent technique in novel view synthesis. The superior performance of 3DGS has catalyzed an increasing number of 3DGS- based applications in edge scenarios, where 3DGS is utilized for various purposes, such as scene representation, comprehension, and generation. Meanwhile, these edge applications also serve as primary sources of scene observations for producing 3DGS models. However, the intensive computation involved in 3DGS training and the massive number of 3D Gaussian primitives required for high-resolution scene repre-sentation hinder the effectiveness of in-situ 3DGS training on off-the-shelf edge devices, whether using standalone training or Data-Distributed-Parallel (DDP) training. To address this issue, this work proposes MIX3D, a novel mixed representation for communication-efficient distributed 3DGS training in edge scenarios. MIX3D features a global sparse sub-model and various local dense sub-models, where the sparse sub-model encodes coarse-grained appearance for the entire scene, and each dense sub-model targets fine-grained details for a specific region of the scene. Extensive evaluations on a four-device edge cluster demonstrate the effectiveness of our developed distributed 3DGS training workflow based on MIX3D, achieving reductions in training time up to 86.6% compared to vanilla DDP training and an average speedup of 3.767x over standalone training. Ke Luo 0001, Kongyange Zhao, Shengyuan Ye, Tao Ouyang, Xu Chen 0004 |
MSN | 2 |
| 2024 | Taming Serverless Cold Start of Cloud Model Inference With Edge ComputingabstractServerless computing is envisioned as the de-facto standard for next-generation cloud computing. However, the cold start dilemma has impeded its adoption by delay-sensitive and burst applications. In this paper, we propose to tame serverless cold start in a cloud inference system with edge computing. Specifically, the proposed solution smooths the serverless cloud workload with user-owned edge computing, reducing the number of cold starts. Leveraging the configurability of requests and serverless functions, the proposed solution further reduces the transmission latency and serverless cost by adapting request configuration (e.g., image resolution) and function configuration (e.g., memory). To alleviate the potential inference accuracy degradation incurred by configuration adaption, we aim to strike a nice balance between inference latency, cost, and accuracy. However, achieving this goal is non-trivial since the underlying optimization is non-convex and involves future uncertain information. To simultaneously address dual challenges, the presented cold-start-aware online algorithms apply the regularization technique to decompose the problem into separate convex subproblems. Then, it applies lazy switching to smooth the number of provisioned functions and thus reduces the cold start. Through rigorous theoretical analysis, realistic prototype evaluations on AWS Lambda, and trace-driven simulations, we comprehensively validate the theoretical and empirical performance of our proposed solution. Kongyange Zhao, Zhi Zhou 0006, Lei Jiao 0002, Shen Cai, Fei Xu 0009, Xu Chen 0004 |
IEEE Trans. Mob. Comput. | 1 |
| 2023 | Cost-Efficient Cloud-Edge Video Analytics with Hybird IaaS and FaaS ResourcesabstractVideo analytics services are extensively employed in various real-time applications, including crime monitoring, business intelligence, and traffic flow control. These applications typically depend on cloud data centers to aggregate video stream tasks using Infrastructure as a Service (IaaS). However, the conventional approach of employing limited-term leased virtual machines (VMs) often leads to leads to high costs and idle time. Serverless computing or Function as a Service (FaaS) offers a more adaptable, pay-as-you-go solution but at a higher price, introducing latency challenges. Edge servers can reduce latency and costs, but limited physical resources affect accuracy. Opting for high-quality analytic services with better frame rates and resolutions can maintain accuracy but may not control costs. Therefore, the key challenge in video analytics is choosing the right configuration and computing methods for high accuracy, low cost, and low latency in edge, IaaS and FaaS environments. To address the challenge, we consider a scenario where the video analytics task is offloaded by combining the above three placements (i.e., VM, serverless, and edge node) to optimize the cost and latency while keeping the accuracy. The optimization problem can be formulated as an Integer Programming (IP) problem of which the NP-hardness is proved. To further deal with it, we propose an efficient online algorithm, which can optimize cost and latency using the Lyapunov optimization analysis while ensuring accuracy over long periods. Finally, the multi-angle simulation experiment results show that OCPA can effectively reduce the cost and delay compared with the benchmarks. At the same time, the achieved accuracy is close to the set long-term time-averaged value. Zhi Zhou 0006, Kongyange Zhao, Huirong Ma, Xu Chen 0004 |
ICPADS | 3 |
| 2023 | Dynamic Edge-centric Resource Provisioning for Online and Offline Services Co-locationabstractDue to the penetration of edge computing, a wide variety of workloads are sunk down to the network edge to alleviate huge pressure of the cloud. With the presence of high input workload dynamics and intensive edge resource contention, it is highly non-trivial for an edge proxy to optimize the scheduling of heterogeneous services with diverse QoS requirements. In general, online services should be quickly completed in a quite stable running environment to meet their tight latency constraint, while offline services can be processed in a loose manner for their elastic soft deadlines. To well coordinate such services at the resource-limited edge cluster, in this paper, we study an edge-centric resource provisioning optimization for dynamic online and offline services co-location, where the proxy seeks to maximize timely online service performances while maintaining satisfactory long-term offline service performances. However, intricate hybrid couplings for provisioning decisions arise due to heterogeneous constraints of the co-located services and their different time-scale performances. We hence first propose a reactive provisioning approach without requiring a prior knowledge of future system dynamics, which leverages a Lagrange relaxation for devising constraint-aware stochastic subgradient algorithm to deal with the challenge of hybrid couplings. To further boost the performance by integrating the powerful machine learning techniques, we also advocate a predictive provisioning approach, where the future request arrivals can be estimated accurately. With rigorous theoretical analysis and extensive trace-driven evaluations, we show the superior performance of our proposed algorithms for online and offline services co-location at the edge. Tao Ouyang, Kongyange Zhao, Xiaoxi Zhang 0001, Zhi Zhou 0006, Xu Chen 0004 |
INFOCOM | 2 |
| 2023 | QoS-aware Resource Optimization for Hierarchical Cross-Edge Video AnalyticsabstractAs the killer application of edge computing, video analytics typically involves multiple vision components in the pipeline, which together determine the quality of service (QoS) for users. By exploiting diverse resource demands of different components, a fine-grained cross-layer orchestration with QoS-aware configuration adaptation can further boost the system efficiency of heterogeneous resources. Thus, we study a video analytics pipeline system with vertical and horizontal resource collaboration across device-edge-cloud hierarchy to achieve QoS-aware cost optimization. To judiciously match the component diversity and the resource heterogeneity, we explore smooth configuration adaptation to model a mixed-integer nonlinear problem, which jointly optimizes long-term resource cost and QoS (including accuracy and latency). However, it is nontrivial to efficiently solve such a NP-hard problem in an online manner without the future information as a prior knowledge due to the time-coupling deployment cost caused by fluctuating input traffic. To address the above challenges, we decouple the intractable problem according to the traffic routing constraints. By leveraging the lazy-switching method, we derive the component orchestration decisions for the decoupled subproblems in each slot and further design a dependent rounding scheme to obtain an efficient feasible solution while guaranteeing the knapsack resource constraints. We rigorously analyze the performance guarantee of our online algorithms by a parameterized competitive ratio, and further verify the empirical performance of our approach through extensive trace-driven experiments. Kongyange Zhao, Zhi Zhou 0006, Tao Ouyang, Mingliao Zhao, Xu Chen 0004 |
SECON | 1 |
| 2023 | EdgeAdaptor: Online Configuration Adaption, Model Selection and Resource Provisioning for Edge DNN Inference Serving at ScaleabstractThe accelerating convergence of artificial intelligence and edge computing has sparked a recent wave of interest in edge intelligence. While pilot efforts focused on edge DNN inference serving for a single user or DNN application, scaling edge DNN inference serving to multiple users and applications is however nontrivial. In this paper, we propose an online optimization framework EdgeAdaptor for multi-user and multi-application edge DNN inference serving at scale, which aims to navigate the three-way trade-off between inference accuracy, latency, and resource cost via jointly optimizing the application configuration adaption, DNN model selection and edge resource provisioning on-the-fly. The underlying long-term optimization problem is difficult since it is NP-hard and involves future uncertain information. To address these dual challenges, we fuse the power of online optimization and approximate optimization into a joint optimization framework, via i) decomposing the long-term problem into a series of single-shot fractional problems with a regularization technique, and ii) rounding the fractional solution to a near-optimal integral solution with a randomized dependent scheme. Rigorous theoretical analysis derives a parameterized competition ratio of our online algorithms, and extensive trace-driven simulations verify that its empirical value is no larger than 1.4 in typical scenarios. Kongyange Zhao, Zhi Zhou 0006, Xu Chen 0004, Ruiting Zhou, Xiaoxi Zhang 0001, Shuai Yu 0001, Di Wu 0001 |
IEEE Trans. Mob. Comput. | 1 |