EDBT 2026 Demo / reviewers in the wild / expert
Andong Zhu 0001
dblp:286/6658-1
· DBLP profile ↗
19ranked-venue papers
6as first author
18since 2021 · last 2026
0009-0002-8233-329XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 13 · 5 first-author · 13 since 2021Systems, architecture and hardware · 4 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Octopus: Accuracy-aware resource scheduling for multi-video streaming inference at the edge
Zhuzhong Qian, Andong Zhu 0001, Hesheng Sun, Lingkun Meng |
Comput. Networks | 3 |
| 2026 | Edge-cloud co-optimized 3D video analytics with synergistic neural codecs
Hebin Sun, Xiaohang Shi 0001, Xiaokun Wang 0002, Sheng Zhang 0001, Lingkun Meng, Andong Zhu 0001, Zhuzhong Qian |
Comput. Networks | 7 |
| 2025 | Bridging the Prediction-Decision Gap: Enhancing Model Deployment and Online Service Request Forecasting in Edge Inference SystemsabstractIn edge inference systems, efficient model deployment and accurate forecasting of online service requests are critical for maintaining optimal performance and resource utilization. This study introduces a novel approach that integrates predictive enhancement techniques to improve model deployment strategies and online service request forecasting. By addressing the existing gap between prediction and decision-making processes, our method enables more responsive and adaptive edge computing environments. Experimental results demonstrate significant improvements in both deployment efficiency and forecasting accuracy, highlighting the potential of our approach to advance the state-of-the-art in edge inference system management. Hesheng Sun, Zhuzhong Qian, Andong Zhu 0001 |
IWQoS | 3 |
| 2025 | VidIQ: Inference-Aware Neural Codecs for Quality-Enhanced, Real-Time Video AnalyticsabstractVideo analytics pipelines migrating to edge deployments are facing performance bottlenecks under limited bandwidth. Non-uniform intra-frame encoding emerges to further compress pixels without affecting the output of the server deep neural network (DNN), while it is inefficient in high-resolution video streaming at low bandwidth. The detail enhancement capability of neural super-resolution (SR) permits resolution downsampling and aggressive compression on edge devices for low-latency transmission. To exploit its accuracy potential, DNN-oriented non-uniform encoding is expected to be additionally aware of SR models. However, traditional codecs struggle to cope with both quality optimization for SR and global semantic features for DNN. We advocate neural codecs for coordinated encoding and enhancement, enabling analytic-oriented video streaming with optimal accuracy-delay tradeoffs. Our system, VidIQ, achieves quality-enhanced real-time video analytics by 1) improving the network architecture of neural codecs (at two granularity) to integrate SR models into a DNN-oriented analytics pipeline, and 2) adapting the multi-scale encoder and SR-decoder to scene dynamics (i.e., content and bandwidth variations) with the help of the monolithic controller to hold a performance advantage. Extensive evaluations showcase that VidIQ reduces end-to-end delay by 35.8% and improves analytics accuracy by 21.2% compared to the recent video compression, enhancement, and streaming baselines. Andong Zhu 0001, Sheng Zhang 0001, Xiaohang Shi 0001, Hesheng Sun, Yu Liang 0001, Zhuzhong Qian, Xiaokun Wang 0002 |
ACM Multimedia | 1 |
| 2025 | Decode-What-Matters: Frame-Level Parallel Generative Decoding to Accelerate Large-Scale Video AnalyticsabstractVideo analytics pipelines (VAPs) have been a paradigm for large-scale video analytics. Due to temporal redundancy in video, frame filtering is widely used in VAPs to reduce analysis workload. However, existing works overlook a limitation: while inference operates only on selected frames, decoders must still process many redundant frames due to codec dependencies, leading to over-decoding trap. This limitation stems from the reference-based design in modern codecs, which require decoding preceding frames to reconstruct any selected one. As a result, over-decoding has become the practical bottleneck in VAPs using modern decoders, highlighting a critical but under-explored problem. To address this issue, we propose ParaDeco, a high-throughput video analytics framework featuring a novel frame-level parallel generative decoder. Unlike traditional decoders, ParaDeco adopts a decode-what-matters approach with decoupled frame dependencies. To decode arbitrary frames independently, ParaDeco generates frame-wise features as standalone skeletons using compressed video metadata, then predicts pseudo frames maintaining semantic consistency with original frames. Moreover, ParaDeco identifies which frames truly matter for analysis via delicate contribution-based frame filtering. We implement ParaDeco on a cloud server and evaluate it on large-scale real-world video datasets. Our experimental results show that ParaDeco achieves a 2.76× speedup on average compared to state-of-the-art VAPs. Xiaokun Wang 0002, Sheng Zhang 0001, Andong Zhu 0001, Ning Chen 0010, Yu Chen 0038, Zhuzhong Qian, Sanglu Lu, Yu Liang 0001 |
ACM Multimedia | 4 |
| 2025 | ABUV: Adaptive bitrate and upsampling for video streaming on mobile devices
Ji Qi 0005, Sheng Zhang 0001, Gangyi Luo, Andong Zhu 0001, Jie Wu 0001, Zhuzhong Qian |
Comput. Networks | 5 |
| 2025 | Provisioning high precision edge inference with runtime model reconfiguration
Hesheng Sun, Zhuzhong Qian, Andong Zhu 0001, Sheng Zhang 0001, Sanglu Lu, Gangyi Luo |
Comput. Networks | 3 |
| 2025 | Adaptive scheduling of online inference pipelines at the edge: A post-hoc request-oriented approach
Hesheng Sun, Zhuzhong Qian, Andong Zhu 0001, Sheng Zhang 0001, Sanglu Lu, Lingkun Meng |
J. Syst. Archit. | 3 |
| 2025 | Mystique: User-Level Adaptation for Real-Time Video Analytics in Edge Networks via Meta-RLabstractDeep neural network (DNN)-based real-time video analytics service, as a core module for numerous crucial applications such as augmented reality (AR), has garnered increasing research attention, where mobile edge computing (MEC) is often leveraged to mitigate its real-time processing burden on resource-constrained user devices. For Quality of Experience (QoE) optimization, latest works employ reinforcement learning (RL)-based methods to adaptively adjust configurations (e.g., resolution and frame rate), yet still presenting significant challenges. Firstly, we observe a substantial diversity in QoE patterns among users. Given that existing methods integrate a fixed QoE pattern in parameter training, it is intuitive to customize a policy network for each user. However, this necessitates significant training investment, failing to support on-the-fly deployment for new users. Secondly, given the dual dynamics from both the network and video content in edge video analytics system, existing methods often fall into the dilemma of fitting newly emerged and diverse system states with offline-trained fixed parameters. While it is promising to employ online learning algorithms, most of them struggle to catch up with the high dynamics. We hence proposeMystique. In real-time edge video analytics domain, it is the first meta-RL-based user-level configuration adaptation framework. Mystique establishes an initial model in offline meta training with model-agnostic meta-learning (MAML), enabling swift online adaptation to new users and system states through limited gradient updates from initial parameters. Comprehensive experiments illustrate that Mystique can improve QoE by 42% on average compared to prior works. Xiaohang Shi 0001, Sheng Zhang 0001, Meizhao Liu, Lingkun Meng, Liu Wei, Yingcheng Gu, Kai Liu 0043, Andong Zhu 0001, Ning Chen 0010, Zhuzhong Qian |
IEEE Trans. Mob. Comput. | 11 |
| 2025 | End-to-End Coordinated Spatio-Temporal Redundancy Elimination for Fast Video AnalyticsabstractEdge video analytics typically rely on conventional encoding standards to transmit device visual data for server-side inference. Unfortunately, general-purpose compression solutions retain unnecessary visual data that does not contribute to accuracy, resulting in significant latency throughout Video Analytics Pipeline (VAP). While previous approaches have made partial progress, they cannot systematically eliminate VAP redundancy due to uncoordinated subsystem-level optimization. Achieving complete redundancy elimination presents a major challenge, as a lack of spatio-temporal coordination risks offsetting latency gains with computational overhead (associated with redundancy elimination).Crucioovercomes these limitations with an end-to-edge framework that integrates temporally adaptive frame filtering and coordinated video compression. It leverages redesigned asymmetric autoencoders to synchronize inter-frame temporal compression with intra-frame spatial feature extraction. Additionally,Crucioemploys a one-pass decoding mechanism for encoded critical frames and dynamically adjusts batching scales to minimize latency. Empirical results demonstrateCrucio's superiority, outperforming existing solutions (e.g., DDS, Reducto, and STAC) by over a 31% reduction in end-to-end latency at 0.9 accuracy thresholds. Andong Zhu 0001, Sheng Zhang 0001, Lingkun Meng, Xiaohang Shi 0001, Hesheng Sun, Sanglu Lu, Jie Wu 0001, Yu Liang 0001 |
IEEE Trans. Mob. Comput. | 1 |
| 2025 | Machine-Centric High-Accuracy Multi-Video Analytics With Adaptive Neural CodecsabstractIncreased videos captured by widely deployed cameras are being analyzed by computer vision-based Deep Neural Networks (DNNs) on servers rather than being streamed for humans. Unfortunately, the conventional codecs (e.g., H.26x and MPEG-x) originally designed for video streaming lack content-aware feature extraction and hinder machine-centric video analytics, making it difficult to achieve the required high accuracy with tolerable delay. Neural codecs (e.g., autoencoder) now hold impressive compression performance and have been widely advocated in video streaming. While autoencoder shows transformative potential, the application in video analytics is hampered by low accuracy in detecting small objects of high-resolution videos and the serious challenges posed by multi-video streaming. To this end, we propose AdaStreamer with adaptive neural codecs to enable real machine-centric high-accuracy multi-video analytics. We also investigate how to achieve optimal accuracy under delay constraints via careful scheduling in Compression Ratios (CRs, the ratio of the compressed size to the original data size) and bandwidth allocation, and further propose a Markov-based Adaptive Compression and Bandwidth Allocation algorithm (MACBA). We have practically developed a prototype of AdaStreamer, based on which extensive experiments verify its accuracy improvement (up to 15%) compared to state-of-the-art coding and streaming solutions. Andong Zhu 0001, Ji Qi 0005, Sheng Zhang 0001, Gangyi Luo, Xiaohang Shi 0001, Zhuzhong Qian, Sanglu Lu |
IEEE Trans. Netw. | 1 |
| 2024 | AdaStreamer: Machine-Centric High-Accuracy Multi-Video Analytics with Adaptive Neural CodecsabstractIncreased videos captured by widely deployed cameras are being analyzed by computer vision-based Deep Neural Networks (DNNs) on servers rather than being streamed for humans. Unfortunately, the conventional codecs (e.g., H.26x and MPEG-x) originally designed for video streaming lack content-aware feature extraction and hinder machine-centric video analytics, making it difficult to achieve the required high accuracy with tolerable delay. Neural codecs (e.g., autoencoder) now hold impressive compression performance and have been widely advocated in video streaming. While autoencoder shows transformative potential, the application in video analytics is hampered by low accuracy in detecting small objects of highresolution videos and the serious challenges posed by multivideo streaming. To this end, we propose AdaStreamer with adaptive neural codecs to enable real machine-centric highaccuracy multi-video analytics. We also investigate how to achieve optimal accuracy under delay constraints via careful scheduling in Compression Ratios (CRs, the ratio of the compressed size to the original data size) and bandwidth allocation, and further propose a Markov-based Adaptive Compression and Bandwidth Allocation algorithm (MACBA). We have practically developed a prototype of AdaStreamer, based on which extensive experiments verify its accuracy improvement (up to 15%) compared to stateof-the-art coding and streaming solutions. Andong Zhu 0001, Sheng Zhang 0001, Xiaohang Shi 0001, Zhuzhong Qian, Sanglu Lu |
INFOCOM | 1 |
| 2024 | Crucio: End-to-End Coordinated Spatio-Temporal Redundancy Elimination for Fast Video AnalyticsabstractVideo Analytics Pipeline (VAP) usually relies on traditional codecs to stream video content from clients to servers. However, such analytics-agnostic codecs preserve considerable pixels not relevant to achieving high analytics accuracy, incurring a large end-to-end delay. Despite the significant efforts of pioneers, they fall short as they resisted complete redundancy elimination. Achieving such a goal is extremely challenging, and naive design without coordination can result in the benefits of redundancy elimination being counterbalanced by intolerable delays introduced. We present CRUCIO, an end-to-end coordinated spatio-temporal redundancy elimination system for edge video analytics. CRUCIO leverages reshaped asymmetric autoencoders for end-to-end frame filtering (temporally) and coordinated intra-frame (spatially), inter-frame (temporally) compression. Furthermore, CRUCIO can decode the compressed key frames all in one go and support adaptive VAP batch size for delay optimization. Extensive evaluations reveal significant end-to-end delay reductions (at least 31% under an accuracy target of 0.9) in CRUCIO compared to the state-of-the-art VAP redundancy elimination methods (e.g., DDS, Reducto, STAC, etc). Andong Zhu 0001, Sheng Zhang 0001, Xiaohang Shi 0001, Hesheng Sun, Sanglu Lu |
INFOCOM | 1 |
| 2024 | GeoScale: Microservice Autoscaling With Cost Budget in Geo-Distributed Edge CloudsabstractDeploying microservice instances in geo-distributed edge clouds which are located at the network edge and in proximity to end-users can provide on-site processing, thereby improving the quality of service (QoS). To accommodate the time-varying request arrival rate of each edge cloud, the deployment scheme of microservice instances is dynamically adapted, which is called microservice autoscaling. However, existing studies on microservice autoscaling at the edge either only optimize the QoS without considering the cost of deploying microservice instances or simply focus on the cost per individual timeslot, and thus always severely violate the long-term budget constraint. To solve this problem, in this article, we propose GeoScale, a novel method that aims to optimize the average request response time under the long-term cost budget constraint. GeoScale first utilizes the Lyapunov optimization framework to decompose the long-term optimization problem into a series of per-timeslot sub-problems and then applies a signomial geometric programming (SGP)-based algorithm to obtain a near-optimal solution to each NP-hard sub-problem. Through extensive trace-driven experiments, we validate the superiority of GeoScale. The experimental results show that compared with existing strategies and designed baselines, GeoScale can improve QoS by reducing the average request response time up to 87.8% while significantly mitigating the violation of the long-term cost budget constraint. Sheng Zhang 0001, Meizhao Liu, Yingcheng Gu, Liu Wei, Kai Liu 0043, Xiaohang Shi 0001, Andong Zhu 0001 |
IEEE Trans. Parallel Distributed Syst. | 10 |
| 2023 | OSCA: Online User-managed Server Selection and Configuration Adaptation for Interactive MARabstractInteractive mobile augmented reality (MAR) applications such as Connected Lens are becoming popular, which often rely on deep neural network (NN)-based video analytics techniques to understand the real world. However, performing computation-intensive NN inference on resource-constrained mobile devices is impractical. It is thus proposed to offload the workloads to edge servers with the help of mobile edge computing (MEC). Existing works often focus on system-wide offloading solutions, optimizing the personalized user experience for interactive applications in dynamic environments is yet rarely studied, where multiple challenges remain to be solved. First, the user has to decide the configuration for video analytics, where the inherent accuracy-cost trade-off exists. Second, it is intractable to decide the target server for offloading, since each server supports limited configurations, and a user needs to balance the experience of analytics service and the quality of interaction with others at the same time. Third, the fluctuating network information is often undisclosed to the users, and the candidate servers also vary over time. Therefore, in this paper, we propose an online user-managed server selection and configuration adaptation scheme (OSCA). Via Lyapunov optimization, we aim to maximize the long-term service experience, under the interactive quality constraint with other users. Besides, volatile multi-armed bandit (MAB) is utilized to handle the network fluctuation and the variance of the candidate servers. We conduct rigorous theoretical analysis, and the deviations of both the service experience and the interactive quality are bounded. Through extensive trace-driven experiments, we demonstrate the superior performance of OSCA. Xiaohang Shi 0001, Sheng Zhang 0001, Yu Chen 0038, Andong Zhu 0001, Sanglu Lu |
IWQoS | 5 |
| 2023 | On Efficient Packet Batching and Resource Allocation for GPU based NFV AccelerationabstractNetwork Function Virtualization (NFV) has already become an essential technology for improving the scalability and flexibility of modern computer networks. The performance gap has become the main issue that impedes the development of NFV. GPUs, with massive parallel processors, are advocated to accelerate the Virtualized Network Functions (VNFs). However, the special architecture and workflow of GPUs introduce new challenges, especially on the batched processing, and resource allocation. In this paper, we propose GPU-based NFV Acceleration framework (GNFA) with an efficient packet batching and resource allocation solution. Considering the increased latency caused by the accumulation of the GPU kernel invoking overhead, we first invent a latency reduction mechanism called SM Performance Compensation (SPC). A Partition and Adjustment based Batching and Resource Allocation (PABARA) algorithm that jointly considers batch size tuning and GPU thread allocation is also proposed. We have practically implemented GNFA and extensively evaluated its performance on some well-known VNFs. The experiment results show that GNFA can effectively promote the GPU resource utilization and improve the NFV performance in terms of per-packet latency. Deze Zeng, Andong Zhu 0001, Lin Gu 0002, Quan Chen 0002, Minyi Guo |
IWQoS | 2 |
| 2023 | Enabling Efficient Spatio-Temporal GPU Sharing for Network Function VirtualizationabstractBy leveraging standard IT virtualization technology and Commercial-Off-The-Shelf (COTS) servers, Network Function Virtualization (NFV) decouples network functions from proprietary hardware devices for flexible service provisioning. But the potential of NFV is significantly limited by its performance inefficiency. With the unparalleled advantages of multi-core parallelism and high memory bandwidth, Graphics Processing Units (GPUs) are regarded as a promising way to accelerate Virtualized Network Functions (VNF). However, the special architecture of GPU brings new challenges to task scheduling and resource allocation. To this end, we propose aGPUorientedspatio-temporal sharing framework for NFV calledGost, aiming for GPU based VNF performance promotion. The execution order and GPU resource allocation (i.e., the number of threads) are considered in task scheduling to minimize the end-to-end latency for VNF flows. First, we formulate the task scheduling problem into a nonlinear programming form, and then transform it into an equivalent Integer Linear Programming (ILP) form. The problem is proved as NP-hard. We customize the classical list scheduling algorithm and propose a List Scheduling based Spatio-Temporal GPU sharing strategy (LSSTG), whose achievable worst-case performance is also formally analyzed. We practically implementGostprototype, based on which extensive experiments verify the high performance efficiency of LSSTG compared to state-of-the-art in terms of latency and throughput. Deze Zeng, Andong Zhu 0001, Lin Gu 0002, Peng Li 0017, Quan Chen 0002, Minyi Guo |
IEEE Trans. Computers | 2 |
| 2021 | Gost: Enabling Efficient Spatio-Temporal GPU Sharing for Network Function VirtualizationabstractNetwork Function Virtualization (NFV) enables network functions to run on general-purpose servers, thus alleviates the reliance on dedicated hardware and significantly improves the scalability and flexibility in networking service provisioning. Meanwhile, it is recognized that Virtualized Network Functions (VNFs) suffer from serious performance problem. Graphics Processing Unit (GPU), with massive processing cores, has been advocated as a potential accelerator for improving the performance efficiency of VNFs. However, the special architecture of GPU makes existing CPU-oriented task scheduling strategies fail to be applied, limiting the acceleration potential of GPUs. To this end, we propose a GPU-oriented spatio-temporal sharing framework as Gost to improve the performance of GPU-accelerated VNFs. We also study how to minimize the end-to-end latency of VNF flows via careful scheduling on the execution order and the GPU resource allocation (i.e., the number of threads). We first formally describe the problem as a non-linear integer programming problem, which is then equivalently transformed into an integer linear programming (ILP) form. Considering the high computation complexity of solving ILP, we further propose a customized list scheduling based spatio-temporal GPU sharing strategy (LSSTG). We have practically implemented a prototype of Gost, based on which we also verify the high efficiency of LSSTG by extensive experiments. Andong Zhu 0001, Deze Zeng, Lin Gu 0002, Peng Li 0017, Quan Chen 0002 |
IWQoS | 1 |
| 2020 | Task Offloading in Trusted Execution Environment empowered Edge ComputingabstractTo tackle the computation resource poorness on the end devices, task offloading is developed to reduce the task completion time and improve the Quality-of-Service (QoS). Edge computing facilitates such offloading by provisioning resources at the proximity of the end devices. Nowadays, many tasks on end devices have an urgent demand for the security of execution environment. To address this problem, we introduce trusted execution environment (TEE) to empower edge computing for secure task offloading. To explore TEE, the offloading process should be redesigned with the introduction of data encryption and decryption. This makes traditional offloading optimization policy fail to be applied directly. To address this issue, we are motivated to take the data encryption and decryption into the offloading scheduling algorithm. In particular, we propose a Customized List Scheduling based Offloading (CLSO) algorithm, aiming at minimizing the total completion time with the consideration of energy budget limitations on the end devices. The experiment results show that our approximation algorithm can effectively reduce the total completion time and significantly outperforms existing state-of-the-art offloading strategy. Yuepeng Li, Deze Zeng, Lin Gu 0002, Andong Zhu 0001, Quan Chen 0002 |
ICPADS | 4 |