Xiaohang Shi 0001

dblp:343/2108-1 · DBLP profile ↗
← Back
12ranked-venue papers
3as first author
12since 2021 · last 2026
0009-0002-9796-238XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 9 · 3 first-author · 9 since 2021Systems, architecture and hardware · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Edge-cloud co-optimized 3D video analytics with synergistic neural codecs
Hebin Sun, Xiaohang Shi 0001, Xiaokun Wang 0002, Sheng Zhang 0001, Lingkun Meng, Andong Zhu 0001, Zhuzhong Qian
Comput. Networks2
2025 VidIQ: Inference-Aware Neural Codecs for Quality-Enhanced, Real-Time Video Analytics
abstract
Video analytics pipelines migrating to edge deployments are facing performance bottlenecks under limited bandwidth. Non-uniform intra-frame encoding emerges to further compress pixels without affecting the output of the server deep neural network (DNN), while it is inefficient in high-resolution video streaming at low bandwidth. The detail enhancement capability of neural super-resolution (SR) permits resolution downsampling and aggressive compression on edge devices for low-latency transmission. To exploit its accuracy potential, DNN-oriented non-uniform encoding is expected to be additionally aware of SR models. However, traditional codecs struggle to cope with both quality optimization for SR and global semantic features for DNN. We advocate neural codecs for coordinated encoding and enhancement, enabling analytic-oriented video streaming with optimal accuracy-delay tradeoffs. Our system, VidIQ, achieves quality-enhanced real-time video analytics by 1) improving the network architecture of neural codecs (at two granularity) to integrate SR models into a DNN-oriented analytics pipeline, and 2) adapting the multi-scale encoder and SR-decoder to scene dynamics (i.e., content and bandwidth variations) with the help of the monolithic controller to hold a performance advantage. Extensive evaluations showcase that VidIQ reduces end-to-end delay by 35.8% and improves analytics accuracy by 21.2% compared to the recent video compression, enhancement, and streaming baselines.
Andong Zhu 0001, Sheng Zhang 0001, Xiaohang Shi 0001, Hesheng Sun, Yu Liang 0001, Zhuzhong Qian, Xiaokun Wang 0002
ACM Multimedia3
2025 Mystique: User-Level Adaptation for Real-Time Video Analytics in Edge Networks via Meta-RL
abstract
Deep neural network (DNN)-based real-time video analytics service, as a core module for numerous crucial applications such as augmented reality (AR), has garnered increasing research attention, where mobile edge computing (MEC) is often leveraged to mitigate its real-time processing burden on resource-constrained user devices. For Quality of Experience (QoE) optimization, latest works employ reinforcement learning (RL)-based methods to adaptively adjust configurations (e.g., resolution and frame rate), yet still presenting significant challenges. Firstly, we observe a substantial diversity in QoE patterns among users. Given that existing methods integrate a fixed QoE pattern in parameter training, it is intuitive to customize a policy network for each user. However, this necessitates significant training investment, failing to support on-the-fly deployment for new users. Secondly, given the dual dynamics from both the network and video content in edge video analytics system, existing methods often fall into the dilemma of fitting newly emerged and diverse system states with offline-trained fixed parameters. While it is promising to employ online learning algorithms, most of them struggle to catch up with the high dynamics. We hence proposeMystique. In real-time edge video analytics domain, it is the first meta-RL-based user-level configuration adaptation framework. Mystique establishes an initial model in offline meta training with model-agnostic meta-learning (MAML), enabling swift online adaptation to new users and system states through limited gradient updates from initial parameters. Comprehensive experiments illustrate that Mystique can improve QoE by 42% on average compared to prior works.
Xiaohang Shi 0001, Sheng Zhang 0001, Meizhao Liu, Lingkun Meng, Liu Wei, Yingcheng Gu, Kai Liu 0043, Andong Zhu 0001, Ning Chen 0010, Zhuzhong Qian
IEEE Trans. Mob. Comput.1
2025 End-to-End Coordinated Spatio-Temporal Redundancy Elimination for Fast Video Analytics
abstract
Edge video analytics typically rely on conventional encoding standards to transmit device visual data for server-side inference. Unfortunately, general-purpose compression solutions retain unnecessary visual data that does not contribute to accuracy, resulting in significant latency throughout Video Analytics Pipeline (VAP). While previous approaches have made partial progress, they cannot systematically eliminate VAP redundancy due to uncoordinated subsystem-level optimization. Achieving complete redundancy elimination presents a major challenge, as a lack of spatio-temporal coordination risks offsetting latency gains with computational overhead (associated with redundancy elimination).Crucioovercomes these limitations with an end-to-edge framework that integrates temporally adaptive frame filtering and coordinated video compression. It leverages redesigned asymmetric autoencoders to synchronize inter-frame temporal compression with intra-frame spatial feature extraction. Additionally,Crucioemploys a one-pass decoding mechanism for encoded critical frames and dynamically adjusts batching scales to minimize latency. Empirical results demonstrateCrucio's superiority, outperforming existing solutions (e.g., DDS, Reducto, and STAC) by over a 31% reduction in end-to-end latency at 0.9 accuracy thresholds.
Andong Zhu 0001, Sheng Zhang 0001, Lingkun Meng, Xiaohang Shi 0001, Hesheng Sun, Sanglu Lu, Jie Wu 0001, Yu Liang 0001
IEEE Trans. Mob. Comput.4
2025 Machine-Centric High-Accuracy Multi-Video Analytics With Adaptive Neural Codecs
abstract
Increased videos captured by widely deployed cameras are being analyzed by computer vision-based Deep Neural Networks (DNNs) on servers rather than being streamed for humans. Unfortunately, the conventional codecs (e.g., H.26x and MPEG-x) originally designed for video streaming lack content-aware feature extraction and hinder machine-centric video analytics, making it difficult to achieve the required high accuracy with tolerable delay. Neural codecs (e.g., autoencoder) now hold impressive compression performance and have been widely advocated in video streaming. While autoencoder shows transformative potential, the application in video analytics is hampered by low accuracy in detecting small objects of high-resolution videos and the serious challenges posed by multi-video streaming. To this end, we propose AdaStreamer with adaptive neural codecs to enable real machine-centric high-accuracy multi-video analytics. We also investigate how to achieve optimal accuracy under delay constraints via careful scheduling in Compression Ratios (CRs, the ratio of the compressed size to the original data size) and bandwidth allocation, and further propose a Markov-based Adaptive Compression and Bandwidth Allocation algorithm (MACBA). We have practically developed a prototype of AdaStreamer, based on which extensive experiments verify its accuracy improvement (up to 15%) compared to state-of-the-art coding and streaming solutions.
Andong Zhu 0001, Ji Qi 0005, Sheng Zhang 0001, Gangyi Luo, Xiaohang Shi 0001, Zhuzhong Qian, Sanglu Lu
IEEE Trans. Netw.6
2024 AdaStreamer: Machine-Centric High-Accuracy Multi-Video Analytics with Adaptive Neural Codecs
abstract
Increased videos captured by widely deployed cameras are being analyzed by computer vision-based Deep Neural Networks (DNNs) on servers rather than being streamed for humans. Unfortunately, the conventional codecs (e.g., H.26x and MPEG-x) originally designed for video streaming lack content-aware feature extraction and hinder machine-centric video analytics, making it difficult to achieve the required high accuracy with tolerable delay. Neural codecs (e.g., autoencoder) now hold impressive compression performance and have been widely advocated in video streaming. While autoencoder shows transformative potential, the application in video analytics is hampered by low accuracy in detecting small objects of highresolution videos and the serious challenges posed by multivideo streaming. To this end, we propose AdaStreamer with adaptive neural codecs to enable real machine-centric highaccuracy multi-video analytics. We also investigate how to achieve optimal accuracy under delay constraints via careful scheduling in Compression Ratios (CRs, the ratio of the compressed size to the original data size) and bandwidth allocation, and further propose a Markov-based Adaptive Compression and Bandwidth Allocation algorithm (MACBA). We have practically developed a prototype of AdaStreamer, based on which extensive experiments verify its accuracy improvement (up to 15%) compared to stateof-the-art coding and streaming solutions.
Andong Zhu 0001, Sheng Zhang 0001, Xiaohang Shi 0001, Zhuzhong Qian, Sanglu Lu
INFOCOM4
2024 Crucio: End-to-End Coordinated Spatio-Temporal Redundancy Elimination for Fast Video Analytics
abstract
Video Analytics Pipeline (VAP) usually relies on traditional codecs to stream video content from clients to servers. However, such analytics-agnostic codecs preserve considerable pixels not relevant to achieving high analytics accuracy, incurring a large end-to-end delay. Despite the significant efforts of pioneers, they fall short as they resisted complete redundancy elimination. Achieving such a goal is extremely challenging, and naive design without coordination can result in the benefits of redundancy elimination being counterbalanced by intolerable delays introduced. We present CRUCIO, an end-to-end coordinated spatio-temporal redundancy elimination system for edge video analytics. CRUCIO leverages reshaped asymmetric autoencoders for end-to-end frame filtering (temporally) and coordinated intra-frame (spatially), inter-frame (temporally) compression. Furthermore, CRUCIO can decode the compressed key frames all in one go and support adaptive VAP batch size for delay optimization. Extensive evaluations reveal significant end-to-end delay reductions (at least 31% under an accuracy target of 0.9) in CRUCIO compared to the state-of-the-art VAP redundancy elimination methods (e.g., DDS, Reducto, STAC, etc).
Andong Zhu 0001, Sheng Zhang 0001, Xiaohang Shi 0001, Hesheng Sun, Sanglu Lu
INFOCOM3
2024 AdaPyramid: Adaptive Pyramid for Accelerating High-Resolution Object Detection on Edge Devices
abstract
Deep convolutional neural network (NN)-based object detectors are not appropriate for straightforward inference on high-resolution videos at edge devices, as maintaining high accuracy often brings about prohibitively long latency. Although existing solutions have attempted to reduce on-device inference latency by selecting a cheaper configuration (e.g., choosing a more lightweight NN or scaling a frame to a smaller size before inference) or eliminating a background containing no object, they often ignore various high-resolution features and fail to optimize for those videos. We thus present AdaPyramid, a framework to reduce as much on-device inference latency as possible, especially for high-resolution videos, while achieving the accuracy demand approximately. We observe that the cheapest configuration to achieve the accuracy demand varies significantly across both different frames and different regions in a frame. The underlying reason is that object features (e.g., the location, size and category of objects) are more uneven in high-resolution videos, both temporally and spatially. Moreover, we observe that the object size presents a prominent hierarchical distribution in high-resolution frames. AdaPyramid thus partitions each frame hierarchically just like a pyramid and chooses a content-aware configuration for each region, which is adapted online based on the feedback. We evaluate the performance of AdaPyramid on a public dataset and our collected real-world videos. The obtained results show that under comparable accuracy to the state-of-the-art solutions, AdaPyramid can decrease inference latency by 40% on average, with up to 2.5× speed-up.
Xiaohang Shi 0001, Sheng Zhang 0001, Jie Wu 0001, Ning Chen 0010, Yu Liang 0001, Sanglu Lu
IEEE Trans. Mob. Comput.1
2024 GeoScale: Microservice Autoscaling With Cost Budget in Geo-Distributed Edge Clouds
abstract
Deploying microservice instances in geo-distributed edge clouds which are located at the network edge and in proximity to end-users can provide on-site processing, thereby improving the quality of service (QoS). To accommodate the time-varying request arrival rate of each edge cloud, the deployment scheme of microservice instances is dynamically adapted, which is called microservice autoscaling. However, existing studies on microservice autoscaling at the edge either only optimize the QoS without considering the cost of deploying microservice instances or simply focus on the cost per individual timeslot, and thus always severely violate the long-term budget constraint. To solve this problem, in this article, we propose GeoScale, a novel method that aims to optimize the average request response time under the long-term cost budget constraint. GeoScale first utilizes the Lyapunov optimization framework to decompose the long-term optimization problem into a series of per-timeslot sub-problems and then applies a signomial geometric programming (SGP)-based algorithm to obtain a near-optimal solution to each NP-hard sub-problem. Through extensive trace-driven experiments, we validate the superiority of GeoScale. The experimental results show that compared with existing strategies and designed baselines, GeoScale can improve QoS by reducing the average request response time up to 87.8% while significantly mitigating the violation of the long-term cost budget constraint.
Sheng Zhang 0001, Meizhao Liu, Yingcheng Gu, Liu Wei, Kai Liu 0043, Xiaohang Shi 0001, Andong Zhu 0001
IEEE Trans. Parallel Distributed Syst.9
2023 Adaptive Provisioning In-band Network Telemetry at Computing Power Network [invited]
abstract
In-band Network Telemetry (INT) is proposed to detect networks via injecting specific probes to collect the hop-by-hop metadata within programmable switches. But there exist multiple challenges to conducting INT at Computing Power Network, such as control decisions of different INT frequencies, and the unforeseeable INT query workloads. In this study, we formulate an online non-linear time-varying integer programming problem that aims to maximize the overall quality of service through both frequency selection and INT query workload distribution. To achieve this, we propose an online learning, INTService, which utilizes a primal-dual mechanism to make fractional decisions. At last, extensive evaluations show that our proposed INTService exhibits up-lift performance 40% on average over other state-of-the-art algorithms.
Mingtao Ji, Chenwei Su, Zhuzhong Qian, Sheng Zhang 0001, Yu Chen 0038, Tuo Cao, Xiaohang Shi 0001, Luis Vasquez
IWQoS8
2023 OSCA: Online User-managed Server Selection and Configuration Adaptation for Interactive MAR
abstract
Interactive mobile augmented reality (MAR) applications such as Connected Lens are becoming popular, which often rely on deep neural network (NN)-based video analytics techniques to understand the real world. However, performing computation-intensive NN inference on resource-constrained mobile devices is impractical. It is thus proposed to offload the workloads to edge servers with the help of mobile edge computing (MEC). Existing works often focus on system-wide offloading solutions, optimizing the personalized user experience for interactive applications in dynamic environments is yet rarely studied, where multiple challenges remain to be solved. First, the user has to decide the configuration for video analytics, where the inherent accuracy-cost trade-off exists. Second, it is intractable to decide the target server for offloading, since each server supports limited configurations, and a user needs to balance the experience of analytics service and the quality of interaction with others at the same time. Third, the fluctuating network information is often undisclosed to the users, and the candidate servers also vary over time. Therefore, in this paper, we propose an online user-managed server selection and configuration adaptation scheme (OSCA). Via Lyapunov optimization, we aim to maximize the long-term service experience, under the interactive quality constraint with other users. Besides, volatile multi-armed bandit (MAB) is utilized to handle the network fluctuation and the variance of the candidate servers. We conduct rigorous theoretical analysis, and the deviations of both the service experience and the interactive quality are bounded. Through extensive trace-driven experiments, we demonstrate the superior performance of OSCA.
Xiaohang Shi 0001, Sheng Zhang 0001, Yu Chen 0038, Andong Zhu 0001, Sanglu Lu
IWQoS1
2023 ProScale: Proactive Autoscaling for Microservice With Time-Varying Workload at the Edge
abstract
Deploying microservice instances on the edge device close to end users can provide on-site processing thus reducing request response time. Each microservice has multiple instances that can process requests in parallel. To achieve high processing efficiency, the number of these instances is scaled according to the workload, which is also known as autoscaling. Previous studies of microservice autoscaling in the edge computing environment lack in-depth consideration of time-varying workload, they assume that the workload of each microservice always depends on that of its upstream. However, through an analysis of Alibaba's microservice trace with hundreds of millions of records, we find that the assumption is impractical thus hurting autoscaling effectiveness. To solve this problem, we propose ProScale, a prediction-driven proactive autoscaling framework for microservices at the edge. ProScale proactively forecasts the workload for each individual microservice per timeslot. Then it utilizes an efficient online algorithm to leverage the predicting results to determine the instance number for each microservice jointly with making placement decisions. For each microservice instance deployed on the edge device, ProScale handles burst requests using a designed offloading strategy. In addition, ProScale can also balance the load for multiple instances of each microservice. Extensive trace-driven experiments show that ProScale has great scalability. It can reduce average response time by 96.7% and resource usage by 96.5% compared with existing strategies and designed baselines.
Sheng Zhang 0001, Chenghong Tu, Xiaohang Shi 0001, Zhaoheng Yin, Sanglu Lu, Yu Liang 0001, Qing Gu 0001
IEEE Trans. Parallel Distributed Syst.4