Hesheng Sun

dblp:284/7272 · DBLP profile ↗
← Back
11ranked-venue papers
4as first author
11since 2021 · last 2026
0009-0005-5246-6018ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 7 · 2 first-author · 7 since 2021Systems, architecture and hardware · 3 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Octopus: Accuracy-aware resource scheduling for multi-video streaming inference at the edge
Zhuzhong Qian, Andong Zhu 0001, Hesheng Sun, Lingkun Meng
Comput. Networks4
2025 Bridging the Prediction-Decision Gap: Enhancing Model Deployment and Online Service Request Forecasting in Edge Inference Systems
abstract
In edge inference systems, efficient model deployment and accurate forecasting of online service requests are critical for maintaining optimal performance and resource utilization. This study introduces a novel approach that integrates predictive enhancement techniques to improve model deployment strategies and online service request forecasting. By addressing the existing gap between prediction and decision-making processes, our method enables more responsive and adaptive edge computing environments. Experimental results demonstrate significant improvements in both deployment efficiency and forecasting accuracy, highlighting the potential of our approach to advance the state-of-the-art in edge inference system management.
Hesheng Sun, Zhuzhong Qian, Andong Zhu 0001
IWQoS1
2025 VidIQ: Inference-Aware Neural Codecs for Quality-Enhanced, Real-Time Video Analytics
abstract
Video analytics pipelines migrating to edge deployments are facing performance bottlenecks under limited bandwidth. Non-uniform intra-frame encoding emerges to further compress pixels without affecting the output of the server deep neural network (DNN), while it is inefficient in high-resolution video streaming at low bandwidth. The detail enhancement capability of neural super-resolution (SR) permits resolution downsampling and aggressive compression on edge devices for low-latency transmission. To exploit its accuracy potential, DNN-oriented non-uniform encoding is expected to be additionally aware of SR models. However, traditional codecs struggle to cope with both quality optimization for SR and global semantic features for DNN. We advocate neural codecs for coordinated encoding and enhancement, enabling analytic-oriented video streaming with optimal accuracy-delay tradeoffs. Our system, VidIQ, achieves quality-enhanced real-time video analytics by 1) improving the network architecture of neural codecs (at two granularity) to integrate SR models into a DNN-oriented analytics pipeline, and 2) adapting the multi-scale encoder and SR-decoder to scene dynamics (i.e., content and bandwidth variations) with the help of the monolithic controller to hold a performance advantage. Extensive evaluations showcase that VidIQ reduces end-to-end delay by 35.8% and improves analytics accuracy by 21.2% compared to the recent video compression, enhancement, and streaming baselines.
Andong Zhu 0001, Sheng Zhang 0001, Xiaohang Shi 0001, Hesheng Sun, Yu Liang 0001, Zhuzhong Qian, Xiaokun Wang 0002
ACM Multimedia4
2025 Provisioning high precision edge inference with runtime model reconfiguration
Hesheng Sun, Zhuzhong Qian, Andong Zhu 0001, Sheng Zhang 0001, Sanglu Lu, Gangyi Luo
Comput. Networks1
2025 Towards strong continuous consistency in edge-assisted VR-SGs: Delay-differences sensitive online task redistribution
Yunqi Sun, Hesheng Sun, Tuo Cao, Mingtao Ji, Zhuzhong Qian, Lingkun Meng
Comput. Networks2
2025 Adaptive scheduling of online inference pipelines at the edge: A post-hoc request-oriented approach
Hesheng Sun, Zhuzhong Qian, Andong Zhu 0001, Sheng Zhang 0001, Sanglu Lu, Lingkun Meng
J. Syst. Archit.1
2025 End-to-End Coordinated Spatio-Temporal Redundancy Elimination for Fast Video Analytics
abstract
Edge video analytics typically rely on conventional encoding standards to transmit device visual data for server-side inference. Unfortunately, general-purpose compression solutions retain unnecessary visual data that does not contribute to accuracy, resulting in significant latency throughout Video Analytics Pipeline (VAP). While previous approaches have made partial progress, they cannot systematically eliminate VAP redundancy due to uncoordinated subsystem-level optimization. Achieving complete redundancy elimination presents a major challenge, as a lack of spatio-temporal coordination risks offsetting latency gains with computational overhead (associated with redundancy elimination).Crucioovercomes these limitations with an end-to-edge framework that integrates temporally adaptive frame filtering and coordinated video compression. It leverages redesigned asymmetric autoencoders to synchronize inter-frame temporal compression with intra-frame spatial feature extraction. Additionally,Crucioemploys a one-pass decoding mechanism for encoded critical frames and dynamically adjusts batching scales to minimize latency. Empirical results demonstrateCrucio's superiority, outperforming existing solutions (e.g., DDS, Reducto, and STAC) by over a 31% reduction in end-to-end latency at 0.9 accuracy thresholds.
Andong Zhu 0001, Sheng Zhang 0001, Lingkun Meng, Xiaohang Shi 0001, Hesheng Sun, Sanglu Lu, Jie Wu 0001, Yu Liang 0001
IEEE Trans. Mob. Comput.8
2024 Crucio: End-to-End Coordinated Spatio-Temporal Redundancy Elimination for Fast Video Analytics
abstract
Video Analytics Pipeline (VAP) usually relies on traditional codecs to stream video content from clients to servers. However, such analytics-agnostic codecs preserve considerable pixels not relevant to achieving high analytics accuracy, incurring a large end-to-end delay. Despite the significant efforts of pioneers, they fall short as they resisted complete redundancy elimination. Achieving such a goal is extremely challenging, and naive design without coordination can result in the benefits of redundancy elimination being counterbalanced by intolerable delays introduced. We present CRUCIO, an end-to-end coordinated spatio-temporal redundancy elimination system for edge video analytics. CRUCIO leverages reshaped asymmetric autoencoders for end-to-end frame filtering (temporally) and coordinated intra-frame (spatially), inter-frame (temporally) compression. Furthermore, CRUCIO can decode the compressed key frames all in one go and support adaptive VAP batch size for delay optimization. Extensive evaluations reveal significant end-to-end delay reductions (at least 31% under an accuracy target of 0.9) in CRUCIO compared to the state-of-the-art VAP redundancy elimination methods (e.g., DDS, Reducto, STAC, etc).
Andong Zhu 0001, Sheng Zhang 0001, Xiaohang Shi 0001, Hesheng Sun, Sanglu Lu
INFOCOM5
2024 Walking on two legs: Joint service placement and computation configuration for provisioning containerized services at edges
Tuo Cao, Qinhui Wang, Zhuzhong Qian, Yue Zeng 0002, Mingtao Ji, Hesheng Sun
Comput. Networks7
2023 BIRP: Batch-aware Inference Workload Redistribution and Parallel Scheme for Edge Collaboration
abstract
The inference workload redistribution is a technique for evacuating inference requests from hot edges to idle edges in edge collaborative systems, thereby achieving inference workload balancing for inference on different edges. However, with the continuous development of edge accelerators, the resource utilization of edge accelerators in executing inference requests in series is often low, and when executing multiple inference requests in parallel, it faces uncertain execution delays, different response-time Service Level Objectives (SLOs), and the generality of inference workloads in heterogeneous edge collaborative systems. To address these issues, for the first time in the domain of inference workload redistribution, we propose a Batch-aware Inference workload Redistribution and Parallel execution scheme, called BIRP, to reduce the additional latency caused by waiting for a single inference task during serial execution, thereby improving the overall inference accuracy. BIRP uses the Multi-Armed Bandit (MAB) algorithm to adjust hyperparameters of the Throughput Improvement Ratio (TIR) function online for improving the overall inference accuracy. For nonlinear terms in the problem, BIRP uses a piecewise linear approximation to convert it into a Quadratic Programming (QP) problem, ensuring the effectiveness of BIRP in theory. We prototype BIRP on an edge collaborative system composed of three heterogeneous edges. Based on real inference workload trace, we validate the superiority of our algorithm compared to the state-of-the-art model selection-based inference workload redistribution algorithm, with an overall inference loss reduction of at least 32.9% and the failure rate of SLO has been reduced to 19.8% of alternatives.
Hesheng Sun, Zhuzhong Qian, Zengji Li, Ning Chen 0010, Tuo Cao, Suwei Xu
ICPP1
2022 Inference replication at edges via combinatorial multi-armed bandit
Hesheng Sun, Yibo Jin 0001, Yanfang Zhu, Zhuzhong Qian, Sheng Zhang 0001, Sanglu Lu
J. Syst. Archit.2