VLDB 2026 Research / reviewers in the wild / expert
Zhaowu Huang
dblp:257/5755
· DBLP profile ↗
18ranked-venue papers
5as first author
17since 2021 · last 2026
0000-0002-3941-6412ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 8 · 2 first-author · 8 since 2021Systems, architecture and hardware · 7 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Adaptive region encoding for efficient video object detection in edge computing
Lisha Gao, Zhenxuan Xu, Zhaowu Huang, Fang Dong 0001 |
Knowl. Based Syst. | 5 |
| 2026 | Popularity-Aware Layer-Wise Caching and Function Scheduling for Dynamic Workflow at the EdgeabstractServerless Edge Computing (SEC) has emerged as a promising paradigm for delivering low-latency, resource-efficient services for edge-native applications, which are implemented as dependent functions, forming Directed Acyclic Graph (DAG) workflows. Unfortunately, the application's performance is hindered by the notorious issue of cold startup, especially in resource-constrained SEC environments. Layer- wise container caching has been proven to be an effective startup acceleration solution in SEC, due to its fine granularity and flexibility. However, due to the dynamic nature of call graphs and the skewness in function popularity in DAG workflows, as well as the heterogeneity of container layer cold start time and edge computing environments, the performance of existing layer- wise caching mechanisms degrades significantly. To solve this problem, we propose an efficient DAG workflow deployment method in SEC to minimize the application completion time (ACT) in the long term. We model the problem as a joint optimization of container layer- wise caching and function scheduling, which is a Time-coupled Integer Nonlinear Programming (TINLP) problem. To solve it, we first convert it to an Integer Linear Programming (ILP) problem and propose an online algorithm with theoretical performance guarantees. Extensive experiments demonstrate that our method achieves up to$2.92\times$speedup in ACT reduction. Xiaolin Guo, Fang Dong 0001, Dian Shen, Zhaowu Huang |
IEEE Trans. Mob. Comput. | 4 |
| 2025 | FaSei: Fast Serverless Edge Inference with Synergistic Lazy Loading and Layer-wise Caching
Zhaowu Huang, Fang Dong 0001, Xiaolin Guo, Daheng Yin |
INFOCOM | 1 |
| 2025 | ADPTD: Adaptive Data Partition With Unbiased Task Dispatching for Video Analytics at the EdgeabstractRecently, edge-assisted methods have been proposed as a promising technique to deliver fast and accurate on-device video analytics by partitioning frame data and dispatching them to edge servers for parallel execution. However, the data partition (DP) reduces the detection latency but decreases accuracy since objects may cross the boundaries of adjacent blocks. The effect of DP on the accuracy and latency depends on multiple vital parameters (e.g., target size, density, network, and computing resources) in an unknown and time-varying fashion. Moreover, these parameters are determined by the application scenarios and edge environment, which are uncertain and heterogeneous at the edge. Hence, how to partition frames to strike a balance between accuracy and latency is a nontrivial and intractable problem. To this end, we propose an online learning-based device-edge–cloud collaboration framework, ADPTD, to guide DP at the edge. We propose an optimal task dispatching algorithm (OTD) to minimize detection latency. Then, we propose a multiarmed bandit-based algorithm to pick a DP strategy and invoke OTD to dispatch tasks in each time slot. Theoretical analysis reveals that ADPTD achieves sublinear regret. Extensive experimental results show that ADPTD outperforms the state-of-the-art methods, achieving a latency reduction of up to$2.53\times $and improving accuracy by up to 49.4%. Zhaowu Huang, Fang Dong 0001, Haopeng Zhu, Mengyang Liu, Dian Shen, Ruiting Zhou, Xiaolin Guo, Baijun Chen |
IEEE Internet Things J. | 1 |
| 2025 | Resource-Efficient DNN Inference With Early Exiting in Serverless Edge ComputingabstractServerless Edge Computing (SEC) has gained widespread adoption in improving resource utilization due to its triggered event-driven model. However, deploying deep neural network (DNN) inference services directly in SEC leads to resource inefficiencies, which stem from two key factors. First, existing methods adopt model-wise function encapsulation, which requires the entire DNN model to occupy memory throughout its execution lifecycle. This increases both memory footprint and occupancy time. Second, uniform DNN inference for diversity input leads to redundant computations and additional inference time. To this end, we propose REDI, a novel framework that leverages fine-grained block-wise function encapsulation and progressive inference to provide resource-efficient DNN inference while ensuring latency requirements. REDI enables the release of memory from already inferred shallow networks and allows each request to exit early based on input data complexity, eliminating redundant computations. To fully unleash the potential, REDI jointly considers resource heterogeneity, data diversity, and environment dynamics to investigate the block-wise function placement problem. We introduce an uncertainty-aware online learning-driven algorithm with bounded regret. Finally, we conduct extensive trace-driven experiments to evaluate our methods, demonstrating that REDI achieves a significant speedup of up to$6.52\times$in terms of resource usage cost compared to state-of-the-art methods. Xiaolin Guo, Fang Dong 0001, Dian Shen, Zhaowu Huang, Jinghui Zhang 0001 |
IEEE Trans. Mob. Comput. | 4 |
| 2024 | Rendering Super Resolution Video Streaming Efficiently with in-Network ComputingabstractEmerging live video streaming applications, e.g., Ultra High Definition videos and interactive video streaming, have put forward new demands for ultra-high bandwidth and reduced delay to match the desired quality of experience. Since current on-device Super-Resolution (SR) approaches are hindered by the limited end-device capabilities, we are motivated to take advantage of Mobile Edge Computing (MEC) and Computing in the Network technologies, such that SR videos can be processed on a more powerful infrastructure by integrating the resources from end-devices through the edge and all the way to the cloud. However, the integration of SR and MEC is non-trivial due to the challenges introduced by the features of SR tasks, and the heterogeneous nature of MEC resources. In this paper, we endeavor to explore and solve these challenges by presenting AVSA, which renders SR live video streaming efficiently with in-network computing. AVSA can adaptively allocate SR work-loads under heterogeneous resources and yield a cost-effective workload allocation with theoretical performance guarantees. Simulation results show that, compared with the state-of-the-art methods, our method achieves up to 13.06× speedup in terms of end-to-end latency. Xiaolin Guo, Fang Dong 0001, Dian Shen, Zhaowu Huang, Baijun Chen, Daheng Yin |
HPCC | 4 |
| 2024 | SkewCache: Skewed Layer-wise Caching for Function Chains in Serverless Edge ComputingabstractIn serverless edge computing (SEC), traditional monolithic applications are encapsulated in multiple dependent functions, in which event-driven service provisioning introduces the cold start problem. Existing works adopt uniform container caching for each function to mitigate the cold start problem, which remains with relatively low efficiency due to ignoring the skewed invocation frequency and the heterogeneity of cold-start behaviors of functions. In this paper, we propose SkewCache, an efficient layer-wise container caching framework for frequency-skewed function chains in SEC. Based on the container’s layer structure, SkewCache enables fine-grained and balanced container caching on multiple edge servers, which takes into account both the invocation frequency skewness and the cold start latency of each function. The problem is modeled as an integer nonlinear programming (INLP) to minimize the application completion time (ACT). To solve the INLP, we first convert it to an equivalent integer linear programming (ILP) form. Then, we propose an approximation algorithm to solve the ILP with a guaranteed approximation ratio. To evaluate the performance of the proposed algorithm, we conduct intensive simulations and the results show that our algorithms outperform baselines, achieving up to 1.94× speedup in terms of ACT reduction. Xiaolin Guo, Fang Dong 0001, Dian Shen, Zhaowu Huang, Haodong Tian |
HPCC | 4 |
| 2024 | Joint Layer-wise Caching and Request Routing for Serverless Inference Acceleration at the EdgeabstractIn recent years, the serverless paradigm has been introduced into edge computing, enhancing resource utilization. However, it has also introduced the issue of cold start, which hinders low-latency responses. This cold start problem is significantly exacerbated when it comes to Artificial Intelligence (AI) applications, due to the necessity of loading bulky deep learning (DL) model files. Existing methods typically address the cold start issue by caching complete containers, which is inefficient in resource-constrained edge servers. This is because massive DL models may not fit in the limited cache space or lead to fragmentation of the cache memory. To solve the above problem, we seek a layer-wise model caching method that regards the DL model as a chain structure of multiple layers and caches a part of the model layers. However, heterogeneity in cold starts and edge computing’s complex environments pose challenges in selecting model layers and servers. In this paper, we investigate the joint layer-wise caching and request routing problem to minimize application completion time (ACT). We propose an approximation algorithm with a provable performance guarantee. Finally, intensive trace-driven simulations show that our algorithm achieves 3.03× speedup in terms of ACT reduction. Zhaowu Huang, Fang Dong 0001, Xiaolin Guo, Haodong Tian |
HPCC | 1 |
| 2024 | FSVFG: Towards Immersive Full-Scene Volumetric Video Streaming with Adaptive Feature GridabstractGiven the truly immersive viewing experiences, full-scene volumetric videos have received increasing attention from both academia and industry. Their vast data volumes, however, present significant challenges for real-time streaming over today's bandwidth-limited Internet. Considering the vast amount of full-scene volumetric data to be streamed and the limited bandwidth on the Internet, achieving adaptive full-scene volumetric video streaming over the Internet presents a significant challenge. Inspired by the advantages offered by neural fields, especially the feature grid method, we propose FSVFG, a novel full-scene volumetric video streaming system integrated feature grids as the representation of volumetric content. FSVFG employs an incremental training approach for feature grids and stores the features and residuals between adjacent grids as frames. To support adaptive streaming, we delve into the data structure and rendering processes of feature grids and propose bandwidth adaptation mechanisms. The mechanisms involve a coarse ray-marching for the selection of features and residuals to be sent, and achieve variable bitrate streaming by Level-of-Detail (LoD) and residual filtering. Based on these mechanisms, FSVFG achieves adaptive streaming by adaptively balancing the transmission of feature and residual according to the available bandwidth. Our preliminary results demonstrate the effectiveness of FSVFG, demonstrating its ability to improve visual quality and reduce bandwidth requirements of full-scene volumetric video streaming. Daheng Yin, Jianxin Shi 0005, Miao Zhang 0003, Zhaowu Huang, Jiangchuan Liu, Fang Dong 0001 |
ACM Multimedia | 4 |
| 2024 | Joint Optimization of Device Selection and Resource Allocation for Multiple Federations in Federated Edge LearningabstractFederated edge learning (FEEL) is a promising collaborative paradigm, which employs edge devices (EDs) to train machine learning models for a federation. It opens countless opportunities to enable edge intelligence. The increasingly diversified demands for intelligent services are driving the deployment of various federations at the edge. Existing works on FEEL focus on a single federation and ignore inter-federation device competition and intra-device resource allocation, which hinders the applications of FEEL. To address this issue, this article first investigates the bottlenecks of executing multiple federations and builds a joint optimization model as a two-stage Stackelberg game involving device selection and resource allocation. To tackle the problem efficiently, we present a game-theoretical approach namedDeviceSelection andResourceAllocation forMultipleFederationsGame (DSRAMF-G). First, following the arbitrary device selection of leaders (i.e., federations), the time cost minimization of followers (i.e., EDs) is modeled as a convex problem to obtain the optimal resource allocation. Then, based on followers’ optimal responses, device selection is modeled as a congestion game. We prove the existence of the Nash equilibrium and propose a decentralized mechanism. Finally, extensive experiments show that DSRAMF-G significantly outperforms the state-of-the-art methods, achieving up to 5.9x training speedup and 2.8x resource-savings. Shucun Fu, Fang Dong 0001, Dian Shen, Jinghui Zhang 0001, Zhaowu Huang, Qiang He 0001 |
IEEE Trans. Serv. Comput. | 5 |
| 2023 | Adaptive Overlap Padding and Resolution Selection for Frame Split-based Edge Video AnalyticsabstractFor providing accurate and fast on-device high-resolution video analytics, edge-assisted methods are widely proposed by using lower-resolution frames and splitting them with overlap padding. The use of low-resolution frames can significantly reduce the computational workload of video analytics. Dividing a frame into multiple overlapping parts simultaneously improves parallelism and maintains processing accuracy. However, accuracy and latency serve as a pair of tradeoff metrics, and prioritizing one to optimize overlap padding or resolution selection will compromise the other metric. Fortunately, we have discovered that in real-world video analytics scenarios with varying object sizes, it is not necessary to simultaneously achieve a high overlap padding size and high resolution. Hence, how to set appropriate overlap padding size and resolution to strike a balance between accuracy and latency in practical scenarios is a nontrivial and intractable problem. To this end, we propose an online learning-based method to achieve adaptive overlap padding and resolution selection, called APR. We model the problem as an integer programming and propose a Muli-armed bandit (MAB) theory-based algorithm to solve it. We discretize the continuum overlap padding size into a finite set to narrow explore space and set the frame split strategy as context information to achieve fast convergence. Theoretical analysis reveals APR achieves sub-linear regret. Extensive experimental results show APR outperforms the benchmark methods, achieving up to 2.06 × speedup in terms of latency and 0.18× increase in accuracy. Haopeng Zhu, Zhaowu Huang, Xiaolin Guo, Mengyang Liu, Baijun Chen, Fang Dong 0001 |
ICPADS | 2 |
| 2023 | WAEVSR: Enabling Collaborative Live Video Super-Resolution in Wide-Area MEC EnvironmentabstractLive video streaming is increasingly popular for its rich content and real-time interactions, but its demand for bandwidth has put a heavy burden on backbone networks. To save bandwidth, recent studies have proposed neural-enhanced live video streaming that deploys deep neural networks (DNNs) for video super-resolution (VSR) on end devices or nearby edge devices to enhance video quality by taking low-resolution frames as input and producing high-resolution output frames. In this solution, the high computational demands of high-quality VSR DNNs make them difficult to support on single end or edge device, necessitating the use of distributed resources in edge facilities. However, the distributed deployment of high-quality VSR DNNs for low-latency inference remains challenging due to the inherent data dependencies of VSR DNNs and the heterogeneity and dynamics of edge facilities. In this paper, we present WAEVSR, a novel collaborative neural-enhanced live video super-resolution system that enables effective leverage of distributed resources to maximize the latency-bounded quality in wide-area MEC environments. WAEVSR consists of two key components: 1) It deploys a parallel-friendly video super-resolution DNN among edge devices, 2) with an inference controller based on the variable-size sliding window to balance the latency and quality of distributed inference in the heterogeneous and dynamics MEC environment. Prototype-based evaluation shows that WAEVSR can achieve 2.5 × lower end-to-end latency than traditional super-resolution serving with a 0.01 drop in SSIM score. The case study also demonstrates its higher stability on latency than vanilla distributed MEC deployment. Daheng Yin, Fang Dong 0001, Baijun Chen, Dian Shen, Ruiting Zhou, Xiaolin Guo, Zhaowu Huang |
IWQoS | 7 |
| 2023 | ROIAdaptor: Adaptive Task Offloading of ROI-Encoded Videos for Edge Video AnalyticsabstractReal-time analytics on video data demands intensive computation resources and high bandwidth consumption. Edge computing enables us to offload resource-intensive analytics tasks to nearby edge servers, effectively reducing the extended latency. Numerous studies have applied ROI encoding technology to decrease video data size, thereby decreasing transmission latency. However, ROI encoding will reduce inference accuracy and indirectly affect inference latency. Existing works have ignored these impacts, leading to significant performance degradation. In this paper, we first identified the necessity of considering video QP settings when offloading, and investigated the impacts of changing QP settings on accuracy and latency to demonstrate that adaptive offloading based on QP settings can improve the efficiency of video analytics. We formulate the problem of minimizing average latency to meet real-time requirements under accuracy constraints, and propose an online offloading algorithm called ROIAdaptor, based on a contextual multi-armed bandit method. Our algorithm is developed based on LinUCB, a contextual multi-armed bandit method, and operates online with historical information, achieving a provable performance bound. Simulation results show that ROIAdaptor can reduce the overall latency by an average of 35.7% and meet the requirements for real-time video analysis, with virtually no loss in accuracy. Zhenxuan Xu, Xiaolin Guo, Zhaowu Huang, Shucun Fu, Fang Dong 0001 |
MSN | 3 |
| 2023 | Multi-Exit DNN Inference Acceleration Based on Multi-Dimensional Optimization for Edge IntelligenceabstractEdge intelligence, as a prospective paradigm for accelerating DNN inference, is mostly implemented by model partitioning which inevitably incurs the large transmission overhead of DNN's intermediate data. A popular solution introduces multi-exit DNNs to reduce latency by enabling early exits. However, existing work ignores the correlation between exit settings and synergistic inference, causing incoordination of device-to-edge. To address this issue, this paper first investigates the bottlenecks of executing multi-exit DNNs in edge computing and builds a novel model for inference acceleration with exit selection, model partition, and resource allocation. To tackle the intractable coupling subproblems, we propose a Multi-exit DNN inference Acceleration framework based on Multi-dimensional Optimization (MAMO). In MAMO, the exit selection subproblem is first extracted from the original problem. Then, bidirectional dynamic programming is employed to determine the optimal exit setting for an arbitrary multi-exit DNN. Finally, based on the optimal exit setting, a DRL-based policy is developed to learn joint decisions of model partition and resource allocation. We deploy MAMO on a real-world testbed and evaluate its performance in various scenarios. Extensive experiments show that it can adapt to heterogeneous tasks and dynamic networks, and accelerate DNN inference by up to 13.7x compared with the state-of-the-art. Fang Dong 0001, Huitian Wang, Dian Shen, Zhaowu Huang, Qiang He 0001, Jinghui Zhang 0001, Liangsheng Wen |
IEEE Trans. Mob. Comput. | 4 |
| 2022 | Enabling Latency-Sensitive DNN Inference via Joint Optimization of Model Surgery and Resource Allocation in Heterogeneous EdgeabstractNowadays, edge computing is widely adopted to resolve the emerging deep neural networks (DNNs)-driven intelligence scenarios with the requirement of low-latency and high-accuracy, which includes heterogeneous end devices and DNNs. In such scenarios, the influx of data and computation into a shared edge server incurs prohibitive latency. Thus, we exploit the advantage of Multi-exit DNNs (ME-DNNs) that tasks can exit early at appropriate depths to save inference time. However, naively using ME-DNNs in the heterogeneous edge still fails to deliver fast inference due to improper model surgery and resource allocation. Zhaowu Huang, Fang Dong 0001, Dian Shen, Huitian Wang, Xiaolin Guo, Shucun Fu |
ICPP | 1 |
| 2022 | Exploiting the Computational Path Diversity with In-network Computing for MECabstractWith Computing in the Network technologies, Mobile Edge Computing (MEC) has expanded the resource distribution and tightly integrated computing-network capabilities from the end-devices, through the edge, to the cloud infrastructure, including at points in between. Thus, edge computing is able to deliver a more collaborative processing, better service responding to the increasing application needs in low latency processing. In the presence of integrated computing-network resources and their increased capacity, current proximity-to-data methods in edge computing lead to sub-optimal performance in terms of processing latency. Addressing this issue, this paper presents a Low-latency Adaptive Workload Allocation framework (LAWA) to harness the growing in-network computing resources to deliver low latency processing capabilities for emerging latency-constrained applications. LAWA defines an application by its computational source and destination. Considering the diversity of computing and network resources, we try to find an optimal computational path and its workload allocation. We model the problem as a mixed integer programming problem. To solve this problem, we propose the computational pathfinding and workload allocation algorithms with optimality guarantees. Experimental results show that, comparing with the state-of-the-art methods, our method achieves up to 8.04× speedup, in terms of end-to-end latency. Xiaolin Guo, Fang Dong 0001, Dian Shen, Zhaowu Huang, Zhenyang Ni, Yulong Jiang, Daheng Yin |
SECON | 4 |
| 2021 | Enabling Low Latency Edge Intelligence based on Multi-exit DNNs in the WildabstractIn recent years, deep neural networks (DNNs) have witnessed a booming of artificial intelligence Internet of Things applications with stringent demands across high accuracy and low latency. A widely adopted solution is to process such computation-intensive DNNs inference tasks with edge computing. Nevertheless, existing edge-based DNN processing methods still cannot achieve acceptable performance due to the intensive transmission data and unnecessary computation. To address the above limitations, we take the advantage of Multi-exit DNNs (ME-DNNs) that allows the tasks to exit early at different depths of the DNN during inference, based on the input complexity. However, naively deploying ME-DNNs in edge still fails to deliver fast and consistent inference in the wild environment. Specifically, 1) at the model-level, unsuitable exit settings will increase additional computational overhead and will lead to excessive queuing delay; 2) at the computation-level, it is hard to sustain high performance consistently in the dynamic edge computing environment. In this paper, we present a Low Latency Edge Intelligence Scheme based on Multi-Exit DNNs (LEIME) to tackle the aforementioned problem. At the model-level, we propose an exit setting algorithm to automatically build optimal ME-DNNs with lower time complexity; At the computation-level, we present a distributed offloading mechanism to fine-tune the task dispatching at runtime to sustain high performance in the dynamic environment, which has the property of close-to-optimal performance guarantee. Finally, we implement a prototype system and extensively evaluate it through testbed and large-scale simulation experiments. Experimental results demonstrate that LEIME significantly improves applications' performance, achieving 1.1–18.7 × speedup in different situations. Zhaowu Huang, Fang Dong 0001, Dian Shen, Junxue Zhang 0001, Huitian Wang, Guangxing Cai, Qiang He 0001 |
ICDCS | 1 |
| 2019 | ADDA: Adaptive Distributed DNN Inference Acceleration in Edge Computing EnvironmentabstractImplementing intelligent mobile applications on IoT devices with DNN technology has become an inevitable trend. Due to the limitations of the size of DNN model deployed onto end devices and the instability of wide-area network transmission, either End-only mode or Cloud-only mode cannot guarantee the reasonable latency and recognition accuracy simultaneously. A better solution is to exploit the edge computing, where the existing edge computing execution framework and offloading mechanism for DNN inference suffer unnecessary computational overheads and underutilized computing capacity of end and edge. To address these shortcomings, an adaptive distributed DNN inference acceleration framework for edge computing environment is proposed in this paper, where DNN computation path optimization and DNN computation partition optimization are taken into consideration. The evaluations demonstrate that our method can effectively accelerate the DNN inference compared to the state-of-the-art methods. Huitian Wang, Guangxing Cai, Zhaowu Huang, Fang Dong 0001 |
ICPADS | 3 |