EDBT 2026 Demo / reviewers in the wild / expert
Miao Zhang 0003
dblp:60/7041-3
· DBLP profile ↗
24ranked-venue papers
7as first author
19since 2021 · last 2026
0000-0001-6126-6142ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 17 · 5 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 8 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | 4DGStream: Variable Bitrate Dynamic Gaussian Splatting StreamingabstractWhile 3D Gaussian Splatting (3DGS) has revolutionized static scene representation, the extension to dynamic scene, i.e., 3DGS video (GSV), faces challenges related to reconstruction quality, rendering speed, and storage requirements. The substantial data volume of current GSV poses significant hurdles for streaming applications, particularly in the realm of AR, VR and MR. To tackle these challenges, we introduce 4DGStream, a novel framework that integrates an efficient GSV compression method, Light4D, and a bitrate adaptation streaming strategy, QoSmooth, to ensure smooth playback while maintaining high visual quality. Light4D employs a binarizationassisted spatiotemporal deformation network to model the deformation of Gaussian primitive attributes over time, while a spatiotemporal-aware masking module prunes trivial Gaussians, further enhancing long-term reconstruction quality. To reduce storage, Light4D uses a binary hash grid to model the entropy of attributes for arithmetic coding, with its binary nature allowing efficient entropy modeling via a Bernoulli distribution. These components enable Light4D to improve the FPS/Storage metric by up to 12.4× over SpacetimeGS and 26.4× over 4DGS on the Neu3D dataset, with performance gains exceeding 3× orders of magnitude compared to other NeRF-based state-of-the-art (SOTA) methods. Here, FPS/Storage reflects the balance between rendering speed and data storage. Despite significant model size reductions, Light4D maintains or surpasses the reconstruction quality of 4DGS. Furthermore, QoSmooth provides effective rate control to enhance playback smoothness, reducing bitrate level switches by 61.6% and increasing time-average utility by 26.2%. All these improvements make 4DGStream highly suited for GSV streaming, improving QoE by 36.7% compared to SOTA methods. Zhicheng Liang, Dayou Zhang, Linfeng Shen, Miao Zhang 0003, Jian Zhang 0054, Bin Ju, Mallesham Dasari, Fangxin Wang 0001, Jiangchuan Liu |
IEEE Trans. Multim. | 4 |
| 2026 | Implicit Representation-based Volumetric Video Streaming for Photorealistic Full-scene ExperienceabstractThe widespread integration of the Internet of Things with sensors like depth-of-field cameras, LiDAR scanners, and eye-tracking infrared sensors, in head-mounted devices, has ushered in a new era of immersive digital experiences. Full-scene volumetric video (VV), a key innovation in this integration, provides a deeply immersive experience by capturing the richness and detail of the 3D world. However, its massive data volume presents significant streaming challenges. While 3D tile-based viewport approaches have been proposed, they struggle to full-scene VV given the small video buffer limitation, high tile segmentation overhead, and lack of full-scene consideration. In this work, inspired by the advancements of implicit neural radiance field (NeRF), we present \({\mathsf{V}^{2}\mathsf{NeRF}}\) , a novel full-scene VV streaming system featured by layered representation. It harmonizes the NeRF with explicit point clouds to represent the static background and dynamic foreground, thereby avoiding large data transfers and achieving photorealistic content representation. To tackle the issues of intensive computation requirements and multiscale adaptation scheduling within \({\mathsf{V}^{2}\mathsf{NeRF}}\) system, we propose a lightweight non-visible background removal method and a two-stage decoupled architecture. In addition, an efficient buffer-aware simulated annealing algorithm is developed, alongside the utilization of a perceptually learned metric, to enhance user experience. We further discuss the concerns about practical development and deployment. Extensive prototype evaluations demonstrate \({\mathsf{V}^{2}\mathsf{NeRF}}\) ’s superior streaming and viewing performance on a wide variety of networks, viewing motions, and scenes. For instance, compared to state-of-the-art approaches, it achieves a 24% increment in perceptual quality, an 83% reduction in rebuffering time, and a 54% enhancement in user experience on average. Jianxin Shi 0005, Miao Zhang 0003, Linfeng Shen, Jiangchuan Liu, Yuan Zhang 0013, Lingjun Pu, Jingdong Xu |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2025 | Commercial Dishes Can Be My Ladder: Sustainable and Collaborative Data Offloading in LEO Satellite Networks
Yi Ching Chou, Long Chen 0025, Hengzhi Wang, Feng Wang 0001, Hao Fang 0012, Haoyuan Zhao, Miao Zhang 0003, Xiaoyi Fan 0001 |
INFOCOM | 7 |
| 2025 | MANSY: Generalizing Neural Adaptive Immersive Video Streaming With Ensemble and Representation LearningabstractThe popularity of immersive videos has prompted extensive research into neural adaptive tile-based streaming to optimize video transmission over networks with limited bandwidth. However, the diversity of users’ viewing patterns and Quality of Experience (QoE) preferences has not been fully addressed yet by existing neural adaptive approaches for viewport prediction and bitrate selection. Their performance can significantly deteriorate when users’ actual viewing patterns and QoE preferences differ considerably from those observed during the training phase, resulting in poor generalization. In this paper, we proposeMANSY, a novel streaming system that embraces user diversity to improve generalization. Specifically, to accommodate users’ diverse viewing patterns, we design a Transformer-based viewport prediction model with an efficient multi-viewport trajectory input output architecture based on implicit ensemble learning. Besides, we for the first time combine the advanced representation learning and deep reinforcement learning to train the bitrate selection model to maximize diverse QoE objectives, enabling the model to generalize across users with diverse preferences. Extensive experiments demonstrate thatMANSYoutperforms state-of-the-art approaches in viewport prediction accuracy and QoE improvement on both trained and unseen viewing patterns and QoE preferences, achieving better generalization. Duo Wu, Panlong Wu, Miao Zhang 0003, Fangxin Wang 0001 |
IEEE Trans. Mob. Comput. | 3 |
| 2025 | Towards Neural Codec-Empowered 360$^\circ$ Video Streaming: A Saliency-Aided Synergistic ApproachabstractNetworked 360$^\circ$video has become increasingly popular. Despite the immersive experience for users, its sheer data volume, even with the latest H.266 coding and viewport adaptation, remains a significant challenge to today's networks. Recent studies have shown that integrating deep learning into video coding can significantly enhance compression efficiency, providing new opportunities for high-quality video streaming. In this work, we conduct a comprehensive analysis of the potential and issues in applying neural codecs to 360$^\circ$video streaming. We accordingly present$\mathsf {NETA}$, a synergistic streaming scheme that merges neural compression with traditional coding techniques, seamlessly implemented within an edge intelligence framework. To address the non-trivial challenges in the short viewport prediction window and time-varying viewing directions, we propose implicit-explicit buffer-based prefetching grounded in content visual saliency and bitrate adaptation with smart model switching around viewports. A novel Lyapunov-guided deep reinforcement learning algorithm is developed to maximize user experience and ensure long-term system stability. We further discuss the concerns towards practical development and deployment and have built a working prototype that verifies$\mathsf {NETA}$’s excellent performance. For instance, it achieves a 27% increment in viewing quality, a 90% reduction in rebuffering time, and a 64% decrease in quality variation on average, compared to state-of-the-art approaches. Jianxin Shi 0005, Miao Zhang 0003, Linfeng Shen, Jiangchuan Liu, Lingjun Pu, Jingdong Xu |
IEEE Trans. Multim. | 2 |
| 2024 | OAVS: Efficient Online Learning of Streaming Policies for Drone-sourced Live Video AnalyticsabstractDrone-sourced live video analytics has extensive applications across diverse domains. Adaptive video streaming is a pivotal technique in these applications that targets at effectively delivering video content to servers under varying network conditions, enabling complex analytics afterward. However, our thorough data analysis reveals that conventional offline video streaming policies cannot effectively adapt to highly fluctuating drone network environments and dynamic changes in aerial view scenes. This results in suboptimal analytic performance and necessitates online adaptation for policy models. Yet, obtaining ground-truth analytics results directly from drones is infeasible due to their limited capacity. Furthermore, naively streaming original videos to the server for online adaption is greatly challenged by the scarce and dynamic networks, leading to decreased accuracy performance and escalated transmission cost if not properly designed. In this paper, we present OAVS, a novel online learning-enabled adaptive streaming framework for drone-sourced video analytics. To facilitate cost-effective online retraining, we design a hierarchical reinforcement learning approach in which the upper-level module intelligently determines the timing for online retraining, balancing machine-perceived quality of experience (QoE) improvement and transmission cost. Meanwhile, the lower-level module dynamically allocates bitrate to maximize machine-perceived QoE. Extensive experiments based on real-world drone video and aerial network datasets demonstrate that our proposed framework achieves a 17.7% mean accuracy increase, a 37.5% decrease in the mean failure rate of video uploading, and a 5.2% mean latency decrease compared to state-of-the-art solutions. Miao Zhang 0003, Yifei Zhu 0001 |
IWQoS | 2 |
| 2024 | Robust Live Streaming over LEO Satellite Constellations: Measurement, Analysis, and Handover-Aware AdaptationabstractLive streaming has experienced significant growth recently. Yet this rise in popularity contrasts with the reality that a substantial segment of the global population still lacks Internet access. The emergence of Low Earth orbit Satellite Networks (LSNs), such as SpaceX's Starlink and Amazon's Project Kuiper, presents a promising solution to fill this gap. Nevertheless, our measurement study reveals that existing live streaming platforms may not be able to deliver a smooth viewing experience on LSNs due to frequent satellite handovers, which lead to frequent video rebuffering events. Current state-of-the-art learning-based Adaptive Bitrate (ABR) algorithms, even when trained on LSNs' network traces, fail to manage the abrupt network variations associated with satellite handovers effectively. To address these challenges, for the first time, we introduce Satellite-Aware Rate Adaptation (SARA), a versatile and lightweight middleware that can seamlessly integrate with various ABR algorithms to enhance the performance of live streaming over LSNs. SARA intelligently modulates video playback speed and furnishes ABR algorithms with insights derived from the distinctive network characteristics of LSNs, thereby aiding ABR algorithms in making informed bitrate selections and effectively minimizing rebuffering events that occur during satellite handovers. Our extensive evaluation shows that SARA can effectively reduce the rebuffering time by an average of 39.41% and slightly improve latency by 0.65% while only introducing an overall loss in bitrate by 0.13%. Hao Fang 0012, Haoyuan Zhao, Jianxin Shi 0005, Miao Zhang 0003, Guanzhen Wu, Yi Ching Chou, Feng Wang 0001, Jiangchuan Liu |
ACM Multimedia | 4 |
| 2024 | FSVFG: Towards Immersive Full-Scene Volumetric Video Streaming with Adaptive Feature GridabstractGiven the truly immersive viewing experiences, full-scene volumetric videos have received increasing attention from both academia and industry. Their vast data volumes, however, present significant challenges for real-time streaming over today's bandwidth-limited Internet. Considering the vast amount of full-scene volumetric data to be streamed and the limited bandwidth on the Internet, achieving adaptive full-scene volumetric video streaming over the Internet presents a significant challenge. Inspired by the advantages offered by neural fields, especially the feature grid method, we propose FSVFG, a novel full-scene volumetric video streaming system integrated feature grids as the representation of volumetric content. FSVFG employs an incremental training approach for feature grids and stores the features and residuals between adjacent grids as frames. To support adaptive streaming, we delve into the data structure and rendering processes of feature grids and propose bandwidth adaptation mechanisms. The mechanisms involve a coarse ray-marching for the selection of features and residuals to be sent, and achieve variable bitrate streaming by Level-of-Detail (LoD) and residual filtering. Based on these mechanisms, FSVFG achieves adaptive streaming by adaptively balancing the transmission of feature and residual according to the available bandwidth. Our preliminary results demonstrate the effectiveness of FSVFG, demonstrating its ability to improve visual quality and reduce bandwidth requirements of full-scene volumetric video streaming. Daheng Yin, Jianxin Shi 0005, Miao Zhang 0003, Zhaowu Huang, Jiangchuan Liu, Fang Dong 0001 |
ACM Multimedia | 3 |
| 2024 | StarStream: Live Video Analytics over Space NetworkingabstractStreaming videos from resource-constrained front-end devices over networks to resource-rich cloud servers has long been a common practice for surveillance and analytics. Most existing live video analytics (LVA) systems, however, have been built over terrestrial networks, limiting their applications during natural disasters and in remote areas that desperately call for real-time visual data delivery and scene analysis. With the recent advent of space networking, in particular, Low Earth Orbit (LEO) satellite constellations such as Starlink, high-speed truly global Internet access is becoming available and affordable. This paper examines the challenges and potentials of LVA over modern LEO satellite networking (LSN). Using Starlink as the testbed, we have carried out extensive in-the-wild measurements to gain insights into its achievable performance for LVA. The results reveal that the uplink bottleneck in today's LSN, together with the volatile network conditions, can significantly affect the service quality of LVA and necessitate prompt adaptation. We accordingly develop StarStream, a novel LSN-adaptive streaming framework for LVA. At its core, StarStream is empowered by a Transformer-based network performance predictor tailored for LSN and a content-aware configuration optimizer. We discuss a series of key design and implementation issues of StarStream and demonstrate its effectiveness and superiority through trace-driven experiments with real-world network and video processing data. Miao Zhang 0003, Jiaxing Li 0006, Haoyuan Zhao, Linfeng Shen, Jiangchuan Liu |
ACM Multimedia | 1 |
| 2024 | You Only Look Once in Panorama: Object Detection for 360° Videos with MLaaSabstract360° videos are gaining popularity, but immersive analytics, particularly in object detection, confront challenges from complex scenes and high data volume. This imposes significant burdens on individual users and resource-limited edge devices. Fortunately, Machine Learning as a Service (MLaaS) offers an economical solution for quick deployment without specific hardware or expertise. However, current MLaaS are mostly 2D image-designated and not optimized for the distinctive characteristics of raw 360° video frames. In this paper, we propose a novel MLaaS-based system to address this challenge. Our solution partitions 360° frames into distortion-free 2D regions with dynamic region of interest prediction. We then present an image-stitching algorithm featuring Skyline representation, seamlessly combining all the 2D regions into a unified frame. This frame is then transmitted to the MLaaS platform, with the detected objects being back-projected to yield the final results. Our experiments demonstrate the superiority of this system over baselines, proving its effectiveness in 360° video object detection tasks. Linfeng Shen, Miao Zhang 0003, Cong Zhang 0002, Jiangchuan Liu |
NOSSDAV | 2 |
| 2024 | Towards Full-scene Volumetric Video Streaming via Spatially Layered Representation and NeRF GenerationabstractImmersive full-scene volumetric video (VV) showcases the richness and detail of the 3D world, yet poses significant streaming challenges given its massive data volume. Existing 3D tile-based viewport approaches struggle to effectively adapt to full-scene VV owing to their small video buffer limitation, high tile segmentation overhead, and lack of full-scene consideration. Jianxin Shi 0005, Miao Zhang 0003, Linfeng Shen, Jiangchuan Liu, Yuan Zhang 0013, Lingjun Pu, Jingdong Xu |
NOSSDAV | 2 |
| 2024 | AdaDSR: Adaptive Configuration Optimization for Neural Enhanced Video Analytics StreamingabstractNeural-based super-resolution (SR) has achieved great success in enhancing image or video quality, creating new opportunities for building bandwidth-efficient and high-accuracy video analytics (VAs) systems. Intuitively, with the help of SR techniques, cameras only need to send downsampled low-quality frames to the server in a canonical edge-assisted VAs framework. The server-side SR model then upscales the quality of received frames for the subsequent VAs tasks, incurring thus substantially reduced bandwidth consumption. Nonetheless, as revealed by our measurement results on real-world video clips, higher delivery quality does not necessarily lead to higher analysis accuracy. This motivates us to study the content-adaptive downsampling and upscaling ratio selection problem for VAs streaming. We propose an SR-based VAs framework, named AdaDSR that can dynamically select the optimal downsampling and upscaling ratios so that the system utility can be maximized. AdaSDR is configured to balance the tradeoffs among accuracy, network cost, and computational cost. It further leverages the temporal consistency of videos to skip trivial decisions so that the camera’s processing overhead can be reduced. Experiments on real-world video data sets demonstrate that AdaDSR can improve the average utility by 7.2%–18.4% when compared with state-of-the-art approaches under diverse video scenes. Sheng Cen, Miao Zhang 0003, Yifei Zhu 0001, Jiangchuan Liu |
IEEE Internet Things J. | 2 |
| 2024 | ILCAS: Imitation Learning-Based Configuration- Adaptive Streaming for Live Video Analytics With Cross-Camera CollaborationabstractThe high-accuracy and resource-intensive deep neural networks (DNNs) have been widely adopted by live video analytics (VA), where camera videos are streamed over the network to resource-rich edge/cloud servers for DNN inference. Common video encoding configurations (e.g., resolution and frame rate) have been identified with significant impacts on striking the balance between bandwidth consumption and inference accuracy and therefore their adaption scheme has been a focus of optimization. However, previous profiling-based solutions suffer from high profiling cost, while existing deep reinforcement learning (DRL) based solutions may achieve poor performance due to the usage of fixed reward function for training the agent, which fails to craft the application goals in various scenarios. In this paper, we proposeILCAS, the first imitation learning (IL) based configuration-adaptive VA streaming system. Unlike DRL-based solutions,ILCAStrains the agent with demonstrations collected from the expert which is designed as an offline optimal policy that solves the configuration adaption problem through dynamic programming. To tackle the challenge of video content dynamics,ILCASderives motion feature maps based on motion vectors which allowILCASto visually “perceive” video content changes. Moreover,ILCASincorporates a cross-camera collaboration scheme to exploit the spatio-temporal correlations of cameras for more proper configuration selection. Extensive experiments confirm the superiority ofILCAScompared with state-of-the-art solutions, with 2-20.9% improvement of mean accuracy and 19.9–85.3% reduction of chunk upload lag. Duo Wu, Dayou Zhang, Miao Zhang 0003, Fangxin Wang 0001, Shuguang Cui |
IEEE Trans. Mob. Comput. | 3 |
| 2023 | OmniSense: Towards Edge-Assisted Online Analytics for 360-Degree VideosabstractWith the reduced hardware costs of omnidirectional cameras and the proliferation of various extended reality applications, more and more 360° videos are being captured. To fully unleash their potential, advanced video analytics is expected to extract actionable insights and situational knowledge without blind spots from the videos. In this paper, we present OmniSense, a novel edge-assisted framework for online immersive video analytics. OmniSense achieves both low latency and high accuracy, combating the significant computation and network resource challenges of analyzing 360° videos. Motivated by our measurement insights into 360° videos, OmniSense introduces a lightweight spherical region of interest (SRoI) prediction algorithm to prune redundant information in 360° frames. Incorporating the video content and network dynamics, it then smartly scales vision models to analyze the predicted SRoIs with optimized resource utilization. We implement a prototype of OmniSense with commodity devices and evaluate it on diverse real-world collected 360° videos. Extensive evaluation results show that compared to resource-agnostic baselines, it improves the accuracy by 19.8% – 114.6% with similar end-to-end latencies. Meanwhile, it hits 2.0× – 2.4× speedups while keeping the accuracy on par with the highest accuracy of baselines. Miao Zhang 0003, Yifei Zhu 0001, Linfeng Shen, Fangxin Wang 0001, Jiangchuan Liu |
INFOCOM | 1 |
| 2023 | AIoT-Empowered Smart Grid Energy Management with Distributed Control and Non-Intrusive Load MonitoringabstractToday's electrical grid is experiencing a fast transition toward a smart infrastructure. Modern smart grid is expected to integrate Artificial Intelligence of Things (AIoT)-empowered energy management systems (EMS) to sense, analyze, and optimize the power consumption and QoS of diverse end users. Non-Intrusive Load Monitoring (NILM) plays a key role in this transition, particularly considering that many legacy devices/appliances may not have built-in sensors. Yet most of the NILM solutions rely on large (often impractical) datasets for training. In this paper, we address this challenge through a meta learning-inspired approach, which implements a hierarchical architecture with a “meta-learner” to supervise the training of each appliance. Current EMS also relies on a central controller to access long-term information across all participants, which mismatches their distributed nature, and so often with slow responses. To this end, we develop a deep reinforcement learning based controller to make dynamic decisions for each component in the system. The experiment results based on real-world data sets and simulation data show that applying the meta learning approach can greatly improve the performance of NILM and the QoS of the whole system. Linfeng Shen, Feng Wang 0001, Miao Zhang 0003, Jiangchuan Liu, Gaoyang Liu, Xiaoyi Fan 0001 |
IWQoS | 3 |
| 2023 | Understanding User Behavior in Volumetric Video Watching: Dataset, Analysis and PredictionabstractVolumetric video emerges as a new attractive video paradigm in recent years since it provides an immersive and interactive 3D viewing experience with six degree-of-freedom (DoF). Unlike traditional 2D or panoramic videos, volumetric videos require dense point clouds, voxels, meshes, or huge neural models to depict volumetric scenes, which results in a prohibitively high bandwidth burden for video delivery. Users' behavior analysis, especially the viewport and gaze analysis, then plays a significant role in prioritizing the content streaming within users' viewport and degrading the remaining content to maximize user QoE with limited bandwidth. Although understanding user behavior is crucial, to the best of our best knowledge, there are no available 3D volumetric video viewing datasets containing fine-grained user interactivity features, not to mention further analysis and behavior prediction. Kaiyuan Hu, Yili Jin 0001, Junhua Liu 0003, Yongting Chen, Miao Zhang 0003, Fangxin Wang 0001 |
ACM Multimedia | 6 |
| 2022 | CASVA: Configuration-Adaptive Streaming for Live Video AnalyticsabstractThe advent of high-accuracy and resource-intensive deep neural networks (DNNs) has fulled the development of live video analytics, where camera videos need to be streamed over the network to edge or cloud servers with sufficient computational resources. Although it is promising to strike a balance between available bandwidth and server-side DNN inference accuracy by adjusting video encoding configurations, the influences of fine-grained network and video content dynamics on configuration performance should be addressed. In this paper, we propose CASVA, a Configuration-Adaptive Streaming framework designed for live Video Analytics. The design of CASVA is motivated by our extensive measurements on how video configuration affects its bandwidth requirement and inference accuracy. To handle the complicated dynamics in live video analytics streaming, CASVA trains a deep reinforcement learning model which does not make any assumptions about the environment but learns to make configuration choices through its experiences. A variety of real-world network traces are used to drive the evaluation of CASVA. The results on a multitude of video types and video analytics tasks show the advantages of CASVA over state-of-the-art solutions. Miao Zhang 0003, Fangxin Wang 0001, Jiangchuan Liu |
INFOCOM | 1 |
| 2022 | CharmSeeker: Automated Pipeline Configuration for Serverless Video ProcessingabstractVideo processing plays an essential role in a wide range of cloud-based applications. It typically involves multiple pipelined stages, which well fits the latest fine-grained serverless computing paradigm if properly configured to match the cost and delay constraints of video. Existing configuration tools, however, are primarily developed for traditional virtual machine clusters with general workloads. This paper presents CharmSeeker, an automated configuration tuning tool for serverless video processing pipelines. We first carefully examine the key steps and the performance bottlenecks for video processing over modern serverless platforms. Then, we identify the configuration space for processing pipelines and leverage a carefully designed Sequential Bayesian Optimization search scheme to identify promising configurations. We further address the practical challenges toward integrating our solution into real-world systems and develop a prototype with AWS Lambda. Evaluation results show that CharmSeeker can find out the optimal or near-optimal configurations that improve the relative processing time up to 408.77%. It is also more robust and scalable to various video processing pipelines compared with state-of-the-art solutions. Miao Zhang 0003, Yifei Zhu 0001, Jiangchuan Liu, Feng Wang 0001, Fangxin Wang 0001 |
IEEE/ACM Trans. Netw. | 1 |
| 2021 | Towards cloud-edge collaborative online video analytics with fine-grained serverless pipelinesabstractThe ever-growing deployment scale of surveillance cameras and the users' increasing appetite for real-time queries have urged online video analytics. Synergizing the virtually unlimited cloud resources with agile edge processing would deliver an ideal online video analytics system; yet, given the complex interaction and dependency within and across video query pipelines, it is easier said than done. This paper starts with a measurement study to acquire a deep understanding of video query pipelines on real-world camera streams. We identify the potentials and practical challenges towards cloud-edge collaborative video analytics. We then argue that the newly emerged serverless computing paradigm is the key to achieve fine-grained resource partitioning with minimum dependency. We accordingly propose CEVAS, a Cloud-Edge collaborative Video Analytics system empowered by fine-grained Serverless pipelines. It builds flexible serverless-based infrastructures to facilitate fine-grained and adaptive partitioning of cloud-edge workloads for multiple concurrent query pipelines. With the optimized design of individual modules and their integration, CEVAS achieves real-time responses to highly dynamic input workloads. We have developed a prototype of CEVAS over Amazon Web Services (AWS) and conducted extensive experiments with real-world video streams and queries. The results show that by judiciously coordinating the fine-grained serverless resources in the cloud and at the edge, CEVAS reduces 86.9% cloud expenditure and 74.4% data transfer overhead of a pure cloud scheme and improves the analysis throughput of a pure edge scheme by up to 20.6%. Thanks to the fine-grained video content-aware forecasting, CEVAS is also more adaptive than the state-of-the-art cloud-edge collaborative scheme. Miao Zhang 0003, Fangxin Wang 0001, Yifei Zhu 0001, Jiangchuan Liu, Zhi Wang 0001 |
MMSys | 1 |
| 2019 | Rendering multi-party mobile augmented reality from edgeabstractMobile augmented reality (MAR) augments a real-world environment (probably surrounding or close to the mobile user) by computer-generated perceptual information. Utilizing the emerging edge computing paradigm in MAR systems can reduce the power consumption and computation load for the mobile devices and improve responsiveness of the MAR service. Different from existing studies that mainly explored how to better enable the MAR services utilizing edge computing resources, our focus is to optimize the video generation stage of the edge-based MAR services-efficiently using the available edge computing resources to render and encode the augmented reality as video streams to the mobile clients. Specifically, for multi-party AR applications, we identify the advantages and disadvantages of two encoding schemes, namely colocated encoding and spilt encoding, and examine the trade-off between performance and scalability when the rendering and encoding tasks are colocated or split. Towards optimally placing AR video rendering and encoding in the edge, we formulate and solve the rendering and encoding task assignment problem for multi-party edge-based MAR services to maximize the QoS for the users and the edge computing efficiency. The proposed task assignment scheme is proved to be superior through extensive trace-driven simulations and experiments on our prototype system. Lei Zhang 0066, Andy Sun, Ryan Shea, Jiangchuan Liu, Miao Zhang 0003 |
NOSSDAV | 5 |
| 2019 | Video processing with serverless computing: a measurement studyabstractThe growing demand for video processing and the advantages in scalability and cost reduction brought by the emerging serverless computing have attracted significant attention in serverless computing powered video processing. However, how to implement and configure serverless functions to optimize the performance and cost of video processing applications remains unclear. In this paper, we explore the configuration and implementation schemes of typical video processing functions deployed to the serverless platforms and quantify their influence on the execution duration and monetary cost from a developer's perspective. Our measurement reveals that memory configuration is non-trivial. Dynamic profiling of workloads is necessary to find the best memory configuration. Moreover, compared with calling external video processing APIs, implementing these services locally in serverless functions can be competitive. We also find that the performance of video processing applications could be affected by the underlying infrastructure. Our work provides guidelines for further function-level optimization and complements the existing measurement studies for both serverless computing and video processing. Miao Zhang 0003, Yifei Zhu 0001, Cong Zhang 0002, Jiangchuan Liu |
NOSSDAV | 1 |
| 2018 | Delay-Aware Upload Balancing Cross APs Based on User RelaysabstractRecent years have witnessed the rapid growth of user-generated content, produced by end users and uploaded via edge networks, in which 802.11 wireless LANs (Wi-Fi) serve a dominant fraction of such traffic. However, the user experience for content uploading via Wi-Fi networks, especially in today's metropolises, remains unsatisfactory. In this paper, we address this problem using a delay-aware wireless Access Point (AP) upload balancing solution: content items can be carried by users from one overloaded AP to another AP for uploading before their deadlines, to balance the load between Wi-Fi APs and achieve a high overall upload capacity. Our contributions are as follows. First, we conduct measurement studies to validate that the AP upload balancing framework is feasible in practice. Second, we formulate the delay-aware AP upload balancing problem as a Cache Utility Maximization (CUM) model and propose an efficient algorithm to guarantee a high success upload rate and average AP load utilization. Third, trace-driven experiments are conducted to evaluate our solution. The results show that our solution can achieve significant performance improvements. Miao Zhang 0003, Zhi Wang 0001, Yong Jiang 0001 |
ICC | 1 |
| 2017 | Understanding Performance of Edge Content Caching for Mobile Video StreamingabstractToday's Internet has witnessed an increase in the popularity of mobile video streaming, which is expected to exceed 3/4 of the global mobile data traffic by 2019. To satisfy the considerable amount of mobile video requests, video service providers have been pushing their content delivery infrastructure to edge networks-from regional content delivery network (CDN) servers to peer CDN servers (e.g., smartrouters in users' homes)-to cache content and serve users with storage and network resources nearby. Among the edge network content caching paradigms, Wi-Fi access point caching and cellular base station caching have become two mainstream solutions. Thus, understanding the effectiveness and performance of these solutions for large-scale mobile video delivery is important. However, the characteristics and request patterns of mobile video streaming are unclear in practical wireless network. In this paper, we use real-world data sets containing 50 million trace items of nearly 2 million users viewing more than 0.3 million unique videos using mobile devices in a metropolis in China over two weeks, not only to understand the request patterns and user behaviors in mobile video streaming, but also to evaluate the effectiveness of Wi-Fi and cellular-based edge content caching solutions. To understand the performance of edge content caching for mobile video streaming, we first present temporal and spatial video request patterns, and we analyze their impacts on caching performance using frequency-domain and entropy analysis approaches. We then study the behaviors of mobile video users, including their mobility and geographical migration behaviors, which determine the request patterns. Using trace-driven experiments, we compare strategies for edge content caching, including least recently used (LRU) and least frequently used (LFU), in terms of supporting mobile video requests. We reveal that content, location, and mobility factors all affect edge content caching performance. Moreover, we design an efficient caching strategy based on the measurement insights and experimentally evaluate its performance. The results show that our design significantly improves the cache hit rate by up to 30% compared with LRU/LFU. Ge Ma, Zhi Wang 0001, Miao Zhang 0003, Jiahui Ye, Minghua Chen 0001, Wenwu Zhu 0001 |
IEEE J. Sel. Areas Commun. | 3 |
| 2017 | Propagation- and Mobility-Aware D2D Social Content ReplicationabstractMobile online social network services have seen rapid expansion; thus, the corresponding huge amounts of user-generated social media contents propagating between users via social connections have significantly challenged the traditional content delivery paradigm. First, replicating all the contents generated by users to edge servers that well “fit” the receivers becomes difficult due to limited bandwidth and storage capacities. Motivated by device-to-device (D2D) communication, which allows users with smart devices to transfer content directly, we propose replicating bandwidth-intensive social contents in a device-to-device manner. Based on large-scale measurement studies on social content propagation and user mobility patterns in edge-network regions, we observe the following: (1) Device-to-device replication can significantly help users download social contents from neighboring peers. (2) Both social propagation and mobility patterns affect how contents should be replicated. (3) The replication strategies depend on regional characteristics (e.g., how users move across regions). Using these measurement insights, we propose a propagationand mobility-aware content replication strategy for edge-network regions, in which social contents are assigned to users in edge-network regions according to a joint consideration of social graphs, content propagation, and user mobility. We formulate the replication scheduling as an optimization problem and design a distributed algorithm using only historical, local, and partial information to solve it. Trace-driven experiments further verify the superiority of our proposal: compared with conventional pure-movement-based and popularity-based approaches, our design can significantly improve (2 - 4-fold improvement) the amount of social content successfully delivered via device-to-device replication. Zhi Wang 0001, Lifeng Sun, Miao Zhang 0003, Haitian Pang, Erfang Tian, Wenwu Zhu 0001 |
IEEE Trans. Mob. Comput. | 3 |