Yao Liu 0001

dblp:64/424-1 · DBLP profile ↗
← Back
60ranked-venue papers
16as first author
21since 2021 · last 2026
0000-0003-2817-0725ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 29 · 5 first-author · 15 since 2021Computer networks · 20 · 7 first-author · 5 since 2021Systems, architecture and hardware · 11 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Security and privacy · 2 · 2 first-authorDatabases, data management, data science and information retrieval · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 first-authorSoftware engineering, systems software and programming languages · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Egocentric Daily Video Question Answering with Token-Efficient Storyboard Retrieval
Tun-Yuan Chang, Cheng-Hsin Hsu, Yao Liu 0001
NOSSDAV4
2025 A Distributed Framework for Privacy-Enhanced Vision Transformers on the Edge
abstract
Nowadays, visual intelligence tools have become ubiquitous, offering all kinds of convenience and possibilities. However, these tools have high computational requirements that exceed the capabilities of resource-constrained mobile and wearable devices. While offloading visual data to the cloud is a common solution, it introduces significant privacy vulnerabilities during transmission and server-side computation. To address this, we propose a novel distributed, hierarchical offloading framework for Vision Transformers (ViTs) that addresses these privacy challenges by design. Our approach uses a local trusted edge device, such as a mobile phone or an Nvidia Jetson, as the edge orchestrator. This orchestrator partitions the user's visual data into smaller portions and distributes them across multiple independent cloud servers. By design, no single external server possesses the complete image, preventing comprehensive data reconstruction. The final data merging and aggregation computation occurs exclusively on the user's trusted edge device. We apply our framework to the Segment Anything Model (SAM) as a practical case study, which demonstrates that our method substantially enhances content privacy over traditional cloud-based approaches. Evaluations show our framework maintains near-baseline segmentation performance while substantially reducing the risk of content reconstruction and user data exposure. Our framework provides a scalable, privacy-preserving solution for vision tasks in the edge-cloud continuum.
Mufeng Zhu, Zhongze Tang, Sheng Wei 0001, Yao Liu 0001
SEC5
2025 EyeNavGS: A 6-DoF Navigation Dataset and Record-n-Replay Software for Real-World 3DGS Scenes in VR
abstract
3D Gaussian Splatting (3DGS) is an emerging media representation that reconstructs real-world 3D scenes in high fidelity, enabling 6-degrees-of-freedom (6-DoF) navigation in virtual reality (VR). However, developing and evaluating 3DGS-enabled applications and optimizing their rendering performance require realistic user navigation data. Such data is currently unavailable for photorealistic 3DGS reconstructions of real-world scenes. This paper introduces EyeNavGS, the first publicly available 6-DoF navigation dataset featuring traces from 46 participants exploring twelve diverse, real-world 3DGS scenes. The dataset was collected at two sites, using the Meta Quest Pro headsets, recording the head pose and eye gaze data for each rendered frame during free world standing 6-DoF navigation. For each of the twelve scenes, we performed careful scene initialization to correct for scene tilt and scale, ensuring a perceptually-comfortable VR experience. We also release our open-source SIBR viewer software fork with record-and-replay functionalities and a suite of utility tools for data processing, conversion, and visualization. The EyeNavGS dataset and its accompanying software tools provide valuable resources for advancing research in 6-DoF viewport prediction, adaptive streaming, 3D saliency, and foveated rendering for 3DGS scenes. The EyeNavGS dataset is available at: https://symmru.github.io/EyeNavGS/
Cheng-Tse Lee, Mufeng Zhu, Yuan-Chun Sun, Cheng-Hsin Hsu, Yao Liu 0001
ACM Multimedia7
2025 NeRFCompressor: Enhancing Dynamic Scene Representation for Efficient 6-DoF Object Transportation
abstract
3D scene modeling is essential for immersive experiences in Virtual, Augmented, and Mixed Reality (VR/AR/MR) applications. Neural Radiance Fields (NeRF) have emerged as a strong alternative to traditional representations such as meshes and point clouds for 6-DoF rendering. However, maintaining high visual quality while enabling efficient transmission in dynamic environments remains a significant challenge. In this paper, we propose NeRFCompressor, a novel compression framework for dynamic scene representation using NeRF-like models. Building on tensor decomposition-based 3D reconstruction, NeRFCompressor improves transmission efficiency by leveraging existing video codecs to exploit both intra-scene and inter-scene redundancies. It maintains high QoE with minimal degradation in reconstruction quality. Experiments show that NeRFCompressor outperforms state-of-the-art methods in compressing both static and dynamic scene representations.
Jin Zhou 0006, Mufeng Zhu, Yao Liu 0001, Songqing Chen
MMSP3
2025 LTS: A DASH Streaming System for Dynamic Multi-Layer 3D Gaussian Splatting Scenes
abstract
We present a novel DASH-based streaming system for dynamic 3D Gaussian Splatting (3DGS) scenes, addressing the challenges of streaming large amounts of 3DGS data over diverse and dynamic networks. Our Layer, Tile, and Segment Adaptive streaming (LTS) system combines three key features: (i) multi-layer streaming, which adapts to diverse client capabilities while balancing visual quality and bandwidth usage, (ii) tiled streaming, which reduces unnecessary data transmission by focusing on the user's viewport, and (iii) segment streaming, which divides dynamic 3DGS scenes into segments, letting clients request them dynamically to handle network fluctuations. Our experimental results demonstrate that our LTS system achieves superior performance in both live and on-demand streaming of dynamic 3DGS scenes compared to the baselines. For example, in live streaming, LTS could achieve up to 99.70% reduction in missing frames on average and deliver a maximum PSNR (Peak Signal-to-Noise Ratio) improvement of 10.08 dB. In on-demand streaming, LTS could reduce the freeze time by up to 92.01%, and increase the synthesized view quality by up to 5.14 dB in PSNR and 0.11 in SSIM (Structural Similarity Index). Our source codes are available at: https://github.com/AIINS-NTHU/LTS-DASH-Streaming-System-for-3DGS.
Yuan-Chun Sun, Yuang Shi, Cheng-Tse Lee, Mufeng Zhu, Wei Tsang Ooi, Yao Liu 0001, Chun-Ying Huang, Cheng-Hsin Hsu
MMSys6
2025 SGSS: Streaming 6-DoF Navigation of Gaussian Splat Scenes
abstract
3D Gaussian Splatting (3DGS) is an emerging approach for training and representing real-world 3D scenes. Due to its photorealistic novel view synthesis and fast rendering speed (e.g., over 100 FPS), it has the potential to transform how scenes that can be explored in 6 degrees-of-freedom (6-DoF) are represented. However, a limiting factor of 3DGS is its large size, which requires high network bandwidth for streaming reconstructed real-world 3D scenes.
Mufeng Zhu, Mingju Liu, Cunxi Yu, Cheng-Hsin Hsu, Yao Liu 0001
MMSys5
2025 Privacy-Preserving Multimedia Mobile Cloud Computing Using Cost-Effective Protective Perturbation
abstract
Mobile cloud computing has been adopted in many multimedia applications, where resource-constrained mobile devices send multimedia data (e.g., images) to remote cloud servers to request computation intensive multimedia services (e.g., image recognition). Despite the performance improvement, the cloud-based mechanism often causes privacy concerns as the user data is offloaded to untrusted cloud servers. Existing solutions require computation-intensive perturbation generation on resource-constrained mobile devices. Also, the protected images are not compliant with standard image compression algorithms, leading to significant bandwidth consumption. We develop a novel privacy-preserving multimedia mobile cloud computing framework, namely PMC2, to address the resource and bandwidth challenges. PMC2 employs confidential computing on an edge server to deploy the perturbation generator, which addresses the on-device resource challenge. Also, we develop a neural compressor for the protected images to address the bandwidth challenge. Our evaluations of PMC2 demonstrate superior latency, power efficiency, and bandwidth consumption while maintaining high accuracy in the target multimedia service.
Zhongze Tang, Mengmei Ye, Yao Liu 0001, Sheng Wei 0001
NOSSDAV4
2025 EVASR: Edge-Based Salience-Aware Super-Resolution for Enhanced Video Quality and Power Efficiency
abstract
With the rapid growth of video content consumption, it is important to deliver high-quality streaming videos to users even under limited available network bandwidth. In this article, we propose EVASR, a system that performs edge-based video delivery to clients with salience-aware super-resolution. We select patches with higher saliency score to perform super-resolution while applying the simple yet efficient bicubic interpolation for the remaining patches in the same video frame. To efficiently use the computation resources available at the edge server, we introduce a new metric called “saliency visual quality” (SVQ) and formulate patch selection as an optimization problem to achieve the best performance when an edge server is serving multiple users. We implement EVASR based on the FFmpeg framework and deploy it on three different platforms including desktop/laptop computers, mobile phones, and single board computers (SBCs). We conduct extensive experiments for evaluating the visual quality, super-resolution speed, and power savings that can be achieved by EVASR. Results show that EVASR outperforms baseline approaches in both resource efficiency and visual quality metrics including PSNR, SVQ, and VMAF. EVASR can also achieve substantial energy savings compared to baseline approaches MobileSR and JetsonSR on mobile devices.
Na Li 0032, Sheng Wei 0001, Yao Liu 0001
ACM Trans. Multim. Comput. Commun. Appl.4
2024 patchDPCC: A Patchwise Deep Compression Framework for Dynamic Point Clouds
abstract
When compressing point clouds, point-based deep learning models operate points in a continuous space, which has a chance to minimize the geometric fidelity loss introduced by voxelization in preprocessing. But these methods could hardly scale to inputs with arbitrary points. Furthermore, the point cloud frames are individually compressed, failing the conventional wisdom of leveraging inter-frame similarity. In this work, we propose a patchwise compression framework called patchDPCC, which consists of a patch group generation module and a point-based compression model. Algorithms are developed to generate patches from different frames representing the same object, and more importantly, these patches are regulated to have the same number of points. We also incorporate a feature transfer module in the compression model, which refines the feature quality by exploiting the inter-frame similarity. Our model generates point-wise features for entropy coding, which guarantees the reconstruction speed. The evaluation on the MPEG 8i dataset shows that our method improves the compression ratio by 47.01% and 85.22% when compared to PCGCv2 and V-PCC with the same reconstruction quality, which is 9% and 16% better than that D-DPCC does. Our method also achieves the fastest decoding speed among the learning-based compression models.
Zirui Pan, Mengbai Xiao, Xu Han 0016, Dongxiao Yu, Yao Liu 0001
AAAI6
2024 RoIRTC: Toward Region-of-Interest Reinforced Real-Time Video Communication
abstract
In this paper, we propose a region-of-interest (RoI) reinforced real-time communication system, RoIRTC, for improving the quality of videos delivered in real-time communication. RoIRTC uses a novel RoI magnification transformation for spatially adapting the camera-captured video frame. To automatically detect the RoI, it intelligently leverages a deep-learning-based saliency prediction model without affecting the video collector’s processing throughput or the encoder’s efficiency. Evaluation results based on actual remote learning videos show that RoIRTC that performs RoI magnification can improve the median PSNR by 2.6 dB compared to the naive WebRTC implementation. Compared to an approach that mimics the "background blur" scheme used in many real-time communication systems, RoIRTC can also improve the median PSNR by 4.2 dB.
Shuoqian Wang, Mengbai Xiao, Yao Liu 0001
ICME3
2024 Dynamic 6-DoF Volumetric Video Generation: Software Toolkit and Dataset
abstract
Volumetric video streaming has become increasingly popular in recent years due to its support of 6 degrees-of-freedom (6-DoF) exploration. There is, however, a shortage of dynamic 6-DoF content suitable for comparing the performance among heterogeneous volumetric video representations. This paper introduces a software toolkit for creating both a dataset of dynamic 6-DoF content in point clouds and a dataset for training and testing neural-based representations such as neural radiance fields (NeRF). Starting with freely available 3D assets online, our software toolkit uses the Blender Python API to generate training and testing datasets for neural-based dynamic volumetric model training. The created datasets are compliant with existing neural-based model training and rendering frameworks. The software can also construct point cloud sequences derived from synthetic dynamic 3D meshes. This further facilitates comparing point clouds and neural-based methods for volumetric video representation. We release the software toolkit along with a rich set of sequence datasets generated in compliance with the permissions granted by the original 3D asset creators. With our toolkit and dataset, we aim to facilitate research from the multimedia systems community to support practical volumetric streaming. Our software toolkit and dataset are available at: https://6-dof-dynamic-content-software.github.io/.
Mufeng Zhu, Yuan-Chun Sun, Na Li 0032, Jin Zhou 0006, Songqing Chen, Cheng-Hsin Hsu, Yao Liu 0001
MMSP7
2024 A GPU-Enabled Real-Time Framework for Compressing and Rendering Volumetric Videos
abstract
Nowadays, volumetric videos have emerged as an attractive multimedia application providing highly immersive watching experiences since viewers could adjust their viewports at 6 degrees-of-freedom. However, the point cloud frames composing the video are prohibitively large, and effective compression techniques should be developed. There are two classes of compression methods. One suggests exploiting the conventional video codecs (2D-based methods) and the other proposes to compress the points in 3D space directly (3D-based methods). Though the 3D-based methods feature fast coding speeds, their compression ratios are low since the failure of leveraging inter-frame redundancy. To resolve this problem, we design a patch-wise compression framework working in the 3D space. Specifically, we search rigid moves of patches via the iterative closest point algorithm and construct a common geometric structure, which is followed by color compensation. We implement our decoder on a GPU platform so that real-time decoding and rendering are realized. We compare our method with GROOT, the state-of-the-art 3D-based compression method, and it reduces the bitrate by up to 5.98$\times$. Moreover, by trimming invisible content, our scheme achieves comparable bandwidth demand of V-PCC, the representative 2D-based method, in FoV-adaptive streaming.
Dongxiao Yu, Ruopeng Chen, Xin Li 0078, Mengbai Xiao, Yao Liu 0001
IEEE Trans. Computers6
2024 VertexShuffle-Based Spherical Super-Resolution for 360-Degree Videos
abstract
360-degree video is an emerging form of media that encodes information about all directions surrounding a camera, offering an immersive experience to the users. Unlike traditional 2D videos, visual information in 360-degree videos can be naturally represented as pixels on a sphere. Inspired by state-of-the-art deep-learning-based 2D image super-resolution models and spherical CNNs, in this article, we design a novel spherical super-resolution (SSR) approach for 360-degree videos. To support viewport-adaptive and bandwidth-efficient transmission/streaming of 360-degree video data and save computation, we propose the Focused Icosahedral Mesh to represent a small area on the sphere. We further construct matrices to rotate spherical content over the entire sphere to the focused mesh area, allowing us to use the focused mesh to represent any area on the sphere. Motivated by the PixelShuffle operation for 2D super-resolution, we also propose a novel VertexShuffle operation on the mesh and an improved version VertexShuffle_V2. We compare our SSR approach with state-of-the-art 2D super-resolution models and show that SSR has the potential to achieve significant benefits when applied to spherical signals.
Na Li 0032, Yao Liu 0001
ACM Trans. Multim. Comput. Commun. Appl.2
2023 VQBA: Visual-Quality-Driven Bit Allocation for Low-Latency Point Cloud Streaming
abstract
Video-based Point Cloud Compression (V-PCC) is an emerging standard for encoding dynamic point cloud data. With V-PCC, point cloud data is segmented, projected, and packed on to 2D video frames, which can be compressed using existing video coding standards such as H.264, H.265 and AV1. This makes it possible to support point cloud streaming via reliable video transmission systems. On the other hand, despite recent advances, many issues still remain and prevent V-PCC from being used in low-latency point cloud streaming. For instance, point cloud registration and patch generation can take a long time.
Shuoqian Wang, Mufeng Zhu, Na Li 0032, Mengbai Xiao, Yao Liu 0001
ACM Multimedia5
2023 patchVVC: A Real-time Compression Framework for Streaming Volumetric Videos
abstract
Nowadays, volumetric video has emerged as an attractive multimedia application, which provides highly immersive watching experiences. However, streaming the volumetric video demands prohibitively high bandwidth. Thus, effectively compressing its underlying point cloud frames is essential to deploying the volumetric videos. The existing compression techniques are either 3D-based or 2D-based, but they still have drawbacks when being deployed in practice. The 2D-based methods compress the videos in an effective but slow manner, while the 3D-based methods feature high coding speeds but low compression ratios. In this paper, we propose patchVVC, a 3D-based compression framework that reaches both a high compression ratio and a real-time decoding speed. More importantly, patchVVC is designed based on point cloud patches, which makes it friendly to an field of view adaptive streaming system that further reduces the bandwidth demands. The evaluation shows patchVCC achieves the real-time decoding speed and the comparable compression ratios as the representative 2D-based scheme, V-PCC, in an FoV-adaptive streaming scenario.
Ruopeng Chen, Mengbai Xiao, Dongxiao Yu, Yao Liu 0001
MMSys5
2023 EVASR: Edge-Based Video Delivery with Salience-Aware Super-Resolution
abstract
With the rapid growth of video content consumption, it is important to deliver high-quality streaming videos to users even under limited available network bandwidth. In this paper, we propose EVASR, a system that performs edge-based video delivery to clients with salience-aware super-resolution. We select patches with higher saliency score to perform super-resolution while applying the simple yet efficient bicubic interpolation for the remaining patches in the same video frame. To efficiently use the computation resources available at the edge server, we introduce a new metric called "saliency visual quality" and formulate patch selection as an optimization problem to achieve the best performance when an edge server is serving multiple users. We implement EVASR based on the FFmpeg framework and conduct extensive experiments for evaluation. Results show that EVASR outperforms baseline approaches in both resource efficiency and visual quality metrics including PSNR, saliency visual quality (SVQ), and VMAF.
Na Li 0032, Yao Liu 0001
MMSys2
2023 Security-Preserving Live 3D Video Surveillance
abstract
3D video surveillance has become the new trend in security monitoring with the popularity of 3D depth cameras in the consumer market. While enabling more fruitful surveillance features, the finer-grained 3D videos being captured would raise new security concerns that have not been addressed by existing research. This paper explores the security implications of live 3D surveillance videos in triggering biometrics-related attacks, such as face ID spoofing. We demonstrate that the state-of-the-art face authentication systems can be effectively compromised by the 3D face models presented in the surveillance video. Then, to defend against such face spoofing attacks, we propose to proactively and benignly inject adversarial perturbations to the surveillance video in real time, prior to the exposure to potential adversaries. Such dynamically generated perturbations can prevent the face models from being exploited to bypass deep learning-based face authentications while maintaining the required quality and functionality of the 3D video surveillance. We evaluate the proposed perturbation generation approach on both an RGB-D dataset and a 3D video dataset, which justifies its effective security protection, low quality degradation, and real-time performance.
Zhongze Tang, Huy Phan, Xianglong Feng, Bo Yuan 0001, Yao Liu 0001, Sheng Wei 0001
MMSys5
2022 FFmpegSR: A General Framework Toward Real-Time 4K Super-Resolution
abstract
With the explosive growth of online video content, the demand for high-quality video is ever-rising. To take advantage of recent advances in deep learning, in this paper, we propose and implement a framework, FFmpegSR, that applies deep learning-based super-resolution into an FFmpeg filter to implement real-time 4K video super-resolution. FFmpegSR applies super-resolution to the Y channel only, allowing reduced inference time while maintaining good inference quality. To further improve the inference speed, we also develop a patch-based solution that uses saliency detection to select regions of interest on the video frame. This allows us to achieve faster inference on key patches only instead of full video frames. We used videos from a public dataset for evaluation. Results show that FFmpegSR can achieve real-time super-resolution to 4K with high visual quality.
Na Li 0032, Yao Liu 0001
ISM2
2022 Exploring Spherical Autoencoder for Spherical Video Content Processing
abstract
3D spherical content is increasingly presented in various applications (e.g., AR/MR/VR) for better users' immersiveness experience, yet today processing such spherical 3D content still mainly relies on the traditional 2D approaches after projection, leading to the distortion and/or loss of critical information. This study sets to explore methods to process spherical 3D content directly and more effectively. Using 360-degree videos as an example, we propose a novel approach called Spherical Autoencoder (SAE) for spherical video processing. Instead of projecting to a 2D space, SAE represents the 360-degree video content as a spherical object and employs encoding and decoding on the 360-degree video directly. Furthermore, to support the adoption of SAE on pervasive mobile devices that often have resource constraints, we further propose two optimizations on top of SAE.First, since the FoV (Field of View) prediction is widely studied and leveraged to transport only a portion of the content to the mobile device to save bandwidth and battery consumption, we design p-SAE, a SAE scheme with the partial view support that can utilize such FoV prediction. Second, since machine learning models are often compressed when running on mobile devices in order to reduce the processing load, which usually leads to degradation of output (e.g., video quality in SAE), we propose c-SAE by applying the compressive sensing theory into SAE to maintain the video quality when the model is compressed. Our extensive experiments show that directly incorporating and processing spherical signals is promising, and it outperforms the traditional approaches by a large margin. Both p-SAE and c-SAE show their effectiveness in delivering high quality videos (e.g., PSNR results) when used alone or combined together with model compression.
Jin Zhou 0006, Na Li 0032, Yao Liu 0001, Shuochao Yao, Songqing Chen
ACM Multimedia3
2022 Applying VertexShuffle toward 360-degree video super-resolution
abstract
With the recent successes of deep learning models, the performance of 2D image super-resolution has improved significantly. Inspired by recent state-of-the-art 2D super-resolution models and spherical CNNs, in this paper, we design a novel spherical superresolution (SSR) approach for 360-degree videos. To address the bandwidth waste problem associated with 360-degree video transmission/streaming and save computation, we propose the Focused Icosahedral Mesh to represent a small area on the sphere and construct matrices to rotate spherical content to the focused mesh area. We also propose a novel VertexShuffle operation on the mesh, motivated by the 2D PixelShuffle operation. We compare our SSR approach with state-of-the-art 2D super-resolution models. We show that SSR has the potential to achieve significant benefits when applied to spherical signals.
Na Li 0032, Yao Liu 0001
NOSSDAV2
2021 Learning to Guide Human Attention on Mobile Telepresence Robots with 360° Vision
abstract
Mobile telepresence robots (MTRs) allow people to navigate and interact with a remote environment that is in a place other than the person’s true location. Thanks to the recent advances in 360° vision, many MTRs are now equipped with an all-degree visual perception capability. However, people’s visual field horizontally spans only about 120° of the visual field captured by the robot. To bridge this observability gap toward human-MTR shared autonomy, we have developed a framework, called GHAL360, to enable the MTR to learn a goal-oriented policy from reinforcements for guiding human attention using visual indicators. Three telepresence environments were constructed using datasets that are extracted from Matterport3D and collected from a real robot respectively. Experimental results show that GHAL360 outperformed the baselines from the literature in the efficiency of a human-MTR team completing target search tasks. A demo video is available: https://youtu.be/aGbTxCGJSDM
Kishan Chandan, Jack Albertson, Xiaohan Zhang 0002, Yao Liu 0001, Shiqi Zhang 0001
IROS5
2020 SphericRTC: A System for Content-Adaptive Real-Time 360-Degree Video Communication
abstract
We present the SphericRTC system for real-time 360-degree video communication. 360-degree video allows the viewer to observe the environment in any direction from the camera location. This more-immersive streaming experience allows users to more-efficiently exchange information and can be beneficial in the real-time setting. Our system applies a novel approach to select representations of 360-degree frames to allow efficient, content-adaptive delivery. The system performs joint content and bitrate adaptation in real-time by offloading expensive transformation operations to the GPU via CUDA. The system demonstrates that the multiple sub-components -- viewport feedback, representation selection, and joint content and bitrate adaptation -- can be effectively integrated within a single framework. Compared to a baseline implementation, views in SphericRTC have consistently higher visual quality. The median Viewport-PSNR of such views is 2.25 dB higher than views in the baseline system.
Shuoqian Wang, Mengbai Xiao, Kenneth Chiu, Yao Liu 0001
ACM Multimedia5
2020 AdaP-360: User-Adaptive Area-of-Focus Projections for Bandwidth-Efficient 360-Degree Video Streaming
abstract
360-degree video is an emerging medium that presents an immersive view of the environment to the user. Despite its potential to provide an immersive watching experience, 360-degree video has not achieved widespread popularity. A significant cause of this slow adoption is the high-bandwidth requirements of the format. The primary source of bandwidth inefficiency in 360-degree video streaming, un-addressed in popular transmission methods, is the discrepancy between the pixels sent over the network (typically the full omnidirectional view) and the pixels displayed in the head-mounted display's field of view. At worst, roughly 88% of transmitted pixels remain unviewed.
Chao Zhou 0004, Shuoqian Wang, Mengbai Xiao, Sheng Wei 0001, Yao Liu 0001
ACM Multimedia5
2020 QuRate: power-efficient mobile immersive video streaming
abstract
Smartphones have recently become a popular platform for deploying the computation-intensive virtual reality (VR) applications, such as immersive video streaming (a.k.a., 360-degree video streaming). One specific challenge involving the smartphone-based head mounted display (HMD) is to reduce the potentially huge power consumption caused by the immersive video. To address this challenge, we first conduct an empirical power measurement study on a typical smartphone immersive streaming system, which identifies the major power consumption sources. Then, we develop QuRate, a quality-aware and user-centric frame rate adaptation mechanism to tackle the power consumption issue in immersive video streaming. QuRate optimizes the immersive video power consumption by modeling the correlation between the perceivable video quality and the user behavior. Specifically, QuRate builds on top of the user's reduced level of concentration on the video frames during view switching and dynamically adjusts the frame rate without impacting the perceivable video quality. We evaluate QuRate with a comprehensive set of experiments involving 5 smartphones, 21 users, and 6 immersive videos using empirical user head movement traces. Our experimental results demonstrate that QuRate is capable of extending the smartphone battery life by up to 1.24X while maintaining the perceivable video quality during immersive video streaming. Also, we conduct an Institutional Review Board (IRB)-approved subjective user study to further validate the minimum video quality impact caused by QuRate.
Nan Jiang 0020, Yao Liu 0001, Tian Guo 0001, Wenyao Xu, Viswanathan (Vishy) Swaminathan, Lisong Xu, Sheng Wei 0001
MMSys2
2020 LiveDeep: Online Viewport Prediction for Live Virtual Reality Streaming Using Lifelong Deep Learning
abstract
Live virtual reality (VR) streaming has become a popular and trending video application in the consumer market providing users with 360-degree, immersive viewing experiences. To provide premium quality of experience, VR streaming faces unique challenges due to the significantly increased bandwidth consumption. To address the bandwidth challenge, VR video viewport prediction has been proposed as a viable solution, which predicts and streams only the user’s viewport of interest with high quality to the VR device. However, most of the existing viewport prediction approaches target only the video-on-demand (VOD) use cases, requiring offline processing of the historical video and/or user data that are not available in the live streaming scenario. In this work, we develop a novel viewport prediction approach for live VR streaming, which only requires video content and user data in the current viewing session. To address the challenges of insufficient training data and real-time processing, we propose a live VR-specific deep learning mechanism, namely LiveDeep, to create the online viewport prediction model and conduct real-time inference. LiveDeep employs a hybrid approach to address the unique challenges in live VR streaming, involving (1) an alternate online data collection, labeling, training, and inference schedule with controlled feedback loop to accommodate for the sparse training data; and (2) a mixture of hybrid neural network models to accommodate for the inaccuracy caused by a single model. We evaluate LiveDeep using 48 users and 14 VR videos of various types obtained from a public VR user head movement dataset. The results indicate around 90% prediction accuracy, around 40% bandwidth savings, and premium processing time, which meets the bandwidth and real-time requirements of live VR streaming.
Xianglong Feng, Yao Liu 0001, Sheng Wei 0001
VR2
2020 Understanding the Ecosystem and Addressing the Fundamental Concerns of Commercial MVNO
abstract
Recent years have witnessed the rapid growth of mobile virtual network operators (MVNOs), which operate on top of existing cellular infrastructures of base carriers, while offering cheaper or more flexible data plans compared to those of the base carriers. In this paper, we present a two-year measurement study towards understanding various fundamental aspects of today's MVNO ecosystem, including its architecture, customers, performance, economics, and the complex interplay with the base carrier. Our study focuses on a large commercial MVNO with one million customers, operating atop a nation-wide base carrier. Our measurements clarify several key concerns raised by MVNO customers, such as inaccurate billing and potential performance discrimination with the base carrier. We also leverage big data analytics, statistical modeling, and machine learning to address the MVNO's key concerns with regard to data usage prediction, data plan reselling, customer churn mitigation, and billing delay reduction. Our proposed techniques can help achieve higher revenues and improved services for commercial MVNOs.
Yang Li 0092, Jianwei Zheng 0003, Zhenhua Li 0001, Yunhao Liu 0001, Feng Qian 0001, Sen Bai, Yao Liu 0001, Xianlong Xin
IEEE/ACM Trans. Netw.7
2019 The Cask Effect of Multi-source Content Delivery: Measurement and Mitigation
abstract
With the explosive growth of Internet traffic, multi-source content delivery has been introduced for improving the performance and quality-of-experience (QoE) of Internet services. Upgrading from single-source content delivery to multi-source content delivery, however, may not always lead to a better performance. Instead, a decline in terms of delivery speed often occurs. By conducting a comprehensive study, we show that the underlying reason of this counter-intuitive phenomenon is actually due to the cask effect of data sources at both macro and micro level. Specifically, at the macro level, data sources with different types are highly heterogeneous in terms of delivery performance, which means data sources with certain types are particularly easy to become the "short boards". At the micro level, for the data sources chosen by a client, the high diversity of participation time (DPT) of the sources could impair the acceleration effect. Motivated by the above findings, we design MDR (Multi-source Delivery Redirector), a middleware that contains two optimizations to improve the acceleration effect. One is the feature-greedy selection algorithm which can avoid selecting data sources with inferior types, and the other is the DPT-driven shuffle strategy which can avoid using unstable data sources. Simulation-based experiments show that the MDR outperforms existing approaches in terms of overall downloading performance.
Minghao Zhao 0001, Xinlei Yang, Zhenhua Li 0001, Yao Liu 0001, Zhenyu Li 0001, Yunhao Liu 0001
ICDCS5
2019 Companion Paper for
abstract
This artifact includes source code, scripts and datasets required to reproduce the experimental figures in the evaluation of the MM'18 paper, which is entitled "MiniView Layout for Bandwidth-Efficient 360-Degree Video". The artifact reports the comparison results among the standard cube layout (CUBE), the equi-angular layout (EAC), and the MiniView layout (MVL) in terms of compressed video size, visual quality of views and decoding and rendering time.
Mengbai Xiao, Shuoqian Wang, Chao Zhou 0004, Li Liu 0045, Zhenhua Li 0001, Yao Liu 0001, Songqing Chen, Lucile Sassatelli, Gwendal Simon
ACM Multimedia6
2019 An In-depth Study of Commercial MVNO: Measurement and Optimization
abstract
Recent years have witnessed the rapid growth of mobile virtual network operators (MVNOs), which operate on top of the existing cellular infrastructures of base carriers while offering cheaper or more flexible data plans compared to those of the base carriers. In this paper, we present a nearly two-year measurement study towards understanding various key aspects of today's MVNO ecosystem, including its architecture, performance, economics, customers, and the complex interplay with the base carrier. Our study focuses on a large commercial MVNO with \reviseabout 1 million customers, operating atop a nation-wide base carrier. Our measurements clarify several key concerns raised by MVNO customers, such as inaccurate billing and potential performance discrimination with the base carrier. We also leverage big data analytics and machine learning to optimize an MVNO's key businesses such as data plan reselling and customer churn mitigation. Our proposed techniques can help achieve %will lead to higher revenues and improved services for commercial MVNOs.
Ao Xiao, Yunhao Liu 0001, Yang Li 0092, Feng Qian 0001, Zhenhua Li 0001, Sen Bai, Yao Liu 0001, Tianyin Xu, Xianlong Xin
MobiSys7
2018 Towards Web-based Delta Synchronization for Cloud Storage Services
Zhenhua Li 0001, Ennan Zhai, Tianyin Xu, Yang Li 0092, Yunhao Liu 0001, Quanlu Zhang, Yao Liu 0001
FAST8
2018 BAS-360°: Exploring Spatial and Temporal Adaptability in 360-degree Videos over HTTP/2
abstract
Today, 360-degree video streaming has become a popular Internet service with the rise of affordable virtual reality (VR) technologies. However, streaming 360-degree videos suffers from the prohibitive bandwidth demand. Existing bandwidth-efficient solutions mainly focus on exploiting the inherent spatial adaptability of 360-degree videos, delivering only video content (spatially-cut tiles) in the viewer's region of interest (ROI) with higher quality. Temporal adaptability, which has been widely leveraged in HTTP streaming, has not been well exploited to select proper quality for video segments according to the bandwidth variations. When these two dimensions of adaptability are jointly considered, bitrate selection for the tiles become more complicated and challenging. The importance of a tile with a spatial coordination played at a specific time should be quantified so that we can determine how to allocate bandwidth for improving the viewer's quality of experience. Furthermore, viewer's head orientation prediction is highly variable, which makes the determination of important tiles highly dynamic. In addition, network fluctuations are very common on the Internet. To overcome these challenges, we propose Bi-Adaptive Streaming for 360-degree videos (BAS-360°). In BAS-360°, both spatial and temporal adaptabilities are explored in the bitrate selection for different tiles. The objective is to minimize the bandwidth waste by allocating bandwidth to more important tiles (the tiles that are more likely to be watched). To tackle the high variability of visual region prediction and the unpredictable network fluctuations, we employ two features provided by HTT P /2: stream termination and stream priority, to efficiently organize tile delivery. Evaluation results show that BAS-360° outperforms naive tile-based 360-degree video streaming strategies when network fluctuations or errors in viewport predictions occur.
Mengbai Xiao, Chao Zhou 0004, Viswanathan (Vishy) Swaminathan, Yao Liu 0001, Songqing Chen
INFOCOM4
2018 ClusTile: Toward Minimizing Bandwidth in 360-degree Video Streaming
abstract
360-degree video has the potential to transform the video streaming experience by providing a more-immersive environment for users to interact with than standard streaming video. This experience is hampered, however, by high bandwidth requirements resulting from the extra information associated with the 360-degree frames. Because users cannot see this full 360-degree view, but the full view is transmitted in the majority of 360-degree streaming systems, there is much potential to reduce wasted bandwidth in this domain. We propose ClusTile, a tiling approach formulated to select a set of tiles that allows minimal bandwidth needed to be used when streaming 360-degree video over an expected set of views. These tiles are selected by solving a set of integer linear programs (ILPs) independently on clusters of collected user views. The clustering approach reduces computation requirements of the ILPs to practical levels. Tilings computed from ClusTile can save up to 76% bandwidth compared to standard 360-degree streaming and up to 52% bandwidth compared to best-performing fixed tiling schemes.
Chao Zhou 0004, Mengbai Xiao, Yao Liu 0001
INFOCOM3
2018 Minimizing the Cask Effect of Multi-Source Content Delivery
abstract
This paper reveals the performance anomaly (i.e., the decline of delivery speed) when the client upgrades a task from single-source content delivery to multi-source content delivery. This anomaly is mainly caused by two aspects: (1) data sources with different types vary greatly in terms of acceleration reward (AR), and data sources with certain types are particularly easy to become inferior; (2) When the data sources remain fixed for a period of time, the large diversity of participant time (DPT) of data sources disturb the acceleration and the data sources with less participant time are inferior. Combing these insights, we figure out that the multi-source content delivery is limited by the so-called cask effect, i.e., the acceleration effect mainly depends on the inferior data sources.
Zhenhua Li 0001, Zhenyu Li 0001, Tianyin Xu, Ennan Zhai, Yao Liu 0001, Minghao Zhao 0001, Yunhao Liu 0001
IWQoS6
2018 MiniView Layout for Bandwidth-Efficient 360-Degree Video
abstract
With the recent increase in popularity of VR devices, 360-degree video has become increasingly popular. As more users experience this new medium, it will likely see further increases in popularity as users experience its greater immersiveness compared to traditional video streams. 360-degree video streams must encode the omnidirectional view, and, with current encoding techniques, these views require significantly higher bandwidth than traditional video streams. These larger bandwidth requirements comprise the main barrier toward wider adoption by video streaming services.
Mengbai Xiao, Shuoqian Wang, Chao Zhou 0004, Li Liu 0045, Zhenhua Li 0001, Yao Liu 0001, Songqing Chen
ACM Multimedia6
2018 On the Effectiveness of Offset Projections for 360-Degree Video Streaming
abstract
A new generation of video streaming technology, 360-degree video, promises greater immersiveness than standard video streams. This level of immersiveness is similar to that produced by virtual reality devices—users can control the field of view using head movements rather than needing to manipulate external devices. Although 360-degree video could revolutionize the streaming experience, its large-scale adoption is hindered by a number of factors: 360-degree video streams have larger bandwidth requirements and require faster responsiveness to user inputs, and users may be more sensitive to lower quality streams. In this article, we review standard approaches toward 360-degree video encoding and compare these to families of approaches that distort the spherical surface to allow oriented concentrations of the 360-degree view. We refer to these distorted projections as offset projections. Our measurement studies show that most types of offset projections produce rendered views with better quality than their nonoffset equivalents when view orientations are within 40 or 50 degrees of the offset orientation. Offset projections complicate adaptive 360-degree video streaming because they require a combination of bitrate and view orientation adaptations. We estimate that this combination of streaming adaptation in two dimensions can cause over 57% extra segments to be downloaded compared to an ideal downloading strategy, wasting 20% of the total downloading bandwidth.
Chao Zhou 0004, Zhenhua Li 0001, Joe Osgood, Yao Liu 0001
ACM Trans. Multim. Comput. Commun. Appl.4
2018 On the Synchronization Bottleneck of OpenStack Swift-Like Cloud Storage Systems
abstract
As one type of the most popular cloud storage services, OpenStack Swift and its follow-up systems replicate each object across multiple storage nodes and leverageobject sync protocolsto achieve high reliability andeventual consistency. The performance of object sync protocols heavily relies on two key parameters:$r$(number of replicas for each object) and$n$(number of objects hosted by each storage node). In existing tutorials and demos, the configurations are usually$r=3$and$n<1,000$by default, and the sync process seems to perform well. However, we discover in data-intensive scenarios, e.g., when$r>3$and$n\gg 1,000$, the sync process is significantly delayed and produces massive network overhead, referred to as thesync bottleneck problem. By reviewing the source code of OpenStack Swift, we find that its object sync protocol utilizes a fairly simple and network-intensive approach to check the consistency among replicas of objects. Hence in a sync round, the number of exchanged hash values per node is$\Theta (n\times r)$. To tackle the problem, we propose a lightweight and practical object sync protocol,LightSync, which not only remarkably reduces the sync overhead, but also preserves high reliability and eventual consistency. LightSync derives this capability from three novel building blocks: 1)Hashing of Hashes, which aggregates all the$h$hash values of each data partition into a single but representative hash value with the Merkle tree; 2)Circular Hash Checking, which checks the consistency of different partition replicas by only sending the aggregated hash value to the clockwise neighbor; and 3)Failed Neighbor Handling, which properly detects and handles node failures with moderate overhead to effectively strengthen the robustness of LightSync. The design of LightSync offers provable guarantee on reducing the per-node network overhead from$\Theta (n\times r)$to$\Theta (\frac{n}{h})$. Furthermore, we have implemented LightSync as an open-source patch and adopted it to OpenStack Swift, thus reducing the sync delay by up to 879$\times$and the network overhead by up to 47.5$\times$.
Mingkang Ruan, Thierry Titcheu Chekam, Ennan Zhai, Zhenhua Li 0001, Yao Liu 0001, Jinlong E, Yong Cui 0001, Hong Xu 0001
IEEE Trans. Parallel Distributed Syst.5
2017 Buffer-Based Reinforcement Learning for Adaptive Streaming
abstract
Adaptive streaming improves user-perceived quality by altering the streaming bitrate depending on network conditions, trading reduced video bitrates for reduced stall times. Existing adaptation approaches, e.g., rate-based, buffer-based, either rely heavily on accurate bandwidth prediction or can be overly-conservative about video bitrates. In this work, we propose a reinforcement learning approach to choose the segment quality during playback. This approach uses only the buffer state information and optimizes for a measure of user-perceived streaming quality. Simulation results show that our proposed approach achieves better QoE than rate-, buffer-based approaches, as well as other reinforcement learning approaches.
Yao Liu 0001
ICDCS2
2017 OpTile: Toward Optimal Tiling in 360-degree Video Streaming
abstract
360-degree videos are encoded for adaptive streaming by first projecting the spherical surface onto two-dimensional frames, then encoding these as standard video segments. During playback of these 360-degree videos, the video player renders the portion of the spherical surface in the direction of the user's view. These user viewports typically cover only a small portion of the 360 degree surface, causing much of the downloaded bandwidth to be wasted. Tile-based approaches can reduce the wasted bandwidth by cutting video spatially into motion-constrained rectangles. Streaming logic then only needs to download the tiles necessary to render the viewport seen by the user. Existing tile-based approaches cut 360-degree videos into tiles of fixed sizes. These fixed-size tiling approaches, however, suffer from reduced encoding efficiency. Tiling cuts away portions of the video that can be copied by the encoder from adjacent frames or within the current frame that are needed for effective video compression.
Mengbai Xiao, Chao Zhou 0004, Yao Liu 0001, Songqing Chen
ACM Multimedia3
2017 A Measurement Study of Oculus 360 Degree Video Streaming
abstract
360 degree video is anew generation of video streaming technology that promises greater immersiveness than standard video streams. This level of immersiveness is similar to that produced by virtual reality devices -- users can control the field of view using head movements rather than needing to manipulate external devices. Although 360 degree video could revolutionize streaming technology, large scale adoption is hindered by a number of factors. 360 degree video streams have larger bandwidth requirements, require faster responsiveness to user inputs, and users may be more sensitive to lower quality streams.; [email protected] this paper, we review standard approaches toward 360 degree video encoding and compare these to a new, as yet unpublished, approach by Oculus which we refer to as the offset cubic projection. Compared to the standard cubic encoding, the offset cube encodes a distorted version of the spherical surface, devoting more information (i.e., pixels) to the view in a chosen direction. We estimate that the offset cube representation can produce better or similar visual quality while using less than 50% pixels under reasonable assumptions about user behavior, resulting in 5.6% to 16.4% average savings in video bitrate. During 360 degree video streaming, Oculus uses a combination of quality level adaptation and view orientation adaptation. We estimate that this combination of streaming adaptation in two dimensions can cause over 57% extra segments to be downloaded compared to an ideal downloading strategy, wasting 20% of the total downloading bandwidth.
Chao Zhou 0004, Zhenhua Li 0001, Yao Liu 0001
MMSys3
2016 GoCAD: GPU-Assisted Online Content-Adaptive Display Power Saving for Mobile Devices in Internet Streaming
abstract
During Internet streaming, a significant portion of the battery power is always consumed by the display panel on mobile devices. To reduce the display power consumption, backlight scaling, a scheme that intelligently dims the backlight has been proposed. To maintain perceived video appearance in backlight scaling, a computationally intensive luminance compensation process is required. However, this step, if performed by the CPU as existing schemes suggest, could easily offset the power savings gained from backlight scaling. Furthermore, computing the optimal backlight scaling values requires per-frame luminance information, which is typically too energy intensive for mobile devices to compute. Thus, existing schemes require such information to be available in advance. And such an offline approach makes these schemes impractical. To address these challenges, in this paper, we design and implement GoCAD, a GPU-assisted Online Content-Adaptive Display power saving scheme for mobile devices in Internet streaming sessions. In GoCAD, we employ the mobile device's GPU rather than the CPU to reduce power consumption during the luminance compensation phase. Furthermore, we compute the optimal backlight scaling values for small batches of video frames in an online fashion using a dynamic programming algorithm. Lastly, we make novel use of the widely available video storyboard, a pre-computed set of thumbnails associated with a video, to intelligently decide whether or not to apply our backlight scaling scheme for a given video. For example, when the GPU power consumption would offset the savings from dimming the backlight, no backlight scaling is conducted. To evaluate the performance of GoCAD, we implement a prototype within an Android application and use a Monsoon power monitor to measure the real power consumption. Experiments are conducted on more than 460 randomly selected YouTube videos. Results show that GoCAD can effectively produce power savings without affecting rendered video quality.
Yao Liu 0001, Mengbai Xiao, Xin Li 0078, Mian Dong, Zhenhua Li 0001, Songqing Chen
WWW1
2016 Content-Adaptive Display Power Saving for Internet Video Applications on Mobile Devices
abstract
Backlight scaling is a technique proposed to reduce the display panel power consumption by strategically dimming the backlight. However, for mobile video applications, a computationally intensive luminance compensation step must be performed in combination with backlight scaling to maintain the perceived appearance of video frames. This step, if done by the Central Processing Unit (CPU), could easily offset the power savings via backlight dimming. Furthermore, computing the backlight scaling values requires per-frame luminance information, which is typically too energy intensive to compute on mobile devices. In this article, we propose Content-Adaptive Display (CAD) for two typical Internet mobile video applications: video streaming and real-time video communication. CAD uses the mobile device’s Graphics Processing Unit (GPU) rather than the CPU to perform luminance compensation at reduced power consumption. For video streaming where video frames are available in advance, we compute the backlight scaling schedule using a more efficient dynamic programming algorithm than existing work. For real-time video communication where video frames are generated on the fly, we propose a greedy algorithm to determine the backlight scaling at runtime. We implement CAD in one video streaming application and one real-time video call application on the Android platform and use a Monsoon power meter to measure the real power consumption. Experiment results show that CAD can save more than 10% overall power consumption for up to 55.7% videos during video streaming and up to 31.0% overall power consumption in real-time video calls.
Yao Liu 0001, Mengbai Xiao, Xin Li 0078, Mian Dong, Zhenhua Li 0001, Lei Guo 0004, Songqing Chen
ACM Trans. Multim. Comput. Commun. Appl.1
2015 Do Twin Clouds Make Smoothness for Transoceanic Video Telephony?
abstract
Transoceanic video telephony (TVT) over the Internet is challenging due to 1) longer round-trip delay, 2) larger number of relay hops, and 3) higher packet loss rate. Real-world measurements of Skype, Face time, and QQ confirm that their TVT service quality is mostly unsatisfactory. Recently, when using We Chat to make transoceanic video calls, we are fortunate to find that it achieves stably smooth TVT. To explore how this is possible, we conduct in-depth measurements of We Chat data flow. In particular, we discover that the service provider of We Chat deploys a novel, specially designed "twin clouds" based architecture to deliver transoceanic (UDP) packets. Thus, data delivery between two callers is no longer point-to-point (used by Skype, Face time, and QQ) over the best-effort Internet. Instead, transoceanic video packets are delivered through the privileged backbone formed by twin clouds, which greatly reduces the round-trip delay, number of relay hops, and packet loss rate. Besides, whenever a packet is found lost, multiple duplicate packets are instantly sent to aggressively make up for the loss. On the other hand, we notice two-fold shortcomings of twin clouds. First, due to the sophisticated resource provisioning inside the twin clouds, the video start up time is considerably extended. Second, due to the high cost of deploying twin clouds, the capacity of the privileged backbone is limited and sometimes in shortage, and thus We Chat has to deliver data via a detour path with degraded performance. Ultimately, we believe that the twin clouds based data delivery solution will arouse a new direction of Internet video telephony research while still deserves optimization efforts.
Zhenhua Li 0001, Yao Liu 0001, Zhi-Li Zhang
ICPP3
2015 Offline Downloading in China: A Comparative Study
abstract
Although Internet access has become more ubiquitous in recent years, most users in China still suffer from low-quality connections, especially when downloading large files. To address this issue, hundreds of millions of China's users have resorted to technologies that allow for ``offline downloading'', where a proxy is employed to pre-download the user's requested file and then deliver the file at her convenience.
Zhenhua Li 0001, Christo Wilson, Tianyin Xu, Yao Liu 0001, Yinlong Wang
Internet Measurement Conference4
2015 Reducing display power consumption for real-time video calls on mobile devices
abstract
The display subsystem of a mobile device usually consumes 38%-68% [1] of the total battery power in video streaming. Therefore, a few schemes have been designed to reduce the display power consumption. The basic idea is to dim the backlight level while properly compensating the pixel luminance to maintain image fidelity. The luminance compensation and proper backlight level calculation are computation intensive and demand per-frame luminance information. For these reasons, existing schemes only work for video-on-demand where each frame (and thus the luminance information) is available in advance. In addition, they demand additional computing resource support. Otherwise, if the computation is conducted on the mobile device, the power consumption due to such computation can easily offset the power savings from dimming the backlight. In this work, we set to investigate power saving for real-time video calls on mobile devices. Different from video-on-demand, real-time video calls are highly delay sensitive and the frame luminance information is not known in advance. Moreover, video calls often involve multiple streaming sources from multiple (≥2) participants, making it more difficult. Because there are few background changes and the frame rate is usually small in video calls, we design a Greedy Display Power saving scheme, called LCD-GDP, which utilizes the commonly available GPU on mobile devices without demanding additional support. Our design is implemented on WebRTC, a popular real-time web browser based video call standard. Experiments show that our scheme can save up to 33% power consumption in video calls without affecting the video call quality.
Mengbai Xiao, Yao Liu 0001, Lei Guo 0004, Songqing Chen
ISLPED2
2015 Content-adaptive display power saving in internet mobile streaming
abstract
Backlight scaling is a technique proposed to reduce the display panel power consumption by strategically dimming the backlight. However, for Internet streaming to mobile devices, a computationally intensive luminance compensation step must be performed in combination with backlight scaling to maintain the perceived appearance of video frames. This step, if done by the CPU, could easily offset the power savings via backlight dimming. Furthermore, computing the backlight scaling values requires per-frame luminance information, which is typically too energy intensive to compute on mobile devices.
Yao Liu 0001, Mengbai Xiao, Xin Li 0078, Mian Dong, Zhenhua Li 0001, Songqing Chen
NOSSDAV1
2015 A Quantitative Study of Video Duplicate Levels in YouTube
Yao Liu 0001, Sam Blasiak, Weijun Xiao, Zhenhua Li 0001, Songqing Chen
PAM1
2014 Towards Network-level Efficiency for Cloud Storage Services
abstract
Cloud storage services such as Dropbox, Google Drive, and Microsoft OneDrive provide users with a convenient and reliable way to store and share data from anywhere, on any device, and at any time. The cornerstone of these services is the data synchronization (sync) operation which automatically maps the changes in users' local filesystems to the cloud via a series of network communications in a timely manner. If not designed properly, however, the tremendous amount of data sync traffic can potentially cause (financial) pains to both service providers and users.
Zhenhua Li 0001, Cheng Jin 0008, Tianyin Xu, Christo Wilson, Yao Liu 0001, Linsong Cheng, Yunhao Liu 0001, Yafei Dai, Zhi-Li Zhang
Internet Measurement Conference5
2014 An Empirical Study of Video Messaging Services on Smartphones
abstract
With the advancements in wireless networks and the pervasive adoptions of smartphones with camera capabilities, hundreds of millions mobile users are attracted to video messaging services on their smartphones. Unlike text messaging, transmitting large video messages (with a size of ~ 10 MBytes vs. a 140 Bytes limit for text messages) demands effective network resource provisioning.
Yao Liu 0001, Lei Guo 0004
NOSSDAV1
2014 Investigating Redundant Internet Video Streaming Traffic on iOS Devices: Causes and Solutions
abstract
The Internet has witnessed rapidly increasing streaming traffic to various mobile devices. In this paper, through analysis of a server-side workload and experiments in a controlled lab environment, we find that current practice has introduced a significant amount of redundant traffic. In particular, for the popular iOS based mobile devices, accessing popular Internet streaming services typically involves about 10%-70% redundant traffic. Such a practice not only over-utilizes and wastes resources on the server side and the network (cellular or Internet), but also consumes additional battery power on user's mobile devices and leads to possible monetary cost. To alleviate such a situation without changing the server side or the client side, we design and implement CStreamer that can transparently work between existing mobile clients and servers. We have implemented a prototype and installed on Amazon EC2. Experiments conducted based on this prototype show that CStreamer can completely eliminate the redundant traffic without degrading user's QoS.
Yao Liu 0001, Qi Wei 0003, Lei Guo 0004, Bo Shen 0003, Songqing Chen, Yingjie Lan
IEEE Trans. Multim.1
2013 Effectively minimizing redundant Internet streaming traffic to iOS devices
abstract
The Internet has witnessed rapidly increasing streaming traffic to various mobile devices. In this paper, we find that for the popular iOS based mobile devices, accessing popular Internet streaming services typically involves about 10% - 70% unnecessary redundant traffic. Such a practice not only overutilizes and wastes resources on the server side and the network (cellular or Internet), but also consumes additional battery power on users' mobile devices and leads to possible monetary cost. To alleviate such a situation without changing the server side or the iOS, we design and implement a CStreamer prototype that can transparently work between existing iOS devices and media servers. We also build a CStreamer iOS App to enable end users to access Internet streaming services via CStreamer. Experiments conducted based on this prototype running on Amazon EC2 show that CStreamer can completely eliminate the redundant traffic without degrading user's QoS.
Yao Liu 0001, Fei Li 0001, Lei Guo 0004, Bo Shen 0003, Songqing Chen
INFOCOM1
2013 Efficient Batched Synchronization in Dropbox-Like Cloud Storage Services
Zhenhua Li 0001, Christo Wilson, Zhefu Jiang, Yao Liu 0001, Ben Y. Zhao, Cheng Jin 0008, Zhi-Li Zhang, Yafei Dai
Middleware4
2013 A Comparative Study of Android and iOS for Accessing Internet Streaming Services
Yao Liu 0001, Fei Li 0001, Lei Guo 0004, Bo Shen 0003, Songqing Chen
PAM1
2013 Measurement and Analysis of an Internet Streaming Service to Mobile Devices
abstract
Receiving Internet streaming services on various mobile devices is getting increasingly popular, and cloud platforms have also been gradually employed for delivering streaming services to mobile devices. While a number of studies have been conducted at the client side to understand and characterize Internet mobile streaming delivery, little is known about the server side, particularly for the recent cloud-based Internet mobile streaming delivery. In this work, we aim to investigate the Internet mobile streaming service at the server side. For this purpose, we have collected a 4-month server-side log on the cloud (with 1,002 TB delivered video traffic) from a top Internet mobile streaming service provider serving worldwide mobile users. Through trace analysis, we find that 1) a major challenge for providing Internet mobile streaming services is rooted from the mobile device hardware and software heterogeneity. In this workload, we find over 3,400 different hardware models with more than 100 different screen resolutions running 14 different mobile OS and three audio codecs and four video codecs. 2) To deal with the device heterogeneity, CPU-intensive transcoding is used on the cloud to customize the video to the appropriate versions at runtime for different devices. A video clip could be transcoded into more than 40 different versions to serve requests from different devices. 3) Compared to videos in traditional Internet streaming, mobile streaming videos are typically of much smaller size (a median of 1.68 MBytes) and shorter duration (a median of 2.7 minutes). Furthermore, the daily mobile user accesses are more skewed following a Zipf-like distribution but users' interests also quickly shift. Considering the huge demand of CPU cycles for online transcoding, we further examine server-side caching to reduce the total CPU cycle demand from the cloud. We show that a policy considering different versions of a video altogether outperforms other intuitive ones when the cache size is limited.
Yao Liu 0001, Fei Li 0001, Lei Guo 0004, Bo Shen 0003, Songqing Chen, Yingjie Lan
IEEE Trans. Parallel Distributed Syst.1
2012 A server's perspective of Internet streaming delivery to mobile devices
abstract
Receiving Internet streaming services on various mobile devices is getting more and more popular. To understand and better support Internet streaming delivery to mobile devices, a number of studies have been conducted. However, existing studies have mainly focused on the client side resource consumption and streaming quality. So far, little is known about the server side, which is the key for providing successful mobile streaming services. In this work, we set to investigate the Internet mobile streaming service at the server side. For this purpose, we have collected a one-month server log (with 212 TB delivered video traffic) from a top Internet mobile streaming service provider serving worldwide mobile users. Through trace analysis, we find that (1) a major challenge for providing Internet mobile streaming services is rooted from the mobile device hardware and software heterogeneity. In this workload, we find over 2800 different hardware models with about 100 different screen resolutions running 14 different mobile OS and 3 audio codecs and 4 video codecs. (2) To deal with the device heterogeneity, transcoding is used to customize the video to the appropriate versions at runtime for different devices. A video clip could be transcoded into more than 40 different versions in order to serve requests from different devices. (3) Compared to videos in traditional Internet streaming, mobile streaming videos are typically of much smaller size (a median of 1.68 MBytes) and shorter duration (a median of 2.7 minutes). Furthermore, the daily mobile user accesses are more skewed following a Zipf-like distribution but users' interests also quickly shift. Considering the huge demand of CPU cycles for online transcoding, we further examine server-side caching in order to reduce CPU cycle demand. We show that a policy considering different versions of a video altogether outperforms other intuitive ones when the cache size is limited.
Yao Liu 0001, Fei Li 0001, Lei Guo 0004, Bo Shen 0003, Songqing Chen
INFOCOM1
2011 Towards efficient resource utilization in internet mobile streaming
abstract
Internet video streaming to mobile devices is challenging because of device heterogeneity, resource constraints, and limited battery power supply of mobile devices. From mobile users' perspective, we examine the power efficiency of existing streaming protocols, and propose to reduce the power consumed by data transmission. Moreover, from the service provider's perspective, we propose to efficiently utilize resources at server side to better serve mobile users.
Yao Liu 0001
ACM Multimedia1
2011 An empirical evaluation of battery power consumption for streaming data transmission to mobile devices
abstract
Internet streaming applications are becoming increasingly popular on mobile devices. However, receiving streaming services on mobile devices is often constrained by their limited battery power supply. Various techniques have been proposed to save battery power consumption on mobile devices, mainly focusing on how much data to transmit and how to transmit.
Yao Liu 0001, Lei Guo 0004, Fei Li 0001, Songqing Chen
ACM Multimedia1
2011 BlueStreaming: towards power-efficient internet P2P streaming to mobile devices
abstract
P2P streaming applications are very popular on the Internet today. However, a mobile device in P2P streaming not only needs to continuously receive streaming data from other peers for its playback, but also needs to continuously exchange control information (e.g., buffermaps and file chunk requests) with neighboring peers and upload the downloaded streaming data to them. These lead to excessive battery power consumption on the mobile device.
Yao Liu 0001, Fei Li 0001, Lei Guo 0004, Yang Guo 0001, Songqing Chen
ACM Multimedia1
2011 A measurement study of resource utilization in internet mobile streaming
abstract
The pervasive usage of mobile devices and wireless networking support have enabled more and more Internet stream- ing services to all kinds of heterogeneous mobile devices. However, Internet mobile streaming services are challenged by the inherently limited on-device resources, device heterogeneity, and the bulk amount of streaming data.
Yao Liu 0001, Fei Li 0001, Lei Guo 0004, Songqing Chen
NOSSDAV1
2010 Reducing data request contentions for improved streaming quality
abstract
In P2P assisted multi-channel live streaming systems, it is commonly believed that in unpopular channels, quality degradation is due to the small number of participating peers with almost-the-same set of available data; this phenomena prevents effective data exchanges among peers themselves and automatically leads to data request contentions once a new data chunk becomes available. In popular programs, our measurement on PPLive for a continuous three-month period at various locations also shows numerous occurrences of quality degradation because of the even higher ratio (up to 190%) of repetitive data requests for the same data chunks.
Yao Liu 0001, Fei Li 0001, Lei Guo 0004, Songqing Chen
NOSSDAV1
2009 A Case Study of Traffic Locality in Internet P2P Live Streaming Systems
abstract
With the ever-increasing P2P Internet traffic, recently much attention has been paid to the topology mismatch between the P2P overlay and the underlying network due to the large amount of cross-ISP traffic. Mainly focusing on BitTorrent-like file sharing systems, several recent studies have demonstrated how to efficiently bridge the overlay and the underlying network by leveraging the existing infrastructure, such as CDN services or developing new application-ISP interfaces, such as P4P. However, so far the traffic locality in existing P2P live streaming systems has not been well studied. In this work, taking PPLive as an example, we examine traffic locality in Internet P2P streaming systems. Our measurement results on both popular and unpopular channels from various locations show that current PPLive traffic is highly localized at the ISP level. In particular, we find: (1) a PPLive peer mainly obtains peer lists referred by its connected neighbors (rather than tracker servers) and up to 90% of listed peers are from the same ISP as the requesting peer; (2) the major portion of the streaming traffic received by a requesting peer (up to 88% in popular channels) is served by peers in the same ISP as the requestor; (3) the top 10\% of the connected peers provide most (about 70%) of the requested streaming data and these top peers have smaller RTT to the requesting peer. Our study reveals that without using any topology information or demanding any infrastructure support, PPLive achieves such high ISP level traffic locality spontaneously with its decentralized, latency based, neighbor referral peer selection strategy. These findings provide some new insights for better understanding and optimizing the network- and user-level performance in practical P2P live streaming systems.
Yao Liu 0001, Lei Guo 0004, Fei Li 0001, Songqing Chen
ICDCS1