EDBT 2026 Demo / reviewers in the wild / expert
Zhisheng Yan
dblp:28/10126
· DBLP profile ↗
54ranked-venue papers
12as first author
26since 2021 · last 2026
0000-0002-6722-5517ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 33 · 8 first-author · 15 since 2021Computer networks · 16 · 2 first-author · 12 since 2021Artificial intelligence and machine learning · 4 · 3 since 2021Security and privacy · 3 · 2 since 2021Systems, architecture and hardware · 1Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DySy-Det: A Synergistic Framework with Dynamic Reconstruction-Path Consistency for AI-Generated Image DetectionabstractAdvanced image generative models have led to concerns about malicious use, underscoring the necessity for generalizable detection methods. However, existing approaches tend to overfit to domain-specific forgery patterns, while overlooking complementary cues from different domains. Therefore, we introduce DySy-Det (Dynamic Synergy Detector), a novel framework that mines collaborative and robust forgery artifacts from multiple evidence domains. First, DySy-Det fine-tunes a CLIP vision transformer to extract high-level semantics for identifying conceptual inconsistencies, while generating attention maps that pinpoint key discriminative regions. Then, this semantic guidance, in the form of a mask, directs a targeted reconstruction process. By focusing on these salient areas, our approach effectively extracts localized reconstruction errors, thereby filtering out irrelevant background noise. Furthermore, inspired by the intrinsic generative mechanics of diffusion models, we introduce the concept of Reconstruction-Path Consistency (RPC), which quantifies the temporal stability of the denoising trajectory to expose dynamic generative artifacts. We capture this by computing noise alignment scores across multiple timesteps and encode them via a lightweight network. Extensive evaluations on GenImage and UniversalFakeDetect benchmarks demonstrate that DySy-Det outperforms the state-of-the-art detector by 6.14% and 1.57% in mean accuracy, respectively. Fanli Jin, Feng Lin 0004, Gaojian Wang, Zhisheng Yan |
AAAI | 5 |
| 2026 | Layer-Wise Diagnostic Probing to Enhance Selectivity in Machine Unlearning
Anudeep Vurity, Zhisheng Yan, Massimiliano Albanese |
ICPR (14) | 2 |
| 2026 | Evaluating Machine Unlearning in Fingerphoto Presentation Attack Detection
Anudeep Vurity, Zhisheng Yan, Massimiliano Albanese |
ICPR (14) | 2 |
| 2025 | MSMA'2025: The 1st International Workshop on Multi-Sensorial Media and ApplicationsabstractToday, truly immersive multimedia systems demand the integration of emerging multi-sensorial media, which go beyond traditional audiovisual signals to include haptics, olfaction, motion capture, electroencephalograms, and other novel media forms. To effectively incorporate these modalities into cutting-edge multimedia systems, advances are needed across the entire pipeline, from processing and encoding to seamless integration. In addition, human-centric factors such as ergonomics and user experience must be considered to ensure practical implementation. Our workshop, the International Workshop on Multi-Sensorial Media and Applications (MSMA'2025), seeks to attract contributions related to multi-sensorial media systems, including system design, evaluation, coding, delivery, media analysis, multi-modal interaction, human factors, ergonomics, and related areas. By fostering collaboration among researchers, MSMA aims to bridge existing work in the field, spark innovation, and push the boundaries of multimedia technology. Tiesong Zhao, Qian Liu 0001, Zhisheng Yan |
ACM Multimedia | 3 |
| 2025 | Orbis: Redesigning Neural-enhanced Video Streaming for Live Immersive ViewingabstractEmerging live immersive viewing systems require streaming large 360 videos to users via limited wireless bandwidth. Neural-enhanced video streaming offers a promising solution by streaming down-scaled videos and enhancing them by client computation. However, prior systems treated video downscaling and enhancement separately, overlaying existing enhancement techniques onto current video infrastructure to accommodate legacy downscaling methods. This supplemental client design has led to spatial information loss and prohibitive model overheads in 360 video streaming systems. This paper presents Orbis, a redesigned, holistic neural-enhanced video streaming framework that integrates complementary down-scaling and enhancement for live immersive viewing. Orbis is empowered by an enhancement-driven interleaved downscaling approach, an inpainting-based enhancement model tailored to interleaved data, and a multi-scale tile adaptation scheme that optimizes immersive viewing experience in dynamic environments. Experimental results show that Orbis improves viewing experience by up to 60% and reduces wireless bandwidth by up to 49% compared to the best-performing baseline. Zhengguan Wu, Jingwei Liao, Anh Nguyen 0011, Feng Lin 0004, Zhisheng Yan |
SenSys | 5 |
| 2025 | ST-360: Spatial-Temporal Filtering-Based Low-Latency 360-Degree Video Analytics FrameworkabstractRecent advances in computer vision algorithms and video streaming technologies have facilitated the development of edge-server-based video analytics systems, enabling them to process sophisticated real-world tasks, such as traffic surveillance and workspace monitoring. Meanwhile, due to their omnidirectional recording capability, 360-degree cameras have been proposed to replace traditional cameras in video analytics systems to offer enhanced situational awareness. Yet, we found that providing an efficient 360-degree video analytics framework is a non-trivial task. Due to the higher resolution and geometric distortion in 360-degree videos, existing video analytics pipelines fail to meet the performance requirements for end-to-end latency and query accuracy. To address these challenges, we introduce the innovative ST-360 framework specifically designed for 360-degree video analytics. This framework features a spatial–temporal filtering algorithm that optimizes both data transmission and computational workloads. Evaluation of the ST-360 framework on a unique dataset of 360-degree first-responders videos reveals that it yields accurate query results with a 50% reduction in end-to-end latency compared to state-of-the-art methods. Jingwei Liao, Bo Chen 0025, Anh Nguyen 0011, Aditi Tiwari, Qian Zhou 0008, Zhisheng Yan, Klara Nahrstedt |
ACM Trans. Multim. Comput. Commun. Appl. | 7 |
| 2025 | A Patch Can Disrupt Live Video Streaming: Physical Adversarial Attacks on Deep Learning CompressionabstractDeep learning (DL)-based compression has achieved outstanding performance compared to traditional compression. However, due to the vulnerability of adversarial attacks on DL models, understanding the security of DL-based compression is crucial. Previous attacks have demonstrated the feasibility of failing DL-based compression models in the digital domain. However, these attacks rely on perfect digital modification of the whole image and internal access to camera/server files, preventing their usage in the physical world where such assumptions do not hold. In this article, we unveil the first physical adversarial attack targeting DL-based compression in the context of live video streaming. The proposed attack, namely CamHack , places a small-sized physical patch in the visual scene covered by the live camera to manipulate the bitrate of compressed content and disrupt live streaming. The patch is crafted to address color and geometric transformations in diverse streaming scenes while remaining inconspicuous. Extensive experiments in various streaming scenes and network conditions show that CamHack increases bit consumption by 276.58%, 384.74%, and 942.28% over the clean scene without a patch on three representative victim models. CamHack is also robust under various practical impacts such as patch location, lighting conditions, and patch size. Anh Nguyen 0011, Zhisheng Yan |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2024 | Vesper: Learning to Manage Uncertainty in Video StreamingabstractVideo codecs are crucial in video streaming systems. However, the quantization operation in existing codecs introduces irreversible jitters. Moreover, the common practice of fitting a single codec to diverse video content lacks the flexibility to adapt the parameters of a codec for specific content. They lead to the problem of quantization and content uncertainty. Our preliminary study shows an ideal codec without uncertainty gains a significant advantage over the conventional codec with uncertainty. However, realizing the ideal codec presents tremendous challenges in the generalizability and the costs of computation, transmission, and delay. In this paper, we present Vesper, a video streaming system that innovatively tackles uncertainty with two learning-based components, super-precision and self-evolution. The super-precision module builds a neural network that predicts original feature values from quantized feature values, which effectively mitigates the impact of quantization without inducing generalizability issues. The self-evolution module performs content-aware adaptation on the encoder and replaces non-content-aware video segments with content-aware ones on the fly, which addresses content uncertainty without adding significant costs to on-demand streaming. Evaluations demonstrate Vesper's superior Quality of Experience compared to streaming systems built with state-of-the-art codecs. Bo Chen 0025, Mingyuan Wu, Hongpeng Guo, Zhisheng Yan, Klara Nahrstedt |
MMSys | 4 |
| 2024 | Cross-content User Authentication in Virtual RealityabstractVirtual reality (VR) presents a rapidly evolving virtual world with varying types of visual content. Authenticating users periodically across different types of VR content to safeguard private data is a cornerstone in many sensitive VR applications. Previous studies have performed authentication by verifying users' actions or involuntary responses to certain visual stimuli. However, these methods tie users to either a predefined task or a single type of VR content, failing to support periodic authentication in dynamic VR interactions. In this paper, we propose a cross-content authentication framework that verifies users across different VR content by jointly analyzing eye movements and VR content. Specifically, we decouple displayed VR content from eye movements and generate a content-agnostic embedding determined by cognitive and physiological attributes unique to each user. Experiments with video viewing and text reading in VR show that our system achieves an F-score of 0.92 in cross-content user authentication. Wenda Shao, Shiqing Luo, Zhisheng Yan |
MobiCom | 3 |
| 2024 | NeRFHub: A Context-Aware NeRF Serving Framework for Mobile Immersive ApplicationsabstractNeural Radiance Fields (NeRF) are recognized for their exceptional photo-realism quality and superior modeling capabilities compared to traditional methods. NeRF empowers a novel application, termed NeRF serving. It delivers data from a server to a mobile client and renders 3D scenes on the client, facilitating a broad spectrum of mobile immersive applications. Towards a satisfactory user experience, we must serve NeRF with low latency while meeting constraints of high visual quality and real-time smoothness. Existing NeRF variants easily violate the constraints or cause an unnecessarily high latency when the diverse applications, mobile devices, and 3D scenes, termed the contexts, change in real life. In this paper, we present NeRFHub, a novel context-aware NeRF serving framework for mobile immersive applications. NeRFHub adeptly manages storage and computation costs, scales to diverse contexts, and swiftly navigates the vast design space inherent in NeRF serving. The evaluation results show that NeRFHub serves synthetic objects with 56%-66% reduced latency and realistic scenes with 26%-55% reduced latency when compared to the baseline without compromising quality or smoothness. Bo Chen 0025, Zhisheng Yan, Bo Han 0001, Klara Nahrstedt |
MobiSys | 2 |
| 2024 | Eavesdropping on Controller Acoustic Emanation for Keystroke Inference Attack in Virtual Reality
Shiqing Luo, Anh Nguyen 0011, Hafsa Farooq, Kun Sun 0001, Zhisheng Yan |
NDSS | 5 |
| 2024 | LiFteR: Unleash Learned Codecs in Video Streaming with Loose Frame Referencing
Bo Chen 0025, Zhisheng Yan, Yinjie Zhang, Zhe Yang 0010, Klara Nahrstedt |
NSDI | 2 |
| 2024 | ImmerScope: Multi-view Video Aggregation at Edge towards Immersive Content ServicesabstractThe multi-camera capture system is an emerging visual sensing modality. It facilitates the production of various immersive contents ranging from regular to neural videos. Although the delivery of immersive content is popular and promising, it suffers from the bandwidth bottleneck when streaming multi-view videos to the cloud (i.e., multi-view video aggregation). Existing works fail to provide a bandwidth-efficient and content-generic solution. Even the closest effort to ours based on the SOTA multi-view video codecs suffers from issues of underutilized dependency and content distortion. In this paper, we present ImmerScope, a multi-view video aggregation framework at the edge with a neural multi-view video codec. It outperforms existing solutions with highly-utilized dependency via neuron connections and distortion awareness via end-to-end training. Evaluations on diverse multi-camera setups show that ImmerScope outperforms single-view codecs by at least 64% bandwidth savings in peak-signal-to-noise ratio with a frame rate of 50 fps. Bo Chen 0025, Hongpeng Guo, Mingyuan Wu, Zhe Yang 0010, Zhisheng Yan, Klara Nahrstedt |
SenSys | 5 |
| 2024 | Penetration Vision through Virtual Reality Headsets: Identifying 360-degree Videos from Head Movements
Anh Nguyen 0011, Xiaokuan Zhang, Zhisheng Yan |
USENIX Security Symposium | 3 |
| 2024 | Shadow Based Non-Line-of-Sight Pedestrian Rushing Detection for Automated DrivingabstractAmong the foremost contributors to compromised driving safety is the abrupt emergence of obstacles or pedestrians within drivers’ non-line-of-sight regions. Previous investigations into non-line-of-sight imaging have predominantly depended on costly apparatus or have been confined to controlled laboratory settings (e.g., extensive planar reflectors and regulated illumination). Consequently, these technological approaches prove impractical within intricate driving environments. In this paper, we introduce a shadow based non-line-of-sight moving obstacle detection system devised to augment Advanced Driver Assistance Systems (ADAS), ensuring adequate time for safe response and halting. Our approach incorporates a shadow signal discriminator tailored to evaluate faint shadows generated by moving obstacles, such as pedestrians within blind spots. Note that we merely use commercial onboard sensors and our system is robust to various lighting scenarios and planar reflectors. We comprehensively assess our methodology's adaptability by employing datasets acquired from real-world driving scenarios encompassing diverse road surfaces and lighting conditions. The results substantiate the system's efficacy in detecting pedestrians in motion within NLOS regions, showcasing an impressive detection range of 22 meters. This proficiency enables the system to pre-emptively forewarn the ADAS, facilitating the maintenance of a safe distance from the pedestrian. Feng Lin 0004, Jin Li 0033, Meng Zhang 0022, Zhisheng Yan, Jian Xiao 0002, Kui Ren 0001 |
IEEE Trans. Mob. Comput. | 5 |
| 2024 | Context-aware Optimization for Bandwidth-Efficient Image Analytics OffloadingabstractConvolutional Neural Networks (CNN) have given rise to numerous visual analytics applications at the edge of the Internet. The image is typically captured by cameras and then live-streamed to edge servers for analytics due to the prohibitive cost of running CNN on computation-constrained end devices. A critical component to ensure low-latency and accurate visual analytics offloading over low bandwidth networks is image compression which minimizes the amount of visual data to offload and maximizes the decoding quality of salient pixels for analytics. Despite the wide adoption, JPEG standards and traditional image compression techniques do not address the accuracy of analytics tasks, leading to ineffective compression for visual analytics offloading. Although recent machine-centric image compression techniques leverage sophisticated neural network models or hardware architecture to support the accuracy-bandwidth tradeoff, they introduce excessive latency in the visual analytics offloading pipeline. This article presents CICO, a Context-aware Image Compression Optimization framework to achieve low-bandwidth and low-latency visual analytics offloading. CICO contextualizes image compression for offloading by employing easily-computable low-level image features to understand the importance of different image regions for a visual analytics task. Accordingly, CICO can optimize the tradeoff between compression size and analytics accuracy. Extensive real-world experiments demonstrate that CICO reduces the bandwidth consumption of existing compression methods by up to 40% under comparable analytics accuracy. Regarding the low-latency support, CICO achieves up to a 2× speedup over state-of-the-art compression techniques. Bo Chen 0025, Zhisheng Yan, Klara Nahrstedt |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2023 | Latency-Aware 360-Degree Video Analytics Framework for First Responders Situational AwarenessabstractFirst responders operate in hazardous working conditions with unpredictable risks. To better prepare for demands of the job, first responder trainees conduct training exercises that are being recorded and reviewed by the instructors, who check for objects indicating risks within the video recordings (e.g., firefighter with an unfastened gas mask). However, the traditional reviewing process is inefficient due to unanalyzed video recordings and limited situational awareness. For better reviewing experience, a latency-aware Viewing and Query Service (VQS) should be provided. The VQS should support object searching, which can be achieved using the video object detection algorithms. Meanwhile, the application of 360-degree cameras facilitates an unlimited field of view of the training environment. Yet, this medium represents a major challenge because low-latency high-accuracy 360-degree object detection is difficult due to higher resolution and geometric distortion. In this paper, we present the Responders-360 system architecture designed for 360-degree object detection. We propose a Dynamic Selection algorithm that optimizes computation resources while yielding accurate 360-degree object inference. The results, using a unique dataset collected from a firefighting training institute, show that the Responders-360 framework achieves 4x speedup and 25% memory usage reduction compared with the state-of-the-art methods. Jingwei Liao, Bo Chen 0025, Anh Nguyen 0011, Aditi Tiwari, Qian Zhou 0008, Zhisheng Yan, Klara Nahrstedt |
NOSSDAV | 7 |
| 2023 | Ghost-Probe: NLOS Pedestrian Rushing Detection with Monocular Camera for Automated DrivingabstractOne of the most serious factors compromising driving safety is when people in drivers' non-line-of-sight areas rush out suddenly. Existing studies on non-line-of-sight imaging rely on expensive equipment or are limited to severe laboratory conditions (e.g., massive planar reflectors and controlled illumination), rendering these technologies inapplicable in complex driving scenarios. In this paper, we propose a non-line-of-sight moving obstacle detection system Ghost-Probe, which can provide an advanced driver assistance system (ADAS) with sufficient time to respond and stop safely. We design a shadow signal discriminator to assess the weak shadows created by a moving obstacle, such as pedestrians in the blind area, while simultaneously filtering out the impacts of other complicated illumination. Note that we merely use commercial monocular cameras and our system is robust to a wide range of lighting scenarios and planar reflectors. We evaluate the generalizability of our approach using the datasets collected in real-world driving scenarios with a variety of road surface and lighting circumstances. The results indicate that our system can detect the moving pedestrian in the non-line-of-sight area at a distance of 20 meters and offer the ADAS system advance warning to keep a safe distance. Feng Lin 0004, Jin Li 0033, Meng Zhang 0022, Zhisheng Yan, Jian Xiao 0002, Kui Ren 0001 |
SenSys | 5 |
| 2023 | No-reference shadow detection quality assessment via reference learning and multi-mode exploring
Housheng Wei, Yanli Liu 0002, Guanyu Xing, Zhisheng Yan, Yanci Zhang |
Comput. Graph. | 4 |
| 2022 | DAO: Dynamic Adaptive Offloading for Video AnalyticsabstractOffloading videos from end devices to edge or cloud servers is the key to enabling computation-intensive video analytics. To ensure the analytics accuracy at the server, the video quality for offloading must be configured based on the specific content and the available network bandwidth. While adaptive video streaming for user viewing has been widely studied, none of the existing works can guarantee the analytics accuracy at the server in bandwidth- and content-adaptive way. To fill in this gap, this paper presents DAO, a dynamic adaptive offloading framework for video analytics that jointly considers the dynamics of network bandwidth and video content. DAO is able to maximize the analytics accuracy at the server by adapting the video bitrate and resolution dynamically. In essence, we shift the context of adaptive video transport from traditional DASH systems to a new dynamic adaptive offloading framework tailored for video analytics. DAO is empowered by some new discoveries about the inherent relationship between analytics accuracy, video content, bitrate, and resolution, as well as by an optimization formulation to adapt the bitrate and resolution dynamically. Results from the real-world implementation of object detection tasks show that DAO's performance is close to the theoretical bound, achieving 20% bandwidth saving and 59% category-wise mAP improvement compared to conventional DASH schemes. Taslim Murad, Anh Nguyen 0011, Zhisheng Yan |
ACM Multimedia | 3 |
| 2022 | Context-aware image compression optimization for visual analytics offloadingabstractConvolutional Neural Networks (CNN) have given rise to numerous visual analytics applications at the edge of the Internet. The image is typically captured by cameras and then live-streamed to edge servers for analytics due to the prohibitive cost of running CNN on computation-constrained end devices. A critical component to ensure low-latency and accurate visual analytics offloading over low bandwidth networks is image compression that minimizes the amount of visual data to offload and maximizes the decoding quality of salient pixels for analytics. Despite the wide adoption, JPEG standard and traditional image compression do not address the accuracy of analytics tasks, leading to ineffective compression for visual analytics offloading. Although recent machine-centric image compression techniques leverage sophisticated neural network models or hardware architecture to support the accuracy-bandwidth trade-off, they introduce excessive latency in the visual analytics offloading pipeline. This paper presents CICO, a Context-aware Image Compression Optimization framework to achieve low-bandwidth and low-latency visual analytics offloading. CICO contextualizes image compression for offloading by employing easily-computable low-level image features to understand the importance of different image regions for a visual analytics task. Accordingly, CICO can optimize the trade-off between compression size and analytics accuracy. Extensive real-world experiments demonstrate that CICO reduces the bandwidth consumption of existing compression methods by up to 40% under a comparable analytics accuracy. In terms of the low-latency support, CICO achieves up to a 2x speedup over state-of-the-art compression techniques. Bo Chen 0025, Zhisheng Yan, Klara Nahrstedt |
MMSys | 2 |
| 2022 | HoloLogger: Keystroke Inference on Mixed Reality Head Mounted DisplaysabstractWhen using personal computing services in mixed reality (MR) such as online payment and social media, sensitive information and account passwords must be typed in MR. To design secure MR systems and build up user trust, it is imperative to first understand the security threat to the sensitive MR input. Although keystroke inference attacks by analyzing human-computer interaction in videos or via wireless signals have been successful, they require placing extra hardware near the user which is easily noticeable in practice. In this paper, we expose a more dangerous malware-based attack through the vulnerability that no permission is required for accessing MR motion data. We aim to monitor MR headset motion and infer the user input through a benign App. Realizing the attack system requires addressing unique challenges in MR such as six-degree-of-freedom (6DoF) device motion and no explicit motion signal for keystroke identification. To this end, we present HoloLogger, the first malware-based keystroke inference attack system on HoloLens. HoloLogger is empowered by a 6DoF-head-motion-driven key tracking scheme and an air-tap-pattern-based keystroke inference framework. Extensive evaluations with 25 users and 750 inference trials of passwords consisting of 4–8 lowercase English letters demonstrate that HoloLogger successfully achieves a top-5 accuracy of 93%. HoloLogger is also robust in various environments such as different user positions and input categories. Shiqing Luo, Zhisheng Yan |
VR | 3 |
| 2022 | A novel transmission approach based on video content for 360-degree streaming
Fan Li 0003, Zhisheng Yan |
Multim. Tools Appl. | 3 |
| 2022 | Embedding Pose Information for Multiview Vehicle Model RecognitionabstractVehicle model recognition is a typical fine-grained classification task that has a wide range of application prospects in safe cities and constitutes a research hotspot in the field of computer vision. Vehicles in images can appear at various angles, resulting in large differences in appearance. The existence of “multiviews” renders vehicle model recognition challenging. Recent research on vehicle model recognition has not fully explored the pose information of vehicles in different images, resulting in low model performance. In this study, we use vehicle pose information to solve the multiview vehicle model recognition (MV-VMR) problem and design a convolutional neural network (CNN) model with embedded vehicle pose information, known as the embedding pose CNN (EP-CNN). The proposed model includes two subnetworks: the pose estimation subnetwork (PE-SubNet) and vehicle model classification subnetwork (VMC-SubNet). PE-SubNet extracts the vehicle pose information, including the pose features and vehicle viewpoint. In VMC-SubNet, considering the scale variation of vehicles, an improved squeeze-and-excitation (SE) block, named the MultiSE block is implemented. We embed the vehicle viewpoint into the MultiSE block, which reweighs each channel such that the extracted features elicit different responses to different viewpoints. Subsequently, the pose features and classification features are integrated for classification. Experiments are conducted on the benchmark CompCars web-nature and Stanford Cars datasets. The results demonstrate that the proposed EP-CNN method can achieve higher recognition accuracy than most classic CNN models and several state-of-the-art fine-grained vehicle model classification algorithms. Code has been made available at:https://github.com/HFUT-CV/EP-CNN. Yuanzi Fu, Wei Jia 0001, Jun Yu 0001, Zhisheng Yan |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2021 | Deep Contextualized Compressive Offloading for ImagesabstractRecent years have witnessed sensors becoming an indispensable part of our life with the camera being one of the most popular and widely deployed sensors. The camera gives rise to numerous vision-based IoT applications that generate high-level understandings of a live video stream by performing analysis on end devices like mobile or embedded devices. Typically, these applications are built with deep learning (DL) models to conduct complex vision tasks, e.g., image classification and object detection. Due to the prohibitive cost of running DL models on end devices close to the camera and with limited computation capabilities, it is widely adopted to offload the computation to a nearby powerful edge server. However, there is a gap between the restricted offloading bandwidth of the end device and the large volume of image data incurred by the live video stream. In this paper, we present Deep Contextualized Compressive Offloading for Images (DCCOI), a lightweight, context-aware, and bandwidth-efficient offloading framework for images. DCCOI consists of the spatial-adaptive encoder, a lightweight neural network, to spatial-adaptively compress the image, and the generative decoder for reconstructing the image from the compressed data. In contrast to existing DL-based encoders, the spatial-adaptive encoder allows an image region to be encoded into different numbers of feature values based on the information in it. This offers a variable-length coding method for image compression, which is a more optimal way for compression than the fix-length coding method took by existing DL-based compression approaches and demonstrates superior accuracy-compression rate trade-offs. We evaluate DCCOI against several baseline compression techniques while serving an object detection-based application. The results show that DCCOI roughly reduces the offloading size of JPEG by a factor of 9 and DeepCOD, the state-of-the-art offloading approach, by 20% with similar accuracy and a compression overhead less than 50ms. Bo Chen 0025, Zhisheng Yan, Hongpeng Guo, Zhe Yang 0010, Ahmed Ali-Eldin, Prashant J. Shenoy, Klara Nahrstedt |
SenSys | 2 |
| 2021 | Feature fusion quality assessment model for DASH video streamingabstractAbstract Dynamic Adaptive Streaming over HTTP (DASH) employs the flexible rate adaptation scheme to combat with time‐varying channel conditions. In addition to compression impairment, DASH video streaming suffers from transmission impairment, such as rate switches and stalling events. Both of them severely degrade users' Quality of Experience (QoE). Herein, an assessment model is established for DASH video streaming by directly jointing multiple QoE influential factors, which quantify impairments resulting from the compression and transmission. To demonstrate the influence of video content characteristics on users' QoE, spatio‐temporal content perceptual features are employed to represent the compression impairment. When reflecting temporal characteristics, a novel motion vector padding method is proposed to quantify the influence of intra macroblock on the human visual system. The proposed model is evaluated on a newly public Waterloo SQoE‐III database, which is available for DASH video streaming. Experimental results demonstrate that the authors' model outperforms the comparative models and owns the strong generalization ability to different video contents. Moreover, the proposed model is statistically superior to the existing models. Fan Li 0003, Zhisheng Yan |
IET Image Process. | 3 |
| 2020 | An Analysis of Delay in Live 360° Video Streaming SystemsabstractWhile live 360° video streaming provides an enriched viewing experience, it is challenging to guarantee the user experience against the negative effects introduced by start-up delay, event-to-eye delay, and low frame rate. It is therefore imperative to understand how different computing tasks of a live 360° streaming system contribute to these three delay metrics. Although prior works have studied commercial live 360° video streaming systems, none of them has dug into the end-to-end pipeline and explored how the task-level time consumption affects the user experience. In this paper, we conduct the first in-depth measurement study of task-level time consumption for five system components in live 360° video streaming. We first identify the subtle relationship between the time consumption breakdown across the system pipeline and the three delay metrics. We then build a prototype Zeus to measure this relationship. Our findings indicate the importance of CPU-GPU transfer at the camera and the server initialization as well as the negligible effect of 360° video stitching on the delay metrics. We finally validate that our results are representative of real world systems by comparing them with those obtained with a commercial system. Md Reazul Islam, Shivang Aggarwal, Dimitrios Koutsonikolas, Y. Charlie Hu, Zhisheng Yan |
ACM Multimedia | 6 |
| 2020 | OcuLock: Exploring Human Visual System for Authentication in Virtual Reality Head-mounted Display
Shiqing Luo, Anh Nguyen 0011, Chen Song 0001, Feng Lin 0004, Wenyao Xu, Zhisheng Yan |
NDSS | 6 |
| 2020 | Integrating Mobile Display Energy Saving into Cloud-Based Video Streaming via Rate-Distortion-Display Energy ProfilingabstractMobile displays have been recognized as the major contributors to the energy consumption of contemporary mobile video services. Existing display energy reduction (DER) algorithms focus on local video/image processing by utilizing the computation resources on the mobile devices. As such, a per-device DER strategy is highly inefficient from the systematic perspective since the same computation is repeated among millions of individual mobile devices. In this new era of cloud-based video streaming, a natural question to ask is can ubiquitous cloud resources be exploited to overcome these drawbacks. It is based on this motivation that we design a paradigm-shifting strategy to integrate the display energy saving engine into cloud-based video streaming in order to simultaneously benefit massive mobile devices with a one-time video processing in the cloud. By taking full advantage of the computational and storage resources in the cloud, a family of Rate-Distortion-Display Energy (R-D-DE) profiles can be created for a given video source and for different types of mobile devices. The preliminary experimental results with OLED displays prove that we can achieve the desired DER performance through jointly managing the display energy and user experience by implementing the proposed R-D-DE integrated video encoding engine in the cloud. Qian Liu 0001, Zhisheng Yan, Chang Wen Chen |
IEEE Trans. Cloud Comput. | 2 |
| 2019 | Event-driven stitching for tile-based live 360 video streamingabstract360 video streaming is gaining popularity because of the new type of experience it creates. Tile-based approaches have been widely used in VoD 360 video streaming to save the network bandwidth. However, they cannot be extended to the case of live streaming because they assume the 360 videos stitched offline before streaming. Instead, stitching has to be done in real-time in live 360 video streaming. More importantly, the stitching speed as shown in our experiments is one order of magnitude lower than the network transmission speed, making stitching more of a deciding factor of the overall frame rate than the network transmission speed. In this paper, we design a stitching algorithm for tile-based live 360 video streaming that adapts stitching quality to make the best use of the timing budget. There are two main challenges. First, existing tile-based approaches do not consider various semantic information in different scenarios. Second, the decision of tiling schemes for tile-based stitching is non-trivial. To solve the above two challenges, we present an event-driven stitching algorithm for tile-based 360 video live streaming, which consists of such an event-driven model to abstract various semantic information as events and a tile actuator to make tiling scheme decisions. We implement a streaming system based on event-driven stitching called LiveTexture. To evaluate the proposed algorithm, we compare LiveTexture with other baseline systems and show that LiveTexture adapts well to various timing budgets by meeting 89.4% of the timing constraints. We also demonstrate that LiveTexture utilizes the timing budget more efficiently than others. Bo Chen 0025, Zhisheng Yan, Haiming Jin, Klara Nahrstedt |
MMSys | 2 |
| 2019 | A saliency dataset for 360-degree videosabstractDespite the increasing popularity, realizing 360-degree videos in everyday applications is still challenging. Considering the unique viewing behavior in head-mounted display (HMD), understanding the saliency of 360-degree videos becomes the key to various 360-degree video research. Unfortunately, existing saliency datasets are either irrelevant to 360-degree videos or too small to support saliency modeling. In this paper, we introduce a large saliency dataset for 360-degree videos with 50,654 saliency maps from 24 diverse videos. The dataset is created by a new methodology supported by psychology studies in HMD viewing. We describe an open-source software implementing this methodology that can generate saliency maps from any head tracking data. Evaluation of the dataset shows that the generated saliency is highly correlated with the actual user fixation and that the saliency data can provide useful insight on user attention in 360-degree video viewing. The dataset and the program used to extract saliency are both made publicly available to facilitate future research. Anh Nguyen 0011, Zhisheng Yan |
MMSys | 2 |
| 2019 | A measurement study of YouTube 360° live video streamingabstract360° live video streaming is becoming increasingly popular. While providing viewers with enriched experience, 360° live video streaming is challenging to achieve since it requires a significantly higher bandwidth and a powerful computation infrastructure. A deeper understanding of this emerging system would benefit both viewers and system designers. Although prior works have extensively studied regular video streaming and 360° video on demand streaming, we for the first time investigate the performance of 360° live video streaming. We conduct a systematic measurement of YouTube's 360° live video streaming using various metrics in multiple practical settings. Our key findings suggest that viewers are advised not to live stream 4K 360° video, even when dynamic adaptive streaming over HTTP (DASH) is enabled. Instead, 1080p 360° live video can be played smoothly. However, the extremely large one-way video delay makes it only feasible for delay-tolerant broadcasting applications rather than real-time interactive applications. More importantly, we have concluded from our results that the primary design weakness of current systems lies in inefficient server processing, non-optimal rate adaptation, and conservative buffer management. Our research insight will help to build a clear understanding of today's 360° live video streaming and lay a foundation for future research on this emerging yet relatively unexplored area. Shiqing Luo, Zhisheng Yan |
NOSSDAV | 3 |
| 2019 | Toward Guaranteed Video Experience: Service-Aware Downlink Resource Allocation in Mobile Edge NetworksabstractVideo delivery has been playing an essential role in video services over edge networks. Although HTTP segment-based streaming, e.g., Dynamic Adaptive Streaming over HTTP (DASH), has become the prevailing technique, it cannot provide guaranteed video playback in terms of bitrate to mobile users. In essence, HTTP streaming downloads the video segments in a best effort fashion, i.e., passively responding to the channel dynamics. This can cause unstable playback with frequent rebuffer and multi-client competition that degrades a network-wide performance. In this paper, we present a network-assisted streaming framework for Guaranteed Playback-Experience Streaming over HTTP (GESH) that leverages the proactive control of network resources and joint coordination among multiple clients for service-aware network resource allocation. Specifically, GESH is empowered by a new weighted proportional fair scheduling without modifying existing cellular infrastructure, a per-segment channel variation model, and a suite of algorithms to seek the optimal weights for the scheduling. Extensive evaluations show that GESH can maximally guarantee the video playback of multiple users, as well as significantly outperforming conventional HTTP streaming and current DASH systems. Zhisheng Yan, Miao Zhao, Cédric Westphal, Chang Wen Chen |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2019 | Scalable Access Control For Privacy-Aware Media SharingabstractThe prevalence of social networks has made it easier than ever for users to share their photos, videos, and other media content with anybody from anywhere. However, the easy access of user-generated media content also brings about privacy concerns. Traditional access control mechanisms, where a single access policy is made for a specific piece of content, cannot satisfy the user privacy requirements in large-scale media sharing systems. Instead, configuring multiple levels of access privileges for the shared media content is desired. On one hand, it conforms to the principle of social networks in information propagation. On the other hand, it accords with the diverse and complex social relationship among social network users. In this paper, we propose a scalable media access control (SMAC) system to enable such a configuration in a secure and efficient manner. The proposed SMAC system is empowered by the scalable ciphertext policy attribute-based encryption algorithm as well as a comprehensive key-management scheme. We provide formal security proof to prove the security of the proposed SMAC system. In addition, we conduct extensive experiments on mobile devices to demonstrate its efficiency. Changsha Ma, Zhisheng Yan, Chang Wen Chen |
IEEE Trans. Multim. | 2 |
| 2019 | SSPA-LBS: Scalable and Social-Friendly Privacy-Aware Location-Based ServicesabstractPrivacy-aware location-based service (PA-LBS) preserves LBS users' privacy but undesirably sacrifices service quality. In order to balance the two factors with satisfactory user experience, existing frameworks are faced with two barriers, that is, scalability and social-friendliness. First, existing schemes do not enable LBS users to flexibly scale their privacy level on service provision. Such a lack of scalability easily results in either unacceptable service-quality degradation or insufficient privacy protection and fails to meet dynamic user requirements. Second, existing schemes handle privacy protection by merely considering the trust relationship between users and servers but ignore the complex trust relationships among users. As a result, users cannot preserve privacy in location-based social services that involve user-to-user interactions. In this paper, we present the first scalable and social-friendly PA-LBS system. In particular, we propose a novel camouflage algorithm with a formal privacy guarantee that enables LBS users to expose their location information by scaling two privacy related factors, that is, camouflage range and place type. Furthermore, we apply the scalable ciphertext policy attribute-based encryption algorithm to enable LBS users to effectively control the access from other users to their location information. Moreover, we also demonstrated the operational efficiency of the proposed system through successful implementations on Android devices. Changsha Ma, Zhisheng Yan, Chang Wen Chen |
IEEE Trans. Multim. | 2 |
| 2018 | Your Attention is Unique: Detecting 360-Degree Video Saliency in Head-Mounted Display for Head Movement PredictionabstractHead movement prediction is the key enabler for the emerging 360-degree videos since it can enhance both streaming and rendering efficiency. To achieve accurate head movement prediction, it becomes imperative to understand user's visual attention on 360-degree videos under head-mounted display (HMD). Despite the rich history of saliency detection research, we observe that traditional models are designed for regular images/videos fixed at a single viewport and would introduce problems such as central bias and multi-object confusion when applied to the multi-viewport 360-degree videos switched by user interaction. To fill in this gap, this paper shifts the traditional single-viewport saliency models that have been extensively studied for decades to a fresh panoramic saliency detection specifically tailored for 360-degree videos, and thus maximally enhances the head movement prediction performance. The proposed head movement prediction framework is empowered by a newly created dataset for 360-degree video saliency, a panoramic saliency detection model and an integration of saliency and head tracking history for the ultimate head movement prediction. Experimental results demonstrate the measurable gain of both the proposed panoramic saliency detection and head movement prediction over traditional models for regular images/videos. Anh Nguyen 0011, Zhisheng Yan, Klara Nahrstedt |
ACM Multimedia | 2 |
| 2018 | Session details: Experience-1 (Multimedia Entertainment and Experience)
Zhisheng Yan |
ACM Multimedia | 1 |
| 2018 | CrowdDBS: A Crowdsourced Brightness Scaling Optimization for Display Energy Reduction in Mobile VideoabstractMobile display has become one of the most power-hungry components in mobile video viewing. Currently, mobile devices can reduce the display energy by performing dynamic brightness scaling (DBS) under the distortion constraint of video signals. We observe that there is a pitfall preventing current practice from systematic display energy reduction. In particular, existing objective DBS schemes lack direct connection to the subjective human perception on DBS-enabled videos, which is the key to achieving human-centered energy-experience optimization. To overcome this pitfall, we present CrowdDBS, a crowdsourced display energy reduction framework for mobile video viewing. CrowdDBS is empowered by a set of crowdsourcing studies that uncover the relationship between human perception and DBS frequency, magnitude, and temporal consistency, respectively. Motivated by the insights obtained from these studies, CrowdDBS employs a suit of designs and a DBS optimization framework to optimize the energy-experience tradeoff in mobile video viewing. Comprehensive experimental results and user evaluations under a variety of practical settings show that CrowdDBS can achieve 37 percent device energy reduction on average while guaranteeing satisfactory user experience in mobile video viewing. Zhisheng Yan, Qian Liu 0001, Tong Zhang 0002, Chang Wen Chen |
IEEE Trans. Mob. Comput. | 1 |
| 2017 | LARM: A Lifetime Aware Regression Model for Predicting YouTube Video PopularityabstractOnline content popularity prediction provides substantial value to a broad range of applications in the end-to-end social media systems, from network resource allocation to targeted advertising. While using historical popularity can predict the near-term popularity with a reasonable accuracy, the bursty nature of online content popularity evolution makes it difficult to capture the correlation between historical data and future data in the long term. Although various existing efforts have been made toward long-term prediction, they need to accumulate a long enough historical data before the prediction and their model assumptions cannot be applied to the complex YouTube networks with inherent unpredictability. Changsha Ma, Zhisheng Yan, Chang Wen Chen |
CIKM | 2 |
| 2017 | Too Many Pixels to Perceive: Subpixel Shutoff for Display Energy Reduction on OLED SmartphonesabstractOrganic light-emitting diode (OLED) has been widely recognized as the next-generation mobile display. Recently, smartphone manufacturers have been pushing up the pixel density of OLED display. Unfortunately, such an effort does not necessarily improve the everyday viewing because of the limitation in human visual acuity. Instead, high pixel density OLED can drain the battery power even more quickly since the power dissipation of OLED is determined by the number of displayed pixels and their RGB values, or subpixels. This paper presents a new design dimension to remedy this prevailing issue by leveraging the intuition that shutting off redundant subpixels of the display content on OLED can reduce power consumption without impacting viewing perception. We introduce ShutPix, a power-saving display system for OLED smartphones that can optimally shut off the redundant subpixels before the content is displayed. Inspired by the motivational studies, ShutPix is empowered by a suite of designs based on visual acuity, human perception, and content redundancy. Experimental results show that ShutPix can, on average, reduce 21% of display power and 15% of system power without degrading user viewing experience. Zhisheng Yan, Chang Wen Chen |
ACM Multimedia | 1 |
| 2017 | Prius: Hybrid Edge Cloud and Client Adaptation for HTTP Adaptive Streaming in Cellular NetworksabstractIn this paper, we present Prius, a hybrid edge cloud and client adaptation framework for HTTP adaptive streaming (HAS) by taking advantage of the new capabilities empowered by recent advances in edge cloud computing. In particular, emerging edge clouds are capable of accessing an application layer and radio access networks (RANs) information in real time. Coupled with powerful computation support, an edge cloud-assisted strategy is expected to significantly enrich mobile services. Meanwhile, although HAS has established itself as the dominant technology for video streaming, one key challenge for adapting HAS to mobile cellular networks is in overcoming the inaccurate bandwidth estimation and unfair bitrate adaptation under the highly dynamic cellular links. Edge cloud-assisted HAS presents a new opportunity to resolve these issues and achieve systematic enhancement of quality of experience (QoE) and QoE fairness in cellular networks. To explore this new opportunity, Prius overlays a layer of adaptation intelligence at the edge cloud to finalize the adaptation decisions while considering the initial bandwidth-irrelevant bitrate selection at the clients. Prius is able to exploit RAN channel status, client device characteristics, and application-layer information in order to jointly adapt the bitrate of multiple clients. Prius also adopts a QoE continuum model to track the cumulative viewing experience and an exponential smoothing estimation to accurately estimate a future channel under different moving patterns. Extensive trace-driven simulation results show that Prius with hybrid edge cloud and client adaptation is promising under both slow and fast-moving environments. Furthermore, the Prius adaptation algorithm achieves a near-optimal performance that outperforms the exiting strategies. Zhisheng Yan, Jingteng Xue, Chang Wen Chen |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2017 | Medium Access Control for Wireless Body Area Networks with QoS Provisioning and Energy Efficient DesignabstractWith the promising applications in e-Health and entertainment services, wireless body area network (WBAN) has attracted significant interest. One critical challenge for WBAN is to track and maintain the quality of service (QoS), e.g., delivery probability and latency, under the dynamic environment dictated by human mobility. Another important issue is to ensure the energy efficiency within such a resource-constrained network. In this paper, a new medium access control (MAC) protocol is proposed to tackle these two important challenges. We adopt a TDMA-based protocol and dynamically adjust the transmission order and transmission duration of the nodes based on channel status and application context of WBAN. The slot allocation is optimized by minimizing energy consumption of the nodes, subject to the delivery probability and throughput constraints. Moreover, we design a new synchronization scheme to reduce the synchronization overhead. Through developing an analytical model, we analyze how the protocol can adapt to different latency requirements in the healthcare monitoring service. Simulations results show that the proposed protocol outperforms CA-MAC and IEEE 802.15.6 MAC in terms of QoS and energy efficiency under extensive conditions. It also demonstrates more effective performance in highly heterogeneous WBAN. Bin Liu 0016, Zhisheng Yan, Chang Wen Chen |
IEEE Trans. Mob. Comput. | 2 |
| 2016 | Cloud-based video streaming with systematic mobile display energy saving: Rate-distortion-display energy profilingabstractMobile display has been considered as the major contributor to the energy consumption of the ever-increasing mobile video services. Current practices in display energy reduction (DER) utilize local computing resources to analyze the video content before DER strategies can be applied in a per-device fashion. For a given video, same analytical computations are repeated in millions of individual devices. In this paper, we demonstrate that a paradigm shifting framework can be designed to systematically move the common DER local processing to the streaming server with the emergence of cloud-based video services. This framework has the potential to replace the massive per-device DER computation by a one-time global video processing in the cloud. To accomplish this ultimate goal in DER, a new family of video rate-distortion (R-D) profile with embedded DER strategies shall be properly generated. This family of rate-distortion-display energy (R-D-DE) profiles contains a set of common DER parameters to be directly extracted and employed by individual mobile devices to achieve desired display energy saving without repeated local computation. Performance evaluations of the proposed design are carried out to show that this family of R-D-DE profiles is indeed able to command the systematic DER design based on the intrinsic relationships among bitrate, video quality and display energy saving. Qian Liu 0001, Zhisheng Yan, Chang Wen Chen |
ICIP | 2 |
| 2016 | Forecasting initial popularity of just-uploaded user-generated videosabstractUser-generated videos (UGVs) have dominated contemporary social networking sites (SNSs). Forecasting their popularity is of great relevance to a broad range of online services. All existing studies forecast popularity of UGVs using their popularity statistics that are accumulated for a period of time after they are uploaded. Hence, there is always a substantial time lag (days to weeks) before popularity forecast can take effects. However, such a time lag is undesirable for timely popularity forecast as forecasting initial popularity during UGVs' lifetime is vitally important. In fact, we have found in our measurement that the most popular UGVs usually precede others starting from the beginning days and UGVs generally receive the highest attentions during the first few days. In this paper, we present the first exploration on forecasting initial popularity for UGVs at their uploading moment without accumulating their popularity statistics. Specifically, we first design an effective crawler framework to collect the publicly observable features of videos at their uploading moment. We then collect a representative and large YouTube video data set with 318,627 videos. Based on the data set, we select the most relevant features as predictors and design a neural network-based learning model to forecast initial popularity of just-uploaded UGVs. Experimental results validate the effectiveness of the proposed forecasting model and demonstrate the model's benefits for online services such as in-video advertising and video caching. Changsha Ma, Zhisheng Yan, Chang Wen Chen |
ICIP | 2 |
| 2016 | Attribute-based multi-dimension scalable access control for social media sharingabstractSocial media sharing is one of the most popular social interactions in online social networks (OSNs). Due to the diverse networking conditions and various privacy requirements of OSN users, scalable media sharing has become a promising paradigm. It allows a media data distributor to share a media content of different qualities with different data consumers. To guarantee user privacy in scalable media sharing, it is essential to design an effective scalable media access control (SMAC) mechanism. However, all existing schemes cannot support scalable media streams with more than two dimensions, which significantly limits the flexibility of OSN services. In this paper, we present the first multi-dimension SMAC (MD-SMAC) system for social media sharing, based on the proposed scalable ciphertext policy attribute-based encryption (SCP-ABE) algorithm. In the MD-SMAC system, secure and reliable access control can be performed on multidimension scalable media streams according to data consumers' attributes. Through the security analysis, we prove the security and reliability of the MD-SMAC system. We also conduct experiments on mobile devices to demonstrate the computation efficiency of the proposed system. Changsha Ma, Zhisheng Yan, Chang Wen Chen |
ICME | 2 |
| 2016 | RnB: rate and brightness adaptation for rate-distortion-energy tradeoff in HTTP adaptive streaming over mobile devicesabstractVideo streaming is a prevalent mobile service that drains a significant amount of battery power. While various efforts have been made toward saving both video transfer and display energy, they are independently designed in an ad-hoc way and thereby can cause some non-apparent yet critical performance issues. To fill in this gap, this paper presents a fundamentally new design by jointly considering the end-to-end pipeline from the initial video encoding to the final mobile display. In essence, we shift the classic R-D tradeoff that has governed streaming system designs for decades to a fresh rate-distortion-energy (R-D-E) tradeoff specifically tailored for mobile devices. We present RnB, a video bitrate and display brightness adaptation platform that is standard-compliant, backward compatible, and device-neutral in order to achieve the proposed R-D-E tradeoff. RnB is empowered by some new discovery about the inherent relationship among bitrate, display brightness, and video quality as well as by an control-theoretic formulation to dynamically adapt the bitrate and scale the display brightness. Experimental results based on real-time implementation show that RnB can achieve an average of 19% energy reduction with final video quality comparable to conventional R-D based schemes. Zhisheng Yan, Chang Wen Chen |
MobiCom | 1 |
| 2015 | Service provisioning and profit maximization in network-assisted adaptive HTTP streamingabstractMobile adaptive HTTP streaming with centralized consideration of multiple streams has gained increasing interest. It poses a special challenge that the interests of both content provider and network operator need to be deliberately balanced. More importantly, the adaptation is required to be flexible enough to be ported to various systems that work under different network environments, QoE levels, and economic objectives. To address these challenges, we propose a Markov Decision Process (MDP) based network-assisted adaptation framework, wherein cost of buffering, significant playback variation, bandwidth management and income of playback are jointly investigated. We then demonstrate its promising service provisioning and maximal profit for a mobile network in which fair or differentiated service is required. Zhisheng Yan, Cédric Westphal, Chang Wen Chen |
ICIP | 1 |
| 2015 | Exploring QoE for Power Efficiency: A Field Study on Mobile Videos with LCD DisplaysabstractDisplay power consumption has become a major concern for both mobile users and design engineers, especially considering the prevalence of today's video-rich mobile services. The power consumption of liquid crystal display (LCD), a dominant mobile display technology, can be reduced by dynamic backlight scaling (DBS). However, such dynamic changes of screen brightness may degrade users' quality of experience (QoE) in viewing videos. How would QoE be impacted by different DBS strategies has not yet been understood clearly and thus obscures the way to achieve systematic power saving. In this paper, we take a first step to explore the QoE of DBS on smartphones and aim at maximally enhancing the display power performance without negatively impacting users' QoE. In particular, we conduct three motivational studies to uncover the inherent relationship between QoE and backlight scaling frequency, magnitude, and temporal consistency, respectively. Motivated by the findings of these studies, we design a suite of techniques to implement a comprehensive DBS strategy. We demonstrate an example application of the proposed DBS designs in a mobile video streaming system. Measurements and user evaluations show that more than 40% system power reduction, or equivalently, 20% more power savings than the non-QoE approaches, can be achieved without QoE impairment. Zhisheng Yan, Qian Liu 0001, Tong Zhang 0002, Chang Wen Chen |
ACM Multimedia | 1 |
| 2014 | QoE continuum driven HTTP adaptive streaming over multi-client wireless networksabstractDifferent from traditional HTTP adaptive streaming (HAS) in which only one client is considered, HAS over multi-client wireless networks faces new challenges. The Quality of Experience (QoE) of users becomes unstable due to users' competition for shared bandwidth. It is thus important to accurately estimate the perceived experience of users and then adapt the streaming process accordingly. Furthermore, the QoE fairness among multiple clients subscribing to the same services shall also be addressed. In this research, we propose a QoE continuum driven HAS adaptation algorithm to address these challenges. We model the QoE continuum as an integrated consideration of cumulative playback quality and playback smoothness. Based on this model, we jointly optimize the quality adaptation of multiple users by considering both QoE history and channel status. Moreover, we propose to use quantization parameter and segment size to represent the video files in a fine-grained fashion, in order to more effectively capture the bandwidth fluctuation. The results from extensive simulations show that the proposed scheme can provide balanced and satisfactory QoE among multiple clients. Zhisheng Yan, Jingteng Xue, Chang Wen Chen |
ICME | 1 |
| 2014 | Admission Control for Wireless Adaptive HTTP Streaming: An Evidence Theory Based ApproachabstractIn this research, we propose an evidence theory based admission control scheme for wireless cellular adaptive HTTP streaming systems. This novel scheme allows us to effectively address the uncertainty and inaccuracy in QoE management and network estimation, and seamlessly grant or deny the access requests. Specifically, based on recent work of QoE continuum model and QoE continuum driven adaptation algorithm, we utilize Dempster-Shafer evidence theory to assign proper degree of belief to admission, rejection and an uncertainty decision for each user's evidence. We then can strategically combine the weighted evidence of multiple users and make the final decision. The evaluation results show that the proposed scheme can provide satisfactory QoE for both existing and new users while still achieving comparable bandwidth efficiency. Zhisheng Yan, Chang Wen Chen, Bin Liu 0016 |
ACM Multimedia | 1 |
| 2013 | Prediction-based dynamic relay transmission scheme for Wireless Body Area NetworksabstractTo support long-term pervasive healthcare services, communications in Wireless Body Area Networks (WBANs) need to be both reliable and energy-efficient. As a cooperative transmission method, relay transmission scheme works effectively in resisting shadowing effect and improving reliability in WBANs. However, the extra energy consumption introduced by relay transmission is very high, which can shorten the lifetime of the whole network. In this paper, temporal and spatial correlation models for on-body channels are first presented to better characterize the slow fading effect of on-body channels. Then a prediction-based dynamic relay transmission (PDRT) scheme that makes full use of the correlation characteristics of on-body channels is proposed. In the PDRT scheme, “when to relay” and “who to relay” are decided in an optimal way based on the last known channel states. Moreover, neither extra signaling procedure nor dedicated channel sensing period is needed. Simulation results show that the PDRT scheme achieves significant performance improvement in energy efficiency, as well as ensuring the transmission reliability. Bin Liu 0016, Zhisheng Yan, Chi Zhang 0001, Chang Wen Chen |
PIMRC | 3 |
| 2012 | QoS-driven scheduling approach using optimal slot allocation for Wireless Body Area NetworksabstractWireless Body Area Network (WBAN) is a promising type of networks that mainly targets at applications in ubiquitous communication and e-Health services. Different from other types of networks, one important challenge for WBAN is that its quality of service (QoS) requirement, in terms of delivery probability and data rate, will be time varying since human body is a highly dynamic physical environment. Another significant challenge for WBAN is that energy efficiency needs to be guaranteed in such a resource-limited network. In this paper, a QoS-driven scheduling approach is proposed to address these challenges. We model the WBAN channel as a Markov model as suggested by the emerging IEEE 802.15.6 BAN standard and propose a threshold-based scheme to adjust the transmission order of nodes. The number of slots for each node is optimally assigned according to the QoS requirement while minimizing the energy consumption of nodes. The results from extensive simulations show that the proposed approach can provide high QoS and energy efficiency under different network conditions, especially in highly heterogeneous ones in WBAN. Zhisheng Yan, Bin Liu 0016, Chang Wen Chen |
Healthcom | 1 |
| 2012 | Block-based variable density compressed image samplingabstractCompressed sampling (CS) is a technique that enables signal reconstruction at sub-Nyquist sampling rate. A key problem in CS is how to design the sampling scheme. In this paper, we propose a novel sampling method for compressed image sampling, which exploits a priori information and uses a block-based strategy to improve image reconstruction. Our block-based sampling scheme assigns more samples to blocks with more high-frequency contents while making sure that important coefficients of each block are sampled. Simulation results show that our proposed method outperforms existing methods on both reconstruction quality and running time. Bin Liu 0016, Zixiang Xiong, Gonzalo R. Arce, Javier Garcia-Frías, Wenwu Zhu 0001, Zhisheng Yan |
ICIP | 7 |
| 2011 | A context aware MAC protocol for medical Wireless Body Area NetworkabstractDuring long-term medical monitoring in Wireless Body Area Networks (WBAN), network requirements (i.e. traffic loads and latency) of various data sources may be different at different time. High traffic loads may lead to data overload and unacceptable latency, which makes potential danger of patients undiagnosed. It is important that real-time transmission of life-critical data can be always guaranteed. To address this problem, a context-aware MAC protocol is presented in this paper. According to analysis of collected life parameters, the protocol can switch between normal state and emergency state. As a result, data rate and duty cycle of sensor nodes are dynamically changed to meet the requirement of latency and traffic loads in a contexta-ware way. To save the power consumption, a TDMA-based MAC frame structure is used. Moreover, a novel optional synchronization scheme is proposed to decrease the overhead caused by traditional TDMA synchronization scheme. Simulation results show significant improvements of our design on latency and power consumption. Zhisheng Yan, Bin Liu 0016 |
IWCMC | 1 |