Bo Chen 0025

dblp:89/5615-25 · DBLP profile ↗
← Back
28ranked-venue papers
12as first author
23since 2021 · last 2026
0009-0006-3392-4834ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 14 · 6 first-author · 10 since 2021Computer networks · 13 · 6 first-author · 13 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 AquaScope: Reliable Underwater Image Transmission on Mobile Devices
abstract
Underwater communication is essential for both recreational and scientific activities, such as scuba diving. However, existing methods remain highly constrained by environmental challenges and often require specialized hardware, driving research into more accessible underwater communication solutions. While recent acoustic-based communication systems support text messaging on mobile devices, their low data rates severely limit broader applications. We present AquaScope, the first acoustic communication system capable of underwater image transmission on commodity mobile devices. To address the key challenges of underwater environments -- limited bandwidth and high transmission errors -- AquaScope employs and enhances generative image compression to improve compression efficiency, and integrates it with reliability-enhancement techniques at the physical layer to strengthen error resilience. We implemented AquaScope on the Android platform and demonstrated its feasibility for underwater image transmission. Experimental results show that AquaScope enables reliable, low-latency image transmission while preserving perceptual image quality, across various bandwidth-constrained and error-prone underwater conditions.
Beitong Tian, Bo Chen 0025, Mingyuan Wu, Haozhen Zheng, Deepak Vasisht, Francis Y. Yan, Klara Nahrstedt
MobiSys3
2025 Cache-of-Thought: Master-Apprentice Framework for Cost-Effective Vision Language Model Reasoning
abstract
Mingyuan Wu, Jize Jiang, Haozhen Zheng, Meitang Li, Zhaoheng Li, Beitong Tian, Bo Chen, Yongjoo Park, Minjia Zhang, ChengXiang Zhai, Klara Nahrstedt. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Mingyuan Wu, Jize Jiang, Haozhen Zheng, Meitang Li, Zhaoheng Li, Beitong Tian, Bo Chen 0025, Yongjoo Park, Minjia Zhang, ChengXiang Zhai, Klara Nahrstedt
EMNLP7
2025 Anywhere Avatar: 3D Telepresence with Just a Phone and a Laptop
abstract
We present Anywhere Avatar, a telepresence system that enables full-body and facial avatar reconstruction using a smartphone and a laptop. Users record short videos to generate personalized avatars, which are animated in real time during teleconferencing using webcam-based tracking. Built on pre-trained FLAME and SMPL models, the avatars are rendered in high fidelity using Gaussian splatting. The system runs at near real-time with minimal bandwidth, making expressive 3D telepresence accessible without specialized hardware.
Ruifan Ji, Mingyuan Wu, Bo Chen 0025, Michael Zink, Ramesh K. Sitaraman, Jacob Chakareski, Klara Nahrstedt
ACM Multimedia3
2025 NeVo: Advancing Volumetric Video Streaming with Neural Content Representation
abstract
Offering high-quality immersive content is the ultimate goal of volumetric video streaming. Although point clouds and meshes are dominant volumetric representations, their limitations in depicting photo-realistic content often undermine user experience. The recent advent of neural radiance fields (NeRF) offers a promising alternative content representation with superior photo-realism. However, streaming NeRF-based volumetric videos over wireless networks to mobile headsets faces significant challenges, including substantial bandwidth usage because of the large frame size, degraded visual quality due to even a low packet loss rate, and content artifacts caused by performance optimizations (e.g., remote rendering at the network edge). To address these challenges, in this paper, we introduce NeVo, a next-generation volumetric video streaming system for efficient delivery of neural content such as NeRF. NeVo incorporates the following innovations into a holistic system: (1) a novel method to model visibility of implicitly encoded neural content, thereby avoiding non-essential transmission to drastically reduce network data usage, (2) a lightweight, learning-based model for real-time content reconstruction after packet loss with carefully chosen data, and (3) judicious identification and selective delivery of intermediate data in edge-based NeRF rendering to effectively mitigate artifacts. Our extensive experiments indicate that compared with the state-of-the-art, NeVo saves up to 68.3% of bandwidth usage, maintains high visual quality despite packet loss, and enhances user experience by reducing artifacts.
Nan Wu 0012, Bo Chen 0025, Ruizhi Cheng, Klara Nahrstedt, Bo Han 0001
MobiCom2
2025 Intelligent Network Infrastructure for Extended Reality
abstract
This extended abstract outlines our research on designing intelligent network infrastructures for Extended Reality (XR). The research aims to leverage Artificial Intelligence (AI) in overcoming the limitations of traditional network infrastructures for XR in discrete content representation, handcrafted compression, and redundancy-based system resilience. To build a practical AI-based XR system, we introduce our AI-system co-design methodology, featuring system optimizations driven by AI models' in-depth measurements and AI algorithms inspired by the XR system's context analysis. We design and implement systems that showcase AI-system co-design advances photo-realism, efficiency, and resilience of network infrastructures for XR practically.
Bo Chen 0025
MobiSys1
2025 NeRFlow: Towards Adaptive Streaming for NeRF Videos
Rui-Xiao Zhang, Tianchi Huang, Bo Chen 0025, Klara Nahrstedt
MobiSys3
2025 EcoLens: Leveraging Multi-Objective Bayesian Optimization for Energy-Efficient Video Processing on Edge Devices
abstract
Video processing for real-time analytics in resource-constrained environments presents a significant challenge in balancing energy consumption and video semantics. This paper addresses the problem of energy-efficient video processing by proposing a system that dynamically optimizes processing configurations to minimize energy usage on the edge, while preserving essential video features for deep learning inference. We first gather an extensive offline profile of various configurations consisting of device CPU frequencies, frame filtering features, difference thresholds, and video bitrates, to establish apriori knowledge of their impact on energy consumption and inference accuracy. Leveraging this insight, we introduce an online system that employs multi-objective Bayesian optimization to intelligently explore and adapt configurations in real time. Our approach continuously refines processing settings to meet a target inference accuracy with minimal edge device energy expenditure. Experimental results demonstrate the system's effectiveness in reducing video processing energy use while maintaining high analytical performance, offering a practical solution for smart devices and edge computing applications.
Benjamin Civjan, Bo Chen 0025, Klara Nahrstedt
SMARTCOMP2
2025 ST-360: Spatial-Temporal Filtering-Based Low-Latency 360-Degree Video Analytics Framework
abstract
Recent advances in computer vision algorithms and video streaming technologies have facilitated the development of edge-server-based video analytics systems, enabling them to process sophisticated real-world tasks, such as traffic surveillance and workspace monitoring. Meanwhile, due to their omnidirectional recording capability, 360-degree cameras have been proposed to replace traditional cameras in video analytics systems to offer enhanced situational awareness. Yet, we found that providing an efficient 360-degree video analytics framework is a non-trivial task. Due to the higher resolution and geometric distortion in 360-degree videos, existing video analytics pipelines fail to meet the performance requirements for end-to-end latency and query accuracy. To address these challenges, we introduce the innovative ST-360 framework specifically designed for 360-degree video analytics. This framework features a spatial–temporal filtering algorithm that optimizes both data transmission and computational workloads. Evaluation of the ST-360 framework on a unique dataset of 360-degree first-responders videos reveals that it yields accurate query results with a 50% reduction in end-to-end latency compared to state-of-the-art methods.
Jingwei Liao, Bo Chen 0025, Anh Nguyen 0011, Aditi Tiwari, Qian Zhou 0008, Zhisheng Yan, Klara Nahrstedt
ACM Trans. Multim. Comput. Commun. Appl.3
2024 FedCore: Straggler-Free Federated Learning with Distributed Coresets
abstract
Federated learning (FL) is a machine learning paradigm that allows multiple clients to collaboratively train a shared model while keeping their data on-premise. However, the straggler issue, due to slow clients, often hinders the efficiency and scalability of FL. This paper presents FedCore, an algorithm that innovatively tackles the straggler problem via the decentralized selection of coresets, representative subsets of a dataset. Contrary to existing centralized coreset methods, FedCore creates coresets directly on each client in a distributed manner, ensuring privacy preservation in FL. FedCore translates the coreset optimization problem into a more tractable k-medoids clustering problem and operates distributedly on each client. Theoretical analysis confirms FedCore's convergence, and practical evaluations demonstrate an 8x reduction in FL training time, without compromising model accuracy. Our extensive evaluations also show that FedCore generalizes well to existing FL frameworks11Code: https://github.com/hongpeng-guo/PedCore.
Hongpeng Guo, Haotian Gu, Bo Chen 0025, Tamar Eilam, Deming Chen, Klara Nahrstedt
ICC4
2024 Scene Graph Driven Hybrid Interactive VR Teleconferencing
abstract
We propose an interactive and intelligent hybrid teleconferencing system compatible with Virtual Reality devices. Our system understands meeting contexts and leverages user interactions to enhance better system configuration. Employing interactive scene graphs [11], the system extracts and transmits essential meeting context to users while relaying user interactions back to the streaming systems for user-involved adaptive streaming and foveated rendering. We demonstrate the system's real-time performance and compatibility with commercial VR devices such as the Meta Quest 3.
Mingyuan Wu, Ruifan Ji, Haozhen Zheng, Beitong Tian, Bo Chen 0025, Jacob Chakareski, Michael Zink, Ramesh K. Sitaraman, Klara Nahrstedt
ACM Multimedia6
2024 Vesper: Learning to Manage Uncertainty in Video Streaming
abstract
Video codecs are crucial in video streaming systems. However, the quantization operation in existing codecs introduces irreversible jitters. Moreover, the common practice of fitting a single codec to diverse video content lacks the flexibility to adapt the parameters of a codec for specific content. They lead to the problem of quantization and content uncertainty. Our preliminary study shows an ideal codec without uncertainty gains a significant advantage over the conventional codec with uncertainty. However, realizing the ideal codec presents tremendous challenges in the generalizability and the costs of computation, transmission, and delay. In this paper, we present Vesper, a video streaming system that innovatively tackles uncertainty with two learning-based components, super-precision and self-evolution. The super-precision module builds a neural network that predicts original feature values from quantized feature values, which effectively mitigates the impact of quantization without inducing generalizability issues. The self-evolution module performs content-aware adaptation on the encoder and replaces non-content-aware video segments with content-aware ones on the fly, which addresses content uncertainty without adding significant costs to on-demand streaming. Evaluations demonstrate Vesper's superior Quality of Experience compared to streaming systems built with state-of-the-art codecs.
Bo Chen 0025, Mingyuan Wu, Hongpeng Guo, Zhisheng Yan, Klara Nahrstedt
MMSys1
2024 NeRFHub: A Context-Aware NeRF Serving Framework for Mobile Immersive Applications
abstract
Neural Radiance Fields (NeRF) are recognized for their exceptional photo-realism quality and superior modeling capabilities compared to traditional methods. NeRF empowers a novel application, termed NeRF serving. It delivers data from a server to a mobile client and renders 3D scenes on the client, facilitating a broad spectrum of mobile immersive applications. Towards a satisfactory user experience, we must serve NeRF with low latency while meeting constraints of high visual quality and real-time smoothness. Existing NeRF variants easily violate the constraints or cause an unnecessarily high latency when the diverse applications, mobile devices, and 3D scenes, termed the contexts, change in real life. In this paper, we present NeRFHub, a novel context-aware NeRF serving framework for mobile immersive applications. NeRFHub adeptly manages storage and computation costs, scales to diverse contexts, and swiftly navigates the vast design space inherent in NeRF serving. The evaluation results show that NeRFHub serves synthetic objects with 56%-66% reduced latency and realistic scenes with 26%-55% reduced latency when compared to the baseline without compromising quality or smoothness.
Bo Chen 0025, Zhisheng Yan, Bo Han 0001, Klara Nahrstedt
MobiSys1
2024 LiFteR: Unleash Learned Codecs in Video Streaming with Loose Frame Referencing
Bo Chen 0025, Zhisheng Yan, Yinjie Zhang, Zhe Yang 0010, Klara Nahrstedt
NSDI1
2024 ImmerScope: Multi-view Video Aggregation at Edge towards Immersive Content Services
abstract
The multi-camera capture system is an emerging visual sensing modality. It facilitates the production of various immersive contents ranging from regular to neural videos. Although the delivery of immersive content is popular and promising, it suffers from the bandwidth bottleneck when streaming multi-view videos to the cloud (i.e., multi-view video aggregation). Existing works fail to provide a bandwidth-efficient and content-generic solution. Even the closest effort to ours based on the SOTA multi-view video codecs suffers from issues of underutilized dependency and content distortion. In this paper, we present ImmerScope, a multi-view video aggregation framework at the edge with a neural multi-view video codec. It outperforms existing solutions with highly-utilized dependency via neuron connections and distortion awareness via end-to-end training. Evaluations on diverse multi-camera setups show that ImmerScope outperforms single-view codecs by at least 64% bandwidth savings in peak-signal-to-noise ratio with a frame rate of 50 fps.
Bo Chen 0025, Hongpeng Guo, Mingyuan Wu, Zhe Yang 0010, Zhisheng Yan, Klara Nahrstedt
SenSys1
2024 Context-aware Optimization for Bandwidth-Efficient Image Analytics Offloading
abstract
Convolutional Neural Networks (CNN) have given rise to numerous visual analytics applications at the edge of the Internet. The image is typically captured by cameras and then live-streamed to edge servers for analytics due to the prohibitive cost of running CNN on computation-constrained end devices. A critical component to ensure low-latency and accurate visual analytics offloading over low bandwidth networks is image compression which minimizes the amount of visual data to offload and maximizes the decoding quality of salient pixels for analytics. Despite the wide adoption, JPEG standards and traditional image compression techniques do not address the accuracy of analytics tasks, leading to ineffective compression for visual analytics offloading. Although recent machine-centric image compression techniques leverage sophisticated neural network models or hardware architecture to support the accuracy-bandwidth tradeoff, they introduce excessive latency in the visual analytics offloading pipeline. This article presents CICO, a Context-aware Image Compression Optimization framework to achieve low-bandwidth and low-latency visual analytics offloading. CICO contextualizes image compression for offloading by employing easily-computable low-level image features to understand the importance of different image regions for a visual analytics task. Accordingly, CICO can optimize the tradeoff between compression size and analytics accuracy. Extensive real-world experiments demonstrate that CICO reduces the bandwidth consumption of existing compression methods by up to 40% under comparable analytics accuracy. Regarding the low-latency support, CICO achieves up to a 2× speedup over state-of-the-art compression techniques.
Bo Chen 0025, Zhisheng Yan, Klara Nahrstedt
ACM Trans. Multim. Comput. Commun. Appl.1
2023 Interactive Scene Graph Analysis for Future Intelligent Teleconferencing Systems
abstract
In a real-life meeting environment, individuals often demonstrate a remarkable ability to selectively focus their attention on specific visual information. This ability allows them to naturally concentrate on a specific region of interest while tuning out others. Understanding and exploiting such selective attention remains unexplored in a user-centric teleconferencing system, where there is a potential to customize video streaming and foveated rendering based on the viewer’s attention. This paper proposes a novel user-centric scene analysis module that fully leverages the power of selective attention for online meeting scenarios and recognizes the unequal importance of individual pixels in the videos. The module determines the user’s selective attention through the meeting contexts. The contextual representation of the meeting is modeled as a combination of two primary components: proactive user interaction within the system and passive real-time analysis of high-level visual semantics from the scenes. As the meeting progresses, the interactive scene analysis module dynamically updates its contextual representation, offering a dual advantage: (a) Videos can be selectively and adaptively streamed within a user’s attention, resulting in bandwidth savings of up to 78 percent. (b) The module enhances the overall quality of the user experience by facilitating higher user interactivity, particularly in meeting-related tasks such as screen sharing, privacy-preserving user blocking, background removal, automatic user attention shift detection, etc. Our interactive scene analysis module makes significant progress toward enabling an efficient, immersive, and intelligent teleconferencing system.
Mingyuan Wu, Yuhan Lu, Shiv Trivedi, Bo Chen 0025, Qian Zhou 0008, Lingdong Wang, Simran Singh, Michael Zink, Ramesh K. Sitaraman, Jacob Chakareski, Klara Nahrstedt
ISM4
2023 SAVG360: Saliency-aware Viewport-guidance-enabled 360-video Streaming System
abstract
The emergence of 360-video streaming systems has brought about new possibilities for immersive video experiences while requiring significantly higher bandwidth than traditional 2D video streaming. Viewport prediction is used to address this problem, but interesting storylines outside the viewport are ignored. To address this limitation, we present SAVG360, a novel viewport guidance system that utilizes global content information available on the server side to enhance streaming with the best saliency-captured storyline of 360-videos. The saliency analysis is performed offline on the media server with powerful GPU, and the saliency-aware guidance information is encoded and shared with clients through the Saliency-aware Guidance Descriptor. This enables the system to proactively guide users to switch between storylines of the video and allow users to follow or break guided storylines through a novel user interface. Additionally, we present a viewing mode prediction algorithms to enhance video delivery in SAVG360. Evaluation of user viewport traces in 360-videos demonstrate that SAVG360 outperforms existing tiled streaming solutions in terms of overall viewport prediction accuracy and the ability to stream high-quality 360 videos under bandwidth constraints. Furthermore, a user study highlights the advantages of our proactive guidance approach over predicting and streaming of where users look.
Yinjie Zhang, Mingyuan Wu, Beitong Tian, Bo Chen 0025, Qian Zhou 0008, Klara Nahrstedt
ISM5
2023 Latency-Aware 360-Degree Video Analytics Framework for First Responders Situational Awareness
abstract
First responders operate in hazardous working conditions with unpredictable risks. To better prepare for demands of the job, first responder trainees conduct training exercises that are being recorded and reviewed by the instructors, who check for objects indicating risks within the video recordings (e.g., firefighter with an unfastened gas mask). However, the traditional reviewing process is inefficient due to unanalyzed video recordings and limited situational awareness. For better reviewing experience, a latency-aware Viewing and Query Service (VQS) should be provided. The VQS should support object searching, which can be achieved using the video object detection algorithms. Meanwhile, the application of 360-degree cameras facilitates an unlimited field of view of the training environment. Yet, this medium represents a major challenge because low-latency high-accuracy 360-degree object detection is difficult due to higher resolution and geometric distortion. In this paper, we present the Responders-360 system architecture designed for 360-degree object detection. We propose a Dynamic Selection algorithm that optimizes computation resources while yielding accurate 360-degree object inference. The results, using a unique dataset collected from a firefighting training institute, show that the Responders-360 framework achieves 4x speedup and 25% memory usage reduction compared with the state-of-the-art methods.
Jingwei Liao, Bo Chen 0025, Anh Nguyen 0011, Aditi Tiwari, Qian Zhou 0008, Zhisheng Yan, Klara Nahrstedt
NOSSDAV3
2022 Context-aware image compression optimization for visual analytics offloading
abstract
Convolutional Neural Networks (CNN) have given rise to numerous visual analytics applications at the edge of the Internet. The image is typically captured by cameras and then live-streamed to edge servers for analytics due to the prohibitive cost of running CNN on computation-constrained end devices. A critical component to ensure low-latency and accurate visual analytics offloading over low bandwidth networks is image compression that minimizes the amount of visual data to offload and maximizes the decoding quality of salient pixels for analytics. Despite the wide adoption, JPEG standard and traditional image compression do not address the accuracy of analytics tasks, leading to ineffective compression for visual analytics offloading. Although recent machine-centric image compression techniques leverage sophisticated neural network models or hardware architecture to support the accuracy-bandwidth trade-off, they introduce excessive latency in the visual analytics offloading pipeline. This paper presents CICO, a Context-aware Image Compression Optimization framework to achieve low-bandwidth and low-latency visual analytics offloading. CICO contextualizes image compression for offloading by employing easily-computable low-level image features to understand the importance of different image regions for a visual analytics task. Accordingly, CICO can optimize the trade-off between compression size and analytics accuracy. Extensive real-world experiments demonstrate that CICO reduces the bandwidth consumption of existing compression methods by up to 40% under a comparable analytics accuracy. In terms of the low-latency support, CICO achieves up to a 2x speedup over state-of-the-art compression techniques.
Bo Chen 0025, Zhisheng Yan, Klara Nahrstedt
MMSys1
2022 CAVE: caching 360° videos at the edge
abstract
While 360° videos are gaining popularity due to the emergence of VR technologies, storing and streaming such videos can incur up to 20X higher overheads than traditional HD content. Edge caching, which involves caching and serving 360° videos from edge servers, is one possible approach for addressing these overheads. Prior work on 360° video caching has been based on using past history to cache tiles that are likely to be in a viewer's field of view and has not considered methods to intelligently share a limited edge cache across a set of videos that exhibit large variations in their popularity, size, content, and user abandonment patterns. Towards this end, we present CAVE, an adaptive edge caching framework that intelligently optimizes cache allocation across a set of videos taking into account video content, size, and popularity. Our experiments using realistic video workloads shows CAVE improves cache hit-rates, and thus network saving, by up to 50% over state-of-the-art approaches, while also scaling to up to two thousand videos per edge cache. In addition, in terms of scalability, our developed algorithm is embarrassingly parallel, allowing CAVE to scale beyond state-of-the-art solutions that typically do not support parallelization.
Ahmed Ali-Eldin, Chirag Goel, Mayank Jha, Bo Chen 0025, Klara Nahrstedt, Prashant J. Shenoy
NOSSDAV4
2021 360ViewPET: View Based Pose EsTimation for Ultra-Sparse 360-Degree Cameras
abstract
Immersive virtual tours based on 360-degree cameras, showing famous outdoor scenery, are becoming more and more desirable due to travel costs, pandemics and other constraints. To feel immersive, a user must receive the view accurately corresponding to her position and orientation in the virtual space when she moves inside, and this requires cameras’ orientations to be known. Outdoor tour contexts have numerous, ultra-sparse cameras deployed across a wide area, making camera pose estimation challenging. As a result, pose estimation techniques like SLAM, which require mobile or dense cameras, are not applicable. In this paper we present a novel strategy called 360ViewPET, which automatically estimates the relative poses of two stationary, ultra-sparse (15 meters apart) 360-degree cameras using one equirectangular image taken by each camera. Our experiments show that it achieves accurate pose estimation, with a mean error as low as 0.9 degree.
Qian Zhou 0008, Bo Chen 0025, Zhe Yang 0010, Hongpeng Guo, Klara Nahrstedt
ISM2
2021 EScALation: a framework for efficient and scalable spatio-temporal action localization
abstract
Spatio-temporal action localization aims to detect the spatial location and the start/end time of the action in a video. The state-of-the-art approach uses convolutional neural networks to extract possible bounding boxes for the action in each frame and then link bounding boxes into action tubes based on the location and the class-specific score of each bounding box. Though this approach has been successful at achieving a good localization accuracy, it is computation-intensive. High-end GPUs are usually demanded for it to achieve real-time performance. In addition, this approach does not scale well on a large number of action classes. In this work, we present a framework, EScALation, for making spatio-temporal action localization efficient and scalable. Our framework involves two main strategies. One is the frame sampling technique that utilizes the temporal correlation between frames and selects key frame(s) from a temporally correlated set of frames to perform bounding box detection. The other is the class filtering technique that exploits bounding box information to predict the action class prior to linking bounding boxes. We compare EScALation with the state-of-the-art approach on UCF101-24 and J-HMDB-21 datasets. One of our experiments shows EScALation is able to save 72.2% of the time with only 6.1% loss of mAP. In addition, we show that EScALation scales better to a large number of action classes than the state-of-the-art approach.
Bo Chen 0025, Klara Nahrstedt
MMSys1
2021 Deep Contextualized Compressive Offloading for Images
abstract
Recent years have witnessed sensors becoming an indispensable part of our life with the camera being one of the most popular and widely deployed sensors. The camera gives rise to numerous vision-based IoT applications that generate high-level understandings of a live video stream by performing analysis on end devices like mobile or embedded devices. Typically, these applications are built with deep learning (DL) models to conduct complex vision tasks, e.g., image classification and object detection. Due to the prohibitive cost of running DL models on end devices close to the camera and with limited computation capabilities, it is widely adopted to offload the computation to a nearby powerful edge server. However, there is a gap between the restricted offloading bandwidth of the end device and the large volume of image data incurred by the live video stream. In this paper, we present Deep Contextualized Compressive Offloading for Images (DCCOI), a lightweight, context-aware, and bandwidth-efficient offloading framework for images. DCCOI consists of the spatial-adaptive encoder, a lightweight neural network, to spatial-adaptively compress the image, and the generative decoder for reconstructing the image from the compressed data. In contrast to existing DL-based encoders, the spatial-adaptive encoder allows an image region to be encoded into different numbers of feature values based on the information in it. This offers a variable-length coding method for image compression, which is a more optimal way for compression than the fix-length coding method took by existing DL-based compression approaches and demonstrates superior accuracy-compression rate trade-offs. We evaluate DCCOI against several baseline compression techniques while serving an object detection-based application. The results show that DCCOI roughly reduces the offloading size of JPEG by a factor of 9 and DeepCOD, the state-of-the-art offloading approach, by 20% with similar accuracy and a compression overhead less than 50ms.
Bo Chen 0025, Zhisheng Yan, Hongpeng Guo, Zhe Yang 0010, Ahmed Ali-Eldin, Prashant J. Shenoy, Klara Nahrstedt
SenSys1
2020 Real-time Spatio-Temporal Action Localization in 360 Videos
abstract
Spatio-temporal action localization of human actions in a video has been a popular topic over the past few years. It tries to localize the bounding boxes, the time span and the class of one action, which summarizes information in the video and helps humans understand it. Though many approaches have been proposed to solve this problem, these efforts have only focused on perspective videos. Unfortunately, perspective videos only cover a small field-of-view (FOV), which limits the capability of action localization. In this paper, we develop a comprehensive approach to real-time spatio-temporal localization that can be used to detect actions in 360 videos. We create two datasets named UCF-101-24-360 and JHMDB-21-360 for our evaluation. Our experiments show that our method consistently outperforms other competing approaches and achieves a real-time processing speed of 15fps for 360 videos.
Bo Chen 0025, Ahmed Ali-Eldin, Prashant J. Shenoy, Klara Nahrstedt
ISM1
2020 SEAWARE: Semantic Aware View Prediction System for 360-degree Video Streaming
abstract
Future view prediction for a 360-degree video streaming system is important to save the network bandwidth and improve the Quality of Experience (QoE). Historical view data of a single viewer and multiple viewers have been used for future view prediction. Video semantic information is also useful to predict the viewer's future behavior. However, extracting video semantic information requires powerful computing hardware and large memory space to perform deep learning-based video analysis. It is not a desirable condition for most of client devices, such as small mobile devices or Head Mounted Display (HMD). Therefore, we develop an approach where video semantic analysis is executed on the media server, and the analysis results are shared with clients via the Semantic Flow Descriptor (SFD) and View-Object State Machine (VOSM). SFD and VOSM become new descriptive additions of the Media Presentation Description (MPD) and Spatial Relation Description (SRD) to support 360-degree video streaming. Using the semantic-based approach, we design the Semantic-Aware View Prediction System (SEAWARE) to improve the overall view prediction performance. The evaluation results of 360-degree videos and real HMD view traces show that the SEAWARE system improves the view prediction performance and streams high-quality video with limited network bandwidth.
Jounsup Park, Mingyuan Wu, Kuan-Ying Lee, Bo Chen 0025, Klara Nahrstedt, Michael Zink, Ramesh K. Sitaraman
ISM4
2019 Event-driven stitching for tile-based live 360 video streaming
abstract
360 video streaming is gaining popularity because of the new type of experience it creates. Tile-based approaches have been widely used in VoD 360 video streaming to save the network bandwidth. However, they cannot be extended to the case of live streaming because they assume the 360 videos stitched offline before streaming. Instead, stitching has to be done in real-time in live 360 video streaming. More importantly, the stitching speed as shown in our experiments is one order of magnitude lower than the network transmission speed, making stitching more of a deciding factor of the overall frame rate than the network transmission speed. In this paper, we design a stitching algorithm for tile-based live 360 video streaming that adapts stitching quality to make the best use of the timing budget. There are two main challenges. First, existing tile-based approaches do not consider various semantic information in different scenarios. Second, the decision of tiling schemes for tile-based stitching is non-trivial. To solve the above two challenges, we present an event-driven stitching algorithm for tile-based 360 video live streaming, which consists of such an event-driven model to abstract various semantic information as events and a tile actuator to make tiling scheme decisions. We implement a streaming system based on event-driven stitching called LiveTexture. To evaluate the proposed algorithm, we compare LiveTexture with other baseline systems and show that LiveTexture adapts well to various timing budgets by meeting 89.4% of the timing constraints. We also demonstrate that LiveTexture utilizes the timing budget more efficiently than others.
Bo Chen 0025, Zhisheng Yan, Haiming Jin, Klara Nahrstedt
MMSys1
2018 ReSPonSe: Real-time, Secure, and Privacy-aware Video Redaction System
abstract
Nowadays the camera has developed into an indispensable and ubiquitous part of our life. It ensures the safety of people and their belongings, keeps records of special moments, or logs daily life. However, the ever-increasing amount of cameras surrounding us raised privacy concerns among people, who find themselves easily captured by a camera without themselves acknowledging it. To make matters worse, cameras, especially those on smart phones, are now more pervasive than ever before and can hardly be regulated as the recorders have full control of their cameras. Motivated by the privacy challenges originated from the ever-increasing and wide-spreading cameras, this paper presents the Real-time, Secure, and Privacy-aware Video Redaction System (ReSPonSe), which aims at protecting private information in personal videos according to permissions of people-in-video for other viewers to view them in the video. This system innovatively separates the production of videos into two stages: Encapsulation and Decapsulation. The first stage produces neutral videos in real-time while the second stage provides privacy-aware video to the viewer revealing private content of people-in-video who grants access rights to that viewer. The evaluation demonstrates the capability of this system to protect private information in videos with high efficiency and accuracy.
Bo Chen 0025, Klara Nahrstedt, Carl A. Gunter
MobiQuitous1
2017 Teleconsultant: Communication and Analysis of Wearable Videos in Emergency Medical Environments
abstract
Telehealth is a healthcare service that relies on exchanging information from one place to another to improve a patient's health status. In this demonstration, we aim to provide similar benefits to instantly bringing a doctor in the field to provide the right treatment at the right time for time-sensitive injuries. We present a telehealth system called Teleconsultant that enables near real-time communication between paramedics and doctors via videos captured from wearable cameras, this is crucial in the acute situations when the paramedic needs immediate assistance from the remote doctor that could help saving patients' lives. Teleconsultant includes capturing the video through body cameras worn by the paramedics, we refer to this video as wearable video. The video is transmitted over a heterogeneous wireless network to the remote doctor. Along the network path, video is analyzed in real-time to: (1) enhance video quality (e.g., video stabilization), and (2) detect time-sensitive injuries (e.g., stroke) so that remote doctors can be alerted and prepared when patient arrives via ambulance to the hospital. We demonstrate an end-to-end system to enable streaming of wearable video from the incident site to the hospital using body cameras worn by paramedics. Additionally, we demonstrate a framework for in-stream processing of the wearable video and we show two real-time video processing functions: stroke detection, and video stabilization.
Tarek Elgamal, Bo Chen 0025, Klara Nahrstedt
ACM Multimedia2