Anh Nguyen 0011

dblp:52/5285-11 · DBLP profile ↗
← Back
10ranked-venue papers
3as first author
7since 2021 · last 2025
0000-0001-9858-5729ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 3 since 2021Computer networks · 4 · 4 since 2021Security and privacy · 3 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2025 Orbis: Redesigning Neural-enhanced Video Streaming for Live Immersive Viewing
abstract
Emerging live immersive viewing systems require streaming large 360 videos to users via limited wireless bandwidth. Neural-enhanced video streaming offers a promising solution by streaming down-scaled videos and enhancing them by client computation. However, prior systems treated video downscaling and enhancement separately, overlaying existing enhancement techniques onto current video infrastructure to accommodate legacy downscaling methods. This supplemental client design has led to spatial information loss and prohibitive model overheads in 360 video streaming systems. This paper presents Orbis, a redesigned, holistic neural-enhanced video streaming framework that integrates complementary down-scaling and enhancement for live immersive viewing. Orbis is empowered by an enhancement-driven interleaved downscaling approach, an inpainting-based enhancement model tailored to interleaved data, and a multi-scale tile adaptation scheme that optimizes immersive viewing experience in dynamic environments. Experimental results show that Orbis improves viewing experience by up to 60% and reduces wireless bandwidth by up to 49% compared to the best-performing baseline.
Zhengguan Wu, Jingwei Liao, Anh Nguyen 0011, Feng Lin 0004, Zhisheng Yan
SenSys3
2025 ST-360: Spatial-Temporal Filtering-Based Low-Latency 360-Degree Video Analytics Framework
abstract
Recent advances in computer vision algorithms and video streaming technologies have facilitated the development of edge-server-based video analytics systems, enabling them to process sophisticated real-world tasks, such as traffic surveillance and workspace monitoring. Meanwhile, due to their omnidirectional recording capability, 360-degree cameras have been proposed to replace traditional cameras in video analytics systems to offer enhanced situational awareness. Yet, we found that providing an efficient 360-degree video analytics framework is a non-trivial task. Due to the higher resolution and geometric distortion in 360-degree videos, existing video analytics pipelines fail to meet the performance requirements for end-to-end latency and query accuracy. To address these challenges, we introduce the innovative ST-360 framework specifically designed for 360-degree video analytics. This framework features a spatial–temporal filtering algorithm that optimizes both data transmission and computational workloads. Evaluation of the ST-360 framework on a unique dataset of 360-degree first-responders videos reveals that it yields accurate query results with a 50% reduction in end-to-end latency compared to state-of-the-art methods.
Jingwei Liao, Bo Chen 0025, Anh Nguyen 0011, Aditi Tiwari, Qian Zhou 0008, Zhisheng Yan, Klara Nahrstedt
ACM Trans. Multim. Comput. Commun. Appl.4
2025 A Patch Can Disrupt Live Video Streaming: Physical Adversarial Attacks on Deep Learning Compression
abstract
Deep learning (DL)-based compression has achieved outstanding performance compared to traditional compression. However, due to the vulnerability of adversarial attacks on DL models, understanding the security of DL-based compression is crucial. Previous attacks have demonstrated the feasibility of failing DL-based compression models in the digital domain. However, these attacks rely on perfect digital modification of the whole image and internal access to camera/server files, preventing their usage in the physical world where such assumptions do not hold. In this article, we unveil the first physical adversarial attack targeting DL-based compression in the context of live video streaming. The proposed attack, namely CamHack , places a small-sized physical patch in the visual scene covered by the live camera to manipulate the bitrate of compressed content and disrupt live streaming. The patch is crafted to address color and geometric transformations in diverse streaming scenes while remaining inconspicuous. Extensive experiments in various streaming scenes and network conditions show that CamHack increases bit consumption by 276.58%, 384.74%, and 942.28% over the clean scene without a patch on three representative victim models. CamHack is also robust under various practical impacts such as patch location, lighting conditions, and patch size.
Anh Nguyen 0011, Zhisheng Yan
ACM Trans. Multim. Comput. Commun. Appl.2
2024 Eavesdropping on Controller Acoustic Emanation for Keystroke Inference Attack in Virtual Reality
Shiqing Luo, Anh Nguyen 0011, Hafsa Farooq, Kun Sun 0001, Zhisheng Yan
NDSS2
2024 Penetration Vision through Virtual Reality Headsets: Identifying 360-degree Videos from Head Movements
Anh Nguyen 0011, Xiaokuan Zhang, Zhisheng Yan
USENIX Security Symposium1
2023 Latency-Aware 360-Degree Video Analytics Framework for First Responders Situational Awareness
abstract
First responders operate in hazardous working conditions with unpredictable risks. To better prepare for demands of the job, first responder trainees conduct training exercises that are being recorded and reviewed by the instructors, who check for objects indicating risks within the video recordings (e.g., firefighter with an unfastened gas mask). However, the traditional reviewing process is inefficient due to unanalyzed video recordings and limited situational awareness. For better reviewing experience, a latency-aware Viewing and Query Service (VQS) should be provided. The VQS should support object searching, which can be achieved using the video object detection algorithms. Meanwhile, the application of 360-degree cameras facilitates an unlimited field of view of the training environment. Yet, this medium represents a major challenge because low-latency high-accuracy 360-degree object detection is difficult due to higher resolution and geometric distortion. In this paper, we present the Responders-360 system architecture designed for 360-degree object detection. We propose a Dynamic Selection algorithm that optimizes computation resources while yielding accurate 360-degree object inference. The results, using a unique dataset collected from a firefighting training institute, show that the Responders-360 framework achieves 4x speedup and 25% memory usage reduction compared with the state-of-the-art methods.
Jingwei Liao, Bo Chen 0025, Anh Nguyen 0011, Aditi Tiwari, Qian Zhou 0008, Zhisheng Yan, Klara Nahrstedt
NOSSDAV4
2022 DAO: Dynamic Adaptive Offloading for Video Analytics
abstract
Offloading videos from end devices to edge or cloud servers is the key to enabling computation-intensive video analytics. To ensure the analytics accuracy at the server, the video quality for offloading must be configured based on the specific content and the available network bandwidth. While adaptive video streaming for user viewing has been widely studied, none of the existing works can guarantee the analytics accuracy at the server in bandwidth- and content-adaptive way. To fill in this gap, this paper presents DAO, a dynamic adaptive offloading framework for video analytics that jointly considers the dynamics of network bandwidth and video content. DAO is able to maximize the analytics accuracy at the server by adapting the video bitrate and resolution dynamically. In essence, we shift the context of adaptive video transport from traditional DASH systems to a new dynamic adaptive offloading framework tailored for video analytics. DAO is empowered by some new discoveries about the inherent relationship between analytics accuracy, video content, bitrate, and resolution, as well as by an optimization formulation to adapt the bitrate and resolution dynamically. Results from the real-world implementation of object detection tasks show that DAO's performance is close to the theoretical bound, achieving 20% bandwidth saving and 59% category-wise mAP improvement compared to conventional DASH schemes.
Taslim Murad, Anh Nguyen 0011, Zhisheng Yan
ACM Multimedia2
2020 OcuLock: Exploring Human Visual System for Authentication in Virtual Reality Head-mounted Display
Shiqing Luo, Anh Nguyen 0011, Chen Song 0001, Feng Lin 0004, Wenyao Xu, Zhisheng Yan
NDSS2
2019 A saliency dataset for 360-degree videos
abstract
Despite the increasing popularity, realizing 360-degree videos in everyday applications is still challenging. Considering the unique viewing behavior in head-mounted display (HMD), understanding the saliency of 360-degree videos becomes the key to various 360-degree video research. Unfortunately, existing saliency datasets are either irrelevant to 360-degree videos or too small to support saliency modeling. In this paper, we introduce a large saliency dataset for 360-degree videos with 50,654 saliency maps from 24 diverse videos. The dataset is created by a new methodology supported by psychology studies in HMD viewing. We describe an open-source software implementing this methodology that can generate saliency maps from any head tracking data. Evaluation of the dataset shows that the generated saliency is highly correlated with the actual user fixation and that the saliency data can provide useful insight on user attention in 360-degree video viewing. The dataset and the program used to extract saliency are both made publicly available to facilitate future research.
Anh Nguyen 0011, Zhisheng Yan
MMSys1
2018 Your Attention is Unique: Detecting 360-Degree Video Saliency in Head-Mounted Display for Head Movement Prediction
abstract
Head movement prediction is the key enabler for the emerging 360-degree videos since it can enhance both streaming and rendering efficiency. To achieve accurate head movement prediction, it becomes imperative to understand user's visual attention on 360-degree videos under head-mounted display (HMD). Despite the rich history of saliency detection research, we observe that traditional models are designed for regular images/videos fixed at a single viewport and would introduce problems such as central bias and multi-object confusion when applied to the multi-viewport 360-degree videos switched by user interaction. To fill in this gap, this paper shifts the traditional single-viewport saliency models that have been extensively studied for decades to a fresh panoramic saliency detection specifically tailored for 360-degree videos, and thus maximally enhances the head movement prediction performance. The proposed head movement prediction framework is empowered by a newly created dataset for 360-degree video saliency, a panoramic saliency detection model and an integration of saliency and head tracking history for the ultimate head movement prediction. Experimental results demonstrate the measurable gain of both the proposed panoramic saliency detection and head movement prediction over traditional models for regular images/videos.
Anh Nguyen 0011, Zhisheng Yan, Klara Nahrstedt
ACM Multimedia1