VLDB 2026 Research / reviewers in the wild / expert
Nan Wu 0012
dblp:58/2484-12
· DBLP profile ↗
13ranked-venue papers
4as first author
13since 2021 · last 2026
0000-0003-4875-2941ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 13 · 4 first-author · 13 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LITE: Loss-resilient Immersive Telepresence with Multi-modal SemanticsabstractImmersive telepresence has the potential to transform real-time communication through highly interactive and engaging experiences. Despite recent advances in reducing communication and computation costs, existing systems largely overlook packet loss, which can severely degrade the quality of experience (QoE). Recovering lost immersive content is considerably more challenging than in 2D video due to the complexity of dense 3D representations. Recovery must be both accurate and timely while minimizing the communication and computation overhead it incurs. To address these challenges, we present LITE, the first loss-resilient immersive telepresence system. LITE incorporates three key design principles: (1) leveraging semantic communication to transmit compact motion and audio semantics, which can be reconstructed into the remote user's immersive representation and voice, enabling fast semantic-level recovery and remaining robust to congestion-control-induced rate reductions under loss; (2) fusing audio and motion semantics via a lightweight multimodal model to achieve accurate, real-time recovery of motion semantics; and (3) encoding audio semantics from multiple past frames into succinct neural redundancy to enable robust recovery. We prototype LITE using a well-known parametric facial motion representation and extensively evaluate its performance across diverse networks. Our results demonstrate that LITE improves QoE by up to 109% compared with existing schemes, while sustaining real-time streaming at 30 frames per second and preserving high visual fidelity (structural similarity index measure above 0.9, where 1 indicates perfect similarity). Ruizhi Cheng, Harshvardhan C. Takawale, Nan Wu 0012, Nirupam Roy, Sennur Ulukus, Matteo Varvello, Eugene Chai, Bo Han 0001 |
SIGCOMM | 3 |
| 2025 | From WebGL to WebGPU: A Reality Check of Browser-Based GPU AccelerationabstractWith the rising demand for cost-effective and privacy-preserving deep learning and visualization services, service providers are increasingly turning to in-browser solutions. General-purpose computations, leveraging the graphics processing unit (GPU), are foundational to executing algorithms that power these services. Web graphics library (WebGL) is a widely adopted GPU-access application programming interface (API) designed for multidimensional rendering in browsers, while WebGPU is a newer API developed with compute-specific capabilities. Although WebGPU is a promising standard, its performance has not been systematically evaluated for general-purpose computation. This paper investigates WebGPU and WebGL for accelerating client-side computation in web browsers. By benchmarking key computational GPU kernels for 16 PolyBench and 2 CHStone functions, we measure the performance of WebGPU and WebGL across varying input sizes and algorithmic complexities. Our results show that: 1) both WebGPU and WebGL exhibit poorer performance than central processing unit (CPU)-based execution for small input data due to setup and CPU-GPU synchronization overheads, but they outperform CPU execution as the input data size increases; 2) WebGL performs better than WebGPU for small inputs, except for CPU-driven loop functions, due to its lower initial setup overhead; and 3) WebGPU outperforms WebGL for large inputs through optimized GPU thread utilization and achieves better performance for loop-driven algorithms across all input sizes by minimizing CPU-GPU data exchange. Overall, our results indicate that WebGPU is a competitive option for enhancing the execution performance of large-scale web applications. Sthitadhi Sengupta, Nan Wu 0012, Matteo Varvello, Krish Jana, Songqing Chen, Bo Han 0001 |
IMC | 2 |
| 2025 | NeVo: Advancing Volumetric Video Streaming with Neural Content RepresentationabstractOffering high-quality immersive content is the ultimate goal of volumetric video streaming. Although point clouds and meshes are dominant volumetric representations, their limitations in depicting photo-realistic content often undermine user experience. The recent advent of neural radiance fields (NeRF) offers a promising alternative content representation with superior photo-realism. However, streaming NeRF-based volumetric videos over wireless networks to mobile headsets faces significant challenges, including substantial bandwidth usage because of the large frame size, degraded visual quality due to even a low packet loss rate, and content artifacts caused by performance optimizations (e.g., remote rendering at the network edge). To address these challenges, in this paper, we introduce NeVo, a next-generation volumetric video streaming system for efficient delivery of neural content such as NeRF. NeVo incorporates the following innovations into a holistic system: (1) a novel method to model visibility of implicitly encoded neural content, thereby avoiding non-essential transmission to drastically reduce network data usage, (2) a lightweight, learning-based model for real-time content reconstruction after packet loss with carefully chosen data, and (3) judicious identification and selective delivery of intermediate data in edge-based NeRF rendering to effectively mitigate artifacts. Our extensive experiments indicate that compared with the state-of-the-art, NeVo saves up to 68.3% of bandwidth usage, maintains high visual quality despite packet loss, and enhances user experience by reducing artifacts. Nan Wu 0012, Bo Chen 0025, Ruizhi Cheng, Klara Nahrstedt, Bo Han 0001 |
MobiCom | 1 |
| 2025 | PIPE: Privacy-preserving 6DoF Pose Estimation for Immersive ApplicationsabstractImage-based mapping and localization offer six degrees of freedom (6DoF) pose estimation for immersive applications. This is achieved by matching, on a server, 2D visual features extracted from a mobile device's camera view and 3D features stored in a map. While effective, this process may lead to privacy breaches (e.g., exposure of sensitive information captured by camera views). To tackle this crucial issue, we present PIPE, a first-of-its-kind Privacy-preserving Image-based 6DoF Pose Estimation system. The design of PIPE is motivated by our key observation that uploading only a small amount of features extracted from camera views for pose estimation could reduce privacy leakage. However, trade-offs exist between privacy preservation, system utility (i.e., pose estimation accuracy), and system performance (e.g., end-to-end latency). To balance the trade-offs, PIPE deliberately explores the feature-detection space to reduce computation latency, designs an efficient feature ranking method by judiciously utilizing map data, and optimizes feature selection by jointly considering the features' ranking and spatial distribution to improve pose estimation accuracy. Moreover, we construct a learning-based metric to quantify the extent of privacy leakage in images. Our extensive performance evaluation reveals that PIPE can effectively preserve privacy and reduce end-to-end latency by up to 22.6%, while marginally affecting pose estimation accuracy (e.g., as low as 2.7%). Nan Wu 0012, Ruizhi Cheng, Songqing Chen, Bo Han 0001 |
SenSys | 1 |
| 2024 | A First Look at Immersive Telepresence on Apple Vision ProabstractDue to the widespread adoption of "work-from-home" policies, videoconferencing applications (e.g., Zoom) have become indispensable for remote communication. However, they often lack immersiveness, leading to "Zoom fatigue" and degrading communication efficiency. The recent debut of Apple Vision Pro, a mobile headset that supports "spatial personas", offers an immersive telepresence experience. In this paper, we conduct a first-of-its-kind in-depth and empirical study to analyze the performance of immersive telepresence with FaceTime, Webex, Teams, and Zoom on Vision Pro. We find that only FaceTime provides a truly immersive experience with spatial personas, whereas others still operate 2D personas. Our measurements reveal that (1) FaceTime delivers semantic data to optimize bandwidth consumption, which is even lower than that of 2D personas for other applications, and (2) it employs visibility-aware optimizations to reduce rendering overhead. However, the scalability of FaceTime remains limited, with a simple server-allocation strategy that potentially leads to high network delay for users. Ruizhi Cheng, Nan Wu 0012, Matteo Varvello, Eugene Chai, Songqing Chen, Bo Han 0001 |
IMC | 2 |
| 2024 | Theia: Gaze-driven and Perception-aware Volumetric Content Delivery for Mixed Reality HeadsetsabstractMinimizing bandwidth consumption while maintaining satisfactory visual quality becomes the holy grail of volumetric content delivery. However, due to the huge amount of 3D data to stream, the stringent latency requirement, and the high computational workload, achieving this ambitious goal could be challenging for mobile mixed reality headsets, which can naturally enable viewers' motion with six degrees of freedom but have limited computing power. Motivated by our critical observations from a user study of eye movements with 50+ participants, in this paper, we present Theia, a first-of-its-kind gaze-driven and perception-aware volumetric content delivery system that effectively incorporates the following innovations into a holistic system: (1) real-time creation of foveated volumetric content to reduce network data usage; (2) efficient augmentation of foveal content to boost user experience; and (3) adaptive omission of peripheral content for further bandwidth savings based on eye movements. We implement a prototype of Theia using Microsoft HoloLens 2 headsets and extensively evaluate its performance. Our results reveal that compared to the state-of-the-art, Theia can drastically reduce bandwidth consumption by up to 67.0% and enhance visual quality by up to 92.5%. Nan Wu 0012, Kaiyan Liu, Ruizhi Cheng, Bo Han 0001, Puqi Zhou |
MobiSys | 1 |
| 2024 | MagicStream: Bandwidth-conserving Immersive Telepresence via Semantic CommunicationabstractImmersive telepresence has the potential to revolutionize remote communication by offering a highly interactive and engaging user experience. However, state-of-the-art exchanges large volumes of 3D content to achieve satisfactory visual quality, resulting in substantial Internet bandwidth consumption. To tackle this challenge, we introduce MagicStream, a first-of-its-kind semantic-driven immersive telepresence system that effectively extracts and delivers compact semantic details of captured 3D representation of users, instead of traditional bit-by-bit communication of raw content. To minimize bandwidth consumption while maintaining low end-to-end latency and high visual quality, MagicStream incorporates the following key innovations: (1) efficient extraction of user's skin/cloth color and motion semantics based on lighting characteristics and body keypoints, respectively; (2) novel, real-time human body reconstruction from motion semantics; and (3) on-the-fly neural rendering of users' immersive representation with color semantics. We implement a prototype of MagicStream and extensively evaluate its performance through both controlled experiments and user trials. Our results show that, compared to existing schemes, MagicStream can drastically reduce Internet bandwidth usage by up to 1195X while maintaining good visual quality. Ruizhi Cheng, Nan Wu 0012, Eugene Chai, Matteo Varvello, Bo Han 0001 |
SenSys | 2 |
| 2023 | Enriching Telepresence with Semantic-driven Holographic CommunicationabstractAchieving the optimal balance of minimizing bandwidth consumption and end-to-end latency while preserving a satisfactory level of visual quality becomes the ultimate goal of live, interactive holographic communication, a fundamental building block of immersive telepresence envisioned for 6G. Nevertheless, achieving this ambitious goal poses significant challenges for mobile devices with limited computing power, considering the substantial amount of 3D data to stream, the demanding latency requirements, and the high computation workload involved. Instead of distributing immersive content bit by bit, in this position paper, we propose to deliver semantic information extracted from telepresence participants to drastically reduce Internet bandwidth usage for task-oriented applications such as remote collaboration. We contribute a taxonomy by categorizing related semantics into three different types (i.e., keypoints, 2D images, and text), pinpoint the open research challenges associated with developing a practical system for each category in our comprehensive research agenda, and delve into the potential solutions for overcoming these challenges. The preliminary results from our proof-of-concept implementation that harnesses keypoint-based semantics (partially) validate the feasibility of our research agenda. Ruizhi Cheng, Kaiyan Liu, Nan Wu 0012, Bo Han 0001 |
HotNets | 3 |
| 2023 | Demystifying Web-based Mobile Extended Reality Accelerated by WebAssemblyabstractBy combining various emerging technologies, mobile extended reality (XR) blends the real world with virtual content to create a spectrum of immersive experiences. Although Web-based XR can offer attractive features such as better accessibility, cross-platform compatibility, and instant updates, its performance may not be on par with its standalone counterpart. As a low-level bytecode, WebAssembly has the potential to drastically accelerate Web-based XR by enabling near-native execution speed. However, little has been known about how well Web-based XR performs with WebAssembly acceleration. To bridge this crucial gap, we conduct a first-of-its-kind systematic and empirical study to analyze the performance of Web-based XR expedited by WebAssembly on four diverse platforms with five different browsers. Our measurement results reveal that although WebAssemlby can accelerate different XR tasks in various contexts, there remains a substantial performance disparity between Web-based and standalone XR. We hope our findings can foster the realization of an immersive Web that is accessible to a wider audience with various emerging technologies. Kaiyan Liu, Nan Wu 0012, Bo Han 0001 |
IMC | 2 |
| 2023 | MetaStream: Live Volumetric Content Capture, Creation, Delivery, and Rendering in Real TimeabstractWhile recent work explored streaming volumetric content on-demand, there is little effort on live volumetric video streaming that bears the potential of bringing more exciting applications than its on-demand counterpart. To fill this critical gap, in this paper, we propose MetaStream, which is, to the best of our knowledge, the first practical live volumetric content capture, creation, delivery, and rendering system for immersive applications such as virtual, augmented, and mixed reality. To address the key challenge of the stringent latency requirement for processing and streaming a huge amount of 3D data, MetaStream integrates several innovations into a holistic system, including dynamic camera calibration, edge-assisted object segmentation, cross-camera redundant point removal, and foveated volumetric content rendering. We implement a prototype of MetaStream using commodity devices and extensively evaluate its performance. Our results demonstrate that MetaStream achieves low-latency live volumetric video streaming at close to 30 frames per second on WiFi networks. Compared to state-of-the-art systems, MetaStream reduces end-to-end latency by up to 31.7% while improving visual quality by up to 12.5%. Yongjie Guan, Xueyu Hou, Nan Wu 0012, Bo Han 0001, Tao Han 0002 |
MobiCom | 3 |
| 2022 | Are we ready for metaverse?: a measurement study of social virtual reality platformsabstractSocial virtual reality (VR) has the potential to gradually replace traditional online social media, thanks to recent advances in consumer-grade VR devices and VR technology itself. As the vital foundation for building the Metaverse, social VR has been extensively examined by the computer graphics and HCI communities. However, there has been little systematic study dissecting the network performance of social VR, other than hype in the industry. To fill this critical gap, we conduct an in-depth measurement study of five popular social VR platforms: AltspaceVR, Horizon Worlds, Mozilla Hubs, Rec Room, and VRChat. Our experimental results reveal that all these platforms are still in their early stage and face fundamental technical challenges to realize the grand vision of Metaverse. For example, their throughput, end-to-end latency, and on-device computation resource utilization increase almost linearly with the number of users, leading to potential scalability issues. We identify the platform servers' direct forwarding of avatar data for embodying users without further processing as the main reason for the poor scalability and discuss potential solutions to address this problem. Moreover, while the visual quality of the current avatar embodiment is low and fails to provide a truly immersive experience, improving the avatar embodiment will consume more network bandwidth and further increase computation overhead and latency, making the scalability issues even more pressing. Ruizhi Cheng, Nan Wu 0012, Matteo Varvello, Songqing Chen, Bo Han 0001 |
IMC | 2 |
| 2022 | DeepMix: mobility-aware, lightweight, and hybrid 3D object detection for headsetsabstractMobile headsets should be capable of understanding 3D physical environments to offer a truly immersive experience for augmented/mixed reality (AR/MR). However, their small form-factor and limited computation resources make it extremely challenging to execute in real-time 3D vision algorithms, which are known to be more compute-intensive than their 2D counterparts. In this paper, we propose DeepMix, a mobility-aware, lightweight, and hybrid 3D object detection framework for improving the user experience of AR/MR on mobile headsets. Motivated by our analysis and evaluation of state-of-the-art 3D object detection models, DeepMix intelligently combines edge-assisted 2D object detection and novel, on-device 3D bounding box estimations that leverage depth data captured by headsets. This leads to low end-to-end latency and significantly boosts detection accuracy in mobile scenarios. A unique feature of DeepMix is that it fully exploits the mobility of headsets to fine-tune detection results and boost detection accuracy. To the best of our knowledge, DeepMix is the first 3D object detection that achieves 30 FPS (i.e., an end-to-end latency much lower than the 100 ms stringent requirement of interactive AR/MR). We implement a prototype of DeepMix on Microsoft HoloLens and evaluate its performance via both extensive controlled experiments and a user study with 30+ participants. DeepMix not only improves detection accuracy by 9.1--37.3% but also reduces end-to-end latency by 2.68--9.15×, compared to the baseline that uses existing 3D object detection models. Yongjie Guan, Xueyu Hou, Nan Wu 0012, Bo Han 0001, Tao Han 0002 |
MobiSys | 3 |
| 2022 | Preserving privacy in mobile spatial computingabstractMapping and localization are the key components in mobile spatial computing to facilitate interactions between users and the digital model of the physical world. To enable localization, mobile devices keep capturing images of the real-world surroundings and uploading them to a server with spatial maps for localization. This leads to privacy concerns on the potential leakage of sensitive information in both spatial maps and localization images (e.g., when used in confidential industrial settings or our homes). Motivated by the above issues, we present a holistic research agenda in this paper for designing principled approaches to preserve privacy in spatial mapping and localization. We introduce our ongoing research, including learning-assisted noise generation to shield spatial maps, distributed architecture with intelligent aggregation to protect localization images, and end-to-end privacy preservation with fully homomorphic encryption. We also discuss the technical challenges, our preliminary results, and open research problems in those areas. Nan Wu 0012, Ruizhi Cheng, Songqing Chen, Bo Han 0001 |
NOSSDAV | 1 |