EDBT 2026 Demo / reviewers in the wild / expert
Ruizhi Cheng
dblp:312/6541
· DBLP profile ↗
15ranked-venue papers
9as first author
15since 2021 · last 2026
0000-0002-3942-5167ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 12 · 7 first-author · 12 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LITE: Loss-resilient Immersive Telepresence with Multi-modal SemanticsabstractImmersive telepresence has the potential to transform real-time communication through highly interactive and engaging experiences. Despite recent advances in reducing communication and computation costs, existing systems largely overlook packet loss, which can severely degrade the quality of experience (QoE). Recovering lost immersive content is considerably more challenging than in 2D video due to the complexity of dense 3D representations. Recovery must be both accurate and timely while minimizing the communication and computation overhead it incurs. To address these challenges, we present LITE, the first loss-resilient immersive telepresence system. LITE incorporates three key design principles: (1) leveraging semantic communication to transmit compact motion and audio semantics, which can be reconstructed into the remote user's immersive representation and voice, enabling fast semantic-level recovery and remaining robust to congestion-control-induced rate reductions under loss; (2) fusing audio and motion semantics via a lightweight multimodal model to achieve accurate, real-time recovery of motion semantics; and (3) encoding audio semantics from multiple past frames into succinct neural redundancy to enable robust recovery. We prototype LITE using a well-known parametric facial motion representation and extensively evaluate its performance across diverse networks. Our results demonstrate that LITE improves QoE by up to 109% compared with existing schemes, while sustaining real-time streaming at 30 frames per second and preserving high visual fidelity (structural similarity index measure above 0.9, where 1 indicates perfect similarity). Ruizhi Cheng, Harshvardhan C. Takawale, Nan Wu 0012, Nirupam Roy, Sennur Ulukus, Matteo Varvello, Eugene Chai, Bo Han 0001 |
SIGCOMM | 1 |
| 2025 | Hello, GenAI? Dissecting Human to Generative AI CallingabstractThe rise of generative artificial intelligence (GenAI), powered by large language models, has led to the emergence of real-time, voice-based conversational applications that enable dynamic, multi-modal interactions for everyday tasks such as checking the weather or planning a trip. These human-to-GenAI calling applications blend speech processing, generative intelligence, and real-time communication, presenting new challenges in latency optimization, network infrastructure design, and resilience under load. Despite their growing popularity, little is known about the operational characteristics and performance of these applications. This paper conducts an empirical measurement of six human-to-GenAI calling applications from Google, Meta, Microsoft, and OpenAI, focusing on their input/output modalities, network behavior, latency metrics, and robustness. Our findings reveal key design choices and performance bottlenecks in these emerging applications. For example, the conversational latency often reaches several seconds, far exceeding the typical sub-second delays of human-to-human voice communication and potentially impairing interactivity. Moreover, voice-based GenAI traffic is inherently asymmetric: the uplink, carrying real-time human speech, benefits from streaming-based transmission, while the typically large downlink GenAI responses are better served through batch-based delivery. Ruizhi Cheng, Surendra Pathak, Guowu Xie, Matteo Varvello, Songqing Chen, Bo Han 0001 |
IMC | 1 |
| 2025 | The Decentralization Dilemma: Performance Trade-Offs in IPFS and BreakpointsabstractWeb 3.0 is redefining the current Web (Web 2.0) with a focus on data and governance decentralization. The InterPlanetary File System (IPFS) exemplifies this shift. However, it faces a trade-off between decentralization and performance: prior studies have shown IPFS's performance degradations but fail to diagnose root causes or deliver actionable fixes. Ruizhe Shi, Yuqi Fu, Ruizhi Cheng, Bo Han 0001, Yue Cheng 0001, Songqing Chen |
IMC | 3 |
| 2025 | NeVo: Advancing Volumetric Video Streaming with Neural Content RepresentationabstractOffering high-quality immersive content is the ultimate goal of volumetric video streaming. Although point clouds and meshes are dominant volumetric representations, their limitations in depicting photo-realistic content often undermine user experience. The recent advent of neural radiance fields (NeRF) offers a promising alternative content representation with superior photo-realism. However, streaming NeRF-based volumetric videos over wireless networks to mobile headsets faces significant challenges, including substantial bandwidth usage because of the large frame size, degraded visual quality due to even a low packet loss rate, and content artifacts caused by performance optimizations (e.g., remote rendering at the network edge). To address these challenges, in this paper, we introduce NeVo, a next-generation volumetric video streaming system for efficient delivery of neural content such as NeRF. NeVo incorporates the following innovations into a holistic system: (1) a novel method to model visibility of implicitly encoded neural content, thereby avoiding non-essential transmission to drastically reduce network data usage, (2) a lightweight, learning-based model for real-time content reconstruction after packet loss with carefully chosen data, and (3) judicious identification and selective delivery of intermediate data in edge-based NeRF rendering to effectively mitigate artifacts. Our extensive experiments indicate that compared with the state-of-the-art, NeVo saves up to 68.3% of bandwidth usage, maintains high visual quality despite packet loss, and enhances user experience by reducing artifacts. Nan Wu 0012, Bo Chen 0025, Ruizhi Cheng, Klara Nahrstedt, Bo Han 0001 |
MobiCom | 3 |
| 2025 | PIPE: Privacy-preserving 6DoF Pose Estimation for Immersive ApplicationsabstractImage-based mapping and localization offer six degrees of freedom (6DoF) pose estimation for immersive applications. This is achieved by matching, on a server, 2D visual features extracted from a mobile device's camera view and 3D features stored in a map. While effective, this process may lead to privacy breaches (e.g., exposure of sensitive information captured by camera views). To tackle this crucial issue, we present PIPE, a first-of-its-kind Privacy-preserving Image-based 6DoF Pose Estimation system. The design of PIPE is motivated by our key observation that uploading only a small amount of features extracted from camera views for pose estimation could reduce privacy leakage. However, trade-offs exist between privacy preservation, system utility (i.e., pose estimation accuracy), and system performance (e.g., end-to-end latency). To balance the trade-offs, PIPE deliberately explores the feature-detection space to reduce computation latency, designs an efficient feature ranking method by judiciously utilizing map data, and optimizes feature selection by jointly considering the features' ranking and spatial distribution to improve pose estimation accuracy. Moreover, we construct a learning-based metric to quantify the extent of privacy leakage in images. Our extensive performance evaluation reveals that PIPE can effectively preserve privacy and reduce end-to-end latency by up to 22.6%, while marginally affecting pose estimation accuracy (e.g., as low as 2.7%). Nan Wu 0012, Ruizhi Cheng, Songqing Chen, Bo Han 0001 |
SenSys | 2 |
| 2025 | Centralization in the Decentralized Web: Challenges and Opportunities in IPFS Data ManagementabstractThe InterPlanetary File System (IPFS) is a pioneering effort for Web 3.0, well-known for its decentralized infrastructure. However, some recent studies have shown that IPFS exhibits a high degree of centralization and has integrated centralized components for improved performance. While this change contradicts the core decentralized ethos of IPFS and introduces risks of hurting the data replication level and thus availability, it also opens some opportunities for better data management and cost savings through deduplication. Ruizhe Shi, Ruizhi Cheng, Yuqi Fu, Bo Han 0001, Yue Cheng 0001, Songqing Chen |
WWW | 2 |
| 2025 | Dissecting User Experience of Social Virtual Reality: A Tale of Five PlatformsabstractSocial virtual reality (VR) has the potential to replace conventional online social media by offering quasi-real-world social experiences. As such, it has been extensively examined by the research community. However, existing studies fall short of providing a comprehensive understanding of how different aspects of social VR platforms interact to affect user experience. Motivated by this limitation, we conduct a user study with Oculus Quest 2 headsets and dissect the user experience on five social VR platforms. We evenly and randomly divide 42 participants into short-term (spending 10-30 minutes/platform) and long-term (spending at least 120 minutes/platform) groups. Besides employing surveys and interviews, we measure the frame rate and resolution of these platforms and explore how various factors interplay to influence the user experience of social VR. Our findings reveal that the frame rate, resolution, and interactive events of social VR platforms have a more significant impact on the experience of long-term users compared to short-term users. The scalability limitations of these platforms, as evidenced by decreased frame rates with the increasing number of concurrent users, result in an increased prevalence of motion sickness among long-term users, negatively impacting their overall experience. Moreover, the absence of highly interactive events also deteriorates their overall experience, and the low resolution combined with the lack of interactive events further decreases their sense of social presence. Additionally, our study demonstrates several common limitations negatively affecting the experience of both long-term and short-term users. For example, the harassment prevention mechanisms on all five platforms are inadequate, and being harassed has a detrimental effect on users' overall experience and sense of social presence. The avatar embodiment of investigated platforms has limited contribution to users' sense of social presence, mainly due to the lack of realism and full-body tracking. Our findings call for more research in scalability support, motion sickness relief, interactive event design, harassment prevention, and avatar development for improving social VR platforms in the future. Ruizhi Cheng, Jie Li 0064, Songqing Chen, Bo Han 0001 |
Proc. ACM Hum. Comput. Interact. | 1 |
| 2024 | A First Look at Immersive Telepresence on Apple Vision ProabstractDue to the widespread adoption of "work-from-home" policies, videoconferencing applications (e.g., Zoom) have become indispensable for remote communication. However, they often lack immersiveness, leading to "Zoom fatigue" and degrading communication efficiency. The recent debut of Apple Vision Pro, a mobile headset that supports "spatial personas", offers an immersive telepresence experience. In this paper, we conduct a first-of-its-kind in-depth and empirical study to analyze the performance of immersive telepresence with FaceTime, Webex, Teams, and Zoom on Vision Pro. We find that only FaceTime provides a truly immersive experience with spatial personas, whereas others still operate 2D personas. Our measurements reveal that (1) FaceTime delivers semantic data to optimize bandwidth consumption, which is even lower than that of 2D personas for other applications, and (2) it employs visibility-aware optimizations to reduce rendering overhead. However, the scalability of FaceTime remains limited, with a simple server-allocation strategy that potentially leads to high network delay for users. Ruizhi Cheng, Nan Wu 0012, Matteo Varvello, Eugene Chai, Songqing Chen, Bo Han 0001 |
IMC | 1 |
| 2024 | Theia: Gaze-driven and Perception-aware Volumetric Content Delivery for Mixed Reality HeadsetsabstractMinimizing bandwidth consumption while maintaining satisfactory visual quality becomes the holy grail of volumetric content delivery. However, due to the huge amount of 3D data to stream, the stringent latency requirement, and the high computational workload, achieving this ambitious goal could be challenging for mobile mixed reality headsets, which can naturally enable viewers' motion with six degrees of freedom but have limited computing power. Motivated by our critical observations from a user study of eye movements with 50+ participants, in this paper, we present Theia, a first-of-its-kind gaze-driven and perception-aware volumetric content delivery system that effectively incorporates the following innovations into a holistic system: (1) real-time creation of foveated volumetric content to reduce network data usage; (2) efficient augmentation of foveal content to boost user experience; and (3) adaptive omission of peripheral content for further bandwidth savings based on eye movements. We implement a prototype of Theia using Microsoft HoloLens 2 headsets and extensively evaluate its performance. Our results reveal that compared to the state-of-the-art, Theia can drastically reduce bandwidth consumption by up to 67.0% and enhance visual quality by up to 92.5%. Nan Wu 0012, Kaiyan Liu, Ruizhi Cheng, Bo Han 0001, Puqi Zhou |
MobiSys | 3 |
| 2024 | MetaFL: Privacy-preserving User Authentication in Virtual Reality with Federated LearningabstractThe increasing popularity of virtual reality (VR) has stressed the importance of authenticating VR users while preserving their privacy. Behavioral biometrics, owing to their robustness and ease of collection, compared to traditional modes such as passwords, have become a favored authentication choice. While current approaches that utilize behavioral biometrics to train classifiers for authentication yield promising accuracy, they cause privacy breaches by sharing sensitive data with a server to train a central model. In this paper, we present MetaFL, a first-of-its-kind privacy-preserving VR authentication framework that leverages federated learning (FL) on multi-modal motion data. The design of MetaFL is motivated by our key insight that various modalities of motion data uniquely affect authentication performance for individual users and among different users. It is attributed to the fundamental challenge of privacy-preserving user authentication: users can access only their own data with limited global knowledge. To tackle this issue, MetaFL judiciously selects the most suitable modalities for each user, which is decomposed into within-user ordering and between-user selection to eliminate the complex interplay between various conflicting factors. Moreover, we develop a personalized strategy to initialize FL models, further improving authentication accuracy. Our extensive performance evaluation on six public datasets shows that MetaFL outperforms state-of-the-art FL-based models (e.g., 17--28% higher authentication accuracy), and its accuracy gap with the non-privacy-preserving central model is small (i.e., only <2%). Ruizhi Cheng, Yuetong Wu, Ashish Kundu, Hugo Latapie, Myungjin Lee, Songqing Chen, Bo Han 0001 |
SenSys | 1 |
| 2024 | MagicStream: Bandwidth-conserving Immersive Telepresence via Semantic CommunicationabstractImmersive telepresence has the potential to revolutionize remote communication by offering a highly interactive and engaging user experience. However, state-of-the-art exchanges large volumes of 3D content to achieve satisfactory visual quality, resulting in substantial Internet bandwidth consumption. To tackle this challenge, we introduce MagicStream, a first-of-its-kind semantic-driven immersive telepresence system that effectively extracts and delivers compact semantic details of captured 3D representation of users, instead of traditional bit-by-bit communication of raw content. To minimize bandwidth consumption while maintaining low end-to-end latency and high visual quality, MagicStream incorporates the following key innovations: (1) efficient extraction of user's skin/cloth color and motion semantics based on lighting characteristics and body keypoints, respectively; (2) novel, real-time human body reconstruction from motion semantics; and (3) on-the-fly neural rendering of users' immersive representation with color semantics. We implement a prototype of MagicStream and extensively evaluate its performance through both controlled experiments and user trials. Our results show that, compared to existing schemes, MagicStream can drastically reduce Internet bandwidth usage by up to 1195X while maintaining good visual quality. Ruizhi Cheng, Nan Wu 0012, Eugene Chai, Matteo Varvello, Bo Han 0001 |
SenSys | 1 |
| 2024 | Understanding Online Education in Metaverse: Systems and User Experience PerspectivesabstractThanks to recent advances in immersive technologies, virtual reality (VR) is becoming increasingly popular in online education, particularly in light of the rise of the Metaverse. However, there is currently no in-depth investigation of the user experience of VR-based online education and the comparison of it with video-conferencing-based counterparts. To fill these critical gaps, we conduct multiple sessions of two courses in a university with 10 and 37 participants on Mozilla Hubs (Hubs for short), a social VR platform that is deemed as one of the early prototypes of the Metaverse, and let them compare the classroom experience on Hubs with Zoom, a popular video-conferencing application. In addition to employing traditional analytical methods to understand user experience, we benefit from an end-to-end measurement study of Hubs to corroborate our findings and systematically detect its performance bottlenecks. Our study leads to the following key observations. First, the scalability issue of Hubs makes it inadequate for accommodating large courses. Second, compared to Zoom, Hubs can offer a better sense of place presence and social presence to students, thanks to its avatar-based interactions and the hand and head tracking enabled by headsets. Third, even though VR headsets help students concentrate in class, effectively utilizing learning tools through them remains a challenge. Ruizhi Cheng, Erdem Murat, Lap-Fai Yu, Songqing Chen, Bo Han 0001 |
VR | 1 |
| 2023 | Enriching Telepresence with Semantic-driven Holographic CommunicationabstractAchieving the optimal balance of minimizing bandwidth consumption and end-to-end latency while preserving a satisfactory level of visual quality becomes the ultimate goal of live, interactive holographic communication, a fundamental building block of immersive telepresence envisioned for 6G. Nevertheless, achieving this ambitious goal poses significant challenges for mobile devices with limited computing power, considering the substantial amount of 3D data to stream, the demanding latency requirements, and the high computation workload involved. Instead of distributing immersive content bit by bit, in this position paper, we propose to deliver semantic information extracted from telepresence participants to drastically reduce Internet bandwidth usage for task-oriented applications such as remote collaboration. We contribute a taxonomy by categorizing related semantics into three different types (i.e., keypoints, 2D images, and text), pinpoint the open research challenges associated with developing a practical system for each category in our comprehensive research agenda, and delve into the potential solutions for overcoming these challenges. The preliminary results from our proof-of-concept implementation that harnesses keypoint-based semantics (partially) validate the feasibility of our research agenda. Ruizhi Cheng, Kaiyan Liu, Nan Wu 0012, Bo Han 0001 |
HotNets | 1 |
| 2022 | Are we ready for metaverse?: a measurement study of social virtual reality platformsabstractSocial virtual reality (VR) has the potential to gradually replace traditional online social media, thanks to recent advances in consumer-grade VR devices and VR technology itself. As the vital foundation for building the Metaverse, social VR has been extensively examined by the computer graphics and HCI communities. However, there has been little systematic study dissecting the network performance of social VR, other than hype in the industry. To fill this critical gap, we conduct an in-depth measurement study of five popular social VR platforms: AltspaceVR, Horizon Worlds, Mozilla Hubs, Rec Room, and VRChat. Our experimental results reveal that all these platforms are still in their early stage and face fundamental technical challenges to realize the grand vision of Metaverse. For example, their throughput, end-to-end latency, and on-device computation resource utilization increase almost linearly with the number of users, leading to potential scalability issues. We identify the platform servers' direct forwarding of avatar data for embodying users without further processing as the main reason for the poor scalability and discuss potential solutions to address this problem. Moreover, while the visual quality of the current avatar embodiment is low and fails to provide a truly immersive experience, improving the avatar embodiment will consume more network bandwidth and further increase computation overhead and latency, making the scalability issues even more pressing. Ruizhi Cheng, Nan Wu 0012, Matteo Varvello, Songqing Chen, Bo Han 0001 |
IMC | 1 |
| 2022 | Preserving privacy in mobile spatial computingabstractMapping and localization are the key components in mobile spatial computing to facilitate interactions between users and the digital model of the physical world. To enable localization, mobile devices keep capturing images of the real-world surroundings and uploading them to a server with spatial maps for localization. This leads to privacy concerns on the potential leakage of sensitive information in both spatial maps and localization images (e.g., when used in confidential industrial settings or our homes). Motivated by the above issues, we present a holistic research agenda in this paper for designing principled approaches to preserve privacy in spatial mapping and localization. We introduce our ongoing research, including learning-assisted noise generation to shield spatial maps, distributed architecture with intelligent aggregation to protect localization images, and end-to-end privacy preservation with fully homomorphic encryption. We also discuss the technical challenges, our preliminary results, and open research problems in those areas. Nan Wu 0012, Ruizhi Cheng, Songqing Chen, Bo Han 0001 |
NOSSDAV | 2 |