Xincheng Huang

dblp:214/8714 · DBLP profile ↗
← Back
11ranked-venue papers
4as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 5 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 One Body, Two Minds: Alternating VR Perspective During Remote Teleoperation of Supernumerary Limbs
abstract
Remote VR teleoperation with supernumerary robotic limbs enables distant users to operate in another’s local space. While a shared first-person view aids hand-eye coordination, locking the guest’s camera to the host’s head can degrade comfort, embodiment, and coordination. Based on a formative study (N=10) using a virtual supernumerary robotic limbs configuration to stress-test coordination, we propose guest-driven perspective switching from a shared first-person baseline (Shared Embodied View) to two alternatives: (a) a stabilized view with guest-controlled rotation (Embedded Anchored View), and (b) a fully decoupled third-person view (Out-of-body View). We ran a user study with 24 pairs (N=48), who switched between the baseline and proposed views as task demands changed. We measured performance, embodiment, fatigue, physiological arousal, and switching behaviors. Our results reveal role-dependent trade-offs: Out-of-body View improves navigation efficiency and reduces errors, while Embedded Anchored View supports embodiment. We conclude with guidelines: use Embedded Anchored View for hand-centric adjustments, Out-of-body View for navigation and object placement, and ensure smooth transitions.
Xincheng Huang, Winston Wijaya, Yi Fei Cheng 0001, David Lindlbauer, Eduardo Velloso, Andrea Bianchi, Zhanna Sarsenbayeva, Anusha Withana
CHI2
2026 Look2React: Making VR NPCs Come Alive with Dynamic Vision-Guided Reactions
abstract
A central promise of virtual reality (VR) games is the increased control players have over their character through pose and body language. However, many non-player character (NPC) systems fail to respond convincingly to these poses, user intent, and situational context, limiting immersion. We present Look2React, an interaction system that captures what NPCs see, using a vision-based reasoning model to select pose and text responses. Look2React endows NPCs with the ability to react dynamically and appropriately to player interactions. Through a gaze and proximity-based detection system inspired by stealth games, we trigger our system intuitively and only when intended, while also reducing resource costs. We invited 20 participants to play two versions of an RPG game: one with NPCs based on contemporary games and the other with Look2React NPCs. Our results demonstrate that Look2React increases engagement, leading to more frequent and repeated interactions with NPCs. Participants reported more satisfactory play sessions, significantly increased feelings of social presence, and felt that the dynamic reactions gave the NPCs more depth and personality - ultimately making them feel more human.
Ritik Vatsal, Xincheng Huang, Robert Xiao
IEEE Trans. Vis. Comput. Graph.2
2025 HaloTouch: Using IR Multi-Path Interference to Support Touch Interactions with General Surfaces
abstract
Sensing touch on arbitrary surfaces has long been a goal of ubiquitous computing, but often requires instrumenting the surface.Depth camera-based systems have emerged as a promising solution for minimizing instrumentation, but at the cost of high touch-down detection error rates, high touch latency, and high minimum hover distance, limiting them to basic tasks.We developed HaloTouch, a vision-based system which exploits a multipath interference effect from an off-the-shelf time-of-flight depth camera to enable fast, accurate touch interactions on general surfaces.HaloTouch achieves a 99.2% touch-down detection accuracy across various materials, with a motion-to-photon latency of 150 ms.With a brief (20s) userspecific calibration, HaloTouch supports millimeter-accurate hover sensing as well as continuous pressure sensing.We conducted a user study with 12 participants, including a typing task demonstrating text input at 26.3 AWPM.HaloTouch shows promise for more robust, dynamic touch interactions without instrumenting surfaces or adding hardware to users.
Ziyi Xia, Xincheng Huang, Sidney S. Fels, Robert Xiao
CHI2
2025 GaussianNexus: Room-Scale Real-Time AR/VR Telepresence with Gaussian Splatting
Xincheng Huang, Dieter Frehlich, Ziyi Xia, Peyman Gholami, Robert Xiao
UIST1
2025 TangiAR: Markerless Tangible Input for Immersive Augmented Reality with Everyday Objects
abstract
Tangible interactions with everyday objects have been shown to be fast, accurate, and natural, and have shown promise when combined with immersive augmented reality. However, implementing tangible controls presents considerable challenges. Previous works in the field either rely on additional tracking markers on objects, inadvertently shifting the difficulty to users, or are too computationally demanding for real-time operation on a head-mounted display (HMD). We propose TangiAR, a tangible control system which tracks everyday objects without the need for fiducial trackers, enabling them as passive controllers and virtual proxies in AR applications. TangiAR additionally enables hand and finger proximity interactions with tangibles, further expanding the interaction space. TangiAR can run on an unmodified Microsoft HoloLens 2, making it immediately practical. We evaluated the performance of TangiAR through a technical evaluation, including occlusion robustness and tracking accuracy tests, and a user study which examined the usability of our markerless object tracking system in various AR interactions.
Neil Xu Fan, Xincheng Huang, Robert Xiao
VRST2
2025 MultiSphere: Latency Optimized Multi-User 360° VR Telepresence with Edge-Assisted Viewport Adaptive IPv6 Multicast
abstract
360° video telepresence with VR enables immersive remote collaboration, but scaling to multiple users is subject to bandwidth and latency constraints. We present MultiSphere, a multi-user edge-assited 360° VR telepresence system, that combines viewport-adaptive IPv6 multicast tiling with a novel dual keyframe interval (KeyInt) streaming technique. Our approach addresses the latency bottleneck inherent in joining live streams of video using standard video codecs while maintaining visual quality through strategic use of low and high KeyInt streams. Our system achieves 75-94% bandwidth savings and an average request-to-decode latency of 56 ms, a 79% reduction compared to using a regular single-KeyInt stream.
Dieter Frehlich, Xincheng Huang, Robert Xiao
VRST2
2025 Adaptive nested sampling for multi-subregion feature learning
Xincheng Huang
Knowl. Based Syst.1
2025 VibRing: A Wearable Vibroacoustic Sensor for Single-Handed Gesture Recognition
abstract
Single-handed gestures offer rapid and intuitive interactions for input in interactive applications ranging from smartwatches and phones to augmented reality. Past research has explored using computer vision or inertial measurement units (IMUs) to sense such gestures, but these sensing modalities can be variously subject to occlusion, high power consumption, or sensitivity to random motion. In this work, we explore passively detecting the vibroacoustic signature of subtle single-handed gestures through a wearable piezoelectric sensor, providing a robust, low-power sensing modality. We present (1) a hand-gesture design framework encompassing a large set of subtle, rapid single-handed gestures which balance comfort and vibroacoustic distinguishability, (2) VibRing, a lightweight wireless hand-gesture sensing platform, leveraging a single finger-worn vibroacoustic sensor, and (3) a multifaceted system evaluation where we consider several aspects - general usability, tolerance to variance, user adaptability, and extended usage. Our results demonstrate that VibRing can support an 11-gesture set with a general accuracy of \(94.2\%\) and low-performance variance across multiple days ( \(90.2\%\) accuracy in cross-day validation). To support a new user, VibRing requires only 10 minutes of training data to achieve an accuracy of \(92.7\%\) . We also tested the extended use of VibRing in an office study where users performed periodic gesture inputs during typical office tasks with real-time classification, achieving a true-positive rate of \(90.9\%\) . Finally, to demonstrate the utility of VibRing, we present three examples of applications which benefit from our subtle gesture interactions.
Bu Li, Xincheng Huang, Robert Xiao
Proc. ACM Hum. Comput. Interact.2
2024 VirtualNexus: Enhancing 360-Degree Video AR/VR Collaboration with Environment Cutouts and Virtual Replicas
abstract
Asymmetric AR/VR collaboration systems bring a remote VR user to a local AR user’s physical environment, allowing them to communicate and work within a shared virtual/physical space. Such systems often display the remote environment through 3D reconstructions or 360° videos. While 360° cameras stream an environment in higher quality, they lack spatial information, making them less interactable. We present VirtualNexus, an AR/VR collaboration system that enhances 360° video AR/VR collaboration with environment cutouts and virtual replicas. VR users can define cutouts of the remote environment to interact with as a world-in-miniature, and their interactions are synchronized to the local AR perspective. Furthermore, AR users can rapidly scan and share 3D virtual replicas of physical objects using neural rendering. We demonstrated our system’s utility through 3 example applications and evaluated our system in a dyadic usability test. VirtualNexus extends the interaction space of 360° telepresence systems, offering improved physical presence, versatility, and clarity in interactions.
Xincheng Huang, Michael Yin, Ziyi Xia, Robert Xiao
UIST1
2023 Virtual Reality Telepresence: 360-Degree Video Streaming with Edge-Compute Assisted Static Foveated Compression
abstract
Real-time communication with immersive 360° video can enable users to be telepresent within a remotely streamed environment. Increasingly, users are shifting to mobile devices and connecting to the Internet via mobile-cellular networks. As the ideal media for 360° videos, some VR headsets now also come with cellular capacity, giving them potential for mobile applications. However, streaming high-quality 360° live video poses challenges for network bandwidth, particularly on cellular connections. To reduce bandwidth requirements, videos can be compressed using viewport-adaptive streaming or foveated rendering techniques. Such approaches require very low latency in order to be effective, which has previously limited their applications on traditional cellular networks. In this work, we demonstrate an end-to-end virtual reality telepresence system that streams ∼6K 360° video over 5G millimeter-wave (mmW) radio. Our use of 5G technologies, in conjunction with mobile edge compute nodes, substantially reduces latency when compared with existing 4G networks, enabling high-efficiency foveated compression over modern cellular networks on par with WiFi. We performed a technical evaluation of our system's visual quality post-compression with peak signal-to-noise ratio (PSNR) and FOVVideoVDP. We also conducted a user study to evaluate users' sensitivity to compressed video. Our findings demonstrate that our system achieves visually indistinguishable video streams while using up to 80% less data when compared with un-foveated video. We demonstrate our video compression system in the context of an immersive, telepresent video calling application.
Xincheng Huang, James Riddell, Robert Xiao
IEEE Trans. Vis. Comput. Graph.1
2022 Self-Adaptive Clustering of Dynamic Multi-Graph Learning
Yangding Li, Xincheng Huang, Jiaye Li 0001
Neural Process. Lett.3