Robert Xiao

dblp:57/9306 · DBLP profile ↗
← Back
59ranked-venue papers
13as first author
30since 2021 · last 2026
0000-0003-4306-8825ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 51 · 12 first-author · 23 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 2 first-author · 10 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Re-Envisioning Instant Photography using Generative AI: An Exploratory Design Probe Using the UnReality Camera
abstract
Generative AI has increasingly been used for artistic creation, but little work has explored how it shapes the experiential meaning of practice. We consider how generative AI might transform the embodied and tangible process of instant photography through the UnReality Camera, an AI-mediated instant camera. The UnReality Camera prints a photo of the environment augmented by a user's spoken words as generative input. In a design probe, we explored how generative AI shapes people's perceptions of both photographic output and the broader photographic process. Although users valued artistic control, they also appreciated the creativity afforded by stochastic unpredictability. The waiting period for an unpredictable output elicited anticipatory suspense, and the camera's physical form evoked ownership and connection despite artificial generation. We discuss how people make sense of instant photography's experiential qualities when generative AI is embedded, and how their opposing affordances reshape interpretations of each other's experiential meaning.
Michael Yin, Angela Chiang, Robert Xiao
DIS3
2026 Roomify: Spatially-Grounded Style Transformation for Immersive Virtual Environments
abstract
We present Roomify, a spatially-grounded transformation system that generates themed virtual environments anchored to users’ physical rooms while maintaining spatial structure and functional semantics. Current VR approaches face a fundamental trade-off: full immersion sacrifices spatial awareness, while passthrough solutions break presence. Roomify addresses this through spatially-grounded transformation—treating physical spaces as “spatial containers” that preserve key functional and geometric properties of furniture while enabling radical stylistic changes. Our pipeline combines in-situ 3D scene understanding, AI-driven spatial reasoning, and style-aware generation to create personalized virtual environments grounded in physical reality. We introduce a cross-reality authoring tool enabling fine-grained user control through MR editing and VR preview workflows. Two user studies validate our approach: one with 18 VR users demonstrates a 63% improvement in presence over passthrough and 26% over fully virtual baselines while maintaining spatial awareness; another with 8 design professionals confirms the system’s creative expressiveness (scene quality: 5.95/7; creativity support: 6.08/7) and professional workflow value across diverse environments.
Qinxuan Cen, Weitao Bi, Yunxiang Ma, Xin Yi 0001, Robert Xiao, Xinyi Fu 0003, Hewu Li
CHI6
2026 The Words That Can't Be Shared: Exploring the Design of Unsent Messages
abstract
People often have things they want to say but hold back in conversations, fearing being vulnerable or facing social consequences. Online, this restraint can take a distinctive form: even when such thoughts are written out — in moments of anger, guilt, or longing — people may choose to withhold them, leaving them unsent. This process is underexamined; we investigate the experience of writing such messages within people’s digital communications. We find that unsent messages become expressive containers for suppressed feelings, where the act of writing creates a pause for reflection on the relationship and oneself. Building on these insights, we probed into how the design of the writing platforms of unsent messages affects people’s experiences and motivations. Speculating with participants on nine evocative variants of a note-taking platform, we highlight how design shapes the emotional, temporal, and ritualistic qualities of unsent messages, revealing tensions between people’s social desires and communicative actions.
Michael Yin, Robert Xiao
CHI2
2026 Reflective Motion and a Physical Canvas: Exploring Embodied Journaling in Virtual Reality
abstract
In traditional journaling practices, authors express and process their thoughts by writing them down. We propose a somaesthetic-inspired alternative that uses the human body, rather than written words, as the medium of expression. We coin this embodied journaling, as people’s isolated body movements and spoken words become the canvas of reflection. We implemented embodied journaling in virtual reality and conducted a within-subject user study (N = 20) to explore the emergent behaviours from the process, comparing its expressive and reflective qualities to those of written journaling. When writing-based norms and affordances were absent, we found that participants defaulted towards unfiltered emotional expression, often forgoing words altogether. Rather, subconscious body motion and paralinguistic acoustic qualities unveiled deeper, sometimes hidden feelings, prompting reflection that happens after emotional expression rather than during it. We discuss both the capabilities and pitfalls of embodied journaling, ultimately challenging the idea that reflection culminates in linguistic reasoning.
Michael Yin, Robert Xiao, Nadine Wagener
CHI2
2026 Dissolving a Digital Relationship: A Critical Examination of Digital Severance Behaviours in Close Relationships CSCW014
abstract
Fulfilling social connections are crucial for human well-being and belonging, but not all relationships last forever. As interactions increasingly move online, the act of digitally severing a relationship — e.g. through blocking or unfriending — has become progressively more common as well. This study considers actions of “digital severance” through interviews with 30 participants with experience as the initiator and/or recipient of such situations. Through a critical interpretative lens, we explore how people perceive and interpret their severance experience and how the online setting of social media shapes these dynamics. We develop themes that position digital severance as being intertwined with power and control, and we highlight (im)balances between an individual’s desires that can lead to feelings of disempowerment and ambiguous loss for both parties. We discuss the implications of our research, outlining three key tensions and four open questions regarding digital relationships, meaning-making, and design outcomes for future exploration.
Michael Yin, Angela Chiang, Robert Xiao
Proc. ACM Hum. Comput. Interact.3
2026 Look2React: Making VR NPCs Come Alive with Dynamic Vision-Guided Reactions
abstract
A central promise of virtual reality (VR) games is the increased control players have over their character through pose and body language. However, many non-player character (NPC) systems fail to respond convincingly to these poses, user intent, and situational context, limiting immersion. We present Look2React, an interaction system that captures what NPCs see, using a vision-based reasoning model to select pose and text responses. Look2React endows NPCs with the ability to react dynamically and appropriately to player interactions. Through a gaze and proximity-based detection system inspired by stealth games, we trigger our system intuitively and only when intended, while also reducing resource costs. We invited 20 participants to play two versions of an RPG game: one with NPCs based on contemporary games and the other with Look2React NPCs. Our results demonstrate that Look2React increases engagement, leading to more frequent and repeated interactions with NPCs. Participants reported more satisfactory play sessions, significantly increased feelings of social presence, and felt that the dynamic reactions gave the NPCs more depth and personality - ultimately making them feel more human.
Ritik Vatsal, Xincheng Huang, Robert Xiao
IEEE Trans. Vis. Comput. Graph.3
2025 PatternTrack: Multi-Device Tracking Using Infrared, Structured-Light Projections from Built-in LiDAR
Daehwa Kim, Robert Xiao, Chris Harrison 0001
CHI2
2025 HaloTouch: Using IR Multi-Path Interference to Support Touch Interactions with General Surfaces
abstract
Sensing touch on arbitrary surfaces has long been a goal of ubiquitous computing, but often requires instrumenting the surface.Depth camera-based systems have emerged as a promising solution for minimizing instrumentation, but at the cost of high touch-down detection error rates, high touch latency, and high minimum hover distance, limiting them to basic tasks.We developed HaloTouch, a vision-based system which exploits a multipath interference effect from an off-the-shelf time-of-flight depth camera to enable fast, accurate touch interactions on general surfaces.HaloTouch achieves a 99.2% touch-down detection accuracy across various materials, with a motion-to-photon latency of 150 ms.With a brief (20s) userspecific calibration, HaloTouch supports millimeter-accurate hover sensing as well as continuous pressure sensing.We conducted a user study with 12 participants, including a typing task demonstrating text input at 26.3 AWPM.HaloTouch shows promise for more robust, dynamic touch interactions without instrumenting surfaces or adding hardware to users.
Ziyi Xia, Xincheng Huang, Sidney S. Fels, Robert Xiao
CHI4
2025 TravelGalleria: Supporting Remembrance and Reflection of Travel Experiences through Digital Storytelling in Virtual Reality
Michael Yin, Robert Xiao
CHI2
2025 Streamlining Image Editing with Layered Diffusion Brushes
abstract
Denoising diffusion models have emerged as powerful tools for image manipulation, yet interactive, localized editing workflows remain underdeveloped. We introduce Layered Diffusion Brushes (LDB), a novel training-free framework that enables interactive, layer-based editing using standard diffusion models. LDB defines each "layer" as a self-contained set of parameters guiding the generative process, enabling independent, non-destructive, and fine-grained prompt-guided edits, even in overlapping regions. LDB leverages a unique intermediate latent caching approach to reduce each edit to only a few denoising steps, achieving 140~ms per edit on consumer GPUs. An editor implementing LDB, incorporating familiar layer concepts, was evaluated via user study and quantitative metrics. Results demonstrate LDB's superior speed alongside comparable or improved image quality, background preservation, and edit fidelity relative to state-of-the-art methods across various sequential image manipulation tasks. The findings highlight LDB's ability to significantly enhance creative workflows by providing an intuitive and efficient approach to diffusion-based image editing and its potential for expansion into related subdomains, such as video editing.
Peyman Gholami, Robert Xiao
ICCV2
2025 Entertainers Between Real and Virtual - Investigating Viewer Interaction, Engagement, and Relationships with Avatarized Virtual Livestreamers
abstract
Virtual YouTubers (VTubers) are avatar-based livestreamers that are voiced and played by human actors. VTubers have been popular in East Asia for years and have more recently seen widespread international growth. Despite their emergent popularity, research has been scarce into the interactions and relationships that exist between avatarized VTubers and their viewers, particularly in contrast to non-avatarized streamers. To address this gap, we performed in-depth interviews with self-reported VTuber viewers (n=21). Our findings first reveal that the avatarized nature of VTubers fosters new forms of theatrical engagement, as factors of the virtual blend with the real to create a mixture of fantasy and realism in possible livestream interactions. Avatarization furthermore results in a unique audience perception regarding the identity of VTubers - an identity which comprises a dynamic, distinct mix of the real human (the voice actor/actress) and the virtual character. Our findings suggest that each of these dual identities both individually and symbiotically affect viewer interactions and relationships with VTubers. Whereas the performer's identity mediates social factors such as intimacy, relatability, and authenticity, the virtual character's identity offers feelings of escapism, novelty in interactions, and a sense of continuity beyond the livestream. We situate our findings within existing livestreaming literature to highlight how avatarization drives unique, character-based interactions as well as reshapes the motivations and relationships that viewers form with livestreamers. Finally, we provide suggestions and recommendations for areas of future exploration to address the challenges involved in present livestreamed avatarized entertainment.
Michael Yin, Chenxinran Shen, Robert Xiao
IMX3
2025 VIBES: Exploring Viewer Spatial Interactions as Direct Input for Livestreamed Content
abstract
Livestreaming has rapidly become a popular online pastime, with real-time interaction between streamer and viewer being a key motivating feature. However, viewers have traditionally had limited opportunity to directly influence the streamed content; even when such interactions are possible, it has been reliant on text-based chat. We investigate the potential of spatial interaction on the livestreamed video content as a form of direct, real-time input for livestreamed applications. We developed VIBES, a flexible digital system that registers viewers' mouse interactions on the streamed video, i.e., clicks or movements, and transmits it directly into the streamed application. We used VIBES as a technology probe; first designing possible demonstrative interactions and using these interactions to explore streamers' perception of viewer influence and possible challenges and opportunities. We then deployed applications built using VIBES in two livestreams to explore its effects on audience engagement and investigate their relationships with the stream, the streamer, and fellow audience members. The use of spatial interactions enhances engagement and participation and opens up new avenues for both streamer-viewer and viewer-viewer participation. We contextualize our findings around a broader understanding of motivations and engagement in livestreaming, and we propose design guidelines and extensions for future research.
Michael Yin, Robert Xiao
IMX2
2025 GaussianNexus: Room-Scale Real-Time AR/VR Telepresence with Gaussian Splatting
Xincheng Huang, Dieter Frehlich, Ziyi Xia, Peyman Gholami, Robert Xiao
UIST5
2025 NFCGest: Contactless Gestural Interactions with NFC Devices
Bu Li, Robert Xiao
UIST2
2025 TangiAR: Markerless Tangible Input for Immersive Augmented Reality with Everyday Objects
abstract
Tangible interactions with everyday objects have been shown to be fast, accurate, and natural, and have shown promise when combined with immersive augmented reality. However, implementing tangible controls presents considerable challenges. Previous works in the field either rely on additional tracking markers on objects, inadvertently shifting the difficulty to users, or are too computationally demanding for real-time operation on a head-mounted display (HMD). We propose TangiAR, a tangible control system which tracks everyday objects without the need for fiducial trackers, enabling them as passive controllers and virtual proxies in AR applications. TangiAR additionally enables hand and finger proximity interactions with tangibles, further expanding the interaction space. TangiAR can run on an unmodified Microsoft HoloLens 2, making it immediately practical. We evaluated the performance of TangiAR through a technical evaluation, including occlusion robustness and tracking accuracy tests, and a user study which examined the usability of our markerless object tracking system in various AR interactions.
Neil Xu Fan, Xincheng Huang, Robert Xiao
VRST3
2025 MultiSphere: Latency Optimized Multi-User 360° VR Telepresence with Edge-Assisted Viewport Adaptive IPv6 Multicast
abstract
360° video telepresence with VR enables immersive remote collaboration, but scaling to multiple users is subject to bandwidth and latency constraints. We present MultiSphere, a multi-user edge-assited 360° VR telepresence system, that combines viewport-adaptive IPv6 multicast tiling with a novel dual keyframe interval (KeyInt) streaming technique. Our approach addresses the latency bottleneck inherent in joining live streams of video using standard video codecs while maintaining visual quality through strategic use of low and high KeyInt streams. Our system achieves 75-94% bandwidth savings and an average request-to-decode latency of 56 ms, a 79% reduction compared to using a regular single-KeyInt stream.
Dieter Frehlich, Xincheng Huang, Robert Xiao
VRST3
2025 VibRing: A Wearable Vibroacoustic Sensor for Single-Handed Gesture Recognition
abstract
Single-handed gestures offer rapid and intuitive interactions for input in interactive applications ranging from smartwatches and phones to augmented reality. Past research has explored using computer vision or inertial measurement units (IMUs) to sense such gestures, but these sensing modalities can be variously subject to occlusion, high power consumption, or sensitivity to random motion. In this work, we explore passively detecting the vibroacoustic signature of subtle single-handed gestures through a wearable piezoelectric sensor, providing a robust, low-power sensing modality. We present (1) a hand-gesture design framework encompassing a large set of subtle, rapid single-handed gestures which balance comfort and vibroacoustic distinguishability, (2) VibRing, a lightweight wireless hand-gesture sensing platform, leveraging a single finger-worn vibroacoustic sensor, and (3) a multifaceted system evaluation where we consider several aspects - general usability, tolerance to variance, user adaptability, and extended usage. Our results demonstrate that VibRing can support an 11-gesture set with a general accuracy of \(94.2\%\) and low-performance variance across multiple days ( \(90.2\%\) accuracy in cross-day validation). To support a new user, VibRing requires only 10 minutes of training data to achieve an accuracy of \(92.7\%\) . We also tested the extended use of VibRing in an office study where users performed periodic gesture inputs during typical office tasks with real-time classification, achieving a true-positive rate of \(90.9\%\) . Finally, to demonstrate the utility of VibRing, we present three examples of applications which benefit from our subtle gesture interactions.
Bu Li, Xincheng Huang, Robert Xiao
Proc. ACM Hum. Comput. Interact.3
2024 Lies, Deceit, and Hallucinations: Player Perception and Expectations Regarding Trust and Deception in Games
abstract
Lying and deception are important parts of social interaction; when applied to storytelling mediums such as video games, such elements can add complexity and intrigue. We developed a game, “AlphaBetaCity”, in which non-playable characters (NPCs) made various false statements, and used this game to investigate perceptions of deceptive behaviour. We used a mix of human-written dialogue incorporating deliberate falsehoods and LLM-written scripts with (human-approved) hallucinated responses. The degree of falsehoods varied between believable but untrue statements to outright fabrications. 29 participants played the game and were interviewed about their experiences. Participants discussed methods for developing trust and gauging NPC truthfulness. Whereas perceived intentional false statements were often attributed towards narrative and gameplay effects, seemingly unintentional false statements generally mismatched participants’ mental models and lacked inherent meaning. We discuss how the perception of intentionality, the audience demographic, and the desire for meaning are major considerations when designing video games with falsehoods.
Michael Yin, Emi Wang, Chuoxi Ng, Robert Xiao
CHI4
2024 VirtualNexus: Enhancing 360-Degree Video AR/VR Collaboration with Environment Cutouts and Virtual Replicas
abstract
Asymmetric AR/VR collaboration systems bring a remote VR user to a local AR user’s physical environment, allowing them to communicate and work within a shared virtual/physical space. Such systems often display the remote environment through 3D reconstructions or 360° videos. While 360° cameras stream an environment in higher quality, they lack spatial information, making them less interactable. We present VirtualNexus, an AR/VR collaboration system that enhances 360° video AR/VR collaboration with environment cutouts and virtual replicas. VR users can define cutouts of the remote environment to interact with as a world-in-miniature, and their interactions are synchronized to the local AR perspective. Furthermore, AR users can rapidly scan and share 3D virtual replicas of physical objects using neural rendering. We demonstrated our system’s utility through 3 example applications and evaluated our system in a dyadic usability test. VirtualNexus extends the interaction space of 360° telepresence systems, offering improved physical presence, versatility, and clarity in interactions.
Xincheng Huang, Michael Yin, Ziyi Xia, Robert Xiao
UIST4
2024 Press A or Wave: User Expectations for NPC Interactions and Nonverbal Behaviour in Virtual Reality
abstract
Non-playable characters (NPCs) are important in games, as they can provide guidance to the player, create social engagement, and advance the game's narrative. Although much research exists regarding NPC interactions for traditional gaming environments, e.g. desktop or console, fewer works have considered this from a virtual reality (VR) perspective. Our work first uncovers the salient and unique dimensions of VR NPC interactions through observations of 47 existing games. We find that VR NPC interactions have an extended set of interaction mechanisms due to two key factors - interaction triggers and player constraints within the game, driven by the unique qualities of physical motion and immersion afforded by the medium. We augment these findings through a user study performed on 18 participants in a VR environment. Participant interactions with a responsive NPC allow us to delve deeper into understanding player perception and expectations of NPC behaviour and interactions. Our findings outline player expectations for NPC realism, player agency during NPC interaction, and NPC expected behaviour and feedback. We tie our findings into discussions on player agency within VR, highlighting design suggestions to develop NPCs to better fit within social behaviour expectations.
Michael Yin, Robert Xiao
Proc. ACM Hum. Comput. Interact.2
2024 How We See Changes How We Feel: Investigating the Effect of Visual Point-of-View on Decision-Making in VR Environments
abstract
Virtual reality (VR) can immerse users into engaging experiences, affording opportunities to study behaviour in simulated contexts such as decision-making processes. However, methodological research into designing meaningful VR experiences - experiences that promote appreciation and deeper understanding of a work - is still underdeveloped. In this two-part study, we investigate how visual point-of-view (POV) in VR impacts feelings of meaningfulness and empathy as well as objective decision-making processes. Our study revolves around a VR application that situates users in moral dilemmas from three different POVs. Data from the choices made is augmented with self-reported subjective data. We find that, from different POVs, users' subjective feelings do show change; users show greater empathy for virtual agents and have an increasingly meaningful experience from a first-person perspective, even if this is not always reflected in changes in their decisions. Finally, we discuss the implications of our findings in the context of VR application design.
Michael Yin, Robert Xiao
Proc. ACM Hum. Comput. Interact.2
2023 Drifting Off in Paradise: Why People Sleep in Virtual Reality
abstract
Sleep is important for humans, and past research has considered methods of improving sleep through technologies such as virtual reality (VR). However, there has been limited research on how such VR technology may affect the experiential and practical aspects of sleep, especially outside of a clinical lab setting. We consider this research gap through the lens of individuals that voluntarily engage in the practice of sleeping in VR. Semi-structured interviews with 14 participants that have slept in VR reveal insights regarding the motivations, actions, and experiential factors that uniquely define this practice. We find that participant motives can be largely categorized through either the experiential or social affordances of VR. We tie these motives into findings regarding the unique customs of sleeping in VR, involving set-up both within the physical and virtual space. Finally, we identify current and future challenges for sleeping in VR, and propose prospective design directions.
Michael Yin, Robert Xiao
CHI2
2023 GestureCanvas: A Programming by Demonstration System for Prototyping Compound Freehand Interaction in VR
abstract
As the use of hand gestures becomes increasingly prevalent in virtual reality (VR) applications, prototyping Compound Freehand Interactions (CFIs) effectively and efficiently has become a critical need in the design process. Compound Freehand Interaction (CFI) is a sequence of freehand interactions where each sub-interaction in the sequence conditions the next. Despite the need for interactive prototypes of CFI in the early design stage, creating them is effortful and remains a challenge for designers since it requires a highly technical workflow that involves programming the recognizers, system responses and conditionals for each sub-interaction. To bridge this gap, we present GestureCanvas, a freehand interaction-based immersive prototyping system that enables a rapid, end-to-end, and code-free workflow for designing, testing, refining, and subsequently deploying CFI by leveraging the three pillars of interaction models: event-driven state machine, trigger-action authoring, and programming by demonstration. The design of GestureCanvas includes three novel design elements — (i) appropriating the multimodal recording of freehand interaction into a CFI authoring workspace called Design Canvas, (ii) semi-automatic identification of the input trigger logic from demonstration to reduce the manual effort of setting up triggers for each sub-interaction, (iii) on the fly testing for independently validating the input conditionals in-situ. We validate the workflow enabled by GestureCanvas through an interview study with professional designers and evaluate its usability through a user study with non-experts. Our work lays the foundation for advancing research on immersive prototyping systems allowing even highly complex gestures to be easily prototyped and tested within VR environments.
Anika Sayara, Emily Lynn Chen, Cuong Nguyen 0003, Robert Xiao, Dongwook Yoon
UIST4
2023 Virtual Reality Telepresence: 360-Degree Video Streaming with Edge-Compute Assisted Static Foveated Compression
abstract
Real-time communication with immersive 360° video can enable users to be telepresent within a remotely streamed environment. Increasingly, users are shifting to mobile devices and connecting to the Internet via mobile-cellular networks. As the ideal media for 360° videos, some VR headsets now also come with cellular capacity, giving them potential for mobile applications. However, streaming high-quality 360° live video poses challenges for network bandwidth, particularly on cellular connections. To reduce bandwidth requirements, videos can be compressed using viewport-adaptive streaming or foveated rendering techniques. Such approaches require very low latency in order to be effective, which has previously limited their applications on traditional cellular networks. In this work, we demonstrate an end-to-end virtual reality telepresence system that streams ∼6K 360° video over 5G millimeter-wave (mmW) radio. Our use of 5G technologies, in conjunction with mobile edge compute nodes, substantially reduces latency when compared with existing 4G networks, enabling high-efficiency foveated compression over modern cellular networks on par with WiFi. We performed a technical evaluation of our system's visual quality post-compression with peak signal-to-noise ratio (PSNR) and FOVVideoVDP. We also conducted a user study to evaluate users' sensitivity to compressed video. Our findings demonstrate that our system achieves visually indistinguishable video streams while using up to 80% less data when compared with un-foveated video. We demonstrate our video compression system in the context of an immersive, telepresent video calling application.
Xincheng Huang, James Riddell, Robert Xiao
IEEE Trans. Vis. Comput. Graph.3
2022 How Should I Respond to "Good Morning?": Understanding Choice in Narrative-Rich Games
abstract
Narrative-rich video games provide opportunities for players to make choices at key points in the game, generating malleability within the game world and its characters. In this study, we explore the types of choices that exist in such games, how choices affect player experience, and how players make decisions when presented with choice. We first conduct interviews with game developers and perform a video observation analysis of existing choices to develop an initial classification system. We then perform a series of semi-structured interviews with video game players to understand how different choices impact player experience. Our findings reveal that choices influence player experience at several levels of meta-gameplay, having impacts on the game itself, the player-game relationship, and the player outside the game. Furthermore, we identify several key factors that affect player decision-making when faced with choice. Finally, we discuss the potential of choice in developing impactful virtual experiences.
Michael Yin, Robert Xiao
Conference on Designing Interactive Systems2
2022 The Reward for Luck: Understanding the Effect of Random Reward Mechanisms in Video Games on Player Experience
abstract
Random Reward Mechanisms (RRMs) in video games are systems in which rewards are issued probabilistically upon certain trigger conditions, such as completing gameplay tasks, exceeding a playtime quota, or making in-game purchases. We investigated the relationship between RRM implementations and user experience. Video analysis of 35 RRM systems allowed for the creation of a classification system based on contrasting observed dimensions. Interviews with 14 video game players provided insights into how factors such as the affordances of non-optimal rewards and the trade-off between random luck and skill impact player perception and interaction with RRMs. We additionally investigated the relationship between auditory, visual, and gameplay design decisions and player expectations for RRM reward presentations, finding that the resources required to obtain the reward and the relative value of the reward impact its expected presentation. Finally, we applied our findings to propose design methodologies for creating engaging and significant RRM systems.
Michael Yin, Robert Xiao
CHI2
2022 Learned Acoustic Reconstruction Using Synthetic Aperture Focusing
abstract
Many algorithmic approaches to 3D acoustic imaging have been devised which rely on a large abundance of receiving elements to produce images with delay-and-sum techniques, but these have found little use in air due to hardware complexity and low accuracy. Recent learning-based approaches to one-shot in-air acoustic reconstruction attempt to overcome these limitations using simple hardware and large datasets of geometry and echo pairs to train neural networks. How-ever, existing learned models use spatially-dense representations and attempt to predict entire scenes at once, requiring an abundance of data to truly generalize.We train an implicit neural network with no spatial awareness to predict the distance to the nearest obstacle at a single location from only time-delayed echoes. Using acoustic wave simulation, we show that our method yields better generalization and behaves more intuitively than competing methods while requiring only a fraction of the training data. Our code and data is available at https://timstr.github.io/learned-acoustic-reconstruction/.
Tim Straubinger, Robert Xiao, Helge Rhodin
ICASSP2
2022 Reducing the Latency of Touch Tracking on Ad-hoc Surfaces
abstract
Touch sensing on ad-hoc surfaces has the potential to transform everyday surfaces in the environment - desks, tables and walls - into tactile, touch-interactive surfaces, creating large, comfortable interactive spaces without the cost of large touch sensors. Depth sensors are a promising way to provide touch sensing on arbitrary surfaces, but past systems have suffered from high latency and poor touch detection accuracy. We apply a novel state machine-based approach to analyzing touch events, combined with a machine-learning approach to predictively classify touch events from depth data with lower latency and higher touch accuracy than previous approaches. Our system can reduce end-to-end touch latency to under 70ms, comparable to conventional capacitive touchscreens. Additionally, we open-source our dataset of over 30,000 touch events recorded in depth, infrared and RGB for the benefit of future researchers.
Neil Xu Fan, Robert Xiao
Proc. ACM Hum. Comput. Interact.2
2021 FoldMold: Automating Papercraft for Fast DIY Casting of Scalable Curved Shapes
abstract
Rapid iteration is crucial to effective prototyping; yet making certain objects - large, smoothly curved and/or of specific material - requires specialized equipment or considerable time. To improve access to casting such objects, we developed FoldMold: a low-cost, simply-resourced and eco-friendly technique for creating scalable, curved mold shapes (any developable surface) with wax-stiffened paper. Starting with a 3D digital shape, we define seams, add bending, joinery and mold-strengthening features, and "unfold" the shape into a 2D pattern, which is then cut, assembled, wax-dipped and cast with materials like silicone, plaster, or ice. To access the concept's full power, we facilitated digital pattern creation with a custom Blender add-on. We assessed FoldMold's viability, first with several molding challenges in which it produced smooth, curved shapes far faster than 3D printing would; then with a small user study that confirmed automation usability. Finally, we describe a range of opportunities for further development.
Hanieh Shakeri, Hannah Elbaggari, Paul Bucci, Robert Xiao, Karon E. MacLean
Graphics Interface4
2021 PAIR: Phone as an Augmented Immersive Reality Controller
abstract
Immersive head-mounted augmented reality allows users to overlay 3D digital content on a user’s view of the world. Current-generation devices primarily support interaction modalities such as gesture, gaze and voice, which are readily available to most users yet lack precision and tactility, rendering them fatiguing for extended interactions. We propose using smartphones, which are also readily available, as companion devices complementing existing AR interaction modalities. We leverage user familiarity with smartphone interactions, coupled with their support for precise, tactile touch input, to unlock a broad range of interaction techniques and applications - for instance, turning the phone into an interior design palette, touch-enabled catapult or AR-rendered sword. We describe a prototype implementation of our interaction techniques using an off-the-shelf AR headset and smartphone, demonstrate applications, and report on the results of a positional accuracy study.
Arda Ege Unlu, Robert Xiao
VRST2
2020 Phasking on Paper: Accessing a Continuum of PHysically Assisted SKetchING
abstract
When sketching, we must choose between paper (expressive ease, ruler and eraser) and computational assistance (parametric support, a digital record). PHysically Assisted SKetching provides both, with a pen that displays force constraints with which the sketcher interacts as they draw on paper. Phasking provides passive, "bound" constraints (like a ruler); or actively "brings" the sketcher along a commanded path (e.g., a curve), which they can violate for creative variation. The sketcher modulates constraint strength (control sharing) by bearing down on the pen-tip. Phasking requires untethered, graded force-feedback, achieved by modifying a ballpoint drive that generates force through rolling surface contact. To understand phasking's viability, we implemented its interaction concepts, related them to sketching tasks and measured device performance. We assessed the experience of 10 sketchers, who could understand, use and delight in phasking, and who valued its control-sharing and digital twinning for productivity, creative control and learning to draw.
Soheil Kianzad, Robert Xiao, Karon E. MacLean
CHI3
2020 VibroComm: Using Commodity Gyroscopes for Vibroacoustic Data Reception
abstract
Inertial Measurement Units (IMUs) with gyroscopic sensors are standard in today's mobile devices. We show that these sensors can be co-opted for vibroacoustic data reception. Our approach, called VibroComm, requires direct physical contact to a transmitting (i.e., vibrating) surface. This makes interactions targeted and explicit in nature, making it well suited for contexts with many targets or requiring and intent. It also offers an orthogonal dimension of physical security to wireless technologies like Blue-tooth and NFC. Using our implementation, we achieve a transfer rate over 2000 bits/sec with less than 5% packet loss – an order of magnitude faster than prior IMU-based approaches at a quarter of the loss rate, opening new, powerful and practical use cases that could be enabled on mobile devices with a simple software update.
Robert Xiao, Sven Mayer, Chris Harrison 0001
MobileHCI1
2019 MeCap: Whole-Body Digitization for Low-Cost VR/AR Headsets
abstract
Low-cost, smartphone-powered VR/AR headsets are becoming more popular. These basic devices - little more than plastic or cardboard shells - lack advanced features, such as controllers for the hands, limiting their interactive capability. Moreover, even high-end consumer headsets lack the ability to track the body and face. For this reason, interactive experiences like social VR are underdeveloped. We introduce MeCap, which enables commodity VR headsets to be augmented with powerful motion capture ("MoCap") and user-sensing capabilities at very low cost (under $5). Using only a pair of hemi-spherical mirrors and the existing rear-facing camera of a smartphone, MeCap provides real-time estimates of a wearer's 3D body pose, hand pose, facial expression, physical appearance and surrounding environment - capabilities which are either absent in contemporary VR/AR systems or which require specialized hardware and controllers. We evaluate the accuracy of each of our tracking features, the results of which show imminent feasibility.
Karan Ahuja, Chris Harrison 0001, Mayank Goel, Robert Xiao
UIST4
2019 LightAnchors: Appropriating Point Lights for Spatially-Anchored Augmented Reality Interfaces
abstract
Augmented reality requires precise and instant overlay of digital information onto everyday objects. We present our work on LightAnchors, a new method for displaying spatially-anchored data. We take advantage of pervasive point lights - such as LEDs and light bulbs - for both in-view anchoring and data transmission. These lights are blinked at high speed to encode data. We built a proof-of-concept ap-plication that runs on iOS without any hardware or software modifications. We also ran a study to characterize the performance of LightAnchors and built eleven example demos to highlight the potential of our approach.
Karan Ahuja, Sujeath Pareddy, Robert Xiao, Mayank Goel, Chris Harrison 0001
UIST3
2018 LumiWatch: On-Arm Projected Graphics and Touch Input
abstract
Compact, worn computers with projected, on-skin touch interfaces have been a long-standing yet elusive goal, largely written off as science fiction. Such devices offer the potential to mitigate the significant human input/output bottleneck inherent in worn devices with small screens. In this work, we present the first fully functional and self-contained projection smartwatch implementation, containing the requisite compute, power, projection and touch-sensing capabilities. Our watch offers roughly 40 sq. cm of interactive surface area -- more than five times that of a typical smartwatch display. We demonstrate continuous 2D finger tracking with interactive, rectified graphics, transforming the arm into a touchscreen. We discuss our hardware and software implementation, as well as evaluation results regarding touch accuracy and projection visibility.
Robert Xiao, Teng Cao, Jun Zhuo, Yang Zhang 0041, Chris Harrison 0001
CHI1
2018 MRTouch: Adding Touch Input to Head-Mounted Mixed Reality
abstract
We present MRTouch, a novel multitouch input solution for head-mounted mixed reality systems. Our system enables users to reach out and directly manipulate virtual interfaces affixed to surfaces in their environment, as though they were touchscreens. Touch input offers precise, tactile and comfortable user input, and naturally complements existing popular modalities, such as voice and hand gesture. Our research prototype combines both depth and infrared camera streams together with real-time detection and tracking of surface planes to enable robust finger-tracking even when both the hand and head are in motion. Our technique is implemented on a commercial Microsoft HoloLens without requiring any additional hardware nor any user or environmental calibration. Through our performance evaluation, we demonstrate high input accuracy with an average positional error of 5.4 mm and 95% button size of 16 mm, across 17 participants, 2 surface orientations and 4 surface materials. Finally, we demonstrate the potential of our technique to enable on-world touch interactions through 5 example applications.
Robert Xiao, Julia Schwarz, Nick Throm, Andrew D. Wilson, Hrvoje Benko
IEEE Trans. Vis. Comput. Graph.1
2017 Deus EM Machina: On-Touch Contextual Functionality for Smart IoT Appliances
abstract
Homes, offices and many other environments will be increasingly saturated with connected, computational appliances, forming the "Internet of Things" (IoT). At present, most of these devices rely on mechanical inputs, webpages, or smartphone apps for control. However, as IoT devices proliferate, these existing interaction methods will become increasingly cumbersome. Will future smart-home owners have to scroll though pages of apps to select and dim their lights? We propose an approach where users simply tap a smartphone to an appliance to discover and rapidly utilize contextual functionality. To achieve this, our prototype smartphone recognizes physical contact with uninstrumented appliances, and summons appliance-specific interfaces. Our user study suggests high accuracy 98.8% recognition accuracy among 17 appliances. Finally, to underscore the immediate feasibility and utility of our system, we built twelve example applications, including six fully functional end-to-end demonstrations.
Robert Xiao, Gierad Laput, Yang Zhang 0041, Chris Harrison 0001
CHI1
2017 Supporting Responsive Cohabitation Between Virtual Interfaces and Physical Objects on Everyday Surfaces
abstract
Systems for providing mixed physical-virtual interaction on desktop surfaces have been proposed for decades, though no such systems have achieved widespread use. One major factor contributing to this lack of acceptance may be that these systems are not designed for the variety and complexity of actual work surfaces, which are often in flux and cluttered with physical objects. In this paper, we use an elicitation study and interviews to synthesize a list of ten interactive behaviors that desk-bound, digital interfaces should implement to support responsive cohabitation with physical objects. As a proof of concept, we implemented these interactive behaviors in a working augmented desk system, demonstrating their imminent feasibility.
Robert Xiao, Scott E. Hudson, Chris Harrison 0001
Proc. ACM Hum. Comput. Interact.1
2016 Augmenting the Field-of-View of Head-Mounted Displays with Sparse Peripheral Displays
abstract
In this paper, we explore the concept of a sparse peripheral display, which augments the field-of-view of a head-mounted display with a lightweight, low-resolution, inexpensively produced array of LEDs surrounding the central high-resolution display. We show that sparse peripheral displays expand the available field-of-view up to 190º horizontal, nearly filling the human field-of-view. We prototyped two proof-of-concept implementations of sparse peripheral displays: a virtual reality headset, dubbed SparseLightVR, and an augmented reality headset, called SparseLightAR. Using SparseLightVR, we conducted a user study to evaluate the utility of our implementation, and a second user study to assess different visualization schemes in the periphery and their effect on simulator sickness. Our findings show that sparse peripheral displays are useful in conveying peripheral information and improving situational awareness, are generally preferred, and can help reduce motion sickness in nausea-susceptible people.
Robert Xiao, Hrvoje Benko
CHI1
2016 DIRECT: Making Touch Tracking on Ordinary Surfaces Practical with Hybrid Depth-Infrared Sensing
abstract
Several generations of inexpensive depth cameras have opened the possibility for new kinds of interaction on everyday surfaces. A number of research systems have demonstrated that depth cameras, combined with projectors for output, can turn nearly any reasonably flat surface into a touch-sensitive display. However, even with the latest generation of depth cameras, it has been difficult to obtain sufficient sensing fidelity across a table-sized surface to get much beyond a proof-of-concept demonstration. In this paper we present DIRECT, a novel touch-tracking algorithm that merges depth and infrared imagery captured by a commodity sensor. This yields significantly better touch tracking than from depth data alone, as well as any prior system. Further extending prior work, DIRECT supports arbitrary user orientation and requires no prior calibration or background capture. We describe the implementation of our system and quantify its accuracy through a comparison study of previously published, depth-based touch-tracking algorithms. Results show that our technique boosts touch detection accuracy by 15% and reduces positional error by 55% compared to the next best-performing technique.
Robert Xiao, Scott E. Hudson, Chris Harrison 0001
ISS1
2016 CapCam: Enabling Rapid, Ad-Hoc, Position-Tracked Interactions Between Devices
abstract
We present CapCam, a novel technique that enables smartphones (and similar devices) to establish quick, ad-hoc connections with a host touchscreen device, simply by pressing a device to the screen's surface. Pairing data, used to bootstrap a conventional wireless connection, is transmitted optically to the phone's rear camera. This approach utilizes the near-ubiquitous rear camera on smart devices, making it applicable to a wide range of devices, both new and old. CapCam also tracks phones' physical positions on the host capacitive touchscreen without any instrumentation, enabling a wide range of targeted interactions. We quantify the communication performance of our pairing approach and demonstrate data transmission rates up to four times faster than prior camera-based techniques. To demonstrate the unique capability and utility of our system, we built a series of example applications, highlighting different interaction techniques CapCam enables.
Robert Xiao, Scott E. Hudson, Chris Harrison 0001
ISS1
2016 ViBand: High-Fidelity Bio-Acoustic Sensing Using Commodity Smartwatch Accelerometers
abstract
Smartwatches and wearables are unique in that they reside on the body, presenting great potential for always-available input and interaction. Their position on the wrist makes them ideal for capturing bio-acoustic signals. We developed a custom smartwatch kernel that boosts the sampling rate of a smartwatch's existing accelerometer to 4 kHz. Using this new source of high-fidelity data, we uncovered a wide range of applications. For example, we can use bio-acoustic data to classify hand gestures such as flicks, claps, scratches, and taps, which combine with on-device motion tracking to create a wide range of expressive input modalities. Bio-acoustic sensing can also detect the vibrations of grasped mechanical or motor-powered objects, enabling passive object recognition that can augment everyday experiences with context-aware functionality. Finally, we can generate structured vibrations using a transducer, and show that data can be transmitted through the human body. Overall, our contributions unlock user interface techniques that previously relied on special-purpose and/or cumbersome instrumentation, making such interactions considerably more feasible for inclusion in future consumer devices.
Gierad Laput, Robert Xiao, Chris Harrison 0001
UIST2
2016 Advancing Hand Gesture Recognition with High Resolution Electrical Impedance Tomography
abstract
Electrical Impedance Tomography (EIT) was recently employed in the HCI domain to detect hand gestures using an instrumented smartwatch. This prior work demonstrated great promise for non-invasive, high accuracy recognition of gestures for interactive control. We introduce a new system that offers improved sampling speed and resolution. In turn, this enables superior interior reconstruction and gesture recognition. More importantly, we use our new system as a vehicle for experimentation ' we compare two EIT sensing methods and three different electrode resolutions. Results from in-depth empirical evaluations and a user study shed light on the future feasibility of EIT for sensing human input.
Yang Zhang 0041, Robert Xiao, Chris Harrison 0001
UIST2
2015 Zensors: Adaptive, Rapidly Deployable, Human-Intelligent Sensor Feeds
abstract
The promise of "smart" homes, workplaces, schools, and other environments has long been championed. Unattractive, however, has been the cost to run wires and install sensors. More critically, raw sensor data tends not to align with the types of questions humans wish to ask, e.g., do I need to restock my pantry? Although techniques like computer vision can answer some of these questions, it requires significant effort to build and train appropriate classifiers. Even then, these systems are often brittle, with limited ability to handle new or unexpected situations, including being repositioned and environmental changes (e.g., lighting, furniture, seasons). We propose Zensors, a new sensing approach that fuses real-time human intelligence from online crowd workers with automatic approaches to provide robust, adaptive, and readily deployable intelligent sensors. With Zensors, users can go from question to live sensor feed in less than 60 seconds. Through our API, Zensors can enable a variety of rich end-user applications and moves us closer to the vision of responsive, intelligent environments.
Gierad Laput, Walter S. Lasecki, Jason Wiese, Robert Xiao, Jeffrey P. Bigham, Chris Harrison 0001
CHI4
2015 Gaze+Gesture: Expressive, Precise and Targeted Free-Space Interactions
abstract
Humans rely on eye gaze and hand manipulations extensively in their everyday activities. Most often, users gaze at an object to perceive it and then use their hands to manipulate it. We propose applying a multimodal, gaze plus free-space gesture approach to enable rapid, precise and expressive touch-free interactions. We show the input methods are highly complementary, mitigating issues of imprecision and limited expressivity in gaze-alone systems, and issues of targeting speed in gesture-alone systems. We extend an existing interaction taxonomy that naturally divides the gaze+gesture interaction space, which we then populate with a series of example interaction techniques to illustrate the character and utility of each method. We contextualize these interaction techniques in three example scenarios. In our user study, we pit our approach against five contemporary approaches; results show that gaze+gesture can outperform systems using gaze or gesture alone, and in general, approach the performance of "gold standard" input systems, such as the mouse and trackpad.
Ishan Chatterjee, Robert Xiao, Chris Harrison 0001
ICMI2
2015 EM-Sense: Touch Recognition of Uninstrumented, Electrical and Electromechanical Objects
abstract
Most everyday electrical and electromechanical objects emit small amounts of electromagnetic (EM) noise during regular operation. When a user makes physical contact with such an object, this EM signal propagates through the user, owing to the conductivity of the human body. By modifying a small, low-cost, software-defined radio, we can detect and classify these signals in real-time, enabling robust on-touch object detection. Unlike prior work, our approach requires no instrumentation of objects or the environment; our sensor is self-contained and can be worn unobtrusively on the body. We call our technique EM-Sense and built a proof-of-concept smartwatch implementation. Our studies show that discrimination between dozens of objects is feasible, independent of wearer, time and local environment.
Gierad Laput, Chouchang Yang, Robert Xiao, Alanson P. Sample, Chris Harrison 0001
UIST3
2014 TouchTools: leveraging familiarity and skill with physical tools to augment touch interaction
abstract
The average person can skillfully manipulate a plethora of tools, from hammers to tweezers. However, despite this remarkable dexterity, gestures on today's touch devices are simplistic, relying primarily on the chording of fingers: one-finger pan, two-finger pinch, four-finger swipe and similar. We propose that touch gesture design be inspired by the manipulation of physical tools from the real world. In this way, we can leverage user familiarity and fluency with such tools to build a rich set of gestures for touch interaction. With only a few minutes of training on a proof-of-concept system, users were able to summon a variety of virtual tools by replicating their corresponding real-world grasps.
Chris Harrison 0001, Robert Xiao, Julia Schwarz, Scott E. Hudson
CHI2
2014 Probabilistic palm rejection using spatiotemporal touch features and iterative classification
abstract
Tablet computers are often called upon to emulate classical pen-and-paper input. However, touchscreens typically lack the means to distinguish between legitimate stylus and finger touches and touches with the palm or other parts of the hand. This forces users to rest their palms elsewhere or hover above the screen, resulting in ergonomic and usability problems. We present a probabilistic touch filtering approach that uses the temporal evolution of touch contacts to reject palms. Our system improves upon previous approaches, reducing accidental palm inputs to 0.016 per pen stroke, while correctly passing 98% of stylus inputs.
Julia Schwarz, Robert Xiao, Jennifer Mankoff, Scott E. Hudson, Chris Harrison 0001
CHI2
2014 Expanding the input expressivity of smartwatches with mechanical pan, twist, tilt and click
abstract
Smartwatches promise to bring enhanced convenience to common communication, creation and information retrieval tasks. Due to their prominent placement on the wrist, they must be small and otherwise unobtrusive, which limits the sophistication of interactions we can perform. This problem is particularly acute if the smartwatch relies on a touchscreen for input, as the display is small and our fingers are relatively large. In this work, we propose a complementary input approach: using the watch face as a multi-degree-of-freedom, mechanical interface. We developed a proof of concept smartwatch that supports continuous 2D panning and twist, as well as binary tilt and click. To illustrate the potential of our approach, we developed a series of example applications, many of which are cumbersome -- or even impossible -- on today's smartwatch devices.
Robert Xiao, Gierad Laput, Chris Harrison 0001
CHI1
2014 Toffee: enabling ad hoc, around-device interaction with acoustic time-of-arrival correlation
abstract
The simple fact that human fingers are large and mobile devices are small has led to the perennial issue of limited surface area for touch-based interactive tasks. In response, we have developed Toffee, a sensing approach that extends touch interaction beyond the small confines of a mobile device and onto ad hoc adjacent surfaces, most notably tabletops. This is achieved using a novel application of acoustic time differences of arrival (TDOA) correlation. Previous time-of-arrival based systems have required semi-permanent instrumentation of the surface and were too large for use in mobile devices. Our approach requires only a hard tabletop and gravity -- the latter acoustically couples mobile devices to surfaces. We conducted an evaluation, which shows that Toffee can accurately resolve the bearings of touch events (mean error of 4.3° with a laptop prototype). This enables radial interactions in an area many times larger than a mobile device; for example, virtual buttons that lie above, below and to the left and right.
Robert Xiao, Greg Lew, James Marsanico, Divya Hariharan, Scott E. Hudson, Chris Harrison 0001
Mobile HCI1
2014 Skin buttons: cheap, small, low-powered and clickable fixed-icon laser projectors
abstract
Smartwatches are a promising new interactive platform, but their small size makes even basic actions cumbersome. Hence, there is a great need for approaches that expand the interactive envelope around smartwatches, allowing human input to escape the small physical confines of the device. We propose using tiny projectors integrated into the smartwatch to render icons on the user's skin. These icons can be made touch sensitive, significantly expanding the interactive region without increasing device size. Through a series of experiments, we show that these 'skin buttons' can have high touch accuracy and recognizability, while being low cost and power-efficient.
Gierad Laput, Robert Xiao, Xiang 'Anthony' Chen, Scott E. Hudson, Chris Harrison 0001
UIST2
2013 HomeProxy: exploring a physical proxy for video communication in the home
abstract
HomeProxy is a research prototype that explores supporting video communication in the home among distributed family members through a physical proxy. It leverages a physical artifact dedicated to representing remote family members to make it easier to share activities with them. HomeProxy combines a form factor designed for the home environment with a "no-touch" user experience and an interface that responsively transitions between recorded and live video messages. We designed and implemented a prototype and conducted a pilot study with eight pairs of users. Our study demonstrated the challenges of a no-touch interface and the promise of offering quick video messaging in the home.
John C. Tang, Robert Xiao, Aaron Hoff, Gina Venolia, Patrick Therien, Asta Roseway
CHI2
2013 WorldKit: rapid and easy creation of ad-hoc interactive applications on everyday surfaces
abstract
Instant access to computing, when and where we need it, has long been one of the aims of research areas such as ubiquitous computing. In this paper, we describe the WorldKit system, which makes use of a paired depth camera and projector to make ordinary surfaces instantly interactive. Using this system, touch-based interactivity can, without prior calibration, be placed on nearly any unmodified surface literally with a wave of the hand, as can other new forms of sensed interaction. From a user perspective, such interfaces are easy enough to instantiate that they could, if desired, be recreated or modified "each time we sat down" by "painting" them next to us. From the programmer's perspective, our system encapsulates these capabilities in a simple set of abstractions that make the creation of interfaces quick and easy. Further, it is extensible to new, custom interactors in a way that closely mimics conventional 2D graphical user interfaces, hiding much of the complexity of working in this new domain. We detail the hardware and software implementation of our system, and several example applications built using the library.
Robert Xiao, Chris Harrison 0001, Scott E. Hudson
CHI1
2013 Lumitrack: low cost, high precision, high speed tracking with projected m-sequences
abstract
We present Lumitrack, a novel motion tracking technology that uses projected structured patterns and linear optical sensors. Each sensor unit is capable of recovering 2D location within the projection area, while multiple sensors can be combined for up to six degree of freedom (DOF) tracking. Our structured light approach is based on special patterns, called m-sequences, in which any consecutive sub-sequence of m bits is unique. Lumitrack can utilize both digital and static projectors, as well as scalable embedded sensing configurations. The resulting system enables high-speed, high precision, and low-cost motion tracking for a wide range of interactive applications. We detail the hardware, operation, and performance characteristics of our approach, as well as a series of example applications that highlight its immediate feasibility and utility.
Robert Xiao, Chris Harrison 0001, Karl D. D. Willis, Ivan Poupyrev, Scott E. Hudson
UIST1
2012 Acoustic barcodes: passive, durable and inexpensive notched identification tags
abstract
We present acoustic barcodes, structured patterns of physical notches that, when swiped with e.g., a fingernail, produce a complex sound that can be resolved to a binary ID. A single, inexpensive contact microphone attached to a surface or object is used to capture the waveform. We present our method for decoding sounds into IDs, which handles variations in swipe velocity and other factors. Acoustic barcodes could be used for information retrieval or to triggering interactive functions. They are passive, durable and inexpensive to produce. Further, they can be applied to a wide range of materials and objects, including plastic, wood, glass and stone. We conclude with several example applications that highlight the utility of our approach, and a user study that explores its feasibility.
Chris Harrison 0001, Robert Xiao, Scott E. Hudson
UIST2
2012 Analysis and comparison of target assistance techniques for relative ray-cast pointing
Scott Bateman, Regan L. Mandryk, Carl Gutwin, Robert Xiao
Int. J. Hum. Comput. Stud.4
2011 Chalk sounds: the effects of dynamic synthesized audio on workspace awareness in distributed groupware
abstract
Awareness of other people's activity is an important part of shared-workspace collaboration, and is typically supported using visual awareness displays such as radar views. These visual presentations are limited in that the user must be able to see and attend to the view in order to gather awareness information. Using audio to convey awareness information does not suffer from these limitations, and previous research has shown that audio can provide valuable awareness in distributed settings. In this paper we evaluate the effectiveness of synthesized dynamic audio information, both on its own and as an adjunct to a visual radar view. We developed a granular-synthesis engine that produces realistic chalk sounds for off-screen activity in a groupware workspace, and tested the audio awareness in two ways. First, we measured people's ability to identify off-screen activities using only sound, and found that people are almost as accurate with synthesized sounds as with real sounds. Second, we tested dynamic audio awareness in a realistic groupware scenario, and found that adding audio to a radar view significantly improved awareness of off-screen activities in situations where it was difficult to see or attend to the visual display. Our work provides new empirical evidence about the value of dynamic synthesized audio in distributed groupware.
Carl Gutwin, Oliver Schneider 0006, Robert Xiao, Stephen A. Brewster
CSCW3
2011 Effects of view, input device, and track width on video game driving
Scott Bateman, Andre Doucette, Robert Xiao, Carl Gutwin, Regan L. Mandryk, Andy Cockburn
Graphics Interface3
2011 Ubiquitous cursor: a comparison of direct and indirect pointing feedback in multi-display environments
Robert Xiao, Miguel A. Nacenta, Regan L. Mandryk, Andy Cockburn, Carl Gutwin
Graphics Interface1