EDBT 2026 Demo / reviewers in the wild / expert
Andrew D. Wilson
dblp:77/3558 · also Andrew David Wilson, Andy D. Wilson
· DBLP profile ↗
85ranked-venue papers
27as first author
15since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 67 · 14 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 8 first-author · 2 since 2021Artificial intelligence and machine learning · 12 · 10 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 first-authorSystems, architecture and hardware · 2 · 2 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Storycaster: An AI System for Immersive Room-based StorytellingabstractWhile Cave Automatic Virtual Environment (CAVE) systems have long enabled room-scale virtual reality and various kinds of interactivity, their content has largely remained predetermined. We present Storycaster, a generative AI CAVE system that transforms physical rooms into responsive storytelling environments. Unlike headset-based VR, Storycaster preserves spatial awareness, using live camera feeds to augment the walls with cylindrical projections, allowing users to create worlds that blend with their physical surroundings. Additionally, our system enables object-level editing, where physical items in the room can be transformed to their virtual counterparts in a story. A narrator agent guides participants, enabling them to co-create stories that evolve in response to voice commands, with each scene enhanced by generated ambient audio, dialogue, and imagery. Participants in our study (n = 13) found the system highly immersive and engaging, identifying the narrator and audio as the most impactful elements, while also highlighting areas of improvement in latency and image resolution. Naisha Agarwal, Judith Amores, Andrew D. Wilson |
CHI | 3 |
| 2026 | Uncertain Pointer: Situated Feedforward Visualizations for Ambiguity-Aware AR Target SelectionabstractTarget disambiguation is crucial in resolving input ambiguity in augmented reality (AR), especially for queries over distant objects or cluttered scenes on the go. Yet, visual feedforward techniques that support this process remain underexplored. We present Uncertain Pointer, a systematic exploration of feedforward visualizations that annotate multiple candidate targets before user confirmation, either by adding distinct visual identities (e.g., colors) to support disambiguation or by modulating visual intensity (e.g., opacity) to convey system uncertainty. First, we construct a pointer space of 25 pointers by analyzing existing placement strategies and visual signifiers used in target visualizations across 30 years of relevant literature. We then evaluate them through two online experiments (n = 60 and 40), measuring user preference, confidence, mental ease, target visibility, and identifiability across varying object distances and sparsities. Finally, from the results, we derive design recommendations in choosing different Uncertain Pointers based on AR context and disambiguation techniques. Ching-Yi Tsai, Nicole Tacconi, Andrew D. Wilson, Parastoo Abtahi |
CHI | 3 |
| 2025 | Sonora: Human-AI Co-Creation of 3D Audio Worlds and its Impact on Anxiety and Cognitive LoadabstractCHI ’25, Yokohama, Japan Fernanda De La Torre, Javier Hernandez, Andrew D. Wilson, Judith Amores |
CHI | 3 |
| 2025 | Out of Sight, Not Out of Context? Egocentric Spatial Reasoning in VLMs Across Disjoint FramesabstractAn embodied AI assistant operating on egocentric video must integrate spatial cues across time - for instance, determining where an object A, glimpsed a few moments ago lies relative to an object B encountered later. We introduce Disjoint-3DQA , a generative QA benchmark that evaluates this ability of VLMs by posing questions about object pairs that are not co-visible in the same frame. We evaluated seven state-of-the-art VLMs and found that models lag behind human performance by 28%, with steeper declines in accuracy (60% → 30 %) as the temporal gap widens. Our analysis further reveals that providing trajectories or bird’s-eye-view projections to VLMs results in only marginal improvements, whereas providing oracle 3D coordinates leads to a substantial 20% performance increase. This highlights a core bottleneck of multi-frame VLMs in constructing and maintaining 3D scene representations over time from visual signals. Disjoint-3DQA therefore sets a clear, measurable challenge for long-horizon spatial reasoning and aims to catalyze future research at the intersection of vision, language, and embodied AI. Sahithya Ravi, Gabriel Sarch, Vibhav Vineet, Andrew D. Wilson, Balasaravanan Thoravi Kumaravel |
EMNLP | 4 |
| 2025 | ImaginationVellum: Generative-AI Ideation Canvas with Spatial Prompts, Generative Strokes, and Ideation History
Nicolai Marquardt, Asta Roseway, Hugo Romat, Payod Panda, Michel Pahud, Gonzalo A. Ramos, Steven Mark Drucker, Andrew D. Wilson, Ken Hinckley, Nathalie Henry Riche |
UIST | 8 |
| 2024 | SharedNeRF: Leveraging Photorealistic and View-dependent Rendering for Real-time and Remote CollaborationabstractCollaborating around physical objects necessitates examining different aspects of design or hardware in detail when reviewing or inspecting physical artifacts or prototypes. When collaborators are remote, coordinating the sharing of views of their physical environment becomes challenging. Video-conferencing tools often do not provide the desired viewpoints for a remote viewer. While RGB-D cameras offer 3D views, they lack the necessary fidelity. We introduce SharedNeRF, designed to enhance synchronous remote collaboration by leveraging the photorealistic and view-dependent nature of Neural Radiance Field (NeRF). The system complements the higher visual quality of the NeRF rendering with the instantaneity of a point cloud and combines them through carefully accommodating the dynamic elements within the shared space, such as hand gestures and moving objects. The system employs a head-mounted camera for data collection, creating a volumetric task space on the fly and updating it as the task space changes. In our preliminary study, participants successfully completed a flower arrangement task, benefiting from SharedNeRF’s ability to render the space in high fidelity from various viewpoints. Mose Sakashita, Balasaravanan Thoravi Kumaravel, Nicolai Marquardt, Andrew D. Wilson |
CHI | 4 |
| 2024 | SpaceBlender: Creating Context-Rich Collaborative Spaces Through Generative 3D Scene BlendingabstractThere is increased interest in using generative AI to create 3D spaces for Virtual Reality (VR) applications. However, today’s models produce artificial environments, falling short of supporting collaborative tasks that benefit from incorporating the user’s physical context. To generate environments that support VR telepresence, we introduce SpaceBlender, a novel pipeline that utilizes generative AI techniques to blend users’ physical surroundings into unified virtual spaces. This pipeline transforms user-provided 2D images into context-rich 3D environments through an iterative process consisting of depth estimation, mesh alignment, and diffusion-based space completion guided by geometric priors and adaptive text prompts. In a preliminary within-subjects study, where 20 participants performed a collaborative VR affinity diagramming task in pairs, we compared SpaceBlender with a generic virtual environment and a state-of-the-art scene generation framework, evaluating its ability to create virtual spaces suitable for collaboration. Participants appreciated the enhanced familiarity and context provided by SpaceBlender but also noted complexities in the generative environments that could detract from task focus. Drawing on participant feedback, we propose directions for improving the pipeline and discuss the value and design of blended spaces for different scenarios. Nels Numan, Shwetha Rajaram, Balasaravanan Thoravi Kumaravel, Nicolai Marquardt, Andrew D. Wilson |
UIST | 5 |
| 2024 | BlendScape: Enabling End-User Customization of Video-Conferencing Environments through Generative AIabstractToday’s video-conferencing tools support a rich range of professional and social activities, but their generic meeting environments cannot be dynamically adapted to align with distributed collaborators’ needs. To enable end-user customization, we developed BlendScape, a rendering and composition system for video-conferencing participants to tailor environments to their meeting context by leveraging AI image generation techniques. BlendScape supports flexible representations of task spaces by blending users’ physical or digital backgrounds into unified environments and implements multimodal interaction techniques to steer the generation. Through an exploratory study with 15 end-users, we investigated whether and how they would find value in using generative AI to customize video-conferencing environments. Participants envisioned using a system like BlendScape to facilitate collaborative activities in the future, but required further controls to mitigate distracting or unrealistic visual elements. We implemented scenarios to demonstrate BlendScape’s expressiveness for supporting environment design strategies from prior work and propose composition techniques to improve the quality of environments. Shwetha Rajaram, Nels Numan, Balasaravanan Thoravi Kumaravel, Nicolai Marquardt, Andrew D. Wilson |
UIST | 5 |
| 2024 | Hybridge: Bridging Spatiality for Inclusive and Equitable Hybrid MeetingsabstractHybrid meetings limit inclusion for remote participants. The Hybridge experimental system provides different interfaces for remote and room endpoints, focusing on improving inclusion via shared spatiality and remote agency. In-room participants see remotes on displays around a table, and remotes see video integrated into a digital twin. Remotes can choose where to appear and from where they view the room. We tested Hybridge in a within-subjects study of group survival tasks. An in-person condition was followed by a counterbalanced order of hybrid traditional videoconferencing ("Gallery") and Hybridge. We found that co-presence and agency differences between in-room and remotes were alleviated in Hybridge but remained in Gallery. Physical presence for remotes was higher in Hybridge than Gallery. Conversation flow was better in Hybridge than Gallery, but ease of awareness was not different. We argue that asymmetry should be embraced when designing hybrid meeting systems, with inclusivity achieved by tailoring features for the needs of different endpoints. Payod Panda, Lev Tankelevitch, Becky Spittle, Kori Inkpen, John C. Tang, Sasa Junuzovic, Qianqian Qi 0008, Pat Sweeney, Andrew D. Wilson, William Buxton, Abigail Sellen, Sean Rintel |
Proc. ACM Hum. Comput. Interact. | 9 |
| 2023 | Spatialized Audio and Hybrid Video Conferencing: Where Should Voices be Positioned for People in the Room and Remote Headset Users?abstractHybrid video calls include attendees in a conference room with loudspeakers and remote attendees using headsets, each with different options for rendering sound spatially. Two studies explored the listener experience with spatial audio in video calls. One study examined the in-room experience using loudspeakers, comparing among spatialization algorithms spreading voices out horizontally. A second study compared varying degrees of horizontal separation of binaurally rendered voices for a remote participant using a headset. In-room participants preferred the widest spatialization over monophonic, stereo, and stereo-binary audio in metrics related to intelligibility and helpfulness. Remote participants preferred different widths of the audio stage depending on the number of voices. In both studies, rendering sound spatially increased performance in speech stream identification. Results indicate spatial audio benefits for in-room and remote attendees in video calls, although the in-room attendees accepted a wider audio stage than remote users. Jeremy Hyrkas, Andrew D. Wilson, John C. Tang, Hannes Gamper, Hong Sodoma, Lev Tankelevitch, Kori Inkpen, Shreya Chappidi, Brennan Jones |
CHI | 2 |
| 2023 | Embodying Physics-Aware Avatars in Virtual RealityabstractEmbodiment toward an avatar in virtual reality (VR) is generally stronger when there is a high degree of alignment between the user’s and self-avatar’s motion. However, one-to-one mapping between the two is not always ideal when user interacts with the virtual environment. On these occasions, the user input often leads to unnatural behavior without physical realism (e.g., objects penetrating virtual body, body unmoved by hitting stimuli). We investigate how adding physics correction to self-avatar motion impacts embodiment. Physics-aware self-avatar preserves the physical meaning of the movement but introduces discrepancies between the user’s and self-avatar’s motion, whose contingency is a determining factor for embodiment. To understand its impact, we conducted an in-lab study (n = 20) where participants interacted with obstacles on their upper bodies in VR with and without physics correction. Our results showed that, rather than compromising embodiment level, physics-responsive self-avatar improved embodiment compared to no-physics condition in both active and passive interactions. Yujie Tao, Cheng Yao Wang, Andrew D. Wilson, Eyal Ofek, Mar González-Franco |
CHI | 3 |
| 2023 | Reality Distortion Room: A Study of User Locomotion Responses to Spatial Augmented Reality EffectsabstractReality Distortion Room (RDR) is a proof-of-concept augmented reality system using projection mapping and unencumbered interaction with the Microsoft RoomAlive system to study a user’s locomotive response to visual effects that seemingly transform the physical room the user is in. This study presents five effects that augment the appearance of a physical room to subtly encourage user motion. Our experiment demonstrates users’ reactions to the different distortion and augmentation effects in a standard living room, with the distortion effects projected as wall grids, furniture holograms, and small particles in the air. The augmented living room can give the impression of becoming elongated, wrapped, shifted, elevated, and enlarged. The study results support the implementation of AR experiences in limited physical spaces by providing an initial understanding of how users can be subtly encouraged to move throughout a room. You-Jin Kim, Andrew D. Wilson, Jennifer Jacobs 0001, Tobias Höllerer |
ISMAR | 2 |
| 2023 | Perspectives: Creating Inclusive and Equitable Hybrid Meeting ExperiencesabstractWith the shift to hybrid meetings in work spaces, there is an increasing need to create a more inclusive hybrid meeting experience where people meeting together in a room interact with those joining remotely. This paper describes a design exploration, implementation, and evaluation of Perspectives, a novel hybrid meeting system that aimed to create an inclusive and equitable space for hybrid meetings. Perspectives digitally composites everyone into a virtual room so that each person has a unique but spatially consistent viewpoint into the meeting. The user study compared Perspectives with three commercially available UX designs for hybrid meetings: Gallery, Together Mode, and Front Row. Results from this study revealed key benefits of Perspectives, including supporting natural interactions, creating a strong sense of co-presence, and reducing cognitive load. Results from the study also helped iterate on the design principles of Perspectives, which offer important insights on supporting hybrid meetings. John C. Tang, Kori Inkpen, Sasa Junuzovic, Keri Mallari, Andrew D. Wilson, Sean Rintel, Shiraz Cupala, Tony Carbary, Abigail Sellen, William Buxton |
Proc. ACM Hum. Comput. Interact. | 5 |
| 2022 | DreamStream: Immersive and Interactive Spectating in VRabstractToday spectating and streaming virtual reality (VR) activities typically involves spectators viewing a 2D stream of the VR user’s view. Streaming 2D videos of the game play is popular and well-supported by platforms such as Twitch. However, the generic streaming of full 3D representations is less explored. Thus, while the VR player’s experience may be fully immersive, spectators are limited to 2D videos. This asymmetry lessens the overall experience for spectators, who themselves may be eager to spectate in VR. DreamStream puts viewers in the virtual environment of the VR application, allowing them to look “over the shoulder” of the VR player. Spectators can view streamed VR content immersively in 3D, independently explore the VR scene beyond what the VR player sees and ultimately cohabit the virtual environment alongside the VR player. For the VR player, DreamStream provides a spatial awareness of all their spectators. DreamStream retrofits and works with existing VR applications. We discuss the design and implementation of DreamStream, and carry out three qualitative informal evaluations. These evaluations shed light on the strengths and weakness of using DreamStream for the purpose of interactive spectating. Our participants found that DreamStream’s VR viewer interface offered increased immersion, and made it easier to communicate and interact with the VR player. Balasaravanan Thoravi Kumaravel, Andrew D. Wilson |
CHI | 2 |
| 2021 | Decoding Music Attention from "EEG Headphones": A User-Friendly Auditory Brain-Computer InterfaceabstractPeople enjoy listening to music as part of their life. This makes music an excellent choice for designing a user-friendly brain-computer interface (BCI) for long-term use. We propose a novel BCI system using music stimuli that relies on brain signals collected via Smartfones, an EEG recording device integrated into a pair of headphones. In a user study of the proposed system, participants were asked to pay attention to one of three musical instruments playing simultaneously from separate spatial directions. We used a stimulus reconstruction method to decode attention from EEG signals. Results show that the proposed system can achieve good decoding accuracy (>70%) while providing superior user-friendliness compared to a traditional EEG setup. Winko W. An, Barbara G. Shinn-Cunningham, Hannes Gamper, Dimitra Emmanouilidou, David Johnston, Mihai Jalobeanu, Edward Cutrell, Andrew D. Wilson, Kuan-Jung Chiang, Ivan Tashev |
ICASSP | 8 |
| 2020 | Combating the Spread of Coronavirus by Modeling Fomites with Depth CamerasabstractCoronavirus is thought to spread through close contact from person to person. While it is believed that the primary means of spread is by inhaling respiratory droplets or aersols, it may also be spread by touching inanimate objects such as doorknobs and handrails that have the virus on it ("fomites''). The Centers for Disease Control and Prevention (CDC) therefore recommends individuals maintain "social distance'' of more than six feet between one another. It further notes that an individual may be infected by touching a fomite and then touching their own mouth, nose or possibly their eyes. We propose the use of computer vision techniques to combat the spread of coronavirus by sounding an audible alarm when an individual touches their own face, or when multiple individuals come within six feet of one another or shake hands. We further propose using depth cameras to track where people touch parts of their physical environment throughout the day, and a simple model of disease spread among potential fomites. Projection mapping techniques can be used to display likely fomites in realtime, while headworn augmented reality systems can be used by custodial staff to perform more effective cleaning of surfaces. Such techniques may find application in particularly vulnerable settings such as schools, long-term care facilities and physician offices. Andrew D. Wilson |
Proc. ACM Hum. Comput. Interact. | 1 |
| 2019 | RealityCheck: Blending Virtual Environments with Situated Physical RealityabstractToday's virtual reality (VR) systems offer chaperone rendering techniques that prevent the user from colliding with physical objects. Without a detailed geometric model of the physical world, these techniques offer limited possibility for more advanced compositing between the real world and the virtual. We explore this using a realtime 3D reconstruction of the real world that can be combined with a virtual environment. RealityCheck allows users to freely move, manipulate, observe, and communicate with people and objects situated in their physical space without losing the sense of immersion or presence inside their virtual world. We demonstrate RealityCheck with seven existing VR titles, and describe compositing approaches that address the potential conflicts when rendering the real world and a virtual environment together. A study with frequent VR users demonstrate the affordances provided by our system and how it can be used to enhance current VR experiences. Jeremy Hartmann, Christian Holz 0001, Eyal Ofek, Andrew D. Wilson |
CHI | 4 |
| 2019 | SeeingVR: A Set of Tools to Make Virtual Reality More Accessible to People with Low VisionabstractCurrent virtual reality applications do not support people who have low vision, i.e., vision loss that falls short of complete blindness but is not correctable by glasses. We present SeeingVR, a set of 14 tools that enhance a VR application for people with low vision by providing visual and audio augmentations. A user can select, adjust, and combine different tools based on their preferences. Nine of our tools modify an existing VR application post hoc via a plugin without developer effort. The rest require simple inputs from developers using a Unity toolkit we created that allows integrating all 14 of our low vision support tools during development. Our evaluation with 11 participants with low vision showed that SeeingVR enabled users to better enjoy VR and complete tasks more quickly and accurately. Developers also found our Unity toolkit easy and convenient to use. Yuhang Zhao 0001, Edward Cutrell, Christian Holz 0001, Meredith Ringel Morris, Eyal Ofek, Andrew D. Wilson |
CHI | 6 |
| 2019 | Mise-Unseen: Using Eye Tracking to Hide Virtual Reality Scene Changes in Plain SightabstractCreating or arranging objects at runtime is needed in many virtual reality applications, but such changes are noticed when they occur inside the user's field of view. We present Mise-Unseen, a software system that applies such scene changes covertly inside the user's field of view. Mise-Unseen leverages gaze tracking to create models of user attention, intention, and spatial memory to determine if and when to inject a change. We present seven applications of Mise-Unseen to unnoticeably modify the scene within view (i) to hide that task difficulty is adapted to the user, (ii) to adapt the experience to the user's preferences, (iii) to time the use of low fidelity effects, (iv) to detect user choice for passive haptics even when lacking physical props, (v) to sustain physical locomotion despite a lack of physical space, (vi) to reduce motion sickness during virtual locomotion, and (vii) to verify user understanding during story progression. We evaluated Mise-Unseen and our applications in a user study with 15 participants and find that while gaze data indeed supports obfuscating changes inside the field of view, a change is rendered unnoticeably by using gaze in combination with common masking techniques. Sebastian Marwecki, Andrew D. Wilson, Eyal Ofek, Mar González-Franco, Christian Holz 0001 |
UIST | 2 |
| 2019 | DreamWalker: Substituting Real-World Walking Experiences with a Virtual RealityabstractWe explore a future in which people spend considerably more time in virtual reality, even during moments when they transition between locations in the real world. In this paper, we present DreamWalker, a VR system that enables such real-world walking while users explore and stay fully immersed inside large virtual environments in a headset. Provided with a real-world destination, DreamWalker finds a similar path in a pre-authored VR environment and guides the user while real-walking the virtual world. To keep the user from colliding with objects and people in the real-world, DreamWalker's tracking system fuses GPS locations, inside-out tracking, and RGBD frames to 1) continuously and accurately position the user in the real world, 2) sense walkable paths and obstacles in real time, and 3) represent paths through a dynamically changing scene in VR to redirect the user towards the chosen destination. We demonstrate DreamWalker's versatility by enabling users to walk three paths across the large Microsoft campus while enjoying pre-authored VR worlds, supplemented with a variety of obstacle avoidance and redirection techniques. In our evaluation, 8 participants walked across campus along a 15-minute route, experiencing a lively virtual Manhattan that was full of animated cars, people, and other objects. Jackie Yang, Christian Holz 0001, Eyal Ofek, Andrew D. Wilson |
UIST | 4 |
| 2019 | VRoamer: Generating On-The-Fly VR Experiences While Walking inside Large, Unknown Real-World Building EnvironmentsabstractProcedural generation in virtual reality (VR) has been used to adapt the virtual world to various indoor environments, fitting different geometries and interiors with virtual environments. However, such applications require that the physical environment be known or pre-scanned prior to use to then generate the corresponding virtual scene, thus restricting the virtual experience to a controlled space. In this paper, we present VRoamer, which enables users to walk unseen physical spaces for which VRoamer procedurally generates a virtual scene on-the-fly. Scaling to the size of office buildings, VRoamer extracts walkable areas and detects physical obstacles in real time, instantiates pre-authored virtual rooms if their sizes fit physically walkable areas or otherwise generates virtual corridors and doors that lead to undiscovered physical areas. The use of these virtual structures allows VRoamer to (1) temporarily block users' passage, thus slowing them down while increasing VRoamer's insight into newly discovered physical areas, (2) prevent users from seeing changes beyond the current virtual scene, and (3) obfuscate the appearance of physical environments. VRoamer animates virtual objects to reflect dynamically discovered changes of the physical environment, such as people walking by or obstacles that become apparent. In our proof-of-concept study, participants were able to walk long distances through a procedurally generated dungeon experience and reported high levels of immersion. Lung-Pan Cheng, Eyal Ofek, Christian Holz 0001, Andrew D. Wilson |
VR | 4 |
| 2018 | Remixed Reality: Manipulating Space and Time in Augmented RealityabstractWe present Remixed Reality, a novel form of mixed reality. In contrast to classical mixed reality approaches where users see a direct view or video feed of their environment, with Remixed Reality they see a live 3D reconstruction, gathered from multiple external depth cameras. This approach enables changing the environment as easily as geometry can be changed in virtual reality, while allowing users to view and interact with the actual physical world as they would in augmented reality. We characterize a taxonomy of manipulations that are possible with Remixed Reality: spatial changes such as erasing objects; appearance changes such as changing textures; temporal changes such as pausing time; and viewpoint changes that allow users to see the world from different points without changing their physical location. We contribute a method that uses an underlying voxel grid holding information like visibility and transformations, which is applied to live geometry in real time. David Lindlbauer, Andrew D. Wilson |
CHI | 2 |
| 2018 | Autopager: exploiting change blindness for gaze-assisted readingabstractA novel gaze-assisted reading technique uses the fact that in linear reading, the looking behavior of the reader is readily predicted. We introduce the AutoPager "page turning" technique, where the next bit of unread text is rendered in the periphery, ready to be read. This approach enables continuous gaze-assisted reading without requiring manual input to scroll: the reader merely saccades to the top of the page to read on. We demonstrate that when the new text is introduced with a gradual cross-fade effect, users are often unaware of the change: the user's impression is of reading the same page over and over again, yet the content changes. We present a user evaluation that compares AutoPager to previous gaze-assisted scrolling techniques. AutoPager may offer some advantages over previous gaze-assisted reading techniques, and is a rare example of exploiting "change blindness" in user interfaces. Andrew D. Wilson, Shane Williams |
ETRA | 1 |
| 2018 | MRTouch: Adding Touch Input to Head-Mounted Mixed RealityabstractWe present MRTouch, a novel multitouch input solution for head-mounted mixed reality systems. Our system enables users to reach out and directly manipulate virtual interfaces affixed to surfaces in their environment, as though they were touchscreens. Touch input offers precise, tactile and comfortable user input, and naturally complements existing popular modalities, such as voice and hand gesture. Our research prototype combines both depth and infrared camera streams together with real-time detection and tracking of surface planes to enable robust finger-tracking even when both the hand and head are in motion. Our technique is implemented on a commercial Microsoft HoloLens without requiring any additional hardware nor any user or environmental calibration. Through our performance evaluation, we demonstrate high input accuracy with an average positional error of 5.4 mm and 95% button size of 16 mm, across 17 participants, 2 surface orientations and 4 surface materials. Finally, we demonstrate the potential of our technique to enable on-world touch interactions through 5 example applications. Robert Xiao, Julia Schwarz, Nick Throm, Andrew D. Wilson, Hrvoje Benko |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2017 | Sparse Haptic Proxy: Touch Feedback in Virtual Environments Using a General Passive PropabstractWe propose a class of passive haptics that we call Sparse Haptic Proxy: a set of geometric primitives that simulate touch feedback in elaborate virtual reality scenes. Unlike previous passive haptics that replicate the virtual environment in physical space, a Sparse Haptic Proxy simulates a scene's detailed geometry by redirecting the user's hand to a matching primitive of the proxy. To bridge the divergence of the scene from the proxy, we augment an existing Haptic Retargeting technique with an on-the-fly target remapping: We predict users' intentions during interaction in the virtual space by analyzing their gaze and hand motions, and consequently redirect their hand to a matching part of the proxy. We conducted three user studies on haptic retargeting technique and implemented a system from three main results: 1) The maximum angle participants found acceptable for retargeting their hand is 40°, with an average rating of 4.6 out of 5. 2) Tracking participants' eye gaze reliably predicts their touch intentions (97.5%), even while simultaneously manipulating the user's hand-eye coordination for retargeting. 3) Participants preferred minimized retargeting distances over better-matching surfaces of our Sparse Haptic Proxy when receiving haptic feedback for single-finger touch input. We demonstrate our system with two virtual scenes: a flight cockpit and a room quest game. While their scene geometries differ substantially, both use the same sparse haptic proxy to provide haptic feedback to the user during task completion. Lung-Pan Cheng, Eyal Ofek, Christian Holz 0001, Hrvoje Benko, Andrew D. Wilson |
CHI | 5 |
| 2017 | Hybrid HFR Depth: Fusing Commodity Depth and Color Cameras to Achieve High Frame Rate, Low Latency Depth Camera InteractionsabstractThe low frame rate and high latency of consumer depth cameras limits their use in interactive applications. We propose combining the Kinect depth camera with an ordinary color camera to synthesize a high frame rate and low latency depth image. We exploit common CMOS camera region of interest (ROI) functionality to obtain a high frame rate image over a small ROI. Motion in the ROI is computed by a fast optical flow implementation. The resulting flow field is used to extrapolate Kinect depth images to achieve high frame rate and low latency depth, and optionally predict depth to further reduce latency. Our "Hybrid HFR Depth" prototype generates useful depth images at maximum 500Hz with minimum 20ms latency. We demonstrate Hybrid HFR Depth in tracking fast moving objects, handwriting in the air, and projecting onto moving hands. Based on commonly available cameras and image processing implementations, Hybrid HFR Depth may be useful to HCI practitioners seeking to create fast, fluid depth camera-based interactions. Jiajun Lu, Hrvoje Benko, Andrew D. Wilson |
CHI | 3 |
| 2017 | MeetAlive: Room-Scale Omni-Directional Display System for Multi-User Content and Control SharingabstractMeetAlive combines multiple depth cameras and projectors to create a room-scale omni-directional display surface designed to support collaborative face-to-face group meetings. With MeetAlive, all participants may simultaneously display and share content from their personal laptop wirelessly anywhere in the room. MeetAlive gives each participant complete control over displayed content in the room. This is achieved by a perspective corrected mouse cursor that transcends the boundary of the laptop screen to position, resize, and edit their own and others' shared content. MeetAlive includes features to replicate content views to ensure that all participants may see the actions of other participants even as they are seated around a conference table. We report on observing six groups of three participants who worked on a collaborative task with minimal assistance. Participants' feedback highlighted the value of MeetAlive features for multi-user engagement in meetings involving brainstorming and content creation. Andreas Rene Fender, Hrvoje Benko, Andrew D. Wilson |
ISS | 3 |
| 2017 | Fast Lossless Depth Image CompressionabstractA lossless image compression technique for 16-bit single channel images typical of depth cameras such as Microsoft Kinect is presented. The proposed "RVL" algorithm achieves similar or better compression rates as existing lossless techniques, yet is much faster. Furthermore, the algorithm's implementation can be very simple; a prototype implementation of less than one hundred lines of C is provided. The algorithm's balance of speed and compression make it especially useful in interactive applications of multiple depth cameras on local area networks. RVL is compared to a variety of existing lossless techniques, and demonstrated in a network of eight Kinect v2 cameras. Andrew D. Wilson |
ISS | 1 |
| 2017 | Dynamic Task Execution Using Active Parameter Identification With the Baxter Research RobotabstractThis paper presents experimental results from the real-time parameter estimation of a system model and subsequent trajectory optimization for a dynamic task using the Baxter Research Robot from Rethink Robotics. An active estimator maximizing Fisher information is used in real time with a closed-loop, nonlinear control technique known as sequential action control. Baxter is tasked with estimating the length of a string connected to a load suspended from the gripper with a load cell providing the single source of feedback to the estimator. Following the active estimation, a trajectory is generated using the trep software package that controls Baxter to dynamically swing a suspended load into a box. Several trials are presented with varying initial estimates showing that the estimation is required to obtain adequate open-loop trajectories to complete the prescribed task. The result of one trial with and without the active estimation is also shown in the accompanying video. Andrew D. Wilson, Jarvis A. Schultz, Alex Ansari, Todd D. Murphey |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2016 | Haptic Retargeting: Dynamic Repurposing of Passive Haptics for Enhanced Virtual Reality ExperiencesabstractManipulating a virtual object with appropriate passive haptic cues provides a satisfying sense of presence in virtual reality. However, scaling such experiences to support multiple virtual objects is a challenge as each one needs to be accompanied with a precisely-located haptic proxy object. We propose a solution that overcomes this limitation by hacking human perception. We have created a framework for repurposing passive haptics, called haptic retargeting, that leverages the dominance of vision when our senses conflict. With haptic retargeting, a single physical prop can provide passive haptics for multiple virtual objects. We introduce three approaches for dynamically aligning physical and virtual objects: world manipulation, body manipulation and a hybrid technique which combines both world and body manipulation. Our study results indicate that all our haptic retargeting techniques improve the sense of presence when compared to typical wand-based 3D control of virtual objects. Furthermore, our hybrid haptic retargeting achieved the highest satisfaction and presence scores while limiting the visible side-effects during interaction. Mahdi Azmandian, Mark S. Hancock, Hrvoje Benko, Eyal Ofek, Andrew D. Wilson |
CHI | 5 |
| 2016 | SnapToReality: Aligning Augmented Reality to the Real WorldabstractAugmented Reality (AR) applications may require the precise alignment of virtual objects to the real world. We propose automatic alignment of virtual objects to physical constraints calculated from the real world in real time ("snapping to reality"). We demonstrate SnapToReality alignment techniques that allow users to position, rotate, and scale virtual content to dynamic, real world scenes. Our proof-of-concept prototype extracts 3D edge and planar surface constraints. We furthermore discuss the unique design challenges of snapping in AR, including the user's limited field of view, noise in constraint extraction, issues with changing the view in AR, visualizing constraints, and more. We also report the results of a user study evaluating SnapToReality, confirming that aligning objects to the real world is significantly faster when assisted by snapping to dynamically extracted constraints. Perhaps more importantly, we also found that snapping in AR enables a fresh and expressive form of AR content creation. Benjamin Nuernberger, Eyal Ofek, Hrvoje Benko, Andrew D. Wilson |
CHI | 4 |
| 2016 | Room2Room: Enabling Life-Size Telepresence in a Projected Augmented Reality EnvironmentabstractRoom2Room is a telepresence system that leverages projected augmented reality to enable life-size, co-present interaction between two remote participants. Our solution recreates the experience of a face-to-face conversation by performing 3D capture of the local user with color + depth cameras and projecting their life-size virtual copy into the remote space. This creates an illusion of the remote person's physical presence in the local space, as well as a shared understanding of verbal and non-verbal cues (e.g., gaze, pointing.) In addition to the technical details of two prototype implementations, we contribute strategies for projecting remote participants onto physically plausible locations, such that they form a natural and consistent conversational formation with the local participant. We also present observations and feedback from an evaluation with 7 pairs of participants on the usability of our solution for solving a collaborative, physical task. Tomislav Pejsa, Julian Kantor, Hrvoje Benko, Eyal Ofek, Andrew D. Wilson |
CSCW | 5 |
| 2016 | A Demonstration of Haptic Retargeting: Dynamic Repurposing of Passive Haptics for Enhanced Virtual Reality ExperiencesabstractManipulating a virtual object with appropriate passive haptic cues provides a satisfying sense of presence in virtual reality. However, scaling such experiences to support multiple virtual objects is a challenge as each one needs to be accompanied with a precisely-located haptic proxy object. We showcase a solution that overcomes this limitation by hacking human perception. Our framework for repurposing passive haptics, called haptic retargeting, leverages the dominance of vision when our senses conflict. With haptic retargeting, a single physical prop can provide passive haptics for multiple virtual objects. We introduce three approaches for dynamically aligning physical and virtual objects: body manipulation, world manipulation and a hybrid technique which combines both world and body manipulation. This demo has been presented previously at CHI 2016. Mahdi Azmandian, Mark S. Hancock, Hrvoje Benko, Eyal Ofek, Andrew D. Wilson |
ISS | 5 |
| 2016 | Projected Augmented Reality with the RoomAlive ToolkitabstractThe RoomAlive Toolkit is an open source SDK that enables developers to create interactive projection mapping applications. The toolkit focuses on calibrating a network of multiple Kinect sensors and video projectors. It also provides a simple projection mapping sample that can be used as a basis to develop new immersive augmented reality experiences similar to those of the IllumiRoom [2] and RoomAlive [3] research projects. In this tutorial we will describe the RoomAlive Toolkit in detail, demonstrate its basic use and discuss integration with Unity3D. Andrew D. Wilson, Hrvoje Benko |
ISS | 1 |
| 2015 | Maximizing fisher information using discrete mechanics and projection-based trajectory optimizationabstractThis paper reformulates an optimization algorithm previously presented in continuous-time to one using structured integration and structured linearization methods from discrete mechanics. The objective is to synthesize trajectories for dynamic robotic systems that improve the estimation of model parameters by using a metric on Fisher information in a nonlinear projection-based trajectory optimization algorithm. A simulation of a robot with a suspended double pendulum is used as an example system to illustrate the algorithm. Results from the simulation show that the change to a discrete mechanics formulation reduces the computation time by a factor of 19 when compared to the continuous algorithm while maintaining the same two orders of magnitude improvement in the Fisher information from the continuous-time formulation. Through the Cramer-Rao bound, the improvement in the Fisher information results in a maximum expected error reduction of the parameter estimates by up to a factor of 102. Andrew D. Wilson, Todd D. Murphey |
ICRA | 1 |
| 2015 | Real-time trajectory synthesis for information maximization using Sequential Action Control and least-squares estimationabstractThis paper presents the details and experimental results from an implementation of real-time trajectory generation and parameter estimation of a dynamic model using the Baxter Research Robot from Rethink Robotics. Trajectory generation is based on the maximization of Fisher information in real-time and closed-loop using a form of Sequential Action Control. On-line estimation is performed with a least-squares estimator employing a nonlinear state observer model computed with trep, a dynamics simulation package. Baxter is tasked with estimating the length of a string connected to a load suspended from the gripper with a load cell providing the single source of feedback to the estimator. Several trials are presented with varying initial estimates showing convergence to the actual length within a 6 second time-frame. Andrew D. Wilson, Jarvis A. Schultz, Alex Ansari, Todd D. Murphey |
IROS | 1 |
| 2015 | FoveAR: Combining an Optically See-Through Near-Eye Display with Projector-Based Spatial Augmented RealityabstractOptically see-through (OST) augmented reality glasses can overlay spatially-registered computer-generated content onto the real world. However, current optical designs and weight considerations limit their diagonal field of view to less than 40 degrees, making it difficult to create a sense of immersion or give the viewer an overview of the augmented reality space. We combine OST glasses with a projection-based spatial augmented reality display to achieve a novel display hybrid, called FoveAR, capable of greater than 100 degrees field of view, view dependent graphics, extended brightness and color, as well as interesting combinations of public and personal data display. We contribute details of our prototype implementation and an analysis of the interactive design space that our system enables. We also contribute four prototype experiences showcasing the capabilities of FoveAR as well as preliminary user feedback providing insights for enhancing future FoveAR experiences. Hrvoje Benko, Eyal Ofek, Andrew D. Wilson |
UIST | 4 |
| 2015 | Trajectory Optimization for Well-Conditioned Parameter EstimationabstractWhen attempting to estimate parameters in a dynamical system, it is often beneficial to strategically design experimental trajectories that facilitate the estimation process. This paper presents an optimization algorithm which improves conditioning of estimation problems by modifying the experimental trajectory. An objective function which minimizes the condition number of the Hessian of the least-squares identification method is derived and a least-squares method is used to estimate parameters of the nonlinear system. A software-simulated example demonstrates that an arbitrarily designed trajectory can lead to an ill-conditioned least-squares estimation problem, which in turn leads to slower convergence to the best estimate and, in the presence of experimental uncertainties, may lead to no convergence at all. A physical experiment with a robot-controlled suspended mass also shows improved estimation results in practice in the presence of noise and uncertainty using the optimized trajectory. Andrew D. Wilson, Jarvis A. Schultz, Todd D. Murphey |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2014 | CrossMotion: Fusing Device and Image Motion for User Identification, Tracking and Device AssociationabstractIdentifying and tracking people and mobile devices indoors has many applications, but is still a challenging problem. We introduce a cross-modal sensor fusion approach to track mobile devices and the users carrying them. The CrossMotion technique matches the acceleration of a mobile device, as measured by an onboard internal measurement unit, to similar acceleration observed in the infrared and depth images of a Microsoft Kinect v2 camera. This matching process is conceptually simple and avoids many of the difficulties typical of more common appearance-based approaches. In particular, CrossMotion does not require a model of the appearance of either the user or the device, nor in many cases a direct line of sight to the device. We demonstrate a real time implementation that can be applied to many ubiquitous computing scenarios. In our experiments, CrossMotion found the person's body 99% of the time, on average within 7cm of a reference device position. Andrew D. Wilson, Hrvoje Benko |
ICMI | 1 |
| 2014 | Dyadic projected spatial augmented realityabstractMano-a-Mano is a unique spatial augmented reality system that combines dynamic projection mapping, multiple perspective views and device-less interaction to support face to face, or dyadic, interaction with 3D virtual objects. Its main advantage over more traditional AR approaches, such as handheld devices with composited graphics or see-through head worn displays, is that users are able to interact with 3D virtual objects and each other without cumbersome devices that obstruct face to face interaction. We detail our prototype system and a number of interactive experiences. We present an initial user experiment that shows that participants are able to deduce the size and distance of a virtual projected object. A second experiment shows that participants are able to infer which of a number of targets the other user indicates by pointing. Hrvoje Benko, Andrew D. Wilson, Federico Zannier |
UIST | 2 |
| 2014 | Sensing techniques for tablet+stylus interactionabstractWe explore grip and motion sensing to afford new techniques that leverage how users naturally manipulate tablet and stylus devices during pen + touch interaction. We can detect whether the user holds the pen in a writing grip or tucked between his fingers. We can distinguish bare-handed inputs, such as drag and pinch gestures produced by the nonpreferred hand, from touch gestures produced by the hand holding the pen, which necessarily impart a detectable motion signal to the stylus. We can sense which hand grips the tablet, and determine the screen's relative orientation to the pen. By selectively combining these signals and using them to complement one another, we can tailor interaction to the context, such as by ignoring unintentional touch inputs while writing, or supporting contextually-appropriate tools such as a magnifier for detailed stroke work that appears when the user pinches with the pen tucked between his fingers. These and other techniques can be used to impart new, previously unanticipated subtleties to pen + touch interaction on tablets. Ken Hinckley, Michel Pahud, Hrvoje Benko, Pourang Irani, François Guimbretière, Marcel Gavriliu, Xiang 'Anthony' Chen, Fabrice Matulic, William Buxton, Andrew D. Wilson |
UIST | 10 |
| 2014 | RoomAlive: magical experiences enabled by scalable, adaptive projector-camera unitsabstractRoomAlive is a proof-of-concept prototype that transforms any room into an immersive, augmented entertainment experience. Our system enables new interactive projection mapping experiences that dynamically adapts content to any room. Users can touch, shoot, stomp, dodge and steer projected content that seamlessly co-exists with their existing physical environment. The basic building blocks of RoomAlive are projector-depth camera units, which can be combined through a scalable, distributed framework. The projector-depth camera units are individually auto-calibrating, self-localizing, and create a unified model of the room with no user intervention. We investigate the design space of gaming experiences that are possible with RoomAlive and explore methods for dynamically mapping content based on room layout and user position. Finally we showcase four experience prototypes that demonstrate the novel interactive experiences that are possible with RoomAlive and discuss the design challenges of adapting any game to any room. Brett R. Jones, Rajinder Sodhi, Michael Murdock, Ravish Mehra, Hrvoje Benko, Andrew D. Wilson, Eyal Ofek, Blair MacIntyre, Nikunj Raghuvanshi, Lior Shapira |
UIST | 6 |
| 2014 | Trajectory Synthesis for Fisher Information MaximizationabstractEstimation of model parameters in a dynamic system can be significantly improved with the choice of experimental trajectory. For general nonlinear dynamic systems, finding globally "best" trajectories is typically not feasible; however, given an initial estimate of the model parameters and an initial trajectory, we present a continuous-time optimization method that produces a locally optimal trajectory for parameter estimation in the presence of measurement noise. The optimization algorithm is formulated to find system trajectories that improve a norm on the Fisher information matrix (FIM). A double-pendulum cart apparatus is used to numerically and experimentally validate this technique. In simulation, the optimized trajectory increases the minimum eigenvalue of the FIM by three orders of magnitude, compared with the initial trajectory. Experimental results show that this optimized trajectory translates to an order-of-magnitude improvement in the parameter estimate error in practice. Andrew D. Wilson, Jarvis A. Schultz, Todd D. Murphey |
IEEE Trans. Robotics | 1 |
| 2013 | IllumiRoom: peripheral projected illusions for interactive experiencesabstractIllumiRoom is a proof-of-concept system that augments the area surrounding a television with projected visualizations to enhance traditional gaming experiences. We investigate how projected visualizations in the periphery can negate, include, or augment the existing physical environment and complement the content displayed on the television screen. Peripheral projected illusions can change the appearance of the room, induce apparent motion, extend the field of view, and enable entirely new physical gaming experiences. Our system is entirely self-calibrating and is designed to work in any room. We present a detailed exploration of the design space of peripheral projected illusions and we demonstrate ways to trigger and drive such illusions from gaming content. We also contribute specific feedback from two groups of target users (10 gamers and 15 game designers); providing insights for enhancing game experiences through peripheral projected illusions. Brett R. Jones, Hrvoje Benko, Eyal Ofek, Andrew D. Wilson |
CHI | 4 |
| 2013 | InfraStructs: fabricating information inside physical objects for imaging in the terahertz regionabstractWe introduce InfraStructs , material-based tags that embed information inside digitally fabricated objects for imaging in the Terahertz region. Terahertz imaging can safely penetrate many common materials, opening up new possibilities for encoding hidden information as part of the fabrication process. We outline the design, fabrication, imaging, and data processing steps to fabricate information inside physical objects. Prototype tag designs are presented for location encoding, pose estimation, object identification, data storage, and authentication. We provide detailed analysis of the constraints and performance considerations for designing InfraStruct tags. Future application scenarios range from production line inventory, to customized game accessories, to mobile robotics. Karl D. D. Willis, Andrew D. Wilson |
ACM Trans. Graph. | 2 |
| 2012 | MirageTable: freehand interaction on a projected augmented reality tabletopabstractInstrumented with a single depth camera, a stereoscopic projector, and a curved screen, MirageTable is an interactive system designed to merge real and virtual worlds into a single spatially registered experience on top of a table. Our depth camera tracks the user's eyes and performs a real-time capture of both the shape and the appearance of any object placed in front of the camera (including user's body and hands). This real-time capture enables perspective stereoscopic 3D visualizations to a single user that account for deformations caused by physical objects on the table. In addition, the user can interact with virtual objects through physically-realistic freehand actions without any gloves, trackers, or instruments. We illustrate these unique capabilities through three application examples: virtual 3D model creation, interactive gaming with real and virtual objects, and a 3D teleconferencing experience that not only presents a 3D view of a remote person, but also a seamless 3D shared task space. We also evaluated the user's perception of projected 3D objects in our system, which confirmed that the users can correctly perceive such objects even when they are projected over different background colors and geometries (e.g., gaps, drops). Hrvoje Benko, Ricardo Jota, Andrew D. Wilson |
CHI | 3 |
| 2012 | HoloDesk: direct 3d interactions with a situated see-through displayabstractHoloDesk is an interactive system combining an optical see through display and Kinect camera to create the illusion that users are directly interacting with 3D graphics. A virtual image of a 3D scene is rendered through a half silvered mirror and spatially aligned with the real-world for the viewer. Users easily reach into an interaction volume displaying the virtual image. This allows the user to literally get their hands into the virtual display and to directly interact with an spatially aligned 3D virtual world, without the need for any specialized head-worn hardware or input device. We introduce a new technique for interpreting raw Kinect data to approximate and track rigid (e.g., books, cups) and non-rigid (e.g., hands, paper) physical objects and support a variety of physics-inspired interactions between virtual and real. In particular the algorithm models natural human grasping of virtual objects with more fidelity than previously demonstrated. A qualitative study highlights rich emergent 3D interactions, using hands and real-world objects. The implementation of HoloDesk is described in full, and example application scenarios explored. Finally, HoloDesk is quantitatively evaluated in a 3D target acquisition task, comparing the system with indirect and glasses-based variants. Otmar Hilliges, David Kim 0002, Shahram Izadi, Malte Weiss, Andrew D. Wilson |
CHI | 5 |
| 2012 | Your phone or mine?: fusing body, touch and device sensing for multi-user device-display interactionabstractDetermining who is interacting with a multi-user interactive touch display is challenging. We describe a technique for associating multi-touch interactions to individual users and their accelerometer-equipped mobile devices. Real-time device accelerometer data and depth camera-based body tracking are compared to associate each phone with a particular user, while body tracking and touch contacts positions are compared to associate a touch contact with a specific user. It is then possible to associate touch contacts with devices, allowing for more seamless device-display multi-user interactions. We detail the technique and present a user study to validate and demonstrate a content exchange application using this approach. Mahsan Rofouei, Andrew D. Wilson, A. J. Bernheim Brush, Stewart Tansley |
CHI | 2 |
| 2012 | Phone as a pixel: enabling ad-hoc, large-scale displays using mobile devicesabstractWe present Phone as a Pixel: a scalable, synchronization-free, platform-independent system for creating large, ad-hoc displays from a collection of smaller devices. In contrast to most tiled-display systems, the only requirement for participation is for devices to have an internet connection and a web browser. Thus, most smartphones, tablets, laptops and similar devices can be used. Phone as a Pixel uses a color-transition encoding scheme to identify and locate displays. This approach has several advantages: devices can be arbitrarily arranged (i.e., not in a grid) and infrastructure consists of a single conventional camera. Further, additional devices can join at any time without re-calibration. These are desirable properties to enable collective displays in contexts like sporting events, concerts and political rallies. In this paper we describe our system, show results from proof-of-concept setups, and quantify the performance of our approach on hundreds of displays. Julia Schwarz, David Klionsky, Chris Harrison 0001, Paul H. Dietz, Andrew D. Wilson |
CHI | 5 |
| 2012 | LightGuide: projected visualizations for hand movement guidanceabstractLightGuide is a system that explores a new approach to gesture guidance where we project guidance hints directly on a user's body. These projected hints guide the user in completing the desired motion with their body part which is particularly useful for performing movements that require accuracy and proper technique, such as during exercise or physical therapy. Our proof-of-concept implementation consists of a single low-cost depth camera and projector and we present four novel interaction techniques that are focused on guiding a user's hand in mid-air. Our visualizations are designed to incorporate both feedback and feedforward cues to help guide users through a range of movements. We quantify the performance of LightGuide in a user study comparing each of our on-body visualizations to hand animation videos on a computer display in both time and accuracy. Exceeding our expectations, participants performed movements with an average error of 21.6mm, nearly 85% more accurately than when guided by video. Rajinder Sodhi, Hrvoje Benko, Andrew D. Wilson |
CHI | 3 |
| 2012 | Steerable augmented reality with the beamatronabstractSteerable displays use a motorized platform to orient a projector to display graphics at any point in the room. Often a camera is included to recognize markers and other objects, as well as user gestures in the display volume. Such systems can be used to superimpose graphics onto the real world, and so are useful in a number of augmented reality and ubiquitous computing scenarios. We contribute the Beamatron, which advances steerable displays by drawing on recent progress in depth camera-based interactions. The Beamatron consists of a computer-controlled pan and tilt platform on which is mounted a projector and Microsoft Kinect sensor. While much previous work with steerable displays deals primarily with projecting corrected graphics onto a discrete set of static planes, we describe computational techniques that enable reasoning in 3D using live depth data. We show two example applications that are enabled by the unique capabilities of the Beamatron: an augmented reality game in which a player can drive a virtual toy car around a room, and a ubiquitous computing demo that uses speech and gesture to move projected graphics throughout the room. Andrew D. Wilson, Hrvoje Benko, Shahram Izadi, Otmar Hilliges |
UIST | 1 |
| 2011 | Data miming: inferring spatial object descriptions from human gestureabstractSpeakers often use hand gestures when talking about or describing physical objects. Such gesture is particularly useful when the speaker is conveying distinctions of shape that are difficult to describe verbally. We present data miming---an approach to making sense of gestures as they are used to describe concrete physical objects. We first observe participants as they use gestures to describe real-world objects to another person. From these observations, we derive the data miming approach, which is based on a voxel representation of the space traced by the speaker's hands over the duration of the gesture. In a final proof-of-concept study, we demonstrate a prototype implementation of matching the input voxel representation to select among a database of known physical objects. Christian Holz 0001, Andrew D. Wilson |
CHI | 2 |
| 2011 | OmniTouch: wearable multitouch interaction everywhereabstractOmniTouch is a wearable depth-sensing and projection system that enables interactive multitouch applications on everyday surfaces. Beyond the shoulder-worn system, there is no instrumentation of the user or environment. Foremost, the system allows the wearer to use their hands, arms and legs as graphical, interactive surfaces. Users can also transiently appropriate surfaces from the environment to expand the interactive area (e.g., books, walls, tables). On such surfaces - without any calibration - OmniTouch provides capabilities similar to that of a mouse or touchscreen: X and Y location in 2D interfaces and whether fingers are "clicked" or hovering, enabling a wide variety of interactions. Reliable operation on the hands, for example, requires buttons to be 2.3cm in diameter. Thus, it is now conceivable that anything one can do on today's mobile devices, they could do in the palm of their hand. Chris Harrison 0001, Hrvoje Benko, Andrew D. Wilson |
UIST | 3 |
| 2010 | Pictionaire: supporting collaborative design work by integrating physical and digital artifactsabstractThis paper introduces an interactive tabletop system that enhances creative collaboration across physical and digital artifacts. Pictionaire offers capture, retrieval, annotation, and collection of visual material. It enables multiple designers to fluidly move imagery from the physical to the digital realm; work with found, drawn and captured imagery; organize items into functional collections; and record meeting histories. These benefits are made possible by a large interactive table augmented with high-resolution overhead image capture. Summative evaluations with 16 professionals and four student pairs validated discoverability and utility of interactions, uncovered emergent functionality, and suggested opportunities for transitioning content to and from the table. Björn Hartmann, Meredith Ringel Morris, Hrvoje Benko, Andrew D. Wilson |
CSCW | 4 |
| 2010 | Design and evaluation of interaction models for multi-touch mice
Hrvoje Benko, Shahram Izadi, Andrew D. Wilson, Dan Rosenfeld, Ken Hinckley |
Graphics Interface | 3 |
| 2010 | Understanding users' preferences for surface gestures
Meredith Ringel Morris, Jacob O. Wobbrock, Andrew D. Wilson |
Graphics Interface | 3 |
| 2010 | Pen + touch = new toolsabstractWe describe techniques for direct pen+touch input. We observe people's manual behaviors with physical paper and notebooks. These serve as the foundation for a prototype Microsoft Surface application, centered on note-taking and scrapbooking of materials. Based on our explorations we advocate a division of labor between pen and touch: the pen writes, touch manipulates, and the combination of pen + touch yields new tools. This articulates how our system interprets unimodal pen, unimodal touch, and multimodal pen+touch inputs, respectively. For example, the user can hold a photo and drag off with the pen to create and place a copy; hold a photo and cross it in a freeform path with the pen to slice it in two; or hold selected photos and tap one with the pen to staple them all together. Touch thus unifies object selection with mode switching of the pen, while the muscular tension of holding touch serves as the "glue" that phrases together all the inputs into a unitary multimodal gesture. This helps the UI designer to avoid encumbrances such as physical buttons, persistent modes, or widgets that detract from the user's focus on the workspace. Ken Hinckley, Koji Yatani, Michel Pahud, Nicole Coddington, Jenny Rodenhouse, Andrew D. Wilson, Hrvoje Benko, William Buxton |
UIST | 6 |
| 2010 | A framework for robust and flexible handling of inputs with uncertaintyabstractNew input technologies (such as touch), recognition based input (such as pen gestures) and next-generation interactions (such as inexact interaction) all hold the promise of more natural user interfaces. However, these techniques all create inputs with some uncertainty. Unfortunately, conventional infrastructure lacks a method for easily handling uncertainty, and as a result input produced by these technologies is often converted to conventional events as quickly as possible, leading to a stunted interactive experience. We present a framework for handling input with uncertainty in a systematic, extensible, and easy to manipulate fashion. To illustrate this framework, we present several traditional interactors which have been extended to provide feedback about uncertain inputs and to allow for the possibility that in the end that input will be judged wrong (or end up going to a different interactor). Our six demonstrations include tiny buttons that are manipulable using touch input, a text box that can handle multiple interpretations of spoken input, a scrollbar that can respond to inexactly placed input, and buttons which are easier to click for people with motor impairments. Our framework supports all of these interactions by carrying uncertainty forward all the way through selection of possible target interactors, interpretation by interactors, generation of (uncertain) candidate actions to take, and a mediation process that decides (in a lazy fashion) which actions should become final. Julia Schwarz, Scott E. Hudson, Jennifer Mankoff, Andrew D. Wilson |
UIST | 4 |
| 2010 | Combining multiple depth cameras and projectors for interactions on, above and between surfacesabstractInstrumented with multiple depth cameras and projectors, LightSpace is a small room installation designed to explore a variety of interactions and computational strategies related to interactive displays and the space that they inhabit. LightSpace cameras and projectors are calibrated to 3D real world coordinates, allowing for projection of graphics correctly onto any surface visible by both camera and projector. Selective projection of the depth camera data enables emulation of interactive displays on un-instrumented surfaces (such as a standard table or office desk), as well as facilitates mid-air interactions between and around these displays. For example, after performing multi-touch interactions on a virtual object on the tabletop, the user may transfer the object to another display by simultaneously touching the object and the destination display. Or the user may "pick up" the object by sweeping it into their hand, see it sitting in their hand as they walk over to an interactive wall display, and "drop" the object onto the wall by touching it with their other hand. We detail the interactions and algorithms unique to LightSpace, discuss some initial observations of use and suggest future directions. Andrew D. Wilson, Hrvoje Benko |
UIST | 1 |
| 2009 | User-defined gestures for surface computingabstractMany surface computing prototypes have employed gestures created by system designers. Although such gestures are appropriate for early investigations, they are not necessarily reflective of user behavior. We present an approach to designing tabletop gestures that relies on eliciting gestures from non-technical users by first portraying the effect of a gesture, and then asking users to perform its cause. In all, 1080 gestures from 20 participants were logged, analyzed, and paired with think-aloud data for 27 commands performed with 1 and 2 hands. Our findings indicate that users rarely care about the number of fingers they employ, that one hand is preferred to two, that desktop idioms strongly influence users' mental models, and that some commands elicit little gestural agreement, suggesting the need for on-screen widgets. We also present a complete user-defined gesture set, quantitative agreement scores, implications for surface technology, and a taxonomy of surface gestures. Our results will help designers create better gesture sets informed by user behavior. Jacob O. Wobbrock, Meredith Ringel Morris, Andrew D. Wilson |
CHI | 3 |
| 2009 | Augmenting interactive tables with mice & keyboardsabstractThis note examines the role traditional input devices can play in surface computing. Mice and keyboards can enhance tabletop technologies since they support high fidelity input, facilitate interaction with distant objects, and serve as a proxy for user identity and position. Interactive tabletops, in turn, can enhance the functionality of traditional input devices: they provide spatial sensing, augment devices with co-located visual content, and support connections among a plurality of devices. We introduce eight interaction techniques for a table with mice and keyboards, and we discuss the design space of such interactions. Björn Hartmann, Meredith Ringel Morris, Hrvoje Benko, Andrew D. Wilson |
UIST | 4 |
| 2009 | Interactions in the air: adding further depth to interactive tabletopsabstractAlthough interactive surfaces have many unique and compelling qualities, the interactions they support are by their very nature bound to the display surface. In this paper we present a technique for users to seamlessly switch between interacting on the tabletop surface to above it. Our aim is to leverage the space above the surface in combination with the regular tabletop display to allow more intuitive manipulation of digital content in three-dimensions. Our goal is to design a technique that closely resembles the ways we manipulate physical objects in the real-world; conceptually, allowing virtual objects to be 'picked up' off the tabletop surface in order to manipulate their three dimensional position or orientation. We chart the evolution of this technique, implemented on two rear projection-vision tabletops. Both use special projection screen materials to allow sensing at significant depths beyond the display. Existing and new computer vision techniques are used to sense hand gestures and postures above the tabletop, which can be used alongside more familiar multi-touch interactions. Interacting above the surface in this way opens up many interesting challenges. In particular it breaks the direct interaction metaphor that most tabletops afford. We present a novel shadow-based technique to help alleviate this issue. We discuss the strengths and limitations of our technique based on our own observations and initial user feedback, and provide various insights from comparing, and contrasting, our tabletop implementations Otmar Hilliges, Shahram Izadi, Andrew D. Wilson, Steve Hodges 0001, Armando Garcia-Mendoza, Andreas Butz |
UIST | 3 |
| 2009 | Synchronous Gestures in Multi-Display EnvironmentsabstractSynchronous gestures are patterns of sensed user or users' activity, spanning a distributed system that take on a new meaning when they occur together in time. Synchronous gestures draw inspiration from real-world social rituals such as toasting by tapping two drinking glasses together. In this article, we explore several interactions based on synchronous gestures, including bumping devices together, drawing corresponding pen gestures on touch-sensitive displays, simultaneously pressing a button on multiple smart-phones, or placing one or more devices on the sensing surface of a tabletop computer. These interactions focus on wireless composition of physically colocated devices, where users perceive one another and coordinate their actions through social protocol. We demonstrate how synchronous gestures may be phrased together with surrounding interactions. Such connection-action phrases afford a rich syntax of cross-device commands, operands, and one-to-one or one-to-many associations with a flexible physical arrangement of devices. Synchronous gestures enable colocated users to combine multiple devices into a heterogeneous display environment, where the users may establish a transient network connection with other select colocated users to facilitate the pooling of input capabilities, display resources, and the digital contents of each device. For example, participants at a meeting may bring mobile devices including tablet computers, PDAs, and smart-phones, and the meeting room infrastructure may include fixed interactive displays, such as a tabletop computer. Our techniques facilitate creation of an ad hoc display environment for tasks such as viewing a large document across multiple devices, presenting information to another user, or offering files to others. The interactions necessary to establish such ad hoc display environments must be rapid and minimally demanding of attention: during face-to-face communication, a pause of even 5 sec is socially awkward and disrupts collaboration. Current devices may associate using a direct transport such as Infrared Data Association ports, or the emerging Near Field Communication standard. However, such transports can only support one-to-one associations between devices and require close physical proximity as well as a specific relative orientation to connect the devices (e.g., the devices may be linked when touching head-to-head but not side-to-side). By contrast, sociology research in proxemics (the study of how people use the “personal space” surrounding their bodies) demonstrates that people carefully select physical distance as well as relative body orientation to suit the task, mood, and social relationship with other persons. Wireless networking can free device-to-device connections from the limitations of direct transports but results in a potentially large number of candidate devices. Synchronous gestures address these problems by allowing users to express naturally a spontaneous wireless connection between specific proximal (collocated) interactive displays. Gonzalo A. Ramos, Ken Hinckley, Andrew D. Wilson, Raman Sarin |
Hum. Comput. Interact. | 3 |
| 2008 | SurfaceFusion: unobtrusive tracking of everyday objects in tangible user interfaces
Alex Olwal, Andrew D. Wilson |
Graphics Interface | 2 |
| 2008 | Sphere: multi-touch interactions on a spherical displayabstractSphere is a multi-user, multi-touch-sensitive spherical display in which an infrared camera used for touch sensing shares the same optical path with the projector used for the display. This novel configuration permits: (1) the enclosure of both the projection and the sensing mechanism in the base of the device, and (2) easy 360-degree access for multiple users, with a high degree of interactivity without shadowing or occlusion. In addition to the hardware and software solution, we present a set of multi-touch interaction techniques and interface concepts that facilitate collaborative interactions around Sphere. We designed four spherical application concepts and report on several important observations of collaborative activity from our initial Sphere installation in three high-traffic locations. Hrvoje Benko, Andrew D. Wilson, Ravin Balakrishnan |
UIST | 2 |
| 2008 | Bringing physics to the surfaceabstractThis paper explores the intersection of emerging surface technologies, capable of sensing multiple contacts and of-ten shape information, and advanced games physics engines. We define a technique for modeling the data sensed from such surfaces as input within a physics simulation. This affords the user the ability to interact with digital objects in ways analogous to manipulation of real objects. Our technique is capable of modeling both multiple contact points and more sophisticated shape information, such as the entire hand or other physical objects, and of mapping this user input to contact forces due to friction and collisions within the physics simulation. This enables a variety of fine-grained and casual interactions, supporting finger-based, whole-hand, and tangible input. We demonstrate how our technique can be used to add real-world dynamics to interactive surfaces such as a vision-based tabletop, creating a fluid and natural experience. Our approach hides from application developers many of the complexities inherent in using physics engines, allowing the creation of applications without preprogrammed interaction behavior or gesture recognition. Andrew D. Wilson, Shahram Izadi, Otmar Hilliges, Armando Garcia-Mendoza, David S. Kirk |
UIST | 1 |
| 2007 | BlueTable: connecting wireless mobile devices on interactive surfaces using vision-based handshakingabstractAssociating and connecting mobile devices for the wireless transfer of data is often a cumbersome process. We present a technique of associating a mobile device to an interactive surface using a combination of computer vision and Bluetooth technologies. Users establish the connection of a mobile device to the system by simply placing the device on a table surface. When the computer vision process detects a phone-like object on the surface, the system follows a handshaking procedure using Bluetooth and vision techniques to establish that the phone on the surface and the wirelessly connected phone are the same device. The connection is broken simply by removing the device. Furthermore, the vision-based handshaking procedure determines the precise position of the device on the interactive surface, thus permitting a variety of interactive scenarios which rely on the presentation of graphics co-located with the device. As an example, we present a prototype interactive system which allows the exchange of automatically downloaded photos by selecting and dragging photos from one cameraphone device to another. Andrew D. Wilson, Raman Sarin |
Graphics Interface | 1 |
| 2007 | Gestures without libraries, toolkits or training: a $1 recognizer for user interface prototypesabstractAlthough mobile, tablet, large display, and tabletop computers increasingly present opportunities for using pen, finger, and wand gestures in user interfaces, implementing gesture recognition largely has been the privilege of pattern matching experts, not user interface prototypers. Although some user interface libraries and toolkits offer gesture recognizers, such infrastructure is often unavailable in design-oriented environments like Flash, scripting environments like JavaScript, or brand new off-desktop prototyping environments. To enable novice programmers to incorporate gestures into their UI prototypes, we present a "$1 recognizer" that is easy, cheap, and usable almost anywhere in about 100 lines of code. In a study comparing our $1 recognizer, Dynamic Time Warping, and the Rubine classifier on user-supplied gestures, we found that $1 obtains over 97% accuracy with only 1 loaded template and 99% accuracy with 3+ loaded templates. These results were nearly identical to DTW and superior to Rubine. In addition, we found that medium-speed gestures, in which users balanced speed and accuracy, were recognized better than slow or fast gestures for all three recognizers. We also discuss the effect that the number of templates or training examples has on recognition, the score falloff along recognizers' N-best lists, and results for individual gestures. We include detailed pseudocode of the $1 recognizer to aid development, inspection, extension, and testing. Jacob O. Wobbrock, Andrew D. Wilson |
UIST | 2 |
| 2006 | Precise selection techniques for multi-touch screensabstractThe size of human fingers and the lack of sensing precision can make precise touch screen interactions difficult. We present a set of five techniques, called Dual Finger Selections, which leverage the recent development of multi-touch sensitive displays to help users select very small targets. These techniques facilitate pixel-accurate targeting by adjusting the control-display ratio with a secondary finger while the primary finger controls the movement of the cursor. We also contribute a "clicking" technique, called SimPress, which reduces motion errors during clicking and allows us to simulate a hover state on devices unable to sense proximity. We implemented our techniques on a multi-touch tabletop prototype that offers computer vision-based tracking. In our formal user study, we tested the performance of our three most promising techniques (Stretch, X-Menu, and Slider) against our baseline (Offset), on four target sizes and three input noise levels. All three chosen techniques outperformed the control technique in terms of error rate reduction and were preferred by our participants, with Stretch being the overall performance and preference winner. Hrvoje Benko, Andrew D. Wilson, Patrick Baudisch |
CHI | 2 |
| 2006 | Text entry using a dual joystick game controllerabstractWe present a new bimanual text entry technique designed for today's dual-joystick game controllers. The left and right joysticks are used to independently select characters from the corresponding (left/right) half of an on-screen se-lection keyboard. Our dual-stick approach is analogous to typing on a standard keyboard, where each hand (left/right) presses keys on the corresponding side of the keyboard. We conducted a user study showing that our technique supports keyboarding skills transfer and is thereby readily learnable. Our technique increases entry speed significantly compared to the status quo single stick selection keyboard technique. Andrew D. Wilson, Maneesh Agrawala |
CHI | 1 |
| 2006 | Robust computer vision-based detection of pinching for one and two-handed gesture inputabstractWe present a computer vision technique to detect when the user brings their thumb and forefinger together (a pinch gesture) for close-range and relatively controlled viewing circumstances. The technique avoids complex and fragile hand tracking algorithms by detecting the hole formed when the thumb and forefinger are touching; this hole is found by simple analysis of the connected components of the background segmented against the hand. Our Thumb and Fore-Finger Interface (TAFFI) demonstrates the technique for cursor control as well as map navigation using one and two-handed interactions. Andrew D. Wilson |
UIST | 1 |
| 2005 | FlowMouse: A Computer Vision-Based Pointing and Gesture Input Device
Andrew D. Wilson, Edward Cutrell |
INTERACT | 1 |
| 2005 | PlayAnywhere: a compact interactive tabletop projection-vision systemabstractWe introduce PlayAnywhere, a front-projected computer vision-based interactive table system which uses a new commercially available projection technology to obtain a compact, self-contained form factor. PlayAnywhere's configuration addresses installation, calibration, and portability issues that are typical of most vision-based table systems, and thereby is particularly motivated in consumer applications. PlayAnywhere also makes a number of contributions related to image processing techniques for front-projected vision-based table systems, including a shadow-based touch detection algorithm, a fast, simple visual bar code scheme tailored to projection-vision table systems, the ability to continuously track sheets of paper, and an optical flow-based algorithm for the manipulation of onscreen objects that does not rely on fragile tracking algorithms. Andrew D. Wilson |
UIST | 1 |
| 2004 | Toward universal mobile interaction for shared displaysabstractResearchers have noted conflicting trends in collaboration technologies between delivering more information on larger displays and exploiting mobility on smaller devices. Large, shared displays provide greater choice in the presentation of information, but mobile devices offer greater flexibility in the access of information. We describe a platform that leverages the best of both worlds by allowing multiple users to access and interact with a large, shared display using their own personal mobile devices, such as a cell phone, laptop, or wireless PDA. We highlight three applications built on top of the platform that demonstrate its generality and utility in a variety of group settings: namely, web browsing, polling, and entertainment. Tim Paek, Maneesh Agrawala, Sumit Basu, Steven Mark Drucker, Trausti T. Kristjansson, Ron Logan, Kentaro Toyama, Andrew D. Wilson |
CSCW | 8 |
| 2004 | TouchLight: an imaging touch screen and display for gesture-based interactionabstractA novel touch screen technology is presented. TouchLight uses simple image processing techniques to combine the output of two video cameras placed behind a semi-transparent plane in front of the user. The resulting image shows objects that are on the plane. This technique is well suited for application with a commercially available projection screen material (DNP HoloScreen) which permits projection onto a transparent sheet of acrylic plastic in normal indoor lighting conditions. The resulting touch screen display system transforms an otherwise normal sheet of acrylic plastic into a high bandwidth input/output surface suitable for gesture-based interaction. Image processing techniques are detailed, and several novel capabilities of the system are outlined. Andrew D. Wilson |
ICMI | 1 |
| 2001 | Hidden Markov Models for Modeling and Recognizing Gesture Under VariationabstractConventional application of hidden Markov models to the task of recognizing human gesture may suffer from multiple sources of systematic variation in the sensor outputs. We present two frameworks based on hidden Markov models which are designed to model and recognize gestures that vary in systematic ways. In the first, the systematic variation is assumed to be communicative in nature, and the input gesture is assumed to belong to gesture family. The variation across the family is modeled explicitly by the parametric hidden Markov model (PHMM). In the second framework, variation in the signal is overcome by relying on online learning rather than conventional offline, batch learning. Andrew D. Wilson, Aaron F. Bobick |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2000 | Realtime Online Adaptive Gesture RecognitionabstractWe introduce an online adaptive algorithm for learning gesture models. By learning gesture models in an online fashion, the gesture recognition process is made more robust, and the need to train on a large training ensemble is obviated. Hidden Markov models are used to represent the spatial and temporal structure of the gesture. The usual output probability distributions-typically representing appearance-are trained at runtime exploiting the temporal structure (Markov model) that is either trained off-line or is explicitly hand-coded. In the early stages of runtime adaptation, contextural information derived from the application is used to bias the expectation as to which Markov state the system is in at any given time. We describe the Watch and Learn system, a computer vision system which is able to learn simple gestures online for interactive control. Andrew D. Wilson, Aaron F. Bobick |
ICPR | 1 |
| 1999 | Sympathetic Interfaces: Using a Plush Toy to Direct Synthetic CharactersabstractWe introduce the concept of a sympathetic inter$ace for controlling an animated synthetic character in a 3D virtual environment. A plush doll embedded with wireless sensors is used to manipulate the virtual character in an iconic and intentional manner. The interface extends from the novel physical input device through interpretation of sensor data to the behavioral brain of the virtual character. We discuss the design of the interface and focus on its latest instantiation in the Swamped! exhibit at SIGGRAPH 98. We also present what we learned from hundreds of casual users, who ranged from young children to adults. Michael Patrick Johnson, Andrew D. Wilson, Bruce Blumberg, Christopher Kline, Aaron F. Bobick |
CHI | 2 |
| 1999 | Parametric Hidden Markov Models for Gesture RecognitionabstractA method for the representation, recognition, and interpretation of parameterized gesture is presented. By parameterized gesture we mean gestures that exhibit a systematic spatial variation; one example is a point gesture where the relevant parameter is the two-dimensional direction. Our approach is to extend the standard hidden Markov model method of gesture recognition by including a global parametric variation in the output probabilities of the HMM states. Using a linear model of dependence, we formulate an expectation-maximization (EM) method for training the parametric HMM. During testing, a similar EM algorithm simultaneously maximizes the output likelihood of the PHMM for the given sequence and estimates the quantifying parameters. Using visually derived and directly measured three-dimensional hand position measurements as input, we present results that demonstrate the recognition superiority of the PHMM over standard HMM techniques, as well as greater robustness in parameter estimation with respect to noise in the input features. Finally, we extend the PHMM to handle arbitrary smooth (nonlinear) dependencies. The nonlinear formulation requires the use of a generalized expectation-maximization (GEM) algorithm for both training and the simultaneous recognition of the gesture and estimation of the value of the parameter. We present results on a pointing gesture, where the nonlinear approach permits the natural spherical coordinate parameterization of pointing direction. Andrew D. Wilson, Aaron F. Bobick |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1998 | Nonlinear PHMMs for the Interpretation of Parameterized GestureabstractRecently we modified the hidden Markov model (HMM) framework to incorporate a global parametric variation in the output probabilities of the states of the HMM. Development of the parametric hidden Markov model (PHMM) was motivated by the task of simultaneously recognizing and interpreting gestures that exhibit meaningful variation. With standard HMMs, such global variation confounds the recognition process. The original PHMM approach assumes a linear dependence of output density means on the global parameter. In this paper we extend the PHMM to handle arbitrary smooth (nonlinear) dependencies. We show a generalized expectation-maximization (GEM) algorithm for training the PHMM and a GEM algorithm to simultaneously recognize the gesture and estimate the value of the parameter. We present results on a pointing gesture, where the nonlinear approach permits the natural azimuth/elevation parameterization of pointing direction. Andrew D. Wilson, Aaron F. Bobick |
CVPR | 1 |
| 1998 | Recognition and Interpretation of Parametric GestureabstractA new method for the representation, recognition, and interpretation of parameterized gesture is presented. By parameterized gesture. We mean gestures that exhibit a meaningful variation; one example is a point gesture where the important parameter is the 2-dimensional direction. Our approach is to extend the standard hidden Markov model method of gesture recognition by including a global parametric variation in the output probabilities of the states of the HMM. Using a linear model to derive the theory, we formulated an expectation-maximization (EM) method for training the parametric HMM. During testing, the parametric HMM simultaneously recognizes the gesture and estimates the quantifying parameters. Using visually derived and directly measured 3-dimensional hand position measurements as input, we present results on two. Different movements-a size gesture and a point gesture-and show robustness with respect to noise in the input features. Andrew D. Wilson, Aaron F. Bobick |
ICCV | 1 |
| 1997 | Temporal Classification of Natural Gesture and Application to Video CodingabstractA method for the temporal classification of natural gesture from video imagery is presented. The work is motivated by recent developments in the theory of natural gesture which have identified several key temporal aspects of gesture important to communication. In particular gesticulation during conversation can be coarsely characterized as periods of bi-phasic or tri-phasic gesture separated by a rest state. We first present an automatic procedure for hypothesizing plausible rest state configurations of a speaker. Second, we develop a state-based parsing algorithm used to both select among candidate rest states and to parse an incoming video stream into bi-phasic and tri-phasic gestures. Finally, we demonstrate the use of the bi-phasic/tri-phasic labeling to select semantically significant static images for low bandwidth coding of video of story-telling speakers. Andrew D. Wilson, Aaron F. Bobick, Justine Cassell |
CVPR | 1 |
| 1997 | A State-Based Approach to the Representation and Recognition of GestureabstractA state-based technique for the representation and recognition of gesture is presented. We define a gesture to be a sequence of states in a measurement or configuration space. For a given gesture, these states are used to capture both the repeatability and variability evidenced in a training set of example trajectories. Using techniques for computing a prototype trajectory of an ensemble of trajectories, we develop methods for defining configuration states along the prototype and for recognizing gestures from an unsegmented, continuous stream of sensor data. The approach is illustrated by application to a range of gesture-related sensory data: the two-dimensional movements of a mouse input device, the movement of the hand measured by a magnetic spatial position and orientation sensor, and, lastly, the changing eigenvector projection coefficients computed from an image sequence. Aaron F. Bobick, Andrew D. Wilson |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1996 | Recovering the Temporal Structure of Natural GestureabstractA method for the recovery of the temporal structure and phases in natural gesture is presented. The work is motivated by recent developments in the theory of natural gesture which have identified several key aspects of gesture important to communication. In particular, gesticulation during conversation can be coarsely characterized as periods of bi-phasic or tri-phasic gesture separated by a rest state. We first present an automatic procedure for hypothesizing plausible rest state configurations of a speaker; the method uses the repetition of subsequences to indicate potential rest states. Second, we develop a state-based parsing algorithm used to both select among candidate rest stares and to parse an incoming video stream into bi-phasic and multi-phasic gestures. We present results from examples of story-telling speakers. Andrew D. Wilson, Aaron F. Bobick, Justine Cassell |
FG | 1 |
| 1995 | A State-Based Technique for the Summarization and Recognition of GestureabstractWe define a gesture to be a sequence of states in a measurement or configuration space. For a given gesture, these states are used to capture both the repeatability and variability evidenced in a training set of example trajectories. The states are positioned along a prototype of the gesture, and shaped such that they are narrow in the directions in which the ensemble of examples is tightly constrained, and wide in directions in which a great deal of variability is observed. We develop techniques for computing a prototype trajectory of an ensemble of trajectories, for defining configuration states along the prototype, and for recognizing gestures from an unsegmented, continuous stream of sensor data. The approach is illustrated by application to a range of gesture-related sensory data: the two-dimensional movements of a mouse input device, the movement of the hand measured by a magnetic spatial position and orientation sensor, and, lastly, the changing eigenvector projection coefficients computed from an image sequence.> Andrew D. Wilson, Aaron F. Bobick |
ICCV | 1 |