EDBT 2026 Demo / reviewers in the wild / expert
Anhong Guo
dblp:151/0002
· DBLP profile ↗
48ranked-venue papers
6as first author
30since 2021 · last 2026
0000-0002-4447-7818ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 45 · 6 first-author · 30 since 2021Artificial intelligence and machine learning · 4 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TouchScribe: Augmenting Non-Visual Hand-Object Interactions with Automated Live Visual DescriptionsabstractPeople who are blind or have low vision regularly use their hands to interact with the physical world to gain access to objects’ shape, size, weight, and texture. However, many rich visual features remain inaccessible through touch alone, making it difficult to distinguish similar objects, interpret visual affordances, and form a complete understanding of objects. In this work, we present TouchScribe, a system that augments hand-object interactions with automated live visual descriptions. We trained a custom egocentric hand interaction model to recognize both common gestures (e.g., grab to inspect, hold side-by-side to compare) and unique ones by blind people (e.g., point to explore color, or swipe to read available texts). Furthermore, TouchScribe provides real-time and adaptive feedback based on hand movement, from hand interaction states, to object labels, and to visual details. Our user study and technical evaluations demonstrate that TouchScribe can provide rich and useful descriptions to support object understanding. Finally, we discuss the implications of making live visual descriptions responsive to users’ physical reach. Ruei-Che Chang, Rosiana Natalie, Jovan Zheng Feng Yap, Tiange Luo, Venkatesh Potluri, Anhong Guo |
CHI | 7 |
| 2026 | Auditorily Embodied Conversational Agents: Effects of Spatialization and Situated Audio Cues on Presence and Social PerceptionabstractEmbodiment can enhance conversational agents, such as increasing their perceived presence. This is typically achieved through visual representations of a virtual body; however, visual modalities are not always available, such as when users interact with agents using headphones or display-less glasses. In this work, we explore auditory embodiment. By introducing auditory cues of bodily presence - through spatially localized voice and situated Foley audio from environmental interactions - we investigate how audio alone can convey embodiment and influence perceptions of a conversational agent. We conducted a 2 (spatialization: monaural vs. spatialized) x 2 (Foley: none vs. Foley) within-subjects study, where participants (n=24) engaged in conversations with agents. Our results show that spatialization and Foley increase co-presence, but reduce users' perceptions of the agent's attention and other social attributes. Yi Fei Cheng 0001, Jarod Bloch, Alexander Wang, Andrea Bianchi, Anusha Withana, Anhong Guo, Laurie M. Heller, David Lindlbauer |
CHI | 6 |
| 2026 | A11yExtensions: Accessibility Extensions to Augment Mobile AI Assistive Technology In-SituabstractExisting visual AI assistive technologies have usability gaps, and may need additional adaptations and features to serve users’ needs. We propose A11yExtensions, in-situ interventions that augment existing mobile AI assistive technology with add-on services. Add-ons include features that have been researched but are not yet deployed (e.g., cross-checking AI results), or that are only available in certain applications (e.g., camera aiming assistance). Through co-design sessions with two blind accessibility professionals, we designed and implemented three exemplar extensions, leveraging mobile automation tools to invoke add-ons, enabling just-in-time interventions for adaptability. We found that A11yExtensions provide opportunities to test new features and a new degree of flexibility and customization, though they introduce additional onboarding and communication challenges. We also derived a design space of accessibility extensions as a basis for future extension designs. Overall, A11yExtensions is a demonstration of the effectiveness of deploying new features in-situ via automation, with the technologies people actually use in their day-to-day lives. Jaylin Herskovitz, Margaret Ellen Seehorn, Ather Jammoa, Jason Meddaugh, Anhong Guo |
CHI | 5 |
| 2026 | A11y-CUA Dataset: Characterizing the Accessibility Gap in Computer Use AgentsabstractComputer Use Agents (CUAs) operate interfaces by pointing, clicking, and typing—mirroring interactions of sighted users (SUs) who can thus monitor CUAs and share control. CUAs do not reflect interactions by blind and low-vision users (BLVUs) who use assistive technology (AT). BLVUs thus cannot easily collaborate with CUAs. To characterize the accessibility gap of CUAs, we present A11y-CUA, a dataset of BLVUs and SUs performing 60 everyday tasks with 40.4 hours and 158,325 events. Our dataset analysis reveals that our collected interaction traces quantitatively confirm distinct interaction styles between SU and BLVU groups (mouse- vs. keyboard-dominant) and demonstrate interaction diversity within each group (sequential vs. shortcut navigation for BLVUs). We then compare collected traces to state-of-the-art CUAs under default and AT conditions (keyboard-only, magnifier). The default CUA executed 78.3% of tasks successfully. But with the AT conditions, CUA’s performance dropped to 41.67% and 28.3% with keyboard-only and magnifier conditions respectively, and did not reflect nuances of real AT use. With our open A11y-CUA dataset, we aim to promote collaborative and accessible CUAs for everyone. Ananya Gubbi Mohanbabu, Rosiana Natalie, Brandon Kim, Anhong Guo, Amy Pavel |
CHI | 4 |
| 2025 | Rubikon: Intelligent Tutoring for Rubik's Cube Learning Through AR-enabled Physical Task ReconfigurationabstractFigure 1: Rubikon is an intelligent tutoring system for Rubik's Cube learning.(a) The foundational design of Rubikon is an AR setup, where learners manipulate a physical cube with ArUco markers attached to each square, and pose a camera towards the cube to enable tracking and rendering.With this setup, learners see a rendered Rubik's Cube on a display while manipulating the physical cube in their hands.(b) Through AR rendering, Rubikon automatically generates new configurations of the Rubik's Cube for the user to practice unmastered skills.Rubikon detects the status of the cube to infer user behavior and provide immediate feedback and hints.(c) Rubikon supports the learning of a 3D physical task by integrating key design principles of cognitive tutors which have seen success in tutoring math and programming. Haocheng Ren, Muzhe Wu, Gregory Thomas Croisdale, Anhong Guo, Xu Wang 0016 |
Conference on Designing Interactive Systems | 4 |
| 2025 | Probing the Gaps in ChatGPT's Live Video Chat for Real-World Assistance for People who are Blind or Visually Impaired
Ruei-Che Chang, Rosiana Natalie, Jovan Zheng Feng Yap, Anhong Guo |
ASSETS | 5 |
| 2025 | How Well Can Vision Language Models Simulate the Vision Perception of People with Low Vision?abstractAdvances in Vision Language Models (VLMs) have enabled the simulation of general human behavior through their reasoning and problem solving capabilities.In the accessibility domain, such simulations may support the initial piloting of inclusive design processes, without replacing real human input, and can facilitate the personalization of AI-based application outcomes based on individual profiles.In this work, we conducted a preliminary examination of the extent to which VLMs can simulate the visual perception of people with low vision when interpreting images.We conducted a survey study with 40 low vision participants, collecting their brief and detailed vision information, and both open-ended and multiplechoice image perception and recognition responses to up to 25 images.Using these responses, we constructed prompts for VLMs to create simulated agents of each participant, varying the included information on vision information and example image responses.We evaluated the agreement between LLM-generated responses and participants' original answers.The agreement between the agent' and participants' responses remained low when only either the vision profile (0.59) or example image responses (0.59) were provided, whereas a combination of both significantly increase the agreement (0.70, p < 0.0001).Notably, a single example combining both open-ended and multiple-choice responses, offered significant performance improvements over either alone (p < 0.0001), while additional examples provided minimal benefits (p > 0.05). CCS Concepts• Human-centered computing → Accessibility. Rosiana Natalie, Ruei-Che Chang, Anhong Guo |
ASSETS | 4 |
| 2025 | A11yShape: AI-Assisted 3-D Modeling for Blind and Low-Vision ProgrammersabstractFigure 1: With A11yShape, (A) a blind or low-vision (BLV) user can create, interpret, and verify 3-D models through (B) a user interface composed of three parts: Code Editor Panel, AI Assistant Panel, and Model Panel.These panels are linked by a cross-representation highlighting mechanism that connects code, textual descriptions, hierarchical model abstractions, and 3-D visual renderings.The system supports the creation of (C) diverse, customized 3-D models created by BLV users. Zhuohao (Jerry) Zhang, Haichang Li, Chun Meng Yu, Faraz Faruqi, Junan Xie, Gene S.-H. Kim, Mingming Fan 0001, Angus G. Forbes, Jacob O. Wobbrock, Anhong Guo, Liang He 0005 |
ASSETS | 10 |
| 2025 | DeckFlow: Specification Decomposition on a Multimodal Generative Canvas
Gregory Thomas Croisdale, Emily Huang, John Joon Young Chung, Anhong Guo, Xu Wang 0016, Austin Z. Henley, Cyrus Omar |
VL/HCC | 4 |
| 2024 | SoundShift: Exploring Sound Manipulations for Accessible Mixed-Reality AwarenessabstractMixed-reality (MR) soundscapes blend real-world sound with virtual audio from hearing devices, presenting intricate auditory information that is hard to discern and differentiate. This is particularly challenging for blind or visually impaired individuals, who rely on sounds and descriptions in their everyday lives. To understand how complex audio information is consumed, we analyzed online forum posts within the blind community, identifying prevailing challenges, needs, and desired solutions. We synthesized the results and propose SoundShift for increasing MR sound awareness, which includes six sound manipulations: Transparency Shift, Envelope Shift, Position Shift, Style Shift, Time Shift, and Sound Append. To evaluate the effectiveness of SoundShift, we conducted a user study with 18 blind participants across three simulated MR scenarios, where participants identified specific sounds within intricate soundscapes. We found that SoundShift increased MR sound awareness and minimized cognitive load. Finally, we developed three real-world example applications to demonstrate the practicality of SoundShift. Ruei-Che Chang, Chia-Sheng Hung, Bing-Yu Chen 0004, Dhruv Jain, Anhong Guo |
Conference on Designing Interactive Systems | 5 |
| 2024 | EditScribe: Non-Visual Image Editing with Natural Language Verification LoopsabstractImage editing is an iterative process that requires precise visual evaluation and manipulation for the output to match the editing intent. However, current image editing tools do not provide accessible interaction nor sufficient feedback for blind and low vision individuals to achieve this level of control. To address this, we developed EditScribe, a prototype system that makes object-level image editing actions accessible using natural language verification loops powered by large multimodal models. Using EditScribe, the user first comprehends the image content through initial general and object descriptions, then specifies edit actions using open-ended natural language prompts. EditScribe performs the image edit, and provides four types of verification feedback for the user to verify the performed edit, including a summary of visual changes, AI judgement, and updated general and object descriptions. The user can ask follow-up questions to clarify and probe into the edits or verification feedback, before performing another edit. In a study with ten blind or low-vision users, we found that EditScribe supported participants to perform and verify image edit actions non-visually. We observed different prompting strategies from participants, and their perceptions on the various types of verification feedback. Finally, we discuss the implications of leveraging natural language verification loops to make visual authoring non-visually accessible. Ruei-Che Chang, Yuxuan Liu 0016, Lotus Hanzi Zhang, Anhong Guo |
ASSETS | 4 |
| 2024 | Audio Description CustomizationabstractBlind and low-vision (BLV) people use audio descriptions (ADs) to access videos. However, current ADs are unalterable by end users, thus are incapable of supporting BLV individuals’ potentially diverse needs and preferences. This research investigates if customizing AD could improve how BLV individuals consume videos. We conducted an interview study (Study 1) with fifteen BLV participants, which revealed desires for customizing properties like length, emphasis, speed, voice, format, tone, and language. At the same time, concerns like interruptions and increased interaction load due to customization emerged. To examine AD customization’s effectiveness and tradeoffs, we designed CustomAD, a prototype that enables BLV users to customize AD content and presentation. An evaluation study (Study 2) with twelve BLV participants showed using CustomAD significantly enhanced BLV people’s video understanding, immersion, and information navigation efficiency. Our work illustrates the importance of AD customization and offers a design that enhances video accessibility for BLV individuals. Rosiana Natalie, Ruei-Che Chang, Smitha Sheshadri, Anhong Guo, Kotaro Hara |
ASSETS | 4 |
| 2024 | InteractOut: Leveraging Interaction Proxies as Input Manipulation Strategies for Reducing Smartphone OveruseabstractSmartphone overuse poses risks to people’s physical and mental health. However, current intervention techniques mainly focus on explicitly changing screen content (i.e., output) and often fail to persistently reduce smartphone overuse due to being over-restrictive or over-flexible. We present the design and implementation of InteractOut, a suite of implicit input manipulation techniques that leverage interaction proxies to weakly inhibit the natural execution of common user gestures on mobile devices. We present a design space for input manipulations and demonstrate 8 Android implementations of input interventions. We first conducted a pilot lab study (N=30) to evaluate the usability of these interventions. Based on the results, we then performed a 5-week within-subject field experiment (N=42) to evaluate InteractOut in real-world scenarios. Compared to the traditional and common timed lockout technique, InteractOut significantly reduced the usage time by an additional 15.6% and opening frequency by 16.5% on participant-selected target apps. InteractOut also achieved a 25.3% higher user acceptance rate, and resulted in less frustration and better user experience according to participants’ subjective feedback. InteractOut demonstrates a new direction for smartphone overuse intervention and serves as a strong complementary set of techniques with existing methods. Tao Lu 0013, Hongxiao Zheng, Tianying Zhang, Xuhai Xu, Anhong Guo |
CHI | 5 |
| 2024 | WorldScribe: Towards Context-Aware Live Visual DescriptionsabstractAutomated live visual descriptions can aid blind people in understanding their surroundings with autonomy and independence. However, providing descriptions that are rich, contextual, and just-in-time has been a long-standing challenge in accessibility. In this work, we develop WorldScribe, a system that generates automated live real-world visual descriptions that are customizable and adaptive to users’ contexts: (i) WorldScribe’s descriptions are tailored to users’ intents and prioritized based on semantic relevance. (ii) WorldScribe is adaptive to visual contexts, e.g., providing consecutively succinct descriptions for dynamic scenes, while presenting longer and detailed ones for stable settings. (iii) WorldScribe is adaptive to sound contexts, e.g., increasing volume in noisy environments, or pausing when conversations start. Powered by a suite of vision, language, and sound recognition models, WorldScribe introduces a description generation pipeline that balances the tradeoffs between their richness and latency to support real-time use. The design of WorldScribe is informed by prior work on providing visual descriptions and a formative study with blind participants. Our user study and subsequent pipeline evaluation show that WorldScribe can provide real-time and fairly accurate visual descriptions to facilitate environment understanding that is adaptive and customized to users’ contexts. Finally, we discuss the implications and further steps toward making live visual descriptions more context-aware and humanized. Ruei-Che Chang, Yuxuan Liu 0016, Anhong Guo |
UIST | 3 |
| 2024 | ProgramAlly: Creating Custom Visual Access Programs via Multi-Modal End-User ProgrammingabstractExisting visual assistive technologies are built for simple and common use cases, and have few avenues for blind people to customize their functionalities. Drawing from prior work on DIY assistive technology, this paper investigates end-user programming as a means for users to create and customize visual access programs to meet their unique needs. We introduce ProgramAlly, a system for creating custom filters for visual information, e.g., ‘find NUMBER on BUS’, leveraging three end-user programming approaches: block programming, natural language, and programming by example. To implement ProgramAlly, we designed a representation of visual filtering tasks based on scenarios encountered by blind people, and integrated a set of on-device and cloud models for generating and running these programs. In user studies with 12 blind adults, we found that participants preferred different programming modalities depending on the task, and envisioned using visual access programs to address unique accessibility challenges that are otherwise difficult with existing applications. Through ProgramAlly, we present an exploration of how blind end-users can create visual access programs to customize and control their experiences. Jaylin Herskovitz, Andi Xu, Rahaf Alharbi, Anhong Guo |
UIST | 4 |
| 2024 | VRCopilot: Authoring 3D Layouts with Generative AI Models in VRabstractImmersive authoring provides an intuitive medium for users to create 3D scenes via direct manipulation in Virtual Reality (VR). Recent advances in generative AI have enabled the automatic creation of realistic 3D layouts. However, it is unclear how capabilities of generative AI can be used in immersive authoring to support fluid interactions, user agency, and creativity. We introduce VRCopilot, a mixed-initiative system that integrates pre-trained generative AI models into immersive authoring to facilitate human-AI co-creation in VR. VRCopilot presents multimodal interactions to support rapid prototyping and iterations with AI, and intermediate representations such as wireframes to augment user controllability over the created content. Through a series of user studies, we evaluated the potential and challenges in manual, scaffolded, and automatic creation in immersive authoring. We found that scaffolded creation using wireframes enhanced the user agency compared to automatic creation. We also found that manual creation via multimodal specification offers the highest sense of creativity and agency. Lei Zhang 0216, Jacob Gettig, Steve Oney, Anhong Guo |
UIST | 5 |
| 2023 | Deploying VizLens: Characterizing User Needs, Preferences, and Challenges of Physical Interfaces Usage in the WildabstractBlind or Visually Impaired (BVI) people often encounter flat, inaccessible interfaces. Current solutions lack cost-effectiveness, portability, and robustness in real-world settings. We introduce VizLens, a fully-automated, full-stack mobile application powered by computer vision algorithms. The system is deployed and publicly available through the Apple App Store (https://vizlens.org/). From May to August 2023, we had 665 users, who uploaded 1,320 interface images. We aim to use it to study usage patterns and possible challenges BVI users may encounter with flat interfaces through a large-scale study in real-world settings. With in-depth analysis of user data and activity logs, our study will provide insights into BVI users’ interface interests, preferred assistance modes, and potential challenges due to system limitations or users’ diverse abilities. Our goal is to enhance the understanding of how BVI users interact with inaccessible, flat interfaces, and inform future assistive technology design. Andi Xu, Mahdi Qazwini, Anhong Guo |
ASSETS | 4 |
| 2023 | Hacking, Switching, Combining: Understanding and Supporting DIY Assistive Technology Design by Blind PeopleabstractExisting assistive technologies (AT) often fail to support the unique needs of blind and visually impaired (BVI) people. Thus, BVI people have become domain experts in customizing and ‘hacking’ AT, creatively suiting their needs. We aim to understand this behavior in depth, and how BVI people envision creating future DIY personalized AT. We conducted a multi-part qualitative study with 12 blind participants: an interview on unique uses of AT, a two-week diary study to log use cases, and a scenario-based design session to imagine creating future technologies. We found that participants work to design new AT both implicitly through creative use cases, and explicitly through regular ideation and development. Participants envisioned creating a variety of new technologies, and we summarize expected benefits and concerns of using a DIY technology approach. From our results, we present design considerations for future DIY technology systems to support existing customization and ‘hacking’ behaviors. Jaylin Herskovitz, Andi Xu, Rahaf Alharbi, Anhong Guo |
CHI | 4 |
| 2023 | VRGit: A Version Control System for Collaborative Content Creation in Virtual RealityabstractImmersive authoring tools allow users to intuitively create and manipulate 3D scenes while immersed in Virtual Reality (VR). Collaboratively designing these scenes is a creative process that involves numerous edits, explorations of design alternatives, and frequent communication with collaborators. Version Control Systems (VCSs) help users achieve this by keeping track of the version history and creating a shared hub for communication. However, most VCSs are unsuitable for managing the version history of VR content because their underlying line differencing mechanism is designed for text and lacks the semantic information of 3D content; and the widely adopted commit model is designed for asynchronous collaboration rather than real-time awareness and communication in VR. We introduce VRGit, a new collaborative VCS that visualizes version history as a directed graph composed of 3D miniatures, and enables users to easily navigate versions, create branches, as well as preview and reuse versions directly in VR. Beyond individual uses, VRGit also facilitates synchronous collaboration in VR by providing awareness of users’ activities and version history through portals and shared history visualizations. In a lab study with 14 participants (seven groups), we demonstrate that VRGit enables users to easily manage version history both individually and collaboratively in VR. Lei Zhang 0216, Ashutosh Agrawal, Steve Oney, Anhong Guo |
CHI | 4 |
| 2023 | Human-Centered Deferred Inference: Measuring User Interactions and Setting Deferral Criteria for Human-AI TeamsabstractAlthough deep learning holds the promise of novel and impactful interfaces, realizing such promise in practice remains a challenge: since dataset-driven deep-learned models assume a one-time human input, there is no recourse when they do not understand the input provided by the user. Works that address this via deferred inference—soliciting additional human input when uncertain—show meaningful improvement, but ignore key aspects of how users and models interact. In this work, we focus on the role of users in deferred inference and argue that the deferral criteria should be a function of the user and model as a team, not simply the model itself. In support of this, we introduce a novel mathematical formulation, validate it via an experiment analyzing the interactions of 25 individuals with a deep learning-based visiolinguistic model, and identify user-specific dependencies that are under-explored in prior work. We conclude by demonstrating two human-centered procedures for setting deferral criteria that are simple to implement, applicable to a wide variety of tasks, and perform equal to or better than equivalent procedures that use much larger datasets. Stephan J. Lemmer, Anhong Guo, Jason J. Corso |
IUI | 2 |
| 2023 | BrushLens: Hardware Interaction Proxies for Accessible Touchscreen Interface ActuationabstractTouchscreen devices, designed with an assumed range of user abilities and interaction patterns, often present challenges for individuals with diverse abilities to operate independently. Prior efforts to improve accessibility through tools or algorithms necessitated alterations to touchscreen hardware or software, making them inapplicable for the large number of existing legacy devices. In this paper, we introduce BrushLens, a hardware interaction proxy that performs physical interactions on behalf of users while allowing them to continue utilizing accessible interfaces, such as screenreaders and assistive touch on smartphones, for interface exploration and command input. BrushLens maintains an interface model for accurate target localization and utilizes exchangeable actuators for physical actuation across a variety of device types, effectively reducing user workload and minimizing the risk of mistouch. Our evaluations reveal that BrushLens lowers the mistouch rate and empowers visually and motor impaired users to interact with otherwise inaccessible physical touchscreens more effectively. Yasha Iravantchi, Thomas Krolikowski, Ruijie Geng, Alanson P. Sample, Anhong Guo |
UIST | 6 |
| 2022 | CustomizAR: Facilitating Interactive Exploration and Measurement of Adaptive 3D DesignsabstractOnline 3D model repositories such as Thingiverse offer millions of open source designs that are shared for reuse and remix. Many of the designs are customizable to adapt to real-world objects upon personal needs of varying tasks and physical dimensions. However, it is challenging for novices to discover such designs using text-based search queries, comprehend what each parameter means for customization, locate these parameters on the target objects for measurement, and conduct measurements correctly. These challenges may cause the designs to be incorrectly adjusted, thus failing to function as expected and requiring users to start over, which costs additional time and material. We present CustomizAR, a pipeline for facilitating the interactive exploration of adaptive designs and the measurement of real-world constraints to fabricate them correctly. CustomizAR supports the search and discovery of adaptive 3D designs using an object-centric graph-based data structure, and guides users through an interactive measurement process leveraging computer vision techniques. Our technical evaluations and user studies demonstrate that CustomizAR facilitates effective discovery, adjustment, and reuse of adaptive designs that are shared online. Anhong Guo, Jeeeun Kim |
Conference on Designing Interactive Systems | 2 |
| 2022 | ImageExplorer: Multi-Layered Touch Exploration to Encourage Skepticism Towards Imperfect AI-Generated Image CaptionsabstractBlind users rely on alternative text (alt-text) to understand an image; however, alt-text is often missing. AI-generated captions are a more scalable alternative, but they often miss crucial details or are completely incorrect, which users may still falsely trust. In this work, we sought to determine how additional information could help users better judge the correctness of AI-generated captions. We developed ImageExplorer, a touch-based multi-layered image exploration system that allows users to explore the spatial layout and information hierarchies of images, and compared it with popular text-based (Facebook) and touch-based (Seeing AI) image exploration systems in a study with 12 blind participants. We found that exploration was generally successful in encouraging skepticism towards imperfect captions. Moreover, many participants preferred ImageExplorer for its multi-layered and spatial information presentation, and Facebook for its summary and ease of use. Finally, we identify design improvements for effective and explainable image exploration systems for blind users. Jaewook Lee 0005, Jaylin Herskovitz, Yi-Hao Peng, Anhong Guo |
CHI | 4 |
| 2022 | CollabAlly: Accessible Collaboration Awareness in Document EditingabstractCollaborative document editing tools are widely used in professional and academic workplaces. While these tools provide basic accessibility support, it is challenging for blind users to gain collaboration awareness that sighted people can easily obtain using visual cues (e.g., who is editing where and what). Through a series of co-design sessions with a blind coauthor, we identified the current practices and challenges in collaborative editing, and iteratively designed CollabAlly, a system that makes collaboration awareness in document editing accessible to blind users. CollabAlly extracts collaborator, comment, and text-change information and their context from a document and presents them in a dialog box to provide easy access and navigation. CollabAlly uses earcons to communicate background events unobtrusively, voice fonts to differentiate collaborators, and spatial audio to convey the location of document activity. In a study with 11 blind participants, we demonstrate that CollabAlly provides improved access to collaboration awareness by centralizing scattered information, sonifying visual information, and simplifying complex operations. Cheuk Yin Phipson Lee, Zhuohao (Jerry) Zhang, Jaylin Herskovitz, Jooyoung Seo, Anhong Guo |
CHI | 5 |
| 2022 | OmniScribe: Authoring Immersive Audio Descriptions for 360° VideosabstractBlind people typically access videos via audio descriptions (AD) crafted by sighted describers who comprehend, select, and describe crucial visual content in the videos. 360° video is an emerging storytelling medium that enables immersive experiences that people may not possibly reach in everyday life. However, the omnidirectional nature of 360° videos makes it challenging for describers to perceive the holistic visual content and interpret spatial information that is essential to create immersive ADs for blind people. Through a formative study with a professional describer, we identified key challenges in describing 360° videos and iteratively designed OmniScribe, a system that supports the authoring of immersive ADs for 360° videos. OmniScribe uses AI-generated content-awareness overlays for describers to better grasp 360° video content. Furthermore, OmniScribe enables describers to author spatial AD and immersive labels for blind users to consume the videos immersively with our mobile prototype. In a study with 11 professional and novice describers, we demonstrated the value of OmniScribe in the authoring workflow; and a study with 8 blind participants revealed the promise of immersive AD over standard AD for 360° videos. Finally, we discuss the implications of promoting 360° video accessibility. Ruei-Che Chang, Chao-Hsien Ting, Chia-Sheng Hung, Wan-Chen Lee, Liang-Jin Chen, Yu-Tzu Chao, Bing-Yu Chen 0004, Anhong Guo |
UIST | 8 |
| 2022 | XSpace: An Augmented Reality Toolkit for Enabling Spatially-Aware Distributed CollaborationabstractAugmented Reality (AR) has the potential to leverage environmental information to better facilitate distributed collaboration, however, such applications are difficult to develop. We present XSpace, a toolkit for creating spatially-aware AR applications for distributed collaboration. Based on a review of existing applications and developer tools, we design XSpace to support three methods for creating shared virtual spaces, each emphasizing a different aspect: shared objects, user perspectives, and environmental meshes. XSpace implements these methods in a developer toolkit, and also provides a set of complimentary visual authoring tools to allow developers to preview a variety of configurations for a shared virtual space. We present five example applications to illustrate that XSpace can support the development of a rich set of collaborative AR experiences that are difficult to produce with current solutions. Through XSpace, we discuss implications for future application design, including user space customization and privacy and safety concerns when sharing users' environments. Jaylin Herskovitz, Yi Fei Cheng 0001, Anhong Guo, Alanson P. Sample, Michael Nebeling |
Proc. ACM Hum. Comput. Interact. | 3 |
| 2021 | Designing Disaggregated Evaluations of AI Systems: Choices, Considerations, and TradeoffsabstractDisaggregated evaluations of AI systems, in which system performance is assessed and reported separately for different groups of people, are conceptually simple. However, their design involves a variety of choices. Some of these choices influence the results that will be obtained, and thus the conclusions that can be drawn; others influence the impacts---both beneficial and harmful---that a disaggregated evaluation will have on people, including the people whose data is used to conduct the evaluation. We argue that a deeper understanding of these choices will enable researchers and practitioners to design careful and conclusive disaggregated evaluations. We also argue that better documentation of these choices, along with the underlying considerations and tradeoffs that have been made, will help others when interpreting an evaluation's results and conclusions. Solon Barocas, Anhong Guo, Ece Kamar, Jacquelyn Krones, Meredith Ringel Morris, Jennifer Wortman Vaughan, W. Duncan Wadsworth, Hanna M. Wallach |
AIES | 2 |
| 2021 | Image Explorer: Multi-Layered Touch Exploration to Make Images AccessibleabstractBlind or visually impaired (BVI) individuals often rely on alternative text (alt-text) in order to understand an image; however, alt-text is often missing or incomplete. Automatically-generated captions are a more scalable alternative, but they are also often missing crucial details, and, sometimes, are completely incorrect, which may still be falsely trusted by BVI users. We hypothesize that additional information could help BVI users better judge the correctness of an auto-generated caption. To achieve this, we present Image Explorer, a touch-based multi-layered image exploration system that enables users to explore the spatial layout and information hierarchies in an image. Image Explorer leverages several off-the-shelf deep learning models to generate segmentation and labeling results for an image, combines and filters the generated information, and presents the resulted information in hierarchical layers. In a pilot study with three BVI users, participants used Image Explorer, Seeing AI, and Facebook to explore images with auto-generated captions of diverging quality, and judge the correctness of the captions. Preliminary results show that participants made more accurate judgements about the correctness of the captions when using Image Explorer, although they were highly confident about their judgement regardless of the tool used. Overall, Image Explorer is a novel touch exploration system that makes images more accessible for BVI users by potentially encouraging skepticism and enabling users to independently validate auto-generated captions. Jaewook Lee 0005, Yi-Hao Peng, Jaylin Herskovitz, Anhong Guo |
ASSETS | 4 |
| 2021 | CollabAlly: Accessible Collaboration Awareness in Document EditingabstractCollaborative document editing tools are widely used in both professional and academic workplaces. While these tools provide some accessibility features, it is still challenging for blind users to gain collaboration awareness that sighted people can easily obtain using visual cues (e.g., who edited or commented where and what in the document). To address this gap, we present CollabAlly, a browser extension that makes extractable collaborative and contextual information in document editing accessible for blind users. With CollabAlly, blind users can easily access collaborators’ information, track real-time or asynchronous content and comment changes, and navigate through these elements. In order to convey this complex information through audio, CollabAlly uses voice fonts and spatial audio to enhance users’ collaboration awareness in shared documents. Through a series of pilot studies with a coauthor who is blind, CollabAlly’s design was refined to include more information and to be more compatible with existing screen readers. Cheuk Yin Phipson Lee, Zhuohao (Jerry) Zhang, Jaylin Herskovitz, Jooyoung Seo, Anhong Guo |
ASSETS | 5 |
| 2021 | "It's Complicated": Negotiating Accessibility and (Mis)Representation in Image Descriptions of Race, Gender, and DisabilityabstractContent creators are instructed to write textual descriptions of visual content to make it accessible; yet existing guidelines lack specifics on how to write about people’s appearance, particularly while remaining mindful of consequences of (mis)representation. In this paper, we report on interviews with screen reader users who were also Black, Indigenous, People of Color, Non-binary, and/or Transgender on their current image description practices and preferences, and experiences negotiating theirs and others’ appearances non-visually. We discuss these perspectives, and the ethics of humans and AI describing appearance characteristics that may convey the race, gender, and disabilities of those photographed. In turn, we share considerations for more carefully describing appearance, and contexts in which such information is perceived salient. Finally, we offer tensions and questions for accessibility research to equitably consider politics and ecosystems in which technologies will embed, such as potential risks of human and AI biases amplifying through image descriptions. Cynthia L. Bennett, Cole Gleason, Morgan Klaus Scheuerman, Jeffrey P. Bigham, Anhong Guo, Alexandra To |
CHI | 5 |
| 2020 | Disability and the COVID-19 Pandemic: Using Twitter to Understand Accessibility during Rapid Societal TransitionabstractThe COVID-19 pandemic has forced institutions to rapidly alter their behavior, which typically has disproportionate negative effects on people with disabilities as accessibility is overlooked. To investigate these issues, we analyzed Twitter data to examine accessibility problems surfaced by the crisis. We identified three key domains at the intersection of accessibility and technology: (i) the allocation of product delivery services, (ii) the transition to remote education, and (iii) the dissemination of public health information. We found that essential retailers expanded their high-risk customer shopping hours and pick-up and delivery services, but individuals with disabilities still lacked necessary access to goods and services. Long-experienced access barriers to online education were exacerbated by the abrupt transition of in-person to remote instruction. Finally, public health messaging has been inconsistent and inaccessible, which is unacceptable during a rapidly-evolving crisis. We argue that organizations should create flexible, accessible technology and policies in calm times to be adaptable in times of crisis to serve individuals with diverse needs. Cole Gleason, Stephanie Valencia, Lynn Kirabo, Jason Wu 0001, Anhong Guo, Elizabeth J. Carter, Jeffrey P. Bigham, Cynthia L. Bennett, Amy Pavel |
ASSETS | 5 |
| 2020 | Making Mobile Augmented Reality Applications AccessibleabstractAugmented Reality (AR) technology creates new immersive experiences in entertainment, games, education, retail, and social media. AR content is often primarily visual and it is challenging to enable access to it non-visually due to the mix of virtual and real-world content. In this paper, we identify common constituent tasks in AR by analyzing existing mobile AR applications for iOS, and characterize the design space of tasks that require accessible alternatives. For each of the major task categories, we create prototype accessible alternatives that we evaluate in a study with 10 blind participants to explore their perceptions of accessible AR. Our study demonstrates that these prototypes make AR possible to use for blind users and reveals a number of insights to move forward. We believe our work sets forth not only exemplars for developers to create accessible AR applications, but also a roadmap for future research to make AR comprehensively accessible. Jaylin Herskovitz, Jason Wu 0001, Samuel White, Amy Pavel, Gabriel Reyes, Anhong Guo, Jeffrey P. Bigham |
ASSETS | 6 |
| 2020 | Sense and Accessibility: Understanding People with Physical Disabilities' Experiences with Sensing SystemsabstractSensing technologies that implicitly and explicitly mediate digital experiences are an increasingly pervasive part of daily living; it is vital to ensure that these technologies work appropriately for people with physical disabilities. We conducted on online survey with 40 adults with physical disabilities, gathering open-ended descriptions about respondents’ experiences with a variety of sensing systems, including motion sensors, biometric sensors, speech input, as well as touch and gesture systems. We present findings regarding the many challenges status quo sensing systems present for people with physical disabilities, as well as the ways in which our participants responded to these challenges. We conclude by reflecting on the significance of these findings for defining a future research agenda for creating more inclusive sensing systems. Shaun K. Kane, Anhong Guo, Meredith Ringel Morris |
ASSETS | 2 |
| 2019 | Supporting Older Adults in Using Complex User Interfaces with Augmented RealityabstractUsing complex interfaces has been shown to be challenging for older adults. Existing tutorial systems can be cumbersome, and sometimes difficult to use. To solve this problem, we present a system to support older adults in using visual interfaces by providing step-by-step visual guidance with augmented reality. Using the Apple ARKit platform, our system detects the interface in a phone camera view, and provides visual guidance for users to access the interface following a generated sequence of interactions based on pre-specified tasks and prior knowledge of the interface. Junhan Kong, Anhong Guo, Jeffrey P. Bigham |
ASSETS | 2 |
| 2019 | X-Ray: Screenshot Accessibility via Embedded MetadataabstractScreenshots are frequently shared on social media, via personal communications, and in academic papers. Unfortunately, existing screenshot tools strip away semantics useful for making the content accessible, leaving only pixels. For example, a screenshot of a table removes the structural information useful for conveying it. We quantify the scale of the problem via a study of academic papers, showing that a large number of images included in academic papers are screenshots, and validate this via qualitative interviews with researchers about their figure generation process. We then introduce X-Ray, a system that captures and embeds the semantics of the underlying content into images. Using the X-Ray screenshot tool, semantic information is captured and stored in the Exif data of the resulting image, allowing it to "tag along" as the image is shared and reposted. We demonstrate that our approach retains accessibility for screen reader users via a study with five blind participants. More generally, our approach suggests a method for embedding accessibility metadata into otherwise inaccessible formats, enabling them to retain the more accessible representations that are present at capture time. Sujeath Pareddy, Anhong Guo, Jeffrey P. Bigham |
ASSETS | 2 |
| 2019 | VizWiz-Priv: A Dataset for Recognizing the Presence and Purpose of Private Visual Information in Images Taken by Blind PeopleabstractWe introduce the first visual privacy dataset originating from people who are blind in order to better understand their privacy disclosures and to encourage the development of algorithms that can assist in preventing their unintended disclosures. It includes 8,862 regions showing private content across 5,537 images taken by blind people. Of these, 1,403 are paired with questions and 62\% of those directly ask about the private content. Experiments demonstrate the utility of this data for predicting whether an image shows private information and whether a question asks about the private content in an image. The dataset is publicly-shared at http://vizwiz.org/data/. Danna Gurari, Qing Li 0003, Chi Lin 0001, Anhong Guo, Abigale Stangl, Jeffrey P. Bigham |
CVPR | 5 |
| 2019 | StateLens: A Reverse Engineering Solution for Making Existing Dynamic Touchscreens AccessibleabstractBlind people frequently encounter inaccessible dynamic touchscreens in their everyday lives that are difficult, frustrating, and often impossible to use independently. Touchscreens are often the only way to control everything from coffee machines and payment terminals, to subway ticket machines and in-flight entertainment systems. Interacting with dynamic touchscreens is difficult non-visually because the visual user interfaces change, interactions often occur over multiple different screens, and it is easy to accidentally trigger interface actions while exploring the screen. To solve these problems, we introduce StateLens - a three-part reverse engineering solution that makes existing dynamic touchscreens accessible. First, StateLens reverse engineers the underlying state diagrams of existing interfaces using point-of-view videos found online or taken by users using a hybrid crowd-computer vision pipeline. Second, using the state diagrams, StateLens automatically generates conversational agents to guide blind users through specifying the tasks that the interface can perform, allowing the StateLens iOS application to provide interactive guidance and feedback so that blind users can access the interface. Finally, a set of 3D-printed accessories enable blind people to explore capacitive touchscreens without the risk of triggering accidental touches on the interface. Our technical evaluation shows that StateLens can accurately reconstruct interfaces from stationary, hand-held, and web videos; and, a user study of the complete system demonstrates that StateLens successfully enables blind users to access otherwise inaccessible dynamic touchscreens. Anhong Guo, Junhan Kong, Michael L. Rivera, Frank F. Xu, Jeffrey P. Bigham |
UIST | 1 |
| 2018 | Investigating Cursor-based Interactions to Support Non-Visual Exploration in the Real WorldabstractThe human visual system processes complex scenes to focus attention on relevant items. However, blind people cannot visually skim for an area of interest. Instead, they use a combination of contextual information, knowledge of the spatial layout of their environment, and interactive scanning to find and attend to specific items. In this paper, we define and compare three cursor-based interactions to help blind people attend to items in a complex visual scene: window cursor (move their phone to scan), finger cursor (point their finger to read), and touch cursor (drag their finger on the touchscreen to explore). We conducted a user study with 12 participants to evaluate the three techniques on four tasks, and found that: window cursor worked well for locating objects on large surfaces, finger cursor worked well for accessing control panels, and touch cursor worked well for helping users understand spatial layouts. A combination of multiple techniques will likely be best for supporting a variety of everyday tasks for blind users. Anhong Guo, Saige McVea, Xu Wang 0016, Patrick Clary, Kenneth J. Goldman, Yang Li 0058, Jeffrey P. Bigham |
ASSETS | 1 |
| 2018 | VizWiz Grand Challenge: Answering Visual Questions From Blind PeopleabstractThe study of algorithms to automatically answer visual questions currently is motivated by visual question answering (VQA) datasets constructed in artificial VQA settings. We propose VizWiz, the first goal-oriented VQA dataset arising from a natural VQA setting. VizWiz consists of over 31,000 visual questions originating from blind people who each took a picture using a mobile phone and recorded a spoken question about it, together with 10 crowdsourced answers per visual question. VizWiz differs from the many existing VQA datasets because (1) images are captured by blind photographers and so are often poor quality, (2) questions are spoken and so are more conversational, and (3) often visual questions cannot be answered. Evaluation of modern algorithms for answering visual questions and deciding if a visual question is answerable reveals that VizWiz is a challenging dataset. We introduce this dataset to encourage a larger community to develop more generalized algorithms that can assist blind people. Danna Gurari, Qing Li 0003, Abigale Stangl, Anhong Guo, Chi Lin 0001, Kristen Grauman, Jiebo Luo 0001, Jeffrey P. Bigham |
CVPR | 4 |
| 2017 | Understanding Uncertainty in Measurement and Accommodating its Impact in 3D Modeling and PrintingabstractThe growing accessibility of 3D printing to everyday users has led to the rapid adoption, sharing of 3D models on sites such as Thingiverse.com, and visions of a future in which customization is a norm and 3D printing can solve a variety of real-world problems. However, in practice, creating models is difficult and many end users simply print models created by others. In this paper, we explore a specific area of model design that is a challenge for end users' measurement. When a model must conform to a specific real world goal once printed, it is important that that goal is precisely specified. We demonstrate that measurement errors are a significant (yet often overlooked) challenge for end users through a systematic study of the sources and types of measurement errors. We argue for a new design principle--accommodating measurement error--that designers, as well as novice modelers, should to use at design time. We offer two strategies--buffer insertion and replacement of minimal parts--to help designers, as well as novice modelers, to build models that are robust to measurement error. We argue that these strategies can reduce the need for and costs of iteration and demonstrate their use in a series of printed objects. Jeeeun Kim, Anhong Guo, Tom Yeh, Scott E. Hudson, Jennifer Mankoff |
Conference on Designing Interactive Systems | 2 |
| 2017 | Facade: Auto-generating Tactile Interfaces to AppliancesabstractCommon appliances have shifted toward flat interface panels, making them inaccessible to blind people. Although blind people can label appliances with Braille stickers, doing so generally requires sighted assistance to identify the original functions and apply the labels. We introduce Facade - a crowdsourced fabrication pipeline to help blind people independently make physical interfaces accessible by adding a 3D printed augmentation of tactile buttons overlaying the original panel. Facade users capture a photo of the appliance with a readily available fiducial marker (a dollar bill) for recovering size information. This image is sent to multiple crowd workers, who work in parallel to quickly label and describe elements of the interface. Facade then generates a 3D model for a layer of tactile and pressable buttons that fits over the original controls. Finally, a home 3D printer or commercial service fabricates the layer, which is then aligned and attached to the interface by the blind person. We demonstrate the viability of Facade in a study with 11 blind participants. Anhong Guo, Jeeeun Kim, Xiang 'Anthony' Chen, Tom Yeh, Scott E. Hudson, Jennifer Mankoff, Jeffrey P. Bigham |
CHI | 1 |
| 2016 | VizMap: Accessible Visual Information Through Crowdsourced Map ReconstructionabstractWhen navigating indoors, blind people are often unaware of key visual information, such as posters, signs, and exit doors. Our VizMap system uses computer vision and crowdsourcing to collect this information and make it available non-visually. VizMap starts with videos taken by on-site sighted volunteers and uses these to create a 3D spatial model. These video frames are semantically labeled by remote crowd workers with key visual information. These semantic labels are located within and embedded into the reconstructed 3D model, forming a query-able spatial representation of the environment. VizMap can then localize the user with a photo from their smartphone, and enable them to explore the visual elements that are nearby. We explore a range of example applications enabled by our reconstructed spatial representation. With VizMap, we move towards integrating the strengths of the end user, on-site crowd, online crowd, and computer vision to solve a long-standing challenge in indoor blind exploration. Cole Gleason, Anhong Guo, Gierad Laput, Kris Makoto Kitani, Jeffrey P. Bigham |
ASSETS | 2 |
| 2016 | Facade: Auto-generating Tactile Interfaces to AppliancesabstractDigital keypads have proliferated on common appliances, from microwaves and refrigerators to printers and remote controls. For blind people, such interfaces are inaccessible. We conducted a formative study with 6 blind people which demonstrated a need for custom designs for tactile labels without dependence on sighted assistance. To address this need, we introduce Facade - a crowdsourced fabrication pipeline to make physical interfaces accessible by adding a 3D printed augmentation of tactile buttons overlaying the original panel. Blind users capture a photo of an inaccessible interface with a standard marker for absolute measurements using perspective transformation. Then this image is sent to multiple crowd workers, who work in parallel to quickly label and describe elements of the interface. These labels are then used to generate 3D models for a layer of tactile and pressable buttons that fits over the original controls. Users can customize the shape and labels of the buttons using a web interface. Finally, a consumer-grade 3D printer fabricates the layer, which is then attached to the interface using adhesives. Such fabricated overlay is an inexpensive ($10) and more general solution to making physical interfaces accessible. Anhong Guo, Jeeeun Kim, Xiang 'Anthony' Chen, Tom Yeh, Scott E. Hudson, Jennifer Mankoff, Jeffrey P. Bigham |
ASSETS | 1 |
| 2016 | WearWrite: Crowd-Assisted Writing from SmartwatchesabstractThe physical constraints of smartwatches limit the range and complexity of tasks that can be completed. Despite interface improvements on smartwatches, the promise of enabling productive work remains largely unrealized. This paper presents WearWrite, a system that enables users to write documents from their smartwatches by leveraging a crowd to help translate their ideas into text. WearWrite users dictate tasks, respond to questions, and receive notifications of major edits on their watch. Using a dynamic task queue, the crowd receives tasks issued by the watch user and generic tasks from the system. In a week-long study with seven smartwatch users supported by approximately 29 crowd workers each, we validate that it is possible to manage the crowd writing process from a watch. Watch users captured new ideas as they came to mind and managed a crowd during spare moments while going about their daily routine. WearWrite represents a new approach to getting work done from wearables using the crowd. Michael Nebeling, Alexandra To, Anhong Guo, Adrian A. de Freitas, Jaime Teevan, Steven Dow, Jeffrey P. Bigham |
CHI | 3 |
| 2016 | Exploring tilt for no-touch, wrist-only interactions on smartwatchesabstractBecause smartwatches are worn on the wrist, they do not require users to hold the device, leaving at least one hand free to engage in other activities. Unfortunately, this benefit is thwarted by the typical interaction model of smartwatches; for interactions beyond glancing at information or using speech, users must utilize their other hand to manipulate a touchscreen and/or hardware buttons. In order to enable no-touch, wrist-only smartwatch interactions so that users can, for example, hold a cup of coffee while controlling their device, we explore two tilt-based interaction techniques for menu selection and navigation: AnglePoint, which directly maps the position of a virtual pointer to the tilt angle of the smartwatch, and ObjectPoint, which objectifies the underlying virtual pointer as an object imbued with a physics model. In a user study, we found that participants were able to perform menu selection and continuous selection of menu items as well as navigation through a menu hierarchy more quickly and accurately with ObjectPoint, even though previous research on tilt for other mobile devices suggested that AnglePoint would be more effective. We provide an explanation of our results and discuss the implications for more "hands-free" smartwatch interactions. Anhong Guo, Tim Paek |
MobileHCI | 1 |
| 2016 | VizLens: A Robust and Interactive Screen Reader for Interfaces in the Real WorldabstractThe world is full of physical interfaces that are inaccessible to blind people, from microwaves and information kiosks to thermostats and checkout terminals. Blind people cannot independently use such devices without at least first learning their layout, and usually only after labeling them with sighted assistance. We introduce VizLens - an accessible mobile application and supporting backend that can robustly and interactively help blind people use nearly any interface they encounter. VizLens users capture a photo of an inaccessible interface and send it to multiple crowd workers, who work in parallel to quickly label and describe elements of the interface to make subsequent computer vision easier. The VizLens application helps users recapture the interface in the field of the camera, and uses computer vision to interactively describe the part of the interface beneath their finger (updating 8 times per second). We show that VizLens provides accurate and usable real-time feedback in a study with 10 blind participants, and our crowdsourcing labeling workflow was fast (8 minutes), accurate (99.7%), and cheap ($1.15). We then explore extensions of VizLens that allow it to (i) adapt to state changes in dynamic interfaces, (ii) combine crowd labeling with OCR technology to handle dynamic displays, and (iii) benefit from head-mounted cameras. VizLens robustly solves a long-standing challenge in accessibility by deeply integrating crowdsourcing and computer vision, and foreshadows a future of increasingly powerful interactive applications that would be currently impossible with either alone. Anhong Guo, Xiang 'Anthony' Chen, Samuel White, Chieko Asakawa, Jeffrey P. Bigham |
UIST | 1 |
| 2016 | Beyond the Touchscreen: An Exploration of Extending Interactions on Commodity SmartphonesabstractMost smartphones today have a rich set of sensors that could be used to infer input (e.g., accelerometer, gyroscope, microphone); however, the primary mode of interaction is still limited to the front-facing touchscreen and several physical buttons on the case. To investigate the potential opportunities for interactions supported by built-in sensors, we present the implementation and evaluation of BeyondTouch, a family of interactions to extend and enrich the input experience of a smartphone. Using only existing sensing capabilities on a commodity smartphone, we offer the user a wide variety of additional inputs on the case and the surface adjacent to the smartphone. Although most of these interactions are implemented with machine learning methods, compact and robust rule-based detection methods can also be applied for recognizing some interactions by analyzing physical characteristics of tapping events on the phone. This article is an extended version of Zhang et al. [2015], which solely covered gestures implemented by machine learning methods. We extended our previous work by adding gestures implemented with rule-based methods, which works well with different users across devices without collecting any training data. We outline the implementation of both machine learning and rule-based methods for these interaction techniques and demonstrate empirical evidence of their effectiveness and usability. We also discuss the practicality of BeyondTouch for a variety of application scenarios and compare the two different implementation methods. Cheng Zhang 0011, Anhong Guo, Dingtian Zhang, Yang Li 0058, Caleb Southern, Rosa I. Arriaga, Gregory D. Abowd |
ACM Trans. Interact. Intell. Syst. | 2 |
| 2015 | BeyondTouch: Extending the Input Language with Built-in Sensors on Commodity SmartphonesabstractWhile most smartphones today have a rich set of sensors that could be used to infer input (.e.g., accelerometer, gyroscope, microphone), the primary mode of interaction is still limited to the front-facing touchscreen and several physical buttons on the case. To investigate the potential opportunities for interactions supported by built-in sensors, we present the implementation and evaluation of BeyondTouch, a family of interactions to extend and enrich the input experience of a smartphone. Using only existing sensing capabilities on a commodity smartphone, we offer the user a wide variety of additional tapping and sliding inputs on the case of and the surface adjacent to the smartphone. We outline the implementation of these interaction techniques and demonstrate empirical evidence of their effectiveness and usability. We also discuss the practicality of BeyondTouch for a variety of application scenarios. Cheng Zhang 0011, Anhong Guo, Dingtian Zhang, Caleb Southern, Rosa I. Arriaga, Gregory D. Abowd |
IUI | 2 |