EDBT 2026 Demo / reviewers in the wild / expert
Amy Pavel
dblp:141/4106
· DBLP profile ↗
48ranked-venue papers
5as first author
33since 2021 · last 2026
0000-0002-3908-4366ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 43 · 5 first-author · 30 since 2021Artificial intelligence and machine learning · 4 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HistoryPalette: Supporting Exploration and Reuse of Past Alternatives in Image Generation and EditingabstractCreative tasks require creators to iteratively produce, select, and discard potentially useful ideas. Now, creativity tools include generative AI features (e.g., Photoshop Generative Fill) that increase the number of alternatives creators consider through rapid experiments with prompts and random generations. Creators use tedious manual systems for organizing their prior ideas by saving file versions or hiding layers, but they lack the support they want for reusing prior alternatives in personal work or in communication with others. We present HistoryPalette, a system that supports exploration and reuse of prior designs in generative image creation and editing. Using HistoryPalette, creators and their collaborators explore a “palette” of prior design alternatives organized by spatial position, topic category, and creation time. HistoryPalette enables creators to quickly preview and reuse their prior work. In creative professional and client collaborator user studies, participants generated and edited images by exploring and reusing past design alternatives with HistoryPalette. Karim Benharrak, Amy Pavel |
CHI | 2 |
| 2026 | Co-Designing Multimodal Systems for Accessible Asynchronous Dance InstructionabstractVideos make exercise instruction widely available, but they rely on visual demonstrations that blind and low vision (BLV) learners cannot see. While audio descriptions (AD) can make videos accessible, describing movements remains challenging as the AD must convey what to do (mechanics, location, orientation) and how to do it (speed, fluidity, timing). Prior work thus used multimodal instruction to support BLV learners with individual simple movements. However, it is unclear how these approaches scale to dance instruction with unique, complex movements and precise timing constraints. To inform accessible asynchronous dance instruction systems, we conducted three co-design workshops (N=28) with BLV dancers, instructors, and experts in sound, haptics, and AD. Participants designed 8 systems revealing common themes: staged learning to dissect routines, crafting vocabularies for movements, and selectively using modalities—narration for movement structure, sound for expression, and haptics for spatial cues. We conclude with design implications to make learning dance accessible. Ujjaini Das, Shreya Kappala, Meng Chen 0020, Mina Huh, Amy Pavel |
CHI | 5 |
| 2026 | A11y-CUA Dataset: Characterizing the Accessibility Gap in Computer Use AgentsabstractComputer Use Agents (CUAs) operate interfaces by pointing, clicking, and typing—mirroring interactions of sighted users (SUs) who can thus monitor CUAs and share control. CUAs do not reflect interactions by blind and low-vision users (BLVUs) who use assistive technology (AT). BLVUs thus cannot easily collaborate with CUAs. To characterize the accessibility gap of CUAs, we present A11y-CUA, a dataset of BLVUs and SUs performing 60 everyday tasks with 40.4 hours and 158,325 events. Our dataset analysis reveals that our collected interaction traces quantitatively confirm distinct interaction styles between SU and BLVU groups (mouse- vs. keyboard-dominant) and demonstrate interaction diversity within each group (sequential vs. shortcut navigation for BLVUs). We then compare collected traces to state-of-the-art CUAs under default and AT conditions (keyboard-only, magnifier). The default CUA executed 78.3% of tasks successfully. But with the AT conditions, CUA’s performance dropped to 41.67% and 28.3% with keyboard-only and magnifier conditions respectively, and did not reflect nuances of real AT use. With our open A11y-CUA dataset, we aim to promote collaborative and accessible CUAs for everyone. Ananya Gubbi Mohanbabu, Rosiana Natalie, Brandon Kim, Anhong Guo, Amy Pavel |
CHI | 5 |
| 2025 | Surfacing Variations to Calibrate Perceived Reliability of MLLM-generated Image DescriptionsabstractMultimodal large language models (MLLMs) provide new opportunities for blind and low vision (BLV) people to access visual information in their daily lives.However, these models often produce errors that are difficult to detect without sight, posing safety and social risks in scenarios from medication identification to outfit selection.While BLV MLLM users use creative workarounds such as crosschecking between tools and consulting sighted individuals, these approaches are often time-consuming and impractical.We explore how systematically surfacing variations across multiple MLLM responses can support BLV users to detect unreliable information without visually inspecting the image.We contribute a design space for eliciting and presenting variations in MLLM descriptions, a prototype system implementing three variation presentation styles, and findings from a user study with 15 BLV participants.Our results demonstrate that presenting variations significantly increases users' ability to identify unreliable claims (by 4.9x using our approach compared to single descriptions) and significantly decreases perceived reliability of MLLM responses.14 of 15 participants preferred seeing variations of MLLM responses over a single description, and all expressed interest in using our system for tasks from understanding a tornado's path to posting an image on social media. Meng Chen 0020, Akhil Iyer, Amy Pavel |
ASSETS | 3 |
| 2025 | Simultaneously Generating Multiple Mediums of Tactile GraphicsabstractFigure 1: To use our system, users (1) select and upload an image, (2-3) iteratively convert the image to a set of tactile graphics and optionally apply refinements, then (4) fabricate the resulting tactile graphics via embossing and 3D printing. Katherine Adele Clark, Amy Pavel |
ASSETS | 2 |
| 2025 | Task Mode: Dynamic Filtering for Task-Specific Web Navigation using LLMsabstractModern web interfaces are unnecessarily complex to use as they overwhelm users with excess text and visuals unrelated to their current goals.Such interfaces can particularly impact screen reader users (SRUs), who may need to navigate content sequentially and thus spend minutes traversing irrelevant elements compared to vision users (VUs) who visually skim in seconds.We present Task Mode, a system that dynamically filters web content based on userspecified goals using large language models to identify and prioritize relevant elements while minimizing distractions.Our approach preserves page structure while offering multiple viewing modes tailored to different access needs.Our user study with 12 participants (6 VUs, 6 SRUs) demonstrates that our approach halved task completion time for SRUs while maintaining performance for VUs, decreasing the completion time gap between groups from 2x to 1.2x.11 of 12 participants wanted to use Task Mode in the future, reporting that Task Mode supported completing tasks with less effort and fewer distractions.This work demonstrates how designing new interactions simultaneously for visual and non-visual access can reduce rather than reinforce accessibility disparities in future technology created by researchers and practitioners. CCS Concepts• Ananya Gubbi Mohanbabu, Yotam Sechayk, Amy Pavel |
ASSETS | 3 |
| 2025 | VeasyGuide: Personalized Visual Guidance for Low-vision Learners on Instructor Actions in Presentation VideosabstractToggle (a.1) Zoom settings panel (a.2) Highlight settings panel (a) The VeasyGuide player (b) Zoomed-in view (c) VeasyGuide's personalization settings panelsFigure 1: VeasyGuide's interface includes: (a) a video player with highlighted areas (red rectangle and hand pointer), (a.1) zoom settings toggle, (a.2) highlight settings toggle, (b) a zoomed window (toggled with the Z key), and (c) zoom and highlight settings panels.While watching a video, VeasyGuide auto-highlights pointing, marking, and sketching activities.Users can zoom into highlighted portions with the Z key, adjust zoom with arrow keys, and customize highlight and zoom appearance in real time. Yotam Sechayk, Ariel Shamir, Amy Pavel, Takeo Igarashi |
ASSETS | 3 |
| 2025 | VideoDiff: Human-AI Video Co-Creation with AlternativesabstractTo make an engaging video, people sequence interesting moments and add visuals such as B-rolls or text. While video editing requires time and effort, AI has recently shown strong potential to make editing easier through suggestions and automation. A key strength of generative models is their ability to quickly generate multiple variations, but when provided with many alternatives, creators struggle to compare them to find the best fit. We propose VideoDiff, an AI video editing tool designed for editing with alternatives. With VideoDiff, creators can generate and review multiple AI recommendations for each editing process: creating a rough cut, inserting B-rolls, and adding text effects. VideoDiff simplifies comparisons by aligning videos and highlighting differences through timelines, transcripts, and video previews. Creators have the flexibility to regenerate and refine AI suggestions as they compare alternatives. Our study participants (N=12) could easily compare and customize alternatives, creating more satisfying results. Mina Huh, Kim Pimmel, Hijung Shin, Amy Pavel, Mira Dontcheva |
CHI | 5 |
| 2025 | Lotus: Creating Short Videos From Long Videos With Abstractive and Extractive SummarizationabstractShort-form videos are popular on platforms like TikTok and Instagram as they quickly capture viewers' attention. Many creators repurpose their long-form videos to produce short-form videos, but creators report that planning, extracting, and arranging clips from long-form videos is challenging. Currently, creators make extractive short-form videos composed of existing long-form video clips or abstractive short-form videos by adding newly recorded narration to visuals. While extractive videos maintain the original connection between audio and visuals, abstractive videos offer flexibility in selecting content to be included in a shorter time. We present Lotus, a system that combines both approaches to balance preserving the original content with flexibility over the content. Lotus first creates an abstractive short-form video by generating both a short-form script and its corresponding speech, then matching long-form video clips to the generated narration. Creators can then add extractive clips with an automated method or Lotus's editing interface. Lotus's interface can be used to further refine the short-form video. We compare short-form videos generated by Lotus with those using an extractive baseline method. In our user study, we compare creating short-form videos using Lotus to participants' existing practice. Aadit Barua, Karim Benharrak, Meng Chen 0020, Mina Huh, Amy Pavel |
IUI | 5 |
| 2025 | TalkLess: Blending Extractive and Abstractive Summarization for Editing Speech to Preserve Content and StyleabstractRemovalFigure 1: TalkLess's interface lets creators edit their speech by skimming and browsing the outline pane to determine regions to cut, using the compression pane editing the transcript directly or by selecting a global-or paragraph-level compression amount, and listening to the original or generated audio aligned to the transcript in the audio pane. Karim Benharrak, Puyuan Peng, Amy Pavel |
UIST | 3 |
| 2025 | Vid2Coach: Transforming How-To Videos into Task Assistants
Mina Huh, Zihui Xue, Ujjaini Das, Kumar Ashutosh, Kristen Grauman, Amy Pavel |
UIST | 6 |
| 2025 | Morae: Proactively Pausing UI Agents for User Choices
Yi-Hao Peng, Dingzeyu Li, Jeffrey P. Bigham, Amy Pavel |
UIST | 4 |
| 2025 | CoSight: Exploring Viewer Contributions to Online Video Accessibility Through Descriptive CommentingabstractFigure 1: CoSight, a Chrome extension developed as a design probe to explore how lightweight interface nudges might encourage accessibility contributions from sighted video viewers when watching and commenting.The prototype augments YouTube video pages with features inspired by Fogg's Behavior Model [24], including color labels to highlight accessibility gaps (sparks), hints and references to guide contributions (facilitators), and reminders at key moments (signals). Ruolin Wang, Xingyu Liu 0002, Wayne Zhang 0004, Ziqian Liao, Ziwen Li 0001, Amy Pavel, Xiang 'Anthony' Chen |
UIST | 7 |
| 2024 | Context-Aware Image Descriptions for Web AccessibilityabstractBlind and low vision (BLV) internet users access images on the web via text descriptions. New vision-to-language models such as GPT-V, Gemini, and LLaVa can now provide detailed image descriptions on-demand. While prior research and guidelines state that BLV audiences’ information preferences depend on the context of the image, existing tools for accessing vision-to-language models provide only context-free image descriptions by generating descriptions for the image alone without considering the surrounding webpage context. To explore how to integrate image context into image descriptions, we designed a Chrome Extension that automatically extracts webpage context to inform GPT-4V-generated image descriptions. We gained feedback from 12 BLV participants in a user study comparing typical context-free image descriptions to context-aware image descriptions. We then further evaluated our context-informed image descriptions with a technical evaluation. Our user evaluation demonstrates that BLV participants frequently prefer context-aware descriptions to context-free descriptions. BLV participants also rate context-aware descriptions significantly higher in quality, imaginability, relevance, and plausibility. All participants shared that they wanted to use context-aware descriptions in the future and highlighted the potential for use in online shopping, social media, news, and personal interest blogs. Ananya Gubbi Mohanbabu, Amy Pavel |
ASSETS | 2 |
| 2024 | Design considerations for photosensitivity warnings in visual mediaabstractWhen digital content is tested for photosensitive safety and is found to contain seizure-inducing strobes or flashing lights, warnings about photosensitive risk are usually shown to the user prior to viewing the content. These photosensitivity warnings are an important accessibility feature for people with photosensitive epilepsy, allowing them to avoid interacting with content that may trigger seizures. However, little is known about how these warnings should be structured to maximize effectiveness in helping with people PSE navigate visual media safely. The design space for photosensitivity warnings is vast and includes questions such as what details to include about strobing light sequences or the content itself, where to place warnings within an interface, and what methods to use to extract information about the strobing light sequences (e.g., crowdsourced or automated methods). In this work, we contribute a thematic analysis of crowdsourced warnings drawn from the DoesTheDogDie online forum and an interview study with five people who have been diagnosed with photosensitive epilepsy about design considerations for photosensitivity warnings on digital platforms. To guide our interviews, we assembled examples of both crowdsourced and automated warnings about seizure-inducing content in films. Automated warnings were presented in the form of a high fidelity sketch demonstrating what an automated system for photosensitivity warnings might look like when deployed by a film streaming platform. We contribute design suggestions for the structure, content, and data sourcing of photosensitivity warnings for visual media based on the findings of our interviews. The results of this work will enable more effective and informative photosensitivity warnings across all forms of digital visual media. Laura South, Caglar Yildirim, Amy Pavel, Michelle Borkin |
ASSETS | 3 |
| 2024 | Making Short-Form Videos Accessible with Hierarchical Video SummariesabstractShort videos on platforms such as TikTok, Instagram Reels, and YouTube Shorts (i.e. short-form videos) have become a primary source of information and entertainment. Many short-form videos are inaccessible to blind and low vision (BLV) viewers due to their rapid visual changes, on-screen text, and music or meme-audio overlays. In our formative study, 7 BLV viewers who regularly watched short-form videos reported frequently skipping such inaccessible content. We present ShortScribe, a system that provides hierarchical visual summaries of short-form videos at three levels of detail to support BLV viewers in selecting and understanding short-form videos. ShortScribe allows BLV users to navigate between video descriptions based on their level of interest. To evaluate ShortScribe, we assessed description accuracy and conducted a user study with 10 BLV participants comparing ShortScribe to a baseline interface. When using ShortScribe, participants reported higher comprehension and provided more accurate summaries of video content. Tess Van Daele, Akhil Iyer, Jalyn C. Derry, Mina Huh, Amy Pavel |
CHI | 6 |
| 2024 | Barriers to Photosensitive Accessibility in Virtual RealityabstractVirtual reality (VR) systems have grown in popularity as an immersive modality for daily activities such as gaming, socializing, and working. However, this technology is not always accessible for people with photosensitive epilepsy (PSE) who may experience seizures or other adverse symptoms when exposed to certain light stimuli (e.g., flashes or strobes). How can VR be made more inclusive and safer for people with PSE? In this paper, we report on a series of semi-structured interviews about current perceptions of accessibility in VR among people with PSE. We identify 12 barriers to accessibility that fall into four categories: physical VR equipment, VR interfaces and content, specific VR applications, and individual differences in sensitivity. Our findings allow researchers and practitioners to better understand the meaning of photosensitive accessibility in the context of VR, and provide a step towards enabling people with PSE to enjoy the benefits offered by immersive technology. Laura South, Caglar Yildirim, Amy Pavel, Michelle Borkin |
CHI | 3 |
| 2024 | COMPA: Using Conversation Context to Achieve Common Ground in AACabstractGroup conversations often shift quickly from topic to topic, leaving a small window of time for participants to contribute. AAC users often miss this window due to the speed asymmetry between using speech and using AAC devices. AAC users may take over a minute longer to contribute, and this speed difference can cause mismatches between the ongoing conversation and the AAC user’s response. This results in misunderstandings and missed opportunities to participate. We present COMPA, an add-on tool for online group conversations that seeks to support conversation partners in achieving common ground. COMPA uses a conversation’s live transcription to enable AAC users to mark conversation segments they intend to address (Context Marking) and generate contextual starter phrases related to the marked conversation segment (Phrase Assistance) and a selected user intent. We study COMPA in 5 different triadic group conversations, each composed by a researcher, an AAC user and a conversation partner (n=10) and share findings on how conversational context supports conversation partners in achieving common ground. Stephanie Valencia, Jessica Huynh, Emma Y. Jiang, Yufei Wu 0020, Teresa Wan, Zixuan Zheng, Henny Admoni, Jeffrey P. Bigham, Amy Pavel |
CHI | 9 |
| 2024 | DreamStruct: Understanding Slides and User Interfaces via Synthetic Data Generation
Yi-Hao Peng, Faria Huq, Yue Jiang 0002, Jason Wu 0001, Xin Yue Li, Jeffrey P. Bigham, Amy Pavel |
ECCV (24) | 7 |
| 2024 | DesignChecker: Visual Design Support for Blind and Low Vision Web DevelopersabstractBlind and low vision (BLV) developers create websites to share knowledge and showcase their work. A well-designed website can engage audiences and deliver information effectively, yet it remains challenging for BLV developers to review their web designs. We conducted interviews with BLV developers (N=9) and analyzed 20 websites created by BLV developers. BLV developers created highly accessible websites but wanted to assess the usability of their websites for sighted users and follow the design standards of other websites. They also encountered challenges using screen readers to identify illegible text, misaligned elements, and inharmonious colors. We present DesignChecker, a browser extension that helps BLV developers improve their web designs. With DesignChecker, users can assess their current design by comparing it to visual design guidelines, a reference website of their choice, or a set of similar websites. DesignChecker also identifies the specific HTML elements that violate design guidelines and suggests CSS changes for improvements. Our user study participants (N=8) recognized more visual design errors than using their typical workflow and expressed enthusiasm about using DesignChecker in the future. Mina Huh, Amy Pavel |
UIST | 2 |
| 2023 | Exploring Community-Driven Descriptions for Making Livestreams AccessibleabstractPeople watch livestreams to connect with others and learn about their hobbies. Livestreams feature multiple visual streams including the main video, webcams, on-screen overlays, and chat, all of which are inaccessible to livestream viewers with visual impairments. While prior work explores creating audio descriptions for recorded videos, live videos present new challenges: authoring descriptions in real-time, describing domain-specific content, and prioritizing which complex visual information to describe. We explore inviting livestream community members who are domain experts to provide live descriptions. We first conducted a study with 18 sighted livestream community members authoring descriptions for livestreams using three different description methods: live descriptions using text, live descriptions using speech, and asynchronous descriptions using text. We then conducted a study with 9 livestream community members with visual impairments, who shared their current strategies and challenges for watching livestreams and provided feedback on the community-written descriptions. We conclude with implications for improving the accessibility of livestreams. Daniel Killough, Amy Pavel |
ASSETS | 2 |
| 2023 | AVscript: Accessible Video Editing with Audio-Visual ScriptsabstractSighted and blind and low vision (BLV) creators alike use videos to communicate with broad audiences. Yet, video editing remains inaccessible to BLV creators. Our formative study revealed that current video editing tools make it difficult to access the visual content, assess the visual quality, and efficiently navigate the timeline. We present AVscript, an accessible text-based video editor. AVscript enables users to edit their video using a script that embeds the video’s visual content, visual errors (e.g., dark or blurred footage), and speech. Users can also efficiently navigate between scenes and visual errors or locate objects in the frame or spoken words of interest. A comparison study (N=12) showed that AVscript significantly lowered BLV creators’ mental demands while increasing confidence and independence in video editing. We further demonstrate the potential of AVscript through an exploratory study (N=3) where BLV creators edited their own footage. Mina Huh, Saelyne Yang, Yi-Hao Peng, Xiang 'Anthony' Chen, Young-Ho Kim, Amy Pavel |
CHI | 6 |
| 2023 | SlideSpecs: Automatic and Interactive Presentation Feedback CollationabstractPresenters often collect audience feedback through practice talks to refine their presentations. In formative interviews, we find that although text feedback and verbal discussions allow presenters to receive feedback, organizing that feedback into actionable presentation revisions remains challenging. Feedback may lack context, be redundant, and be spread across various emails, notes, and conversations. To collate and contextualize both text and verbal feedback, we present SlideSpecs. SlideSpecs lets audience members provide text feedback (e.g., ‘font too small’) while attaching an automatically detected context, including relevant slides (e.g., ‘Slide 7’) or content tags (e.g., ‘slide design’). SlideSpecs also records and transcribes spoken group discussions that commonly occur after practice talks and facilitates linking text critiques to relevant discussion segments. Finally, presenters can use SlideSpecs to review all text and spoken feedback in a single contextually rich interface (e.g., relevant slides, topics, and follow-up discussions). We demonstrate the effectiveness of SlideSpecs by deploying it in eight practice talks with a range of topics and purposes and reporting our findings. Jeremy Warner, Amy Pavel, Tonya Nguyen, Maneesh Agrawala, Björn Hartmann |
IUI | 2 |
| 2023 | GenAssist: Making Image Generation AccessibleabstractBlind and low vision (BLV) creators use images to communicate with sighted audiences. However, creating or retrieving images is challenging for BLV creators as it is difficult to use authoring tools or assess image search results. Thus, creators limit the types of images they create or recruit sighted collaborators. While text-to-image generation models let creators generate high-fidelity images based on a text description (i.e. prompt), it is difficult to assess the content and quality of generated images. We present GenAssist, a system to make text-to-image generation accessible. Using our interface, creators can verify whether generated image candidates followed the prompt, access additional details in the image not specified in the prompt, and skim a summary of similarities and differences between image candidates. To power the interface, GenAssist uses a large language model to generate visual questions, vision-language models to extract answers, and a large language model to summarize the results. Our study with 12 BLV creators demonstrated that GenAssist enables and simplifies the process of image selection and generation, making visual authoring more accessible to all. Mina Huh, Yi-Hao Peng, Amy Pavel |
UIST | 3 |
| 2023 | GeoLatent: A Geometric Approach to Latent Space Design for Deformable Shape GeneratorsabstractWe study how to optimize the latent space of neural shape generators that map latent codes to 3D deformable shapes. The key focus is to look at a deformable shape generator from a differential geometry perspective. We define a Riemannian metric based on as-rigid-as-possible and as-conformal-as-possible deformation energies. Under this metric, we study two desired properties of the latent space: 1) straight-line interpolations in latent codes follow geodesic curves; 2) latent codes disentangle pose and shape variations at different scales. Strictly enforcing the geometric interpolation property, however, only applies if the metric matrix is a constant. We show how to achieve this property approximately by enforcing that geodesic interpolations are axis-aligned, i.e., interpolations along coordinate axis follow geodesic curves. In addition, we introduce a novel approach that decouples pose and shape variations via generalized eigendecomposition. We also study efficient regularization terms for learning deformable shape generators, e.g., that promote smooth interpolations. Experimental results on benchmark datasets show that our approach leads to interpretable latent codes, improves the generalizability of synthetic shapes, and enhances performance in geodesic interpolation and geodesic shooting. Haitao Yang 0005, Amy Pavel, Qixing Huang |
ACM Trans. Graph. | 4 |
| 2022 | Tech Help Desk: Support for Local Entrepreneurs Addressing the Long Tail of Computing ChallengesabstractEven entrepreneurs whose businesses are not technological (e.g., handmade goods) need to be able to use a wide range of computing technologies in order to achieve their business goals. In this paper, we follow a participatory action research approach and collaborate with various stakeholders at an entrepreneurial co-working space to design “Tech Help Desk”, an on-going technical service for entrepreneurs. Our model for technical assistance is strategic, in how it is designed to fit the context of local entrepreneurs, and responsive, in how it prioritizes emergent needs. From our engagements with 19 entrepreneurs and support personnel, we reflect on the challenges with existing technology support for non-technological entrepreneurs. Our work highlights the importance of ensuring technological support services can adapt based on entrepreneurs’ ever-evolving priorities, preferences and constraints. Furthermore, we find technological support services should maintain broad technical support for entrepreneurs’ long tail of computing challenges. Yasmine Kotturi, Herman T. Johnson, Michael Skirpan, Sarah E. Fox, Jeffrey P. Bigham, Amy Pavel |
CHI | 6 |
| 2022 | CrossA11y: Identifying Video Accessibility Issues via Cross-modal GroundingabstractAuthors make their videos visually accessible by adding audio descriptions (AD), and auditorily accessible by adding closed captions (CC). However, creating AD and CC is challenging and tedious, especially for non-professional describers and captioners, due to the difficulty of identifying accessibility problems in videos. A video author will have to watch the video through and manually check for inaccessible information frame-by-frame, for both visual and auditory modalities. In this paper, we present CrossA11y, a system that helps authors efficiently detect and address visual and auditory accessibility issues in videos. Using cross-modal grounding analysis, CrossA11y automatically measures accessibility of visual and audio segments in a video by checking for modality asymmetries. CrossA11y then displays these segments and surfaces visual and audio accessibility issues in a unified interface, making it intuitive to locate, review, script AD/CC in-place, and preview the described and captioned video immediately. We demonstrate the effectiveness of CrossA11y through a lab study with 11 participants, comparing to existing baseline. Xingyu Liu 0002, Ruolin Wang, Dingzeyu Li, Xiang 'Anthony' Chen, Amy Pavel |
UIST | 5 |
| 2022 | Diffscriber: Describing Visual Design Changes to Support Mixed-Ability Collaborative Presentation AuthoringabstractVisual slide-based presentations are ubiquitous, yet slide authoring tools are largely inaccessible to people who are blind or visually impaired (BVI). When authoring presentations, the 9 BVI presenters in our formative study usually work with sighted collaborators to produce visual slides based on the text content they produce. While BVI presenters valued collaborators’ visual design skill, the collaborators often felt they could not fully review and provide feedback on the visual changes that were made. We present Diffscriber, a system that identifies and describes changes to a slide’s content, layout, and style for presentation authoring. Using our system, BVI presentation authors can efficiently review changes to their presentation by navigating either a summary of high-level changes or individual slide elements. To learn more about changes of interest, presenters can use a generated change hierarchy to navigate to lower-level change details and element styles. BVI presenters using Diffscriber were able to identify slide design changes and provide feedback more easily as compared to using only the slides alone. More broadly, Diffscriber illustrates how advances in detecting and describing visual differences can improve mixed-ability collaboration. Yi-Hao Peng, Jason Wu 0001, Jeffrey P. Bigham, Amy Pavel |
UIST | 4 |
| 2021 | Slidecho: Flexible Non-Visual Exploration of Presentation VideosabstractWe present Slidecho, a system that enables non-visual access of the slide content in a presentation video on-demand. Slidecho automatically extracts slides and their text and image elements from the presentation video and aligns these elements to the presenter’s speech. When listening to the video, Slidecho provides learners with audio notifications about slide changes and slide elements that are not described by the presenter. The learner can pause the video and browse the entire slide, or only the undescribed slide elements, to gain information. A technical evaluation with presentation videos in-the-wild shows that compared to the presenter’s speech alone, Slidecho provides access to an additional 20% of total text elements and 30% of total image elements that were previously not described. Blind and visually impaired participants in our user study reported that it was easier to locate undescribed slide elements with Slidecho’s synchronized interface than when browsing the video and extracted slides separately, and using Slidecho they read fewer slides that were fully redundant with the speech. Yi-Hao Peng, Jeffrey P. Bigham, Amy Pavel |
ASSETS | 3 |
| 2021 | What Makes Videos Accessible to Blind and Visually Impaired People?abstractUser-generated videos are an increasingly important source of information online, yet most online videos are inaccessible to blind and visually impaired (BVI) people. To find videos that are accessible, or understandable without additional description of the visual content, BVI people in our formative studies reported that they used a time-consuming trial-and-error approach: clicking on a video, watching a portion, leaving the video, and repeating the process. BVI people also reported video accessibility heuristics that characterize accessible and inaccessible videos. We instantiate 7 of the identified heuristics (2 audio-related, 2 video-related, and 3 audio-visual) as automated metrics to assess video accessibility. We collected a dataset of accessibility ratings of videos by BVI people and found that our automatic video accessibility metrics correlated with the accessibility ratings (Adjusted R2 = 0.642). We augmented a video search interface with our video accessibility metrics and predictions. BVI people using our augmented video search interface selected an accessible video more efficiently than when using the original search interface. By integrating video accessibility metrics, video hosting platforms could help people surface accessible videos and encourage content creators to author more accessible products, improving video accessibility for all. Xingyu Liu 0002, Patrick Carrington, Xiang 'Anthony' Chen, Amy Pavel |
CHI | 4 |
| 2021 | Say It All: Feedback for Improving Non-Visual Presentation AccessibilityabstractPresenters commonly use slides as visual aids for informative talks. When presenters fail to verbally describe the content on their slides, blind and visually impaired audience members lose access to necessary content, making the presentation difficult to follow. Our analysis of 90 presentation videos revealed that 72% of 610 visual elements (e.g., images, text) were insufficiently described. To help presenters create accessible presentations, we introduce Presentation A11y, a system that provides real-time and post-presentation accessibility feedback. Our system analyzes visual elements on the slide and the transcript of the verbal presentation to provide element-level feedback on what visual content needs to be further described or even removed. Presenters using our system with their own slide-based presentations described more of the content on their slides, and identified 3.26 times more accessibility problems to fix after the talk than when using a traditional slide-based presentation interface. Integrating accessibility feedback into content creation tools will improve the accessibility of informational content for all. Yi-Hao Peng, JiWoong Jang, Jeffrey P. Bigham, Amy Pavel |
CHI | 4 |
| 2021 | Co-designing Socially Assistive Sidekicks for Motion-based AACabstractAugmentative and alternative communication (AAC) devices enable speech-based communication. However, AAC devices do not support nonverbal communication, which allows people to take turns, regulate conversation dynamics, and express intentions. Nonverbal communication requires motion, which is often challenging for AAC users to produce due to motor constraints. In this work, we explore how socially assistive robots, framed as ''sidekicks,'' might provide augmented communicators (ACs) with a nonverbal channel of communication to support their conversational goals. We developed and conducted an accessible co-design workshop that involved two ACs, their caregivers, and three motion experts. We identified goals for conversational support, co-designed prototypes depicting possible sidekick forms, and enacted different sidekick motions and behaviors to achieve speakers' goals. We contribute guidelines for designing sidekicks that support ACs according to three key parameters: attention, precision, and timing. We show how these parameters manifest in appearance and behavior and how they can guide future designs for augmented nonverbal communication. Stephanie Valencia, Michal Luria, Amy Pavel, Jeffrey P. Bigham, Henny Admoni |
HRI | 3 |
| 2021 | Controlling Dialogue Generation with Semantic ExemplarsabstractPrakhar Gupta, Jeffrey Bigham, Yulia Tsvetkov, Amy Pavel. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Prakhar Gupta, Jeffrey P. Bigham, Yulia Tsvetkov, Amy Pavel |
NAACL-HLT | 4 |
| 2020 | Making GIFs AccessibleabstractSocial media platforms feature short animations known as GIFs, but they are inaccessible to people with vision impairments. Unlike static images, GIFs contain action and visual indications of sound, which can be challenging to describe in alternative text descriptions. We examine a large sample of inaccessible GIFs on Twitter to document how they are used and what visual elements they contain. In interviews with 10 blind Twitter users, we discuss what elements of GIF content should be described and their experiences with GIFs online. The participants compared alternative text descriptions with two other alternative audio formats: (i) the original audio from the GIF source video and (ii) a spoken audio description. We recommend that social media platforms automatically include alt text descriptions for popular GIFs (as Twitter has begun to do), and content producers create audio descriptions to ensure everyone has a rich and emotive experience with GIFs online. Cole Gleason, Amy Pavel, Himalini Gururaj, Kris Makoto Kitani, Jeffrey P. Bigham |
ASSETS | 2 |
| 2020 | Disability and the COVID-19 Pandemic: Using Twitter to Understand Accessibility during Rapid Societal TransitionabstractThe COVID-19 pandemic has forced institutions to rapidly alter their behavior, which typically has disproportionate negative effects on people with disabilities as accessibility is overlooked. To investigate these issues, we analyzed Twitter data to examine accessibility problems surfaced by the crisis. We identified three key domains at the intersection of accessibility and technology: (i) the allocation of product delivery services, (ii) the transition to remote education, and (iii) the dissemination of public health information. We found that essential retailers expanded their high-risk customer shopping hours and pick-up and delivery services, but individuals with disabilities still lacked necessary access to goods and services. Long-experienced access barriers to online education were exacerbated by the abrupt transition of in-person to remote instruction. Finally, public health messaging has been inconsistent and inaccessible, which is unacceptable during a rapidly-evolving crisis. We argue that organizations should create flexible, accessible technology and policies in calm times to be adaptable in times of crisis to serve individuals with diverse needs. Cole Gleason, Stephanie Valencia, Lynn Kirabo, Jason Wu 0001, Anhong Guo, Elizabeth J. Carter, Jeffrey P. Bigham, Cynthia L. Bennett, Amy Pavel |
ASSETS | 9 |
| 2020 | Making Mobile Augmented Reality Applications AccessibleabstractAugmented Reality (AR) technology creates new immersive experiences in entertainment, games, education, retail, and social media. AR content is often primarily visual and it is challenging to enable access to it non-visually due to the mix of virtual and real-world content. In this paper, we identify common constituent tasks in AR by analyzing existing mobile AR applications for iOS, and characterize the design space of tasks that require accessible alternatives. For each of the major task categories, we create prototype accessible alternatives that we evaluate in a study with 10 blind participants to explore their perceptions of accessible AR. Our study demonstrates that these prototypes make AR possible to use for blind users and reveals a number of insights to move forward. We believe our work sets forth not only exemplars for developers to create accessible AR applications, but also a roadmap for future research to make AR comprehensively accessible. Jaylin Herskovitz, Jason Wu 0001, Samuel White, Amy Pavel, Gabriel Reyes, Anhong Guo, Jeffrey P. Bigham |
ASSETS | 4 |
| 2020 | Twitter A11y: A Browser Extension to Make Twitter Images AccessibleabstractSocial media platforms are integral to public and private discourse, but are becoming less accessible to people with vision impairments due to an increase in user-posted images. Some platforms (i.e. Twitter) let users add image descriptions (alternative text), but only 0.1% of images include these. To address this accessibility barrier, we created Twitter A11y, a browser extension to add alternative text on Twitter using six methods. For example, screenshots of text are common, so we detect textual images, and create alternative text using optical character recognition. Twitter A11y also leverages services to automatically generate alternative text or reuse them from across the web. We compare the coverage and quality of Twitter A11y's six alt-text strategies by evaluating the timelines of 50 self-identified blind Twitter users. We find that Twitter A11y increases alt-text coverage from 7.6% to 78.5%, before crowdsourcing descriptions for the remaining images. We estimate that 57.5% of returned descriptions are high-quality. We then report on the experiences of 10 participants with visual impairments using the tool during a week-long deployment. Twitter A11y increases access to social media platforms for people with visual impairments by providing high-quality automatic descriptions for user-posted images. Cole Gleason, Amy Pavel, Emma McCamey, Christina Low, Patrick Carrington, Kris Makoto Kitani, Jeffrey P. Bigham |
CHI | 2 |
| 2020 | Conversational Agency in Augmentative and Alternative CommunicationabstractAugmented communicators (ACs) use augmentative and alternative communication (AAC) technologies to speak. Prior work in AAC research has looked to improve efficiency and expressivity of AAC via device improvements and user training. However, ACs also face constraints in communication beyond their device and individual abilities such as when they can speak, what they can say, and who they can address. In this work, we recast and broaden this prior work using conversational agency as a new frame to study AC communication. We investigate AC conversational agency with a study examining different conversational tasks between four triads of expert ACs, their close conversation partners (paid aide or parent), and a third party (experimenter). We define metrics to analyze AAC conversational agency quantitatively and qualitatively. We conclude with implications for future research to enable ACs to easily exercise conversational agency. Stephanie Valencia, Amy Pavel, Jared Santa Maria, Seunga (Gloria) Yu, Jeffrey P. Bigham, Henny Admoni |
CHI | 2 |
| 2020 | Rescribe: Authoring and Automatically Editing Audio DescriptionsabstractAudio descriptions make videos accessible to those who cannot see them by describing visual content in audio. Producing audio descriptions is challenging due to the synchronous nature of the audio description that must fit into gaps of other video content. An experienced audio description author will produce content that fits narration necessary to understand, enjoy, or experience the video content into the time available. This can be especially tricky for novices to do well. In this paper, we introduce a tool, Rescribe, that helps authors create and refine their audio descriptions. Using Rescribe, authors first create a draft of all the content they would like to include in the audio description. Rescribe then uses a dynamic programming approach to optimize between the length of the audio description, available automatic shortening approaches, and source track lengthening approaches. Authors can iteratively visualize and refine the audio descriptions produced by Rescribe, working in concert with the tool. We evaluate the effectiveness of Rescribe through interviews with blind and visually impaired audio description users who give feedback on Rescribe results. In addition, we invite novice users to create audio descriptions with Rescribe and another tool, finding that users produce audio descriptions with fewer placement errors using Rescribe. Amy Pavel, Gabriel Reyes, Jeffrey P. Bigham |
UIST | 1 |
| 2019 | Making Memes AccessibleabstractImages on social media platforms are inaccessible to people with vision impairments due to a lack of descriptions that can be read by screen readers. Providing accurate alternative text for all visual content on social media is not yet feasible, but certain subsets of images, such as internet memes, offer affordances for automatic or semi-automatic generation of alternative text. We present two methods for making memes accessible semi-automatically through (1) the generation of rich alternative text descriptions and (2) the creation of audio macro memes. Meme authors create alternative text templates or audio meme templates, and insert placeholders instead of the meme text. When a meme with the same image is encountered again, it is automatically recognized from a database of meme templates. Text is then extracted and either inserted into the alternative text template or rendered in the audio template using text-to-speech. In our evaluation of meme formats with 10 Twitter users with vision impairments, we found that most users preferred alternative text memes because the description of the visual content conveys the emotional tone of the character. As the preexisting templates can be automatically matched to memes using the same visual image, this combined approach can make a large subset of images on the web accessible, while preserving the emotion and tone inherent in the image memes. Cole Gleason, Amy Pavel, Xingyu Liu 0002, Patrick Carrington, Lydia B. Chilton, Jeffrey P. Bigham |
ASSETS | 2 |
| 2019 | Twitter A11y: A Browser Extension to Describe ImagesabstractTwitter is integral to many people's lives for news, entertainment, and communication. While people increasingly post images to Twitter, a large majority of images remain inaccessible to people with vision impairments due to a lack of image descriptions (i.e. alternative text). We present Twitter A11y (pronounced ally), a browser extension to make images accessible through a set of strategies tailored to the platform. For example, screenshots of text that exceed the Twitter character limit are common, so we detect textual images, and automatically add alternative text using optical character recognition. Tweet images apart from screenshots and link previews receive descriptions from crowd workers. Based on an evaluation of the timelines of 50 self-identified blind Twitter users, Twitter A11y increases automatic alt text coverage from 2.6% to 25.6%, before crowdsourcing the remaining images. Christina Low, Emma McCamey, Cole Gleason, Patrick Carrington, Jeffrey P. Bigham, Amy Pavel |
ASSETS | 6 |
| 2019 | Investigating Evaluation of Open-Domain Dialogue Systems With Human Generated Multiple ReferencesabstractThe aim of this paper is to mitigate the shortcomings of automatic evaluation of open-domain dialog systems through multireference evaluation.Existing metrics have been shown to correlate poorly with human judgement, particularly in open-domain dialog.One alternative is to collect human annotations for evaluation, which can be expensive and time consuming.To demonstrate the effectiveness of multi-reference evaluation, we augment the test set of DailyDialog with multiple references.A series of experiments show that the use of multiple references results in improved correlation between several automatic metrics and human judgement for both the quality and the diversity of system output. Prakhar Gupta, Shikib Mehri, Amy Pavel, Maxine Eskénazi, Jeffrey P. Bigham |
SIGdial | 4 |
| 2018 | Saliency in VR: How Do People Explore Virtual Environments?abstractUnderstanding how people explore immersive virtual environments is crucial for many applications, such as designing virtual reality (VR) content, developing new compression algorithms, or learning computational models of saliency or visual attention. Whereas a body of recent work has focused on modeling saliency in desktop viewing conditions, VR is very different from these conditions in that viewing behavior is governed by stereoscopic vision and by the complex interaction of head orientation, gaze, and other kinematic constraints. To further our understanding of viewing behavior and saliency in VR, we capture and analyze gaze and head orientation data of 169 users exploring stereoscopic, static omni-directional panoramas, for a total of 1980 head and gaze trajectories for three different viewing conditions. We provide a thorough analysis of our data, which leads to several important insights, such as the existence of a particular fixation bias, which we then use to adapt existing saliency predictors to immersive VR conditions. In addition, we explore other applications of our data and analysis, including automatic alignment of VR video cuts, panorama thumbnails, panorama video synopsis, and saliency-basedcompression. Vincent Sitzmann, Ana Serrano, Amy Pavel, Maneesh Agrawala, Diego Gutierrez, Belén Masiá, Gordon Wetzstein |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2017 | Shot Orientation Controls for Interactive Cinematography with 360 VideoabstractVirtual reality filmmakers creating 360-degree video currently rely on cinematography techniques that were developed for traditional narrow field of view film. They typically edit together a sequence of shots so that they appear at a fixed-orientation irrespective of the viewer's field of view. But because viewers set their own camera orientation they may miss important story content while looking in the wrong direction. We present new interactive shot orientation techniques that are designed to help viewers see all of the important content in 360-degree video stories. Our viewpoint-oriented technique reorients the shot at each cut so that the most important content lies in the the viewer's current field of view. Our active reorientation technique, lets the viewer press a button to immediately reorient the shot so that important content lies in their field of view. We present a 360-degree video player which implements these techniques and conduct a user study which finds that users spend 5.2-9.5% more time viewing the important points (manually labelled) of the scene with our techniques compared to the traditional fixed-orientation cuts. In practice, 360-degree video creators may label important content, but we also provide an automatic method for determining important content in existing 360-degree videos. Amy Pavel, Björn Hartmann, Maneesh Agrawala |
UIST | 1 |
| 2016 | VidCrit: Video-based Asynchronous Video ReviewabstractVideo production is a collaborative process in which stakeholders regularly review drafts of the edited video to indicate problems and offer suggestions for improvement. Although practitioners prefer in-person feedback, most reviews are conducted asynchronously via email due to scheduling and location constraints. The use of this impoverished medium is challenging for both providers and consumers of feedback. We introduce VidCrit, a system for providing asynchronous feedback on drafts of edited video that incorporates favorable qualities of an in-person review. This system consists of two separate interfaces: (1) A feedback recording interface captures reviewers' spoken comments, mouse interactions, hand gestures and other physical reactions. (2) A feedback viewing interface transcribes and segments the recorded review into topical comments so that the video author can browse the review by either text or timelines. Our system features novel methods to automatically segment a long review session into topical text comments, and to label such comments with additional contextual information. We interviewed practitioners to inform a set of design guidelines for giving and receiving feedback, and based our system's design on these guidelines. Video reviewers using our system preferred our feedback recording interface over email for providing feedback due to the reduction in time and effort. In a fixed amount of time, reviewers provided 10.9 (σ=5.09) more local comments than when using text. All video authors rated our feedback viewing interface preferable to receiving feedback via e-mail. Amy Pavel, Dan B. Goldman, Björn Hartmann, Maneesh Agrawala |
UIST | 1 |
| 2015 | Structuring, Aggregating, and Evaluating Crowdsourced Design CritiqueabstractFeedback is an important component of the design process, but gaining access to high-quality critique outside a classroom or firm is challenging. We present CrowdCrit, a web-based system that allows designers to receive design critiques from non-expert crowd workers. We evaluated CrowdCrit in three studies focusing on the designer's experience and benefits of the critiques. In the first study, we compared crowd and expert critiques and found evidence that aggregated crowd critique approaches expert critique. In a second study, we found that designers who got crowd feedback perceived that it improved their design process. The third study showed that designers were enthusiastic about crowd critiques and used them to change their designs. We conclude with implications for the design of crowd feedback services. Kurt Luther, Jari-Lee Tolentino, Amy Pavel, Brian P. Bailey, Maneesh Agrawala, Björn Hartmann, Steven Dow |
CSCW | 4 |
| 2015 | SceneSkim: Searching and Browsing Movies Using Synchronized Captions, Scripts and Plot SummariesabstractSearching for scenes in movies is a time-consuming but crucial task for film studies scholars, film professionals, and new media artists. In pilot interviews we have found that such users search for a wide variety of clips---e.g., actions, props, dialogue phrases, character performances, locations---and they return to particular scenes they have seen in the past. Today, these users find relevant clips by watching the entire movie, scrubbing the video timeline, or navigating via DVD chapter menus. Increasingly, users can also index films through transcripts---however, dialogue often lacks visual context, character names, and high level event descriptions. We introduce SceneSkim, a tool for searching and browsing movies using synchronized captions, scripts and plot summaries. Our interface integrates information from such sources to allow expressive search at several levels of granularity: Captions provide access to accurate dialogue, scripts describe shot-by-shot actions and settings, and plot summaries contain high-level event descriptions. We propose new algorithms for finding word-level caption to script alignments, parsing text scripts, and aligning plot summaries to scripts. Film studies graduate students evaluating SceneSkim expressed enthusiasm about the usability of the proposed system for their research and teaching. Amy Pavel, Dan B. Goldman, Björn Hartmann, Maneesh Agrawala |
UIST | 1 |
| 2014 | Video digests: a browsable, skimmable format for informational lecture videosabstractIncreasingly, authors are publishing long informational talks, lectures, and distance-learning videos online. However, it is difficult to browse and skim the content of such videos using current timeline-based video players. Video digests are a new format for informational videos that afford browsing and skimming by segmenting videos into a chapter/section structure and providing short text summaries and thumbnails for each section. Viewers can navigate by reading the summaries and clicking on sections to access the corresponding point in the video. We present a set of tools to help authors create such digests using transcript-based interactions. With our tools, authors can manually create a video digest from scratch, or they can automatically generate a digest by applying a combination of algorithmic and crowdsourcing techniques and then manually refine it as needed. Feedback from first-time users suggests that our transcript-based authoring tools and automated techniques greatly facilitate video digest creation. In an evaluative crowdsourced study we find that given a short viewing time, video digests support browsing and skimming better than timeline-based or transcript-based video players. Amy Pavel, Colorado Reed, Björn Hartmann, Maneesh Agrawala |
UIST | 1 |