VLDB 2026 Research / reviewers in the wild / expert
Jeffrey P. Bigham
dblp:83/6818 · also Jeffrey Philip Bigham
· DBLP profile ↗
163ranked-venue papers
21as first author
42since 2021 · last 2026
0000-0002-2072-0625ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 132 · 17 first-author · 30 since 2021Artificial intelligence and machine learning · 24 · 1 first-author · 11 since 2021Databases, data management, data science and information retrieval · 19 · 3 first-authorGraphics, computer vision, multimedia, augmented reality and games · 15 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 3 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | iTagPDF: Towards Finally Automating PDF Accessibility
Peya Mowar, Aaron Steinfeld, Jeffrey P. Bigham |
CHI | 3 |
| 2026 | "I followed what felt right, not what I was told": Autonomy, Coaching, and Recognizing Bias Through AI-Mediated DialogueabstractAbleist microaggressions remain pervasive in everyday interactions, yet interventions to help people recognize them are limited. We present an experiment testing how AI-mediated dialogue influences recognition of ableism. 160 participants completed a pre-test, intervention, and a post-test across four conditions: AI nudges toward bias (Bias-Directed), inclusion (Neutral-Directed), unguided dialogue (Self-Directed), and a text-only non-dialogue (Reading). Participants rated scenarios on standardness of social experience and emotional impact; those in dialogue-based conditions also provided qualitative reflections. Quantitative results showed dialogue-based conditions produced stronger recognition than Reading, though trajectories diverged: biased nudges improved differentiation of bias from neutrality but increased overall negativity. Inclusive or no nudges remained more balanced, while Reading participants showed weaker gains and even declines. Qualitative findings revealed biased nudges were often rejected, while inclusive nudges were adopted as scaffolding. We contribute a validated vignette corpus, an AI-mediated intervention platform, and design implications highlighting trade-offs conversational systems face when integrating bias-related nudges. Atieh Taheri, Hamza El Alaoui, Patrick Carrington, Jeffrey P. Bigham |
CHI | 4 |
| 2025 | We Write Our Research Papers in WYSIWYM. Why Do We Tag Our PDFs in WYSIWYG?
Peya Mowar, Aaron Steinfeld, Jeffrey P. Bigham |
ASSETS | 3 |
| 2025 | Designing Through Lived Experience: Reflections on Control, Embodiment, and Social Bias in Accessibility ResearchabstractThis paper presents an analytic autoethnography of three accessibility research projects, MouseClicker, Virtual Steps, and Simulated Conversations, led by the first author, a disabled researcher with Spinal Muscular Atrophy.Each project emerged from personal need and embodied experience, and together they explore new possibilities in accessible interaction, sensation, and social engagement.Drawing from feminist HCI, crip technoscience, and design justice, we argue that designing through disability is not simply a methodological stance but a form of epistemic resistance.We show how emotional labor, insider knowledge, and lived specificity can generate design insights that challenge normative assumptions about simplicity, generalizability, and what disabled users should want.Our contributions include: (1) documenting three disabilitycentered design interventions; (2) surfacing cross-cutting themes of agency, emotional labor, and epistemic friction; and (3) offering implications for reframing accessibility research as an inclusive, reflexive, and justice-oriented practice.This report invites the HCI community to recognize lived experience not as anecdotal, but as rigorous situated knowledge essential to equitable design. Atieh Taheri, Misha Sra, Patrick Carrington, Jeffrey P. Bigham |
ASSETS | 4 |
| 2025 | CodeA11y: Making AI Coding Assistants Useful for Accessible Web Development
Peya Mowar, Yi-Hao Peng, Jason Wu 0001, Aaron Steinfeld, Jeffrey P. Bigham |
CHI | 5 |
| 2025 | NoTeeline: Supporting Real-Time, Personalized Notetaking with LLM-Enhanced MicronotesabstractTaking notes quickly while effectively capturing key information can be challenging, especially when watching videos that present simultaneous visual and auditory streams. Manually taken notes often miss crucial details due to the fast-paced nature of the content, while automatically generated notes fail to incorporate user preferences and discourage active engagement with the content. To address this, we propose an interactive system, NoTeeline, for supporting real-time, personalized notetaking. Given micronotes, NoTeeline automatically expands them into full-fledged notes using a Large Language Model (LLM). The generated notes build on the content of micronotes by adding relevant details while maintaining consistency with the user's writing style. In a within-subjects study (n=12), we found that NoTeeline creates high-quality notes that capture the essence of participant micronotes with 93.2% factual correctness and accurately align with participant writing style (8.33% improvement). Using NoTeeline, participants could capture their desired notes with significantly reduced mental effort, writing 47.0% less text and completing their notes in 43.9% less time compared to a manual notetaking baseline. Our results suggest that NoTeeline enables users to integrate LLM assistance in a familiar notetaking workflow while ensuring consistency with their preferences - providing an example of how to address broader challenges in designing AI-assisted tools to augment human capabilities without compromising user autonomy and personalization. Faria Huq, Abdus Samee, David Chuan-En Lin, Alice Xiaodi Tang, Jeffrey P. Bigham |
IUI | 5 |
| 2025 | Position: Towards Bidirectional Human-AI AlignmentabstractRecent advances in general-purpose AI underscore the urgent need to align AI systems with human goals and values. Yet, the lack of a clear, shared understanding of what constitutes "alignment" limits meaningful progress and cross-disciplinary collaboration. In this position paper, we argue that the research community should explicitly define and critically reflect on "alignment" to account for the bidirectional and dynamic relationship between humans and AI. Through a systematic review of over 400 papers spanning HCI, NLP, ML, and more, we examine how alignment is currently defined and operationalized. Building on this analysis, we introduce the Bidirectional Human-AI Alignment framework, which not only incorporates traditional efforts to align AI with human values but also introduces the critical, underexplored dimension of aligning humans with AI – supporting cognitive, behavioral, and societal adaptation to rapidly advancing AI technologies. Our findings reveal significant gaps in current literature, especially in long-term interaction design, human value modeling, and mutual understanding. We conclude with three central challenges and actionable recommendations to guide future research toward more nuanced, reciprocal, and human-AI alignment approaches. Hua Shen 0005, Tiffany Knearem, Reshmi Ghosh, Kenan Alkiek, Kundan Krishna, Yachuan Liu, Savvas Petridis, Yi-Hao Peng, Li Qiwei, Chenglei Si, Yutong Xie 0007, Jeffrey P. Bigham, Frank Bentley, Joyce Y. Chai, Zachary C. Lipton, Qiaozhu Mei, Michael Terry, Diyi Yang, Meredith Ringel Morris, Paul Resnick, David Jurgens |
NeurIPS | 12 |
| 2025 | StepWrite: Adaptive Planning for Speech-Driven Text Generation
Hamza El Alaoui, Atieh Taheri, Yi-Hao Peng, Jeffrey P. Bigham |
UIST | 4 |
| 2025 | Policy Maps: Tools for Guiding the Unbounded Space of LLM BehaviorsabstractFigure 1: Policy maps chart LLM policy coverage over an unbounded space of model behaviors.Here, an AI practitioner is designing a policy for how an LLM should summarize violent text.Policy map abstractions (right) allow the policy designer to interactively author and test policies that govern a model's behavior using if-then rules over concepts.The designer can create any desired concept by providing a simple text definition to capture cases of model behavior.Our Policy Projector tool (center) renders cases, concepts, and policies as visual map layers to aid iterative policy design. Michelle S. Lam, Fred Hohman, Dominik Moritz, Jeffrey P. Bigham, Kenneth Holstein, Mary Beth Kery |
UIST | 4 |
| 2025 | Morae: Proactively Pausing UI Agents for User Choices
Yi-Hao Peng, Dingzeyu Li, Jeffrey P. Bigham, Amy Pavel |
UIST | 3 |
| 2024 | "This really lets us see the entire world: " Designing a conversational telepresence robot for homebound older adultsabstractIn this paper, we explore the design and use of conversational telepresence robots to help homebound older adults interact with the external world. An initial needfinding study (N=8) using video vignettes revealed older adults’ experiential needs for robot-mediated remote experiences such as exploration, reminiscence and social participation. We then designed a prototype system to support these goals and conducted a technology probe study (N=11) to garner a deeper understanding of user preferences for remote experiences. The study revealed user interactive patterns in each desired experience, highlighting the need of robot guidance, social engagements with the robot and the remote bystanders. Our work identifies a novel design space where conversational telepresence robots can be used to foster meaningful interactions in the remote physical environment. We offer design insights into the robot’s proactive role in providing guidance and using dialogue to create personalized, contextualized and meaningful experiences. Yaxin Hu 0002, Laura Stegner, Yasmine Kotturi, Caroline Zhang, Yi-Hao Peng, Faria Huq, Yuhang Zhao 0001, Jeffrey P. Bigham, Bilge Mutlu |
Conference on Designing Interactive Systems | 8 |
| 2024 | Tab to Autocomplete: The Effects of AI Coding Assistants on Web AccessibilityabstractA long-standing challenge in accessible computing has been to get developers to produce the accessible UI code necessary for assistive technologies to work properly. AI coding assistants (e.g., Github Copilot) potentially offer a new opportunity to make UI code more accessible automatically, but it is unclear how their use impacts code accessibility and what developers need to know in order to use them effectively. In this paper, we report on a study where developers untrained in accessibility were tasked with building web UI components with and without an AI coding assistant. Our findings suggest that while current AI coding assistants show potential for creating more accessible UIs, they currently require accessibility awareness and expertise, limiting their expected impact. Peya Mowar, Yi-Hao Peng, Aaron Steinfeld, Jeffrey P. Bigham |
ASSETS | 4 |
| 2024 | Talaria: Interactively Optimizing Machine Learning Models for Efficient InferenceabstractOn-device machine learning (ML) moves computation from the cloud to personal devices, protecting user privacy and enabling intelligent user experiences. However, fitting models on devices with limited resources presents a major technical challenge: practitioners need to optimize models and balance hardware metrics such as model size, latency, and power. To help practitioners create efficient ML models, we designed and developed Talaria : a model visualization and optimization system. Talaria enables practitioners to compile models to hardware, interactively visualize model statistics, and simulate optimizations to test the impact on inference metrics. Since its internal deployment two years ago, we have evaluated Talaria using three methodologies: (1) a log analysis highlighting its growth of 800+ practitioners submitting 3,600+ models; (2) a usability survey with 26 users assessing the utility of 20 Talaria features; and (3) a qualitative interview with the 7 most active users about their experience using Talaria. Fred Hohman, Chaoqun Wang 0002, Jinmook Lee, Jochen Görtler, Dominik Moritz, Jeffrey P. Bigham, Zhile Ren, Cecile Foret, Qi Shan, Xiaoyi Zhang 0006 |
CHI | 6 |
| 2024 | Deconstructing the Veneer of Simplicity: Co-Designing Introductory Generative AI Workshops with Local EntrepreneursabstractGenerative AI platforms and features are permeating many aspects of work. Entrepreneurs from lean economies in particular are well positioned to outsource tasks to generative AI given limited resources. In this paper, we work to address a growing disparity in use of these technologies by building on a four-year partnership with a local entrepreneurial hub dedicated to equity in tech and entrepreneurship. Together, we co-designed an interactive workshops series aimed to onboard local entrepreneurs to generative AI platforms. Alongside four community-driven and iterative workshops with entrepreneurs across five months, we conducted interviews with 15 local entrepreneurs and community providers. We detail the importance of communal and supportive exposure to generative AI tools for local entrepreneurs, scaffolding actionable use (and supporting non-use), demystifying generative AI technologies by emphasizing entrepreneurial power, while simultaneously deconstructing the veneer of simplicity to address the many operational skills needed for successful application. Yasmine Kotturi, Angel Anderson, Glenn Ford, Michael Skirpan, Jeffrey P. Bigham |
CHI | 5 |
| 2024 | COMPA: Using Conversation Context to Achieve Common Ground in AACabstractGroup conversations often shift quickly from topic to topic, leaving a small window of time for participants to contribute. AAC users often miss this window due to the speed asymmetry between using speech and using AAC devices. AAC users may take over a minute longer to contribute, and this speed difference can cause mismatches between the ongoing conversation and the AAC user’s response. This results in misunderstandings and missed opportunities to participate. We present COMPA, an add-on tool for online group conversations that seeks to support conversation partners in achieving common ground. COMPA uses a conversation’s live transcription to enable AAC users to mark conversation segments they intend to address (Context Marking) and generate contextual starter phrases related to the marked conversation segment (Phrase Assistance) and a selected user intent. We study COMPA in 5 different triadic group conversations, each composed by a researcher, an AAC user and a conversation partner (n=10) and share findings on how conversational context supports conversation partners in achieving common ground. Stephanie Valencia, Jessica Huynh, Emma Y. Jiang, Yufei Wu 0020, Teresa Wan, Zixuan Zheng, Henny Admoni, Jeffrey P. Bigham, Amy Pavel |
CHI | 8 |
| 2024 | DreamStruct: Understanding Slides and User Interfaces via Synthetic Data Generation
Yi-Hao Peng, Faria Huq, Yue Jiang 0002, Jason Wu 0001, Xin Yue Li, Jeffrey P. Bigham, Amy Pavel |
ECCV (24) | 6 |
| 2024 | UICoder: Finetuning Large Language Models to Generate User Interface Code through Automated FeedbackabstractJason Wu, Eldon Schoop, Alan Leung, Titus Barik, Jeffrey Bigham, Jeffrey Nichols. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Jason Wu 0001, Eldon Schoop, Alan Leung, Titus Barik, Jeffrey P. Bigham, Jeffrey Nichols 0001 |
NAACL-HLT | 5 |
| 2024 | UIClip: A Data-driven Model for Assessing User Interface DesignabstractUser interface (UI) design is a difficult yet important task for ensuring the usability, accessibility, and aesthetic qualities of applications. In our paper, we develop a machine-learned model, UIClip, for assessing the design quality and visual relevance of a UI given its screenshot and natural language description. To train UIClip, we used a combination of automated crawling, synthetic augmentation, and human ratings to construct a large-scale dataset of UIs, collated by description and ranked by design quality. Through training on the dataset, UIClip implicitly learns properties of good and bad designs by i) assigning a numerical score that represents a UI design’s relevance and quality and ii) providing design suggestions. In an evaluation that compared the outputs of UIClip and other baselines to UIs rated by 12 human designers, we found that UIClip achieved the highest agreement with ground-truth rankings. Finally, we present three example applications that demonstrate how UIClip can facilitate downstream applications that rely on instantaneous assessment of UI design quality: i) UI code generation, ii) UI design tips generation, and iii) quality-aware UI example search. Jason Wu 0001, Yi-Hao Peng, Xin Yue Amanda Li, Amanda Swearngin, Jeffrey P. Bigham, Jeffrey Nichols 0001 |
UIST | 5 |
| 2024 | Towards Automated Accessibility Report Generation for Mobile AppsabstractMany apps have basic accessibility issues, like missing labels or low contrast. To supplement manual testing, automated tools can help developers and QA testers find basic accessibility issues, but they can be laborious to use or require writing dedicated tests. To motivate our work, we interviewed eight accessibility QA professionals at a large technology company. From these interviews, we synthesized three design goals for accessibility report generation systems. Motivated by these goals, we developed a system to generate whole app accessibility reports by combining varied data collection methods (e.g., app crawling, manual recording) with an existing accessibility scanner. Many such scanners are based on single-screen scanning, and a key problem in whole app accessibility reporting is to effectively de-duplicate and summarize issues collected across an app. To this end, we developed a screen grouping model with 96.9% accuracy (88.8% F1-score) and UI element matching heuristics with 97% accuracy (98.2% F1-score). We combine these technologies in a system to report and summarize unique issues across an app, and enable a unique pixel-based ignore feature to help engineers and testers better manage reported issues across their app’s lifetime. We conducted a user study where 19 accessibility engineers and testers used multiple tools to create lists of prioritized issues in the context of an accessibility audit. Our system helped them create lists they were more satisfied with while addressing key limitations of current accessibility scanning tools. Amanda Swearngin, Jason Wu 0001, Xiaoyi Zhang 0006, Esteban Gomez, Jen Coughenour, Rachel Stukenborg, Bhavya Garg, Greg Hughes, Adriana Hilliard, Jeffrey P. Bigham, Jeffrey Nichols 0001 |
ACM Trans. Comput. Hum. Interact. | 10 |
| 2023 | Downstream Datasets Make Surprisingly Good Pretraining CorporaabstractFor most natural language processing tasks, the dominant practice is to finetune large pretrained transformer models (e.g., BERT) using smaller downstream datasets.Despite the success of this approach, it remains unclear to what extent these gains are attributable to the massive background corpora employed for pretraining versus to the pretraining objectives themselves.This paper introduces a large-scale study of self-pretraining, where the same (downstream) training data is used for both pretraining and finetuning.In experiments addressing both ELECTRA and RoBERTa models and 10 distinct downstream classification datasets, we observe that self-pretraining rivals standard pretraining on the BookWiki corpus (despite using around 10×-500× less data), outperforming the latter on 7 and 5 datasets, respectively.Surprisingly, these task-specific pretrained models often perform well on other tasks, including the GLUE benchmark.Self-pretraining also provides benefits on structured output prediction tasks such as question answering and commonsense inference, often providing more than 50% improvements compared to standard pretraining.Our results hint that often performance gains attributable to pretraining are driven primarily by the pretraining objective itself and are not always attributable to the use of external pretraining data in massive amounts.These findings are especially relevant in light of concerns about intellectual property and offensive content in web-scale pretraining data. 1 Kundan Krishna, Jeffrey P. Bigham, Zachary C. Lipton |
ACL (1) | 3 |
| 2023 | From User Perceptions to Technical Improvement: Enabling People Who Stutter to Better Use Speech RecognitionabstractConsumer speech recognition systems do not work as well for many people with speech differences, such as stuttering, relative to the rest of the general population. However, what is not clear is the degree to which these systems do not work, how they can be improved, or how much people want to use them. In this paper, we first address these questions using results from a 61-person survey from people who stutter and find participants want to use speech recognition but are frequently cut off, misunderstood, or speech predictions do not represent intent. In a second study, where 91 people who stutter recorded voice assistant commands and dictation, we quantify how dysfluencies impede performance in a consumer-grade speech recognition system. Through three technical investigations, we demonstrate how many common errors can be prevented, resulting in a system that cuts utterances off 79.1% less often and improves word error rate from 25.4% to 9.9%. Colin Lea, Zifang Huang, Jaya Narain, Lauren Tooley, Dianna Yee, Tien Dung Tran, Panayiotis G. Georgiou, Jeffrey P. Bigham, Leah Findlater |
CHI | 8 |
| 2023 | WebUI: A Dataset for Enhancing Visual UI Understanding with Web SemanticsabstractModeling user interfaces (UIs) from visual information allows systems to make inferences about the functionality and semantics needed to support use cases in accessibility, app automation, and testing. Current datasets for training machine learning models are limited in size due to the costly and time-consuming process of manually collecting and annotating UIs. We crawled the web to construct WebUI, a large dataset of 400,000 rendered web pages associated with automatically extracted metadata. We analyze the composition of WebUI and show that while automatically extracted data is noisy, most examples meet basic criteria for visual UI modeling. We applied several strategies for incorporating semantics found in web pages to increase the performance of visual UI understanding models in the mobile domain, where less labeled data is available: (i) element detection, (ii) screen classification and (iii) screen similarity. Jason Wu 0001, Siyan Wang, Siman Shen, Yi-Hao Peng, Jeffrey Nichols 0001, Jeffrey P. Bigham |
CHI | 6 |
| 2023 | Latent Phrase Matching for Dysarthric Speech
Dianna Yee, Colin Lea, Jaya Narain, Zifang Huang, Lauren Tooley, Jeffrey P. Bigham, Leah Findlater |
INTERSPEECH | 6 |
| 2023 | Never-ending Learning of User InterfacesabstractMachine learning models have been trained to predict semantic information about user interfaces (UIs) to make apps more accessible, easier to test, and to automate. Currently, most models rely on datasets of static screenshots that are labeled by human annotators, a process that is costly and surprisingly error-prone for certain tasks. For example, workers labeling whether a UI element is “tappable” from a screenshot must guess using visual signifiers, and do not have the benefit of tapping on the UI element in the running app and observing the effects. In this paper, we present the Never-ending UI Learner, an app crawler that automatically installs real apps from a mobile app store and crawls them to infer semantic properties of UIs by interacting with UI elements, discovering new and challenging training examples to learn from, and continually updating machine learning models designed to predict these semantics. The Never-ending UI Learner so far has crawled for more than 5,000 device-hours, performing over half a million actions on 6,000 apps to train three computer vision models for i) tappability prediction, ii) draggability prediction, and iii) screen similarity. Jason Wu 0001, Rebecca Krosnick, Eldon Schoop, Amanda Swearngin, Jeffrey P. Bigham, Jeffrey Nichols 0001 |
UIST | 5 |
| 2022 | Tech Help Desk: Support for Local Entrepreneurs Addressing the Long Tail of Computing ChallengesabstractEven entrepreneurs whose businesses are not technological (e.g., handmade goods) need to be able to use a wide range of computing technologies in order to achieve their business goals. In this paper, we follow a participatory action research approach and collaborate with various stakeholders at an entrepreneurial co-working space to design “Tech Help Desk”, an on-going technical service for entrepreneurs. Our model for technical assistance is strategic, in how it is designed to fit the context of local entrepreneurs, and responsive, in how it prioritizes emergent needs. From our engagements with 19 entrepreneurs and support personnel, we reflect on the challenges with existing technology support for non-technological entrepreneurs. Our work highlights the importance of ensuring technological support services can adapt based on entrepreneurs’ ever-evolving priorities, preferences and constraints. Furthermore, we find technological support services should maintain broad technical support for entrepreneurs’ long tail of computing challenges. Yasmine Kotturi, Herman T. Johnson, Michael Skirpan, Sarah E. Fox, Jeffrey P. Bigham, Amy Pavel |
CHI | 5 |
| 2022 | Anticipate and Adjust: Cultivating Access in Human-Centered MethodsabstractMethods are fundamental to doing research and can directly impact who is included in scientific advances. Given accessibility research's increasing popularity and pervasive barriers to conducting and participating in research experienced by people with disabilities, it is critical to ask how methods are made accessible. Yet papers rarely describe their methods in detail. This paper reports on 17 interviews with accessibility experts about how they include both facilitators and participants with disabilities in popular user research methods. Our findings offer strategies for anticipating access needs while remaining flexible and responsive to unexpected access barriers. We emphasize the importance of considering accessibility at all stages of the research process, and contextualize access work in recent disability and accessibility literature. We explore how technology or processes could reflect a norm of accessibility. Finally, we discuss how various needs intersect and conflict and offer a practical structure for planning accessible research. Kelly Mack, Emma McDonnell, Venkatesh Potluri, Maggie Xu, Jailyn Zabala, Jeffrey P. Bigham, Jennifer Mankoff, Cynthia L. Bennett |
CHI | 6 |
| 2022 | InstructDial: Improving Zero and Few-shot Generalization in Dialogue through Instruction TuningabstractInstruction tuning is an emergent paradigm in NLP wherein natural language instructions are leveraged with language models to induce zeroshot performance on unseen tasks.Dialogue is an especially interesting area in which to explore instruction tuning because dialogue systems perform multiple tasks related to language (e.g., natural language understanding and generation, domain-specific interaction), yet instruction tuning has not been systematically explored for dialogue-related tasks.We introduce INSTRUCTDIAL, an instruction tuning framework for dialogue, which consists of a repository of 48 diverse dialogue tasks in a unified text-to-text format created from 59 openly available dialogue datasets.We explore crosstask generalization ability on models tuned on INSTRUCTDIAL across diverse dialogue tasks.Our analysis reveals that INSTRUCTDIAL enables good zero-shot performance on unseen datasets and tasks such as dialogue evaluation and intent detection, and even better performance in a few-shot setting.To ensure that models adhere to instructions, we introduce novel meta-tasks.We establish benchmark zero-shot and few-shot performance of models trained using the proposed framework on multiple dialogue tasks 1 . Prakhar Gupta, Cathy Jiao, Shikib Mehri, Maxine Eskénazi, Jeffrey P. Bigham |
EMNLP | 6 |
| 2022 | Nonverbal Sound Detection for Disordered SpeechabstractVoice assistants have become an essential tool for people with various disabilities because they enable complex phone-or tablet-based interactions without the need for fine-grained motor control, such as with touchscreens. However, these systems are not tuned for the unique characteristics of individuals with speech disorders, including many of those who have a motor-speech disorder, are deaf or hard of hearing, have a severe stutter, or are minimally verbal. We introduce an alternative voice-based input system which relies on sound event detection using fifteen nonverbal mouth sounds like "pop", "click", or "eh." This system was designed to work regardless of ones’ speech abilities and allows full access to existing technology. In this paper, we describe the design of a dataset, model considerations for real-world deployment, and efforts towards model personalization. Our fully-supervised model achieves segment-level precision and recall of 88.6% and 88.4% on an internal dataset of 710 adults, while achieving 0.31 false positives per hour on aggressors such as speech. Five-shot personalization enables satisfactory performance in 84.5% of cases where the generic model fails. Colin Lea, Zifang Huang, Dhruv Jain, Lauren Tooley, Zeinab Liaghat, Shrinath Thelapurath, Leah Findlater, Jeffrey P. Bigham |
ICASSP | 8 |
| 2022 | DialCrowd 2.0: A Quality-Focused Dialog System Crowdsourcing ToolkitabstractDialog system developers need high-quality data to train, fine-tune and assess their systems. They often use crowdsourcing for this since it provides large quantities of data from many workers. However, the data may not be of sufficiently good quality. This can be due to the way that the requester presents a task and how they interact with the workers. This paper introduces DialCrowd 2.0 to help requesters obtain higher quality data by, for example, presenting tasks more clearly and facilitating effective communication with workers. DialCrowd 2.0 guides developers in creating improved Human Intelligence Tasks (HITs) and is directly applicable to the workflows used currently by developers and researchers. Jessica Huynh, Ting-Rui Chiang, Jeffrey P. Bigham, Maxine Eskénazi |
LREC | 3 |
| 2022 | Diffscriber: Describing Visual Design Changes to Support Mixed-Ability Collaborative Presentation AuthoringabstractVisual slide-based presentations are ubiquitous, yet slide authoring tools are largely inaccessible to people who are blind or visually impaired (BVI). When authoring presentations, the 9 BVI presenters in our formative study usually work with sighted collaborators to produce visual slides based on the text content they produce. While BVI presenters valued collaborators’ visual design skill, the collaborators often felt they could not fully review and provide feedback on the visual changes that were made. We present Diffscriber, a system that identifies and describes changes to a slide’s content, layout, and style for presentation authoring. Using our system, BVI presentation authors can efficiently review changes to their presentation by navigating either a summary of high-level changes or individual slide elements. To learn more about changes of interest, presenters can use a generated change hierarchy to navigate to lower-level change details and element styles. BVI presenters using Diffscriber were able to identify slide design changes and provide feedback more easily as compared to using only the slides alone. More broadly, Diffscriber illustrates how advances in detecting and describing visual differences can improve mixed-ability collaboration. Yi-Hao Peng, Jason Wu 0001, Jeffrey P. Bigham, Amy Pavel |
UIST | 3 |
| 2021 | Accessibility and The Crowded Sidewalk: Micromobility's Impact on Public SpaceabstractOver the past several years, micromobility devices—small-scale, networked vehicles used to travel short distances—have begun to pervade cities, bringing promises of sustainable transportation and decreased congestion. Though proponents herald their role in offering lightweight solutions to disconnected transit, smart scooters and autonomous delivery robots increasingly occupy pedestrian pathways, reanimating tensions around the right to public space. Drawing on interviews with disabled activists, government officials, and commercial representatives, we chart how devices and policies co-evolve to fulfill municipal sustainability goals, while creating obstacles for people with disabilities whose activism has long resisted inaccessible infrastructure. We reflect on efforts to redistribute space, institute tech governance, and offer accountability to those who involuntarily encounter interventions on the ground. In studying micromobility within spatial and political context, we call for the HCI community to consider how innovation transforms as it moves out from centers of development toward peripheries of design consideration. Cynthia L. Bennett, Emily E. Ackerman, Bonnie Fan, Jeffrey P. Bigham, Patrick Carrington, Sarah E. Fox |
Conference on Designing Interactive Systems | 4 |
| 2021 | Generating SOAP Notes from Doctor-Patient Conversations Using Modular Summarization TechniquesabstractKundan Krishna, Sopan Khosla, Jeffrey Bigham, Zachary C. Lipton. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Kundan Krishna, Sopan Khosla, Jeffrey P. Bigham, Zachary C. Lipton |
ACL/IJCNLP (1) | 3 |
| 2021 | Slidecho: Flexible Non-Visual Exploration of Presentation VideosabstractWe present Slidecho, a system that enables non-visual access of the slide content in a presentation video on-demand. Slidecho automatically extracts slides and their text and image elements from the presentation video and aligns these elements to the presenter’s speech. When listening to the video, Slidecho provides learners with audio notifications about slide changes and slide elements that are not described by the presenter. The learner can pause the video and browse the entire slide, or only the undescribed slide elements, to gain information. A technical evaluation with presentation videos in-the-wild shows that compared to the presenter’s speech alone, Slidecho provides access to an additional 20% of total text elements and 30% of total image elements that were previously not described. Blind and visually impaired participants in our user study reported that it was easier to locate undescribed slide elements with Slidecho’s synchronized interface than when browsing the video and extracted slides separately, and using Slidecho they read fewer slides that were fully redundant with the speech. Yi-Hao Peng, Jeffrey P. Bigham, Amy Pavel |
ASSETS | 2 |
| 2021 | Aided Nonverbal Communication through Physical Expressive ObjectsabstractAugmentative and alternative communication (AAC) devices enable speech-based communication, but generating speech is not the only resource needed to have a successful conversation. Being able to signal one wishes to take a turn by raising a hand or providing some other cue is critical in securing a turn to speak. Experienced conversation partners know how to recognize the nonverbal communication an augmented communicator (AC) displays, but these same nonverbal gestures can be hard to interpret by people who meet an AC for the first time. Prior work has identified motion-based AAC as a viable and underexplored modality for increasing ACs’ agency in conversation. We build on this prior work to dig deeper into a particular case study on motion-based AAC by co-designing a physical expressive object to support ACs during conversations. We found that our physical expressive object could support communication with unfamiliar partners. As such, we present our process and resulting lessons on the designed object itself and the co-design process. Stephanie Valencia, Mark Steidl, Michael L. Rivera, Cynthia L. Bennett, Jeffrey P. Bigham, Henny Admoni |
ASSETS | 5 |
| 2021 | "It's Complicated": Negotiating Accessibility and (Mis)Representation in Image Descriptions of Race, Gender, and DisabilityabstractContent creators are instructed to write textual descriptions of visual content to make it accessible; yet existing guidelines lack specifics on how to write about people’s appearance, particularly while remaining mindful of consequences of (mis)representation. In this paper, we report on interviews with screen reader users who were also Black, Indigenous, People of Color, Non-binary, and/or Transgender on their current image description practices and preferences, and experiences negotiating theirs and others’ appearances non-visually. We discuss these perspectives, and the ethics of humans and AI describing appearance characteristics that may convey the race, gender, and disabilities of those photographed. In turn, we share considerations for more carefully describing appearance, and contexts in which such information is perceived salient. Finally, we offer tensions and questions for accessibility research to equitably consider politics and ecosystems in which technologies will embed, such as potential risks of human and AI biases amplifying through image descriptions. Cynthia L. Bennett, Cole Gleason, Morgan Klaus Scheuerman, Jeffrey P. Bigham, Anhong Guo, Alexandra To |
CHI | 4 |
| 2021 | Say It All: Feedback for Improving Non-Visual Presentation AccessibilityabstractPresenters commonly use slides as visual aids for informative talks. When presenters fail to verbally describe the content on their slides, blind and visually impaired audience members lose access to necessary content, making the presentation difficult to follow. Our analysis of 90 presentation videos revealed that 72% of 610 visual elements (e.g., images, text) were insufficiently described. To help presenters create accessible presentations, we introduce Presentation A11y, a system that provides real-time and post-presentation accessibility feedback. Our system analyzes visual elements on the slide and the transcript of the verbal presentation to provide element-level feedback on what visual content needs to be further described or even removed. Presenters using our system with their own slide-based presentations described more of the content on their slides, and identified 3.26 times more accessibility problems to fix after the talk than when using a traditional slide-based presentation interface. Integrating accessibility feedback into content creation tools will improve the accessibility of informational content for all. Yi-Hao Peng, JiWoong Jang, Jeffrey P. Bigham, Amy Pavel |
CHI | 3 |
| 2021 | Screen Recognition: Creating Accessibility Metadata for Mobile Applications from PixelsabstractMany accessibility features available on mobile platforms require applications (apps) to provide complete and accurate metadata describing user interface (UI) components. Unfortunately, many apps do not provide sufficient metadata for accessibility features to work as expected. In this paper, we explore inferring accessibility metadata for mobile apps from their pixels, as the visual interfaces often best reflect an app’s full functionality. We trained a robust, fast, memory-efficient, on-device model to detect UI elements using a dataset of 77,637 screens (from 4,068 iPhone apps) that we collected and annotated. To further improve UI detections and add semantic information, we introduced heuristics (e.g., UI grouping and ordering) and additional models (e.g., recognize UI content, state, interactivity). We built Screen Recognition to generate accessibility metadata to augment iOS VoiceOver. In a study with 9 screen reader users, we validated that our approach improves the accessibility of existing mobile apps, enabling even previously inaccessible apps to be used. Xiaoyi Zhang 0006, Lilian de Greef, Amanda Swearngin, Samuel White, Kyle I. Murray, Lisa Yu, Qi Shan, Jeffrey Nichols 0001, Jason Wu 0001, Chris Fleizach, Aaron Everitt, Jeffrey P. Bigham |
CHI | 12 |
| 2021 | Co-designing Socially Assistive Sidekicks for Motion-based AACabstractAugmentative and alternative communication (AAC) devices enable speech-based communication. However, AAC devices do not support nonverbal communication, which allows people to take turns, regulate conversation dynamics, and express intentions. Nonverbal communication requires motion, which is often challenging for AAC users to produce due to motor constraints. In this work, we explore how socially assistive robots, framed as ''sidekicks,'' might provide augmented communicators (ACs) with a nonverbal channel of communication to support their conversational goals. We developed and conducted an accessible co-design workshop that involved two ACs, their caregivers, and three motion experts. We identified goals for conversational support, co-designed prototypes depicting possible sidekick forms, and enacted different sidekick motions and behaviors to achieve speakers' goals. We contribute guidelines for designing sidekicks that support ACs according to three key parameters: attention, precision, and timing. We show how these parameters manifest in appearance and behavior and how they can guide future designs for augmented nonverbal communication. Stephanie Valencia, Michal Luria, Amy Pavel, Jeffrey P. Bigham, Henny Admoni |
HRI | 4 |
| 2021 | SEP-28k: A Dataset for Stuttering Event Detection from Podcasts with People Who StutterabstractThe ability to automatically detect stuttering events in speech could help speech pathologists track an individual’s fluency over time or help improve speech recognition systems for people with atypical speech patterns. Despite increasing interest in this area, existing public datasets are too small to build generalizable dysfluency detection systems and lack sufficient annotations. In this work, we introduce Stuttering Events in Podcasts (SEP-28k), a dataset containing over 28k clips labeled with five event types including blocks, prolongations, sound repetitions, word repetitions, and interjections. Audio comes from public podcasts largely consisting of people who stutter interviewing other people who stutter. We benchmark a set of acoustic models on SEP-28k and the public FluencyBank dataset and highlight how simply increasing the amount of training data improves relative detection performance by 28% and 24% F1 on each. Annotations from over 32k clips across both datasets will be publicly released. Colin Lea, Vikramjit Mitra, Aparna Joshi, Sachin Kajarekar, Jeffrey P. Bigham |
ICASSP | 5 |
| 2021 | Analysis and Tuning of a Voice Assistant System for Dysfluent SpeechabstractDysfluencies and variations in speech pronunciation can severely degrade speech recognition performance, and for many individuals with moderate-to-severe speech disorders, voice operated systems do not work. Current speech recognition systems are trained primarily with data from fluent speakers and as a consequence do not generalize well to speech with dysfluencies such as sound or word repetitions, sound prolongations, or audible blocks. The focus of this work is on quantitative analysis of a consumer speech recognition system on individuals who stutter and production-oriented approaches for improving performance for common voice assistant tasks (i.e., "what is the weather?"). At baseline, this system introduces a significant number of insertion and substitution errors resulting in intended speech Word Error Rates (isWER) that are 13.64\% worse (absolute) for individuals with fluency disorders. We show that by simply tuning the decoding parameters in an existing hybrid speech recognition system one can improve isWER by 24\% (relative) for individuals with fluency disorders. Tuning these parameters translates to 3.6\% better domain recognition and 1.7\% better intent recognition relative to the default setup for the 18 study participants across all stuttering severities. Vikramjit Mitra, Zifang Huang, Colin Lea, Lauren Tooley, Sarah Wu, Darren Botten, Ashwini Palekar, Shrinath Thelapurath, Panayiotis G. Georgiou, Sachin Kajarekar, Jeffrey P. Bigham |
Interspeech | 11 |
| 2021 | Controlling Dialogue Generation with Semantic ExemplarsabstractPrakhar Gupta, Jeffrey Bigham, Yulia Tsvetkov, Amy Pavel. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Prakhar Gupta, Jeffrey P. Bigham, Yulia Tsvetkov, Amy Pavel |
NAACL-HLT | 2 |
| 2021 | Screen Parsing: Towards Reverse Engineering of UI Models from ScreenshotsabstractAutomated understanding of user interfaces (UIs) from their pixels can improve accessibility, enable task automation, and facilitate interface design without relying on developers to comprehensively provide metadata. A first step is to infer what UI elements exist on a screen, but current approaches are limited in how they infer how those elements are semantically grouped into structured interface definitions. In this paper, we motivate the problem of screen parsing, the task of predicting UI elements and their relationships from a screenshot. We describe our implementation of screen parsing and provide an effective training procedure that optimizes its performance. In an evaluation comparing the accuracy of the generated output, we find that our implementation significantly outperforms current systems (up to 23%). Finally, we show three example applications that are facilitated by screen parsing: (i) UI similarity search, (ii) accessibility enhancement, and (iii) code generation from UI screenshots. Jason Wu 0001, Xiaoyi Zhang 0006, Jeffrey Nichols 0001, Jeffrey P. Bigham |
UIST | 4 |
| 2020 | Making GIFs AccessibleabstractSocial media platforms feature short animations known as GIFs, but they are inaccessible to people with vision impairments. Unlike static images, GIFs contain action and visual indications of sound, which can be challenging to describe in alternative text descriptions. We examine a large sample of inaccessible GIFs on Twitter to document how they are used and what visual elements they contain. In interviews with 10 blind Twitter users, we discuss what elements of GIF content should be described and their experiences with GIFs online. The participants compared alternative text descriptions with two other alternative audio formats: (i) the original audio from the GIF source video and (ii) a spoken audio description. We recommend that social media platforms automatically include alt text descriptions for popular GIFs (as Twitter has begun to do), and content producers create audio descriptions to ensure everyone has a rich and emotive experience with GIFs online. Cole Gleason, Amy Pavel, Himalini Gururaj, Kris Makoto Kitani, Jeffrey P. Bigham |
ASSETS | 5 |
| 2020 | Disability and the COVID-19 Pandemic: Using Twitter to Understand Accessibility during Rapid Societal TransitionabstractThe COVID-19 pandemic has forced institutions to rapidly alter their behavior, which typically has disproportionate negative effects on people with disabilities as accessibility is overlooked. To investigate these issues, we analyzed Twitter data to examine accessibility problems surfaced by the crisis. We identified three key domains at the intersection of accessibility and technology: (i) the allocation of product delivery services, (ii) the transition to remote education, and (iii) the dissemination of public health information. We found that essential retailers expanded their high-risk customer shopping hours and pick-up and delivery services, but individuals with disabilities still lacked necessary access to goods and services. Long-experienced access barriers to online education were exacerbated by the abrupt transition of in-person to remote instruction. Finally, public health messaging has been inconsistent and inaccessible, which is unacceptable during a rapidly-evolving crisis. We argue that organizations should create flexible, accessible technology and policies in calm times to be adaptable in times of crisis to serve individuals with diverse needs. Cole Gleason, Stephanie Valencia, Lynn Kirabo, Jason Wu 0001, Anhong Guo, Elizabeth J. Carter, Jeffrey P. Bigham, Cynthia L. Bennett, Amy Pavel |
ASSETS | 7 |
| 2020 | Making Mobile Augmented Reality Applications AccessibleabstractAugmented Reality (AR) technology creates new immersive experiences in entertainment, games, education, retail, and social media. AR content is often primarily visual and it is challenging to enable access to it non-visually due to the mix of virtual and real-world content. In this paper, we identify common constituent tasks in AR by analyzing existing mobile AR applications for iOS, and characterize the design space of tasks that require accessible alternatives. For each of the major task categories, we create prototype accessible alternatives that we evaluate in a study with 10 blind participants to explore their perceptions of accessible AR. Our study demonstrates that these prototypes make AR possible to use for blind users and reveals a number of insights to move forward. We believe our work sets forth not only exemplars for developers to create accessible AR applications, but also a roadmap for future research to make AR comprehensively accessible. Jaylin Herskovitz, Jason Wu 0001, Samuel White, Amy Pavel, Gabriel Reyes, Anhong Guo, Jeffrey P. Bigham |
ASSETS | 7 |
| 2020 | Towards Recommending Accessibility Features on Mobile DevicesabstractNumerous accessibility features have been developed to increase who and how people can access computing devices. Increasingly, these features are included as part of popular platforms, e.g., Apple iOS, Google Android, and Microsoft Windows. Despite their potential to improve the computing experience, many users are unaware of these features and do not know which combination of them could benefit them. In this work, we first quantified this problem by surveying 100 participants online (including 25 older adults) about their knowledge of accessibility and features that they could benefit from, showing very low awareness. We developed four prototypes spanning numerous accessibility categories (e.g., vision, hearing, motor), that embody signals and detection strategies applicable to accessibility recommendation in general. Preliminary results from a study with 20 older adults show that proactive recommendation is a promising approach for better pairing users with accessibility features they could benefit from. Jason Wu 0001, Gabriel Reyes, Sam C. White, Xiaoyi Zhang 0006, Jeffrey P. Bigham |
ASSETS | 5 |
| 2020 | Twitter A11y: A Browser Extension to Make Twitter Images AccessibleabstractSocial media platforms are integral to public and private discourse, but are becoming less accessible to people with vision impairments due to an increase in user-posted images. Some platforms (i.e. Twitter) let users add image descriptions (alternative text), but only 0.1% of images include these. To address this accessibility barrier, we created Twitter A11y, a browser extension to add alternative text on Twitter using six methods. For example, screenshots of text are common, so we detect textual images, and create alternative text using optical character recognition. Twitter A11y also leverages services to automatically generate alternative text or reuse them from across the web. We compare the coverage and quality of Twitter A11y's six alt-text strategies by evaluating the timelines of 50 self-identified blind Twitter users. We find that Twitter A11y increases alt-text coverage from 7.6% to 78.5%, before crowdsourcing descriptions for the remaining images. We estimate that 57.5% of returned descriptions are high-quality. We then report on the experiences of 10 participants with visual impairments using the tool during a week-long deployment. Twitter A11y increases access to social media platforms for people with visual impairments by providing high-quality automatic descriptions for user-posted images. Cole Gleason, Amy Pavel, Emma McCamey, Christina Low, Patrick Carrington, Kris Makoto Kitani, Jeffrey P. Bigham |
CHI | 7 |
| 2020 | Conversational Agency in Augmentative and Alternative CommunicationabstractAugmented communicators (ACs) use augmentative and alternative communication (AAC) technologies to speak. Prior work in AAC research has looked to improve efficiency and expressivity of AAC via device improvements and user training. However, ACs also face constraints in communication beyond their device and individual abilities such as when they can speak, what they can say, and who they can address. In this work, we recast and broaden this prior work using conversational agency as a new frame to study AC communication. We investigate AC conversational agency with a study examining different conversational tasks between four triads of expert ACs, their close conversation partners (paid aide or parent), and a third party (experimenter). We define metrics to analyze AAC conversational agency quantitatively and qualitatively. We conclude with implications for future research to enable ACs to easily exercise conversational agency. Stephanie Valencia, Amy Pavel, Jared Santa Maria, Seunga (Gloria) Yu, Jeffrey P. Bigham, Henny Admoni |
CHI | 5 |
| 2020 | Automated Class Discovery and One-Shot Interactions for Acoustic Activity RecognitionabstractAcoustic activity recognition has emerged as a foundational element for imbuing devices with context-driven capabilities, enabling richer, more assistive, and more accommodating computational experiences. Traditional approaches rely either on custom models trained in situ, or general models pre-trained on preexisting data, with each approach having accuracy and user burden implications. We present Listen Learner, a technique for activity recognition that gradually learns events specific to a deployed environment while minimizing user burden. Specifically, we built an end-to-end system for self-supervised learning of events labelled through one-shot interaction. We describe and quantify system performance 1) on preexisting audio datasets, 2) on real-world datasets we collected, and 3) through user studies which uncovered system behaviors suitable for this new type of interaction. Our results show that our system can accurately and automatically learn acoustic events across environments (e.g., 97% precision, 87% recall), while adhering to users' preferences for non-intrusive interactive behavior. Jason Wu 0001, Chris Harrison 0001, Jeffrey P. Bigham, Gierad Laput |
CHI | 3 |
| 2020 | The Challenges of Crowd Workers in Rural and Urban AmericaabstractCrowd work has the potential of helping the financial recovery of regions traditionally plagued by a lack of economic opportunities, e.g., rural areas. However, we currently have limited information about the challenges facing crowd workers from rural and super rural areas as they struggle to make a living through crowd work sites. This paper examines the challenges and advantages of rural and super rural Amazon Mechanical Turk (MTurk) crowd workers and contrasts them with those of workers from urban areas. Based on a survey of 421 crowd workers from differing geographic regions in the U.S., we identified how across regions, people struggled with being onboarded into crowd work. We uncovered that despite the inequalities and barriers, rural workers tended to be striving more in micro-tasking than their urban counterparts. We also identified cultural traits, relating to time dimension and individualism, that offer us an insight into crowd workers and the necessary qualities for them to succeed on gig platforms. We finish by providing design implications based on our findings to create more inclusive crowd work platforms and tools. Claudia Flores-Saviaga, Benjamin V. Hanrahan, Jeffrey P. Bigham, Saiph Savage |
HCOMP | 4 |
| 2020 | Rescribe: Authoring and Automatically Editing Audio DescriptionsabstractAudio descriptions make videos accessible to those who cannot see them by describing visual content in audio. Producing audio descriptions is challenging due to the synchronous nature of the audio description that must fit into gaps of other video content. An experienced audio description author will produce content that fits narration necessary to understand, enjoy, or experience the video content into the time available. This can be especially tricky for novices to do well. In this paper, we introduce a tool, Rescribe, that helps authors create and refine their audio descriptions. Using Rescribe, authors first create a draft of all the content they would like to include in the audio description. Rescribe then uses a dynamic programming approach to optimize between the length of the audio description, available automatic shortening approaches, and source track lengthening approaches. Authors can iteratively visualize and refine the audio descriptions produced by Rescribe, working in concert with the tool. We evaluate the effectiveness of Rescribe through interviews with blind and visually impaired audio description users who give feedback on Rescribe results. In addition, we invite novice users to create audio descriptions with Rescribe and another tool, finding that users produce audio descriptions with fewer placement errors using Rescribe. Amy Pavel, Gabriel Reyes, Jeffrey P. Bigham |
UIST | 3 |
| 2020 | Becoming the Super Turker: Increasing Wages via a Strategy from High Earning WorkersabstractCrowd markets have traditionally limited workers by not providing transparency information concerning which tasks pay fairly or which requesters are unreliable. Researchers believe that a key reason why crowd workers earn low wages is due to this lack of transparency. As a result, tools have been developed to provide more transparency within crowd markets to help workers. However, while most workers use these tools, they still earn less than minimum wage. We argue that the missing element is guidance on how to use transparency information. In this paper, we explore how novice workers can improve their earnings by following the transparency criteria of Super Turkers, i.e., crowd workers who earn higher salaries on Amazon Mechanical Turk (MTurk). We believe that Super Turkers have developed effective processes for using transparency information. Therefore, by having novices follow a Super Turker criteria (one that is simple and popular among Super Turkers), we can help novices increase their wages. For this purpose, we: (i) conducted a survey and data analysis to computationally identify a simple yet common criteria that Super Turkers use for handling transparency tools; (ii) deployed a two-week field experiment with novices who followed this Super Turker criteria to find better work on MTurk. Novices in our study viewed over 25,000 tasks by 1,394 requesters. We found that novices who utilized this Super Turkers’ criteria earned better wages than other novices. Our results highlight that tool development to support crowd workers should be paired with educational opportunities that teach workers how to effectively use the tools and their related metrics (e.g., transparency values). We finish with design recommendations for empowering crowd workers to earn higher salaries. Saiph Savage, Chun-Wei Chiang, Carlos Toxtli, Jeffrey P. Bigham |
WWW | 5 |
| 2019 | Making Memes AccessibleabstractImages on social media platforms are inaccessible to people with vision impairments due to a lack of descriptions that can be read by screen readers. Providing accurate alternative text for all visual content on social media is not yet feasible, but certain subsets of images, such as internet memes, offer affordances for automatic or semi-automatic generation of alternative text. We present two methods for making memes accessible semi-automatically through (1) the generation of rich alternative text descriptions and (2) the creation of audio macro memes. Meme authors create alternative text templates or audio meme templates, and insert placeholders instead of the meme text. When a meme with the same image is encountered again, it is automatically recognized from a database of meme templates. Text is then extracted and either inserted into the alternative text template or rendered in the audio template using text-to-speech. In our evaluation of meme formats with 10 Twitter users with vision impairments, we found that most users preferred alternative text memes because the description of the visual content conveys the emotional tone of the character. As the preexisting templates can be automatically matched to memes using the same visual image, this combined approach can make a large subset of images on the web accessible, while preserving the emotion and tone inherent in the image memes. Cole Gleason, Amy Pavel, Xingyu Liu 0002, Patrick Carrington, Lydia B. Chilton, Jeffrey P. Bigham |
ASSETS | 6 |
| 2019 | Supporting Older Adults in Using Complex User Interfaces with Augmented RealityabstractUsing complex interfaces has been shown to be challenging for older adults. Existing tutorial systems can be cumbersome, and sometimes difficult to use. To solve this problem, we present a system to support older adults in using visual interfaces by providing step-by-step visual guidance with augmented reality. Using the Apple ARKit platform, our system detects the interface in a phone camera view, and provides visual guidance for users to access the interface following a generated sequence of interactions based on pre-specified tasks and prior knowledge of the interface. Junhan Kong, Anhong Guo, Jeffrey P. Bigham |
ASSETS | 3 |
| 2019 | Twitter A11y: A Browser Extension to Describe ImagesabstractTwitter is integral to many people's lives for news, entertainment, and communication. While people increasingly post images to Twitter, a large majority of images remain inaccessible to people with vision impairments due to a lack of image descriptions (i.e. alternative text). We present Twitter A11y (pronounced ally), a browser extension to make images accessible through a set of strategies tailored to the platform. For example, screenshots of text that exceed the Twitter character limit are common, so we detect textual images, and automatically add alternative text using optical character recognition. Tweet images apart from screenshots and link previews receive descriptions from crowd workers. Based on an evaluation of the timelines of 50 self-identified blind Twitter users, Twitter A11y increases automatic alt text coverage from 2.6% to 25.6%, before crowdsourcing the remaining images. Christina Low, Emma McCamey, Cole Gleason, Patrick Carrington, Jeffrey P. Bigham, Amy Pavel |
ASSETS | 5 |
| 2019 | X-Ray: Screenshot Accessibility via Embedded MetadataabstractScreenshots are frequently shared on social media, via personal communications, and in academic papers. Unfortunately, existing screenshot tools strip away semantics useful for making the content accessible, leaving only pixels. For example, a screenshot of a table removes the structural information useful for conveying it. We quantify the scale of the problem via a study of academic papers, showing that a large number of images included in academic papers are screenshots, and validate this via qualitative interviews with researchers about their figure generation process. We then introduce X-Ray, a system that captures and embeds the semantics of the underlying content into images. Using the X-Ray screenshot tool, semantic information is captured and stored in the Exif data of the resulting image, allowing it to "tag along" as the image is shared and reposted. We demonstrate that our approach retains accessibility for screen reader users via a study with five blind participants. More generally, our approach suggests a method for embedding accessibility metadata into otherwise inaccessible formats, enabling them to retain the more accessible representations that are present at capture time. Sujeath Pareddy, Anhong Guo, Jeffrey P. Bigham |
ASSETS | 3 |
| 2019 | App Usage Predicts Cognitive Ability in Older AdultsabstractWe have limited understanding of how older adults use smartphones, how their usage differs from younger users, and the causes for those differences. As a result, researchers and developers may miss promising opportunities to support older adults or offer solutions to unimportant problems. To characterize smartphone usage among older adults, we collected iPhone usage data from 84 healthy older adults over three months. We find that older adults use fewer apps, take longer to complete tasks, and send fewer messages. We use cognitive test results from these same older adults to then show that up to 79% of these differences can be explained by cognitive decline, and that we can predict cognitive test performance from smartphone usage with 83% ROCAUC. While older adults differ from younger adults in app usage behavior, the "cognitively young" older adults use smartphones much like their younger counterparts. Our study suggests that to better support all older adults, researchers and developers should consider the full spectrum of cognitive function. Mitchell L. Gordon, Leon A. Gatys, Carlos Guestrin, Jeffrey P. Bigham, Andrew Trister, Kayur Patel |
CHI | 4 |
| 2019 | VizWiz-Priv: A Dataset for Recognizing the Presence and Purpose of Private Visual Information in Images Taken by Blind PeopleabstractWe introduce the first visual privacy dataset originating from people who are blind in order to better understand their privacy disclosures and to encourage the development of algorithms that can assist in preventing their unintended disclosures. It includes 8,862 regions showing private content across 5,537 images taken by blind people. Of these, 1,403 are paired with questions and 62\% of those directly ask about the private content. Experiments demonstrate the utility of this data for predicting whether an image shows private information and whether a question asks about the private content in an image. The dataset is publicly-shared at http://vizwiz.org/data/. Danna Gurari, Qing Li 0003, Chi Lin 0001, Anhong Guo, Abigale Stangl, Jeffrey P. Bigham |
CVPR | 7 |
| 2019 | Investigating Evaluation of Open-Domain Dialogue Systems With Human Generated Multiple ReferencesabstractThe aim of this paper is to mitigate the shortcomings of automatic evaluation of open-domain dialog systems through multireference evaluation.Existing metrics have been shown to correlate poorly with human judgement, particularly in open-domain dialog.One alternative is to collect human annotations for evaluation, which can be expensive and time consuming.To demonstrate the effectiveness of multi-reference evaluation, we augment the test set of DailyDialog with multiple references.A series of experiments show that the use of multiple references results in improved correlation between several automatic metrics and human judgement for both the quality and the diversity of system output. Prakhar Gupta, Shikib Mehri, Amy Pavel, Maxine Eskénazi, Jeffrey P. Bigham |
SIGdial | 6 |
| 2019 | StateLens: A Reverse Engineering Solution for Making Existing Dynamic Touchscreens AccessibleabstractBlind people frequently encounter inaccessible dynamic touchscreens in their everyday lives that are difficult, frustrating, and often impossible to use independently. Touchscreens are often the only way to control everything from coffee machines and payment terminals, to subway ticket machines and in-flight entertainment systems. Interacting with dynamic touchscreens is difficult non-visually because the visual user interfaces change, interactions often occur over multiple different screens, and it is easy to accidentally trigger interface actions while exploring the screen. To solve these problems, we introduce StateLens - a three-part reverse engineering solution that makes existing dynamic touchscreens accessible. First, StateLens reverse engineers the underlying state diagrams of existing interfaces using point-of-view videos found online or taken by users using a hybrid crowd-computer vision pipeline. Second, using the state diagrams, StateLens automatically generates conversational agents to guide blind users through specifying the tasks that the interface can perform, allowing the StateLens iOS application to provide interactive guidance and feedback so that blind users can access the interface. Finally, a set of 3D-printed accessories enable blind people to explore capacitive touchscreens without the risk of triggering accidental touches on the interface. Our technical evaluation shows that StateLens can accurately reconstruct interfaces from stationary, hand-held, and web videos; and, a user study of the complete system demonstrates that StateLens successfully enables blind users to access otherwise inaccessible dynamic touchscreens. Anhong Guo, Junhan Kong, Michael L. Rivera, Frank F. Xu, Jeffrey P. Bigham |
UIST | 5 |
| 2019 | "It's almost like they're trying to hide it": How User-Provided Image Descriptions Have Failed to Make Twitter AccessibleabstractTo make images on Twitter and other social media platforms accessible to screen reader users, image descriptions (alternative text) need to be added that describe the information contained within the image. The lack of alternative text has been an enduring accessibility problem since the “alt” attribute was added in HTML 2.0 over 20 years ago, and the rise of user-generated content has only increased the number of images shared. As of 2016, Twitter provides users the ability to turn on a feature that allows descriptions to be added to images in their tweets, presumably in an effort to combat this accessibility problem. What has remained unknown is whether simply enabling users to provide alternative text has an impact on experienced accessibility. In this paper, we present a study of 1.09 million tweets with images, finding that only 0.1% of those tweets included descriptions. In a separate analysis of the timelines of 94 blind Twitter users, we found that these image tweets included descriptions more often. Even users with the feature turned on only write descriptions for about half of the images they tweet. To better understand why users provide alternative text descriptions (or not), we interviewed 20 Twitter users who have written image descriptions. Users did not remember to add alternative text, did not have time to add it, or did not know what to include when writing the descriptions. Our findings indicate that simply making it possible to provide image descriptions is not enough, and reveal future directions for automated tools that may support users in writing high-quality descriptions. Cole Gleason, Patrick Carrington, Cameron Tyler Cassidy, Meredith Ringel Morris, Kris Makoto Kitani, Jeffrey P. Bigham |
WWW | 6 |
| 2019 | TurkScanner: Predicting the Hourly Wage of MicrotasksabstractWorkers in crowd markets struggle to earn a living. One reason for this is that it is difficult for workers to accurately gauge the hourly wages of microtasks, and they consequently end up performing labor with little pay. In general, workers are provided with little information about tasks, and are left to rely on noisy signals, such as textual description of the task or rating of the requester. This study explores various computational methods for predicting the working times (and thus hourly wages) required for tasks based on data collected from other workers completing crowd work. We provide the following contributions. (i) A data collection method for gathering real-world training data on crowd-work tasks and the times required for workers to complete them; (ii) TurkScanner: a machine learning approach that predicts the necessary working time to complete a task (and can thus implicitly provide the expected hourly wage). We collected 9,155 data records using a web browser extension installed by 84 Amazon Mechanical Turk workers, and explored the challenge of accurately recording working times both automatically and by asking workers. TurkScanner was created using ~ 150 derived features, and was able to predict the hourly wages of 69.6% of all the tested microtasks within a 75% error. Directions for future research include observing the effects of tools on people's working practices, adapting this approach to a requester tool for better price setting, and predicting other elements of work (e.g., the acceptance likelihood and worker task preferences.) Chun-Wei Chiang, Saiph Savage, Teppei Nakano, Tetsunori Kobayashi, Jeffrey P. Bigham |
WWW | 6 |
| 2018 | Exploring the Data Tracking and Sharing Preferences of Wheelchair AthletesabstractSports are increasingly data-driven. Athletes use a variety of physical activity monitors to capture their movements, improve performance, and achieve excellence. To understand how wheelchair athletes want to use and share their activity data, we conducted a study using a prototype wheelchair fitness tracking device, which served as a probe to facilitate discussions. We interviewed 15 wheelchair basketball players about the use of performance data in the context of wheelchair basketball, and we discuss several implications for using and sharing automatically-tracked data. We find that the wheelchair basketball community is less concerned about the privacy of their data, and, in contrast to health data, athletes are motivated by competition. We conclude with a set of design opportunities that leverage digitized performance metrics within wheelchair basketball, which could apply to the broader wheelchair and adaptive athletics community. Patrick Carrington, Gierad Laput, Jeffrey P. Bigham |
ASSETS | 3 |
| 2018 | Investigating Cursor-based Interactions to Support Non-Visual Exploration in the Real WorldabstractThe human visual system processes complex scenes to focus attention on relevant items. However, blind people cannot visually skim for an area of interest. Instead, they use a combination of contextual information, knowledge of the spatial layout of their environment, and interactive scanning to find and attend to specific items. In this paper, we define and compare three cursor-based interactions to help blind people attend to items in a complex visual scene: window cursor (move their phone to scan), finger cursor (point their finger to read), and touch cursor (drag their finger on the touchscreen to explore). We conducted a user study with 12 participants to evaluate the three techniques on four tasks, and found that: window cursor worked well for locating objects on large surfaces, finger cursor worked well for accessing control panels, and touch cursor worked well for helping users understand spatial layouts. A combination of multiple techniques will likely be best for supporting a variety of everyday tasks for blind users. Anhong Guo, Saige McVea, Xu Wang 0016, Patrick Clary, Kenneth J. Goldman, Yang Li 0058, Jeffrey P. Bigham |
ASSETS | 8 |
| 2018 | Jellys: Towards a Videogame that Trains Rhythm and Visual Attention for DyslexiaabstractThis demo describes an ongoing research project that aims to develop a video game for the training of two independent cognitive components involved in reading development: visual attention and auditory rhythm. The video game includes two types of gaming activities for each component. First, a proof of concept was carried out with 10 children with dyslexia. The outcome of this proof of concept study served as foundation for the development of a prototype that has been assessed. Human-computer interaction, usability and engagement were measured in a user study with 22 children with dyslexia and 22 without dyslexia. Significant interaction differences between group were not found. Usability and engagement evaluation was positive and will be used to improve the video game. Its efficacy will be tested with a longitudinal training study in developing readers. A video of Jellys user testing is available in https://youtu.be/T9oO9bZFdmM. Mikel Ostiz-Blanco, Marie Lallier, Sergi Grau Carrión, Luz Rello, Jeffrey P. Bigham, Manuel Carreiras |
ASSETS | 5 |
| 2018 | A Data-Driven Analysis of Workers' Earnings on Amazon Mechanical TurkabstractA growing number of people are working as part of on-line crowd work. Crowd work is often thought to be low wage work. However, we know little about the wage distribution in practice and what causes low/high earnings in this setting. We recorded 2,676 workers performing 3.8 million tasks on Amazon Mechanical Turk. Our task-level analysis revealed that workers earned a median hourly wage of only ~$2/h, and only 4% earned more than $7.25/h. While the average requester pays more than $11/h, lower-paying requesters post much more work. Our wage calculations are influenced by how unpaid work is accounted for, e.g., time spent searching for tasks, working on tasks that are rejected, and working on tasks that are ultimately not submitted. We further explore the characteristics of tasks and working patterns that yield higher hourly wages. Our analysis informs platform design and worker tools to create a more positive future for crowd work. Kotaro Hara, Abi Adams, Kristy Milland, Saiph Savage, Chris Callison-Burch, Jeffrey P. Bigham |
CHI | 6 |
| 2018 | Evorus: A Crowd-powered Conversational Assistant Built to Automate Itself Over TimeabstractCrowd-powered conversational assistants have been shown to be more robust than automated systems, but do so at the cost of higher response latency and monetary costs. A promising direction is to combine the two approaches for high quality, low latency, and low cost solutions. In this paper, we introduce Evorus, a crowd-powered conversational assistant built to automate itself over time by (i) allowing new chatbots to be easily integrated to automate more scenarios, (ii) reusing prior crowd answers, and (iii) learning to automatically approve response candidates. Our 5-month-long deployment with 80 participants and 281 conversations shows that Evorus can automate itself without compromising conversation quality. Crowd-AI architectures have long been proposed as a way to reduce cost and latency for crowd-powered systems; Evorus demonstrates how automation can be introduced successfully in a deployed system. Its architecture allows future researchers to make further innovation on the underlying automated components in the context of a deployed open domain dialog system. Ting-Hao 'Kenneth' Huang, Joseph Chee Chang, Jeffrey P. Bigham |
CHI | 3 |
| 2018 | All (of us) can Help: Inclusive crowdfunding research trends and future challengesabstractThis paper presents an overview of the donation based crowdfunding state of the art, establishing a classification scheme to analyze the major platforms, and discussing current research trends and future challenges of crowdfunding as a social inclusion instrument. In many social exclusion situations crowdfunding is the last stronghold to ensure the access to basic commodities, essential to the daily life and well-being of individuals. Despite the commercial success of many crowdfunding platforms, this study shows future research opportunities in the crowdfunding as an inclusion mechanism that can trigger a broader adoption for social causes. The high social impact of the research contributions in this domain can also contribute to make it a hot topic in the upcoming years. Hugo Paredes, João Barroso 0001, Jeffrey P. Bigham |
CSCWD | 3 |
| 2018 | VizWiz Grand Challenge: Answering Visual Questions From Blind PeopleabstractThe study of algorithms to automatically answer visual questions currently is motivated by visual question answering (VQA) datasets constructed in artificial VQA settings. We propose VizWiz, the first goal-oriented VQA dataset arising from a natural VQA setting. VizWiz consists of over 31,000 visual questions originating from blind people who each took a picture using a mobile phone and recorded a spoken question about it, together with 10 crowdsourced answers per visual question. VizWiz differs from the many existing VQA datasets because (1) images are captured by blind photographers and so are often poor quality, (2) questions are spoken and so are more conversational, and (3) often visual questions cannot be answered. Evaluation of modern algorithms for answering visual questions and deciding if a visual question is answerable reveals that VizWiz is a challenging dataset. We introduce this dataset to encourage a larger community to develop more generalized algorithms that can assist blind people. Danna Gurari, Qing Li 0003, Abigale Stangl, Anhong Guo, Chi Lin 0001, Kristen Grauman, Jiebo Luo 0001, Jeffrey P. Bigham |
CVPR | 8 |
| 2018 | Striving to Earn More: A Survey of Work Strategies and Tool Use Among Crowd WorkersabstractEarning money is a primary motivation for workers on Amazon Mechanical Turk, but earning a good wage is difficult because work that pays well is not easily identified and can be time-consuming to find. We explored the strategies that both low- and high-earning workers use to find and complete tasks via a survey of 360 workers. Nearly all workers surveyed had earning money as their primary goal, and workers used many of the same tools (browser extensions and scripts) and strategies in an attempt to earn more money, regardless of earning level. However, high-earning workers used more tools, were more involved in worker communities, and more heavily used batch completion strategies. A natural next step is to use automated systems to assist workers with finding and completing tasks. Workers found this idea interesting, but expressed concerns about impact on the quality of their work and whether using automated tools to support them would violate platform rules. We conclude with ideas for future work in supporting workers to earn more and design considerations for such tools. Toni Kaplan, Kotaro Hara, Jeffrey P. Bigham |
HCOMP | 4 |
| 2017 | Audience Participation Games: Blurring the Line Between Player and SpectatorabstractAudience Participation Games challenge traditional assumptions about gameplay by blurring the line between audience and player, allowing audience members to impact gameplay in a meaningful way. Their recent rise in popularity has created new opportunities for game research and development. To better understand this design space, we developed several versions of two prototype games as design probes. We livestreamed them to an online audience in order to develop a framework for audience motivations and participation styles, to explore ways in which mechanics can affect audience members' sense of agency, and to identify promising design spaces. Our results show the breadth of opportunities and challenges that designers face in creating engaging Audience Participation Games. Joseph Seering, Saiph Savage, Michael Eagle, Joshua Churchin, Rachel Moeller, Jeffrey P. Bigham, Jessica Hammer |
Conference on Designing Interactive Systems | 6 |
| 2017 | On How Deaf People Might Use Speech to Control DevicesabstractSmart devices connected to the Internet are proliferating.To reduce costs of devices that havetraditionally been inexpensive(toasters, microwaves, printers, etc), manyof these devices have chosen to use a speech interface rather than a visual one. This transition has been hastened by the increasing capabilities of speech interfaces,exemplifiedbyproducts likeAmazon Echo and Apple'sSiri.A consequence of these products moving to voice control is that people who are deaf and hard of hearing (DHH) may be unable to use them. In this paper, we briefly introduce two technical approaches we are pursuingfor enabling DHH people to provide input to these devices: (i) human computationworkflows for understanding "deaf speech," and (ii) mobile interfaces that can be instructed to speak on the user's behalf. Jeffrey P. Bigham, Raja S. Kushalnagar, Ting-Hao 'Kenneth' Huang, Juan Pablo Flores, Saiph Savage |
ASSETS | 1 |
| 2017 | The Effects of "Not Knowing What You Don't Know" on Web Accessibility for Blind Web UsersabstractWeb accessibility and usability have been extensively studied for blind web users. The focus has generally been on making it technically possible for blind users to access content, or on helping to make the web more usable. This paper explores a challenge at the intersection of these two lenses, which is the effects of blind web users not knowing what they don't know. On the web, this often means that the user is having a problem completing a task, but does not know whether the problem is because the information is there and not accessible, whether the information is simply difficult to access, or whether the information is not present at all. We first discuss how this issue has manifested itself in other work in this space. We then present the results of a study with 30 sighted web users and 30 blind web users exploring the phenomenon, demonstrating that not knowing the source of a problem causes frustration and wastes time. We conclude with recommendations for future research to help understand and address this problem, as well as design implications for future technology that may assist non-visual web navigation. Jeffrey P. Bigham, Irene Lin, Saiph Savage |
ASSETS | 1 |
| 2017 | Introducing People with ASD to Crowd WorkabstractAdults with Autism Spectrum Disorders (ASD) are unemployed at a high rate, in part because the constraints and expectations of traditional employment can be difficult for them. In this paper, we report on our work in introducing people with ASD to remote work on a crowdsourcing platform and a prototype tool we developed by working with participants. We conducted a six-week long user-centered design study with three participants with ASD. The early stage of the study focused on assessing the abilities of our participants to search and work on micro-tasks available on the crowdsourcing market. Based on our preliminary findings, we designed, developed, and evaluated a prototype tool to facilitate image transcription tasks that are increasingly popular on crowd labor markets. Our findings suggest that people with ASD have varying levels of ability to work on micro-tasks, but are likely to be able to work on tasks like image transcription. The tool we introduce, Assistive Task Queue (ATQ), facilitated our participants' completion of image transcription tasks by removing ambiguity in finding the next task to work on and in simplifying tasks into discrete steps. ATQ may serve as a general platform for finding and delivering appropriate tasks to workers with autism. Kotaro Hara, Jeffrey P. Bigham |
ASSETS | 2 |
| 2017 | Good Background Colors for Readers: A Study of People with and without DyslexiaabstractThe use of colors to enhance the reading of people with dyslexia have been broadly discussed and is often recommended, but evidence of the effectiveness of this approach is lacking. This paper presents a user study with 341 participants (89 with dyslexia) that measures the effect of using background colors on screen readability. Readability was measured via reading time and distance travelled by the mouse. Comprehension was used as a control variable. The results show that using certain background colors have a significant impact on people with and without dyslexia. Warm background colors, Peach, Orange and Yellow, significantly improved reading performance over cool background colors, Blue, Blue Grey and Green. These results provide evidence to the practice of using colored backgrounds to improve readability; people with and without dyslexia benefit, but people with dyslexia may especially benefit from the practice given the difficulty they have in reading in general. Luz Rello, Jeffrey P. Bigham |
ASSETS | 2 |
| 2017 | DytectiveU: A Game to Train the Difficulties and the Strengths of Children with DyslexiaabstractIn this demo we present DytectiveU, a game with 35,000 exercises to train the cognitive abilities related to dyslexia. To personalize the exercises, the game takes into consideration 25 indicators grouped in performance measures, language skills, working memory, executive functions and perceptual processes. The main contribution of this approach is to train dyslexia from a holistic point of view addressing not only the difficulties in reading and writing but also other cognitive abilities that are related to dyslexia and/or contribute to create coping skills to overcome dyslexia. The game is available for Android, iOS and Web (PC/Mac). Luz Rello, Arturo Macias, Mariía Herrera, Camila de Ros, Enrique Romero, Jeffrey P. Bigham |
ASSETS | 6 |
| 2017 | Facade: Auto-generating Tactile Interfaces to AppliancesabstractCommon appliances have shifted toward flat interface panels, making them inaccessible to blind people. Although blind people can label appliances with Braille stickers, doing so generally requires sighted assistance to identify the original functions and apply the labels. We introduce Facade - a crowdsourced fabrication pipeline to help blind people independently make physical interfaces accessible by adding a 3D printed augmentation of tactile buttons overlaying the original panel. Facade users capture a photo of the appliance with a readily available fiducial marker (a dollar bill) for recovering size information. This image is sent to multiple crowd workers, who work in parallel to quickly label and describe elements of the interface. Facade then generates a 3D model for a layer of tactile and pressable buttons that fits over the original controls. Finally, a home 3D printer or commercial service fabricates the layer, which is then aligned and attached to the interface by the blind person. We demonstrate the viability of Facade in a study with 11 blind participants. Anhong Guo, Jeeeun Kim, Xiang 'Anthony' Chen, Tom Yeh, Scott E. Hudson, Jennifer Mankoff, Jeffrey P. Bigham |
CHI | 7 |
| 2017 | Leveraging Complementary Contributions of Different Workers for Efficient Crowdsourcing of Video CaptionsabstractHearing-impaired people and non-native speakers rely on captions for access to video content, yet most videos remain uncaptioned or have machine-generated captions with high error rates. In this paper, we present the design, implementation and evaluation of BandCaption, a system that combines automatic speech recognition with input from crowd workers to provide a cost-efficient captioning solution for accessible online videos. We consider four stakeholder groups as our source of crowd workers: (i) individuals with hearing impairments, (ii) second-language speakers with low proficiency, (iii) second-language speakers with high proficiency, and (iv) native speakers. Each group has different abilities and incentives, which our workflow leverages. Our findings show that BandCaption enables crowd workers who have different needs and strengths to accomplish micro-tasks and make complementary contributions. Based on our results, we outline opportunities for future research and provide design suggestions to deliver cost-efficient captioning solutions. Yun Huang 0003, Na Xue, Jeffrey P. Bigham |
CHI | 4 |
| 2017 | People with Visual Impairment Training Personal Object Recognizers: Feasibility and ChallengesabstractBlind people often need to identify objects around them, from packages of food to items of clothing. Automatic object recognition continues to provide limited assistance in such tasks because models tend to be trained on images taken by sighted people with different background clutter, scale, viewpoints, occlusion, and image quality than in photos taken by blind users. We explore personal object recognizers, where visually impaired people train a mobile application with a few snapshots of objects of interest and provide custom labels. We adopt transfer learning with a deep learning system for user-defined multi-label k-instance classification. Experiments with blind participants demonstrate the feasibility of our approach, which reaches accuracies over 90% for some participants. We analyze user data and feedback to explore effects of sample size, photo-quality variance, and object shape; and contrast models trained on photos by blind participants to those by sighted participants and generic recognizers. Hernisa Kacorri, Kris Makoto Kitani, Jeffrey P. Bigham, Chieko Asakawa |
CHI | 3 |
| 2017 | Subcontracting MicroworkabstractMainstream crowdwork platforms treat microtasks as indivisible units; however, in this article, we propose that there is value in re-examining this assumption. We argue that crowdwork platforms can improve their value proposition for all stakeholders by supporting subcontracting within microtasks. After describing the value proposition of subcontracting, we then define three models for microtask subcontracting: real-time assistance, task management, and task improvement, and reflect on potential use cases and implementation considerations associated with each. Finally, we describe the outcome of two tasks on Mechanical Turk meant to simulate aspects of subcontracting. We reflect on the implications of these findings for the design of future crowd work platforms that effectively harness the potential of subcontracting workflows. Meredith Ringel Morris, Jeffrey P. Bigham, Robin Brewer, Jonathan Bragg, Anand Kulkarni, Saiph Savage |
CHI | 2 |
| 2017 | A 10-Month-Long Deployment Study of On-Demand Recruiting for Low-Latency CrowdsourcingabstractA number of interactive crowd-powered systems have been developed to solve difficult problems out of reach for automated solutions. To work interactively, such systems need access to on-demand labor. To meet this demand, workers can be (i) recruited when needed directly from the crowd marketplace, or (ii) recruited in advance and asked to wait in a retainer pool until they are needed. Most of the evaluations of these systems have been over a short time period, even though we know that marketplaces change and adapt over time. In this paper, we present the results of a 10-month deployment of a crowd-powered system that uses a hybrid approach to fast recruitment of workers that we call Ignition. We describe the Ignition approach and the observed times required to recruit workers from the marketplace and retainer over this long period of time. Our results demonstrate that it is possible to recruit workers with low latency even over long periods, and suggest a number of opportunities for future work for recruitment strategies and modeling that may further improve on-demand recruitment for deployed systems. Ting-Hao 'Kenneth' Huang, Jeffrey P. Bigham |
HCOMP | 2 |
| 2017 | CrowdMask: Using Crowds to Preserve Privacy in Crowd-Powered Systems via Progressive FilteringabstractCrowd-powered systems leverage human intelligence to go beyond the capabilities of automated systems, but also introduce privacy and security concerns because unknown people must view the data that the system processes. While automated approaches cannot robustly filter private information from these datasets, people have the ability to do so if the risk from them viewing the data can be mitigated. We present a crowd-powered approach to masking private content in data by segmenting and distributing smaller segments to crowd workers so that individual workers can identify potentially private content without being able to fully view it themselves. We introduce a novel pyramid workflow for segmentation that uses segments at multiple levels of granularity to overcome problems with fixed-sized approaches. We implement our approach in CrowdMask, a system that allows images with potentially sensitive content to be masked by appearing in progressively larger, more identifiable segments, and masking portions of the image as soon as a risk is identified. Our experiments with 4134 Mechanical Turk workers show that CrowdMask can effectively mask private content from images without revealing sensitive content to constituent workers, while still enabling future systems to use the filtered result. Harmanpreet Kaur, Mitchell L. Gordon, Yiwei Yang 0004, Jeffrey P. Bigham, Jaime Teevan, Ece Kamar, Walter S. Lasecki |
HCOMP | 4 |
| 2017 | WearMail: On-the-Go Access to Information in Your Email with a Privacy-Preserving Human Computation WorkflowabstractEmail is more than just a communication medium. Email serves as an external memory for people---it contains our reservation numbers, meeting details, phone numbers, and more. Often, people need access to this information while on the go, which is cumbersome from mobile devices with limited I/O bandwidth. In this paper, we introduce WearMail, a conversational interface to retrieve specific information in email. WearMail is mostly automated but is made robust to information extraction tasks via a novel privacy-preserving human computation workflow. In WearMail, crowdworkers never have direct access to emails, but rather (i) generate an email filter to help the system find messages that may contain the desired information, and (ii) generate examples of the requested information that are then used to create custom, low-level information extractors that run automatically within the set of filtered emails. We explore the impact of varying levels of obfuscation on result quality, demonstrating that workers are able to deal with highly-obfuscated information nearly as well as with the original. WearMail introduces general mechanisms that let the crowd search and select private data without having direct access to the data itself. Sai Swaminathan, Raymond Fok, Ting-Hao 'Kenneth' Huang, Irene Lin, Rohan Jadvani, Walter S. Lasecki, Jeffrey P. Bigham |
UIST | 8 |
| 2016 | VizMap: Accessible Visual Information Through Crowdsourced Map ReconstructionabstractWhen navigating indoors, blind people are often unaware of key visual information, such as posters, signs, and exit doors. Our VizMap system uses computer vision and crowdsourcing to collect this information and make it available non-visually. VizMap starts with videos taken by on-site sighted volunteers and uses these to create a 3D spatial model. These video frames are semantically labeled by remote crowd workers with key visual information. These semantic labels are located within and embedded into the reconstructed 3D model, forming a query-able spatial representation of the environment. VizMap can then localize the user with a photo from their smartphone, and enable them to explore the visual elements that are nearby. We explore a range of example applications enabled by our reconstructed spatial representation. With VizMap, we move towards integrating the strengths of the end user, on-site crowd, online crowd, and computer vision to solve a long-standing challenge in indoor blind exploration. Cole Gleason, Anhong Guo, Gierad Laput, Kris Makoto Kitani, Jeffrey P. Bigham |
ASSETS | 5 |
| 2016 | Facade: Auto-generating Tactile Interfaces to AppliancesabstractDigital keypads have proliferated on common appliances, from microwaves and refrigerators to printers and remote controls. For blind people, such interfaces are inaccessible. We conducted a formative study with 6 blind people which demonstrated a need for custom designs for tactile labels without dependence on sighted assistance. To address this need, we introduce Facade - a crowdsourced fabrication pipeline to make physical interfaces accessible by adding a 3D printed augmentation of tactile buttons overlaying the original panel. Blind users capture a photo of an inaccessible interface with a standard marker for absolute measurements using perspective transformation. Then this image is sent to multiple crowd workers, who work in parallel to quickly label and describe elements of the interface. These labels are then used to generate 3D models for a layer of tactile and pressable buttons that fits over the original controls. Users can customize the shape and labels of the buttons using a web interface. Finally, a consumer-grade 3D printer fabricates the layer, which is then attached to the interface using adhesives. Such fabricated overlay is an inexpensive ($10) and more general solution to making physical interfaces accessible. Anhong Guo, Jeeeun Kim, Xiang 'Anthony' Chen, Tom Yeh, Scott E. Hudson, Jennifer Mankoff, Jeffrey P. Bigham |
ASSETS | 7 |
| 2016 | "With most of it being pictures now, I rarely use it": Understanding Twitter's Evolving Accessibility to Blind UsersabstractSocial media is an increasingly important part of modern life. We investigate the use of and usability of Twitter by blind users, via a combination of surveys of blind Twitter users, large-scale analysis of tweets from and Twitter profiles of blind and sighted users, and analysis of tweets containing embedded imagery. While Twitter has traditionally been thought of as the most accessible social media platform for blind users, Twitter's increasing integration of image content and users' diverse uses for images have presented emergent accessibility challenges. Our findings illuminate the importance of the ability to use social media for people who are blind, while also highlighting the many challenges such media currently present this user base, including difficulty in creating profiles, in awareness of available features and settings, in controlling revelations of one's disability status, and in dealing with the increasing pervasiveness of image-based content. We propose changes that Twitter and other social platforms should make to promote fuller access to users with visual impairments. Meredith Ringel Morris, Annuska Z. Perkins, Catherine Yao, Sina Bahram, Jeffrey P. Bigham, Shaun K. Kane |
CHI | 5 |
| 2016 | WearWrite: Crowd-Assisted Writing from SmartwatchesabstractThe physical constraints of smartwatches limit the range and complexity of tasks that can be completed. Despite interface improvements on smartwatches, the promise of enabling productive work remains largely unrealized. This paper presents WearWrite, a system that enables users to write documents from their smartwatches by leveraging a crowd to help translate their ideas into text. WearWrite users dictate tasks, respond to questions, and receive notifications of major edits on their watch. Using a dynamic task queue, the crowd receives tasks issued by the watch user and generic tasks from the system. In a week-long study with seven smartwatch users supported by approximately 29 crowd workers each, we validate that it is possible to manage the crowd writing process from a watch. Watch users captured new ideas as they came to mind and managed a crowd during spare moments while going about their daily routine. WearWrite represents a new approach to getting work done from wearables using the crowd. Michael Nebeling, Alexandra To, Anhong Guo, Adrian A. de Freitas, Jaime Teevan, Steven Dow, Jeffrey P. Bigham |
CHI | 7 |
| 2016 | "Is There Anything Else I Can Help You With?" Challenges in Deploying an On-Demand Crowd-Powered Conversational AgentabstractIntelligent conversational assistants, such as Apple's Siri, Microsoft's Cortana, and Amazon's Echo, have quickly become a part of our digital life. However, these assistants have major limitations, which prevents users from conversing with them as they would with human dialog partners. This limits our ability to observe how users really want to interact with the underlying system. To address this problem, we developed a crowd-powered conversational assistant, Chorus, and deployed it to see how users and workers would interact together when mediated by the system. Chorus sophisticatedly converses with end users over time by recruiting workers on demand, which in turn decide what might be the best response for each user sentence. Up to the first month of our deployment, 59 users have held conversations with Chorus during 320 conversational sessions. In this paper, we present an account of Chorus' deployment, with a focus on four challenges: (i) identifying when conversations are over, (ii) malicious users and workers, (iii) on-demand recruiting, and (iv) settings in which consensus is not enough. Our observations could assist the deployment of crowd-powered conversation systems and crowd-powered systems in general. Ting-Hao 'Kenneth' Huang, Walter S. Lasecki, Amos Azaria, Jeffrey P. Bigham |
HCOMP | 4 |
| 2016 | Questimator: Generating Knowledge Assessments for Arbitrary Topics
Qi Guo 0003, Chinmay Kulkarni 0001, Aniket Kittur, Jeffrey P. Bigham, Emma Brunskill |
IJCAI | 4 |
| 2016 | Manipulating Word Lattices to Incorporate Human Corrections
Yashesh Gaur, Florian Metze, Jeffrey P. Bigham |
INTERSPEECH | 3 |
| 2016 | VizLens: A Robust and Interactive Screen Reader for Interfaces in the Real WorldabstractThe world is full of physical interfaces that are inaccessible to blind people, from microwaves and information kiosks to thermostats and checkout terminals. Blind people cannot independently use such devices without at least first learning their layout, and usually only after labeling them with sighted assistance. We introduce VizLens - an accessible mobile application and supporting backend that can robustly and interactively help blind people use nearly any interface they encounter. VizLens users capture a photo of an inaccessible interface and send it to multiple crowd workers, who work in parallel to quickly label and describe elements of the interface to make subsequent computer vision easier. The VizLens application helps users recapture the interface in the field of the camera, and uses computer vision to interactively describe the part of the interface beneath their finger (updating 8 times per second). We show that VizLens provides accurate and usable real-time feedback in a study with 10 blind participants, and our crowdsourcing labeling workflow was fast (8 minutes), accurate (99.7%), and cheap ($1.15). We then explore extensions of VizLens that allow it to (i) adapt to state changes in dynamic interfaces, (ii) combine crowd labeling with OCR technology to handle dynamic displays, and (iii) benefit from head-mounted cameras. VizLens robustly solves a long-standing challenge in accessibility by deeply integrating crowdsourcing and computer vision, and foreshadows a future of increasingly powerful interactive applications that would be currently impossible with either alone. Anhong Guo, Xiang 'Anthony' Chen, Samuel White, Chieko Asakawa, Jeffrey P. Bigham |
UIST | 7 |
| 2015 | What's Hot in Crowdsourcing and Human ComputationabstractThe focus of HCOMP 2014 was the crowd worker. While crowdsourcing is motivated by the promise of leveraging people's intelligence and diverse skillsets in computational processes, the human aspects of this workforce are all too often overlooked. Instead, workers are frequently viewed as interchangeable components that can be statistically managed to eek out reasonable outputs.We are quickly moving past and rejecting these notions, and beginning to understand that it is sometimes the very abstractions that we introduce to make human computation feasible, e.g., abstracting humans behind APIs or isolating workers from others in order to ensure independent input, that can lead to the problems that we then set about trying to solve, e.g., poor or inconsistent quality work. Creating a brighter future for crowd work will require new socio-technical systems that not only decompose tasks, recruit and coordinate workers, and make sense of results, but also find interesting tasks for people to contribute to, structure tasks so that workers learn from them as they go, and eventually automate mundane parts of work. Research in artificial intelligence will be vital for achieving this future. Jeffrey P. Bigham |
AAAI | 1 |
| 2015 | Dytective: Toward a Game to Detect DyslexiaabstractDetecting dyslexia is crucial so that people who have dyslexia can receive training to avoid associated high rates of academic failure. In this paper we present Dytective, a game designed to detect dyslexia. The results of a within-subjects experiment with 40 children (20 with dyslexia) show significant differences between groups who played Dytective. These differences suggest that Dytective could be used to help identify those likely to have dyslexia. Luz Rello, Abdullah X. Ali, Jeffrey P. Bigham |
ASSETS | 3 |
| 2015 | A Spellchecker for DyslexiaabstractPoor spelling is a challenge faced by people with dyslexia throughout their lives. Spellcheckers are therefore a crucial tool for people with dyslexia, but current spellcheckers do not detect real-word errors, which are a common type of errors made by people with dyslexia. Real-word errors are spelling mistakes that result in an unintended but real word, for instance, form instead of from. Nearly 20% of the errors that people with dyslexia make are real-word errors. In this paper, we introduce a system called Real Check that uses a probabilistic language model, a statistical dependency parser and Google n-grams to detect real-world errors. We evaluated Real Check on text written by people with dyslexia, and showed that it detects more of these errors than widely used spellcheckers. In an experiment with 34 people (17 with dyslexia), people with dyslexia corrected sentences more accurately and in less time with Real Check. Luz Rello, Miguel Ballesteros, Jeffrey P. Bigham |
ASSETS | 3 |
| 2015 | Gauging Receptiveness to Social MicrovolunteeringabstractCrowd-powered systems that help people are difficult to scale and sustain because human labor is expensive and worker pools are difficult to grow. To address this problem we introduce the idea of social microvolunteering, a type of intermediated friendsourcing in which a person can provide access to their friends as potential workers for microtasks supporting causes that they care about. We explore this idea by creating Visual Answers, an exemplar social microvolunteering application for Facebook that posts visual questions from people who are blind. We present results of a survey of 350 participants on the concept of social microvolunteering, and a deployment of the Visual Answers application with 91 participants, which collected 618 high-quality answers to questions asked over 12 days, illustrating the feasibility of the approach. Erin L. Brady, Meredith Ringel Morris, Jeffrey P. Bigham |
CHI | 3 |
| 2015 | Zensors: Adaptive, Rapidly Deployable, Human-Intelligent Sensor FeedsabstractThe promise of "smart" homes, workplaces, schools, and other environments has long been championed. Unattractive, however, has been the cost to run wires and install sensors. More critically, raw sensor data tends not to align with the types of questions humans wish to ask, e.g., do I need to restock my pantry? Although techniques like computer vision can answer some of these questions, it requires significant effort to build and train appropriate classifiers. Even then, these systems are often brittle, with limited ability to handle new or unexpected situations, including being repositioned and environmental changes (e.g., lighting, furniture, seasons). We propose Zensors, a new sensing approach that fuses real-time human intelligence from online crowd workers with automatic approaches to provide robust, adaptive, and readily deployable intelligent sensors. With Zensors, users can go from question to live sensor feed in less than 60 seconds. Through our API, Zensors can enable a variety of rich end-user applications and moves us closer to the vision of responsive, intelligent environments. Gierad Laput, Walter S. Lasecki, Jason Wiese, Robert Xiao, Jeffrey P. Bigham, Chris Harrison 0001 |
CHI | 5 |
| 2015 | Exploring Privacy and Accuracy Trade-Offs in Crowdsourced Behavioral Video CodingabstractCoding behavioral video is an important method used by researchers to understand social phenomenon. Unfortunately, traditional hand-coding approaches can take days or weeks of time to complete. Recent work has shown that these tasks can be completed quickly by leveraging the parallelism of large online crowds, but using the crowd introduces new concerns about accuracy, reliability, privacy, and cost. To explore these issues, we conducted interviews with 12 researchers who frequently code behavioral video, to investigate common practices and challenges with video coding. We find accuracy and privacy to be the researchers' primary concerns. To explore this more concretely, we used sample videos to investigate whether crowds can accurately recognize instances of commonly coded behaviors, and show that the crowd yields accurate results. Then, we demonstrate a method for obfuscating participant identity with a video blur filter, and find, as expected, that workers' ability to identify participants decreases as blur level increases. The workers' ability to accurately and reliably code behaviors also decreases, but not as steeply as the identity test. This trade-off between coding quality and privacy protection suggests that researchers can use online crowds to code for some key behaviors in video without compromising participant identity. We conclude with a discussion of how researchers can balance privacy and accuracy on their own data using a system we introduce called Incognito. Walter S. Lasecki, Mitchell L. Gordon, Winnie Leung, Ellen Lim, Jeffrey P. Bigham, Steven Dow |
CHI | 5 |
| 2015 | Apparition: Crowdsourced User Interfaces that Come to Life as You Sketch ThemabstractPrototyping allows designers to quickly iterate and gather feedback, but the time it takes to create even a Wizard-of-Oz prototype reduces the utility of the process. In this paper, we introduce crowdsourcing techniques and tools for prototyping interactive systems in the time it takes to describe the idea. Our Apparition system uses paid microtask crowds to make even hard-to-automate functions work immediately, allowing more fluid prototyping of interfaces that contain interactive elements and complex behaviors. As users sketch their interface and describe it aloud in natural language, crowd workers and sketch recognition algorithms translate the input into user interface elements, add animations, and provide Wizard-of-Oz functionality. We discuss how design teams can use our approach to reflect on prototypes or begin user studies within seconds, and how, over time, Apparition prototypes can become fully-implemented versions of the systems they simulate. Powering Apparition is the first self-coordinated, real-time crowdsourcing infrastructure. We anchor this infrastructure on a new, lightweight write-locking mechanism that workers can use to signal their intentions to each other. Walter S. Lasecki, Juho Kim 0001, Nick Rafter, Onkur Sen, Jeffrey P. Bigham, Michael S. Bernstein |
CHI | 5 |
| 2015 | The Effects of Sequence and Delay on Crowd WorkabstractA common approach in crowdsourcing is to break large tasks into small microtasks so that they can be parallelized across many crowd workers and so that redundant work can be more easily compared for quality control. In practice, this can result in the microtasks being presented out of their natural order and often introduces delays between individual microtasks. In this paper, we demonstrate in a study of 338 crowd workers that non-sequential microtasks and the introduction of delays significantly decreases worker performance. We show that interruptions where a large delay occurs between two related tasks can cause up to a 102% slowdown in completion time, and interruptions where workers are asked to perform different tasks in sequence can slow down completion time by 57%. We conclude with a set of design guidelines to improve both worker performance and realized pay, and instructions for implementing these changes in existing interfaces for crowd work. Walter S. Lasecki, Jeffrey M. Rzeszotarski, Adam Marcus 0002, Jeffrey P. Bigham |
CHI | 4 |
| 2015 | RegionSpeak: Quick Comprehensive Spatial Descriptions of Complex Images for Blind UsersabstractBlind people often seek answers to their visual questions from remote sources, however, the commonly adopted single-image, single-response model does not always guarantee enough bandwidth between users and sources. This is especially true when questions concern large sets of information, or spatial layout, e.g., where is there to sit in this area, what tools are on this work bench, or what do the buttons on this machine do? Our RegionSpeak system addresses this problem by providing an accessible way for blind users to (i) combine visual information across multiple photographs via image stitching, em (ii) quickly collect labels from the crowd for all relevant objects contained within the resulting large visual area in parallel, and (iii) then interactively explore the spatial layout of the objects that were labeled. The regions and descriptions are displayed on an accessible touchscreen interface, which allow blind users to interactively explore their spatial layout. We demonstrate that workers from Amazon Mechanical Turk are able to quickly and accurately identify relevant regions, and that asking them to describe only one region at a time results in more comprehensive descriptions of complex images. RegionSpeak can be used to explore the spatial layout of the regions identified. It also demonstrates broad potential for helping blind users to answer difficult spatial layout questions. Walter S. Lasecki, Erin L. Brady, Jeffrey P. Bigham |
CHI | 4 |
| 2015 | Accessible Crowdwork?: Understanding the Value in and Challenge of Microtask Employment for People with DisabilitiesabstractWe present the first formal study of crowdworkers who have disabilities via in-depth open-ended interviews of 17 people (disabled crowdworkers and job coaches for people with disabilities) and a survey of 631 adults with disabilities. Our findings establish that people with a variety of disabilities currently participate in the crowd labor marketplace, despite challenges such as crowdsourcing workflow designs that inadvertently prohibit participation by, and may negatively affect the worker reputations of, people with disabilities. Despite such challenges, we find that crowdwork potentially offers different opportunities for people with disabilities relative to the normative office environment, such as job flexibility and lack of a need to rely on public transit. We close by identifying several ways in which crowd labor platform operators and/or individual task requestors could improve the accessibility of this increasingly important form of employment. Kathryn Zyskowski, Meredith Ringel Morris, Jeffrey P. Bigham, Mary L. Gray, Shaun K. Kane |
CSCW | 3 |
| 2015 | Guardian: A Crowd-Powered Spoken Dialog System for Web APIsabstractNatural language dialog is an important and intuitive way for people to access information and services. However, current dialog systems are limited in scope, brittle to the richness of natural language, and expensive to produce. This paper introduces Guardian, a crowd-powered framework that wraps existing Web APIs into immediately usable spoken dialog systems. Guardian takes as input the Web API and desired task, and the crowd determines the parameters necessary to complete it, how to ask for them, and interprets the responses from the API. The system is structured so that, over time, it can learn to take over for the crowd. This hybrid systems approach will help make dialog systems both more general and more robust going forward. Ting-Hao 'Kenneth' Huang, Walter S. Lasecki, Jeffrey P. Bigham |
HCOMP | 3 |
| 2015 | Using keyword spotting to help humans correct captioning fasterabstractAutomatic real-time captioning provides immediate and on de-mand access to spoken content in lectures or talks, and is a cru-cial accommodation for deaf and hard of hearing (DHH) people. However, in the presence of specialized content, like in techni-cal talks, automatic speech recognition (ASR) still makes mis-takes which may render the output incomprehensible. In this paper, we introduce a new approach, which allows audience or crowd workers, to quickly correct errors that they spot in ASR output. Prior approaches required the crowd worker to manu-ally “edit ” the ASR hypothesis by selecting and replacing the text, which is not suitable for real-time scenarios. Our approach is faster and allows the worker to simply type corrections for misrecognized words as soon as he or she spots them. The sys-tem then finds the most likely position for the correction in the ASR output using keyword search (KWS) and stitches the word into the ASR output. Our work demonstrates the potential of computation to incorporate human input quickly enough to be usable in real-time scenarios, and may be a better method for providing this vital accommodation to DHH people. Index Terms: speech recognition, human-computer interac-tion, spoken term detection, real-time crowd sourcing. Yashesh Gaur, Florian Metze, Yajie Miao, Jeffrey P. Bigham |
INTERSPEECH | 4 |
| 2014 | How companies engage customers around accessibility on social mediaabstractSocial media offers a targeted way for mainstream technology companies to communicate with people with disabilities about the accessibility problems that they face. While companies have started to engage with users on social media about accessibility, they differ greatly in terms of their approach and how well they support the ways in which their users want to engage. In this paper, we describe current use patterns of six corporate accessibility teams and their users on Twitter, and present an analysis of these interactions. We find that while many users want to interact directly with companies about accessibility, companies prefer to redirect them to other channels and use Twitter for broadcast messages promoting their accessibility work instead. Our analysis demonstrates that users want to use social media to become part of the process of improving accessibility of mainstream technology, and suggests the extent to which a company is able to leverage this input depends greatly on how they choose to present themselves and interact on social media. Erin L. Brady, Jeffrey P. Bigham |
ASSETS | 2 |
| 2014 | Legion scribe: real-time captioning by non-expertsabstractThe promise of affordable, automatic approaches to real-time captioning imagines a future in which deaf and hard of hearing (DHH) users have immediate access to speech in the world around them my simply picking up their phone or other mobile device. While the challenges of processing highly variable natural language has prevented automated approaches from completing this task reliably enough for use in settings such as classrooms or workplaces [4], recent work in crowd-powered approaches have allowed groups of non-expert captionists to provide a similarly-flexible source of captions for DHH users. This is in contrast to current human-powered approaches, which use highly-trained professional captionists who can type up to 250 words per minute (WPM), but also can cost over $100/hr. In this paper, we describe a real-time demo of Legion:Scribe (or just "Scribe"), a crowd-powered captioning system that allows untrained participants and volunteers to provide reliable captions with less than 5 seconds of latency by computationally merging their input into a single collective answer that is more accurate and more complete than any one worker could have generated alone. Walter S. Lasecki, Raja S. Kushalnagar, Jeffrey P. Bigham |
ASSETS | 3 |
| 2014 | Increasing the bandwidth of crowdsourced visual question answering to better support blind usersabstractMany of the visual questions that blind people ask cannot be easily answered with a single image or a short response, especially when questions are of an exploratory nature, e.g. what is in this area, or what tools are available on this work bench? We introduce RegionSpeak to allow blind users to capture large areas of visual information, identify all of the objects within them, and explore their spatial layout with fewer interactions. RegionSpeak helps blind users capture all of the relevant visual information using an interface designed to support stitching multiple images together. We use a parallel crowdsourcing workflow that asks workers to define and describe regions of interest, allowing even complex images to be described quickly. The regions and descriptions are displayed on an auditory touchscreen interface, allowing users to know what is in a scene and how it is laid out. Walter S. Lasecki, Jeffrey P. Bigham |
ASSETS | 3 |
| 2014 | Crowd storage: storing information on existing memoriesabstractThis paper introduces the concept of crowd storage, the idea that digital files can be stored and retrieved later from the memories of people in the crowd. Similar to human memory, crowd storage is ephemeral, which means that storage is temporary and the quality of the stored information degrades over time. Crowd storage may be preferred over storing information directly in the cloud, or when it is desirable for information to degrade inline with normal human memories. To explore and validate this idea, we created WeStore, a system that stores and then later retrieves digital files in the existing memories of crowd workers. WeStore does not store information directly, but rather encrypts the files using details of the existing memories elicited from individuals within the crowd as cryptographic keys. The fidelity of the retrieved information is tied to how well the crowd remembers the details of the memories they provided. We demonstrate that crowd storage is feasible using an existing crowd marketplace (Amazon Mechanical Turk), explore design considerations important for building systems that use crowd storage, and outline ideas for future research in this area. Jeffrey P. Bigham, Walter S. Lasecki |
CHI | 1 |
| 2014 | Finding dependencies between actions using the crowdabstractActivity recognition can provide computers with the context underlying user inputs, enabling more relevant responses and more fluid interaction. However, training these systems is difficult because it requires observing every possible sequence of actions that comprise a given activity. Prior work has enabled the crowd to provide labels in real-time to train automated systems on-the-fly, but numerous examples are still needed before the system can recognize an activity on its own. To reduce the need to collect this data by observing users, we introduce ARchitect, a system that uses the crowd to capture the dependency structure of the actions that make up activities. Our tests show that over seven times as many examples can be collected using our approach versus relying on direct observation alone, demonstrating that by leveraging the understanding of the crowd, it is possible to more easily train automated systems. Walter S. Lasecki, Leon Weingard, George Ferguson, Jeffrey P. Bigham |
CHI | 4 |
| 2014 | Introducing shared character control to existing video games
Anna Loparev, Walter S. Lasecki, Kyle I. Murray, Jeffrey P. Bigham |
FDG | 4 |
| 2014 | Friendsourcing for the Greater Good: Perceptions of Social MicrovolunteeringabstractPeople with disabilities can be reluctant to friendsource help from their own friends for fear of appearing dependent or annoying. Our social microvolunteering approach has volunteers post friendsourcing tasks on behalf of people with disabilities. We demonstrate this approach via a Facebook application that answers visual questions on behalf of blind users. Erin L. Brady, Meredith Ringel Morris, Jeffrey P. Bigham |
HCOMP | 3 |
| 2014 | Glance Privacy: Obfuscating Personal Identity While Coding Behavioral VideoabstractBehavioral researchers code video to extract systematic meaning from subtle human actions and emotions. While this has traditionally been done by analysts within a research group, recent methods have leveraged online crowds to massively parallelize this task and reduce the time required from days to seconds. However, using the crowd to code video increases the risk that private information will be disclosed because workers who have not been vetted will view the video data in order to code it. In this Work-in-Progress, we discuss techniques for maintaining privacy when using Glance to code video and present initial experimental evidence to support them. Mitchell L. Gordon, Walter S. Lasecki, Winnie Leung, Ellen Lim, Steven Dow, Jeffrey P. Bigham |
HCOMP | 6 |
| 2014 | Combining Non-Expert and Expert Crowd Work to Convert Web APIs to Dialog SystemsabstractThousands of web APIs expose data and services that would be useful to access with natural dialog, from weather and sports to Twitter and movies. The process of adapting each API to a robust dialog system is difficult and time-consuming, as it requires not only programming but also anticipating what is mostly likely to be asked and how it is likely to be asked. We present a crowd-powered system able to generate a natural languageinterface for arbitrary web APIs from scratch without domain-dependent training data or knowledge.Our approach combines two types of crowd workers: non-expert Mechanical Turk workers interpret the functions of the API and elicit information from the user, and expert oDesk workers provide a minimal sufficient scaffolding around the API to allow us to make general queries.We describe our multi-stage process and present results for each stage. Ting-Hao 'Kenneth' Huang, Walter S. Lasecki, Alan L. Ritter, Jeffrey P. Bigham |
HCOMP | 4 |
| 2014 | Tuning the Diversity of Open-Ended Responses From the CrowdabstractCrowdsourcing can solve problems beyond the reach of state-of-the-art fully automated systems. A common pattern found in many such systems is for the workers to discover, in parallel, a number of candidate solutions and then vote on the best one to pass forward, often within a fixed amount of time. We present the propose-vote-abstain mechanism for eliciting from crowd workers the proper balance between solution discovery and selection. Each crowd worker is given a choice among proposing an answer, voting among the answers proposed so far, or abstaining, i.e., doing nothing. When a stopping condition is reached, the mechanism returns the answer with the most votes. Workers are paid a base amount, with bonuses if they propose or vote for the winning answer. Walter S. Lasecki, Christopher Homan, Jeffrey P. Bigham |
HCOMP | 3 |
| 2014 | Low Effort Crowdsourcing: Leveraging Peripheral Attention for Crowd WorkabstractCrowdsourcing systems leverage short bursts of focused attention from many contributors to achieve a goal. By requiring people’s full attention, existing crowdsourcing systems fail to leverage people’s cognitive surplus in the many settings for which they may be distracted, performing or waiting to perform another task, or barely paying attention. In this paper, we study opportunities for low-effort crowdsourcing that enable people to contribute to problem solving in such settings. We discuss the design space for low-effort crowdsourcing, and through a series of prototypes, demonstrate interaction techniques, mechanisms, and emerging principles for enabling low-effort crowdsourcing. Rajan Vaish, Peter Organisciak, Kotaro Hara, Jeffrey P. Bigham |
HCOMP | 4 |
| 2014 | Tracking @stemxcomet: teaching programming to blind students via 3D printing, crisis management, and twitterabstractIntroductory programming activities for students often include graphical user interfaces or other visual media that are inaccessible to students with visual impairments. Digital fabrication techniques such as 3D printing offer an opportunity for students to write programs that produce tactile objects, providing an accessible way of exploring program output. This paper describes the planning and execution of a four-day computer science education workshop in which blind and visually impaired students wrote Ruby programs to analyze data from Twitter regarding a fictional ecological crisis. Students then wrote code to produce accessible tactile visualizations of that data. This paper describes outcomes from our workshop and suggests future directions for integrating data analysis and 3D printing into programming instruction for blind students. Shaun K. Kane, Jeffrey P. Bigham |
SIGCSE | 2 |
| 2014 | Making the web easier to see with opportunistic accessibility improvementabstractMany people would find the Web easier to use if content was a little bigger, even those who already find the Web possible to use now. This paper introduces the idea of opportunistic accessibility improvement in which improvements intended to make a web page easier to access, such as magnification, are automatically applied to the extent that they can be without causing negative side effects. We explore this idea with oppaccess.js, an easily-deployed system for magnifying web pages that iteratively increases magnification until it notices negative side effects, such as horizontal scrolling or overlapping text. We validate this approach by magnifying existing web pages 1.6x on average without introducing negative side effects. We believe this concept applies generally across a wide range of accessibility improvements designed to help people with diverse abilities. Jeffrey P. Bigham |
UIST | 1 |
| 2014 | Glance: rapidly coding behavioral video with the crowdabstractBehavioral researchers spend considerable amount of time coding video data to systematically extract meaning from subtle human actions and emotions. In this paper, we present Glance, a tool that allows researchers to rapidly query, sample, and analyze large video datasets for behavioral events that are hard to detect automatically. Glance takes advantage of the parallelism available in paid online crowds to interpret natural language queries and then aggregates responses in a summary view of the video data. Glance provides analysts with rapid responses when initially exploring a dataset, and reliable codings when refining an analysis. Our experiments show that Glance can code nearly 50 minutes of video in 5 minutes by recruiting over 60 workers simultaneously, and can get initial feedback to analysts in under 10 seconds for most clips. We present and compare new methods for accurately aggregating the input of multiple workers marking the spans of events in video data, and for measuring the quality of their coding in real-time before a baseline is established by measuring the variance between workers. Glance's rapid responses to natural language queries, feedback regarding question ambiguity and anomalies in the data, and ability to build on prior context in followup queries allow users to have a conversation-like interaction with their data - opening up new possibilities for naturally exploring video data. Walter S. Lasecki, Mitchell L. Gordon, Danai Koutra, Malte F. Jung, Steven Dow, Jeffrey P. Bigham |
UIST | 6 |
| 2013 | Crowd Formalization of Action ConditionsabstractTraining intelligent systems is a time consuming and costly process that often limits their application to real-world problems. Prior work in crowdsourcing has attempted to compensate for this challenge by generating sets of labeled training data for machine learning algorithms. In this work, we seek to move beyond collecting just statistical data and explore how to gather structured, relational representations of a scenario using the crowd. We focus on activity recognition because of its broad applicability, high level of variation between individual instances, and difficulty of training systems a priori. We present ARchitect, a system that uses the crowd to ascertain pre and post conditions for actions observed in a video and find relations between actions. Our ultimate goal is to identify multiple valid execution paths from a single set of observations, which suggests one-off learning from the crowd is possible. Walter S. Lasecki, Leon Weingard, Jeffrey P. Bigham, George Ferguson |
AAAI | 3 |
| 2013 | Real-time captioning by non-experts with legion scribeabstractReal-time captioning provides people who are deaf or hard of hearing access to speech in settings such as classrooms and live events. The most reliable approach to provide these captions is to recruit an expert stenographer who is able to type at natural speaking rates, but they charge more than $100 USD per hour and must be scheduled in advance. We introduce Legion Scribe (Scribe), a system that allows 3-5 ordinary people who can hear and type to jointly caption speech in real-time. Each person is unable to type at natural speaking rates, and so is asked only to type part of what they hear. Scribe automatically stitches all of the partial captions together to form a complete caption stream. We have shown that the accuracy of Scribe captions approaches that of a professional stenographer, while its latency and cost is dramatically lower. Walter S. Lasecki, Christopher D. Miller, Raja S. Kushalnagar, Jeffrey P. Bigham |
ASSETS | 4 |
| 2013 | Answering visual questions with conversational crowd assistantsabstractBlind people face a range of accessibility challenges in their everyday lives, from reading the text on a package of food to traveling independently in a new place. Answering general questions about one's visual surroundings remains well beyond the capabilities of fully automated systems, but recent systems are showing the potential of engaging on-demand human workers (the crowd) to answer visual questions. The input to such systems has generally been a single image, which can limit the interaction with a worker to one question; or video streams where systems have paired the end user with a single worker, limiting the benefits of the crowd. In this paper, we introduce Chorus:View, a system that assists users over the course of longer interactions by engaging workers in a continuous conversation with the user about a video stream from the user's mobile device. We demonstrate the benefit of using multiple crowd workers instead of just one in terms of both latency and accuracy, then conduct a study with 10 blind users that shows Chorus:View answers common visual questions more quickly and accurately than existing approaches. We conclude with a discussion of users' feedback and potential future work on interactive crowd support of blind users. Walter S. Lasecki, Phyo Thiha, Erin L. Brady, Jeffrey P. Bigham |
ASSETS | 5 |
| 2013 | Real time object scanning using a mobile phone and cloud-based visual search engineabstractComputer vision and human-powered services can provide blind people access to visual information in the world around them, but their efficacy is dependent on high-quality photo inputs. Blind people often have difficulty capturing the information necessary for these applications to work because they cannot see what they are taking a picture of. In this paper, we present Scan Search, a mobile application that offers a new way for blind people to take high-quality photos to support recognition tasks. To support realtime scanning of objects, we developed a key frame extraction algorithm that automatically retrieves high-quality frames from continuous camera video stream of mobile phones. Those key frames are streamed to a cloud-based recognition engine that identifies the most significant object inside the picture. This way, blind users can scan for objects of interest and hear potential results in real time. We also present a study exploring the tradeoffs in how many photos are sent, and conduct a user study with 8 blind participants that compares Scan Search with a standard photo-snapping interface. Our results show that Scan Search allows users to capture objects of interest more efficiently and is preferred by users to the standard interface. Pierre J. Garrigues, Jeffrey P. Bigham |
ASSETS | 3 |
| 2013 | Visual challenges in the everyday lives of blind peopleabstractThe challenges faced by blind people in their everyday lives are not well understood. In this paper, we report on the findings of a large-scale study of the visual questions that blind people would like to have answered. As part of this year-long study, 5,329 blind users asked 40,748 questions about photographs that they took from their iPhones using an application called VizWiz Social. We present a taxonomy of the types of questions asked, report on a number of features of the questions and accompanying photographs, and discuss how individuals changed how they used VizWiz Social over time. These results improve our understanding of the problems blind people face, and may help motivate new projects more accurately targeted to help blind people live more independently in their everyday lives. Erin L. Brady, Meredith Ringel Morris, Samuel White, Jeffrey P. Bigham |
CHI | 5 |
| 2013 | Warping time for more effective real-time crowdsourcingabstractIn this paper, we introduce the idea of "warping time" to improve crowd performance on the difficult task of captioning speech in real-time. Prior work has shown that the crowd can collectively caption speech in real-time by merging the partial results of multiple workers. Because non-expert workers cannot keep up with natural speaking rates, the task is frustrating and prone to errors as workers buffer what they hear to type later. The TimeWarp approach automatically increases and decreases the speed of speech playback systematically across individual workers who caption only the periods played at reduced speed. Studies with 139 remote crowd workers and 24 local participants show that this approach improves median coverage (14.8%), precision (11.2%), and per-word latency (19.1%). Warping time may also help crowds outperform individuals on other difficult real-time performance tasks. Walter S. Lasecki, Christopher D. Miller, Jeffrey P. Bigham |
CHI | 3 |
| 2013 | Investigating the appropriateness of social network question asking as a resource for blind usersabstractRecent work has shown the potential of having remote humans answer visual questions that blind users have. On the surface social networking sites (SNSs) offer an attractive free source of human-powered answers that can be personalized to the user. In this paper, we explore the potential of blind users asking visual questions to their social networks. We present the first formal study of how blind people use social networking sites via a survey of 191 blind adults. We also explore whether blind users find SNSs an appropriate venue for Q&A through a log analysis of questions asked using VizWiz Social, an iPhone app with over 5,000 users, which lets blind users ask questions to either the crowd or friends. We then report findings of a field experiment with 23 blind VizWiz Social users, which explored question asking on VizWiz Social in the presence of monetary costs for non-social sources. We find that blind people have a large presence on social networking sites, but do not see them as an appropriate venue for asking questions due to high perceived social costs. Erin L. Brady, Meredith Ringel Morris, Jeffrey P. Bigham |
CSCW | 4 |
| 2013 | Real-time crowd labeling for deployable activity recognitionabstractSystems that automatically recognize human activities offer the potential of timely, task-relevant information and support. For example, prompting systems can help keep people with cognitive disabilities on track and surveillance systems can warn of activities of concern. Current automatic systems are difficult to deploy because they cannot identify novel activities, and, instead, must be trained in advance to recognize important activities. Identifying and labeling these events is time consuming and thus not suitable for real-time support of already-deployed activity recognition systems. In this paper, we introduce Legion:AR, a system that provides robust, deployable activity recognition by supplementing existing recognition systems with on-demand, real-time activity identification using input from the crowd. Walter S. Lasecki, Young Chol Song, Henry A. Kautz, Jeffrey P. Bigham |
CSCW | 4 |
| 2013 | Finding action dependencies using the crowdabstractTraining intelligent systems is a time-consuming and costly process that often limits real-world applications. Prior work has attempted to compensate for this challenge by generating sets of labeled training data for machine learning algorithms using affordable human contributors. In this paper, we present ARchitect, a system that uses the crowd to extract context-dependent relational structure. We focus on activity recognition because of its broad applicability, high level of variation, and difficulty of training systems a priority. We demonstrate that using our approach, the crowd can accurately and consistently identify relationships between actions even over sessions containing different workers and varied executions of an activity. This results in the ability to identify multiple valid execution paths from a single observation, suggesting that one-off learning can be facilitated by using the crowd as an on-demand source of human intelligence in the knowledge acquisition process. Walter S. Lasecki, Leon Weingard, George Ferguson, Jeffrey P. Bigham |
K-CAP | 4 |
| 2013 | Text Alignment for Real-Time Crowd Captioning
Iftekhar Naim, Daniel Gildea, Walter S. Lasecki, Jeffrey P. Bigham |
HLT-NAACL | 4 |
| 2013 | Chorus: a crowd-powered conversational assistantabstractDespite decades of research attempting to establish conversational interaction between humans and computers, the capabilities of automated conversational systems are still limited. In this paper, we introduce Chorus, a crowd-powered conversational assistant. When using Chorus, end users converse continuously with what appears to be a single conversational partner. Behind the scenes, Chorus leverages multiple crowd workers to propose and vote on responses. A shared memory space helps the dynamic crowd workforce maintain consistency, and a game-theoretic incentive mechanism helps to balance their efforts between proposing and voting. Studies with 12 end users and 100 crowd workers demonstrate that Chorus can provide accurate, topical responses, answering nearly 93% of user queries appropriately, and staying on-topic in over 95% of responses. We also observed that Chorus has advantages over pairing an end user with a single crowd worker and end users completing their own tasks in terms of speed, quality, and breadth of assistance. Chorus demonstrates a new future in which conversational assistants are made usable in the real world by combining human and machine intelligence, and may enable a useful new way of interacting with the crowds powering other systems. Walter S. Lasecki, Rachel Wesley, Jeffrey Nichols 0001, Anand Kulkarni, James F. Allen, Jeffrey P. Bigham |
UIST | 6 |
| 2012 | Real-Time Collaborative Planning with the CrowdabstractPlanning is vital to a wide range of domains, including robotics, military strategy, logistics, itinerary generation and more, that both humans and computers find difficult. Collaborative planning holds the promise of greatly improving performance on these tasks by leveraging the strengths of both humans and automated planners. However, this requires formalizing the problem domain and input, which must be done by hand, a priori, restricting its use in general real-world domains. We propose using a real-time crowd of workers to simultaneously solve the planning problem, formalize the domain, and train an automated system. As plans are developed, the system is able to learn the domain, and contribute larger segments of work. Walter S. Lasecki, Jeffrey P. Bigham, James F. Allen, George Ferguson |
AAAI | 2 |
| 2012 | Online Sequence Alignment for Real-Time Audio Transcription by Non-ExpertsabstractReal-time transcription provides deaf and hard of hearing people visual access to spoken content, such as classroom instruction, and other live events. Currently, the only reliable source of real-time transcriptions are expensive, highly-trained experts who are able to keep up with speaking rates. Automatic speech recognition is cheaper but produces too many errors in realistic settings. We introduce a new approach in which partial captions from multiple non-experts are combined to produce a high-quality transcription in real-time. We demonstrate the potential of this approach with data collected from 20 non-expert captionists. Walter S. Lasecki, Christopher D. Miller, Donato Borrello, Jeffrey P. Bigham |
AAAI | 4 |
| 2012 | Crowdsourcing subjective fashion advice using VizWiz: challenges and opportunitiesabstractFashion is a language. How we dress signals to others who we are and how we want to be perceived. However, this language is primarily visual, making it inaccessible to people with vision impairments. Someone who is low-vision or completely blind cannot see what others are wearing or readily know what constitutes the norms and extremes of fashion, but most everyone they encounter can see (and judge) their fashion choices. We describe our findings of a diary study with people with vision impairments that revealed the many accessibility barriers fashion presents, and how an online survey revealed that clothing decisions are often made collaboratively, regardless of visual ability. Based on these findings, we identified a need for a collaborative and real-time environment for fashion advice. We have tested the feasibility of providing this advice through crowdsourcing using VizWiz, a mobile phone application where participants receive nearly real-time answers to visual questions. Our pilot study results show that this application has the potential to address a great need within the blind community, but remaining challenges include improving photo capture and assembling a set of crowd workers with the requisite expertise. More broadly our research highlights the feasibility of using crowdsourcing for subjective, opinion-based advice. Michele A. Burton, Erin L. Brady, Robin Brewer, Callie Neylan, Jeffrey P. Bigham, Amy Hurst |
ASSETS | 5 |
| 2012 | A readability evaluation of real-time crowd captions in the classroomabstractDeaf and hard of hearing individuals need accommodations that transform aural to visual information, such as captions that are generated in real-time to enhance their access to spoken information in lectures and other live events. The captions produced by professional captionists work well in general events such as community or legal meetings, but is often unsatisfactory in specialized content events such as higher education classrooms. In addition, it is hard to hire professional captionists, especially those that have experience in specialized content areas, as they are scarce and expensive. The captions produced by commercial automatic speech recognition (ASR) software are far cheaper, but is often perceived as unreadable due to ASR's sensitivity to accents, background noise and slow response time. We ran a study to evaluate the readability of captions generated by a new crowd captioning approach versus professional captionists and ASR. In this approach, captions are typed by classmates into a system that aligns and merges the multiple incomplete caption streams into a single, comprehensive real-time transcript. Our study asked 48 deaf and hearing readers to evaluate transcripts produced by a professional captionist, ASR and crowd captioning software respectively and found the readers preferred crowd captions over professional captions and ASR. Raja S. Kushalnagar, Walter S. Lasecki, Jeffrey P. Bigham |
ASSETS | 3 |
| 2012 | Online quality control for real-time crowd captioningabstractApproaches for real-time captioning of speech are either expensive (professional stenographers) or error-prone (automatic speech recognition). As an alternative approach, we have been exploring whether groups of non-experts can collectively caption speech in real-time. In this approach, each worker types as much as they can and the partial captions are merged together in real-time automatically. This approach works best when partial captions are correct and received within a few seconds of when they were spoken, but these assumptions break down when engaging workers on-demand from existing sources of crowd work like Amazon's Mechanical Turk. In this paper, we present methods for quickly identifying workers who are producing good partial captions and estimating the quality of their input. We evaluate these methods in experiments run on Mechanical Turk in which a total of 42 workers captioned 20 minutes of audio. The methods introduced in this paper were able to raise overall accuracy from 57.8% to 81.22% while keeping coverage of the ground truth signal nearly unchanged. Walter S. Lasecki, Jeffrey P. Bigham |
ASSETS | 2 |
| 2012 | Improving the accessibility of computing enrichment programs (abstract only)abstractMany wonderful enrichment programs have been created to introduce young people to computing, but with little attention to making them accessible to students with disabilities. In this workshop participants will learn from practitioners who have introduced computing and programming to young people with disabilities. They will also learn first-hand from students with disabilities about their needs in learning programming. There will be breakout sessions for participants to apply what they have learned to improve existing enrichment programs such as Alice, Arduino, Scratch, Kodu, App Inventor, Greenfoot, Lego Mindstorms, Processing, and Computer Science Unplugged. Richard E. Ladner, Karen Alkoby, Jeffrey P. Bigham, Stephanie Ludi, Daniela Marghitu, Andreas Stefik |
SIGCSE | 3 |
| 2012 | Real-time captioning by groups of non-expertsabstractReal-time captioning provides deaf and hard of hearing people immediate access to spoken language and enables participation in dialogue with others. Low latency is critical because it allows speech to be paired with relevant visual cues. Currently, the only reliable source of real-time captions are expensive stenographers who must be recruited in advance and who are trained to use specialized keyboards. Automatic speech recognition (ASR) is less expensive and available on-demand, but its low accuracy, high noise sensitivity, and need for training beforehand render it unusable in real-world situations. In this paper, we introduce a new approach in which groups of non-expert captionists (people who can hear and type) collectively caption speech in real-time on-demand. We present Legion:Scribe, an end-to-end system that allows deaf people to request captions at any time. We introduce an algorithm for merging partial captions into a single output stream in real-time, and a captioning interface designed to encourage coverage of the entire audio stream. Evaluation with 20 local participants and 18 crowd workers shows that non-experts can provide an effective solution for captioning, accurately covering an average of 93.2% of an audio stream with only 10 workers and an average per-word latency of 2.9 seconds. More generally, our model in which multiple workers contribute partial inputs that are automatically merged in real-time may be extended to allow dynamic groups to surpass constituent individuals (even experts) on a variety of human performance tasks. Walter S. Lasecki, Christopher D. Miller, Adam Sadilek, Andrew Abumoussa, Donato Borrello, Raja S. Kushalnagar, Jeffrey P. Bigham |
UIST | 7 |
| 2012 | Finding your friends and following them to where you areabstractLocation plays an essential role in our lives, bridging our online and offline worlds. This paper explores the interplay between people's location, interactions, and their social ties within a large real-world dataset. We present and evaluate Flap, a system that solves two intimately related tasks: link and location prediction in online social networks. For link prediction, Flap infers social ties by considering patterns in friendship formation, the content of people's messages, and user location. We show that while each component is a weak predictor of friendship alone, combining them results in a strong model, accurately identifying the majority of friendships. For location prediction, Flap implements a scalable probabilistic model of human mobility, where we treat users with known GPS positions as noisy sensors of the location of their friends. We explore supervised and unsupervised learning scenarios, and focus on the efficiency of both learning and inference. We evaluate Flap on a large sample of highly active users from two distinct geographical areas and show that it (1) reconstructs the entire friendship graph with high accuracy even when no edges are given; and (2) infers people's fine-grained location, even when they keep their data private and we can only access the location of their friends. Our models significantly outperform current comparable approaches to either task. Adam Sadilek, Henry A. Kautz, Jeffrey P. Bigham |
WSDM | 3 |
| 2011 | The design of human-powered access technologyabstractPeople with disabilities have always overcome accessibility problems by enlisting people in their community to help. The Internet has broadened the available community and made it easier to get on-demand assistance remotely. In particular, the past few years have seen the development of technology in both research and industry that uses human power to overcome technical problems too difficult to solve automatically. In this paper, we frame recent developments in human computation in the historical context of accessibility, and outline a framework for discussing new advances in human-powered access technology. Specifically, we present a set of 13 design principles for human-powered access technology motivated both by historical context and current technological developments. We then demonstrate the utility of these principles by using them to compare several existing human-powered access technologies. The power of identifying the 13 principles is that they will inspire new ways of thinking about human-powered access technologies. Jeffrey P. Bigham, Richard E. Ladner, Yevgen Borodin |
ASSETS | 1 |
| 2011 | Supporting blind photographyabstractBlind people want to take photographs for the same reasons as others -- to record important events, to share experiences, and as an outlet for artistic expression. Furthermore, both automatic computer vision technology and human-powered services can be used to give blind people feedback on their environment, but to work their best these systems need high-quality photos as input. In this paper, we present the results of a large survey that shows how blind people are currently using cameras. Next, we introduce EasySnap, an application that provides audio feedback to help blind people take pictures of objects and people and show that blind photographers take better photographs with this feedback. We then discuss how we iterated on the portrait functionality to create a new application called PortraitFramer designed specifically for this function. Finally, we present the results of an in-depth study with 15 blind and low-vision participants, showing that they could pick up how to successfully use the application very quickly. Chandrika Jayant, Hanjie Ji, Samuel White, Jeffrey P. Bigham |
ASSETS | 4 |
| 2011 | Multimodal summarization of complex sentencesabstractIn this paper, we introduce the idea of automatically illustrating complex sentences as multimodal summaries that combine pictures, structure and simplified compressed text. By including text and structure in addition to pictures, multimodal summaries provide additional clues of what happened, who did it, to whom and how, to people who may have difficulty reading or who are looking to skim quickly. We present ROC-MMS, a system for automatically creating multimodal summaries (MMS) of complex sentences by generating pictures, textual summaries and structure. We show that pictures alone are insufficient to help people understand most sentences, especially for readers who are unfamiliar with the domain. An evaluation of ROC-MMS in the Wikipedia domain illustrates both the promise and challenge of automatically creating multimodal summaries. Naushad UzZaman, Jeffrey P. Bigham, James F. Allen |
IUI | 2 |
| 2011 | Real-time crowd control of existing interfacesabstractCrowdsourcing has been shown to be an effective approach for solving difficult problems, but current crowdsourcing systems suffer two main limitations: (i) tasks must be repackaged for proper display to crowd workers, which generally requires substantial one-off programming effort and support infrastructure, and (ii) crowd workers generally lack a tight feedback loop with their task. In this paper, we introduce Legion, a system that allows end users to easily capture existing GUIs and outsource them for collaborative, real-time control by the crowd. We present mediation strategies for integrating the input of multiple crowd workers in real-time, evaluate these mediation strategies across several applications, and further validate Legion by exploring the space of novel applications that it enables. Walter S. Lasecki, Kyle I. Murray, Samuel White, Rob Miller 0001, Jeffrey P. Bigham |
UIST | 5 |
| 2011 | Beyond autocomplete: Automatic function definitionabstractProgrammers have used autocomplete to reduce the cognitive overhead of remembering exhaustive lists of APIs for years. Autocomplete has a primary and obvious point of failure: when a programmer expects a certain method or function name to exist and it does not, the autocompletion list simply stops displaying results and disappears. We describe automatic function definition (AFD), which can succeed where autocomplete fails. It is a novel way to reduce the impact of threadbare libraries, increase the coding speed of primary programming tasks, and distribute work among different types of programmers and automatic tools. Instead of seeing an empty list, users can instead perform automatic function definition, which uses several sources to define the function that the user intended to use. We present three complementary techniques for defining functions based on the information about the function that a user provides while writing code as usual: code search, fellow programmers, and the crowd. Finally, we discuss our implementation of this work in progress and plans for evaluation. Kyle I. Murray, Jeffrey P. Bigham |
VL/HCC | 2 |
| 2010 | Accessibility by demonstration: enabling end users to guide developers to web accessibility solutionsabstractFew web developers have been explicitly trained to create accessible web pages, and are unlikely to recognize subtle accessibility and usability concerns that disabled people face. Evaluating web pages with assistive technology can reveal problems, but this software takes time to install and its complexity can be overwhelming. To address these problems, we introduce a new approach for accessibility evaluation called Accessibility by Demonstration (ABD). ABD lets assistive technology users retroactively record accessibility problems at the time they experience them as human-readable macros and easily send those recordings and the software necessary to replay them to others. This paper describes an implementation of ABD as an extension to the WebAnywhere screen reader, and presents an evaluation with 15 web developers not experienced with accessibility showing that interacting with these recordings helped them understand and fix some subtle accessibility problems better than existing tools. Jeffrey P. Bigham, Jeremy T. Brudvik, Bernie Zhang |
ASSETS | 1 |
| 2010 | Asl-stem forum: enabling sign language to grow through online collaborationabstractAmerican Sign Language (ASL) currently lacks agreed-upon signs for complex terms in scientific fields, causing deaf students to miss or misunderstand course material. Furthermore, the same term or concept may have multiple signs, resulting in inconsistent standards and strained collaboration. The ASL-STEM Forum is an online, collaborative, video forum for sharing ASL signs and discussing them. An initial user study of the Forum has shown its viability and revealed lessons in accommodating varying user types, from lurkers to advanced contributors, until critical mass is achieved. Anna Cavender, Daniel S. Otero, Jeffrey P. Bigham, Richard E. Ladner |
CHI | 3 |
| 2010 | WebTrax: Visualizing Non-visual Web Interactions
Jeffrey P. Bigham, Kyle I. Murray |
ICCHP (2) | 1 |
| 2010 | VizWiz: nearly real-time answers to visual questionsabstractThe lack of access to visual information like text labels, icons, and colors can cause frustration and decrease independence for blind people. Current access technology uses automatic approaches to address some problems in this space, but the technology is error-prone, limited in scope, and quite expensive. In this paper, we introduce VizWiz, a talking application for mobile phones that offers a new alternative to answering visual questions in nearly real-time - asking multiple people on the web. To support answering questions quickly, we introduce a general approach for intelligently recruiting human workers in advance called quikTurkit so that workers are available when new questions arrive. A field deployment with 11 blind participants illustrates that blind people can effectively use VizWiz to cheaply answer questions in their everyday lives, highlighting issues that automatic approaches will need to address to be useful. Finally, we illustrate the potential of using VizWiz as part of the participatory design of advanced tools by using it to build and evaluate VizWiz::LocateIt, an interactive mobile tool that helps blind people solve general visual search problems. Jeffrey P. Bigham, Chandrika Jayant, Hanjie Ji, Greg Little, Rob Miller 0001, Robin Miller, Aubrey Tatarowicz, Brandyn White, Samuel White, Tom Yeh |
UIST | 1 |
| 2010 | A conversational interface to web automationabstractThis paper presents CoCo, a system that automates web tasks on a user's behalf through an interactive conversational interface. Given a short command such as "get road conditions for highway 88," CoCo synthesizes a plan to accomplish the task, executes it on the web, extracts an informative response, and returns the result to the user as a snippet of text. A novel aspect of our approach is that we leverage a repository of previously recorded web scripts and the user's personal web browsing history to determine how to complete each requested task. This paper describes the design and implementation of our system, along with the results of a brief user study that evaluates how likely users are to understand what CoCo does for them. Tessa A. Lau, Julian A. Cerruti, Guillermo Manzato, Mateo N. Bengualid, Jeffrey P. Bigham, Jeffrey Nichols 0001 |
UIST | 5 |
| 2009 | ClassInFocus: enabling improved visual attention strategies for deaf and hard of hearing studentsabstractDeaf and hard of hearing students must juggle their visual attention in current classroom settings. Managing many visual sources of information (instructor, interpreter or captions, slides or whiteboard, classmates, and personal notes) can be a challenge. ClassInFocus automatically notifies students of classroom changes, such as slide changes or new speakers, helping them employ more beneficial observing strategies. A user study of notification techniques shows that students who liked the notifications were more likely to visually utilize them to improve performance. Anna Cavender, Jeffrey P. Bigham, Richard E. Ladner |
ASSETS | 2 |
| 2009 | Evaluating existing audio CAPTCHAs and an interface optimized for non-visual useabstractAudio CAPTCHAs were introduced as an accessible alternative for those unable to use the more common visual CAPTCHAs, but anecdotal accounts have suggested that they may be more difficult to solve. This paper demonstrates in a large study of more than 150 participants that existing audio CAPTCHAs are clearly more difficult and time-consuming to complete as compared to visual CAPTCHAs for both blind and sighted users. In order to address this concern, we developed and evaluated a new interface for solving CAPTCHAs optimized for non-visual use that can be added in-place to existing audio CAPTCHAs. In a subsequent study, the optimized interface increased the success rate of blind participants by 59% on audio CAPTCHAs, illustrating a broadly applicable principle of accessible design: the most usable audio interfaces are often not direct translations of existing visual interfaces. Jeffrey P. Bigham, Anna Cavender |
CHI | 1 |
| 2009 | Trailblazer: enabling blind users to blaze trails through the webabstractFor blind web users, completing tasks on the web can be frustrating. Each step can require a time-consuming linear search of the current web page to find the needed interactive element or piece of information. Existing interactive help systems and the playback components of some programming-by-demonstration tools identify the needed elements of a page as they guide the user through predefined tasks, obviating the need for a linear search on each step. We introduce TrailBlazer, a system that provides an accessible, non-visual interface to guide blind users through existing how-to knowledge. A formative study indicated that participants saw the value of TrailBlazer but wanted to use it for tasks and web sites for which no existing script was available. To address this, TrailBlazer offers suggestion-based help created on-the-fly from a short, user-provided task description and an existing repository of how-to knowledge. In an evaluation on 15 tasks, the correct prediction was contained within the top 5 suggestions 75.9% of the time. Jeffrey P. Bigham, Tessa A. Lau, Jeffrey Nichols 0001 |
IUI | 1 |
| 2009 | Mining web interactions to automatically create mash-upsabstractThe deep web contains an order of magnitude more information than the surface web, but that information is hidden behind the web forms of a large number of web sites. Metasearch engines can help users explore this information by aggregating results from multiple resources, but previously these could only be created and maintained by programmers. In this paper, we explore the automatic creation of metasearch mash-ups by mining the web interactions of multiple web users to find relations between query forms on different web sites. We also present an implemented system called TX2 that uses those connections to search multiple deep web resources simultaneously and integrate the results in context in a single results page. TX2 illustrates the promise of constructing mash-ups automatically and the potential of mining web interactions to explore deep web resources. Jeffrey P. Bigham, Ryan S. Kaminsky, Jeffrey Nichols 0001 |
UIST | 1 |
| 2008 | What's new?: making web page updates accessibleabstractWeb applications facilitated by technologies such as JavaScript, DHTML, AJAX, and Flash use a considerable amount of dynamic web content that is either inaccessible or unusable by blind people. Server side changes to web content cause whole page refreshes, but only small sections of the page update, causing blind web users to search linearly through the page to find new content. The connecting theme is the need to quickly and unobtrusively identify the segments of a web page that have changed and notify the user of them. In this paper we propose Dynamo, a system designed to unify different types of dynamic content and make dynamic content accessible to blind web users. Dynamo treats web page updates uniformly and its methods encompass both web updates enabled through dynamic content and scripting, and updates resulting from static page refreshes, form submissions, and template-based web sites. From an algorithmic and interaction perspective Dynamo detects underlying changes and provides users with a single and intuitive interface for reviewing the changes that have occurred. We report on the quantitative and qualitative results of an evaluation conducted with blind users. These results suggest that Dynamo makes access to dynamic content faster, and that blind web users like it better than existing interfaces. Yevgen Borodin, Jeffrey P. Bigham, Rohit Raman, I. V. Ramakrishnan |
ASSETS | 2 |
| 2008 | Hunting for headings: sighted labeling vs. automatic classification of headingsabstractProper use of headings in web pages can make navigation more efficient for blind web users by indicating semantic divisions in the page. Unfortunately, many web pages do not use proper HTML markup (h1-h6 tags) to indicate headings, instead using visual styling to create headings, thus making the distinction between headings and other page text indistinguishable to blind users. In a user study in which sighted participants labeled headings on a set of web pages, participants did not often agree on which elements on the page should be labeled as headings, suggesting why headings are not used properly on the web today. To address this problem, we have created a system called HeadingHunter that predicts whether web page text semantically functions as a heading by examining visual features of the text as rendered in a web browser. Its performance in labeling headings compares favorably with both a manually-classified set of heading examples and the combined results of the sighted labelers in our study. The resulting system illustrates a general methodology of creating simple scripts operating over visual features that can be directly included in existing tools. Jeremy T. Brudvik, Jeffrey P. Bigham, Anna Cavender, Richard E. Ladner |
ASSETS | 2 |
| 2008 | Slide rule: making mobile touch screens accessible to blind people using multi-touch interaction techniquesabstractRecent advances in touch screen technology have increased the prevalence of touch screens and have prompted a wave of new touch screen-based devices. However, touch screens are still largely inaccessible to blind users, who must adopt error-prone compensatory strategies to use them or find accessible alternatives. This inaccessibility is due to interaction techniques that require the user to visually locate objects on the screen. To address this problem, we introduce Slide Rule, a set of audio-based multi-touch interaction techniques that enable blind users to access touch screen applications. We describe the design of Slide Rule, our interaction techniques, and a user study in which 10 blind people used Slide Rule and a button-based Pocket PC screen reader. Results show that Slide Rule was significantly faster than the button-based system, and was preferred by 7 of 10 users. However, users made more errors when using Slide Rule than when using the more familiar button-based system. Shaun K. Kane, Jeffrey P. Bigham, Jacob O. Wobbrock |
ASSETS | 2 |
| 2008 | Accessibility commons: a metadata infrastructure for web accessibilityabstractResearch projects, assistive technology, and individuals all create metadata in order to improve Web accessibility for visually impaired users. However, since these projects are disconnected from one another, this metadata is isolated in separate tools, stored in disparate repositories, and represented in incompatible formats. Web accessibility could be greatly improved if these individual contributions were merged. An integration method will serve as the bridge between future academic research projects and end users, enabling new technologies to reach end users more quickly. Therefore we introduce Accessibility Commons, a common infrastructure to integrate, store, and share metadata designed to improve Web accessibility. We explore existing tools to show how the metadata that they produce could be integrated into this common infrastructure, we present the design decisions made in order to help ensure that our common repository will remain relevant in the future as new metadata is developed, and we discuss how the common infrastructure component facilitates our broader social approach to improving accessibility. Shinya Kawanaka, Yevgen Borodin, Jeffrey P. Bigham, Darren Lunn, Hironobu Takagi, Chieko Asakawa |
ASSETS | 3 |
| 2008 | Addressing Performance and Security in a Screen Reading Web Application That Enables Accessibility AnywhereabstractThe web provides nearly ubiquitous access to information, but access for blind web users requires the use of expensive, specialized software programs called screen readers unlikely to be installed on most computers. WebAnywhere is a self-voicing, web-browsing web application that makes the web accessible for blind web users from most devices with web access. WebAnywhere requires no special permissions or additional software to be installed on the host machine, enabling it provide a self-voicing interface on almost any web-enabled device. WebAnywhere’s interface is written in Javascript, speech is retrieved from a remote server, and sounds are played using either Flash or existing embedded sound players. This paper describes the performance and security implications of the system’s unique design and how it has been engineered to provide usable access anywhere. Specifically, we present prefetching andcaching strategies developed to make the system responsive even on low-bandwidth connections and security considerations that replicate existing browser security policies. Jeffrey P. Bigham, Craig Prince, Richard E. Ladner |
ICWE | 1 |
| 2008 | ASL-STEM Forum: A Bottom-Up Approach to Enabling American Sign Language to Grow in STEM Fields
Jeffrey P. Bigham, Daniel S. Otero, Jessica N. DeWitt, Anna Cavender, Richard E. Ladner |
ICWSM | 1 |
| 2008 | Transcendence: enabling a personal view of the deep webabstractA wealth of structured, publicly-available information exists in the deep web but is only accessible by querying web forms. As a result, users are restricted by the interfaces provided and lack a convenient mechanism to express novel and independent extractions and queries on the underlying data. Transcendence enables personalized access to the deep web by enabling users to partially reconstruct web databases in order to perform new types of queries. From just a few examples, Transcendence helps users produce a large number of values for form input fields by using unsupervised information extraction and collaborative filtering of user suggestions. Structural and semantic analysis of returned pages finds individual results and identifies relevant fields. Users may revise automated decisions, balancing the power of automation with the errors it can introduce. In a user evaluation, both programmers and non-programmers found Transcendence to be a powerful way to explore deep web resources and wanted to use it in the future. Jeffrey P. Bigham, Anna Cavender, Ryan S. Kaminsky, Craig Prince, Tyler Robison |
IUI | 1 |
| 2008 | Inspiring blind high school students to pursue computer science with instant messaging chatbotsabstractBlind students are an underrepresented group in computer science. In this paper, we describe our experience preparing and leading the computer science track at the National Federation of the Blind Youth Slam. As part of this workshop, fifteen blind high school students created and personalized instant messaging chatbots, a project designed to be completely accessible to blind students. Chatbots enable students to infuse their own personalities into a socially-oriented program that incorporates ideas from artificial intelligence, natural language processing, and web services. We first outline the chatbots project and curriculum, which has wide appeal for all students, and then offer general design principles used to create it that can help ensure the accessibility of future projects. Students created their chatbots using a real programming language and were guided by both blind and sighted mentors. By programming from the start in a supportive environment, our students will gain the confidence to persevere in computer science in the future. Jeffrey P. Bigham, Maxwell B. Aller, Jeremy T. Brudvik, Jessica O. Leung, Lindsay A. Yazzolino, Richard E. Ladner |
SIGCSE | 1 |
| 2008 | Webanywhere: enabling a screen reading interface for the web on any computerabstractPeople often use computers other than their own to access web content, but blind users are restricted to using computers equipped with expensive, special-purpose screen reading programs that they use to access the web. WebAnywhere is a web-based, self-voicing web application that enables blind web users to access the web from almost any computer that can produce sound without installing new software. WebAnywhere could serve as a convenient, low-cost solution for blind users on-the-go, for blind users unable to afford another screen reader and for web developers targeting accessible design. This paper describes the implementation of WebAnywhere, overviews an evaluation of it by blind web users, and summarizes a survey of public terminals that shows it can run on most public computers. Jeffrey P. Bigham, Craig Prince, Richard E. Ladner |
WWW | 1 |
| 2007 | WebinSitu: a comparative analysis of blind and sighted browsing behaviorabstractWeb browsing is inefficient for blind web users because of persistent accessibility problems, but the extent of these problems and their practical effects from the perspective of the user has not been sufficiently examined. We conducted a study in situ to investigate the accessibility of the web as experienced by web users. This remote study used an advanced web proxy that leverages AJAX technology to record both the pages viewed and the actions taken by users on the web pages that they visited. Our study was conducted remotely over the period of one week, and our participants used the assistive technology and software to which they were already accustomed and had already configured according to preference. These advantages allowed us to aggregate observations of many users and to explore the practical effects on and coping strategies employed by our blind participants. Our study reflects web accessibility from the perspective of web users and describes quantitative differences in the browsing behavior of blind and sighted web users. Jeffrey P. Bigham, Anna Cavender, Jeremy T. Brudvik, Jacob O. Wobbrock, Richard E. Ladner |
ASSETS | 1 |
| 2007 | WebAnywhere: a screen reader on-the-goabstractPeople often use computers other than their own to browse the web, but blind web users are limited in where they access the web because they require specialized, expensive programs for access. WebAnywhere is a web-based, self-voicing browser that enables blind web users to accessthe web from almost any computer that can produce sound. The system runs entirely in standard web browsers and requires no additional software to be installed. The system could serve as a convenient, low-cost solution for both web developers targeting accessible design and end users unable to afford a full screen reader. This demonstration will offer visitors the opportunity to try WebAnywhere and learn more about it. Jeffrey P. Bigham, Craig Prince |
ASSETS | 1 |
| 2007 | Increasing web accessibility by automatically judging alternative text qualityabstractThe lack of appropriate alternative text for web images remains a problem for blind users and others accessing the web with non-visual interfaces. The content contained within web images is vital for understanding many web sites but the majority are assigned either inaccurate alternative text or none at all. The capability to automatically judge the quality of alternative text has the promise to dramatically improve the accessibility of the web by bringing intelligence to three categories of interfaces: tools that help web authors verify that they have provided adequate alternative text for web images, systems that automatically produce and insert alternative text for web images, and screen reading software. In this paper we describe a classifier capable of measuring the quality of alternative text given only a few labeled training examples by automatically considering the image context. Jeffrey P. Bigham |
IUI | 1 |
| 2006 | WebInSight: : making web images accessibleabstractImages without alternative text are a barrier to equal web access for blind users. To illustrate the problem, we conducted a series of studies that conclusively show that a large fraction of significant images have no alternative text. To ameliorate this problem, we introduce WebInSight, a system that automatically creates and inserts alternative text into web pages on-the-fly. To formulate alternative text for images, we present three labeling modules based on web context analysis, enhanced optical character recognition (OCR) and human labeling. The system caches alternative text in a local database and can add new labels seamlessly after a web page is downloaded, resulting in minimal impact to the browsing experience. Jeffrey P. Bigham, Ryan S. Kaminsky, Richard E. Ladner, Oscar M. Danielsson, Gordon L. Hempton |
ASSETS | 1 |