Walter S. Lasecki

dblp:04/10300 · also Walter Stephen Lasecki · DBLP profile ↗
← Back
71ranked-venue papers
22as first author
3since 2021 · last 2021
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 58 · 17 first-author · 2 since 2021Artificial intelligence and machine learning · 12 · 5 first-author · 1 since 2021Databases, data management, data science and information retrieval · 11 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-authorApplied, interdisciplinary, general and emerging computing · 2Systems, architecture and hardware · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Human-computer interaction and pervasive computing
33 papers
Collaborative and social computing · 51% Human-AI interaction · 13% Learning and educational technologies · 12%
Artificial intelligence
12 papers
Question answering and dialogue systems · 45% Trustworthy machine learning · 24% Information extraction and text analysis · 20%
Software engineering, system software, and programming languages
4 papers
Software testing · 69% Requirements engineering and software design · 28% Services computing and microservices · 3%

Topics — the 30 heaviest of 58, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Collaborative and social computing
crowdsourcing
3.7192020
Crowdsourced Detection of Emotionally Manipulative Language · CHI 2020
Bolt: Instantaneous Crowdsourcing via Just-in-Time Training · CHI 2018
Kurator: Using The Crowd to Help Families With Personal Curation Tasks · CSCW 2017
Requirements engineering and software design
developer support tools
0.522017
Codeon: On-Demand Software Development Assistance · CHI 2017
Towards Providing On-Demand Expert Support for Software Developers · CHI 2016
Natural language and speech › Question answering and dialogue systems › task-oriented dialogue
dialogue state tracking
0.512021
Overview of the Eighth Dialog System Technology Challenge: DSTC8 · IEEE ACM Trans. Audio Speech Lang. Process. 2021
Learning and educational technologies
knowledge acquisition
0.512021
Think-Aloud Computing: Supporting Rich and Low-Effort Knowledge Capture · CHI 2021
Machine learning › Trustworthy machine learning
robustness
0.412020
Towards Hybrid Human-AI Workflows for Unknown Unknown Detection · WWW 2020
Human-AI interaction
interactive machine learning
0.412020
Towards Hybrid Human-AI Workflows for Unknown Unknown Detection · WWW 2020
Software testing › dynamic testing › manual testing
crowdsourced testing
0.412020
Improving Crowd-Supported GUI Testing with Structural Guidance · CHI 2020
Software testing
GUI testing
0.412020
Improving Crowd-Supported GUI Testing with Structural Guidance · CHI 2020
Software testing
test coverage
0.412020
Improving Crowd-Supported GUI Testing with Structural Guidance · CHI 2020
Natural language and speech › Information extraction and text analysis
dialogue analysis
0.412019
A Large-Scale Corpus for Conversation Disentanglement · ACL (1) 2019
Natural language and speech › Question answering and dialogue systems › multi-party dialogue
dialogue disentanglement
0.412019
A Large-Scale Corpus for Conversation Disentanglement · ACL (1) 2019
Data models and query languages › natural language interface › natural language interface to database
text-to-SQL
0.412019
CoSQL: A Conversational Text-to-SQL Challenge Towards Cross-Domain Natural Language Interfaces to Databases · EMNLP/IJCNLP (1) 2019
Human-AI interaction
AI-assisted decision-making
0.412019
Updates in Human-AI Teams: Understanding and Addressing the Performance/Compatibility Tradeoff · AAAI 2019
Collaborative and social computing
remote collaboration
0.422017
Codeon: On-Demand Software Development Assistance · CHI 2017
Towards Providing On-Demand Expert Support for Software Developers · CHI 2016
Ubiquitous computing and smart environments › context recognition
activity recognition
0.422014
Finding dependencies between actions using the crowd · CHI 2014
Real-time crowd labeling for deployable activity recognition · CSCW 2013
Collaborative and social computing › computer-supported cooperative work › collaborative information seeking
collaborative browsing
0.312018
Arboretum and Arbility: Improving Web Accessibility Through a Shared Browsing Architecture · UIST 2018
Learning and educational technologies
online learning
0.312018
Enhancing Online Problems Through Instructor-Centered Tools for Randomized Experiments · CHI 2018
Accessibility and assistive technology
screen reader accessibility
0.312018
Arboretum and Arbility: Improving Web Accessibility Through a Shared Browsing Architecture · UIST 2018
Accessibility and assistive technology
web accessibility
0.312018
Arboretum and Arbility: Improving Web Accessibility Through a Shared Browsing Architecture · UIST 2018
Natural language and speech › Information extraction and text analysis › document analysis
email information extraction
0.312017
WearMail: On-the-Go Access to Information in Your Email with a Privacy-Preserving Human Computation Workflow · UIST 2017
Collaborative and social computing
human computation
0.312017
WearMail: On-the-Go Access to Information in Your Email with a Privacy-Preserving Human Computation Workflow · UIST 2017
Accessibility and assistive technology › captioning
real-time captioning
0.322012
Real-time captioning by groups of non-experts · UIST 2012
Online Sequence Alignment for Real-Time Audio Transcription by Non-Experts · AAAI 2012
Accessibility and assistive technology › assistive technology for visual impairment
assistive technology for blind users
0.212015
RegionSpeak: Quick Comprehensive Spatial Descriptions of Complex Images for Blind Users · CHI 2015
Collaborative and social computing › crowdsourcing › micro-task crowdsourcing
microtask design
0.212015
The Effects of Sequence and Delay on Crowd Work · CHI 2015
User interface design and tools
prototyping
0.212015
Apparition: Crowdsourced User Interfaces that Come to Life as You Sketch Them · CHI 2015
Learning and educational technologies › instructional design
task sequencing
0.212015
The Effects of Sequence and Delay on Crowd Work · CHI 2015
Design research and methods › research methodology
wizard of oz prototyping
0.212015
Apparition: Crowdsourced User Interfaces that Come to Life as You Sketch Them · CHI 2015
Privacy and data protection › differential privacy
privacy-accuracy tradeoff
0.212015
Exploring Privacy and Accuracy Trade-Offs in Crowdsourced Behavioral Video Coding · CHI 2015
Collaborative and social computing › crowdsourcing
crowd-powered systems
0.212014
Information extraction and manipulation threats in crowd-powered systems · CSCW 2014
Collaborative and social computing › crowdsourcing
crowd work
0.222018
Arboretum and Arbility: Improving Web Accessibility Through a Shared Browsing Architecture · UIST 2018
SketchExpress: Remixing Animations for More Effective Crowd-Powered Prototyping of Interactive Interfaces · UIST 2017

Methods — techniques the papers use, named apart from their topics

crowdsourcing · 1.7interviews · 0.9deception study · 0.9crowd worker pattern identification · 0.9classifier retraining · 0.9anchor comparison · 0.9cross-domain generalization · 0.8think-aloud protocol · 0.5formative study · 0.5end-to-end dialog modeling · 0.5event-flow graph · 0.4event flow graph · 0.4crowd labeling · 0.4retraining objective with error penalty · 0.4performance/compatibility tradeoff analysis · 0.4corpus annotation · 0.4encryption · 0.4crowd work · 0.4
YearPublicationVenuePosition
2021 Think-Aloud Computing: Supporting Rich and Low-Effort Knowledge Capture
abstract
When users complete tasks on the computer, the knowledge they leverage and their intent is often lost because it is tedious or challenging to capture. This makes it harder to understand why a colleague designed a component a certain way or to remember requirements for software you wrote a year ago. We introduce think-aloud computing, a novel application of the think-aloud protocol where computer users are encouraged to speak while working to capture rich knowledge with relatively low effort. Through a formative study we find people shared information about design intent, work processes, problems encountered, to-do items, and other useful information. We developed a prototype that supports think-aloud computing by prompting users to speak and contextualizing speech with labels and application context. Our evaluation shows more subtle design decisions and process explanations were captured in think-aloud than via traditional documentation. Participants reported that think-aloud required similar effort as traditional documentation.
Rebecca Krosnick, Fraser Anderson, Justin Matejka, Steve Oney, Walter S. Lasecki, Tovi Grossman, George W. Fitzmaurice
CHI5
2021 Human-in-the-loop Pose Estimation via Shared Autonomy
abstract
Reliable, efficient shared autonomy requires balancing human operation and robot automation on complex tasks, such as dexterous manipulation. Adding to the difficulty of shared autonomy is a robot’s limited ability to perceive the 6 degree-of-freedom pose of objects, which is essential to perform manipulations those objects afforded. Inspired by Monte Carlo Localization, we propose a generative human-in-the-loop approach to estimating object pose. We characterize the performance of our mixed-initiative 3D registration approach using 2D pointing devices via a user study. Seeking an analog for Fitts’s Law for 3D registration, we introduce a new evaluation framework that takes the entire registration process into account instead of only the outcome. When combined with estimates of registration confidence, we posit that mixed-initiative registration will reduce the human workload while maintaining or even improving final pose estimation accuracy.
Zhefan Ye, Jean Y. Song, Zhiqiang Sui, Stephen Hart, Jorge Vilchis, Walter S. Lasecki, Odest Chadwicke Jenkins
IUI6
2021 Overview of the Eighth Dialog System Technology Challenge: DSTC8
abstract
This paper introduces the Eighth Dialog System Technology Challenge. In line with recent challenges, the eighth edition focuses on applying end-to-end dialog technologies in a pragmatic way for multi-domain task-completion, noetic response selection, audio visual scene-aware dialog, and schema-guided dialog state tracking tasks. This paper describes the task definition, provided datasets, baselines and evaluation set-up for each track. We also summarize the results of the submitted systems to highlight the overall trends of the state-of-the-art technologies for the tasks.
Seokhwan Kim, Michel Galley, R. Chulaka Gunasekara, Adam Atkinson, Baolin Peng, Hannes Schulz, Jianfeng Gao 0001, Jinchao Li, Mahmoud Adada, Minlie Huang, Luis A. Lastras, Jonathan K. Kummerfeld, Walter S. Lasecki, Chiori Hori, Anoop Cherian, Tim K. Marks, Abhinav Rastogi, Xiaoxue Zang, Srinivas Sunkara
IEEE ACM Trans. Audio Speech Lang. Process.14
2020 Improving Crowd-Supported GUI Testing with Structural Guidance
abstract
Crowd testing is an emerging practice in Graphical User Interface (GUI) testing, where developers recruit a large number of crowd testers to test GUI features. It is often easier and faster than a dedicated quality assurance team, and its output is more realistic than that of automated testing. However, crowds of testers working in parallel tend to focus on a small set of commonly-used User Interface (UI) navigation paths, which can lead to low test coverage and redundant effort. In this paper, we introduce two techniques to increase crowd testers' coverage: interactive event-flow graphs and GUI-level guidance. The interactive event-flow graphs track and aggregate every tester's interactions into a single directed graph that visualizes the cases that have already been explored. Crowd testers can interact with the graphs to find new navigation paths and increase the coverage of the created tests. We also use the graphs to augment the GUI (GUI-level guidance) to help testers avoid only exploring common paths. Our evaluation with 30 crowd testers on 11 different test pages shows that the techniques can help testers avoid redundant effort while also increasing untrained testers' coverage by 55%. These techniques can help us develop more robust software that works in more mission-critical settings not only by performing more thorough testing with the same effort that has been put in before but also by integrating them into different parts of the development pipeline to make more reliable software in the early development stage.
Yan Chen 0033, Maulishree Pandey, Jean Y. Song, Walter S. Lasecki, Steve Oney
CHI4
2020 Crowdsourced Detection of Emotionally Manipulative Language
abstract
Detecting rhetoric that manipulates readers' emotions requires distinguishing intrinsically emotional content (IEC; e.g., a parent losing a child) from emotionally manipulative language (EML; e.g., using fear-inducing language to spread anti-vaccine propaganda). However, this remains an open classification challenge for both automatic and crowdsourcing approaches. Machine Learning approaches only work in narrow domains where labeled training data is available, and non-expert annotators tend to conflate IEC with EML. We introduce an approach, anchor comparison, that leverages workers' ability to identify and remove instances of EML in text to create a paraphrased "anchor text", which is then used as a comparison point to classify EML in the original content. We evaluate our approach with a dataset of news-style text snippets and show that precision and recall can be tuned for system builders' needs. Our contribution is a crowdsourcing approach that enables non-expert disentanglement of social references from content.
Jordan S. Huffaker, Jonathan K. Kummerfeld, Walter S. Lasecki, Mark S. Ackerman
CHI3
2020 An Experimental Study of Bias in Platform Worker Ratings: The Role of Performance Quality and Gender
abstract
We study how the ratings people receive on online labor platforms are influenced by their performance, gender, their rater's gender, and displayed ratings from other raters. We conducted a deception study in which participants collaborated on a task with a pair of simulated workers, who varied in gender and performance level, and then rated their performance. When the performance of paired workers was similar, low-performing females were rated lower than their male counterparts. Where there was a clear performance difference between paired workers, low-performing females were preferred over a similarly-performing male peer. Furthermore, displaying an average rating from other raters made ratings more extreme, resulting in high performing workers receiving significantly higher ratings and low performers lower ratings compared to when average ratings were absent. This work contributes an empirical understanding of when biases in ratings manifest, and offers recommendations for how online work platforms can counter these biases.
Farnaz Jahanbakhsh, Justin Cranshaw, Scott Counts, Walter S. Lasecki, Kori Inkpen
CHI4
2020 Using affordances to improve AI support of social media posting decisions
abstract
Intelligent systems are limited in their ability to match the fluid social needs of people. We use affordances---people's perceptions of the utilities of a target system---as a means of creating models that provide intelligent systems with a better understanding of how people make decisions. We study affordance-based models in the context of social network site (SNS) usage, a domain where people have complex social needs often poorly supported by technology. Using data collected via a scenario-based survey (N=674), we build two affordance-based models about people's multi-SNS posting behavior. Our results highlight the feasibility of using affordances to help intelligent systems support people's decision-making behavior: both of our models are ~15% more accurate than a majority-class baseline, and they are ~33% and ~48% more accurate than a random baseline for this task. We contrast our approach with other ways of modeling posting behavior and discuss the implications of using affordances for modeling human behavior for intelligent systems.
Harmanpreet Kaur, Cliff Lampe, Walter S. Lasecki
IUI3
2020 Sifter: A Hybrid Workflow for Theme-based Video Curation at Scale
abstract
User-generated content platforms curate their vast repositories into thematic compilations that facilitate the discovery of high-quality material. Platforms that seek tight editorial control employ people to do this curation, but this process involves time-consuming routine tasks, such as sifting through thousands of videos. We introduce Sifter, a system that improves the curation process by combining automated techniques with a human-powered pipeline that browses, selects, and reaches an agreement on what videos to include in a compilation. We evaluated Sifter by creating 12 compilations from over 34,000 user-generated videos. Sifter was more than three times faster than dedicated curators, and its output was of comparable quality. We reflect on the challenges and opportunities introduced by Sifter to inform the design of content curation systems that need subjective human judgments of videos at scale.
Yan Chen 0033, Andrés Monroy-Hernández, Ian Wehrman, Steve Oney, Walter S. Lasecki, Rajan Vaish
IMX5
2020 Bashon: A Hybrid Crowd-Machine Workflow for Shell Command Synthesis
abstract
Despite advances in machine learning, there has been little progress towards creating automated systems that can reliably solve general purpose tasks, such as programming or scripting. In this paper, we propose techniques for increasing the reliability of automated systems for program synthesis tasks via a hybrid workflow that augments the system with input from crowds of human workers. Unlike previous hybrid workflow systems, which have been focused on less complex tasks that crowd workers can do in their entirety (e.g., image labeling), our proposed workflow handles tasks that untrained crowd workers cannot do alone (i.e., scripting). We evaluate our approach by creating BashOn, a system that increases the performance of an automated program that generates Bash shell commands from natural language descriptions by ~30%. Our approach can not only help people make program synthesis tools more robust, reliable, and trustworthy for end-users to use, but also help lower the cost of downstream data collection for program synthesis when a preliminary model exists.
Yan Chen 0033, Jaylin Herskovitz, Walter S. Lasecki, Steve Oney
VL/HCC3
2020 EdCode: Towards Personalized Support at Scale for Remote Assistance in CS Education
abstract
Programming support methods, like discussion fo-rums and office hours, are important in CS education, but difficult to scale. In this paper, we introduce EdCode, a system that allows students to seek remote instructional support within their IDE in a way that resembles in-person support. It also allows instructors to provide contextualized responses by referencing students' code, and curate and publish their answers for an entire class by selecting only the relevant part of the code referenced, thereby helping to avoid plagiarism. We evaluated EdCode with a series of usability studies and identified benefits and challenges for its use in programming courses. Students found that the perceived quality of support from EdCode was comparable to that of support from in-person office hours, and both students and instructors found publishing and viewing other students' answers helpful.
Yan Chen 0033, Jaylin Herskovitz, Gabriel Matute, April Yi Wang, Sang Won Lee 0002, Walter S. Lasecki, Steve Oney
VL/HCC6
2020 Towards Hybrid Human-AI Workflows for Unknown Unknown Detection
abstract
Predictive models are susceptible to errors called unknown unknowns, in which the model assigns incorrect labels to instances with high confidence. These commonly arise when training data does not represent variations of a class encountered at model deployment. Prior work showed that crowd workers can identify instances of unknown unknowns, but asking the crowd to identify a sufficient number of individual instances can be costly to acquire [2]. Instead, this paper presents an approach that leverages people’s ability to find patterns to retrain classifiers more effectively with fewer examples. We ask crowd workers to suggest and verify patterns in unknown unknowns. We then use these patterns to train an expansion classifier to identify additional examples from existing data that the primary classifier has encountered (and potentially misclassified) in the past. Our experiments show that our approach outperforms existing unknown unknown detection methods at improving classifier performance. This work is the first to leverage crowds to identify error patterns in large datasets to improve ML training.
Anthony Z. Liu, Santiago Guerra, Isaac Fung, Gabriel Matute, Ece Kamar, Walter S. Lasecki
WWW6
2020 Towards Supporting Programming Education at Scale via Live Streaming
abstract
Live streaming, which allows streamers to broadcast their work to live viewers, is an emerging practice for teaching and learning computer programming. Participation in live streaming is growing rapidly, despite several apparent challenges, such as a general lack of training in pedagogy among streamers and scarce signals about a stream's characteristics (e.g., difficulty, style, and usefulness) to help viewers decide what to watch. To understand why people choose to participate in live streaming for teaching or learning programming, and how they cope with both apparent and non-obvious challenges, we interviewed 14 streamers and viewers about their experience with live streaming programming. Among other results, we found that the casual and impromptu nature of live streaming makes it easier to prepare than pre-recorded videos, and viewers have the opportunity to shape the content and learning experience via real-time communication with both the streamer and each other. Nonetheless, we identified several challenges that limit the potential of live streaming as a learning medium. For example, streamers voiced privacy and harassment concerns, and existing streaming platforms do not adequately support viewer-streamer interactions, adaptive learning, and discovery and selection of streaming content. Based on these findings, we suggest specialized tools to facilitate knowledge sharing among people teaching and learning computer programming online, and we offer design recommendations that promote a healthy, safe, and engaging learning environment.
Yan Chen 0033, Walter S. Lasecki
Proc. ACM Hum. Comput. Interact.2
2020 C-Reference: Improving 2D to 3D Object Pose Estimation Accuracy via Crowdsourced Joint Object Estimation
abstract
Converting widely-available 2D images and videos, captured using an RGB camera, to 3D can help accelerate the training of machine learning systems in spatial reasoning domains ranging from in-home assistive robots to augmented reality to autonomous vehicles. However, automating this task is challenging because it requires not only accurately estimating object location and orientation, but also requires knowing currently unknown camera properties (e.g., focal length). A scalable way to combat this problem is to leverage people's spatial understanding of scenes by crowdsourcing visual annotations of 3D object properties. Unfortunately, getting people to directly estimate 3D properties reliably is difficult due to the limitations of image resolution, human motor accuracy, and people's 3D perception (i.e., humans do not "see" depth like a laser range finder). In this paper, we propose a crowd-machine hybrid approach that jointly uses crowds' approximate measurements of multiple in-scene objects to estimate the 3D state of a single target object. Our approach can generate accurate estimates of the target object by combining heterogeneous knowledge from multiple contributors regarding various different objects that share a spatial relationship with the target object. We evaluate our joint object estimation approach with 363 crowd workers and show that our method can reduce errors in the target object's 3D location estimation by over 40%, while requiring only $35$% as much human time. Our work introduces a novel way to enable groups of people with different perspectives and knowledge to achieve more accurate collective performance on challenging visual annotation tasks.
Jean Y. Song, John Joon Young Chung, David F. Fouhey, Walter S. Lasecki
Proc. ACM Hum. Comput. Interact.4
2020 FourEyes: Leveraging Tool Diversity as a Means to Improve Aggregate Accuracy in Crowdsourcing
abstract
Crowdsourcing is a common means of collecting image segmentation training data for use in a variety of computer vision applications. However, designing accurate crowd-powered image segmentation systems is challenging, because defining object boundaries in an image requires significant fine motor skills and hand-eye coordination, which makes these tasks error-prone. Typically, special segmentation tools are created and then answers from multiple workers are aggregated to generate more accurate results. However, individual tool designs can bias how and where people make mistakes, resulting in shared errors that remain even after aggregation. In this article, we introduce a novel crowdsourcing approach that leverages tool diversity as a means of improving aggregate crowd performance. Our idea is that given a diverse set of tools, answer aggregation done across tools can help improve the collective performance by offsetting systematic biases induced by the individual tools themselves. To demonstrate the effectiveness of the proposed approach, we design four different tools and present FourEyes, a crowd-powered image segmentation system that uses aggregation across different tools. We then conduct a series of studies that evaluate different aggregation conditions and show that using multiple tools can significantly improve aggregate accuracy. Furthermore, we investigate the idea of applying post-processing for multi-tool aggregation in terms of correction mechanism. We introduce a novel region-based method for synthesizing more accurate bounds for image segmentation tasks through averaging surrounding annotations. In addition, we explore the effect of adjusting the threshold parameter of an EM-based aggregation method. Our results suggest that not only the individual tool’s design, but also the correction mechanism, can affect the performance of multi-tool aggregation. This article extends a work presented at ACM IUI 2018 [46] by providing a novel region-based error-correction method and additional in-depth evaluation of the proposed approach.
Jean Y. Song, Raymond Fok, Juho Kim 0001, Walter S. Lasecki
ACM Trans. Interact. Intell. Syst.4
2019 Updates in Human-AI Teams: Understanding and Addressing the Performance/Compatibility Tradeoff
abstract
AI systems are being deployed to support human decision making in high-stakes domains such as healthcare and criminal justice. In many cases, the human and AI form a team, in which the human makes decisions after reviewing the AI’s inferences. A successful partnership requires that the human develops insights into the performance of the AI system, including its failures. We study the influence of updates to an AI system in this setting. While updates can increase the AI’s predictive performance, they may also lead to behavioral changes that are at odds with the user’s prior experiences and confidence in the AI’s inferences. We show that updates that increase AI performance may actually hurt team performance. We introduce the notion of the compatibility of an AI update with prior user experience and present methods for studying the role of compatibility in human-AI teams. Empirical results on three high-stakes classification tasks show that current machine learning algorithms do not produce compatible updates. We propose a re-training objective to improve the compatibility of an update by penalizing new errors. The objective offers full leverage of the performance/compatibility tradeoff across different datasets, enabling more compatible yet accurate updates.
Gagan Bansal, Besmira Nushi, Ece Kamar, Daniel S. Weld, Walter S. Lasecki, Eric Horvitz
AAAI5
2019 A Large-Scale Corpus for Conversation Disentanglement
abstract
Jonathan K. Kummerfeld, Sai R. Gouravajhala, Joseph J. Peper, Vignesh Athreya, Chulaka Gunasekara, Jatin Ganhotra, Siva Sankalp Patel, Lazaros C Polymenakos, Walter Lasecki. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019.
Jonathan K. Kummerfeld, Sai R. Gouravajhala, Joseph Peper, Vignesh Athreya, R. Chulaka Gunasekara, Jatin Ganhotra, Siva Sankalp Patel, Lazaros Polymenakos, Walter S. Lasecki
ACL (1)9
2019 The Effect of Social Interaction on Facilitating Audience Participation in a Live Music Performance
abstract
Facilitating audience participation in a music performance brings with it challenges in involving non-expert users in large-scale collaboration. A musical piece needs to be created live, over a short period of time, with limited communication channels. To address this challenge, we propose to incorporate social interaction through mobile music instruments that the audience is given to play with, and examine how this feature sustains and affects the audience involvement. We test this idea with an audience participation music system, Crowd in C. We realized a participation-based musical performance with the system and validated our approach by analyzing the interaction traces of the audience at a performance. The result indicates that the audience members were actively engaged throughout the performance, with multiple layers of social interaction available in the system. We also present how the social interactivity among the audience shaped their interaction in the music making process.
Sang Won Lee 0002, Aaron Willette, Danai Koutra, Walter S. Lasecki
Creativity & Cognition4
2019 CoSQL: A Conversational Text-to-SQL Challenge Towards Cross-Domain Natural Language Interfaces to Databases
abstract
Tao Yu, Rui Zhang, Heyang Er, Suyi Li, Eric Xue, Bo Pang, Xi Victoria Lin, Yi Chern Tan, Tianze Shi, Zihan Li, Youxuan Jiang, Michihiro Yasunaga, Sungrok Shim, Tao Chen, Alexander Fabbri, Zifan Li, Luyao Chen, Yuwen Zhang, Shreya Dixit, Vincent Zhang, Caiming Xiong, Richard Socher, Walter Lasecki, Dragomir Radev. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Tao Yu 0009, Rui Zhang 0037, Heyang Er, Suyi Li 0002, Eric Xue 0001, Bo Pang 0004, Xi Victoria Lin, Yi Chern Tan, Tianze Shi, Youxuan Jiang, Michihiro Yasunaga, Sungrok Shim, Alexander R. Fabbri, Zifan Li, Shreya Dixit, Caiming Xiong, Richard Socher, Walter S. Lasecki, Dragomir R. Radev
EMNLP/IJCNLP (1)23
2019 Beyond Accuracy: The Role of Mental Models in Human-AI Team Performance
abstract
Decisions made by human-AI teams (e.g., AI-advised humans) are increasingly common in high-stakes domains such as healthcare, criminal justice, and finance. Achieving high team performance depends on more than just the accuracy of the AI system: Since the human and the AI may have different expertise, the highest team performance is often reached when they both know how and when to complement one another. We focus on a factor that is crucial to supporting such complementary: the human’s mental model of the AI capabilities, specifically the AI system’s error boundary (i.e. knowing “When does the AI err?”). Awareness of this lets the human decide when to accept or override the AI’s recommendation. We highlight two key properties of an AI’s error boundary, parsimony and stochasticity, and a property of the task, dimensionality. We show experimentally how these properties affect humans’ mental models of AI capabilities and the resulting team performance. We connect our evaluations to related work and propose goals, beyond accuracy, that merit consideration during model selection and optimization to improve overall human-AI team performance.
Gagan Bansal, Besmira Nushi, Ece Kamar, Walter S. Lasecki, Daniel S. Weld, Eric Horvitz
HCOMP4
2019 Popup: reconstructing 3D video using particle filtering to aggregate crowd responses
abstract
Collecting a sufficient amount of 3D training data for autonomous vehicles to handle rare, but critical, traffic events (e.g., collisions) may take decades of deployment. Abundant video data of such events from municipal traffic cameras and video sharing sites (e.g., YouTube) could provide a potential alternative, but generating realistic training data in the form of 3D video reconstructions is a challenging task beyond the current capabilities of computer vision. Crowdsourcing the annotation of necessary information could bridge this gap, but the level of accuracy required to obtain usable reconstructions makes this task nearly impossible for non-experts. In this paper, we propose a novel hybrid intelligence method that combines annotations from workers viewing different instances (video frames) of the same target (3D object), and uses particle filtering to aggregate responses. Our approach can leveraging temporal dependencies between video frames, enabling higher quality through more aggressive filtering. The proposed method results in a 33% reduction in the relative error of position estimation compared to a state-of-the-art baseline. Moreover, our method enables skipping (self-filtering) challenging annotations, reducing the total annotation time for hard-to-annotate frames by 16%. Our approach provides a generalizable means of aggregating more accurate crowd responses in settings where annotation is especially challenging or error-prone.
Jean Y. Song, Stephan J. Lemmer, Michael Xieyang Liu, Shiyan Yan, Juho Kim 0001, Jason J. Corso, Walter S. Lasecki
IUI7
2019 Efficient Elicitation Approaches to Estimate Collective Crowd Answers
abstract
When crowdsourcing the creation of machine learning datasets, statistical distributions that capture diverse answers can represent ambiguous data better than a single best answer. Unfortunately, collecting distributions is expensive because a large number of responses need to be collected to form a stable distribution. Despite this, the efficient collection of answer distributions-that is, ways to use less human effort to collect estimates of the eventual distribution that would be formed by a large group of responses-is an under-studied topic. In this paper, we demonstrate that this type of estimation is possible and characterize different elicitation approaches to guide the development of future systems. We investigate eight elicitation approaches along two dimensions: annotation granularity and estimation perspective. Annotation granularity is varied by annotating i) a single "best" label, ii) all relevant labels, iii) a ranking of all relevant labels, or iv) real-valued weights for all relevant labels. Estimation perspective is varied by prompting workers to either respond with their own answer or an estimate of the answer(s) that they expect other workers would provide. Our study collected ordinal annotations on the emotional valence of facial images from 1,960 crowd workers and found that, surprisingly, the most fine-grained elicitation methods were not the most accurate, despite workers spending more time to provide answers. Instead, the most efficient approach was to ask workers to choose all relevant classes that others would have selected. This resulted in a 21.4% reduction in the human time required to reach the same performance as the baseline (i.e., selecting a single answer with their own perspective). By analyzing cases in which finer-grained annotations degraded performance, we contribute to a better understanding of the trade-offs between answer elicitation approaches. Our work makes it more tractable to use answer distributions in large-scale tasks such as ML training, and aims to spark future work on techniques that can efficiently estimate answer distributions.
John Joon Young Chung, Jean Y. Song, Sindhu Kutty, Sungsoo Ray Hong, Juho Kim 0001, Walter S. Lasecki
Proc. ACM Hum. Comput. Interact.6
2018 Towards More Robust Speech Interactions for Deaf and Hard of Hearing Users
abstract
Mobile, wearable, and other ubiquitous computing devices are increasingly creating a context in which conventional keyboard and screen-based inputs are being replaced in favor of more natural speech-based interactions. Digital personal assistants use speech to control a wide range of functionality, from environmental controls to information access. However, many deaf and hard-of-hearing users have speech patterns that vary from those of hearing users due to incomplete acoustic feedback from their own voices. Because automatic speech recognition (ASR) systems are largely trained using speech from hearing individuals, speech-controlled technologies are typically inaccessible to deaf users. Prior work has focused on providing deaf users access to aural output via real-time captioning or signing, but little has been done to improve users' ability to provide input to these systems' speech-based interfaces. Further, the vocalization patterns of deaf speech often make accurate recognition intractable for both automated systems and human listeners, making traditional approaches to mitigate ASR limitations, such as human captionists, less effective. To bridge this accessibility gap, we investigate the limitations of common speech recognition approaches and techniques---both automatic and human-powered---when applied to deaf speech. We then explore the effectiveness of an iterative crowdsourcing workflow, and characterize the potential for groups to collectively exceed the performance of individuals. This paper contributes a better understanding of the challenges of deaf speech recognition and provides insights for future system development in this space.
Raymond Fok, Harmanpreet Kaur, Skanda Palani, Martez E. Mott, Walter S. Lasecki
ASSETS5
2018 Bolt: Instantaneous Crowdsourcing via Just-in-Time Training
abstract
Real-time crowdsourcing has made it possible to solve problems that are beyond the scope of artificial intelligence (AI) within a matter of seconds, rather than hours or days with traditional crowdsourcing techniques. While this has led to an increase in the potential application domains of crowdsourcing and human computation, problems that require machine-level speeds---on the order of milliseconds, not seconds---have remained out of reach because of the fundamental bounds of human perception and response time. In this paper, we demonstrate that it is possible to exceed these bounds by combining human and machine intelligence. We introduce the look-ahead approach, a hybrid intelligence workflow that enables instantaneous crowdsourcing systems (i.e., those that can return crowd responses within mere milliseconds). The look-ahead approach works by exploring possible future states that may be encountered within a short time horizon (e.g., a few seconds into the future) and prefetching crowd worker responses to these states. We validate the efficacy and explore the limitations of our approach on the Bolt system, which consists of an arcade-style game (Lightning Dodger) that we formally model as a Markov Decision Process (MDP). When the MDP reward function is unspecified---as in many real-world tasks---the look-ahead approach enables just-in-time (JIT) training of the agent's policy function. Through a series of crowd worker experiments, we demonstrate that the look-ahead approach can outperform the fastest individual worker by approximately two orders of magnitude. Our work opens new avenues for hybrid intelligence systems that are as smart as people, but also far faster than humanly possible.
Alan Lundgard, Yiwei Yang 0004, Maya L. Foster, Walter S. Lasecki
CHI4
2018 Enhancing Online Problems Through Instructor-Centered Tools for Randomized Experiments
abstract
Digital educational resources could enable the use of randomized experiments to answer pedagogical questions that instructors care about, taking academic research out of the laboratory and into the classroom. We take an instructor-centered approach to designing tools for experimentation that lower the barriers for instructors to conduct experiments. We explore this approach through DynamicProblem, a proof-of-concept system for experimentation on components of digital problems, which provides interfaces for authoring of experiments on explanations, hints, feedback messages, and learning tips. To rapidly turn data from experiments into practical improvements, the system uses an interpretable machine learning algorithm to analyze students' ratings of which conditions are helpful, and present conditions to future students in proportion to the evidence they are higher rated. We evaluated the system by collaboratively deploying experiments in the courses of three mathematics instructors. They reported benefits in reflecting on their pedagogy, and having a new method for improving online problems for future students.
Joseph Jay Williams, Anna N. Rafferty, Dustin Tingley, Andrew M. Ang, Walter S. Lasecki, Juho Kim 0001
CHI5
2018 EURECA: Enhanced Understanding of Real Environments via Crowd Assistance
abstract
Indoor robots hold the promise of automatically handling mundane daily tasks, helping to improve access for people with disabilities, and providing on-demand access to remote physical environments. Unfortunately, the ability to understand never-before-seen objects in scenes where new items may be added (e.g., purchased) or altered (e.g., damaged) on a regular basis remains an open challenge for robotics. In this paper, we introduce EURECA, a mixed-initiative system that leverages online crowds of human contributors to help robots robustly identify 3D point cloud segments corresponding to user-referenced objects in near real-time. EURECA allows robots to understand multi-object 3D scenes on-the-fly (in ~40 seconds) by providing groups of non-expert crowd workers with intelligent tools that can segment objects more quickly (~70% faster) and more accurately (6% higher F1 score) than individuals. More broadly, EURECA introduces the first real-time crowdsourcing tool that addresses the challenge of learning about new objects in real-world settings, creating a new source of data for training robots online, as well as a platform for studying mixed-initiative crowdsourcing workflows for understanding 3D scenes.
Sai R. Gouravajhala, Jinyeong Yim, Karthik Desingh, Yanda Huang, Odest Chadwicke Jenkins, Walter S. Lasecki
HCOMP6
2018 Plexiglass: Multiplexing Passive and Active Tasks for More Efficient Crowdsourcing
abstract
Efficiently scaling continuous real-time crowdsourcing tasks — which engage crowd workers over long periods of time to complete tasks, such as monitoring video for critical events — is challenging largely because of the cost of keeping people consistently engaged. Worse, for many continuous tasks on which progress cannot be immediately made, which we term passive tasks, this engagement effort is wasted until something becomes true about the environment (e.g., the annotation of a critical event cannot happen until the event is seen). In this paper, we present the idea of passive-active task multiplexing, in which continuous tasks that involve waiting for a change in state are completed concurrently with traditional (offline/non-real time) tasks. We then implement this idea in Plexiglass, a system that answers visual queries made by blind and low-vision users more efficiently by concurrently presenting a (passive) real-time sensing task and an (active) offline visual question answering tasks to workers. We explore different approaches to accomplishing this, and validate that task multiplexing can lead to an improvement in efficiency in these settings of more than 40% in terms of overall worker time taken and 85% decrease in cost of running continuous real-time tasks. This work has implications on how real-time crowdsourcing tasks are presented to workers, and it increases the feasibility of deploying continuous real-time crowdsourcing systems in real-world settings.
Akshay Rao, Harmanpreet Kaur, Walter S. Lasecki
HCOMP3
2018 Two Tools are Better Than One: Tool Diversity as a Means of Improving Aggregate Crowd Performance
abstract
Crowdsourcing is a common means of collecting image segmentation training data for use in a variety of computer vision applications. However, designing accurate crowd-powered image segmentation systems is challenging because defining object boundaries in an image requires significant fine motor skills and hand-eye coordination, which makes these tasks error-prone. Typically, special segmentation tools are created and then answers from multiple workers are aggregated to generate more accurate results. However, individual tool designs can bias how and where people make mistakes, resulting in shared errors that remain even after aggregation. In this paper, we introduce a novel crowdsourcing workflow that leverages multiple tools for the same task to increase output accuracy by reducing systematic error biases introduced by the tools themselves. When a task can no longer be broken down into more-tractable subtasks (the conventional approach taken by microtask crowdsourcing), our multi-tool approach can be used to further improve accuracy by assigning different tools to different workers. We present a series of studies that evaluate our multi-tool approach and show that it can significantly improve aggregate accuracy in semantic image segmentation.
Jean Y. Song, Raymond Fok, Alan Lundgard, Juho Kim 0001, Walter S. Lasecki
IUI6
2018 Arboretum and Arbility: Improving Web Accessibility Through a Shared Browsing Architecture
abstract
Many web pages developed today require navigation by visual interaction-seeing, hovering, pointing, clicking, and dragging with the mouse over dynamic page content. These forms of interaction are increasingly popular as developer trends have moved from static, logically structured pages to dynamic, interactive pages. However, they are also often inaccessible to blind web users who tend to rely on keyboard-based screen readers to navigate the web. Despite existing web accessibility standards, engineering web pages to be equally accessible via both keyboard and visuomotor mouse-based interactions is often not a priority for developers. Improving access to this kind of visual and interactive web content has been a long-standing goal of HCI researchers, but the barriers have proven to be too varied and unpredictable to be overcome by some of the proposed solutions: promoting guidelines and best practices, automatically generating accessible versions of pre-exisiting web pages, or developing human-assisted solutions, such as screen and cursor-sharing, which tend to diminish an end user's agency. In this paper we present a real-time, collaborative approach to helping blind web users overcome inaccessible parts of existing web pages. We introduce *Arboretum*, a new architecture that enables any web user to seamlessly hand off controlled parts of their browsing session to remote users, while maintaining control over the interface via a "propose and accept/reject" mechanism. We illustrate the benefit of Arboretum by using it to implement *Arbility*, a browser that allows blind users to hand off targeted visual interaction tasks to remote crowd workers. We evaluate the entire system in a study with 9 blind web users, showing that Arbility allows them to interact with web content that was previously difficult to access via a screen reader alone.
Steve Oney, Alan Lundgard, Rebecca Krosnick, Michael Nebeling, Walter S. Lasecki
UIST5
2018 Expresso: Building Responsive Interfaces with Keyframes
abstract
Web developers use responsive web design to create user interfaces that can adapt to many form factors. To define responsive pages, developers must use Cascading Style Sheets (CSs) or libraries and tools built on top of it. CSS provides high customizability, but requires significant experience. As a result, non-programmers and novice programmers generally lack a means of easily building custom responsive web pages. In this paper, we present a new approach that allows users to create custom responsive user interfaces without writing program code. We demonstrate the feasibility and effectiveness of the approach through a new system we built, named Expresso. With Expresso, users define “keyframes” - examples of how their VI should look for particular viewport sizes - by simply directly manipulating elements in a WYSIWYG editor. Expresso uses these keyframes to infer rules about the responsive behavior of elements, and automatically renders the appropriate css for a given viewport size. To allow users to create the desired appearance of their page at all viewport sizes, Expresso lets users define either a “smooth” or “jump” transition between adjacent keyframes. We conduct a user study and show that participants are able to effectively use Expresso to build realistic responsive interfaces.
Rebecca Krosnick, Sang Won Lee 0002, Walter S. Lasecki, Steve Oney
VL/HCC3
2018 Exploring Real-Time Collaboration in Crowd-Powered Systems Through a UI Design Tool
abstract
Real-time collaboration between a requester and crowd workers expands the scope of tasks that crowdsourcing can be used for by letting requesters and crowd workers interactively create various artifacts (e.g., a sketch prototype, writing, or program code). In such systems, it is increasingly common to allow requesters to verbally describe their requests, receive responses from workers, and provide immediate and continuous feedback to enhance the overall outcome of the two groups' real-time collaboration. This work is motivated by the lack of a deep understanding of the challenges that end users of such systems face in their communication with workers and the need of design implications that can address such challenges for other similar systems. In this paper, we investigate how requesters verbally communicate and collaborate with crowd workers to solve a complex task. Using a crowd-powered UI design tool, we conducted a qualitative user study to explore how requesters with varying expertise communicate and collaborate with crowd workers. Our work also identifies the unique challenges that collaborative crowdsourcing systems pose: potential expertise differences between requesters and crowd workers, the asymmetry of two-way communication (e.g., speech versus text), and the shared artifact's concurrent modification by two disparate groups. Finally, we make design recommendations that can inform the design of future real-time collaboration processes in crowdsourcing systems.
Sang Won Lee 0002, Rebecca Krosnick, Brandon Keelean, Sach Vaidya, Stephanie D. O'Keefe, Walter S. Lasecki
Proc. ACM Hum. Comput. Interact.7
2018 Creating Better Action Plans for Writing Tasks via Vocabulary-Based Planning
abstract
While having a step-by-step breakdown for a task-an action plan-helps people complete tasks, prior work has shown that people prefer not to make action plans for their own tasks. Getting planning support from others could be beneficial, but it is limited by how much domain knowledge people have about the task and how available they are. Our goal is to incorporate the benefits of having action plans in the complex domain of writing, while mitigating the time and effort costs of creating plans. To mitigate these costs, we introduce a vocabulary-a finite set of functions pertaining to writing tasks-as a cognitive scaffold that enables people with necessary context (e.g. collaborators) to generate action plans for others. We develop this vocabulary by analyzing 264 comments, and compare plans created using it with those created without any aid, in an online study with 768 comments (N=145) and a lab study with 96 comments (N=8). We show that using a vocabulary reduces planning time and effort and improves plan quality compared to unstructured planning, and opens the door for automation and task sharing for complex tasks.
Harmanpreet Kaur, Alex C. Williams, Anne Loomis Thompson, Walter S. Lasecki, Shamsi T. Iqbal, Jaime Teevan
Proc. ACM Hum. Comput. Interact.4
2017 Codeon: On-Demand Software Development Assistance
abstract
Software developers rely on support from a variety of resources---including other developers---but the coordination cost of finding another developer with relevant experience, explaining the context of the problem, composing a specific help request, and providing access to relevant code is prohibitively high for all but the largest of tasks. Existing technologies for synchronous communication (e.g. voice chat) have high scheduling costs, and asynchronous communication tools (e.g. forums) require developers to carefully describe their code context to yield useful responses. This paper introduces Codeon, a system that enables more effective task hand-off between end-user developers and remote helpers by allowing asynchronous responses to on-demand requests. With Codeon, developers can request help by speaking their requests aloud within the context of their IDE. Codeon automatically captures the relevant code context and allows remote helpers to respond with high-level descriptions, code annotations, code snippets, and natural language explanations. Developers can then immediately view and integrate these responses into their code. In this paper, we describe Codeon, the studies that guided its design, and our evaluation that its effectiveness as a support tool. In our evaluation, developers using Codeon completed nearly twice as many tasks as those who used state-of-the-art synchronous video and code sharing tools, by reducing the coordination costs of seeking assistance from other developers.
Yan Chen 0033, Sang Won Lee 0002, Yin Xie, Yiwei Yang 0004, Walter S. Lasecki, Steve Oney
CHI5
2017 Kurator: Using The Crowd to Help Families With Personal Curation Tasks
abstract
People capture photos, audio recordings, video, and more on a daily basis, but organizing all these digital artifacts quickly becomes a daunting task. Automated solutions struggle to help us manage this data because they cannot understand its meaning. In this paper, we introduce Kurator, a hybrid intelligence system leveraging mixed-expertise crowds to help families curate their personal digital content. Kurator produces a refined set of content via a combination of automated systems able to scale to large data sets and human crowds able to understand the data. Our results with 5 families show that Kurator can reduce the amount of effort needed to find meaningful memories within a large collection. This work also suggests that crowdsourcing can be used effectively even in domains where personal preference is key to accurately solving the task.
David Merritt, Jasmine Jones, Mark S. Ackerman, Walter S. Lasecki
CSCW4
2017 CrowdMask: Using Crowds to Preserve Privacy in Crowd-Powered Systems via Progressive Filtering
abstract
Crowd-powered systems leverage human intelligence to go beyond the capabilities of automated systems, but also introduce privacy and security concerns because unknown people must view the data that the system processes. While automated approaches cannot robustly filter private information from these datasets, people have the ability to do so if the risk from them viewing the data can be mitigated. We present a crowd-powered approach to masking private content in data by segmenting and distributing smaller segments to crowd workers so that individual workers can identify potentially private content without being able to fully view it themselves. We introduce a novel pyramid workflow for segmentation that uses segments at multiple levels of granularity to overcome problems with fixed-sized approaches. We implement our approach in CrowdMask, a system that allows images with potentially sensitive content to be masked by appearing in progressively larger, more identifiable segments, and masking portions of the image as soon as a risk is identified. Our experiments with 4134 Mechanical Turk workers show that CrowdMask can effectively mask private content from images without revealing sensitive content to constituent workers, while still enabling future systems to use the filtered result.
Harmanpreet Kaur, Mitchell L. Gordon, Yiwei Yang 0004, Jeffrey P. Bigham, Jaime Teevan, Ece Kamar, Walter S. Lasecki
HCOMP7
2017 SketchExpress: Remixing Animations for More Effective Crowd-Powered Prototyping of Interactive Interfaces
abstract
Low-fidelity prototyping at the early stages of user interface (UI) design can help designers and system builders quickly explore their ideas. However, interactive behaviors in such prototypes are often replaced by textual descriptions because it usually takes even professionals hours or days to create animated interactive elements due to the complexity of creating them. In this paper, we introduce SketchExpress, a crowd-powered prototyping tool that enables crowd workers to create reusable interactive behaviors easily and accurately. With the system, a requester-designers or end-users-describes aloud how an interface should behave and crowd workers make the sketched prototype interactive within minutes using a demonstrate-remix-replay approach. These behaviors are manually demonstrated, refined using remix functions, and then can be replayed later. The recorded behaviors persist for future reuse to help users communicate with the animated prototype. We conducted a study with crowd workers recruited from Mechanical Turk, which demonstrated that workers could create animations using SketchExpress in 2.9 minutes on average with 27% gain in the quality of animations compared to the baseline condition of manual demonstration.
Sang Won Lee 0002, Isabelle Wong, Yiwei Yang 0004, Stephanie D. O'Keefe, Walter S. Lasecki
UIST6
2017 WearMail: On-the-Go Access to Information in Your Email with a Privacy-Preserving Human Computation Workflow
abstract
Email is more than just a communication medium. Email serves as an external memory for people---it contains our reservation numbers, meeting details, phone numbers, and more. Often, people need access to this information while on the go, which is cumbersome from mobile devices with limited I/O bandwidth. In this paper, we introduce WearMail, a conversational interface to retrieve specific information in email. WearMail is mostly automated but is made robust to information extraction tasks via a novel privacy-preserving human computation workflow. In WearMail, crowdworkers never have direct access to emails, but rather (i) generate an email filter to help the system find messages that may contain the desired information, and (ii) generate examples of the requested information that are then used to create custom, low-level information extractors that run automatically within the set of filtered emails. We explore the impact of varying levels of obfuscation on result quality, demonstrating that workers are able to deal with highly-obfuscated information nearly as well as with the original. WearMail introduces general mechanisms that let the crowd search and select private data without having direct access to the data itself.
Sai Swaminathan, Raymond Fok, Ting-Hao 'Kenneth' Huang, Irene Lin, Rohan Jadvani, Walter S. Lasecki, Jeffrey P. Bigham
UIST7
2016 Towards Providing On-Demand Expert Support for Software Developers
abstract
Software development is an expert task that requires complex reasoning and the ability to recall language or API-specific details. In practice, developers often seek support from IDE tools, Web resources, or other developers to help fill in gaps in their knowledge on-demand. In this paper, we present two studies that seek to inform the design of future systems that use remote experts to support developers on demand. The first explores what types of questions developers would ask a hypothetical assistant capable of answering any question they pose. The second study explores the interactions between developers and remote experts in supporting roles. Our results suggest eight key system features needed for on-demand remote developer assistants to be effective, which has implications for future human-powered development tools.
Yan Chen 0033, Steve Oney, Walter S. Lasecki
CHI3
2016 "Is There Anything Else I Can Help You With?" Challenges in Deploying an On-Demand Crowd-Powered Conversational Agent
abstract
Intelligent conversational assistants, such as Apple's Siri, Microsoft's Cortana, and Amazon's Echo, have quickly become a part of our digital life. However, these assistants have major limitations, which prevents users from conversing with them as they would with human dialog partners. This limits our ability to observe how users really want to interact with the underlying system. To address this problem, we developed a crowd-powered conversational assistant, Chorus, and deployed it to see how users and workers would interact together when mediated by the system. Chorus sophisticatedly converses with end users over time by recruiting workers on demand, which in turn decide what might be the best response for each user sentence. Up to the first month of our deployment, 59 users have held conversations with Chorus during 320 conversational sessions. In this paper, we present an account of Chorus' deployment, with a focus on four challenges: (i) identifying when conversations are over, (ii) malicious users and workers, (iii) on-demand recruiting, and (iv) settings in which consensus is not enough. Our observations could assist the deployment of crowd-powered conversation systems and crowd-powered systems in general.
Ting-Hao 'Kenneth' Huang, Walter S. Lasecki, Amos Azaria, Jeffrey P. Bigham
HCOMP2
2016 AXIS: Generating Explanations at Scale with Learnersourcing and Machine Learning
abstract
While explanations may help people learn by providing information about why an answer is correct, many problems on online platforms lack high-quality explanations. This paper presents AXIS (Adaptive eXplanation Improvement System), a system for obtaining explanations. AXIS asks learners to generate, revise, and evaluate explanations as they solve a problem, and then uses machine learning to dynamically determine which explanation to present to a future learner, based on previous learners' collective input. Results from a case study deployment and a randomized experiment demonstrate that AXIS elicits and identifies explanations that learners find helpful. Providing explanations from AXIS also objectively enhanced learning, when compared to the default practice where learners solved problems and received answers without explanations. The rated quality and learning benefit of AXIS explanations did not differ from explanations generated by an experienced instructor.
Joseph Jay Williams, Juho Kim 0001, Anna N. Rafferty, Samuel G. Maldonado, Krzysztof Z. Gajos, Walter S. Lasecki, Neil T. Heffernan
L@S6
2015 Zensors: Adaptive, Rapidly Deployable, Human-Intelligent Sensor Feeds
abstract
The promise of "smart" homes, workplaces, schools, and other environments has long been championed. Unattractive, however, has been the cost to run wires and install sensors. More critically, raw sensor data tends not to align with the types of questions humans wish to ask, e.g., do I need to restock my pantry? Although techniques like computer vision can answer some of these questions, it requires significant effort to build and train appropriate classifiers. Even then, these systems are often brittle, with limited ability to handle new or unexpected situations, including being repositioned and environmental changes (e.g., lighting, furniture, seasons). We propose Zensors, a new sensing approach that fuses real-time human intelligence from online crowd workers with automatic approaches to provide robust, adaptive, and readily deployable intelligent sensors. With Zensors, users can go from question to live sensor feed in less than 60 seconds. Through our API, Zensors can enable a variety of rich end-user applications and moves us closer to the vision of responsive, intelligent environments.
Gierad Laput, Walter S. Lasecki, Jason Wiese, Robert Xiao, Jeffrey P. Bigham, Chris Harrison 0001
CHI2
2015 Exploring Privacy and Accuracy Trade-Offs in Crowdsourced Behavioral Video Coding
abstract
Coding behavioral video is an important method used by researchers to understand social phenomenon. Unfortunately, traditional hand-coding approaches can take days or weeks of time to complete. Recent work has shown that these tasks can be completed quickly by leveraging the parallelism of large online crowds, but using the crowd introduces new concerns about accuracy, reliability, privacy, and cost. To explore these issues, we conducted interviews with 12 researchers who frequently code behavioral video, to investigate common practices and challenges with video coding. We find accuracy and privacy to be the researchers' primary concerns. To explore this more concretely, we used sample videos to investigate whether crowds can accurately recognize instances of commonly coded behaviors, and show that the crowd yields accurate results. Then, we demonstrate a method for obfuscating participant identity with a video blur filter, and find, as expected, that workers' ability to identify participants decreases as blur level increases. The workers' ability to accurately and reliably code behaviors also decreases, but not as steeply as the identity test. This trade-off between coding quality and privacy protection suggests that researchers can use online crowds to code for some key behaviors in video without compromising participant identity. We conclude with a discussion of how researchers can balance privacy and accuracy on their own data using a system we introduce called Incognito.
Walter S. Lasecki, Mitchell L. Gordon, Winnie Leung, Ellen Lim, Jeffrey P. Bigham, Steven Dow
CHI1
2015 Apparition: Crowdsourced User Interfaces that Come to Life as You Sketch Them
abstract
Prototyping allows designers to quickly iterate and gather feedback, but the time it takes to create even a Wizard-of-Oz prototype reduces the utility of the process. In this paper, we introduce crowdsourcing techniques and tools for prototyping interactive systems in the time it takes to describe the idea. Our Apparition system uses paid microtask crowds to make even hard-to-automate functions work immediately, allowing more fluid prototyping of interfaces that contain interactive elements and complex behaviors. As users sketch their interface and describe it aloud in natural language, crowd workers and sketch recognition algorithms translate the input into user interface elements, add animations, and provide Wizard-of-Oz functionality. We discuss how design teams can use our approach to reflect on prototypes or begin user studies within seconds, and how, over time, Apparition prototypes can become fully-implemented versions of the systems they simulate. Powering Apparition is the first self-coordinated, real-time crowdsourcing infrastructure. We anchor this infrastructure on a new, lightweight write-locking mechanism that workers can use to signal their intentions to each other.
Walter S. Lasecki, Juho Kim 0001, Nick Rafter, Onkur Sen, Jeffrey P. Bigham, Michael S. Bernstein
CHI1
2015 The Effects of Sequence and Delay on Crowd Work
abstract
A common approach in crowdsourcing is to break large tasks into small microtasks so that they can be parallelized across many crowd workers and so that redundant work can be more easily compared for quality control. In practice, this can result in the microtasks being presented out of their natural order and often introduces delays between individual microtasks. In this paper, we demonstrate in a study of 338 crowd workers that non-sequential microtasks and the introduction of delays significantly decreases worker performance. We show that interruptions where a large delay occurs between two related tasks can cause up to a 102% slowdown in completion time, and interruptions where workers are asked to perform different tasks in sequence can slow down completion time by 57%. We conclude with a set of design guidelines to improve both worker performance and realized pay, and instructions for implementing these changes in existing interfaces for crowd work.
Walter S. Lasecki, Jeffrey M. Rzeszotarski, Adam Marcus 0002, Jeffrey P. Bigham
CHI1
2015 RegionSpeak: Quick Comprehensive Spatial Descriptions of Complex Images for Blind Users
abstract
Blind people often seek answers to their visual questions from remote sources, however, the commonly adopted single-image, single-response model does not always guarantee enough bandwidth between users and sources. This is especially true when questions concern large sets of information, or spatial layout, e.g., where is there to sit in this area, what tools are on this work bench, or what do the buttons on this machine do? Our RegionSpeak system addresses this problem by providing an accessible way for blind users to (i) combine visual information across multiple photographs via image stitching, em (ii) quickly collect labels from the crowd for all relevant objects contained within the resulting large visual area in parallel, and (iii) then interactively explore the spatial layout of the objects that were labeled. The regions and descriptions are displayed on an accessible touchscreen interface, which allow blind users to interactively explore their spatial layout. We demonstrate that workers from Amazon Mechanical Turk are able to quickly and accurately identify relevant regions, and that asking them to describe only one region at a time results in more comprehensive descriptions of complex images. RegionSpeak can be used to explore the spatial layout of the regions identified. It also demonstrates broad potential for helping blind users to answer difficult spatial layout questions.
Walter S. Lasecki, Erin L. Brady, Jeffrey P. Bigham
CHI2
2015 Guardian: A Crowd-Powered Spoken Dialog System for Web APIs
abstract
Natural language dialog is an important and intuitive way for people to access information and services. However, current dialog systems are limited in scope, brittle to the richness of natural language, and expensive to produce. This paper introduces Guardian, a crowd-powered framework that wraps existing Web APIs into immediately usable spoken dialog systems. Guardian takes as input the Web API and desired task, and the crowd determines the parameters necessary to complete it, how to ask for them, and interprets the responses from the API. The system is structured so that, over time, it can learn to take over for the crowd. This hybrid systems approach will help make dialog systems both more general and more robust going forward.
Ting-Hao 'Kenneth' Huang, Walter S. Lasecki, Jeffrey P. Bigham
HCOMP2
2014 Legion scribe: real-time captioning by non-experts
abstract
The promise of affordable, automatic approaches to real-time captioning imagines a future in which deaf and hard of hearing (DHH) users have immediate access to speech in the world around them my simply picking up their phone or other mobile device. While the challenges of processing highly variable natural language has prevented automated approaches from completing this task reliably enough for use in settings such as classrooms or workplaces [4], recent work in crowd-powered approaches have allowed groups of non-expert captionists to provide a similarly-flexible source of captions for DHH users. This is in contrast to current human-powered approaches, which use highly-trained professional captionists who can type up to 250 words per minute (WPM), but also can cost over $100/hr. In this paper, we describe a real-time demo of Legion:Scribe (or just "Scribe"), a crowd-powered captioning system that allows untrained participants and volunteers to provide reliable captions with less than 5 seconds of latency by computationally merging their input into a single collective answer that is more accurate and more complete than any one worker could have generated alone.
Walter S. Lasecki, Raja S. Kushalnagar, Jeffrey P. Bigham
ASSETS1
2014 Increasing the bandwidth of crowdsourced visual question answering to better support blind users
abstract
Many of the visual questions that blind people ask cannot be easily answered with a single image or a short response, especially when questions are of an exploratory nature, e.g. what is in this area, or what tools are available on this work bench? We introduce RegionSpeak to allow blind users to capture large areas of visual information, identify all of the objects within them, and explore their spatial layout with fewer interactions. RegionSpeak helps blind users capture all of the relevant visual information using an interface designed to support stitching multiple images together. We use a parallel crowdsourcing workflow that asks workers to define and describe regions of interest, allowing even complex images to be described quickly. The regions and descriptions are displayed on an auditory touchscreen interface, allowing users to know what is in a scene and how it is laid out.
Walter S. Lasecki, Jeffrey P. Bigham
ASSETS1
2014 Crowd storage: storing information on existing memories
abstract
This paper introduces the concept of crowd storage, the idea that digital files can be stored and retrieved later from the memories of people in the crowd. Similar to human memory, crowd storage is ephemeral, which means that storage is temporary and the quality of the stored information degrades over time. Crowd storage may be preferred over storing information directly in the cloud, or when it is desirable for information to degrade inline with normal human memories. To explore and validate this idea, we created WeStore, a system that stores and then later retrieves digital files in the existing memories of crowd workers. WeStore does not store information directly, but rather encrypts the files using details of the existing memories elicited from individuals within the crowd as cryptographic keys. The fidelity of the retrieved information is tied to how well the crowd remembers the details of the memories they provided. We demonstrate that crowd storage is feasible using an existing crowd marketplace (Amazon Mechanical Turk), explore design considerations important for building systems that use crowd storage, and outline ideas for future research in this area.
Jeffrey P. Bigham, Walter S. Lasecki
CHI2
2014 Finding dependencies between actions using the crowd
abstract
Activity recognition can provide computers with the context underlying user inputs, enabling more relevant responses and more fluid interaction. However, training these systems is difficult because it requires observing every possible sequence of actions that comprise a given activity. Prior work has enabled the crowd to provide labels in real-time to train automated systems on-the-fly, but numerous examples are still needed before the system can recognize an activity on its own. To reduce the need to collect this data by observing users, we introduce ARchitect, a system that uses the crowd to capture the dependency structure of the actions that make up activities. Our tests show that over seven times as many examples can be collected using our approach versus relying on direct observation alone, demonstrating that by leveraging the understanding of the crowd, it is possible to more easily train automated systems.
Walter S. Lasecki, Leon Weingard, George Ferguson, Jeffrey P. Bigham
CHI1
2014 Information extraction and manipulation threats in crowd-powered systems
abstract
Crowd-powered systems have become a popular way to augment the capabilities of automated systems in real-world settings. Many of these systems rely on human workers to process potentially sensitive data or make important decisions. This puts these systems at risk of unintentionally releasing sensitive data or having their outcomes maliciously manipulated. While almost all crowd-powered approaches account for errors made by individual workers, few factor in active attacks on the system. In this paper, we analyze different forms of threats from individuals and groups of workers extracting information from crowd-powered systems or manipulating these systems' outcomes. Via a set of studies performed on Amazon's Mechanical Turk platform and involving 1,140 unique workers, we demonstrate the viability of these threats. We show that the current system is vulnerable to coordinated attacks on a task based on the requests of another task and that a significant portion of Mechanical Turk workers are willing to contribute to an attack. We propose several possible approaches to mitigating these threats, including leveraging workers who are willing to go above and beyond to help, automatically flagging sensitive content, and using workflows that conceal information from each individual, while still allowing the group to complete a task. Our findings enable the crowd to continue to play an important part in automated systems, even as the data they use and the decisions they support become increasingly important.
Walter S. Lasecki, Jaime Teevan, Ece Kamar
CSCW1
2014 Introducing shared character control to existing video games
Anna Loparev, Walter S. Lasecki, Kyle I. Murray, Jeffrey P. Bigham
FDG2
2014 Glance Privacy: Obfuscating Personal Identity While Coding Behavioral Video
abstract
Behavioral researchers code video to extract systematic meaning from subtle human actions and emotions. While this has traditionally been done by analysts within a research group, recent methods have leveraged online crowds to massively parallelize this task and reduce the time required from days to seconds. However, using the crowd to code video increases the risk that private information will be disclosed because workers who have not been vetted will view the video data in order to code it. In this Work-in-Progress, we discuss techniques for maintaining privacy when using Glance to code video and present initial experimental evidence to support them.
Mitchell L. Gordon, Walter S. Lasecki, Winnie Leung, Ellen Lim, Steven Dow, Jeffrey P. Bigham
HCOMP2
2014 Combining Non-Expert and Expert Crowd Work to Convert Web APIs to Dialog Systems
abstract
Thousands of web APIs expose data and services that would be useful to access with natural dialog, from weather and sports to Twitter and movies. The process of adapting each API to a robust dialog system is difficult and time-consuming, as it requires not only programming but also anticipating what is mostly likely to be asked and how it is likely to be asked. We present a crowd-powered system able to generate a natural languageinterface for arbitrary web APIs from scratch without domain-dependent training data or knowledge.Our approach combines two types of crowd workers: non-expert Mechanical Turk workers interpret the functions of the API and elicit information from the user, and expert oDesk workers provide a minimal sufficient scaffolding around the API to allow us to make general queries.We describe our multi-stage process and present results for each stage.
Ting-Hao 'Kenneth' Huang, Walter S. Lasecki, Alan L. Ritter, Jeffrey P. Bigham
HCOMP2
2014 Tuning the Diversity of Open-Ended Responses From the Crowd
abstract
Crowdsourcing can solve problems beyond the reach of state-of-the-art fully automated systems. A common pattern found in many such systems is for the workers to discover, in parallel, a number of candidate solutions and then vote on the best one to pass forward, often within a fixed amount of time. We present the propose-vote-abstain mechanism for eliciting from crowd workers the proper balance between solution discovery and selection. Each crowd worker is given a choice among proposing an answer, voting among the answers proposed so far, or abstaining, i.e., doing nothing. When a stopping condition is reached, the mechanism returns the answer with the most votes. Workers are paid a base amount, with bonuses if they propose or vote for the winning answer.
Walter S. Lasecki, Christopher Homan, Jeffrey P. Bigham
HCOMP1
2014 Glance: rapidly coding behavioral video with the crowd
abstract
Behavioral researchers spend considerable amount of time coding video data to systematically extract meaning from subtle human actions and emotions. In this paper, we present Glance, a tool that allows researchers to rapidly query, sample, and analyze large video datasets for behavioral events that are hard to detect automatically. Glance takes advantage of the parallelism available in paid online crowds to interpret natural language queries and then aggregates responses in a summary view of the video data. Glance provides analysts with rapid responses when initially exploring a dataset, and reliable codings when refining an analysis. Our experiments show that Glance can code nearly 50 minutes of video in 5 minutes by recruiting over 60 workers simultaneously, and can get initial feedback to analysts in under 10 seconds for most clips. We present and compare new methods for accurately aggregating the input of multiple workers marking the spans of events in video data, and for measuring the quality of their coding in real-time before a baseline is established by measuring the variance between workers. Glance's rapid responses to natural language queries, feedback regarding question ambiguity and anomalies in the data, and ability to build on prior context in followup queries allow users to have a conversation-like interaction with their data - opening up new possibilities for naturally exploring video data.
Walter S. Lasecki, Mitchell L. Gordon, Danai Koutra, Malte F. Jung, Steven Dow, Jeffrey P. Bigham
UIST1
2014 Expert crowdsourcing with flash teams
abstract
We introduce flash teams, a framework for dynamically assembling and managing paid experts from the crowd. Flash teams advance a vision of expert crowd work that accomplishes complex, interdependent goals such as engineering and design. These teams consist of sequences of linked modular tasks and handoffs that can be computationally managed. Interactive systems reason about and manipulate these teams' structures: for example, flash teams can be recombined to form larger organizations and authored automatically in response to a user's request. Flash teams can also hire more people elastically in reaction to task needs, and pipeline intermediate output to accelerate completion times. To enable flash teams, we present Foundry, an end-user authoring platform and runtime manager. Foundry allows users to author modular tasks, then manages teams through handoffs of intermediate work. We demonstrate that Foundry and flash teams enable crowdsourcing of a broad class of goals including design prototyping, course development, and film animation, in half the work time of traditional self-managed teams.
Daniela Retelny, Sébastien Robaszkiewicz, Alexandra To, Walter S. Lasecki, Negar Rahmati, Tulsee Doshi, Melissa A. Valentine, Michael S. Bernstein
UIST4
2013 Crowdsourcing for Deployable Intelligent Systems
abstract
My work aims to create a scaffold for deployable intelligent systems using crowdsourcing. Current approaches in artificial intelligence (AI) typically focus on solving a narrow subset of problems in a given space - for example: automatic speech recognition as part of a conversational assistant, machine vision as part of a question answering service for blind people, or planning as part of a home assistive robot. This approach is necessary to scope the solution, but often results in a large number of systems that are rarely deployed in real-world setting, but instead operate in toy domains, or in situations where other parts of the problem are assumed to be solved. The framework I have developed aims to use the crowd to help in two ways: (i) make it possible to use human intelligence to power parts of a system that automated approaches cannot or do not yet handle, and (ii) provide a means of enabling more effective deployable systems by people to provide reliable training data on-demand. This summary begins with a brief review of prior work, then outlines a number of different system that I have developed to demonstrate the capabilities of this framework, and concludes with future work to be completed as part of my thesis.
Walter S. Lasecki
AAAI1
2013 Crowd Formalization of Action Conditions
abstract
Training intelligent systems is a time consuming and costly process that often limits their application to real-world problems. Prior work in crowdsourcing has attempted to compensate for this challenge by generating sets of labeled training data for machine learning algorithms. In this work, we seek to move beyond collecting just statistical data and explore how to gather structured, relational representations of a scenario using the crowd. We focus on activity recognition because of its broad applicability, high level of variation between individual instances, and difficulty of training systems a priori. We present ARchitect, a system that uses the crowd to ascertain pre and post conditions for actions observed in a video and find relations between actions. Our ultimate goal is to identify multiple valid execution paths from a single set of observations, which suggests one-off learning from the crowd is possible.
Walter S. Lasecki, Leon Weingard, Jeffrey P. Bigham, George Ferguson
AAAI1
2013 Real-time captioning by non-experts with legion scribe
abstract
Real-time captioning provides people who are deaf or hard of hearing access to speech in settings such as classrooms and live events. The most reliable approach to provide these captions is to recruit an expert stenographer who is able to type at natural speaking rates, but they charge more than $100 USD per hour and must be scheduled in advance. We introduce Legion Scribe (Scribe), a system that allows 3-5 ordinary people who can hear and type to jointly caption speech in real-time. Each person is unable to type at natural speaking rates, and so is asked only to type part of what they hear. Scribe automatically stitches all of the partial captions together to form a complete caption stream. We have shown that the accuracy of Scribe captions approaches that of a professional stenographer, while its latency and cost is dramatically lower.
Walter S. Lasecki, Christopher D. Miller, Raja S. Kushalnagar, Jeffrey P. Bigham
ASSETS1
2013 Answering visual questions with conversational crowd assistants
abstract
Blind people face a range of accessibility challenges in their everyday lives, from reading the text on a package of food to traveling independently in a new place. Answering general questions about one's visual surroundings remains well beyond the capabilities of fully automated systems, but recent systems are showing the potential of engaging on-demand human workers (the crowd) to answer visual questions. The input to such systems has generally been a single image, which can limit the interaction with a worker to one question; or video streams where systems have paired the end user with a single worker, limiting the benefits of the crowd. In this paper, we introduce Chorus:View, a system that assists users over the course of longer interactions by engaging workers in a continuous conversation with the user about a video stream from the user's mobile device. We demonstrate the benefit of using multiple crowd workers instead of just one in terms of both latency and accuracy, then conduct a study with 10 blind users that shows Chorus:View answers common visual questions more quickly and accurately than existing approaches. We conclude with a discussion of users' feedback and potential future work on interactive crowd support of blind users.
Walter S. Lasecki, Phyo Thiha, Erin L. Brady, Jeffrey P. Bigham
ASSETS1
2013 Warping time for more effective real-time crowdsourcing
abstract
In this paper, we introduce the idea of "warping time" to improve crowd performance on the difficult task of captioning speech in real-time. Prior work has shown that the crowd can collectively caption speech in real-time by merging the partial results of multiple workers. Because non-expert workers cannot keep up with natural speaking rates, the task is frustrating and prone to errors as workers buffer what they hear to type later. The TimeWarp approach automatically increases and decreases the speed of speech playback systematically across individual workers who caption only the periods played at reduced speed. Studies with 139 remote crowd workers and 24 local participants show that this approach improves median coverage (14.8%), precision (11.2%), and per-word latency (19.1%). Warping time may also help crowds outperform individuals on other difficult real-time performance tasks.
Walter S. Lasecki, Christopher D. Miller, Jeffrey P. Bigham
CHI1
2013 Real-time crowd labeling for deployable activity recognition
abstract
Systems that automatically recognize human activities offer the potential of timely, task-relevant information and support. For example, prompting systems can help keep people with cognitive disabilities on track and surveillance systems can warn of activities of concern. Current automatic systems are difficult to deploy because they cannot identify novel activities, and, instead, must be trained in advance to recognize important activities. Identifying and labeling these events is time consuming and thus not suitable for real-time support of already-deployed activity recognition systems. In this paper, we introduce Legion:AR, a system that provides robust, deployable activity recognition by supplementing existing recognition systems with on-demand, real-time activity identification using input from the crowd.
Walter S. Lasecki, Young Chol Song, Henry A. Kautz, Jeffrey P. Bigham
CSCW1
2013 Finding action dependencies using the crowd
abstract
Training intelligent systems is a time-consuming and costly process that often limits real-world applications. Prior work has attempted to compensate for this challenge by generating sets of labeled training data for machine learning algorithms using affordable human contributors. In this paper, we present ARchitect, a system that uses the crowd to extract context-dependent relational structure. We focus on activity recognition because of its broad applicability, high level of variation, and difficulty of training systems a priority. We demonstrate that using our approach, the crowd can accurately and consistently identify relationships between actions even over sessions containing different workers and varied executions of an activity. This results in the ability to identify multiple valid execution paths from a single observation, suggesting that one-off learning can be facilitated by using the crowd as an on-demand source of human intelligence in the knowledge acquisition process.
Walter S. Lasecki, Leon Weingard, George Ferguson, Jeffrey P. Bigham
K-CAP1
2013 Text Alignment for Real-Time Crowd Captioning
Iftekhar Naim, Daniel Gildea, Walter S. Lasecki, Jeffrey P. Bigham
HLT-NAACL3
2013 Chorus: a crowd-powered conversational assistant
abstract
Despite decades of research attempting to establish conversational interaction between humans and computers, the capabilities of automated conversational systems are still limited. In this paper, we introduce Chorus, a crowd-powered conversational assistant. When using Chorus, end users converse continuously with what appears to be a single conversational partner. Behind the scenes, Chorus leverages multiple crowd workers to propose and vote on responses. A shared memory space helps the dynamic crowd workforce maintain consistency, and a game-theoretic incentive mechanism helps to balance their efforts between proposing and voting. Studies with 12 end users and 100 crowd workers demonstrate that Chorus can provide accurate, topical responses, answering nearly 93% of user queries appropriately, and staying on-topic in over 95% of responses. We also observed that Chorus has advantages over pairing an end user with a single crowd worker and end users completing their own tasks in terms of speed, quality, and breadth of assistance. Chorus demonstrates a new future in which conversational assistants are made usable in the real world by combining human and machine intelligence, and may enable a useful new way of interacting with the crowds powering other systems.
Walter S. Lasecki, Rachel Wesley, Jeffrey Nichols 0001, Anand Kulkarni, James F. Allen, Jeffrey P. Bigham
UIST1
2012 Real-Time Collaborative Planning with the Crowd
abstract
Planning is vital to a wide range of domains, including robotics, military strategy, logistics, itinerary generation and more, that both humans and computers find difficult. Collaborative planning holds the promise of greatly improving performance on these tasks by leveraging the strengths of both humans and automated planners. However, this requires formalizing the problem domain and input, which must be done by hand, a priori, restricting its use in general real-world domains. We propose using a real-time crowd of workers to simultaneously solve the planning problem, formalize the domain, and train an automated system. As plans are developed, the system is able to learn the domain, and contribute larger segments of work.
Walter S. Lasecki, Jeffrey P. Bigham, James F. Allen, George Ferguson
AAAI1
2012 Online Sequence Alignment for Real-Time Audio Transcription by Non-Experts
abstract
Real-time transcription provides deaf and hard of hearing people visual access to spoken content, such as classroom instruction, and other live events. Currently, the only reliable source of real-time transcriptions are expensive, highly-trained experts who are able to keep up with speaking rates. Automatic speech recognition is cheaper but produces too many errors in realistic settings. We introduce a new approach in which partial captions from multiple non-experts are combined to produce a high-quality transcription in real-time. We demonstrate the potential of this approach with data collected from 20 non-expert captionists.
Walter S. Lasecki, Christopher D. Miller, Donato Borrello, Jeffrey P. Bigham
AAAI1
2012 A readability evaluation of real-time crowd captions in the classroom
abstract
Deaf and hard of hearing individuals need accommodations that transform aural to visual information, such as captions that are generated in real-time to enhance their access to spoken information in lectures and other live events. The captions produced by professional captionists work well in general events such as community or legal meetings, but is often unsatisfactory in specialized content events such as higher education classrooms. In addition, it is hard to hire professional captionists, especially those that have experience in specialized content areas, as they are scarce and expensive. The captions produced by commercial automatic speech recognition (ASR) software are far cheaper, but is often perceived as unreadable due to ASR's sensitivity to accents, background noise and slow response time. We ran a study to evaluate the readability of captions generated by a new crowd captioning approach versus professional captionists and ASR. In this approach, captions are typed by classmates into a system that aligns and merges the multiple incomplete caption streams into a single, comprehensive real-time transcript. Our study asked 48 deaf and hearing readers to evaluate transcripts produced by a professional captionist, ASR and crowd captioning software respectively and found the readers preferred crowd captions over professional captions and ASR.
Raja S. Kushalnagar, Walter S. Lasecki, Jeffrey P. Bigham
ASSETS2
2012 Online quality control for real-time crowd captioning
abstract
Approaches for real-time captioning of speech are either expensive (professional stenographers) or error-prone (automatic speech recognition). As an alternative approach, we have been exploring whether groups of non-experts can collectively caption speech in real-time. In this approach, each worker types as much as they can and the partial captions are merged together in real-time automatically. This approach works best when partial captions are correct and received within a few seconds of when they were spoken, but these assumptions break down when engaging workers on-demand from existing sources of crowd work like Amazon's Mechanical Turk. In this paper, we present methods for quickly identifying workers who are producing good partial captions and estimating the quality of their input. We evaluate these methods in experiments run on Mechanical Turk in which a total of 42 workers captioned 20 minutes of audio. The methods introduced in this paper were able to raise overall accuracy from 57.8% to 81.22% while keeping coverage of the ground truth signal nearly unchanged.
Walter S. Lasecki, Jeffrey P. Bigham
ASSETS1
2012 Real-time captioning by groups of non-experts
abstract
Real-time captioning provides deaf and hard of hearing people immediate access to spoken language and enables participation in dialogue with others. Low latency is critical because it allows speech to be paired with relevant visual cues. Currently, the only reliable source of real-time captions are expensive stenographers who must be recruited in advance and who are trained to use specialized keyboards. Automatic speech recognition (ASR) is less expensive and available on-demand, but its low accuracy, high noise sensitivity, and need for training beforehand render it unusable in real-world situations. In this paper, we introduce a new approach in which groups of non-expert captionists (people who can hear and type) collectively caption speech in real-time on-demand. We present Legion:Scribe, an end-to-end system that allows deaf people to request captions at any time. We introduce an algorithm for merging partial captions into a single output stream in real-time, and a captioning interface designed to encourage coverage of the entire audio stream. Evaluation with 20 local participants and 18 crowd workers shows that non-experts can provide an effective solution for captioning, accurately covering an average of 93.2% of an audio stream with only 10 workers and an average per-word latency of 2.9 seconds. More generally, our model in which multiple workers contribute partial inputs that are automatically merged in real-time may be extended to allow dynamic groups to surpass constituent individuals (even experts) on a variety of human performance tasks.
Walter S. Lasecki, Christopher D. Miller, Adam Sadilek, Andrew Abumoussa, Donato Borrello, Raja S. Kushalnagar, Jeffrey P. Bigham
UIST1
2011 Real-time crowd control of existing interfaces
abstract
Crowdsourcing has been shown to be an effective approach for solving difficult problems, but current crowdsourcing systems suffer two main limitations: (i) tasks must be repackaged for proper display to crowd workers, which generally requires substantial one-off programming effort and support infrastructure, and (ii) crowd workers generally lack a tight feedback loop with their task. In this paper, we introduce Legion, a system that allows end users to easily capture existing GUIs and outsource them for collaborative, real-time control by the crowd. We present mediation strategies for integrating the input of multiple crowd workers in real-time, evaluate these mediation strategies across several applications, and further validate Legion by exploring the space of novel applications that it enables.
Walter S. Lasecki, Kyle I. Murray, Samuel White, Rob Miller 0001, Jeffrey P. Bigham
UIST1