Jeffrey M. Rzeszotarski

dblp:18/10300 · DBLP profile ↗
← Back
29ranked-venue papers
6as first author
18since 2021 · last 2026
0000-0002-4317-9501ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 26 · 6 first-author · 15 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Embodying Facts, Figures, and Faiths in Narrative Artistic Performances in Rural Bangladesh
abstract
There is an increasing interest in telling serious stories with data. Designers organize information, construct narratives, and present findings to inform audiences. However, many of these practices emerge from modern information visualization rhetoric and ethical frameworks which may marginalize communities with low digital and media literacy. In a ten-month-long ethnographic study in three Bangladeshi villages, we investigated how these communities use entertainment and cultural practices, namely Puthi, Bhandari Gaan, and Pot music, to instruct, communicate traditional moral lessons and recall history. We found that these communities embrace polyvocality and multiple ethical frameworks in their performances, construct narratives combining factuality, emotionality, and aesthetics, and adapt their performances to changing technology and audience needs. Our findings provide HCI, visualization, and ethical data practitioners with implications for the design of accessible and culturally appropriate ways of presenting data narratives in data-driven systems.
Sharifa Sultana, Jeffrey M. Rzeszotarski, Zinnat Sultana, Syed Ishtiaque Ahmed
CHI2
2026 Fairness-in-the-Workflow: How Machine Learning Practitioners at Big Tech Companies Approach Fairness in Recommender Systems
abstract
Recommender systems (RS), which are widely deployed across high-stakes domains, are susceptible to biases that can cause large-scale societal impacts. Researchers have proposed methods to measure and mitigate such biases—but translating academic theory into practice is inherently challenging. Through a semi-structured interview study (N=11), we map the RS practitioner workflow within large technology companies, focusing on how technical teams consider fairness internally and in collaboration with legal, data, and fairness teams. We identify key challenges to incorporating fairness into existing RS workflows: defining fairness in RS contexts, balancing multi-stakeholder interests, and navigating dynamic environments. We also identify key organization-wide challenges: making time for fairness work and facilitating cross-team communication. Finally, we offer actionable recommendations for the RS community, including practitioners and HCI researchers.
Jing Nathan Yan, Emma Harvey, Junxiong Wang, Jeffrey M. Rzeszotarski, Allison Koenecke
CHI4
2026 EvaluAId: Human-AI Collaborative Evaluation of Open-Ended Student Essays
abstract
Open-ended writing assignments are central to higher education, yet heterogeneous submissions and scale make evaluation difficult. Automated writing evaluation (AWE) promises speed but often trades away transparency and sidelines human judgment. This paper repositions the AI as an on-demand collaborator that can provide specific, targeted support. In a formative study, we expose leverage points in three cognitive dimensions: evidence identification, comparative judgment, and feedback composition. Guided by these insights, we build EvaluAId, which supports interactive rubric-content mapping, adaptive benchmarking and self-calibration, and personalized, rubric-aligned feedback synthesis. Through a within-subjects study with 12 TAs, we evaluate how this approach supports grading compared with a rubric+LLM chatbot and an LLM-based AWE; EvaluAId improved alignment with expert ratings and increased graders’ satisfaction. Finally, interviews with TAs, instructors, and students underscored the value of thoughtfulness supported by EvaluAId while surfacing practical considerations for integration into classroom. Together, our results argue for deliberate, evidence-first, human-in-the-loop evaluation.
Chao Zhang 0082, Kexin Ju 0001, Xinyi Lu 0004, Yu-Chun (Grace) Yen, Jeffrey M. Rzeszotarski
CHI5
2025 Towards Hormone Health: An Autoethnography of Long-Term Holistic Tracking to Manage PCOS
Daye Kang, Jingjin Li, Gilly Leshed, Jeffrey M. Rzeszotarski, Xi Lu 0002
CHI4
2025 Friction: Deciphering Writing Feedback into Writing Revisions through LLM-Assisted Reflection
Chao Zhang 0082, Kexin Ju 0001, Peter Bidoshi, Yu-Chun (Grace) Yen, Jeffrey M. Rzeszotarski
CHI5
2025 What We Talk About When We Talk About LMs: Implicit Paradigm Shifts and the Ship of Language Models
abstract
Shengqi Zhu, Jeffrey Rzeszotarski. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Shengqi Zhu 0002, Jeffrey M. Rzeszotarski
NAACL (Long Papers)2
2025 Synthia: Visually Interpreting and Synthesizing Feedback for Writing Revision
Chao Zhang 0082, Kexin Ju 0001, Zhuolun Han, Yu-Chun (Grace) Yen, Jeffrey M. Rzeszotarski
UIST5
2025 ThemeViz: Understanding the Effect of Human-AI Collaboration in Theme Development with an LLM-enhanced Interactive Visual System
abstract
This paper explores the potential role of AI, e.g., large language models (LLMs), in supporting theme development in thematic analysis. While prior applications of AI in qualitative data analysis have focused on supporting coding, we investigate whether LLMs can effectively contribute as collaborators in the more abstract and conceptual phases of qualitative analysis, specifically theme development. Despite growing interest in AI as a collaborator in theme development, there is limited empirical evidence on designing AI-assisted tools while supporting user autonomy and understanding researcher interaction with AI-assisted theme development. To address this gap, we designed ThemeViz, an interactive system that uses GPT-4 to generate and visualize multiple versions of themes based on user input while allowing researchers to maintain control through manual coding and theme development. Our study examines the effectiveness of this human-AI collaboration approach in iterative theme development and its implications for future designs.
Daye Kang, Zhuolun Han, Muhan Zhang, Jeffrey M. Rzeszotarski
Proc. ACM Hum. Comput. Interact.5
2024 Teachable Facets: A Framework of Interactive Machine Teaching for Information Filtering
abstract
Interactive tools help users filter relevant information from massive online sources, like news feeds and online discussion forums, by enabling them to externalize their preferences. However, users’ information goals and preferences are often complex and are comprised of data attributes and a user’s subjective judgements over these attributes. For instance, when filtering news articles based on their newsworthiness, the system must capture both data attributes like recency and shareability of the article, along with the user’s personal and flexible assessment of news sentiment. While most interactive tools enable users to externalize goals that are expressible as true/false statements, they do not support incorporating subjective, loosely structured judgements of data attributes which fulfill complex goals. In this paper, we introduce Teachable Facets (TF), widgets that users can create on the fly to filter relevant information to improve the sense-making of analysts. These teachable widgets employ a Machine Teaching (MT) framework to enable users to formulate personalized filtering criteria for complex, multi-dimensional, loosely indexed, and unstructured data; teach a filtering criterion using representative samples; apply these filters to new data streams; and assess the relevance of outcomes. Through a user study, we evaluate the performance of these filters based on their ability to discover relevant items and the expressibility they offer to the users in teaching criteria. In our discussion, we identify ways this approach might improve future systems and delineate implications should such systems be deployed broadly.
Swati Mishra 0006, Matthew L. Ryerkerk, Yitzchak Lockerman, David Eis, Jeffrey M. Rzeszotarski
CHIIR5
2024 Challenges and Opportunities for Tool Adoption in Industrial UX Research Collaborations
abstract
UX research practitioners analyze qualitative data to comprehend users' needs and synthesize design implications for software systems. Collaborating with multiple stakeholders is inevitable for these professionals and adds additional pressure to their already laborious data analysis tasks. In this paper, we investigate how these practitioners' multi-stakeholder collaboration affects their data analysis practices. Specifically, we investigate the challenges qualitative UX research (QUXR) practitioners face, limitations of the current qualitative data analysis (QDA) support tools they use, and design implications for future QDA tools to address the limitations. Through semi-structured interviews combined with diagramming activities with thirteen industry QUXR practitioners, we have revealed that collaboration becomes a bigger challenge than data analysis and that it often leads to (1) preferring simple tools despite the QDA specialized tools' beneficial features and (2) wanting to triangulate their work to convince stakeholders. Finally, we have synthesized these results and suggest design implications for QDA support tools.
Daye Kang, Jeffrey M. Rzeszotarski
Proc. ACM Hum. Comput. Interact.2
2023 Communicating Consequences: Visual Narratives, Abstraction, and Polysemy in Rural Bangladesh
abstract
Information communication and visualization practices reflect two centuries of developments of conventions and best practices which may not be reflective of global audiences’ methods for conveying information. Contrasting between rural traditional visual culture and contemporary HCI and data-visualization, we argue that an understanding of traditional practices for information visualization is required for building rich data-narratives and making data-driven systems more accessible and culturally situated. Our ten-month ethnographic study investigates how rural Bangladeshi communities construct narratives through visual media. 1 Our observation, interviews, and FGDs (n=54) expose how participants convey risk management, decision-making, and monetary management practices to their peers. We find that villagers used a rich network of polysemic symbols and abstractions to manifest subjectivity, factuality, consequence, situatedness, and uncertainty; varied visual attributes for constructing narratives; and emphasized material relations among components in visuals. These findings inform the design of future systems for decision support in a culturally situated manner.
Sharifa Sultana, Syed Ishtiaque Ahmed, Jeffrey M. Rzeszotarski
CHI3
2023 Human Expectations and Perceptions of Learning in Machine Teaching
abstract
Interactive interfaces in tandem with Machine Learning (ML) models support user understanding of model uncertainty, build confidence, improve predictive accuracy and enable users to teach application-specific concepts that are difficult for the model to learn otherwise. These systems offer empirically proven benefits due to tightly coupled feedback loops and workflow scaffolding. However, deployment with ML non-experts who cannot manage the complex, expertise-heavy process remains challenging. Through deployment with non-expert users in a common classification task, we investigate the impact of human factors of machine teaching interfaces such as user expectations, their perceptions of the learning process and user engagement with respect to teaching process and outcomes. We measure how affective and performance attributes shape the success or failure of the process. Finally, we reflect on how intelligent user interfaces can be designed to accommodate these factors for successful deployment with a broad spectrum of human adjudicators.
Swati Mishra 0006, Jeffrey M. Rzeszotarski
UMAP2
2023 Understanding Motivational Factors in Social Media News Sharing Decisions
abstract
News sharing has become prevalent on many social media platforms. Users are not only exposed to news shared by others, but also actively share information with a diverse set of motivations. In this work, we propose five news sharing motivations based on the intrinsic and extrinsic factors found in prior literature. Through an online experiment, we further examine how a host of factors, including motivations, influence participants' decision to share news online. We then prompt participants to switch their original decision for extra compensation, observing how different news types, motivational and demographic factors may affect the switch. Our analysis suggests that sharing decisions can be reversed when a strong external stimulus (higher bonus) is presented. Further, there are motivational factors that independently influence participants' reversal decisions. Finally, using our work as an empirical basis, we propose designs for future new sharing systems.
Luping Wang 0006, Jeffrey M. Rzeszotarski
Proc. ACM Hum. Comput. Interact.2
2021 Designing Interactive Transfer Learning Tools for ML Non-Experts
abstract
Interactive machine learning (iML) tools help to make ML accessible to users with limited ML expertise. However, gathering necessary training data and expertise for model-building remains challenging. Transfer learning, a process where learned representations from a model trained on potentially terabytes of data can be transferred to a new, related task, offers the possibility of providing ”building blocks” for non-expert users to quickly and effectively apply ML in their work. However, transfer learning largely remains an expert tool due to its high complexity. In this paper, we design a prototype to understand non-expert user behavior in an interactive environment that supports transfer learning. Our findings reveal a series of data- and perception-driven decision-making strategies non-expert users employ, to (in)effectively transfer elements using their domain expertise. Finally, we synthesize design implications which might inform future interactive transfer learning environments.
Swati Mishra 0006, Jeffrey M. Rzeszotarski
CHI2
2021 Tessera: Discretizing Data Analysis Workflows on a Task Level
abstract
Researchers have investigated a number of strategies for capturing and analyzing data analyst event logs in order to design better tools, identify failure points, and guide users. However, this remains challenging because individual- and session-level behavioral differences lead to an explosion of complexity and there are few guarantees that log observations map to user cognition. In this paper we introduce a technique for segmenting sequential analyst event logs which combines data, interaction, and user features in order to create discrete blocks of goal-directed activity. Using measures of inter-dependency and comparisons between analysis states, these blocks identify patterns in interaction logs coupled with the current view that users are examining. Through an analysis of publicly available data and data from a lab study across a variety of analysis tasks, we validate that our segmentation approach aligns with users’ changing goals and tasks. Finally, we identify several downstream applications for our approach.
Jing Nathan Yan, Ziwei Gu, Jeffrey M. Rzeszotarski
CHI3
2021 DataPrep.EDA: Task-Centric Exploratory Data Analysis for Statistical Modeling in Python
abstract
Exploratory Data Analysis (EDA) is a crucial step in any data science project. However, existing Python libraries fall short in supporting data scientists to complete common EDA tasks for statistical modeling. Their API design is either too low level, which is optimized for plotting rather than EDA, or too high level, which is hard to specify more fine-grained EDA tasks. In response, we propose DataPrep.EDA, a novel task-centric EDA system in Python. DataPrep.EDA allows data scientists to declaratively specify a wide range of EDA tasks in different granularity with a single function call. We identify a number of challenges to implement DataPrep.EDA, and propose effective solutions to improve the scalability, usability, customizability of the system. In particular, we discuss some lessons learned from using Dask to build the data processing pipelines for EDA tasks and describe our approaches to accelerate the pipelines. We conduct extensive experiments to compare DataPrep.EDA with Pandas-profiling, the state-of-the-art EDA system in Python. The experiments show that DataPrep.EDA significantly outperforms Pandas-profiling in terms of both speed and user experience. DataPrep.EDA is open-sourced as an EDA component of DataPrep: https://github.com/sfu-db/dataprep.
Jinglin Peng, Weiyuan Wu, Brandon Lockhart, Song Bian 0002, Jing Nathan Yan, Linghao Xu, Zhixuan Chi, Jeffrey M. Rzeszotarski, Jiannan Wang 0001
SIGMOD Conference8
2021 Understanding User Sensemaking in Machine Learning Fairness Assessment Systems
abstract
A variety of systems have been proposed to assist users in detecting machine learning (ML) fairness issues. These systems approach bias reduction from a number of perspectives, including recommender systems, exploratory tools, and dashboards. In this paper, we seek to inform the design of these systems by examining how individuals make sense of fairness issues as they use different de-biasing affordances. In particular, we consider the tension between de-biasing recommendations which are quick but may lack nuance and ”what-if” style exploration which is time consuming but may lead to deeper understanding and transferable insights. Using logs, think-aloud data, and semi-structured interviews we find that exploratory systems promote a rich pattern of hypothesis generation and testing, while recommendations deliver quick answers which satisfy participants at the cost of reduced information exposure. We highlight design requirements and trade-offs in the design of ML fairness systems to promote accurate and explainable assessments.
Ziwei Gu, Jing Nathan Yan, Jeffrey M. Rzeszotarski
WWW3
2021 Crowdsourcing and Evaluating Concept-driven Explanations of Machine Learning Models
abstract
An important challenge in building explainable artificially intelligent (AI) systems is designing interpretable explanations. AI models often use low-level data features which may be hard for humans to interpret. Recent research suggests that situating machine decisions in abstract, human understandable concepts can help. However, it is challenging to determine the right level of conceptual mapping. In this research, we explore granularity (of data features) and context (of data instances) as dimensions underpinning conceptual mappings. Based on these measures, we explore strategies for designing explanations in classification models. We introduce an end-to-end concept elicitation pipeline that supports gathering high-level concepts for a given data set. Through crowd-sourced experiments, we examine how providing conceptual information shapes the effectiveness of explanations, finding that a balance between coarse and fine-grained explanations help users better estimate model predictions. We organize our findings into systematic themes that can inform design considerations for future systems.
Swati Mishra 0006, Jeffrey M. Rzeszotarski
Proc. ACM Hum. Comput. Interact.2
2020 Silva: Interactively Assessing Machine Learning Fairness Using Causality
abstract
Machine learning models risk encoding unfairness on the part of their developers or data sources. However, assessing fairness is challenging as analysts might misidentify sources of bias, fail to notice them, or misapply metrics. In this paper we introduce Silva, a system for exploring potential sources of unfairness in datasets or machine learning models interactively. Silva directs user attention to relationships between attributes through a global causal view, provides interactive recommendations, presents intermediate results, and visualizes metrics. We describe the implementation of Silva, identify salient design and technical challenges, and provide an evaluation of the tool in comparison to an existing fairness optimization tool.
Jing Nathan Yan, Ziwei Gu, Hubert Lin, Jeffrey M. Rzeszotarski
CHI4
2020 Seeing in Context: Traditional Visual Communication Practices in Rural Bangladesh
abstract
There is a risk that modern practices of information communication and visualization in human-computer interaction can sideline communities due to their prioritization of scientific rationality. Such ideological hegemony can complicate interactions with data and computers, especially for low-literate communities in the global south. Through a six-month long ethnographic study with Nakshi-Katha makers, Hindu Idol makers, and witchcraft practitioners, we investigated how rural practitioners use their own forms of representation and narrative in record keeping, social and religious storytelling, and information mediated decision making. We find that traditionally developed approaches towards presenting and communicating information often make use of concrete units to represent entities and connect to designers' cultural practices and the physical location. Further, we identify how medium has significant influence in meaning-making. Often these strategies and conventions are passed down through generations within the community. In this paper, we discuss how this rural tradition differs from the modern information communication practices, discussing how an understanding of traditional practices for representing information can be useful in developing more accessible, and culturally appropriate modern tools and technologies for the people of rural Bangladesh and similar communities.
Sharifa Sultana, Syed Ishtiaque Ahmed, Jeffrey M. Rzeszotarski
Proc. ACM Hum. Comput. Interact.3
2019 The Tools of Management: Adapting Historical Union Tactics to Platform-Mediated Labor
abstract
At the same time that workers' rights are generally declining in the United States (US), workplace computing systems gather more data about workers and their activities than ever before. The rise of large scale labor analytics raises questions about how and whether workers could use such data to advocate for their own goals. Here, we analyze the historical development of workplace technology design methods in CSCW to show how mid-20th century labor responses to scientific management can inform directions in contemporary digital labor advocacy. First, we demonstrate how specific methodological tendencies from industrial scientific management were adapted to work in CSCW, and then subsequently altered in crowd work and social computing research to more closely resemble industrial approaches. Next, we show how three tactics used by labor unions to strategically engage with industrial scientific management in the mid-20th century can inform data-driven worker advocacy in platform-mediated work. Finally, we discuss how this history shapes our understanding of worker participation and the implications of using worker data for contemporary advocacy goals.
Vera D. Khovanskaya, Lynn Dombrowski, Jeffrey M. Rzeszotarski, Phoebe Sengers
Proc. ACM Hum. Comput. Interact.3
2015 The Effects of Sequence and Delay on Crowd Work
abstract
A common approach in crowdsourcing is to break large tasks into small microtasks so that they can be parallelized across many crowd workers and so that redundant work can be more easily compared for quality control. In practice, this can result in the microtasks being presented out of their natural order and often introduces delays between individual microtasks. In this paper, we demonstrate in a study of 338 crowd workers that non-sequential microtasks and the introduction of delays significantly decreases worker performance. We show that interruptions where a large delay occurs between two related tasks can cause up to a 102% slowdown in completion time, and interruptions where workers are asked to perform different tasks in sequence can slow down completion time by 57%. We conclude with a set of design guidelines to improve both worker performance and realized pay, and instructions for implementing these changes in existing interfaces for crowd work.
Walter S. Lasecki, Jeffrey M. Rzeszotarski, Adam Marcus 0002, Jeffrey P. Bigham
CHI2
2015 And Now for Something Completely Different: Improving Crowdsourcing Workflows with Micro-Diversions
abstract
Crowdsourcing has become a popular and indispensable component of many problem-solving pipelines in the research literature, with crowd workers often treated as computational resources that can reliably solve problems that computers have trouble with, such as image labeling/classification, natural language processing, or document writing. Yet, obviously crowd workers are human, and long sequences of the same monotonous tasks might intuitively reduce the amount of good quality work done by the workers. Here we propose an investigation into how we can use diversions containing small amounts of entertainment to improve crowd workers' experiences. We call these small period of entertainment ``micro-diversions", which we hypothesize to provide timely relief to workers during long sequences of micro-tasks. We hope to improve productivity by retaining workers to work on our tasks longer and to either improve or retain the quality of work. We experimentally test micro-diversions on Amazon's Mechanical Turk, a large paid-crowdsourcing platform. We find that micro-diversions can significantly improve worker retention rate while retaining the same work quality.
Peng Dai 0001, Jeffrey M. Rzeszotarski, Praveen K. Paritosh, Ed H. Chi
CSCW2
2014 Kinetica: naturalistic multi-touch data visualization
abstract
Over the last several years there has been an explosion of powerful, affordable, multi-touch devices. This provides an outstanding opportunity for novel data visualization techniques that leverage new interaction methods and minimize their barriers to entry. In this paper we describe an approach for multivariate data visualization that uses physics-based affordances that are easy to intuit, constraints that are easy to apply and visualize, and a consistent view as data is manipulated in order to promote data exploration and interrogation. We provide a framework for exploring this problem space, and an example proof of concept system called Kinetica. We describe the results of a user study that suggest users of Kinetica were able to explore multiple dimensions of data at once, identify outliers, and discover trends with minimal training.
Jeffrey M. Rzeszotarski, Aniket Kittur
CHI1
2014 Estimating the social costs of friendsourcing
abstract
Every day users of social networking services ask their followers and friends millions of questions. These friendsourced questions not only provide informational benefits, but also may reinforce social bonds. However, there is a limit to how much a person may want to friendsource. They may be uncomfortable asking questions that are too private, might not want to expend others' time or effort, or may feel as though they have already accrued too many social debts. These perceived social costs limit the potential benefits of friendsourcing. In this paper we explore the perceived social costs of friendsourcing on Twitter via a monetary choice. We develop a model of how users value the attention and effort of their social network while friendsourcing, compare and contrast it with paid question answering in a crowdsourced labor market, and provide future design considerations for better supporting friendsourcing.
Jeffrey M. Rzeszotarski, Meredith Ringel Morris
CHI1
2014 Is anyone out there?: unpacking Q&A hashtags on twitter
abstract
In addition to posting news and status updates, many Twitter users post questions that seek various types of subjective and objective information. These questions are often labeled with 'Q&A' hashtags, such as #lazyweb or #twoogle. We surveyed Twitter users and found they employ these Q&A hashtags both as a topical signifier (this tweet needs an answer!) and to reach out to those beyond their immediate followers (a community of helpful tweeters who monitor the hashtag). However, our log analysis of thousands of hashtagged Q&A exchanges reveals that nearly all replies to hashtagged questions come from a user's immediate follower network, contradicting users' beliefs that they are tapping into a larger community by tagging their question tweets. This finding has implications for designing next-generation social search systems that reach and engage a wide audience of answerers.
Jeffrey M. Rzeszotarski, Emma S. Spiro, J. Nathan Matias, Andrés Monroy-Hernández, Meredith Ringel Morris
CHI1
2012 Learning from history: predicting reverted work at the word level in wikipedia
abstract
Wikipedia's remarkable success in aggregating millions of contributions can pose a challenge for current editors, whose hard work may be reverted unless they understand and follow established norms, policies, and decisions and avoid contentious or proscribed terms. We present a machine learning model for predicting whether a contribution will be reverted based on word level features. Unlike previous models relying on editor-level characteristics, our model can make accurate predictions based only on the words a contribution changes. A key advantage of the model is that it can provide feedback on not only whether a contribution is likely to be rejected, but also the particular words that are likely to be controversial, enabling new forms of intelligent interfaces and visualizations. We examine the performance of the model across a variety of Wikipedia articles.
Jeffrey M. Rzeszotarski, Aniket Kittur
CSCW1
2012 CrowdScape: interactively visualizing user behavior and output
abstract
Crowdsourcing has become a powerful paradigm for accomplishing work quickly and at scale, but involves significant challenges in quality control. Researchers have developed algorithmic quality control approaches based on either worker outputs (such as gold standards or worker agreement) or worker behavior (such as task fingerprinting), but each approach has serious limitations, especially for complex or creative work. Human evaluation addresses these limitations but does not scale well with increasing numbers of workers. We present CrowdScape, a system that supports the human evaluation of complex crowd work through interactive visualization and mixed initiative machine learning. The system combines information about worker behavior with worker outputs, helping users to better understand and harness the crowd. We describe the system and discuss its utility through grounded case studies. We explore other contexts where CrowdScape's visualizations might be useful, such as in user studies.
Jeffrey M. Rzeszotarski, Aniket Kittur
UIST1
2011 Instrumenting the crowd: using implicit behavioral measures to predict task performance
abstract
Detecting and correcting low quality submissions in crowdsourcing tasks is an important challenge. Prior work has primarily focused on worker outcomes or reputation, using approaches such as agreement across workers or with a gold standard to evaluate quality. We propose an alternative and complementary technique that focuses on the way workers work rather than the products they produce. Our technique captures behavioral traces from online crowd workers and uses them to predict outcome measures such quality, errors, and the likelihood of cheating. We evaluate the effectiveness of the approach across three contexts including classification, generation, and comprehension tasks. The results indicate that we can build predictive models of task performance based on behavioral traces alone, and that these models generalize to related tasks. Finally, we discuss limitations and extensions of the approach.
Jeffrey M. Rzeszotarski, Aniket Kittur
UIST1