EDBT 2026 Demo / reviewers in the wild / expert
Ming Yin 0001
dblp:89/453-1
· DBLP profile ↗
54ranked-venue papers
3as first author
46since 2021 · last 2026
0000-0002-7364-139XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 34 · 2 first-author · 29 since 2021Artificial intelligence and machine learning · 18 · 18 since 2021Databases, data management, data science and information retrieval · 8 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 8 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 3 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Align When They Want, Complement When They Need! Human-Centered Ensembles for Adaptive Human-AI CollaborationabstractIn human-AI decision making, designing AI that complements human expertise has been a natural strategy to enhance human-AI collaboration, yet it often comes at the cost of decreased AI performance in areas of human strengths. This can inadvertently erode human trust and cause them to ignore AI advice precisely when it is most needed. Conversely, an aligned AI fosters trust yet risks reinforcing suboptimal human behavior and lowering human-AI team performance. In this paper, we start by identifying this fundamental tension between performance-boosting (i.e., complementarity) and trust-building (i.e., alignment) as an inherent limitation of the traditional approach for training a single AI model to assist human decision making. To overcome this, we introduce a novel, human-centered adaptive AI ensemble that strategically toggles between two specialist AI models—the aligned model and the complementary model—based on contextual cues, using an elegantly simple yet provably near-optimal Rational Routing Shortcut mechanism. Comprehensive theoretical analyses elucidate why the adaptive AI ensemble is effective and when it yields maximum benefits. Moreover, experiments on both simulated and real-world data show that when humans are assisted by the adaptive AI ensemble in decision making, they can achieve significantly higher performance than when they are assisted by single AI models that are trained to either optimize for its independent performance or even the human-AI team performance. Syed Hasan Amin Mahmood, Ming Yin 0001, Rajiv Khanna |
AAAI | 2 |
| 2026 | Large Language Model (LLM)-driven Adversarial Social Influences in Online Information Spread: Risks and InterventionsabstractPeople’s online information processing is strongly shaped by social influence, and large language models (LLMs) now enable social bots to manipulate such influence at scale. This paper examines the effects of LLM-driven adversarial social influence—a strategy in which automated agents employ LLMs to distort truth by making misinformation appear credible or by undermining factual news—on how people evaluate and share information. Across two pre-registered, randomized experiments, we first show that exposure to LLM-driven adversarial social influence significantly reduces people’s ability to judge the veracity of news and lowers their discernment between sharing true versus false content. We then test two credibility prompts: AI-generated content detectors and warnings, as potential interventions. Results show that both prompts mitigate some harms such as by improving misinformation detection, though their effectiveness were dependent on the context. We conclude by discussing the risks of LLM-driven adversarial social bots and the implications for designing interventions to combat misinformation. Zhuoran Lu, Gionnieve Lim, Ming Yin 0001 |
CHI | 3 |
| 2026 | Understanding the Effects of AI-Assisted Critical Thinking on Human-AI Decision MakingabstractDespite the growing prevalence of human-AI decision making, the human-AI team’s decision performance often remains suboptimal, partially due to insufficient examination of humans’ own reasoning. In this paper, we explore designing AI systems that directly analyze humans’ decision rationales and encourage critical reflection of their own decisions. We introduce the AI-Assisted Critical Thinking (AACT) framework, which leverages a domain-specific AI model’s counterfactual analysis of human decision to help decision-makers identify potential flaws in their decision argument and support the correction of them. Through a case study on house price prediction, we find that AACT outperforms traditional AI-based decision-support in reducing over-reliance on AI, though also triggering higher cognitive load. Subgroup analysis reveals AACT can be particularly beneficial for some decision-makers such as those very familiar with AI technologies. We conclude by discussing the practical implications of our findings, use cases and design choices of AACT, and considerations for using AI to facilitate critical thinking. Harry Yizhou Tian, Hasan Amin, Ming Yin 0001 |
CHI | 3 |
| 2026 | StarBurst: Aiding Design Ideation Through AI-Generated Remote AssociationsabstractRemote associations play a crucial role in enhancing creativity during design ideation, yet designers face challenges in effectively creating and integrating them. Our formative study (N = 5) shows the potential of AI-generated remote associations to facilitate this process, but there are still challenges to understand and apply them. To address these, we propose StarBurst, which supports design ideation through AI-generated remote associations. It consists of three core components: (1) generating diverse remote associations from images to expand creative possibilities; (2) constructing an attribute map to explore connections between associative elements; and (3) providing suggestions to integrate these associations into the final design idea. Through two forms of user studies (N = 32, N = 16), we found that StarBurst outperformed designers in generating remote associations and provided more effective support for diverse and creative idea development compared to the baseline system. Additionally, we discussed how users’ usage patterns and perceptions influence StarBurst’s effectiveness. Runqi Fang, Fang Liu 0002, Yunfan Ye, Shenglan Cui, Ming Yin 0001 |
Int. J. Hum. Comput. Interact. | 6 |
| 2026 | Hello!AI: An Interactive Rhyme-Based Game for Children AI Literacy EducationabstractAI literacy is critical for young children as AI is rapidly integrated into people’s daily lives. However, the complexity of AI knowledge presents significant learning challenges, and there is currently a lack of effective approaches for converting complex AI concepts into easy-comprehend content. Based on the formative analysis, we propose Hello!AI, an interactive rhyme-based AI literacy education game targeted at children in Grades 2–6 of primary schools. Hello!AI comprises 3 modules: (i) Algorithm Adventure, focusing on basic AI concept learning, (ii) Algorithm Handbook, promoting thinking and reflection on AI algorithms, and (iii) City Builder, emphasizing the application of AI algorithm to solve real-life problems. We developed the prototype system, iterated it through pilot study, and then conducted a user study. The results demonstrate that Hello!AI can effectively engage children and, to a certain extent, improve their ability to understand and apply AI knowledge, as well as their thinking and reflective capabilities regarding AI technologies. Mohan Zhang, Changjuan Ran, Fang Liu 0002, Ming Yin 0001, Shenglan Cui, Chuhan Li, Biyao Li |
Int. J. Hum. Comput. Interact. | 4 |
| 2025 | From Text to Trust: Empowering AI-assisted Decision Making with Adaptive LLM-powered Analysis
Zhuoyan Li, Hangxiao Zhu, Zhuoran Lu, Ziang Xiao, Ming Yin 0001 |
CHI | 5 |
| 2025 | Understanding the Effects of AI-based Credibility Indicators When People Are Influenced By Both Peers and Experts
Zhuoran Lu, Patrick Li, Ming Yin 0001 |
CHI | 4 |
| 2025 | Towards Human-AI Deliberation: Design and Evaluation of LLM-Empowered Deliberative AI for AI-Assisted Decision-MakingabstractTraditional AI-assisted decision-making systems often provide fixed recommendations that users must either accept or reject entirely, limiting meaningful interaction - especially in cases of disagreement. To address this, we introduce Human-AI Deliberation, an approach inspired by human deliberation theories that enables dimension-level opinion elicitation, iterative decision updates, and structured discussions between humans and AI. At the core of this approach is Deliberative AI, an assistant powered by large language models (LLMs) that facilitates flexible, conversational interactions and precise information exchange with domain-specific models. Through a mixed-methods user study, we found that Deliberative AI outperforms traditional explainable AI (XAI) systems by fostering appropriate human reliance and improving task performance. By analyzing participant perceptions, user experience, and open-ended feedback, we highlight key findings, discuss potential concerns, and explore the broader applicability of this approach for future AI-assisted decision-making systems. Shuai Ma 0005, Qiaoyi Chen, Chengbo Zheng, Zhenhui Peng, Ming Yin 0001, Xiaojuan Ma |
CHI | 6 |
| 2025 | On the Support Vector Effect in DNNs: Rethinking Data Selection and AttributionabstractIn Deep Neural Networks (DNNs), manipulating gradients is central to various algorithms, including data subset selection and instance attribution. For better tractability, practitioners often resort to using only the gradients of the last layer as a heuristic, instead of the full gradient across all model parameters, which we show is detrimental due to the Support Vector Effect (SVE). We introduce SVE, a max-margin-like behavior in the last layer(s) of DNNs and employ it to thoroughly scrutinize prevalent data selection and attribution methods relying on last layer gradients. Our investigation exposes limitations in these techniques and not only provides explanations for previously observed pitfalls, like lack of diversity and temporal performance degradation, but also offers fresh insights, including the vulnerability of existing methods to basic poisoning attacks and the potential for competitive performance using much simpler alternatives. Based on insights from SVE, we craft new methods RandE and PAE for data subset selection and instance attribution, respectively, which often outperform the purported state-of-the-art at a fraction of the cost, emphasizing the practical advantages of more efficient and less complex approaches. Syed Hasan Amin Mahmood, Ming Yin 0001, Rajiv Khanna |
KDD (1) | 2 |
| 2025 | Exploring the Cost-Effectiveness of Perspective Taking in Crowdsourcing Subjective Assessment: A Case Study of Toxicity DetectionabstractXiaoni Duan, Zhuoyan Li, Chien-Ju Ho, Ming Yin. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Xiaoni Duan, Zhuoyan Li, Chien-Ju Ho, Ming Yin 0001 |
NAACL (Long Papers) | 4 |
| 2025 | AI Pilot in the Cockpit: An Investigation of Public AcceptanceabstractCrew-reducing exhibits promise for various benefits in the aviation industry. However, there is limited understanding regarding public acceptance. Using an experimental design deployed with a vignette-based online study, we investigated individuals’ negative emotion, trust, risk acceptance, and willingness to ride toward Single Pilot Operations (SPO) and Dual Pilot Operations (DPO). Results established that people preferred DPO flights, relying on affect heuristic. Specifically, people’s negative emotion associated with SPO decreases their trust, which subsequently results in lower levels of willingness to ride and risk acceptance. Furthermore, we observed that people are less likely to accept risk evoked by intelligent autonomous system in DPO flights, which likely due to their psychological model about two pilots in the cockpit. Findings from this research highlight the importance of users’ initially positive affects about intelligent equipment in the cockpit. Other theoretical and practical implications for narrowing the acceptable gap between SPO and DPO are discussed. Shan Gao 0010, Zhuoran Lu, Ming Yin 0001, Lei Wang 0019 |
Int. J. Hum. Comput. Interact. | 4 |
| 2025 | Understanding User Needs and Attitudes for Privacy Protection Tools in Online Visual Content SharingabstractVisual content shared on social media often includes sensitive elements that can threaten personal privacy. While privacy protection tools--some of which are powered by the state-of-the-art generative AI (Gen-AI) technologies--have been increasingly developed to address such visual privacy concerns by identifying sensitive elements in visual content and suggesting or applying modifications to process the visual content, the success of these tools depends on how well they meet users' nuanced needs and preferences. In this study, we conducted semi-structured interviews with 18 individuals who have either experienced or caused privacy violations in shared visual content in the past to gather first-hand perspectives on stakeholders' privacy concerns, their preferences for how to address these concerns, and their attitude toward the use of generative AI for privacy protection. Our findings highlight that sensitive elements are often not limited to direct identifiers but include contextual combinations and external information that can lead to unintended inferences. Decisions about whether and what to modify are shaped by concerns about privacy effectiveness, content value, content meaning, and emotional or social relevance, while choices around how to modify are influenced by recognition difficulty, visual content integrity, contextual consistency, atmosphere, and usability of modification methods. Participants saw Gen-AI as a promising tool for lowering editing barriers and enhancing creative control but also raised concerns about data usage, manipulation, and transparency. Importantly, we identify tensions between uploaders and depicted individuals, emphasizing the need for shared consent mechanisms and user-centered design in privacy protection. We conclude by discussing design implications for context-aware, flexible, and ethically responsible privacy tools. Chun-Wei Chiang, Harry Yizhou Tian, Ming Yin 0001 |
Proc. ACM Hum. Comput. Interact. | 3 |
| 2024 | Decoding AI's Nudge: A Unified Framework to Predict Human Behavior in AI-Assisted Decision MakingabstractWith the rapid development of AI-based decision aids, different forms of AI assistance have been increasingly integrated into the human decision making processes. To best support humans in decision making, it is essential to quantitatively understand how diverse forms of AI assistance influence humans' decision making behavior. To this end, much of the current research focuses on the end-to-end prediction of human behavior using ``black-box'' models, often lacking interpretations of the nuanced ways in which AI assistance impacts the human decision making process. Meanwhile, methods that prioritize the interpretability of human behavior predictions are often tailored for one specific form of AI assistance, making adaptations to other forms of assistance difficult. In this paper, we propose a computational framework that can provide an interpretable characterization of the influence of different forms of AI assistance on decision makers in AI-assisted decision making. By conceptualizing AI assistance as the ``nudge'' in human decision making processes, our approach centers around modelling how different forms of AI assistance modify humans' strategy in weighing different information in making their decisions. Evaluations on behavior data collected from real human decision makers show that the proposed framework outperforms various baselines in accurately predicting human behavior in AI-assisted decision making. Based on the proposed framework, we further provide insights into how individuals with different cognitive styles are nudged by AI assistance differently. Zhuoyan Li, Zhuoran Lu, Ming Yin 0001 |
AAAI | 3 |
| 2024 | The Value, Benefits, and Concerns of Generative AI-Powered Assistance in WritingabstractRecent advances in generative AI technologies like large language models raise both excitement and concerns about the future of human-AI co-creation in writing. To unpack people’s attitude towards and experience with generative AI-powered writing assistants, in this paper, we conduct an experiment to understand whether and how much value people attach to AI assistance, and how the incorporation of AI assistance in writing workflows changes people’s writing perceptions and performance. Our results suggest that people are willing to forgo financial payments to receive writing assistance from AI, especially if AI can provide direct content generation assistance and the writing task is highly creative. Generative AI-powered assistance is found to offer benefits in increasing people’s productivity and confidence in writing. However, direct content generation assistance offered by AI also comes with risks, including decreasing people’s sense of accountability and diversity in writing. We conclude by discussing the implications of our findings. Zhuoyan Li, Jing Peng 0006, Ming Yin 0001 |
CHI | 4 |
| 2024 | "Are You Really Sure?" Understanding the Effects of Human Self-Confidence Calibration in AI-Assisted Decision MakingabstractIn AI-assisted decision-making, it is crucial but challenging for humans to achieve appropriate reliance on AI. This paper approaches this problem from a human-centered perspective, “human self-confidence calibration”. We begin by proposing an analytical framework to highlight the importance of calibrated human self-confidence. In our first study, we explore the relationship between human self-confidence appropriateness and reliance appropriateness. Then in our second study, We propose three calibration mechanisms and compare their effects on humans’ self-confidence and user experience. Subsequently, our third study investigates the effects of self-confidence calibration on AI-assisted decision-making. Results show that calibrating human self-confidence enhances human-AI team performance and encourages more rational reliance on AI (in some aspects) compared to uncalibrated baselines. Finally, we discuss our main findings and provide implications for designing future AI-assisted decision-making interfaces. Shuai Ma 0005, Chuhan Shi, Ming Yin 0001, Xiaojuan Ma |
CHI | 5 |
| 2024 | How Does the Disclosure of AI Assistance Affect the Perceptions of Writing?abstractRecent advances in generative AI technologies like large language models have boosted the incorporation of AI assistance in writing workflows, leading to the rise of a new paradigm of human-AI co-creation in writing.To understand how people perceive writings that are produced under this paradigm, in this paper, we conduct an experimental study to understand whether and how the disclosure of the level and type of AI assistance in the writing process would affect people's perceptions of the writing on various aspects, including their evaluation on the quality of the writing and their ranking of different writings.Our results suggest that disclosing the AI assistance in the writing process, especially if AI has provided assistance in generating new content, decreases the average quality ratings for both argumentative essays and creative stories.This decrease in the average quality ratings often comes with an increased level of variations in different individuals' quality evaluations of the same writing.Indeed, factors such as an individual's writing confidence and familiarity with AI writing assistants are shown to moderate the impact of AI assistance disclosure on their writing quality evaluations.We also find that disclosing the use of AI assistance may significantly reduce the proportion of writings produced with AI's content generation assistance among the top-ranked writings. Independent AI editing AI generation Zhuoyan Li, Jing Peng 0006, Ming Yin 0001 |
EMNLP | 4 |
| 2024 | Designing Behavior-Aware AI to Improve the Human-AI Team Performance in AI-Assisted Decision Making
Syed Hasan Amin Mahmood, Zhuoran Lu, Ming Yin 0001 |
IJCAI | 3 |
| 2024 | Enhancing AI-Assisted Group Decision Making through LLM-Powered Devil's AdvocateabstractGroup decision making plays a crucial role in our complex and interconnected world. The rise of AI technologies has the potential to provide data-driven insights to facilitate group decision making, although it is found that groups do not always utilize AI assistance appropriately. In this paper, we aim to examine whether and how the introduction of a devil’s advocate in the AI-assisted group decision making processes could help groups better utilize AI assistance and change the perceptions of group processes during decision making. Inspired by the exceptional conversational capabilities exhibited by modern large language models (LLMs), we design four different styles of devil’s advocate powered by LLMs, varying their interactivity (i.e., interactive vs. non-interactive) and their target of objection (i.e., challenge the AI recommendation or the majority opinion within the group). Through a randomized human-subject experiment, we find evidence suggesting that LLM-powered devil’s advocates that argue against the AI model’s decision recommendation have the potential to promote groups’ appropriate reliance on AI. Meanwhile, the introduction of LLM-powered devil’s advocate usually does not lead to substantial increases in people’s perceived workload for completing the group decision making tasks, while interactive LLM-powered devil’s advocates are perceived as more collaborating and of higher quality. We conclude by discussing the practical implications of our findings. Chun-Wei Chiang, Zhuoran Lu, Zhuoyan Li, Ming Yin 0001 |
IUI | 4 |
| 2024 | Do Crowdsourced Fairness Preferences Correlate with Risk Perceptions?abstractWith the increasing prevalence of automatic decision-making systems, concerns regarding the fairness of these systems also arise. Without a universally agreed-upon definition of fairness, given an automated decision-making scenario, researchers often adopt a crowdsourced approach to solicit people’s preferences across multiple fairness definitions. However, it is often found that crowdsourced fairness preferences are highly context-dependent, making it intriguing to explore the driving factors behind these preferences. One plausible hypothesis is that people’s fairness preferences reflect their perceived risk levels for different decision-making mistakes, such that the fairness definition that equalizes across groups the type of mistakes that are perceived as most serious will be preferred. To test this conjecture, we conduct a human-subject study (N = 213) to study people’s fairness perceptions in three societal contexts. In particular, these three societal contexts differ on the expected level of risk associated with different types of decision mistakes, and we elicit both people’s fairness preferences and risk perceptions for each context. Our results show that people can often distinguish between different levels of decision risks across different societal contexts. However, we find that people’s fairness preferences do not vary significantly across the three selected societal contexts, except for within a certain subgroup of people (e.g., people with a certain racial background). As such, we observe minimal evidence suggesting that people’s risk perceptions of decision mistakes correlate with their fairness preference. These results highlight that fairness preferences are highly subjective and nuanced, and they might be primarily affected by factors other than the perceived risks of decision mistakes. Chowdhury Mohammad Rakin Haider, Chris Clifton, Ming Yin 0001 |
IUI | 3 |
| 2024 | Utilizing Human Behavior Modeling to Manipulate Explanations in AI-Assisted Decision Making: The Good, the Bad, and the ScaryabstractRecent advances in AI models have increased the integration of AI-based decision aids into the human decision making process. To fully unlock the potential of AI-assisted decision making, researchers have computationally modeled how humans incorporate AI recommendations into their final decisions, and utilized these models to improve human-AI team performance. Meanwhile, due to the ``black-box'' nature of AI models, providing AI explanations to human decision makers to help them rely on AI recommendations more appropriately has become a common practice. In this paper, we explore whether we can quantitatively model how humans integrate both AI recommendations and explanations into their decision process, and whether this quantitative understanding of human behavior from the learned model can be utilized to manipulate AI explanations, thereby nudging individuals towards making targeted decisions. Our extensive human experiments across various tasks demonstrate that human behavior can be easily influenced by these manipulated explanations towards targeted outcomes, regardless of the intent being adversarial or benign. Furthermore, individuals often fail to detect any anomalies in these explanations, despite their decisions being affected by them. Zhuoyan Li, Ming Yin 0001 |
NeurIPS | 2 |
| 2024 | ToneCheck: Unveiling the Impact of Dialects in Privacy PolicyabstractUsers frequently struggle to decipher privacy policies, facing challenges due to the legalese often present in privacy policies, leaving trust and comprehension shrouded in ambiguity. This study dives into the transformative power of language, exploring how different linguistic tones can bridge the gap between legal, technical jargon, and genuine user engagement-through a comparative analysis involving diverse focus groups, immersing them in three distinct policy variations: legalistic, casual, and empathetic. We explored how these tones reshape the user experience and bridge the gap between legal discourse and comprehension. Analysis of the data revealed significant associations between linguistic tone and user trust and comprehension. The adoption of an empathetic tone significantly enhanced user trust, as evidenced by a 40.4% increase compared to alternative language styles. This preference highlights the human desire for genuine connection, even in the intricate domain of data privacy. Furthermore, comprehension indices arise for both empathetic and casual tones, leaving legalistic language lagging far behind. This suggests a clear path towards user-friendly policies, where clarity exceeds complexity. Our exploration goes beyond mere compliance. We illustrate the complex gap between subtle linguistic shifts and user perception. By deciphering the language that resonates with trust and understanding, We plant the seeds for the development of privacy policies that not only meet legal requirements but also enhance user trust and comprehension. Jay Barot, Ali A. Allami, Ming Yin 0001, Dan Lin 0001 |
SACMAT | 3 |
| 2024 | Does More Advice Help? The Effects of Second Opinions in AI-Assisted Decision MakingabstractAI assistance in decision-making has become popular, yet people's inappropriate reliance on AI often leads to unsatisfactory human-AI collaboration performance. In this paper, through three pre-registered, randomized human subject experiments, we explore whether and how the provision of second opinions may affect decision-makers' behavior and performance in AI-assisted decision-making. We find that if both the AI model's decision recommendation and a second opinion are always presented together, decision-makers reduce their over-reliance on AI while increase their under-reliance on AI, regardless whether the second opinion is generated by a peer or another AI model. However, if decision-makers have the control to decide when to solicit a peer's second opinion, we find that their active solicitations of second opinions have the potential to mitigate over-reliance on AI without inducing increased under-reliance in some cases. We conclude by discussing the implications of our findings for promoting effective human-AI collaborations in decision-making. Zhuoran Lu, Dakuo Wang, Ming Yin 0001 |
Proc. ACM Hum. Comput. Interact. | 3 |
| 2023 | Modeling Human Trust and Reliance in AI-Assisted Decision Making: A Markovian ApproachabstractThe increased integration of artificial intelligence (AI) technologies in human workflows has resulted in a new paradigm of AI-assisted decision making, in which an AI model provides decision recommendations while humans make the final decisions. To best support humans in decision making, it is critical to obtain a quantitative understanding of how humans interact with and rely on AI. Previous studies often model humans' reliance on AI as an analytical process, i.e., reliance decisions are made based on cost-benefit analysis. However, theoretical models in psychology suggest that the reliance decisions can often be driven by emotions like humans' trust in AI models. In this paper, we propose a hidden Markov model to capture the affective process underlying the human-AI interaction in AI-assisted decision making, by characterizing how decision makers adjust their trust in AI over time and make reliance decisions based on their trust. Evaluations on real human behavior data collected from human-subject experiments show that the proposed model outperforms various baselines in accurately predicting humans' reliance behavior in AI-assisted decision making. Based on the proposed model, we further provide insights into how humans' trust and reliance dynamics in AI-assisted decision making is influenced by contextual factors like decision stakes and their interaction experiences. Zhuoyan Li, Zhuoran Lu, Ming Yin 0001 |
AAAI | 3 |
| 2023 | How does Value Similarity affect Human Reliance in AI-Assisted Ethical Decision Making?abstractThis paper explores the impact of value similarity between humans and AI on human reliance in the context of AI-assisted ethical decision-making. Using kidney allocation as a case study, we conducted a randomized human-subject experiment where workers were presented with ethical dilemmas in various conditions, including no AI recommendations, recommendations from a similar AI, and recommendations from a dissimilar AI. We found that recommendations provided by a dissimilar AI had a higher overall effect on human decisions than recommendations from a similar AI. However, when humans and AI disagreed, participants were more likely to change their decisions when provided with recommendations from a similar AI. The effect was not due to humans’ perceptions of the AI being similar, but rather due to the AI displaying similar ethical values through its recommendations. We also conduct a preliminary analysis on the relationship between value similarity and trust, and potential shifts in ethical preferences at the population-level. Saumik Narayanan, Chien-Ju Ho, Ming Yin 0001 |
AIES | 4 |
| 2023 | Are Two Heads Better Than One in AI-Assisted Decision Making? Comparing the Behavior and Performance of Groups and Individuals in Human-AI Collaborative Recidivism Risk AssessmentabstractWith the prevalence of AI assistance in decision making, a more relevant question to ask than the classical question of “are two heads better than one?’’ is how groups’ behavior and performance in AI-assisted decision making compare with those of individuals’. In this paper, we conduct a case study to compare groups and individuals in human-AI collaborative recidivism risk assessment along six aspects, including decision accuracy and confidence, appropriateness of reliance on AI, understanding of AI, decision-making fairness, and willingness to take accountability. Our results highlight that compared to individuals, groups rely on AI models more regardless of their correctness, but they are more confident when they overturn incorrect AI recommendations. We also find that groups make fairer decisions than individuals according to the accuracy equality criterion, and groups are willing to give AI more credit when they make correct decisions. We conclude by discussing the implications of our work. Chun-Wei Chiang, Zhuoran Lu, Zhuoyan Li, Ming Yin 0001 |
CHI | 4 |
| 2023 | Who Should I Trust: AI or Myself? Leveraging Human and AI Correctness Likelihood to Promote Appropriate Trust in AI-Assisted Decision-MakingabstractIn AI-assisted decision-making, it is critical for human decision-makers to know when to trust AI and when to trust themselves. However, prior studies calibrated human trust only based on AI confidence indicating AI’s correctness likelihood (CL) but ignored humans’ CL, hindering optimal team decision-making. To mitigate this gap, we proposed to promote humans’ appropriate trust based on the CL of both sides at a task-instance level. We first modeled humans’ CL by approximating their decision-making models and computing their potential performance in similar instances. We demonstrated the feasibility and effectiveness of our model via two preliminary studies. Then, we proposed three CL exploitation strategies to calibrate users’ trust explicitly/implicitly in the AI-assisted decision-making process. Results from a between-subjects experiment (N=293) showed that our CL exploitation strategies promoted more appropriate human trust in AI, compared with only using AI confidence. We further provided practical implications for more human-compatible AI-assisted decision-making. Shuai Ma 0005, Chengbo Zheng, Chuhan Shi, Ming Yin 0001, Xiaojuan Ma |
CHI | 6 |
| 2023 | Watch Out for Updates: Understanding the Effects of Model Explanation Updates in AI-Assisted Decision MakingabstractAI explanations have been increasingly used to help people better utilize AI recommendations in AI-assisted decision making. While AI explanations may change over time due to updates of the AI model, little is known about how these changes may affect people’s perceptions and usage of the model. In this paper, we study how varying levels of similarity between the AI explanations before and after a model update affects people’s trust in and satisfaction with the AI model. We conduct randomized human-subject experiments on two decision making contexts where people have different levels of domain knowledge. Our results show that changes in AI explanation during the model update do not affect people’s tendency to adopt AI recommendations. However, they may change people’s subjective trust in and satisfaction with the AI model via changing both their perceived model accuracy and perceived consistency of AI explanations with their prior knowledge. Ming Yin 0001 |
CHI | 2 |
| 2023 | Synthetic Data Generation with Large Language Models for Text Classification: Potential and LimitationsabstractThe collection and curation of high-quality training data is crucial for developing text classification models with superior performance, but it is often associated with significant costs and time investment.Researchers have recently explored using large language models (LLMs) to generate synthetic datasets as an alternative approach.However, the effectiveness of the LLM-generated synthetic data in supporting model training is inconsistent across different classification tasks.To better understand factors that moderate the effectiveness of the LLMgenerated synthetic data, in this study, we look into how the performance of models trained on these synthetic data may vary with the subjectivity of classification.Our results indicate that subjectivity, at both the task level and instance level, is negatively associated with the performance of the model trained on synthetic data.We conclude by discussing the implications of our work on the potential and limitations of leveraging LLM for synthetic data generation 1 . Zhuoyan Li, Hangxiao Zhu, Zhuoran Lu, Ming Yin 0001 |
EMNLP | 4 |
| 2023 | Strategic Adversarial Attacks in AI-assisted Decision Making to Reduce Human Trust and RelianceabstractWith the increased integration of AI technologies in human decision making processes, adversarial attacks on AI models become a greater concern than ever before as they may significantly hurt humans’ trust in AI models and decrease the effectiveness of human-AI collaboration. While many adversarial attack methods have been proposed to decrease the performance of an AI model, limited attention has been paid on understanding how these attacks will impact the human decision makers interacting with the model, and accordingly, how to strategically deploy adversarial attacks to maximize the reduction of human trust and reliance. In this paper, through a human-subject experiment, we first show that in AI-assisted decision making, the timing of the attacks largely influences how much humans decrease their trust in and reliance on AI—the decrease is particularly salient when attacks occur on decision making tasks that humans are highly confident themselves. Based on these insights, we next propose an algorithmic framework to infer the human decision maker’s hidden trust in the AI model and dynamically decide when the attacker should launch an attack to the model. Our evaluations show that following the proposed approach, attackers deploy more efficient attacks and achieve higher utility than adopting other baseline strategies. Zhuoran Lu, Zhuoyan Li, Chun-Wei Chiang, Ming Yin 0001 |
IJCAI | 4 |
| 2023 | The Effects of AI Biases and Explanations on Human Decision Fairness: A Case Study of Bidding in Rental Housing MarketsabstractThe use of AI-based decision aids in diverse domains has inspired many empirical investigations into how AI models’ decision recommendations impact humans’ decision accuracy in AI-assisted decision making, while explorations on the impacts on humans’ decision fairness are largely lacking despite their clear importance. In this paper, using a real-world business decision making scenario—bidding in rental housing markets—as our testbed, we present an experimental study on understanding how the bias level of the AI-based decision aid as well as the provision of AI explanations affect the fairness level of humans’ decisions, both during and after their usage of the decision aid. Our results suggest that when people are assisted by an AI-based decision aid, both the higher level of racial biases the decision aid exhibits and surprisingly, the presence of AI explanations, result in more unfair human decisions across racial groups. Moreover, these impacts are partly made through triggering humans’ “disparate interactions” with AI. However, regardless of the AI bias level and the presence of AI explanations, when people return to make independent decisions after their usage of the AI-based decision aid, their decisions no longer exhibit significant unfairness across racial groups. Ming Yin 0001 |
IJCAI | 3 |
| 2022 | Understanding Decision Subjects' Fairness Perceptions and Retention in Repeated Interactions with AI-Based Decision SystemsabstractThe wide application of AI-based decision systems in many high-stake domains has raised concerns regarding fairness of these systems. As these systems will lead to real-life consequences to people who are subject to their decisions, understanding what these decision subjects perceive as a fair or unfair system is of vital importance. In this paper, we extend prior work in this direction by taking a perspective of repeated interactions---We ask that when decision subjects interact with an AI-based decision system repeatedly and can strategically respond to the system by determining whether to stay in the system, what factors will affect the decision subjects' fairness perceptions and retention in the system and how. To answer these questions, we conducted two randomized human-subject experiments in the context of an AI-based loan lending system. Our results suggest that in repeated interactions with the AI-based decision system, overall, decision subjects' fairness perceptions and retention in the system are significantly affected by whether the system is in favor of the group that subjects themselves belong to, rather than whether the system treats different groups in an unbiased way. However, decision subjects with different qualification levels have different reactions to the AI system's biased treatment across groups or the AI system's tendency to favor/disfavor their own group. Finally, we also find that while subjects' retention in the AI-based decision system is largely driven by their own prospects of receiving the favorable decision from the system, their fairness perceptions of the system is influenced by the system's treatment to people in other groups in a complex way. Meric Altug Gemalmaz, Ming Yin 0001 |
AIES | 2 |
| 2022 | Towards Better Detection of Biased Language with Scarce, Noisy, and Biased AnnotationsabstractBiased language is prevalent in today's online social media. To reduce the amount of online biased language, one critical first step is to accurately detect such biased language, ideally automatically. This is a challenging problem, however, as the annotated data necessary for training a biased language classifier is either scarce and costly (e.g., when collected from experts), or noisy and potentially biased on their own (e.g., when collected from crowd workers). The biased language classifier built based on these annotations may thus be inaccurate, and sometimes unfair (e.g., have systematic accuracy disparities across texts with different political leanings). In this paper, we propose a novel method, CLEARE, for biased language detection, in which we utilize self-supervised contrastive learning to enhance the biased language classifier---we learn a robust encoder of the textual data through solving a min-max optimization problem, so that the encoder could help achieve the best classification performance even if the worst data augmentation strategy is selected. Extensive evaluations suggest that CLEARE shows substantial improvements compared to the state-of-art biased language detection methods on several benchmark datasets, in terms of improving both the accuracy and the fairness of the detection. Zhuoyan Li, Zhuoran Lu, Ming Yin 0001 |
AIES | 3 |
| 2022 | How Does Predictive Information Affect Human Ethical Preferences?abstractArtificial intelligence (AI) has been increasingly involved in decision making in high-stakes domains, including loan applications, employment screening, and assistive clinical decision making. Meanwhile, involving AI in these high-stake decisions has created ethical concerns on how to balance different trade-offs to respect human values. One approach for aligning AIs with human values is to elicit human ethical preferences and incorporate this information in the design of computer systems. In this work, we explore how human ethical preferences are impacted by the information shown to humans during elicitation. In particular, we aim to provide a contrast between verifiable information (e.g., patient demographics or blood test results) and predictive information (e.g., the probability of organ transplant success). Using kidney transplant allocation as a case study, we conduct a randomized experiment to elicit human ethical preferences on scarce resource allocation to understand how human ethical preferences are impacted by the verifiable and predictive information. We find that the presence of predictive information significantly changes how humans take into account other verifiable information in their ethical preferences. We also find that the source of the predictive information (e.g., whether the predictions are made by AI or human doctors) plays a key role in how humans incorporate the predictive information into their own ethical judgements. Saumik Narayanan, Chien-Ju Ho, Ming Yin 0001 |
AIES | 5 |
| 2022 | When Confidence Meets Accuracy: Exploring the Effects of Multiple Performance Indicators on Trust in Machine Learning ModelsabstractPrevious research shows that laypeople’s trust in a machine learning model can be affected by both performance measurements of the model on the aggregate level and performance estimates on individual predictions. However, it is unclear how people would trust the model when multiple performance indicators are presented at the same time. We conduct an exploratory human-subject experiment to answer this question. We find that while the level of model confidence significantly affects people’s belief in model accuracy, both the model’s stated and observed accuracy generally have a larger impact on people’s willingness to follow the model’s predictions as well as their self-reported levels of trust in the model, especially after observing the model’s performance in practice. We hope the empirical evidence reported in this work could open doors to further studies to advance understanding of how people perceive, process, and react to performance-related information of machine learning. Amy Rechkemmer, Ming Yin 0001 |
CHI | 2 |
| 2022 | Exploring the Effects of Machine Learning Literacy Interventions on Laypeople's Reliance on Machine Learning ModelsabstractToday, machine learning (ML) technologies have penetrated almost every aspect of people’s lives, yet public understandings of these technologies are often limited. This highlights the urgent need of designing effective methods to increase people’s machine learning literacy, as the lack of relevant knowledge may result in people’s inappropriate usage of machine learning technologies. In this paper, we focus on an ML-assisted decision-making setting and conduct a human-subject randomized experiment to explore how providing different types of user tutorials as the machine learning literacy interventions can influence laypeople’s reliance on ML models, on both in-distribution and out-of-distribution examples. We vary the existence, interactivity and scope of the user tutorial across different treatments in our experiment. Our results show that user tutorials, when presented in appropriate forms, can help some people rely on ML models more appropriately. For example, for those individuals who have relatively high ability in solving the decision-making task themselves, receiving a user tutorial that is interactive and addresses the specific ML model to be used allows them to reduce their over-reliance on the ML model when they could outperform the model. In contrast, low-performing individuals’ reliance on the ML model is not affected by the presence or the type of user tutorial. Finally, we also find that people perceive the interactive tutorial to be more understandable and slightly more useful. We conclude by discussing the design implications of our study. Chun-Wei Chiang, Ming Yin 0001 |
IUI | 2 |
| 2022 | The Influences of Task Design on Crowdsourced Judgement: A Case Study of Recidivism Risk EvaluationabstractCrowdsourcing is widely used to solicit judgement from people in diverse applications ranging from evaluating information quality to rating gig worker performance. To encourage the crowd to put in genuine effort in the judgement tasks, various ways to structure and organize these tasks have been explored, though the understandings of how these task design choices influence the crowd’s judgement are still largely lacking. In this paper, using recidivism risk evaluation as an example, we conduct a randomized experiment to examine the effects of two common designs of crowdsourcing judgement tasks—encouraging the crowd to deliberate and providing feedback to the crowd—on the quality, strictness, and fairness of the crowd’s recidivism risk judgements. Our results show that different designs of the judgement tasks significantly affect the strictness of the crowd’s judgements. Moreover, task designs also have the potential to significantly influence how fairly the crowd judges defendants from different racial groups, on those cases where the crowd exhibits substantial in-group bias. Finally, we find that the impacts of task designs on the judgement also vary with the crowd workers’ own characteristics, such as their cognitive reflection levels. Together, these results highlight the importance of obtaining a nuanced understanding on the relationship between task designs and properties of the crowdsourced judgements. Xiaoni Duan, Chien-Ju Ho, Ming Yin 0001 |
WWW | 3 |
| 2022 | Will You Accept the AI Recommendation? Predicting Human Behavior in AI-Assisted Decision MakingabstractInternet users make numerous decisions online on a daily basis. With the rapid advances in AI recently, AI-assisted decision making—in which an AI model provides decision recommendations and confidence, while the humans make the final decisions—has emerged as a new paradigm of human-AI collaboration. In this paper, we aim at obtaining a quantitative understanding of whether and when would human decision makers adopt the AI model’s recommendations. We define a space of human behavior models by decomposing the human decision maker’s cognitive process in each decision-making task into two components: the utility component (i.e., evaluate the utility of different actions) and the selection component (i.e., select an action to take), and we perform a systematic search in the model space to identify the model that fits real-world human behavior data the best. Our results highlight that in AI-assisted decision making, human decision makers’ utility evaluation and action selection are influenced by their own judgement and confidence on the decision-making task. Further, human decision makers exhibit a tendency to distort the decision confidence in utility evaluations. Finally, we also analyze the differences in humans’ adoption behavior of AI recommendations as the stakes of the decisions vary. Zhuoran Lu, Ming Yin 0001 |
WWW | 3 |
| 2022 | The Effects of AI-based Credibility Indicators on the Detection and Spread of Misinformation under Social InfluenceabstractMisinformation on social media has become a serious concern. Marking news stories with credibility indicators, possibly generated by an AI model, is one way to help people combat misinformation. In this paper, we report the results of two randomized experiments that aim to understand the effects of AI-based credibility indicators on people's perceptions of and engagement with the news, when people are under social influence such that their judgement of the news is influenced by other people. We find that the presence of AI-based credibility indicators nudges people into aligning their belief in the veracity of news with the AI model's prediction regardless of its correctness, thereby changing people's accuracy in detecting misinformation. However, AI-based credibility indicators show limited impacts on influencing people's engagement with either real news or fake news when social influence exists. Finally, it is shown that when social influence is present, the effects of AI-based credibility indicators on the detection and spread of misinformation are larger as compared to when social influence is absent, when these indicators are provided to people before they form their own judgements about the news. We conclude by providing implications for better utilizing AI to fight misinformation. Zhuoran Lu, Patrick Li, Ming Yin 0001 |
Proc. ACM Hum. Comput. Interact. | 4 |
| 2022 | Understanding the Microtask Crowdsourcing Experience for Workers with Disabilities: A Comparative ViewabstractMicrotask crowdsourcing holds great potential as an employment opportunity with the flexibility and anonymity that individuals with disability may require. Though prior research has explored the accessibility of crowd work, the lived crowd work experiences of the broader community of workers with disability are still largely under-explored, especially when it comes to how their experiences are similar to or different from the experiences of workers without disability. In this work, we aim to obtain a deeper understanding of the microtask crowdsourcing experience for people with disabilities, especially regarding their financial and social experiences of participating in crowd work, along with the benefits and challenges that they encounter through this work. Specifically, we first surveyed 1,200 crowd workers both with and without disability about their experiences using the Amazon Mechanical Turk platform, and the differences we found inspired the design of a follow-up survey to gain greater understanding of the crowd work experience for workers with disability. Our findings reveal that workers with disability receive unique benefits from performing crowd work, such as a greater sense of purpose, but also encounter many challenges, such as completing tasks on time and earning a livable wage, causing them to turn to online communities for assistance. Although many of the challenges they face are not unique to crowd workers with disability, workers with disability may be disproportionately impacted by these challenges. From our findings, we provide implications for crowd platforms, as well as the gig economy as a whole, that seek to promote greater consideration of workers with a diverse range of conditions to create a more valuable work experience for them. Amy Rechkemmer, Ming Yin 0001 |
Proc. ACM Hum. Comput. Interact. | 2 |
| 2022 | A Computer Vision Approach for Estimating Lifting Load Contributors to Injury RiskabstractSafety practitioners widely use the lifting index (LI) to determine workers’ lifting risk but are hampered by the difficulties of estimating the lifting load without intervention or intrusive sensors. This study proposes a computer vision method for estimating the LI across varying lifting loads. The proposed method can also predict the Brog rating of perceived exertion (RPE), a measure associated with the lifting load. A controlled lifting experiment was conducted to demonstrate the approach. Thirty participants performed 2176 lifting tasks at three LI levels. These levels were controlled by varying the lifting load and fixing other task variables (e.g., the lifting distance). The proposed method combined the pose estimation (OpenPose) and the optical flow estimation (SelFlow) techniques for extracting the participants’ body motion and posture features; a facial expression recognition algorithm (OpenFace) built upon the facial action unit coding system (FACS) was used to extract the participants’ facial features. The extracted features were combined and used to develop prediction models. The best-performing model was an integration of the 1-D convolutional neural network and the long short-term memory network. It achieved an area under curve of 0.890 in classifying the LI and a root mean square of 2.264 in predicting the participants’ RPE. Critical indicators were identified by investigating the contribution of the features through interpretable machine learning techniques. In summary, this study demonstrates a nonintrusive method for lifting risk assessment and discovers behavioral indicators that predict changes in the LI and RPE due to varying loads. Guoyang Zhou, Vaneet Aggarwal, Ming Yin 0001, Denny Yu |
IEEE Trans. Hum. Mach. Syst. | 3 |
| 2022 | Effects of Explanations in AI-Assisted Decision Making: Principles and ComparisonsabstractRecent years have witnessed the growing literature in empirical evaluation of explainable AI (XAI) methods. This study contributes to this ongoing conversation by presenting a comparison on the effects of a set of established XAI methods in AI-assisted decision making. Based on our review of previous literature, we highlight three desirable properties that ideal AI explanations should satisfy — improve people’s understanding of the AI model, help people recognize the model uncertainty, and support people’s calibrated trust in the model. Through three randomized controlled experiments, we evaluate whether four types of common model-agnostic explainable AI methods satisfy these properties on two types of AI models of varying levels of complexity, and in two kinds of decision making contexts where people perceive themselves as having different levels of domain expertise. Our results demonstrate that many AI explanations do not satisfy any of the desirable properties when used on decision making tasks that people have little domain expertise in. On decision making tasks that people are more knowledgeable, the feature contribution explanation is shown to satisfy more desiderata of AI explanations, even when the AI model is inherently complex. We conclude by discussing the implications of our study for improving the design of XAI methods to better support human decision making, and for advancing more rigorous empirical evaluation of XAI methods. Ming Yin 0001 |
ACM Trans. Interact. Intell. Syst. | 2 |
| 2021 | Human Reliance on Machine Learning Models When Performance Feedback is Limited: Heuristics and RisksabstractThis paper addresses an under-explored problem of AI-assisted decision-making: when objective performance information of the machine learning model underlying a decision aid is absent or scarce, how do people decide their reliance on the model? Through three randomized experiments, we explore the heuristics people may use to adjust their reliance on machine learning models when performance feedback is limited. We find that the level of agreement between people and a model on decision-making tasks that people have high confidence in significantly affects reliance on the model if people receive no information about the model’s performance, but this impact will change after aggregate-level model performance information becomes available. Furthermore, the influence of high confidence human-model agreement on people’s reliance on a model is moderated by people’s confidence in cases where they disagree with the model. We discuss potential risks of these heuristics, and provide design implications on promoting appropriate reliance on AI. Zhuoran Lu, Ming Yin 0001 |
CHI | 2 |
| 2021 | Accounting for Confirmation Bias in Crowdsourced Label AggregationabstractCollecting large-scale human-annotated datasets via crowdsourcing to train and improve automated models is a prominent human-in-the-loop approach to integrate human and machine intelligence. However, together with their unique intelligence, humans also come with their biases and subjective beliefs, which may influence the quality of the annotated data and negatively impact the effectiveness of the human-in-the-loop systems. One of the most common types of cognitive biases that humans are subject to is the confirmation bias, which is people's tendency to favor information that confirms their existing beliefs and values. In this paper, we present an algorithmic approach to infer the correct answers of tasks by aggregating the annotations from multiple crowd workers, while taking workers' various levels of confirmation bias into consideration. Evaluations on real-world crowd annotations show that the proposed bias-aware label aggregation algorithm outperforms baseline methods in accurately inferring the ground-truth labels of different tasks when crowd workers indeed exhibit some degree of confirmation bias. Through simulations on synthetic data, we further identify the conditions when the proposed algorithm has the largest advantages over baseline methods. Meric Altug Gemalmaz, Ming Yin 0001 |
IJCAI | 2 |
| 2021 | Exploring the Effects of Goal Setting When Training for Complex Crowdsourcing Tasks (Extended Abstract)abstractTraining is one way of enabling novice workers to work on complex crowdsourcing tasks. Based on goal setting theory in psychology, we conduct a randomized experiment to study whether and how setting different goals---including performance goal, learning goal, and behavioral goal---when training workers for a complex crowdsourcing task affects workers' learning perception, learning gain, and post-training performance. We find that setting different goals during training significantly affects workers' learning perception, but does not have an effect on learning gain or post-training performance. Further, exploratory analysis helps shed light on when and why various goals may or may not work in the crowdsourcing context. Amy Rechkemmer, Ming Yin 0001 |
IJCAI | 2 |
| 2021 | Are Explanations Helpful? A Comparative Study of the Effects of Explanations in AI-Assisted Decision-MakingabstractThis paper contributes to the growing literature in empirical evaluation of explainable AI (XAI) methods by presenting a comparison on the effects of a set of established XAI methods in AI-assisted decision making. Specifically, based on our review of previous literature, we highlight three desirable properties that ideal AI explanations should satisfy—improve people’s understanding of the AI model, help people recognize the model uncertainty, and support people’s calibrated trust in the model. Through randomized controlled experiments, we evaluate whether four types of common model-agnostic explainable AI methods satisfy these properties on two types of decision making tasks where people perceive themselves as having different levels of domain expertise in (i.e., recidivism prediction and forest cover prediction). Our results show that the effects of AI explanations are largely different on decision making tasks where people have varying levels of domain expertise in, and many AI explanations do not satisfy any of the desirable properties for tasks that people have little domain expertise in. Further, for decision making tasks that people are more knowledgeable, feature contribution explanation is shown to satisfy more desiderata of AI explanations, while the explanation that is considered to resemble how human explain decisions (i.e., counterfactual explanation) does not seem to improve calibrated trust. We conclude by discussing the implications of our study for improving the design of XAI methods to better support human decision making. Ming Yin 0001 |
IUI | 2 |
| 2021 | Video-based AI Decision Support System for Lifting Risk AssessmentabstractPhysical injuries induced by lifting are commonly reported in the workplace. Early risk detection is essential for reducing lifting injuries but requires trained observers to perform assessments manually. Machine learning and computer vision techniques have been proposed to aid ergonomists in lifting risk assessments. However, these methods may not bring the practitioners into the decision-making process and frequently not interpretable to practitioners. We conducted a user study with a proposed risk assessment system that consists of a prediction module, explanation module, and prototype user interface. The prediction module consists of a logistics regression model capable of distinguishing the injury risk levels induced by different levels of force exertion in common lifting tasks. The logistics regression model makes predictions based on explainable body motion, posture, and facial features extracted through computer vision techniques. The explanation module makes up of explainable AI techniques. Specifically, a surrogate model provides local explanations for presenting how the system makes each prediction to users. The prototype interface presents the system’s predictions and explanations. A usability study shows that the proposed system increases crowd-workers’ and domain scholars’ performance in assessing workers’ injury risks in lifting. Furthermore, the usability study also shows that the proposed system increases their confidence in the assessment tasks when the system’s evaluations agree with their subjective evaluations. Guoyang Zhou, Vaneet Aggarwal, Ming Yin 0001, Denny Yu |
SMC | 3 |
| 2020 | Does Exposure to Diverse Perspectives Mitigate Biases in Crowdwork? An Explorative StudyabstractEarlier research has shown the promise of enabling worker interactions in crowdwork to mitigate worker biases and improve the quality of crowdwork. In this study, we focus on one characteristic of the interacting workers that may influence the effectiveness of worker interactions in enhancing crowdwork—the diversity of perspectives that the interacting workers bring together—and we explore whether and how interactions between a set of workers holding different perspectives can help mitigate biases in crowdwork. Through two sets of randomized experiments, we find that whether interactions between workers with different perspectives can help mitigate biases in crowdwork depends on task properties. We also find no conclusive evidence in our experimental settings suggesting that interactions among workers with diverse perspectives reduce biases in crowdwork to a larger extent compared to interactions among workers with similar perspectives. Xiaoni Duan, Chien-Ju Ho, Ming Yin 0001 |
HCOMP | 3 |
| 2020 | Motivating Novice Crowd Workers through Goal Setting: An Investigation into the Effects on Complex Crowdsourcing Task TrainingabstractTraining workers within a task is one way of enabling novice workers, who may lack domain knowledge or experience, to work on complex crowdsourcing tasks. Based on goal setting theory in psychology, we conduct a randomized experiment to study whether and how setting different goals—including performance goal, learning goal, and behavioral goal—when training workers for a complex crowdsourcing task affects workers’ learning perception, learning gain, and post-training performance. We find that setting different goals during training significantly affects workers’ learning perception, but overall does not have an effect on learning gain or post-training performance. However, higher levels of learning gain can be obtained when setting learning goals for workers who are highly learning-oriented. Additionally, giving workers a challenging behavioral goal can nudge them to adopt desirable behavior meant to improve learning and performance, though the adoption of such behavior does not lead to as much improvement as when the worker decides to take part in the behavior themselves. We conclude by discussing the lessons we’ve learned on how to effectively utilize goals in complex crowdsourcing task training. Amy Rechkemmer, Ming Yin 0001 |
HCOMP | 2 |
| 2020 | Crowdsourcing Detection of Sampling Biases in Image DatasetsabstractDespite many exciting innovations in computer vision, recent studies reveal a number of risks in existing computer vision systems, suggesting results of such systems may be unfair and untrustworthy. Many of these risks can be partly attributed to the use of a training image dataset that exhibits sampling biases and thus does not accurately reflect the real visual world. Being able to detect potential sampling biases in the visual dataset prior to model development is thus essential for mitigating the fairness and trustworthy concerns in computer vision. In this paper, we propose a three-step crowdsourcing workflow to get humans into the loop for facilitating bias discovery in image datasets. Through two sets of evaluation studies, we find that the proposed workflow can effectively organize the crowd to detect sampling biases in both datasets that are artificially created with designed biases and real-world image datasets that are widely used in computer vision research and system development. Xiao Hu 0004, Anirudh Vegesana, Somesh Dube, Kaiwen Yu, Gore Kao, Shuo-Han Chen, Yung-Hsiang Lu, George K. Thiruvathukal, Ming Yin 0001 |
WWW | 10 |
| 2019 | Understanding the Effect of Accuracy on Trust in Machine Learning ModelsabstractWe address a relatively under-explored aspect of human-computer interaction: people's abilities to understand the relationship between a machine learning model's stated performance on held-out data and its expected performance post deployment. We conduct large-scale, randomized human-subject experiments to examine whether laypeople's trust in a model, measured in terms of both the frequency with which they revise their predictions to match those of the model and their self-reported levels of trust in the model, varies depending on the model's stated accuracy on held-out data and on its observed accuracy in practice. We find that people's trust in a model is affected by both its stated accuracy and its observed accuracy, and that the effect of stated accuracy can change depending on the observed accuracy. Our work relates to recent research on interpretable machine learning, but moves beyond the typical focus on model internals, exploring a different component of the machine learning pipeline. Ming Yin 0001, Jennifer Wortman Vaughan, Hanna M. Wallach |
CHI | 1 |
| 2019 | Leveraging Peer Communication to Enhance CrowdsourcingabstractCrowdsourcing has become a popular tool for large-scale data collection where it is often assumed that crowd workers complete the work independently. In this paper, we relax such independence property and explore the usage of peer communication-a kind of direct interactions between workers-in crowdsourcing. In particular, in the crowdsourcing setting with peer communication, a pair of workers are asked to complete the same task together by first generating their initial answers to the task independently and then freely discussing the task with each other and updating their answers after the discussion. We first experimentally examine the effects of peer communication on individual microtasks. Our results conducted on three types of tasks consistently suggest that work quality is significantly improved in tasks with peer communication compared to tasks where workers complete the work independently. We next explore how to utilize peer communication to optimize the requester's total utility while taking into account higher data correlation and higher cost introduced by peer communication. In particular, we model the requester's online decision problem of whether and when to use peer communication in crowdsourcing as a constrained Markov decision process which maximizes the requester's total utility under budget constraints. Our proposed approach is empirically shown to bring higher total utility compared to baseline approaches. Ming Yin 0001, Chien-Ju Ho |
WWW | 2 |
| 2019 | Understanding the Skill Provision in Gig Economy from A Network Perspective: A Case Study of FiverrabstractThe recent emergence of gig economy facilitates the exchange of skilled labor by allowing workers to showcase and sell their skills to a global market. Despite the recent effort on thoroughly examining who workers in the gig economy are and what their experience in the gig economy are like, our knowledge on how exactly do workers provide their skills in gig economy, and how worker's strategies on skill provision and expansion relate to their success in gig economy is still lacking. In this paper, we conduct a case study on a prominent gig economy platform, Fiverr.com, to better understand the provision of skills on it through large-scale, data-driven analysis. In particular, we propose the concept of "skill space" from a network perspective to characterize the relationship between different skills by measuring how frequently workers provide different skills together. Through our analysis, we reveal interesting patterns in worker's provision of skills on Fiverr. We then show how these patterns change over time and differ across subgroups of workers with different characteristics. In addition, we find that providing a set of skills that are highly related with each other correlates with a better overall performance in gig economy, and when workers expand their skillsets, expanding to a new skill that is highly-related to the existing skills takes less time and is associated with better performance on the new skill. We conclude by discussing the implications of our findings for gig economy workers and platform in general. Keman Huang, Jinhui Yao, Ming Yin 0001 |
Proc. ACM Hum. Comput. Interact. | 3 |
| 2018 | Running Out of Time: The Impact and Value of Flexibility in On-Demand CrowdworkabstractWith a seemingly endless stream of tasks, on-demand labor markets appear to offer workers flexibility in when and how much they work. This research argues that platforms afford workers far less flexibility than widely believed. A large part of the "inflexibility" comes from tight deadlines imposed on tasks, leaving workers little control over their work schedules. We experimentally examined the impact of offering workers control of their time in on-demand crowdwork. We found that granting higher "in-task flexibility" dramatically affected the temporal dynamics of worker behavior and produced a larger amount of work with similar quality. In a second experiment, we measured the compensating differential and found that workers would give up significant compensation to control their time, indicating workers attach substantial value to in-task flexibility. Our results suggest that designing tasks which give workers direct control of their time within tasks benefits both buyers and sellers of on-demand crowdwork. Ming Yin 0001, Siddharth Suri, Mary L. Gray |
CHI | 1 |
| 2016 | The Communication Network Within the CrowdabstractSince its inception, crowdsourcing has been considered a black-box approach to solicit labor from a crowd of workers. Furthermore, the "crowd" has been viewed as a group of independent workers dispersed all over the world. Recent studies based on in-person interviews have opened up the black box and shown that the crowd is not a collection of independent workers, but instead that workers communicate and collaborate with each other. Put another way, prior work has shown the existence of edges between workers. We build on and extend this discovery by mapping the entire communication network of workers on Amazon Mechanical Turk, a leading crowdsourcing platform. We execute a task in which over 10,000 workers from across the globe self-report their communication links to other workers, thereby mapping the communication network among workers. Our results suggest that while a large percentage of workers indeed appear to be independent, there is a rich network topology over the rest of the population. That is, there is a substantial communication network within the crowd. We further examine how online forum usage relates to network topology, how workers communicate with each other via this network, how workers' experience levels relate to their network positions, and how U.S. workers differ from international workers in their network characteristics. We conclude by discussing the implications of our findings for requesters, workers, and platform providers like Amazon. Ming Yin 0001, Mary L. Gray, Siddharth Suri, Jennifer Wortman Vaughan |
WWW | 1 |