VLDB 2026 Research / reviewers in the wild / expert
Qingzi Vera Liao
dblp:01/7985 · also Q. Vera Liao, Vera Liao, Vera Q. Liao
· DBLP profile ↗
58ranked-venue papers
15as first author
32since 2021 · last 2026
0000-0003-4543-7196ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 46 · 13 first-author · 25 since 2021Artificial intelligence and machine learning · 11 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the WildabstractHow do product teams evaluate LLM-powered products? As organizations integrate large language models (LLMs) into digital products, their unpredictable nature makes traditional evaluation approaches inadequate, yet little is known about how practitioners navigate this challenge. Through interviews with nineteen practitioners across diverse sectors, we identify ten evaluation practices spanning informal ‘vibe checks’ to organizational meta-work. Beyond confirming four documented challenges, we introduce a novel fifth we call the results-actionability gap, in which practitioners gather evaluation data but cannot translate findings into concrete improvements. Drawing on patterns from successful teams, we contribute strategies to bridge this gap, supporting practitioners’ formalization journey from ad-hoc interpretive practices (e.g., vibe checks) toward systematic evaluation. Our analysis suggests these interpretive practices are necessary adaptations to LLM characteristics rather than methodological failures. For HCI researchers, this presents a research opportunity to support practitioners in systematizing emerging practices rather than developing new evaluation frameworks. Willem van der Maden, Malak Sadek, Ziang Xiao, Aske Mottelson, Qingzi Vera Liao, Jichen Zhu |
CHI | 5 |
| 2026 | From Use to Oversight: How Mental Models Influence User Behavior and Output in AI Writing AssistantsabstractAI-based writing assistants are ubiquitous, yet little is known about how users’ mental models shape their use. We examine two types of mental models—functional or related to what the system does, and structural or related to how the system works—and how they affect control behavior—how users request, accept, or edit AI suggestions as they write—and writing outcomes. We primed participants (N = 48) with different system descriptions to induce these mental models before asking them to complete a cover letter writing task using a writing assistant that occasionally offered preconfigured ungrammatical suggestions to test whether the mental models affected participants’ critical oversight. We find that while participants in the structural mental model condition demonstrate a better understanding of the system, this can have a backfiring effect: while these participants judged the system as more usable, they also produced letters with more grammatical errors, highlighting a complex relationship between system understanding, trust, and control in contexts that require user oversight of error-prone AI outputs. Shalaleh Rismani, Su Lin Blodgett, Qingzi Vera Liao, Alexandra Olteanu, AJung Moon |
CHI | 3 |
| 2025 | Canvil: Designerly Adaptation for LLM-Powered User Experiences
K. J. Kevin Feng, Qingzi Vera Liao, Ziang Xiao, Jennifer Wortman Vaughan, Amy X. Zhang, David W. McDonald |
CHI | 2 |
| 2025 | Fostering Appropriate Reliance on Large Language Models: The Role of Explanations, Sources, and InconsistenciesabstractLarge language models (LLMs) can produce erroneous responses that sound fluent and convincing, raising the risk that users will rely on these responses as if they were correct. Mitigating such overreliance is a key challenge. Through a think-aloud study in which participants use an LLM-infused application to answer objective questions, we identify several features of LLM responses that shape users' reliance: explanations (supporting details for answers), inconsistencies in explanations, and sources. Through a large-scale, pre-registered, controlled experiment (N=308), we isolate and study the effects of these features on users' reliance, accuracy, and other measures. We find that the presence of explanations increases reliance on both correct and incorrect responses. However, we observe less reliance on incorrect responses when sources are provided or when explanations exhibit inconsistencies. We discuss the implications of these findings for fostering appropriate reliance on LLMs. Sunnie S. Y. Kim, Jennifer Wortman Vaughan, Qingzi Vera Liao, Tania Lombrozo, Olga Russakovsky |
CHI | 3 |
| 2025 | As Confidence Aligns: Understanding the Effect of AI Confidence on Human Self-confidence in Human-AI Decision Making
Yitian Yang, Qingzi Vera Liao, Junti Zhang, Yi-Chieh Lee |
CHI | 3 |
| 2025 | 'It was 80% me, 20% AI': Seeking Authenticity in Co-Writing with Large Language ModelsabstractGiven the rising proliferation and diversity of AI writing assistance tools, especially those powered by large language models (LLMs), both writers and readers may have concerns about the impact of these tools on the authenticity of writing work. We examine whether and how writers want to preserve their authentic voice when co-writing with AI tools and whether personalization of AI writing support could help achieve this goal. We conducted semi-structured interviews with 19 professional writers, during which they co-wrote with both personalized and non-personalized AI writing-support tools. We supplemented writers' perspectives with opinions from 30 avid readers about the written work co-produced with AI collected through an online survey. Our findings illuminate conceptions of authenticity in human-AI co-creation, which focus more on the process and experience of constructing creators' authentic selves. While writers reacted positively to personalized AI writing tools, they believed the form of personalization needs to target writers' growth and go beyond the phase of text production. Overall, readers' responses showed less concern about human-AI co-writing. Readers could not distinguish AI-assisted work, personalized or not, from writers' solo-written work and showed positive attitudes toward writers experimenting with new technology for creative writing. Angel Hwang, Qingzi Vera Liao, Su Lin Blodgett, Alexandra Olteanu, Adam Trischler |
Proc. ACM Hum. Comput. Interact. | 2 |
| 2025 | Generation Probabilities Are Not Enough: Uncertainty Highlighting in AI Code CompletionsabstractLarge-scale generative models have enabled the development of AI-powered code completion tools to assist programmers in writing code. Like all AI-powered tools, these code completion tools are not always accurate and can introduce bugs or even security vulnerabilities into code if not properly detected and corrected by a human programmer. One technique that has been proposed and implemented to help programmers locate potential errors is to highlight uncertain tokens. However, little is known about the effectiveness of this technique. Through a mixed-methods study with 30 programmers, we compare three conditions: providing the AI system's code completion alone, highlighting tokens with the lowest likelihood of being generated by the underlying generative model, and highlighting tokens with the highest predicted likelihood of being edited by a programmer. We find that highlighting tokens with the highest predicted likelihood of being edited leads to faster task completion and more targeted edits, and is subjectively preferred by study participants. In contrast, highlighting tokens according to their probability of being generated does not provide any benefit over the baseline with no highlighting. We further explore the design space of how to convey uncertainty in AI-powered code completion tools and find that programmers prefer highlights that are granular, informative, interpretable, and not overwhelming. This work contributes to building an understanding of what uncertainty means for generative models and how to convey it effectively. Helena Vasconcelos, Gagan Bansal, Adam Fourney, Qingzi Vera Liao, Jennifer Wortman Vaughan |
ACM Trans. Comput. Hum. Interact. | 4 |
| 2025 | Why is AI Not a Panacea for Data Workers? An Interview Study on Human-AI Collaboration in Data StorytellingabstractThis paper explores the potential for human-AI collaboration in the context of data storytelling for data workers. Data storytelling communicates insights and knowledge from data analysis. It plays a vital role in data workers' daily jobs since it boosts team collaboration and public communication. However, to make an appealing data story, data workers need to spend tremendous effort on various tasks, including outlining and styling the story. Recently, a growing research trend has been exploring how to assist data storytelling with advanced artificial intelligence (AI). However, existing studies focus more on individual tasks in the workflow of data storytelling and do not reveal a complete picture of humans' preference for collaborating with AI. To address this gap, we conducted an interview study with 18 data workers to explore their preferences for AI collaboration in the planning, implementation, and communication stages of their workflow. We propose a framework for expected AI collaborators' roles, categorize people's expectations for the level of automation for different tasks, and delve into the reasons behind them. Our research provides insights and suggestions for the design of future AI-powered data storytelling tools. Haotian Li 0001, Yun Wang 0012, Qingzi Vera Liao, Huamin Qu |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2024 | ECBD: Evidence-Centered Benchmark Design for NLPabstractYu Lu Liu, Su Lin Blodgett, Jackie Cheung, Q. Vera Liao, Alexandra Olteanu, Ziang Xiao. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Yu Lu Liu, Su Lin Blodgett, Jackie Chi Kit Cheung, Qingzi Vera Liao, Alexandra Olteanu, Ziang Xiao |
ACL (1) | 4 |
| 2024 | The Who in XAI: How AI Background Shapes Perceptions of AI ExplanationsabstractExplainability of AI systems is critical for users to take informed actions. Understanding who opens the black-box of AI is just as important as opening it. We conduct a mixed-methods study of how two different groups—people with and without AI background—perceive different types of AI explanations. Quantitatively, we share user perceptions along five dimensions. Qualitatively, we describe how AI background can influence interpretations, elucidating the differences through lenses of appropriation and cognitive heuristics. We find that (1) both groups showed unwarranted faith in numbers for different reasons and (2) each group found value in different explanations beyond their intended design. Carrying critical implications for the field of XAI, our findings showcase how AI generated explanations can have negative consequences despite best intentions and how that could lead to harmful manipulation of trust. We propose design interventions to mitigate them. Upol Ehsan, Samir Passi, Qingzi Vera Liao, Larry Chan, I-Hsiang Lee, Michael J. Muller, Mark O. Riedl |
CHI | 3 |
| 2024 | Generative Echo Chamber? Effect of LLM-Powered Search Systems on Diverse Information SeekingabstractLarge language models (LLMs) powered conversational search systems have already been used by hundreds of millions of people, and are believed to bring many benefits over conventional search. However, while decades of research and public discourse interrogated the risk of search systems in increasing selective exposure and creating echo chambers—limiting exposure to diverse opinions and leading to opinion polarization, little is known about such a risk of LLM-powered conversational search. We conduct two experiments to investigate: 1) whether and how LLM-powered conversational search increases selective exposure compared to conventional search; 2) whether and how LLMs with opinion biases that either reinforce or challenge the user’s view change the effect. Overall, we found that participants engaged in more biased information querying with LLM-powered conversational search, and an opinionated LLM reinforcing their views exacerbated this bias. These results present critical implications for the development of LLMs and conversational search systems, and the policy governing these technologies. Qingzi Vera Liao, Ziang Xiao |
CHI | 2 |
| 2024 | Seamful XAI: Operationalizing Seamful Design in Explainable AIabstractMistakes in AI systems are inevitable, arising from both technical limitations and sociotechnical gaps. While black-boxing AI systems can make the user experience seamless, hiding the seams risks disempowering users to mitigate fallouts from AI mistakes. Instead of hiding these AI imperfections, can we leverage them to help the user? While Explainable AI (XAI) has predominantly tackled algorithmic opaqueness, we propose that seamful design can foster AI explainability by revealing and leveraging sociotechnical and infrastructural mismatches. We introduce the concept of Seamful XAI by (1) conceptually transferring "seams" to the AI context and (2) developing a design process that helps stakeholders anticipate and design with seams. We explore this process with 43 AI practitioners and real end-users, using a scenario-based co-design activity informed by real-world use cases. We found that the Seamful XAI design process helped users foresee AI harms, identify underlying reasons (seams), locate them in the AI's lifecycle, learn how to leverage seamful information to improve XAI and user agency. We share empirical insights, implications, and reflections on how this process can help practitioners anticipate and craft seams in AI, how seamfulness can improve explainability, empower end-users, and facilitate Responsible AI. Upol Ehsan, Qingzi Vera Liao, Samir Passi, Mark O. Riedl, Hal Daumé III |
Proc. ACM Hum. Comput. Interact. | 2 |
| 2023 | Designerly Understanding: Information Needs for Model Transparency to Support Design Ideation for AI-Powered User ExperienceabstractDespite the widespread use of artificial intelligence (AI), designing user experiences (UX) for AI-powered systems remains challenging. UX designers face hurdles understanding AI technologies, such as pre-trained language models, as design materials. This limits their ability to ideate and make decisions about whether, where, and how to use AI. To address this problem, we bridge the literature on AI design and AI transparency to explore whether and how frameworks for transparent model reporting can support design ideation with pre-trained models. By interviewing 23 UX practitioners, we find that practitioners frequently work with pre-trained models, but lack support for UX-led ideation. Through a scenario-based design task, we identify common goals that designers seek model understanding for and pinpoint their model transparency information needs. Our study highlights the pivotal role that UX designers can play in Responsible AI and calls for supporting their understanding of AI limitations through model transparency and interrogation. Qingzi Vera Liao, Hariharan Subramonyam, Jennifer Wortman Vaughan |
CHI | 1 |
| 2023 | fAIlureNotes: Supporting Designers in Understanding the Limits of AI Models for Computer Vision TasksabstractTo design with AI models, user experience (UX) designers must assess the fit between the model and user needs. Based on user research, they need to contextualize the model’s behavior and potential failures within their product-specific data instances and user scenarios. However, our formative interviews with ten UX professionals revealed that such a proactive discovery of model limitations is challenging and time-intensive. Furthermore, designers often lack technical knowledge of AI and accessible exploration tools, which challenges their understanding of model capabilities and limitations. In this work, we introduced a failure-driven design approach to AI, a workflow that encourages designers to explore model behavior and failure patterns early in the design process. The implementation of fAIlureNotes, a designer-centered failure exploration and analysis tool, supports designers in evaluating models and identifying failures across diverse user groups and scenarios. Our evaluation with UX practitioners shows that fAIlureNotes outperforms today’s interactive model cards in assessing context-specific model performance. Steven Moore, Qingzi Vera Liao, Hariharan Subramonyam |
CHI | 2 |
| 2023 | Evaluating Evaluation Metrics: A Framework for Analyzing NLG Evaluation Metrics using Measurement TheoryabstractWe address a fundamental challenge in Natural Language Generation (NLG) model evaluation-the design and evaluation of evaluation metrics.Recognizing the limitations of existing automatic metrics and noises from how current human evaluation was conducted, we propose METRICEVAL, a framework informed by measurement theory, the foundation of educational test design, for conceptualizing and evaluating the reliability and validity of NLG evaluation metrics.The framework formalizes the source of measurement error and offers statistical tools for evaluating evaluation metrics based on empirical data.With our framework, one can quantify the uncertainty of the metrics to better interpret the result.To exemplify the use of our framework in practice, we analyzed a set of evaluation metrics for summarization and identified issues related to conflated validity structure in human-eval and reliability in LLM-based metrics.Through METRICEVAL 1 , we aim to promote the design, evaluation, and interpretation of valid and reliable metrics to advance robust and effective NLG models. Ziang Xiao, Susu Zhang, Vivian Lai, Qingzi Vera Liao |
EMNLP | 4 |
| 2023 | Understanding Uncertainty: How Lay Decision-makers Perceive and Interpret Uncertainty in Human-AI Decision MakingabstractDecision Support Systems (DSS) based on Machine Learning (ML) often aim to assist lay decision-makers, who are not math-savvy, in making high-stakes decisions. However, existing ML-based DSS are not always transparent about the probabilistic nature of ML predictions and how uncertain each prediction is. This lack of transparency could give lay decision-makers a false sense of reliability. Growing calls for AI transparency have led to increasing efforts to quantify and communicate model uncertainty. However, there are still gaps in knowledge regarding how and why the decision-makers utilize ML uncertainty information in their decision process. Here, we conducted a qualitative, think-aloud user study with 17 lay decision-makers who interacted with three different DSS: 1) interactive visualization, 2) DSS based on an ML model that provides predictions without uncertainty information, and 3) the same DSS with uncertainty information. Our qualitative analysis found that communicating uncertainty about ML predictions forced participants to slow down and think analytically about their decisions. This in turn made participants more vigilant, resulting in reduction in over-reliance on ML-based DSS. Our work contributes empirical knowledge on how lay decision-makers perceive, interpret, and make use of uncertainty information when interacting with DSS. Such foundational knowledge informs the design of future ML-based DSS that embrace transparent uncertainty communication. Snehal Prabhudesai, Leyao Yang, Sumit Asthana, Xun Huan, Qingzi Vera Liao, Nikola Banovic 0001 |
IUI | 5 |
| 2023 | Powering an AI Chatbot with Expert Sourcing to Support Credible Health Information AccessabstractDuring a public health crisis like the COVID-19 pandemic, a credible and easy-to-access information portal is highly desirable. It helps with disease prevention, public health planning, and misinformation mitigation. However, creating such an information portal is challenging because 1) domain expertise is required to identify and curate credible and intelligible content, 2) the information needs to be updated promptly in response to the fast-changing environment, and 3) the information should be easily accessible by the general public; which is particularly difficult when most people do not have the domain expertise about the crisis. In this paper, we presented an expert-sourcing framework and created Jennifer, an AI chatbot, which serves as a credible and easy-to-access information portal for individuals during the COVID-19 pandemic. Jennifer was created by a team of over 150 scientists and health professionals around the world, deployed in the real world and answered thousands of user questions about COVID-19. We evaluated Jennifer from two key stakeholders’ perspectives, expert volunteers and information seekers. We first interviewed experts who contributed to the collaborative creation of Jennifer to learn about the challenges in the process and opportunities for future improvement. We then conducted an online experiment that examined Jennifer’s effectiveness in supporting information seekers in locating COVID-19 information and gaining their trust. We share the key lessons learned and discuss design implications for building expert-sourced and AI-powered information portals, along with the risks and opportunities of misinformation mitigation and beyond. Ziang Xiao, Qingzi Vera Liao, Michelle X. Zhou, Tyrone Grandison, Yunyao Li 0001 |
IUI | 2 |
| 2023 | Understanding the Role of Human Intuition on Reliance in Human-AI Decision-Making with ExplanationsabstractAI explanations are often mentioned as a way to improve human-AI decision-making, but empirical studies have not found consistent evidence of explanations' effectiveness and, on the contrary, suggest that they can increase overreliance when the AI system is wrong. While many factors may affect reliance on AI support, one important factor is how decision-makers reconcile their own intuition---beliefs or heuristics, based on prior knowledge, experience, or pattern recognition, used to make judgments---with the information provided by the AI system to determine when to override AI predictions. We conduct a think-aloud, mixed-methods study with two explanation types (feature- and example-based) for two prediction tasks to explore how decision-makers' intuition affects their use of AI predictions and explanations, and ultimately their choice of when to rely on AI. Our results identify three types of intuition involved in reasoning about AI predictions and explanations: intuition about the task outcome, features, and AI limitations. Building on these, we summarize three observed pathways for decision-makers to apply their own intuition and override AI predictions. We use these pathways to explain why (1) the feature-based explanations we used did not improve participants' decision outcomes and increased their overreliance on AI, and (2) the example-based explanations we used improved decision-makers' performance over feature-based explanations and helped achieve complementary human-AI performance. Overall, our work identifies directions for further development of AI decision-support systems and explanation methods that help decision-makers effectively apply their intuition to achieve appropriate reliance on AI. Valerie Chen, Qingzi Vera Liao, Jennifer Wortman Vaughan, Gagan Bansal |
Proc. ACM Hum. Comput. Interact. | 2 |
| 2023 | Sensing Wellbeing in the Workplace, Why and For Whom? Envisioning Impacts with Organizational StakeholdersabstractWith the heightened digitization of the workplace, alongside the rise of remote and hybrid work prompted by the pandemic, there is growing corporate interest in using passive sensing technologies for workplace wellbeing. Existing research on these technologies often focus on understanding or improving interactions between an individual user and the technology. Workplace settings can, however, introduce a range of complexities that challenge the potential impact and in-practice desirability of wellbeing sensing technologies. Today, there is an inadequate empirical understanding of how everyday workers---including those who are impacted by, and impact the deployment of workplace technologies--envision its broader socio-ecological impacts. In this study, we conduct storyboard-driven interviews with 33 participants across three stakeholder groups: organizational governors, AI builders, and worker data subjects. Overall, our findings surface how workers envisioned wellbeing sensing technologies may lead to cascading impacts on their broader organizational culture, interpersonal relationships with colleagues, and individual day-to-day lives. Participants anticipated harms arising from ambiguity and misalignment around scaled notions of "worker wellbeing,'' underlying technical limitations to workplace-situated sensing, and assumptions regarding how social structures and relationships may shape the impacts and use of these technologies. Based on our findings, we discuss implications for designing worker-centered data-driven wellbeing technologies. Anna Kawakami, Shreya Chowdhary, Shamsi T. Iqbal, Qingzi Vera Liao, Alexandra Olteanu, Jina Suh, Koustuv Saha |
Proc. ACM Hum. Comput. Interact. | 4 |
| 2023 | Selective Explanations: Leveraging Human Input to Align Explainable AIabstractWhile a vast collection of explainable AI (XAI) algorithms has been developed in recent years, they have been criticized for significant gaps with how humans produce and consume explanations. As a result, current XAI techniques are often found to be hard to use and lack effectiveness. In this work, we attempt to close these gaps by making AI explanations selective ---a fundamental property of human explanations---by selectively presenting a subset of model reasoning based on what aligns with the recipient's preferences. We propose a general framework for generating selective explanations by leveraging human input on a small dataset. This framework opens up a rich design space that accounts for different selectivity goals, types of input, and more. As a showcase, we use a decision-support task to explore selective explanations based on what the decision-maker would consider relevant to the decision task. We conducted two experimental studies to examine three paradigms based on our proposed framework: in Study 1, we ask the participants to provide critique-based or open-ended input to generate selective explanations (self-input). In Study 2, we show the participants selective explanations based on input from a panel of similar users (annotator input). Our experiments demonstrate the promise of selective explanations in reducing over-reliance on AI and improving collaborative decision making and subjective perceptions of the AI system, but also paint a nuanced picture that attributes some of these positive effects to the opportunity to provide one's own input to augment AI explanations. Overall, our work proposes a novel XAI framework inspired by human communication behaviors and demonstrates its potential to encourage future work to make AI explanations more human-compatible. Vivian Lai, Yiming Zhang 0022, Chacha Chen, Qingzi Vera Liao, Chenhao Tan |
Proc. ACM Hum. Comput. Interact. | 4 |
| 2022 | AI Explainability 360: Impact and DesignabstractAs artificial intelligence and machine learning algorithms become increasingly prevalent in society, multiple stakeholders are calling for these algorithms to provide explanations. At the same time, these stakeholders, whether they be affected citizens, government regulators, domain experts, or system developers, have different explanation needs. To address these needs, in 2019, we created AI Explainability 360, an open source software toolkit featuring ten diverse and state-of-the-art explainability methods and two evaluation metrics. This paper examines the impact of the toolkit with several case studies, statistics, and community feedback. The different ways in which users have experienced AI Explainability 360 have resulted in multiple types of impact and improvements in multiple metrics, highlighted by the adoption of the toolkit by the independent LF AI & Data Foundation. The paper also describes the flexible design of the toolkit, examples of its use, and the significant educational material and documentation available to its users. Vijay Arya, Rachel K. E. Bellamy, Amit Dhurandhar, Michael Hind, Samuel C. Hoffman, Stephanie Houde, Qingzi Vera Liao, Ronny Luss, Aleksandra Mojsilovic, Sami Mourad, Pablo Pedemonte, Ramya Raghavendra, John T. Richards, Prasanna Sattigeri, Karthikeyan Shanmugam 0001, Moninder Singh, Kush R. Varshney, Dennis Wei |
AAAI | 8 |
| 2022 | Human-AI Collaboration via Conditional Delegation: A Case Study of Content ModerationabstractDespite impressive performance in many benchmark datasets, AI models can still make mistakes, especially among out-of-distribution examples. It remains an open question how such imperfect models can be used effectively in collaboration with humans. Prior work has focused on AI assistance that helps people make individual high-stakes decisions, which is not scalable for a large amount of relatively low-stakes decisions, e.g., moderating social media comments. Instead, we propose conditional delegation as an alternative paradigm for human-AI collaboration where humans create rules to indicate trustworthy regions of a model. Using content moderation as a testbed, we develop novel interfaces to assist humans in creating conditional delegation rules and conduct a randomized experiment with two datasets to simulate in-distribution and out-of-distribution scenarios. Our study demonstrates the promise of conditional delegation in improving model performance and provides insights into design for this novel paradigm, including the effect of AI explanations. Vivian Lai, Samuel Carton, Rajat Bhatnagar, Qingzi Vera Liao, Chenhao Tan |
CHI | 4 |
| 2022 | Connecting Algorithmic Research and Usage Contexts: A Perspective of Contextualized Evaluation for Explainable AIabstractRecent years have seen a surge of interest in the field of explainable AI (XAI), with a plethora of algorithms proposed in the literature. However, a lack of consensus on how to evaluate XAI hinders the advancement of the field. We highlight that XAI is not a monolithic set of technologies---researchers and practitioners have begun to leverage XAI algorithms to build XAI systems that serve different usage contexts, such as model debugging and decision-support. Algorithmic research of XAI, however, often does not account for these diverse downstream usage contexts, resulting in limited effectiveness or even unintended consequences for actual users, as well as difficulties for practitioners to make technical choices. We argue that one way to close the gap is to develop evaluation methods that account for different user requirements in these usage contexts. Towards this goal, we introduce a perspective of contextualized XAI evaluation by considering the relative importance of XAI evaluation criteria for prototypical usage contexts of XAI. To explore the context dependency of XAI evaluation criteria, we conduct two survey studies, one with XAI topical experts and another with crowd workers. Our results urge for responsible AI research with usage-informed evaluation practices, and provide a nuanced understanding of user requirements for XAI in different usage contexts. Qingzi Vera Liao, Ronny Luss, Finale Doshi-Velez, Amit Dhurandhar |
HCOMP | 1 |
| 2022 | Investigating Explainability of Generative AI for Code through Scenario-based DesignabstractWhat does it mean for a generative AI model to be explainable? The emergent discipline of explainable AI (XAI) has made great strides in helping people understand discriminative models. Less attention has been paid to generative models that produce artifacts, rather than decisions, as output. Meanwhile, generative AI (GenAI) technologies are maturing and being applied to application domains such as software engineering. Using scenario-based design and question-driven XAI design approaches, we explore users’ explainability needs for GenAI in three software engineering use cases: natural language to code, code translation, and code auto-completion. We conducted 9 workshops with 43 software engineers in which real examples from state-of-the-art generative AI models were used to elicit users’ explainability needs. Drawing from prior work, we also propose 4 types of XAI features for GenAI for code and gathered additional design ideas from participants. Our work explores explainability needs for GenAI for code and demonstrates how human-centered approaches can drive the technical development of XAI in novel domains. Jiao Sun, Qingzi Vera Liao, Michael J. Muller, Mayank Agarwal, Stephanie Houde, Kartik Talamadupula, Justin D. Weisz |
IUI | 2 |
| 2022 | Human-AI Collaboration for UX Evaluation: Effects of Explanation and SynchronizationabstractAnalyzing usability test videos is arduous. Although recent research showed the promise of AI in assisting with such tasks, it remains largely unknown how AI should be designed to facilitate effective collaboration between user experience (UX) evaluators and AI. Inspired by the concepts of agency and work context in human and AI collaboration literature, we studied two corresponding design factors for AI-assisted UX evaluation: explanations and synchronization. Explanations allow AI to further inform humans how it identifies UX problems from a usability test session; synchronization refers to the two ways humans and AI collaborate: synchronously and asynchronously. We iteratively designed a tool-AI Assistant-with four versions of UIs corresponding to the two levels of explanations (with/without) and synchronization (sync/async). By adopting a hybrid wizard-of-oz approach to simulating an AI with reasonable performance, we conducted a mixed-method study with 24 UX evaluators identifying UX problems from usability test videos using AI Assistant. Our quantitative and qualitative results show that AI with explanations, regardless of being presented synchronously or asynchronously, provided better support for UX evaluators' analysis and was perceived more positively; when without explanations, synchronous AI better improved UX evaluators' performance and engagement compared to the asynchronous AI. Lastly, we present the design implications for AI-assisted UX evaluation and facilitating more effective human-AI collaboration. Mingming Fan 0001, Xianyou Yang, Tsz Tung Yu, Qingzi Vera Liao, Jian Zhao 0010 |
Proc. ACM Hum. Comput. Interact. | 4 |
| 2022 | Special Issue on Conversational Agents for Healthcare and WellbeingabstractConversational agents (CAs) are systems that interact with humans through natural language user interfaces. They include systems with a range of conversational capabilities and modalities. For example, there are text- or voice-only question-answering interactions such as Apple Siri, Google Assistant, and Amazon Alexa, and there are also multimodal conversational AI agents that can engage users in long-term dialogues. Advances in speech recognition, natural language processing, and computer vision have resulted in a greater acceptance and use of CAs. CAs have already started to play important roles in various healthcare settings, including assisting clinicians during consultations, assisting consumers in changing health behaviours, and helping patients such as the elderly in their living environments. \n \nThere have been several systematic reviews on the use of CAs in health and wellbeing recently. Although the field still appears to be nascent, the emerging evidence has shown user acceptance of CAs in the healthcare domain as well as the early promises in boosting healthcare outcomes in both physical and mental health. Despite the increasing adoption and the benefits of using CAs to support health and wellbeing, the review studies also revealed (i) patient safety was rarely examined, (ii) health outcomes were inadequately measured, and (iii) no standardised evaluation methods were employed. There were also limitations in reporting the technical implementation details of CAs used, making the replicability of prior studies problematic. \n \nIn addition to addressing some of the current challenges and limitations, this special issue features cutting-edge research on designing, developing, and evaluating CAs for health and wellbeing that aim to improve health outcomes and services, and satisfy unique application needs (e.g., safety, trust, and user experience). The seven articles included in this special issue cover many application areas ranging from mental health and social support to information seeking to coaching. Amongst the accepted articles, mental health and social support themes represented the primary research foci. The articles also covered different population groups including older adults, young adults, and homeless people. Based on their foci, we have grouped the articles in this special issue by three areas: mental health, older adult wellbeing, and social support and coaching. Ahmet Baki Kocaballi, Liliana Laranjo, Leigh Clark, Rafal Kocielnik, Robert J. Moore, Qingzi Vera Liao, Timothy W. Bickmore |
ACM Trans. Interact. Intell. Syst. | 6 |
| 2021 | Doc2Bot: Document grounded Bot FrameworkabstractConversational agents, or chatbots, are widely used to provide customer care and other informational support. Currently, the development of chatbots using standard frameworks requires a lot of manual crafting by subject matter experts (SMEs). On the other hand, while learning-based approaches to dialog have made significant advancements, they require training with a large volume of dialog data, which chatbot developers typically do not have access to. To tackle these challenges, we introduce DOC2BOT, a system that supports the automated construction of chatbots by digesting various forms of documents such as business manuals, HowTos, and customer support pages that organizations own. In addition to that, DOC2BOT provides a user-friendly experience to SMEs, and to minimize their effort by supporting intuitive interactions and streamlining their workflow. Kshitij Fadnis, Pankaj Dhoolia, Qingzi Vera Liao, Steven Ross, Nathaniel Mills, Sachindra Joshi, Luis A. Lastras |
AAAI | 4 |
| 2021 | Uncertainty as a Form of Transparency: Measuring, Communicating, and Using UncertaintyabstractAlgorithmic transparency entails exposing system properties to various stakeholders for purposes that include understanding, improving, and contesting predictions. Until now, most research into algorithmic transparency has predominantly focused on explainability. Explainability attempts to provide reasons for a machine learning model's behavior to stakeholders. However, understanding a model's specific behavior alone might not be enough for stakeholders to gauge whether the model is wrong or lacks sufficient knowledge to solve the task at hand. In this paper, we argue for considering a complementary form of transparency by estimating and communicating the uncertainty associated with model predictions. First, we discuss methods for assessing uncertainty. Then, we characterize how uncertainty can be used to mitigate model unfairness, augment decision-making, and build trustworthy systems. Finally, we outline methods for displaying uncertainty to stakeholders and recommend how to collect information required for incorporating uncertainty into existing ML pipelines. This work constitutes an interdisciplinary review drawn from literature spanning machine learning, visualization/HCI, design, decision-making, and fairness. We aim to encourage researchers and practitioners to measure, communicate, and use uncertainty as a form of transparency. Umang Bhatt, Javier Antorán, Qingzi Vera Liao, Prasanna Sattigeri, Riccardo Fogliato, Gabrielle Gauthier Melançon, Ranganath Krishnan, Jason Stanley, Omesh Tickoo, Lama Nachman, Rumi Chunara, Madhulika Srikumar, Adrian Weller, Alice Xiang |
AIES | 4 |
| 2021 | Expanding Explainability: Towards Social Transparency in AI systemsabstractAs AI-powered systems increasingly mediate consequential decision-making, their explainability is critical for end-users to take informed and accountable actions. Explanations in human-human interactions are socially-situated. AI systems are often socio-organizationally embedded. However, Explainable AI (XAI) approaches have been predominantly algorithm-centered. We take a developmental step towards socially-situated XAI by introducing and exploring Social Transparency (ST), a sociotechnically informed perspective that incorporates the socio-organizational context into explaining AI-mediated decision-making. To explore ST conceptually, we conducted interviews with 29 AI users and practitioners grounded in a speculative design scenario. We suggested constitutive design elements of ST and developed a conceptual framework to unpack ST’s effect and implications at the technical, decision-making, and organizational level. The framework showcases how ST can potentially calibrate trust in AI, improve decision-making, facilitate organizational collective actions, and cultivate holistic explainability. Our work contributes to the discourse of Human-Centered XAI by expanding the design space of XAI. Upol Ehsan, Qingzi Vera Liao, Michael J. Muller, Mark O. Riedl, Justin D. Weisz |
CHI | 2 |
| 2021 | Model LineUpper: Supporting Interactive Model Comparison at Multiple Levels for AutoMLabstractAutomated Machine Learning (AutoML) is a rapidly growing set of technologies that automate the model development pipeline by searching model space and generating candidate models. A critical, final step of AutoML is human selection of a final model from dozens of candidates. In current AutoML systems, selection is supported only by performance metrics. Prior work has shown that in practice, people evaluate ML models based on additional criteria, such as the way a model makes predictions. Comparison may happen at multiple levels, from types of errors, to feature importance, to how the model makes predictions of specific instances. We developed Model LineUpper to support interactive model comparison for AutoML by integrating multiple Explainable AI (XAI) and visualization techniques. We conducted a user study in which we both evaluated the system and used it as a technology probe to understand how users perform model comparison in an AutoML system. We discuss design implications for utilizing XAI techniques for model comparison and supporting the unique needs of data scientists in comparing AutoML models. Shweta Narkar, Qingzi Vera Liao, Dakuo Wang, Justin D. Weisz |
IUI | 3 |
| 2021 | Facilitating Knowledge Sharing from Domain Experts to Data Scientists for Building NLP ModelsabstractData scientists face a steep learning curve in understanding a new domain for which they want to build machine learning (ML) models. While input from domain experts could offer valuable help, such input is often limited, expensive, and generally not in a form readily consumable by a model development pipeline. In this paper, we propose Ziva, a framework to guide domain experts in sharing essential domain knowledge to data scientists for building NLP models. With Ziva, experts are able to distill and share their domain knowledge using domain concept extractors and five types of label justification over a representative data sample. The design of Ziva is informed by preliminary interviews with data scientists, in order to understand current practices of domain knowledge acquisition process for ML development projects. To assess our design, we run a mix-method case-study to evaluate how Ziva can facilitate interaction between domain experts and data scientists. Our results highlight that (1) domain experts are able to use Ziva to provide rich domain knowledge, while maintaining low mental load and stress levels; and (2) data scientists find Ziva’s output helpful for learning essential information about the domain, offering scalability of information, and lowering the burden on domain experts to share knowledge. We conclude this work by experimenting with building NLP models using the Ziva output for our case study. Soya Park, April Yi Wang, Ban Kawas, Qingzi Vera Liao, David Piorkowski, Marina Danilevsky |
IUI | 4 |
| 2021 | Why or Why Not? The Effect of Justification Styles on Chatbot RecommendationsabstractChatbots or conversational recommenders have gained increasing popularity as a new paradigm for Recommender Systems (RS). Prior work on RS showed that providing explanations can improve transparency and trust, which are critical for the adoption of RS. Their interactive and engaging nature makes conversational recommenders a natural platform to not only provide recommendations but also justify the recommendations through explanations. The recent surge of interest inexplainable AI enables diverse styles of justification, and also invites questions on how styles of justification impact user perception. In this article, we explore the effect of “why” justifications and “why not” justifications on users’ perceptions of explainability and trust. We developed and tested a movie-recommendation chatbot that provides users with different types of justifications for the recommended items. Our online experiment ( n = 310) demonstrates that the “why” justifications (but not the “why not” justifications) have a significant impact on users’ perception of the conversational recommender. Particularly, “why” justifications increase users’ perception of system transparency, which impacts perceived control, trusting beliefs and in turn influences users’ willingness to depend on the system’s advice. Finally, we discuss the design implications for decision-assisting chatbots. Daricia Wilkinson, Oznur Alkan, Qingzi Vera Liao, Massimiliano Mattetti, Inge Vejsbjerg, Bart P. Knijnenburg, Elizabeth Daly |
ACM Trans. Inf. Syst. | 3 |
| 2020 | Doc2Dial: A Framework for Dialogue Composition Grounded in DocumentsabstractWe introduce Doc2Dial, an end-to-end framework for generating conversational data grounded in given documents. It takes the documents as input and generates the pipelined tasks for obtaining the annotations specifically for producing the simulated dialog flows. Then, the dialog flows are used to guide the collection of the utterances via the integrated crowdsourcing tool. The outcomes include the human-human dialogue data grounded in the given documents, as well as various types of automatically or human labeled annotations that help ensure the quality of the dialog data with the flexibility to (re)composite dialogues. We expect such data can facilitate building automated dialogue agents for goal-oriented tasks. We demonstrate Doc2Dial system with the various domain documents for customer care. Song Feng 0002, Kshitij Fadnis, Qingzi Vera Liao, Luis A. Lastras |
AAAI | 3 |
| 2020 | Questioning the AI: Informing Design Practices for Explainable AI User ExperiencesabstractA surge of interest in explainable AI (XAI) has led to a vast collection of algorithmic work on the topic. While many recognize the necessity to incorporate explainability features in AI systems, how to address real-world user needs for understanding AI remains an open question. By interviewing 20 UX and design practitioners working on various AI products, we seek to identify gaps between the current XAI algorithmic work and practices to create explainable AI products. To do so, we develop an algorithm-informed XAI question bank in which user needs for explainability are represented as prototypical questions users might ask about the AI, and use it as a study probe. Our work contributes insights into the design space of XAI, informs efforts to support design practices in this space, and identifies opportunities for future XAI work. We also provide an extended XAI question bank and discuss how it can be used for creating user-centered XAI. Qingzi Vera Liao, Dan Gruen |
CHI | 1 |
| 2020 | AI Explainability 360: An Extensible Toolkit for Understanding Data and Machine Learning ModelsabstractAs artificial intelligence algorithms make further inroads in high-stakes societal applications, there are increasing calls from multiple stakeholders for these algorithms to explain their outputs. To make matters more challenging, different personas of consumers of explanations have different requirements for explanations. Toward addressing these needs, we introduce AI Explainability 360, an open-source Python toolkit featuring ten diverse and state-of-the-art explainability methods and two evaluation metrics. Equally important, we provide a taxonomy to help entities requiring explanations to navigate the space of interpretation and explanation methods, not only those in the toolkit but also in the broader literature on explainability. For data scientists and other users of the toolkit, we have implemented an extensible software architecture that organizes methods according to their place in the AI modeling pipeline. The toolkit is not only the software, but also guidance material, tutorials, and an interactive web demo to introduce AI explainability to different audiences. Together, our toolkit and taxonomy can help identify gaps where more explainability methods are needed and provide a platform to incorporate them as they are developed. Vijay Arya, Rachel K. E. Bellamy, Amit Dhurandhar, Michael Hind, Samuel C. Hoffman, Stephanie Houde, Qingzi Vera Liao, Ronny Luss, Aleksandra Mojsilovic, Sami Mourad, Pablo Pedemonte, Ramya Raghavendra, John T. Richards, Prasanna Sattigeri, Karthikeyan Shanmugam 0001, Moninder Singh, Kush R. Varshney, Dennis Wei |
J. Mach. Learn. Res. | 8 |
| 2020 | Human-AI Collaboration in a Cooperative Game Setting: Measuring Social Perception and OutcomesabstractHuman-AI interaction is pervasive across many areas of our day to day lives. In this paper, we investigate human-AI collaboration in the context of a collaborative AI-driven word association game with partially observable information. In our experiments, we test various dimensions of subjective social perceptions (rapport, intelligence, creativity and likeability) of participants towards their partners when participants believe they are playing with an AI or with a human. We also test subjective social perceptions of participants towards their partners when participants are presented with a variety of confidence levels. We ran a large scale study on Mechanical Turk (n=164) of this collaborative game. Our results show that when participants believe their partners were human, they found their partners to be more likeable, intelligent, creative and having more rapport and use more positive words to describe their partner's attributes than when they believed they were interacting with an AI partner. We also found no differences in game outcome including win rate and turns to completion. Drawing on both quantitative and qualitative findings, we discuss AI agent transparency, include design implications for tools incorporating or supporting human-AI collaboration, and lay out directions for future research. Our findings lead to implications for other forms of human-AI interaction and communication. Zahra Ashktorab, Qingzi Vera Liao, Casey Dugan, Wei Zhang 0057, Sadhana Kumaravel, Murray Campbell |
Proc. ACM Hum. Comput. Interact. | 2 |
| 2020 | Explainable Active Learning (XAL): Toward AI Explanations as Interfaces for Machine TeachersabstractThe wide adoption of Machine Learning (ML) technologies has created a growing demand for people who can train ML models. Some advocated the term "machine teacher'' to refer to the role of people who inject domain knowledge into ML models. This "teaching'' perspective emphasizes supporting the productivity and mental wellbeing of machine teachers through efficient learning algorithms and thoughtful design of human-AI interfaces. One promising learning paradigm is Active Learning (AL), by which the model intelligently selects instances to query a machine teacher for labels, so that the labeling workload could be largely reduced. However, in current AL settings, the human-AI interface remains minimal and opaque. A dearth of empirical studies further hinders us from developing teacher-friendly interfaces for AL algorithms. In this work, we begin considering AI explanations as a core element of the human-AI interface for teaching machines. When a human student learns, it is a common pattern to present one's own reasoning and solicit feedback from the teacher. When a ML model learns and still makes mistakes, the teacher ought to be able to understand the reasoning underlying its mistakes. When the model matures, the teacher should be able to recognize its progress in order to trust and feel confident about their teaching outcome. Toward this vision, we propose a novel paradigm of explainable active learning (XAL), by introducing techniques from the surging field of explainable AI (XAI) into an AL setting. We conducted an empirical study comparing the model learning outcomes, feedback content and experience with XAL, to that of traditional AL and coactive learning (providing the model's prediction without explanation). Our study shows benefits of AI explanation as interfaces for machine teaching--supporting trust calibration and enabling rich forms of teaching feedback, and potential drawbacks--anchoring effect with the model judgment and additional cognitive workload. Our study also reveals important individual factors that mediate a machine teacher's reception to AI explanations, including task knowledge, AI experience and Need for Cognition. By reflecting on the results, we suggest future directions and design implications for XAL, and more broadly, machine teaching through AI explanations. Bhavya Ghai, Qingzi Vera Liao, Rachel K. E. Bellamy, Klaus Mueller 0001 |
Proc. ACM Hum. Comput. Interact. | 2 |
| 2020 | Tell Me About Yourself: Using an AI-Powered Chatbot to Conduct Conversational Surveys with Open-ended QuestionsabstractThe rise of increasingly more powerful chatbots offers a new way to collect information through conversational surveys, where a chatbot asks open-ended questions, interprets a user’s free-text responses, and probes answers whenever needed. To investigate the effectiveness and limitations of such a chatbot in conducting surveys, we conducted a field study involving about 600 participants. In this study with mostly open-ended questions, half of the participants took a typical online survey on Qualtrics and the other half interacted with an AI-powered chatbot to complete a conversational survey. Our detailed analysis of over 5,200 free-text responses revealed that the chatbot drove a significantly higher level of participant engagement and elicited significantly better quality responses measured by Gricean Maxims in terms of their informativeness, relevance, specificity, and clarity. Based on our results, we discuss design implications for creating AI-powered chatbots to conduct effective surveys and beyond. Ziang Xiao, Michelle X. Zhou, Qingzi Vera Liao, Gloria Mark, Chang Yan Chi, Huahai Yang |
ACM Trans. Comput. Hum. Interact. | 3 |
| 2019 | Bootstrapping Conversational Agents with Weak SupervisionabstractMany conversational agents in the market today follow a standard bot development framework which requires training intent classifiers to recognize user input. The need to create a proper set of training examples is often the bottleneck in the development process. In many occasions agent developers have access to historical chat logs that can provide a good quantity as well as coverage of training examples. However, the cost of labeling them with tens to hundreds of intents often prohibits taking full advantage of these chat logs. In this paper, we present a framework called search, label, and propagate (SLP) for bootstrapping intents from existing chat logs using weak supervision. The framework reduces hours to days of labeling effort down to minutes of work by using a search engine to find examples, then relies on a data programming approach to automatically expand the labels. We report on a user study that shows positive user feedback for this new approach to build conversational agents, and demonstrates the effectiveness of using data programming for autolabeling. While the system is developed for training conversational agents, the framework has broader application in significantly reducing labeling effort for training text classifiers. Neil Mallinar, Abhishek Shah, Rajendra Ugrani, Manikandan Gurusankar, Tin Kam Ho, Qingzi Vera Liao, Rachel K. E. Bellamy, Robert Yates, Chris Desmarais, Blake McGregor |
AAAI | 7 |
| 2019 | Resilient Chatbots: Repair Strategy Preferences for Conversational BreakdownsabstractText-based conversational systems, also referred to as chatbots, have grown widely popular. Current natural language understanding technologies are not yet ready to tackle the complexities in conversational interactions. Breakdowns are common, leading to negative user experiences. Guided by communication theories, we explore user preferences for eight repair strategies, including ones that are common in commercially-deployed chatbots (e.g., confirmation, providing options), as well as novel strategies that explain characteristics of the underlying machine learning algorithms. We conducted a scenario-based study to compare repair strategies with Mechanical Turk workers (N=203). We found that providing options and explanations were generally favored, as they manifest initiative from the chatbot and are actionable to recover from breakdowns. Through detailed analysis of participants' responses, we provide a nuanced understanding on the strengths and weaknesses of each repair strategy. Zahra Ashktorab, Qingzi Vera Liao, Justin D. Weisz |
CHI | 3 |
| 2019 | How Data Science Workers Work with Data: Discovery, Capture, Curation, Design, CreationabstractWith the rise of big data, there has been an increasing need for practitioners in this space and an increasing opportunity for researchers to understand their workflows and design new tools to improve it. Data science is often described as data-driven, comprising unambiguous data and proceeding through regularized steps of analysis. However, this view focuses more on abstract processes, pipelines, and workflows, and less on how data science workers engage with the data. In this paper, we build on the work of other CSCW and HCI researchers in describing the ways that scientists, scholars, engineers, and others work with their data, through analyses of interviews with 21 data science professionals. We set five approaches to data along a dimension of interventions: Data as given; as captured; as curated; as designed; and as created. Data science workers develop an intuitive sense of their data and processes, and actively shape their data. We propose new ways to apply these interventions analytically, to make sense of the complex activities around data practices. Michael J. Muller, Ingrid Lange, Dakuo Wang, David Piorkowski, Jason Tsay, Qingzi Vera Liao, Casey Dugan, Thomas Erickson |
CHI | 6 |
| 2019 | Explaining models: an empirical study of how explanations impact fairness judgmentabstractEnsuring fairness of machine learning systems is a human-in-the-loop process. It relies on developers, users, and the general public to identify fairness problems and make improvements. To facilitate the process we need effective, unbiased, and user-friendly explanations that people can confidently rely on. Towards that end, we conducted an empirical study with four types of programmatically generated explanations to understand how they impact people's fairness judgments of ML systems. With an experiment involving more than 160 Mechanical Turk workers, we show that: 1) Certain explanations are considered inherently less fair, while others can enhance people's confidence in the fairness of the algorithm; 2) Different fairness problems-such as model-wide fairness issues versus case-specific fairness discrepancies-may be more effectively exposed through different styles of explanation; 3) Individual differences, including prior positions and judgment criteria of algorithmic fairness, impact how people react to different styles of explanation. We conclude with a discussion on providing personalized and adaptive explanations to support fairness judgments of ML systems. Jonathan Dodge, Qingzi Vera Liao, Rachel K. E. Bellamy, Casey Dugan |
IUI | 2 |
| 2018 | All Work and No Play?abstractMany conversational agents (CAs) are developed to answer users' questions in a specialized domain. In everyday use of CAs, user experience may extend beyond satisfying information needs to the enjoyment of conversations with CAs, some of which represent playful interactions. By studying a field deployment of a Human Resource chatbot, we report on users' interest areas in conversational interactions to inform the development of CAs. Through the lens of statistical modeling, we also highlight rich signals in conversational interactions for inferring user satisfaction with the instrumental usage and playful interactions with the agent. These signals can be utilized to develop agents that adapt functionality and interaction styles. By contrasting these signals, we shed light on the varying functions of conversational interactions. We discuss design implications for CAs, and directions for developing adaptive agents based on users' conversational behaviors. Qingzi Vera Liao, Muhammed Mas-ud Hussain, Praveen Chandar, Yasaman Khazaeni, Marco Crasso, Dakuo Wang, Michael J. Muller, N. Sadat Shami, Werner Geyer |
CHI | 1 |
| 2018 | Face Value?abstractWe are interested in increasing the ability of groups to collaborate efficiently by leveraging new advances in AI and Conversational Agent (CA) technology. Given the longstanding debate on the necessity of embodiment for CAs, bringing them to groups requires answering the questions of whether and how providing a CA with a face affects its interaction with the humans in a group. We explored these questions by comparing group decision-making sessions facilitated by an embodied agent, versus a voice-only agent. Results of an experiment with 20 user groups revealed that while the embodiment improved various aspects of group's social perception of the agent (e.g., rapport, trust, intelligence, and power), its impact on the group-decision process and outcome was nuanced. Drawing on both quantitative and qualitative findings, we discuss the pros and cons of embodiment, argue that the value of having a face depends on the types of assistance the agent provides, and lay out directions for future research. Ameneh Shamekhi, Qingzi Vera Liao, Dakuo Wang, Rachel K. E. Bellamy, Thomas Erickson |
CHI | 2 |
| 2018 | Towards an Optimal Dialog Strategy for Information Retrieval Using Both Open- and Close-ended QuestionsabstractThe emerging paradigm of dialogue interfaces for information retrieval systems opens new opportunities for interactively narrowing down users' information query and improving search results. Prior research has largely focused on methods that use a set of close-ended questions, such as decision tree, to learn about the user's search target. However, when there is a myriad of documents or items to search, solely relying on close-ended questions can lead to long and undesirable dialogues. We propose an adaptive dialogue strategy framework that incorporates open-ended questions at the optimal timing to reduce the length of the dialogue. We propose a method to estimate the information gain of open-ended questions, and in each dialog turn, we compare it with that of close-ended questions to decide which question to ask. We present experiments using several synthetic datasets designed to explore the behavior of such an adaptive dialogue strategy under different environments, and compare the system's performance with that of a close-ended-questions-only strategy. Qingzi Vera Liao, Biplav Srivastava |
IUI | 2 |
| 2018 | What's in it for me? Self-serving versus other-oriented framing in messages advocating use of prosocial peer-to-peer services
Rajan Vaish, Qingzi Vera Liao, Victoria Bellotti |
Int. J. Hum. Comput. Stud. | 2 |
| 2017 | Leveraging Conversational Systems to Assists New Hires During Onboarding
Praveen Chandar, Yasaman Khazaeni, Michael J. Muller, Marco Crasso, Qingzi Vera Liao, N. Sadat Shami, Werner Geyer |
INTERACT (2) | 6 |
| 2016 | What Can You Do?: Studying Social-Agent Orientation and Agent Proactive Interactions with an Agent for EmployeesabstractPersonal agent software is now in daily use in personal devices and in some organizational settings. While many advocate an agent sociality design paradigm that incorporates human-like features and social dialogues, it is unclear whether this is a good match for professionals who seek productivity instead of leisurely use. We conducted a 17-day field study of a prototype of a personal AI agent that helps employees find work-related information. Using log data, surveys, and interviews, we found individual differences in the preference for humanized social interactions (social-agent orientation), which led to different user needs and requirements for agent design. We also explored the effect of agent proactive interactions and found that they carried the risk of interruption, especially for users who were generally averse to interruptions at work. Further, we found that user differences in social-agent orientation and aversion to agent proactive interactions can be inferred from behavioral signals. Our results inform research into social agent design, proactive agent interaction, and personalization of AI agents. Qingzi Vera Liao, Werner Geyer, Michael J. Muller, N. Sadat Shami |
Conference on Designing Interactive Systems | 1 |
| 2016 | #Snowden: Understanding Biases Introduced by Behavioral Differences of Opinion Groups on Social MediaabstractWe present a study of 10-month Twitter discussions on the controversial topic of Edward Snowden. We demonstrate how behavioral differences of opinion groups can distort the presence of opinions on a social media platform. By studying the differences between a numerical minority (anti-Snowden) and a majority (pro-Snowden) group, we found that the minority group engaged in a "shared audiencing" practice with more persistent production of original tweets, focusing increasingly on inter-personal interactions with like-minded others. The majority group engaged in a "gatewatching" practice by disseminating information from the group, and over time shifted further from making original comments to retweeting others'. The findings show consistency with previous social science research on how social environment shapes majority and minority group behaviors. We also highlight that they can be further distorted by the collective use of social media design features such as the "retweet" button, by introducing the concept of "amplification'" to measure how a design feature biases the voice of an opinion group. Our work presents a warning to not oversimplify analysis of social media data for inferring social opinions. Qingzi Vera Liao, Wai-Tat Fu, Markus Strohmaier |
CHI | 1 |
| 2016 | Improvising Harmony: Opportunities for Technologies to Support Crowd OrchestrationabstractThis paper details the work of a seldom studied but growing population of members of grassroots, offline-project based groups. We aim to understand how these groups self-organize to enable a large number of volunteers to gather and "get things done," and identify design opportunities for technologies to support such work. By studying the work structure, we identified two types of members, regular and episodic participants, who differ in structural role, motivation, and type of work they do. We studied two key tasks: 1) project management, which is mostly done collaboratively by the regular participants; and 2) organization of work events-the project implementation, which involve many episodic participants. For both tasks, we report on common practices and tools that are currently used. We then discuss design implications and user requirements for developing specialized tools to support these tasks. Qingzi Vera Liao, Victoria Bellotti, Michael Youngblood |
GROUP | 1 |
| 2015 | It Is All About Perspective: An Exploration of Mitigating Selective Exposure with Aspect IndicatorsabstractSelective exposure, the preferential seeking of confirmatory information, can potentially exacerbate fragmentation of online opinions and lead to biased decisions. We tested whether features that allowed users to better distinguish information about different issue aspects would encourage them to take different perspectives, thereby moderating the negative influence of pre-existing beliefs on information seeking. Using an information aggregator that provided drug related comments, we conducted an experiment to study the impact of aspect indicators (indicating whether the comment was about effectiveness or side effects) on moderating selective exposure. We found that, when participants were asked to decide between medications for high-risk diseases, and had preexisting biased beliefs in their effectiveness (one medication was less effective than the other), without aspect indicators they exhibited selective exposure to both types of comments (effectiveness and side effects) and were biased to choose the medication in confirmation of their pre-existing beliefs. With aspect indicators, we found reduced selective exposure to information about side effects of the medications, and as a result their overall decision bias was mitigated. However, the effect of aspect indicators in reducing selective exposure was moderated by the decision contexts, including the perceived risk of the diseases and whether the aspect was perceived to be critical to the decision. Qingzi Vera Liao, Wai-Tat Fu, Sri Shilpa Mamidi |
CHI | 1 |
| 2014 | Expert voices in echo chambers: effects of source expertise indicators on exposure to diverse opinionsabstractWe studied how a source expertise indicator impacted users' information seeking behavior when using a system aggregating diverse opinions, and how it interacted with a source position indicator to shape users' selectivity of information. We found that, for both attitude consistent and inconsistent information, the expertise indicator increased the selection of sources indicated to have high expertise and decreased that of low expertise. Moreover, when both source expertise and position indicators were present, users' selective exposure tendency, i.e., preferential selection of attitude consistent sources over inconsistent ones, decreased among expert sources. Moreover, we found that the expertise indicator could benefit encouraging common ground seeking with different others by increasing the agreement with, and perceived expertise of inconsistent sources indicated to be experts. Design implications for moderating selective exposure by highlighting the utility of dissonant information were discussed. Qingzi Vera Liao, Wai-Tat Fu |
CHI | 1 |
| 2014 | Can you hear me now?: mitigating the echo chamber effect by source position indicatorsabstractWe examined how a source position indicator showing both valences (pro/con) and magnitudes (moderate/extreme) of positions on controversial topics influenced users' selection and reception of diverse opinions in online discussions. Results showed that the indicator had differential impact on participants who had varied levels of accuracy motives -- i.e., motivation to accurately learn about the topic, by leading to greater exposure to attitude-challenging information for participants with higher accuracy motives. Further analysis revealed that it was mainly caused by the fact that the presence of position indicator increased the selection of moderately inconsistent sources for participants with high accuracy motives but decreased the selection of them for participants with low accuracy motives. The indicator also helped participants differentiate between sources with moderate and extreme positions, and increased their tendency to agree with attitude-challenging information from sources with moderately inconsistent positions. Participants with high accuracy motives were also found to learn significantly more about the arguments put forward by the opposite side with the help of the position indicator. We discussed the implications of the results for the nature of the echo chamber effect, as well as for designing information systems that encourage seeking of diverse information and common ground seeking. Qingzi Vera Liao, Wai-Tat Fu |
CSCW | 1 |
| 2014 | Age differences in credibility judgments of online health informationabstractOlder adults are a notable group among the exponentially growing population of online health information consumers. In order to better support older adults’ health-related information seeking on the Internet, it is important to understand how they judge the credibility of such information when compared to younger users. We conducted two laboratory studies to explore how the credibility cues in message contents, website features, and user-generated comments differentially impact younger (19 to 26 years of age) and older adults’ (58 to 80 years of age) credibility judgments. Results from the first experiment showed that older adults were less sensitive to the credibility cues in message contents and those in website features than younger adults. Verbal protocol analysis revealed that these differences could be caused by the higher tendency of older adults to passively accept web information, and their lack of deliberation on its quality and attention towards contextual web features (e.g., design look, source identity). In the second experiment, we studied how credibility cues from user reviews might differentially impact older and younger adults’ credibility judgments of online health information. Results showed that consistent credibility cues in user reviews and message contents could facilitate older adults’ credibility judgments. When the two were inconsistent, older adults, as compared to younger ones, were less swayed by highly appraising user reviews given to low credibility information. These results provided important implications for designing health information technologies that better fit the older population. Qingzi Vera Liao, Wai-Tat Fu |
ACM Trans. Comput. Hum. Interact. | 1 |
| 2013 | Beyond the filter bubble: interactive effects of perceived threat and topic involvement on selective exposure to informationabstractWe investigated participants' preferential selection of information and their attitude moderation in an online environment. Results showed that even when opposing views were presented side-to-side, people would still preferentially select information that reinforced their existing attitudes. Preferential selection of information was, however, influenced by both situational (e.g., perceived threat) and personal (e.g., topic involvement) factors. Specifically, perceived threat induced selective exposure to attitude consistent information for topics that participants had low involvement. Participants had a higher tendency to select peer user opinions in topics that they had low than high involvement, but only when there was no perception of threat. Overall, participants' attitudes were moderated after being exposed to diverse views, although high topic involvement led to higher resistance to such moderation. Perceived threat also weakened attitude moderation, especially for low involvement topics. Results have important implication to the potential effects of "information bubble" - selective exposure can be induced by situational and personal factors even when competing views are presented side-by-side. Qingzi Vera Liao, Wai-Tat Fu |
CHI | 1 |
| 2012 | Understanding experts' and novices' expertise judgment of twitter usersabstractJudging topical expertise of micro-blogger is one of the key challenges for information seekers when deciding which information sources to follow. However, it is unclear how useful different types of information are for people to make expertise judgments and to what extent their background knowledge influences their judgments. This study explored differences between experts and novices in inferring expertise of Twitter users. In three conditions, participants rated the level of expertise of users after seeing (1) only the tweets, (2) only the contextual information including short biographical and user list information, and (3) both tweets and contextual information. Results indicated that, in general, contextual information provides more useful information for making expertise judgment of Twitter users than tweets. While the addition of tweets seems to make little difference, or even add nuances to novices' expertise judgment, experts' judgments were improved when both content and contextual information were presented. Qingzi Vera Liao, Claudia Wagner 0001, Peter Pirolli, Wai-Tat Fu |
CHI | 1 |
| 2011 | Effects of Aging and Individual Differences on Credibility Judgment of Online Health Information
Qingzi Vera Liao, Wai-Tat Fu |
CogSci | 1 |
| 2011 | The Impact of User Reviews on Older and Younger Adults' Attitude towards Online Medication Information
Qingzi Vera Liao, Wai-Tat Fu |
CogSci | 1 |