VLDB 2026 Research / reviewers in the wild / expert
Harmanpreet Kaur
dblp:167/6027
· DBLP profile ↗
20ranked-venue papers
8as first author
10since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 19 · 8 first-author · 10 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Opportunities and Barriers for AI Feedback on Meeting Inclusion in Socioorganizational TeamsabstractInclusion is important for meeting effectiveness, which is in turn central to organizational functioning. One way of improving inclusion in meetings is through feedback, but social dynamics make giving feedback difficult. We propose that AI agents can facilitate feedback exchange by being psychologically safer recipients, and we test this through a meeting system with an AI agent feedback mediator. When delivering feedback, the agent uses the Induced Hypocrisy Procedure, a social psychological technique that prompts behavior change by highlighting value-behavior inconsistencies. In a within-subjects lab study (n = 28), the agent made speaking times more balanced and improved meeting quality. However, a field study at a small consulting firm (n = 10) revealed organizational barriers that led to its use for personal reflection rather than feedback exchange. We contribute a novel sociotechnical system for feedback exchange in groups, and empirical findings demonstrating the importance of considering organizational barriers in designing AI tools for organizations. Mo Houtti, Moyan Zhou, Daniel Runningen, Surabhi Sunil, Leor Porat, Harmanpreet Kaur, Loren G. Terveen, Stevie Chancellor |
CHI | 6 |
| 2026 | An Expert Schema for Evaluating Large Language Model Errors in Scholarly Question-Answering SystemsabstractLarge Language Models (LLMs) are transforming scholarly tasks like search and summarization, but their reliability remains uncertain. Current evaluation metrics for testing LLM reliability are primarily automated approaches that prioritize efficiency and scalability, but lack contextual nuance and fail to reflect how scientific domain experts assess LLM outputs in practice. We developed and validated a schema for evaluating LLM errors in scholarly question-answering systems that reflects the assessment strategies of practicing scientists. In collaboration with domain experts, we identified 20 error patterns across seven categories through thematic analysis of 68 question-answer pairs. We validated this schema through contextual inquiries with 10 additional scientists, which showed not only which errors experts naturally identify but also how structured evaluation schemas can help them detect previously overlooked issues. Domain experts use systematic assessment strategies, including technical precision testing, value-based evaluation, and meta-evaluation of their own practices. We discuss implications for supporting expert evaluation of LLM outputs, including opportunities for personalized, schema-driven tools that adapt to individual evaluation patterns and expertise levels. Anna Martin-Boyle, William Humphreys, Martha Brown, Cara Leckey, Harmanpreet Kaur |
CHI | 5 |
| 2026 | PaperTrail: A Claim-Evidence Interface for Grounding Provenance in LLM-based Scholarly Q&AabstractLarge language models (LLMs) are increasingly used in scholarly question-answering (QA) systems to help researchers synthesize vast amounts of literature. However, these systems often produce subtle errors (e.g., unsupported claims, errors of omission), and current provenance mechanisms like source citations are not granular enough for the rigorous verification that scholarly domain requires. To address this, we introduce PaperTrail, a novel interface that decomposes both LLM answers and source documents into discrete claims and evidence, mapping them to reveal supported assertions, unsupported claims, and information omitted from the source texts. We evaluated PaperTrail in a within-subjects study with 26 researchers who performed two scholarly editing tasks using PaperTrail and a baseline interface. Our results show that PaperTrail significantly lowered participants’ trust compared to the baseline. However, this increased caution did not translate to behavioral changes, as people continued to rely on LLM-generated scholarly edits to avoid a cognitively burdensome task. We discuss the value of claim-evidence matching for understanding LLM trustworthiness in scholarly settings, and present design implications for cognition-friendly communication of provenance information. Anna Martin-Boyle, Cara Leckey, Martha Brown, Harmanpreet Kaur |
CHI | 4 |
| 2026 | Unraveling Entangled Feeds: Rethinking Social Media Design to Enhance User Well-beingabstractSocial media platforms have rapidly adopted algorithmic curation with little consideration for the potential harm to users' mental well-being. We present findings from design workshops with 21 participants diagnosed with mental illness about their interactions with social media platforms. We find that users develop cause-and-effect explanations, or folk theories, to understand their experiences with algorithmic curation. These folk theories highlight a breakdown in algorithmic design that we explain using the framework of entanglement, a phenomenon where there is a disconnect between users' actions and platform outcomes on an emotional level. Participants' designs to address entanglement and mitigate harms centered on contextualizing their engagement and restoring explicit user control on social media. The conceptualization of entanglement and the resulting design recommendations have implications for social computing and recommender systems research, particularly in evaluating and designing social media platforms that support users' mental well-being. Ashlee Milton, Daniel Runningen, Loren G. Terveen, Harmanpreet Kaur, Stevie Chancellor |
CHI | 4 |
| 2025 | Catalyst for Creativity or a Hollow Trend?: A Cross-Level Perspective on The Role of Generative AI in Design
Syeda Masooma Naqvi, Ruichen He, Harmanpreet Kaur |
CHI | 3 |
| 2025 | Wisdom of the Crowd, Without the Crowd: A Socratic LLM for Asynchronous Deliberation on Perspectivist DataabstractData annotation underpins the success of modern AI, but the aggregation of crowd-collected datasets can harm the preservation of diverse perspectives in data. Difficult and ambiguous tasks cannot easily be collapsed into unitary labels. Prior work has shown that deliberation and discussion improve data quality and preserve diverse perspectives--however, synchronous deliberation through crowdsourcing platforms is time-intensive and costly. In this work, we create a Socratic dialog system using Large Language Models (LLMs) to act as a deliberation partner in place of other crowdworkers. Against a benchmark of synchronous deliberation on two tasks (Sarcasm and Relation detection), our Socratic LLM encouraged participants to consider alternate annotation perspectives, update their labels as needed (with higher confidence), and resulted in higher annotation accuracy (for the Relation task where ground truth is available). Qualitative findings show that our agent's Socratic approach was effective at encouraging reasoned arguments from our participants, and that the intervention was well-received. Our methodology lays the groundwork for building scalable systems that preserve individual perspectives in generating more representative datasets. Malik Khadar, Daniel Runningen, Julia Tang, Stevie Chancellor, Harmanpreet Kaur |
Proc. ACM Hum. Comput. Interact. | 5 |
| 2024 | Interpretability Gone Bad: The Role of Bounded Rationality in How Practitioners Understand Machine LearningabstractWhile interpretability tools are intended to help people better understand machine learning (ML), we find that they can, in fact, impair understanding. This paper presents a pre-registered, controlled experiment showing that ML practitioners (N=119) spent 5x less time on task, and were 17% less accurate about the data and model, when given access to interpretability tools. We present bounded rationality as the theoretical reason behind these findings. Bounded rationality presumes human departures from perfect rationality, and it is often effectuated by satisficing, i.e., an inclination towards "good enough" understanding. Adding interactive elements---a strategy often employed to promote deliberative thinking and engagement, and tested in our experiment---also does not help. We discuss implications for interpretability designers and researchers related to how cognitive and contextual factors can affect the effectiveness of interpretability tool use. Harmanpreet Kaur, Matthew R. Conrad, Davis Rule, Cliff Lampe, Eric Gilbert |
Proc. ACM Hum. Comput. Interact. | 1 |
| 2022 | "I Didn't Know I Looked Angry": Characterizing Observed Emotion and Reported Affect at WorkabstractWith the growing prevalence of affective computing applications, Automatic Emotion Recognition (AER) technologies have garnered attention in both research and industry settings. Initially limited to speech-based applications, AER technologies now include analysis of facial landmarks to provide predicted probabilities of a common subset of emotions (e.g., anger, happiness) for faces observed in an image or video frame. In this paper, we study the relationship between AER outputs and self-reports of affect employed by prior work, in the context of information work at a technology company. We compare the continuous observed emotion output from an AER tool to discrete reported affect obtained via a one-day combined tool-use and diary study (N = 15). We provide empirical evidence showing that these signals do not completely align, and find that using additional workplace context only improves alignment up to 58.6%. These results suggest affect must be studied in the context it is being expressed, and observed emotion signal should not replace internal reported affect for affective computing applications. Harmanpreet Kaur, Daniel McDuff, Alex C. Williams, Jaime Teevan, Shamsi T. Iqbal |
CHI | 1 |
| 2022 | FeedLens: Polymorphic Lenses for Personalizing Exploratory Search over Knowledge GraphsabstractThe vast scale and open-ended nature of knowledge graphs (KGs) make exploratory search over them cognitively demanding for users. We introduce a new technique, polymorphic lenses, that improves exploratory search over a KG by obtaining new leverage from the existing preference models that KG-based systems maintain for recommending content. The approach is based on a simple but powerful observation: in a KG, preference models can be re-targeted to recommend not only entities of a single base entity type (e.g., papers in the scientific literature KG, products in an e-commerce KG), but also all other types (e.g., authors, conferences, institutions; sellers, buyers). We implement our technique in a novel system, FeedLens, which is built over Semantic Scholar, a production system for navigating the scientific literature KG. FeedLens reuses the existing preference models on Semantic Scholar—people’s curated research feeds—as lenses for exploratory search. Semantic Scholar users can curate multiple feeds/lenses for different topics of interest, e.g., one for human-centered AI and another for document embeddings. Although these lenses are defined in terms of papers, FeedLens re-purposes them to also guide search over authors, institutions, venues, etc. Our system design is based on feedback from intended users via two pilot surveys (n = 17 and n = 13, respectively). We compare FeedLens and Semantic Scholar via a third (within-subjects) user study (n = 15) and find that FeedLens increases user engagement while reducing the cognitive effort required to complete a short literature review task. Our qualitative results also highlight people’s preference for this more effective exploratory search experience enabled by FeedLens. Harmanpreet Kaur, Doug Downey, Amanpreet Singh, Evie Yu-Yen Cheng, Daniel S. Weld, Jonathan Bragg |
UIST | 1 |
| 2021 | From Human Explanation to Model Interpretability: A Framework Based on Weight of EvidenceabstractWe take inspiration from the study of human explanation to inform the design and evaluation of interpretability methods in machine learning. First, we survey the literature on human explanation in philosophy, cognitive science, and the social sciences, and propose a list of design principles for machine-generated explanations that are meaningful to humans. Using the concept of weight of evidence from information theory, we develop a method for generating explanations that adhere to these principles. We show that this method can be adapted to handle high-dimensional, multi-class settings, yielding a flexible framework for generating explanations. We demonstrate that these explanations can be estimated accurately from finite samples and are robust to small perturbations of the inputs. We also evaluate our method through a qualitative user study with machine learning practitioners, where we observe that the resulting explanations are usable despite some participants struggling with background concepts like prior class probabilities. Finally, we conclude by surfacing design implications for interpretability tools in general. David Alvarez-Melis, Harmanpreet Kaur, Hal Daumé III, Hanna M. Wallach, Jennifer Wortman Vaughan |
HCOMP | 2 |
| 2020 | Interpreting Interpretability: Understanding Data Scientists' Use of Interpretability Tools for Machine LearningabstractMachine learning (ML) models are now routinely deployed in domains ranging from criminal justice to healthcare. With this newfound ubiquity, ML has moved beyond academia and grown into an engineering discipline. To that end, interpretability tools have been designed to help data scientists and machine learning practitioners better understand how ML models work. However, there has been little evaluation of the extent to which these tools achieve this goal. We study data scientists' use of two existing interpretability tools, the InterpretML implementation of GAMs and the SHAP Python package. We conduct a contextual inquiry (N=11) and a survey (N=197) of data scientists to observe how they use interpretability tools to uncover common issues that arise when building and evaluating ML models. Our results indicate that data scientists over-trust and misuse interpretability tools. Furthermore, few of our participants were able to accurately describe the visualizations output by these tools. We highlight qualitative themes for data scientists' mental models of interpretability tools. We conclude with implications for researchers and tool designers, and contextualize our findings in the social science literature. Harmanpreet Kaur, Harsha Nori, Samuel Jenkins, Rich Caruana, Hanna M. Wallach, Jennifer Wortman Vaughan |
CHI | 1 |
| 2020 | Optimizing for Happiness and Productivity: Modeling Opportune Moments for Transitions and Breaks at WorkabstractInformation workers perform jobs that demand constant multitasking, leading to context switches, productivity loss, stress, and unhappiness. Systems that can mediate task transitions and breaks have the potential to keep people both productive and happy. We explore a crucial initial step for this goal: finding opportune moments to recommend transitions and breaks without disrupting people during focused states. Using affect, workstation activity, and task data from a three-week field study (N=25), we build models to predict whether a person should continue their task, transition to a new task, or take a break. The R-squared values of our models are as high as 0.7, with only 15% error cases. We ask users to evaluate the timing of recommendations provided by a recommender that relies on these models. Our study shows that users find our transition and break recommendations to be well-timed, rating them as 86% and 77% accurate, respectively. We conclude with a discussion of the implications for intelligent systems that seek to guide task transitions and manage interruptions at work. Harmanpreet Kaur, Alex C. Williams, Daniel McDuff, Mary Czerwinski, Jaime Teevan, Shamsi T. Iqbal |
CHI | 1 |
| 2020 | Using affordances to improve AI support of social media posting decisionsabstractIntelligent systems are limited in their ability to match the fluid social needs of people. We use affordances---people's perceptions of the utilities of a target system---as a means of creating models that provide intelligent systems with a better understanding of how people make decisions. We study affordance-based models in the context of social network site (SNS) usage, a domain where people have complex social needs often poorly supported by technology. Using data collected via a scenario-based survey (N=674), we build two affordance-based models about people's multi-SNS posting behavior. Our results highlight the feasibility of using affordances to help intelligent systems support people's decision-making behavior: both of our models are ~15% more accurate than a majority-class baseline, and they are ~33% and ~48% more accurate than a random baseline for this task. We contrast our approach with other ways of modeling posting behavior and discuss the implications of using affordances for modeling human behavior for intelligent systems. Harmanpreet Kaur, Cliff Lampe, Walter S. Lasecki |
IUI | 1 |
| 2019 | Mercury: Empowering Programmers' Mobile Work Practices with MicroproductivityabstractThere has been considerable research on how software can enhance programmers' productivity within their workspace. In this paper, we instead explore how software might help programmers make productive use of their time while away from their workspace. We interviewed 10 software engineers and surveyed 78 others and found that while programmers often do work while mobile, their existing mobile work practices are primarily exploratory (e.g., capturing thoughts or performing online research). In contrast, they want to be doing work that is more grounded in their existing code (e.g., code review or bug triage). Based on these findings, we introduce Mercury, a system that guides programmers in making progress on-the-go with auto-generated microtasks derived from their source code's current state. A study of Mercury with 20 programmers revealed that they could make meaningful progress with Mercury while mobile with little effort or attention. Our findings suggest an opportunity exists to support the continuation of programming tasks across devices and help programmers resume coding upon returning to their workspace. Alex C. Williams, Harmanpreet Kaur, Shamsi T. Iqbal, Ryen W. White, Jaime Teevan, Adam Fourney |
UIST | 2 |
| 2018 | Towards More Robust Speech Interactions for Deaf and Hard of Hearing UsersabstractMobile, wearable, and other ubiquitous computing devices are increasingly creating a context in which conventional keyboard and screen-based inputs are being replaced in favor of more natural speech-based interactions. Digital personal assistants use speech to control a wide range of functionality, from environmental controls to information access. However, many deaf and hard-of-hearing users have speech patterns that vary from those of hearing users due to incomplete acoustic feedback from their own voices. Because automatic speech recognition (ASR) systems are largely trained using speech from hearing individuals, speech-controlled technologies are typically inaccessible to deaf users. Prior work has focused on providing deaf users access to aural output via real-time captioning or signing, but little has been done to improve users' ability to provide input to these systems' speech-based interfaces. Further, the vocalization patterns of deaf speech often make accurate recognition intractable for both automated systems and human listeners, making traditional approaches to mitigate ASR limitations, such as human captionists, less effective. To bridge this accessibility gap, we investigate the limitations of common speech recognition approaches and techniques---both automatic and human-powered---when applied to deaf speech. We then explore the effectiveness of an iterative crowdsourcing workflow, and characterize the potential for groups to collectively exceed the performance of individuals. This paper contributes a better understanding of the challenges of deaf speech recognition and provides insights for future system development in this space. Raymond Fok, Harmanpreet Kaur, Skanda Palani, Martez E. Mott, Walter S. Lasecki |
ASSETS | 2 |
| 2018 | Supporting Workplace Detachment and Reattachment with Conversational IntelligenceabstractResearch has shown that productivity is mediated by an individual's ability to detach from their work at the end of the day and reattach with it when they return the next day. In this paper we explore the extent to which structured dialogues, focused on individuals' work-related tasks or emotions, can help them with the detachment and reattachment processes. Our inquiry is driven with SwitchBot, a conversational bot which engages with workers at the start and end of their work day. After preliminarily validating the design of a detachment and reattachment dialogue frame-work with 108 crowdworkers, we study SwitchBot's use in-situ for 14 days with 34 information workers. We find that workers send fewer e-mails after work hours and spend a larger percentage of their first hour at work using productivity applications than they normally would when using SwitchBot. Further, we find that productivity gains were better sustained when conversations focused on work-related emotions. Our results suggest that conversational bots can be effective tools for aiding workplace detachment and reattachment and help people make successful use of their time on and off the job. Alex C. Williams, Harmanpreet Kaur, Gloria Mark, Anne Loomis Thompson, Shamsi T. Iqbal, Jaime Teevan |
CHI | 2 |
| 2018 | Plexiglass: Multiplexing Passive and Active Tasks for More Efficient CrowdsourcingabstractEfficiently scaling continuous real-time crowdsourcing tasks — which engage crowd workers over long periods of time to complete tasks, such as monitoring video for critical events — is challenging largely because of the cost of keeping people consistently engaged. Worse, for many continuous tasks on which progress cannot be immediately made, which we term passive tasks, this engagement effort is wasted until something becomes true about the environment (e.g., the annotation of a critical event cannot happen until the event is seen). In this paper, we present the idea of passive-active task multiplexing, in which continuous tasks that involve waiting for a change in state are completed concurrently with traditional (offline/non-real time) tasks. We then implement this idea in Plexiglass, a system that answers visual queries made by blind and low-vision users more efficiently by concurrently presenting a (passive) real-time sensing task and an (active) offline visual question answering tasks to workers. We explore different approaches to accomplishing this, and validate that task multiplexing can lead to an improvement in efficiency in these settings of more than 40% in terms of overall worker time taken and 85% decrease in cost of running continuous real-time tasks. This work has implications on how real-time crowdsourcing tasks are presented to workers, and it increases the feasibility of deploying continuous real-time crowdsourcing systems in real-world settings. Akshay Rao, Harmanpreet Kaur, Walter S. Lasecki |
HCOMP | 2 |
| 2018 | Creating Better Action Plans for Writing Tasks via Vocabulary-Based PlanningabstractWhile having a step-by-step breakdown for a task-an action plan-helps people complete tasks, prior work has shown that people prefer not to make action plans for their own tasks. Getting planning support from others could be beneficial, but it is limited by how much domain knowledge people have about the task and how available they are. Our goal is to incorporate the benefits of having action plans in the complex domain of writing, while mitigating the time and effort costs of creating plans. To mitigate these costs, we introduce a vocabulary-a finite set of functions pertaining to writing tasks-as a cognitive scaffold that enables people with necessary context (e.g. collaborators) to generate action plans for others. We develop this vocabulary by analyzing 264 comments, and compare plans created using it with those created without any aid, in an online study with 768 comments (N=145) and a lab study with 96 comments (N=8). We show that using a vocabulary reduces planning time and effort and improves plan quality compared to unstructured planning, and opens the door for automation and task sharing for complex tasks. Harmanpreet Kaur, Alex C. Williams, Anne Loomis Thompson, Walter S. Lasecki, Shamsi T. Iqbal, Jaime Teevan |
Proc. ACM Hum. Comput. Interact. | 1 |
| 2017 | CrowdMask: Using Crowds to Preserve Privacy in Crowd-Powered Systems via Progressive FilteringabstractCrowd-powered systems leverage human intelligence to go beyond the capabilities of automated systems, but also introduce privacy and security concerns because unknown people must view the data that the system processes. While automated approaches cannot robustly filter private information from these datasets, people have the ability to do so if the risk from them viewing the data can be mitigated. We present a crowd-powered approach to masking private content in data by segmenting and distributing smaller segments to crowd workers so that individual workers can identify potentially private content without being able to fully view it themselves. We introduce a novel pyramid workflow for segmentation that uses segments at multiple levels of granularity to overcome problems with fixed-sized approaches. We implement our approach in CrowdMask, a system that allows images with potentially sensitive content to be masked by appearing in progressively larger, more identifiable segments, and masking portions of the image as soon as a risk is identified. Our experiments with 4134 Mechanical Turk workers show that CrowdMask can effectively mask private content from images without revealing sensitive content to constituent workers, while still enabling future systems to use the filtered result. Harmanpreet Kaur, Mitchell L. Gordon, Yiwei Yang 0004, Jeffrey P. Bigham, Jaime Teevan, Ece Kamar, Walter S. Lasecki |
HCOMP | 1 |
| 2015 | Putting Users in Control of their RecommendationsabstractThe essence of a recommender system is that it can recommend items personalized to the preferences of an individual user. But typically users are given no explicit control over this personalization, and are instead left guessing about how their actions affect the resulting recommendations. We hypothesize that any recommender algorithm will better fit some users' expectations than others, leaving opportunities for improvement. To address this challenge, we study a recommender that puts some control in the hands of users. Specifically, we build and evaluate a system that incorporates user-tuned popularity and recency modifiers, allowing users to express concepts like "show more popular items". We find that users who are given these controls evaluate the resulting recommendations much more positively. Further, we find that users diverge in their preferred settings, confirming the importance of giving control to users. F. Maxwell Harper, Funing Xu, Harmanpreet Kaur, Kyle Condiff, Shuo Chang, Loren G. Terveen |
RecSys | 3 |