VLDB 2026 Research / reviewers in the wild / expert
Saleema Amershi
dblp:51/681
· DBLP profile ↗
35ranked-venue papers
16as first author
6since 2021 · last 2025
0000-0002-3294-7288ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 22 · 12 first-author · 5 since 2021Artificial intelligence and machine learning · 11 · 4 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-authorDatabases, data management, data science and information retrieval · 3 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Interactive Debugging and Steering of Multi-Agent AI SystemsabstractFully autonomous teams of LLM-powered AI agents are emerging that collaborate to perform complex tasks for users. What challenges do developers face when trying to build and debug these AI agent teams? In formative interviews with five AI agent developers, we identify core challenges: difficulty reviewing long agent conversations to localize errors, lack of support in current tools for interactive debugging, and the need for tool support to iterate on agent configuration. Based on these needs, we developed an interactive multi-agent debugging tool, AGDebugger, with a UI for browsing and sending messages, the ability to edit and reset prior agent messages, and an overview visualization for navigating complex message histories. In a two-part user study with 14 participants, we identify common user strategies for steering agents and highlight the importance of interactive message resets for debugging. Our studies deepen understanding of interfaces for debugging increasingly important agentic workflows. Will Epperson, Gagan Bansal, Victor Dibia, Adam Fourney, Jack Gerrits, Erkang Zhu, Saleema Amershi |
CHI | 7 |
| 2023 | Increasing Diversity While Maintaining Accuracy: Text Data Generation with Large Language Models and Human InterventionsabstractLarge language models (LLMs) can be used to generate text data for training and evaluating other models.However, creating highquality datasets with LLMs can be challenging.In this work, we explore human-AI partnerships to facilitate high diversity and accuracy in LLM-based text data generation.We first examine two approaches to diversify text generation: 1) logit suppression, which minimizes the generation of languages that have already been frequently generated, and 2) temperature sampling, which flattens the token sampling probability.We found that diversification approaches can increase data diversity but often at the cost of data accuracy (i.e., text and labels being appropriate for the target domain).To address this issue, we examined two human interventions, 1) label replacement (LR), correcting misaligned labels, and 2) out-of-scope filtering (OOSF), removing instances that are out of the user's domain of interest or to which no considered label applies.With oracle studies, we found that LR increases the absolute accuracy of models trained with diversified datasets by 14.4%.Moreover, we found that some models trained with data generated with LR interventions outperformed LLM-based few-shot classification.In contrast, OOSF was not effective in increasing model accuracy, implying the need for future work in human-in-the-loop text data generation. John Joon Young Chung, Ece Kamar, Saleema Amershi |
ACL (1) | 3 |
| 2023 | Supporting Human-AI Collaboration in Auditing LLMs with LLMsabstractLarge language models (LLMs) are increasingly becoming all-powerful and pervasive via deployment in sociotechnical systems. Yet these language models, be it for classification or generation, have been shown to be biased, behave irresponsibly, causing harm to people at scale. It is crucial to audit these language models rigorously before deployment. Existing auditing tools use either or both humans and AI to find failures. In this work, we draw upon literature in human-AI collaboration and sensemaking, and interview research experts in safe and fair AI, to build upon the auditing tool: AdaTest [36], which is powered by a generative LLM. Through the design process we highlight the importance of sensemaking and human-AI communication to leverage complementary strengths of humans and generative models in collaborative auditing. To evaluate the effectiveness of AdaTest++, the augmented tool, we conduct user studies with participants auditing two commercial language models: OpenAI’s GPT-3 and Azure’s sentiment analysis model. Qualitative analysis shows that AdaTest++ effectively leverages human strengths such as schematization, hypothesis testing. Further, with our tool, users identified a variety of failures modes, covering 26 different topics over 2 tasks, that have been shown in formal audits and also those previously under-reported. Charvi Rastogi, Marco Túlio Ribeiro, Nicholas King, Harsha Nori, Saleema Amershi |
AIES | 5 |
| 2023 | Assessing Human-AI Interaction Early through Factorial Surveys: A Study on the Guidelines for Human-AI InteractionabstractThis work contributes a research protocol for evaluating human-AI interaction in the context of specific AI products. The research protocol enables UX and HCI researchers to assess different human-AI interaction solutions and validate design decisions before investing in engineering. We present a detailed account of the research protocol and demonstrate its use by employing it to study an existing set of human-AI interaction guidelines. We used factorial surveys with a 2 × 2 mixed design to compare user perceptions when a guideline is applied versus violated, under conditions of optimal versus sub-optimal AI performance. The results provided both qualitative and quantitative insights into the UX impact of each guideline. These insights can support creators of user-facing AI systems in their nuanced prioritization and application of the guidelines. Tianyi Li 0008, Mihaela Vorvoreanu, Derek DeBellis, Saleema Amershi |
ACM Trans. Comput. Hum. Interact. | 4 |
| 2022 | HINT: Integration Testing for AI-based features with Humans in the LoopabstractThe dynamic nature of AI technologies makes testing human-AI interaction and collaboration challenging – especially before such features are deployed in the wild. This presents a challenge for designers and AI practitioners as early feedback for iteration is often unavailable in the development phase. In this paper, we take inspiration from integration testing concepts in software development and present HINT (Human-AI INtegration Testing), a crowd-based framework for testing AI-based experiences integrated with a humans-in-the-loop workflow. HINT supports early testing of AI-based features within the context of realistic user tasks and makes use of successive sessions to simulate AI experiences that evolve over-time. Finally, it provides practitioners with reports to evaluate and compare aspects of these experiences. Quan Ze Chen, Tobias Schnabel, Besmira Nushi, Saleema Amershi |
IUI | 4 |
| 2021 | Planning for Natural Language Failures with the AI PlaybookabstractPrototyping AI user experiences is challenging due in part to probabilistic AI models making it difficult to anticipate, test, and mitigate AI failures before deployment. In this work, we set out to support practitioners with early AI prototyping, with a focus on natural language (NL)-based technologies. Our interviews with 12 NL practitioners from a large technology company revealed that, in addition to challenges prototyping AI, prototyping was often not happening at all or focused only on idealized scenarios due to a lack of tools and tight timelines. These findings informed our design of the AI Playbook, an interactive and low-cost tool we developed to encourage proactive and systematic consideration of AI errors before deployment. Our evaluation of the AI Playbook demonstrates its potential to 1) encourage product teams to prioritize both ideal and failure scenarios, 2) standardize the articulation of AI failures from a user experience perspective, and 3) act as a boundary object between user experience designers, data scientists, and engineers. Matthew K. Hong, Adam Fourney, Derek DeBellis, Saleema Amershi |
CHI | 4 |
| 2020 | Toward Responsible AI by Planning to FailabstractThe potential for AI technologies to enhance human capabilities and improve our lives is of little debate; yet, neither is their potential to cause harm and social disruption. While preventing or minimizing AI biases and harms is justifiably the subject of intense study in academic, industrial and even legal communities, an approach centered on acknowledging and planning for AI-based failures has the potential to shed new light on how to develop and deploy responsible AI-based systems. Saleema Amershi |
KDD | 1 |
| 2020 | "Who doesn't like dinosaurs?" Finding and Eliciting Richer Preferences for RecommendationabstractReal-world recommender systems often allow users to adjust the presented content through a variety of preference elicitation techniques such as “liking” or interest profiles. These elicitation techniques trade-off time and effort to users with the richness of the signal they provide to learning component driving the recommendations. In this paper, we explore this trade-off, seeking new ways for people to express their preferences with the goal of improving communication channels between users and the recommender system. Through a need-finding study, we observe the patterns in how people express their preferences during curation task, propose a taxonomy for organizing them, and point out research opportunities. We present a case study that illustrates how using this taxonomy to design an onboarding experience can lead to more accurate machine-learned recommendations while maintaining user satisfaction under low effort. Tobias Schnabel, Gonzalo A. Ramos, Saleema Amershi |
RecSys | 3 |
| 2020 | The Impact of More Transparent Interfaces on Behavior in Personalized RecommendationabstractMany interactive online systems, such as social media platforms or news sites, provide personalized experiences through recommendations or news feed customization based on people's feedback and engagement on individual items (e.g., liking items). In this paper, we investigate how we can support a greater degree of user control in such systems by changing the way the system allows people to gauge the consequences of their feedback actions. To this end, we consider two important aspects of how the system responds to feedback actions: (i) immediacy, i.e., how quickly the system responds with an update, and (ii) visibility, i.e., whether or not changes will get highlighted. We used both an in-lab qualitative study and a large-scale crowd-sourced study to examine the impact of these factors on people's reported preferences and observed behavioral metrics. We demonstrate that UX design which enables people to preview the impact of their actions and highlights changes results in a higher reported transparency, an overall preference for this design, and a greater selectivity in which items are liked. Tobias Schnabel, Saleema Amershi, Paul N. Bennett, Peter Bailey, Thorsten Joachims |
SIGIR | 2 |
| 2019 | Guidelines for Human-AI InteractionabstractAdvances in artificial intelligence (AI) frame opportunities and challenges for user interface design. Principles for human-AI interaction have been discussed in the human-computer interaction community for over two decades, but more study and innovation are needed in light of advances in AI and the growing uses of AI technologies in human-facing applications. We propose 18 generally applicable design guidelines for human-AI interaction. These guidelines are validated through multiple rounds of evaluation including a user study with 49 design practitioners who tested the guidelines against 20 popular AI-infused products. The results verify the relevance of the guidelines over a spectrum of interaction scenarios and reveal gaps in our knowledge, highlighting opportunities for further research. Based on the evaluations, we believe the set of design guidelines can serve as a resource to practitioners working on the design of applications and features that harness AI technologies, and to researchers interested in the further development of human-AI interaction design principles. Saleema Amershi, Daniel S. Weld, Mihaela Vorvoreanu, Adam Fourney, Besmira Nushi, Penny Collisson, Jina Suh, Shamsi T. Iqbal, Paul N. Bennett, Kori Inkpen, Jaime Teevan, Ruth Kikin-Gil, Eric Horvitz |
CHI | 1 |
| 2019 | Will You Accept an Imperfect AI?: Exploring Designs for Adjusting End-user Expectations of AI SystemsabstractAI technologies have been incorporated into many end-user applications. However, expectations of the capabilities of such systems vary among people. Furthermore, bloated expectations have been identified as negatively affecting perception and acceptance of such systems. Although the intelligibility of ML algorithms has been well studied, there has been little work on methods for setting appropriate expectations before the initial use of an AI-based system. In this work, we use a Scheduling Assistant - an AI system for automated meeting request detection in free-text email - to study the impact of several methods of expectation setting. We explore two versions of this system with the same 50% level of accuracy of the AI component but each designed with a different focus on the types of errors to avoid (avoiding False Positives vs. False Negatives). We show that such different focus can lead to vastly different subjective perceptions of accuracy and acceptance. Further, we design expectation adjustment techniques that prepare users for AI imperfections and result in a significant increase in acceptance. Rafal Kocielnik, Saleema Amershi, Paul N. Bennett |
CHI | 2 |
| 2019 | Sketching NLP: A Case Study of Exploring the Right Things To Design with Language IntelligenceabstractThis paper investigates how to sketch NLP-powered user experiences. Sketching is a cornerstone of design innovation. When sketching, designers rapidly experiment with a number of abstract ideas using simple, tangible instruments such as drawings and paper prototypes. Sketching NLP-powered experiences, however, presents challenges, i.e. How to visualize abstract language interaction? How to ideate a broad range of technically feasible intelligent functionalities? As a first step towards understanding these challenges, we present a first-person account of our sketching process when designing intelligent writing assistance. We detail the challenges we encountered and emergent solutions, such as a new format of wireframe for sketching language interactions and a new wizard-of-oz-based NLP rapid prototyping method. Drawing on these findings, we discuss the importance of abstraction in sketching and other implications. Qian Yang 0004, Justin Cranshaw, Saleema Amershi, Shamsi T. Iqbal, Jaime Teevan |
CHI | 3 |
| 2017 | Revolt: Collaborative Crowdsourcing for Labeling Machine Learning DatasetsabstractCrowdsourcing provides a scalable and efficient way to construct labeled datasets for training machine learning systems. However, creating comprehensive label guidelines for crowdworkers is often prohibitive even for seemingly simple concepts. Incomplete or ambiguous label guidelines can then result in differing interpretations of concepts and inconsistent labels. Existing approaches for improving label quality, such as worker screening or detection of poor work, are ineffective for this problem and can lead to rejection of honest work and a missed opportunity to capture rich interpretations about data. We introduce Revolt, a collaborative approach that brings ideas from expert annotation workflows to crowd-based labeling. Revolt eliminates the burden of creating detailed label guidelines by harnessing crowd disagreements to identify ambiguous concepts and create rich structures (groups of semantically related items) for post-hoc label decisions. Experiments comparing Revolt to traditional crowdsourced labeling show that Revolt produces high quality labels without requiring label guidelines in turn for an increase in monetary cost. This up front cost, however, is mitigated by Revolt's ability to produce reusable structures that can accommodate a variety of label boundaries without requiring new data to be collected. Further comparisons of Revolt's collaborative and non-collaborative variants show that collaboration reaches higher label accuracy with lower monetary cost. Joseph Chee Chang, Saleema Amershi, Ece Kamar |
CHI | 2 |
| 2017 | Squares: Supporting Interactive Performance Analysis for Multiclass ClassifiersabstractPerformance analysis is critical in applied machine learning because it influences the models practitioners produce. Current performance analysis tools suffer from issues including obscuring important characteristics of model behavior and dissociating performance from data. In this work, we present Squares, a performance visualization for multiclass classification problems. Squares supports estimating common performance metrics while displaying instance-level distribution information necessary for helping practitioners prioritize efforts and access data. Our controlled study shows that practitioners can assess performance significantly faster and more accurately with Squares than a confusion matrix, a common performance analysis tool in machine learning. Donghao Ren, Saleema Amershi, Bongshin Lee, Jina Suh, Jason D. Williams |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2016 | A Dataset and Evaluation Metrics for Abstractive Compression of Sentences and Short ParagraphsabstractWe introduce a manually-created, multireference dataset for abstractive sentence and short paragraph compression.First, we examine the impact of single-and multi-sentence level editing operations on human compression quality as found in this corpus.We observe that substitution and rephrasing operations are more meaning preserving than other operations, and that compressing in context improves quality.Second, we systematically explore the correlations between automatic evaluation metrics and human judgments of meaning preservation and grammaticality in the compression task, and analyze the impact of the linguistic units used and precision versus recall measures on the quality of the metrics.Multi-reference evaluation metrics are shown to offer significant advantage over single reference-based metrics. Kristina Toutanova, Chris Brockett, Ke M. Tran, Saleema Amershi |
EMNLP | 4 |
| 2016 | The Label Complexity of Mixed-Initiative Classifier TrainingabstractMixed-initiative classifier training, where the human teacher can choose which items to label or to label items chosen by the computer, has enjoyed empirical success but without a rigorous statistical learning theoretical justification. We analyze the label complexity of a simple mixed-initiative training mechanism using teach- ing dimension and active learning. We show that mixed-initiative training is advantageous com- pared to either computer-initiated (represented by active learning) or human-initiated classifier training. The advantage exists across all human teaching abilities, from optimal to completely unhelpful teachers. We further improve classifier training by educating the human teachers. This is done by showing, or explaining, optimal teaching sets to the human teachers. We conduct Mechanical Turk human experiments on two stylistic classifier training tasks to illustrate our approach. Jina Suh, Xiaojin Zhu 0001, Saleema Amershi |
ICML | 3 |
| 2016 | Active Learning with Oracle EpiphanyabstractWe present a theoretical analysis of active learning with more realistic interactions with human oracles. Previous empirical studies have shown oracles abstaining on difficult queries until accumulating enough information to make label decisions. We formalize this phenomenon with an “oracle epiphany model” and analyze active learning query complexity under such oracles for both the realizable and the agnos- tic cases. Our analysis shows that active learning is possible with oracle epiphany, but incurs an additional cost depending on when the epiphany happens. Our results suggest new, principled active learning approaches with realistic oracles. Tzu-Kuo Huang, Lihong Li 0001, Ara Vartanian, Saleema Amershi, Xiaojin Zhu 0001 |
NIPS | 4 |
| 2015 | ModelTracker: Redesigning Performance Analysis Tools for Machine LearningabstractModel building in machine learning is an iterative process. The performance analysis and debugging step typically involves a disruptive cognitive switch from model building to error analysis, discouraging an informed approach to model building. We present ModelTracker, an interactive visualization that subsumes information contained in numerous traditional summary statistics and graphs while displaying example-level performance and enabling direct error examination and debugging. Usage analysis from machine learning practitioners building real models with ModelTracker over six months shows ModelTracker is used often and throughout model building. A controlled experiment focusing on ModelTracker's debugging capabilities shows participants prefer ModelTracker over traditional tools without a loss in model performance. Saleema Amershi, David Maxwell Chickering, Steven Mark Drucker, Bongshin Lee, Patrice Y. Simard, Jina Suh |
CHI | 1 |
| 2014 | Structured labeling for facilitating concept evolution in machine learningabstractLabeling data is a seemingly simple task required for training many machine learning systems, but is actually fraught with problems. This paper introduces the notion of concept evolution, the changing nature of a person's underlying concept (the abstract notion of the target class a person is labeling for, e.g., spam email, travel related web pages) which can result in inconsistent labels and thus be detrimental to machine learning. We introduce two structured labeling solutions, a novel technique we propose for helping people define and refine their concept in a consistent manner as they label. Through a series of five experiments, including a controlled lab study, we illustrate the impact and dynamics of concept evolution in practice and show that structured labeling helps people label more consistently in the presence of concept evolution than traditional labeling. Todd Kulesza, Saleema Amershi, Rich Caruana, Danyel Fisher, Denis Xavier Charles |
CHI | 2 |
| 2013 | LiveAction: Automating Web Task Model GenerationabstractTask automation systems promise to increase human productivity by assisting us with our mundane and difficult tasks. These systems often rely on people to (1) identify the tasks they want automated and (2) specify the procedural steps necessary to accomplish those tasks (i.e., to create task models). However, our interviews with users of a Web task automation system reveal that people find it difficult to identify tasks to automate and most do not even believe they perform repetitive tasks worthy of automation. Furthermore, even when automatable tasks are identified, the well-recognized difficulties of specifying task steps often prevent people from taking advantage of these automation systems. In this research, we analyze real Web usage data and find that people do in fact repeat behaviors on the Web and that automating these behaviors, regardless of their complexity, would reduce the overall number of actions people need to perform when completing their tasks, potentially saving time. Motivated by these findings, we developed LiveAction, a fully-automated approach to generating task models from Web usage data. LiveAction models can be used to populate the task model repositories required by many automation systems, helping us take advantage of automation in our everyday lives. Saleema Amershi, Jalal Mahmud, Jeffrey Nichols 0001, Tessa A. Lau, German Attanasio Ruiz |
ACM Trans. Interact. Intell. Syst. | 1 |
| 2012 | Regroup: interactive machine learning for on-demand group creation in social networksabstractWe present ReGroup, a novel end-user interactive machine learning system for helping people create custom, on demand groups in online social networks. As a person adds members to a group, ReGroup iteratively learns a probabilistic model of group membership specific to that group. ReGroup then uses its currently learned model to suggest additional members and group characteristics for filtering. Our evaluation shows that ReGroup is effective for helping people create large and varied groups, whereas traditional methods (searching by name or selecting from an alphabetical list) are better suited for small groups whose members can be easily recalled by name. By facilitating on demand group creation, ReGroup can enable in-context sharing and potentially encourage better online privacy practices. In addition, applying interactive machine learning to social network group creation introduces several challenges for designing effective end-user interaction with machine learning. We identify these challenges and discuss how we address them in ReGroup. Saleema Amershi, James Fogarty, Daniel S. Weld |
CHI | 1 |
| 2011 | Effective End-User Interaction with Machine LearningabstractEnd-user interactive machine learning is a promising tool for enhancing human productivity and capabilities with large unstructured data sets. Recent work has shown that we can create end-user interactive machine learning systems for specific applications. However, we still lack a generalized understanding of how to design effective end-user interaction with interactive machine learning systems. This work presents three explorations in designing for effective end-user interaction with machine learning in CueFlik, a system developed to support Web image search. These explorations demonstrate that interactions designed to balance the needs of end-users and machine learning algorithms can significantly improve the effectiveness of end-user interactive machine learning. Saleema Amershi, James Fogarty, Ashish Kapoor, Desney S. Tan |
AAAI | 1 |
| 2011 | CueT: human-guided fast and accurate network alarm triageabstractNetwork alarm triage refers to grouping and prioritizing a stream of low-level device health information to help operators find and fix problems. Today, this process tends to be largely manual because existing tools cannot easily evolve with the network. We present CueT, a system that uses interactive machine learning to learn from the triaging decisions of operators. It then uses that learning in novel visualizations to help them quickly and accurately triage alarms. Unlike prior interactive machine learning systems, CueT handles a highly dynamic environment where the groups of interest are not known a-priori and evolve constantly. A user study with real operators and data from a large network shows that CueT significantly improves the speed and accuracy of alarm triage compared to the network's current practice. Saleema Amershi, Bongshin Lee, Ashish Kapoor, Ratul Mahajan, Blaine Christian |
CHI | 1 |
| 2011 | Human-Guided Machine Learning for Fast and Accurate Network Alarm Triage
Saleema Amershi, Bongshin Lee, Ashish Kapoor, Ratul Mahajan, Blaine Christian |
IJCAI | 1 |
| 2010 | Examining multiple potential models in end-user interactive concept learningabstractEnd-user interactive concept learning is a technique for interacting with large unstructured datasets, requiring insights from both human-computer interaction and machine learning. This note re-examines an assumption implicit in prior interactive machine learning research, that interaction should focus on the question "what class is this object?". We broaden interaction to include examination of multiple potential models while training a machine learning system. We evaluate this approach and find that people naturally adopt revision in the interactive machine learning process and that this improves the quality of their resulting models for difficult concepts. Saleema Amershi, James Fogarty, Ashish Kapoor, Desney S. Tan |
CHI | 1 |
| 2010 | Multiple mouse text entry for single-display groupwareabstractA recent trend in interface design for classrooms in developing regions has many students interacting on the same display using mice. Text entry has emerged as an important problem preventing such mouse-based singledisplay groupware systems from offering compelling interactive activities. We explore the design space of mouse-based text entry and develop 13 techniques with novel characteristics suited to the multiple mouse scenario. We evaluated these in a 3-phase study over 14 days with 40 students in 2 developing region schools. The results show that one technique effectively balanced all of our design dimensions, another was most preferred by students, and both could benefit from augmentation to support collaborative interaction. Our results also provide insights into the factors that create an optimal text entry technique for single-display groupware systems. Saleema Amershi, Meredith Ringel Morris, Neema Moraveji, Ravin Balakrishnan, Kentaro Toyama |
CSCW | 1 |
| 2009 | Amplifying community content creation with mixed initiative information extractionabstractAlthough existing work has explored both information extraction and community content creation, most research has focused on them in isolation. In contrast, we see the greatest leverage in the synergistic pairing of these methods as two interlocking feedback cycles. This paper explores the potential synergy promised if these cycles can be made to accelerate each other by exploiting the same edits to advance both community content creation and learning-based information extraction. We examine our proposed synergy in the context of Wikipedia infoboxes and the Kylin information extraction system. After developing and refining a set of interfaces to present the verification of Kylin extractions as a non primary task in the context of Wikipedia articles, we develop an innovative use of Web search advertising services to study people engaged in some other primary task. We demonstrate our proposed synergy by analyzing our deployment from two complementary perspectives: (1) we show we accelerate community content creation by using Kylin's information extraction to significantly increase the likelihood that a person visiting a Wikipedia article as a part of some other primary task will spontaneously choose to help improve the article's infobox, and (2) we show we accelerate information extraction by using contributions collected from people interacting with our designs to significantly improve Kylin's extraction performance. Raphael Hoffmann, Saleema Amershi, Kayur Patel, Fei Wu 0003, James Fogarty, Daniel S. Weld |
CHI | 2 |
| 2009 | Overview based example selection in end user interactive concept learningabstractInteraction with large unstructured datasets is difficult because existing approaches, such as keyword search, are not always suited to describing concepts corresponding to the distinctions people want to make within datasets. One possible solution is to allow end users to train machine learning systems to identify desired concepts, a strategy known as interactive concept learning. A fundamental challenge is to design systems that preserve end user flexibility and control while also guiding them to provide examples that allow the machine learning system to effectively learn the desired concept. This paper presents our design and evaluation of four new overview based approaches to guiding example selection. We situate our explorations within CueFlik, a system examining end user interactive concept learning in Web image search. Our evaluation shows our approaches not only guide end users to select better training examples than the best performing previous design for this application, but also reduce the impact of not knowing when to stop training the system. We discuss challenges for end user interactive concept learning systems and identify opportunities for future research on the effective design of such systems. Saleema Amershi, James Fogarty, Ashish Kapoor, Desney S. Tan |
UIST | 1 |
| 2008 | Intelligence in Wikipedia
Daniel S. Weld, Fei Wu 0003, Eytan Adar, Saleema Amershi, James Fogarty, Raphael Hoffmann, Kayur Patel, Michael Skinner |
AAAI | 4 |
| 2008 | CoSearch: a system for co-located collaborative web searchabstractWeb search is often viewed as a solitary task; however, there are many situations in which groups of people gather around a single computer to jointly search for information online. We present the findings of interviews with teachers, librarians, and developing world researchers that provide details about users' collaborative search habits in shared-computer settings, revealing several limitations of this practice. We then introduce CoSearch, a system we developed to improve the experience of co-located collaborative Web search by leveraging readily available devices such as mobile phones and extra mice. Finally, we present an evaluation comparing CoSearch to status quo collaboration approaches, and show that CoSearch enabled distributed control and division of labor, thus reducing the frustrations associated with shared-computer searches, while still preserving the positive aspects of communication and collaboration associated with joint computer use. Saleema Amershi, Meredith Ringel Morris |
CHI | 1 |
| 2008 | Pedagogy and usability in interactive algorithm visualizations: Designing and evaluating CIspaceabstractInteractive algorithm visualizations (AVs) are powerful tools for teaching and learning concepts that are difficult to describe with static media alone. However, while countless AVs exist, their widespread adoption by the academic community has not occurred due to usability problems and mixed results of pedagogical effectiveness reported in the AV and education literature. This paper presents our experiences designing and evaluating CIspace, a set of interactive AVs for demonstrating fundamental Artificial Intelligence algorithms. In particular, we first review related work on AVs and theories of learning. Then, from this literature, we extract and compile a taxonomy of goals for designing interactive AVs that address key pedagogical and usability limitations of existing AVs. We advocate that differentiating between goals and design features that implement these goals will help designers of AVs make more informed choices, especially considering the abundance of often conflicting and inconsistent design recommendations in the AV literature. We also describe and present the results of a range of evaluations that we have conducted on CIspace that include semi-formal usability studies, usability surveys from actual students using CIspace as a course resource, and formal user studies designed to assess the pedagogical effectiveness of CIspace in terms of both knowledge gain and user preference. Our main results show that (i) studying with our interactive AVs is at least as effective at increasing student knowledge as studying with carefully designed paper-based materials; (ii) students like using our interactive AVs more than studying with the paper-based materials; (iii) students use both our interactive AVs and paper-based materials in practice although they are divided when forced to choose between them; (iv) students find our interactive AVs generally easy to use and useful. From these results, we conclude that while interactive AVs may not be universally preferred by students, it is beneficial to offer a variety of learning media to students to accommodate individual learning preferences. We hope that our experiences will be informative for other developers of interactive AVs, and encourage educators to exploit these potentially powerful resources in classrooms and other learning environments. Saleema Amershi, Giuseppe Carenini, Cristina Conati, Alan K. Mackworth, David Poole 0001 |
Interact. Comput. | 1 |
| 2007 | Using Eye-Tracking Data for High-Level User Modeling in Adaptive Interfaces
Cristina Conati, Christina Merten, Saleema Amershi, Kasia Muldner |
AAAI | 3 |
| 2007 | Unsupervised and supervised machine learning in user modeling for intelligent learning environmentsabstractIn this research, we outline a user modeling framework that uses both unsupervised and supervised machine learning in order to reduce development costs of building user models, and facilitate transferability. We apply the framework to model student learning during interaction with the Adaptive Coach for Exploration (ACE) learning environment (using both interface and eye-tracking data). In addition to demonstrating framework effectiveness, we also compare results from previous research on applying the framework to a different learning environment and data type. Our results also confirm previous research on the value of using eye-tracking data to assess student learning. Saleema Amershi, Cristina Conati |
IUI | 1 |
| 2006 | Automatic Recognition of Learner Groups in Exploratory Learning Environments
Saleema Amershi, Cristina Conati |
Intelligent Tutoring Systems | 1 |
| 2005 | Designing CIspace: pedagogy and usability in a learning environment for AIabstractThis paper describes the design of the CIspace interactive visualization tools for teaching and learning Artificial Intelligence. Our approach to design is to iterate through three phases: identifying pedagogical and usability goals for supporting both educators and students, designing to achieve these goals, and then evaluating our system. We believe identifying these goals is essential in confronting the usability deficiencies and mixed results about the pedagogical effectiveness of interactive visualizations reported in the Education literature. The CIspace tools have been used and positively received in undergraduate and graduate classrooms at the University of British Columbia and internationally. We hope that our experiences can inform other developers of interactive visualizations and encourage their use in classrooms and other learning environments. Saleema Amershi, N. Arksey, Giuseppe Carenini, Cristina Conati, Alan K. Mackworth, Heather Maclaren, David Poole 0001 |
ITiCSE | 1 |