VLDB 2026 Research / reviewers in the wild / expert
Sherry Tongshuang Wu
dblp:179/3791 · also Sherry Wu, Tongshuang Wu
· DBLP profile ↗
46ranked-venue papers
8as first author
35since 2021 · last 2026
0000-0003-1630-0588ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 21 · 4 first-author · 17 since 2021Artificial intelligence and machine learning · 17 · 2 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Practice Less, Explain More: LLM-Supported Self-Explanation Improves Explanation Quality on Transfer Problems in Calculus
Eason Chen, Yvonne Zhao, Meiyi Chen, Meryam Elmir, Elizabeth A. McLaughlin, Mingyu Yuan, Yumo Wang, Shyam Agarwal, Jared Cochrane, Jionghao Lin, Sherry Tongshuang Wu, Kenneth R. Koedinger |
AIED | 12 |
| 2026 | "GenAI Defaults to Bias!" Gamify AI Literacy Through Reflections on Prompts
Qianou Ma, Megan Chai, Yike Tan, Jini Kim, Erik Harpstead, Geoff Kauffman, Sherry Tongshuang Wu |
AIED | 8 |
| 2026 | Evidotes: Integrating Scientific Evidence and Anecdotes to Support Uncertainties Triggered by Peer Health PostsabstractPeer health posts surface new uncertainties, such as questions and concerns for readers. Prior work focused primarily on improving relevance and accuracy fails to address users’ diverse information needs and emotions triggered. Instead, we propose directly addressing these by information augmentation. We introduce Evidotes, an information support system that augments individual posts with relevant scientific and anecdotal information retrieved using three user-selectable lenses (dive deeper, focus on positivity, and big picture). In a mixed-methods study with 17 chronic illness patients, Evidotes improved self-reported information satisfaction (3.2 → 4.6) and reduced self-reported emotional cost (3.4 → 1.9) compared to participants’ baseline browsing. Moreover, by co-presenting sources, Evidotes unlocked information symbiosis: anecdotes made research accessible and contextual, while research helped filter and generalize peer stories. Our work enables an effective integration of scientific evidence and human anecdotes to help users better manage health uncertainty. Shreya Bali, Riku Arakawa, Peace Odiase, Sherry Tongshuang Wu, Mayank Goel |
CHI | 4 |
| 2026 | Behavioral Indicators of Overreliance During Interaction with Conversational Language ModelsabstractLLMs are now embedded in a wide range of everyday scenarios. However, their inherent hallucinations risk hiding misinformation in fluent responses, raising concerns about overreliance on AI. Detecting overreliance is challenging, as it often arises in complex, dynamic contexts and cannot be easily captured by post-hoc task outcomes. In this work, we aim to investigate how users’ behavioral patterns correlate with overreliance. We collected interaction logs from 77 participants working with an LLM injected plausible misinformation across three real-world tasks and we assessed overreliance by whether participants detected and corrected these errors. By semantically encoding and clustering segments of user interactions, we identified five behavioral patterns linked to overreliance: users with low overreliance show careful task comprehension and fine-grained navigation; users with high overreliance show frequent copy-paste, skipping initial comprehension, repeated LLM references, coarse locating, and accepting misinformation despite hesitation. We discuss design implications for mitigation. Chang Liu 0150, Qinyi Zhou, Xinjie Shen, Xingyu Liu 0002, Sherry Tongshuang Wu, Xiang 'Anthony' Chen |
CHI | 5 |
| 2026 | Not Everyone Wins with LLMs: Behavioral Patterns and Pedagogical Implications for AI Literacy in Programmatic Data ScienceabstractLLMs promise to democratize technical work in complex domains like programmatic data analysis, but not everyone benefits equally. We study how students with varied experiences use LLMs to complete Python-based data analysis in computational notebooks in a graduate course. Drawing on homework logs, recordings, and surveys from 36 students, we ask: Which experience matters most, and how does it shape AI use? Our mixed-methods analysis shows that technical experience – not AI familiarity or communication skills – remains a significant predictor of success. Students also vary widely in how they leverage LLMs, struggling at stages of forming intent, expressing inputs, interpreting outputs, and assessing results. We identify success and failure behaviors, such as providing context or decomposing prompts, that distinguish effective use. These findings inform AI literacy interventions, highlighting that lightweight demonstrations improve surface fluency but are insufficient; deeper training and scaffolds are needed to cultivate resilient AI use skills. Qianou Ma, Kenneth R. Koedinger, Sherry Tongshuang Wu |
CHI | 3 |
| 2025 | Evaluating Mathematical Reasoning Beyond AccuracyabstractThe leaderboard of Large Language Models (LLMs) in mathematical tasks has been continuously updated. However, the majority of evaluations focus solely on the final results, neglecting the quality of the intermediate steps. This oversight can mask underlying problems, such as logical errors or unnecessary steps in the reasoning process. To measure reasoning beyond final-answer accuracy, we introduce ReasonEval, a new methodology for evaluating the quality of reasoning steps. ReasonEval employs validity and redundancy to characterize the reasoning quality, as well as accompanying LLMs to assess them automatically. We explore different design options for the LLM-based evaluators and empirically demonstrate that ReasonEval, when instantiated with base models possessing strong mathematical knowledge and trained with high-quality labeled data, consistently outperforms baseline methods in the meta-evaluation datasets. We also highlight the strong generalization capabilities of ReasonEval. By utilizing ReasonEval to evaluate LLMs specialized in math, we find that an increase in final-answer accuracy does not necessarily guarantee an improvement in the overall quality of the reasoning steps for challenging mathematical problems. Additionally, we observe that ReasonEval can play a significant role in data selection. We open-source the best-performing model, meta-evaluation script, and all evaluation results to facilitate future research. Shijie Xia, Xuefeng Li 0003, Yixin Liu 0003, Sherry Tongshuang Wu, Pengfei Liu 0003 |
AAAI | 4 |
| 2025 | MoR: Better Handling Diverse Queries with a Mixture of Sparse, Dense, and Human RetrieversabstractRetrieval-augmented Generation (RAG) is powerful, but its effectiveness hinges on which retrievers we use and how.Different retrievers offer distinct, often complementary signals: BM25 captures lexical matches; dense retrievers, semantic similarity.Yet in practice, we typically fix a single retriever based on heuristics, which fails to generalize across diverse information needs.Can we dynamically select and integrate multiple retrievers for each individual query, without the need for manual selection?In our work, we validate this intuition with quantitative analysis and introduce a mixture of retrievers: a zero-shot, weighted combination of heterogeneous retrievers.Extensive experiments show that such mixtures are effective and efficient: Despite totaling just 0.8B parameters, this mixture outperforms every individual retriever and even larger 7B models-by +10.8% and +3.9% on average, respectively.Further analysis also shows that this mixture framework can help incorporate specialized non-oracle human information sources as retrievers to achieve good collaboration, with a 58.9% relative performance improvement over simulated humans alone. Jushaan Singh Kalra, To Eun Kim, Fengyu Cai, Fernando Diaz 0001, Sherry Tongshuang Wu |
EMNLP | 6 |
| 2025 | How to Teach Programming in the AI Era? Using LLMs as a Teachable Agent for Debugging (Extended Abstract)abstractLarge Language Models (LLMs) excel at generating content at impeccable speeds. However, they are imperfect and still make various mistakes. In Computer Science education, as LLMs are widely recognized as "AI pair programmers," it becomes increasingly important to train students on evaluating and debugging LLM-generated codes. In this work, we introduce HypoCompass, a novel system to facilitate deliberate practice on debugging, where human novices play the role of Teaching Assistants and help LLM-powered teachable agents debug code. We enable effective task delegation between students and LLMs in this learning-by-teaching environment: students focus on hypothesizing the cause of code errors, while adjacent skills like code completion are offloaded to LLM-agents. Our evaluations demonstrate that HypoCompass generates high-quality training materials (e.g., bugs and fixes), outperforming human counterparts fourfold in efficiency, and significantly improves student performance on debugging by 12% in the pre-to-post test. Qianou Ma, Hua Shen 0005, Kenneth R. Koedinger, Sherry Tongshuang Wu |
IJCAI | 4 |
| 2025 | Orbit: A Framework for Designing and Evaluating Multi-objective Rankers
Chenyang Yang 0002, Tesi Xiao, Michael Shavlovsky, Christian Kästner, Sherry Tongshuang Wu |
IUI | 5 |
| 2025 | Checklists Are Better Than Reward Models For Aligning Language ModelsabstractLanguage models must be adapted to understand and follow user instructions. Reinforcement learning is widely used to facilitate this —typically using fixed criteria such as "helpfulness" and "harmfulness". In our work, we instead propose using flexible, instruction-specific criteria as a means of broadening the impact that reinforcement learning can have in eliciting instruction following. We propose "Reinforcement Learning from Checklist Feedback" (RLCF). From instructions, we extract checklists and evaluate how well responses satisfy each item—using both AI judges and specialized verifier programs—then combine these scores to compute rewards for RL. We compare RLCF with other alignment methods on top of a strong instruction following model (Qwen2.5-7B-Instruct) on five widely-studied benchmarks — RLCF is the only method to help on every benchmark, including a 4-point boost in hard satisfaction rate on FollowBench, a 6-point increase on InFoBench, and a 3-point rise in win rate on Arena-Hard. We show that RLCF can also be used off-policy to improve Llama 3.1 8B Instruct and OLMo 2 7B Instruct. These results establish rubrics as a key tool for improving language models' support of queries that express a multitude of needs. We release our our dataset of rubrics (WildChecklists), models, and code to the public. Vijay Viswanathan 0002, Yanchao Sun, Xiang Kong, Graham Neubig, Sherry Tongshuang Wu |
NeurIPS | 6 |
| 2025 | What Should We Engineer in Prompts? Training Humans in Requirement-Driven LLM UseabstractPrompting LLMs for complex tasks (e.g., building a trip advisor chatbot) needs humans to clearly articulate customized requirements (e.g., “start the response with a tl;dr”). However, existing prompt engineering instructions often lack focused training on requirement articulation and instead tend to emphasize increasingly automatable strategies (e.g., tricks like adding role-plays and “think step-by-step”). To address the gap, we introduce Requirement-Oriented Prompt Engineering ( ROPE ), a paradigm that focuses human attention on generating clear, complete requirements during prompting. We implement ROPE through an assessment and training suite that provides deliberate practice with LLM-generated feedback. In a randomized controlled experiment with 30 novices, ROPE significantly outperforms conventional prompt engineering training (20% vs. 1% gains), a gap that automatic prompt optimization cannot close. Furthermore, we demonstrate a direct correlation between the quality of input requirements and LLM outputs. Our work paves the way to empower more end-users to build complex LLM applications. Qianou Ma, Weirui Peng, Chenyang Yang 0002, Hua Shen 0005, Kenneth R. Koedinger, Sherry Tongshuang Wu |
ACM Trans. Comput. Hum. Interact. | 6 |
| 2024 | How to Teach Programming in the AI Era? Using LLMs as a Teachable Agent for Debugging
Qianou Ma, Hua Shen 0005, Kenneth R. Koedinger, Sherry Tongshuang Wu |
AIED (1) | 4 |
| 2024 | Generating Situated Reflection Triggers About Alternative Solution Paths: A Case Study of Generative AI for Computer-Supported Collaborative Learning
Atharva Naik, Jessica Ruhan Yin, Anusha Kamath, Qianou Ma, Sherry Tongshuang Wu, R. Charles Murray, Christopher Bogart, Majd F. Sakr, Carolyn P. Rosé |
AIED (1) | 5 |
| 2024 | Wikibench: Community-Driven Data Curation for AI Evaluation on WikipediaabstractAI tools are increasingly deployed in community contexts. However, datasets used to evaluate AI are typically created by developers and annotators outside a given community, which can yield misleading conclusions about AI performance. How might we empower communities to drive the intentional design and curation of evaluation datasets for AI that impacts them? We investigate this question on Wikipedia, an online community with multiple AI-based content moderation tools deployed. We introduce Wikibench, a system that enables communities to collaboratively curate AI evaluation datasets, while navigating ambiguities and differences in perspective through discussion. A field study on Wikipedia shows that datasets curated using Wikibench can effectively capture community consensus, disagreement, and uncertainty. Furthermore, study participants used Wikibench to shape the overall data curation process, including refining label definitions, determining data inclusion criteria, and authoring data statements. Based on our findings, we propose future directions for systems that support community-driven data curation. Tzu-Sheng Kuo, Aaron Halfaker, Zirui Cheng, Meng-Hsin Wu, Sherry Tongshuang Wu, Kenneth Holstein, Haiyi Zhu |
CHI | 6 |
| 2024 | Selenite: Scaffolding Online Sensemaking with Comprehensive Overviews Elicited from Large Language ModelsabstractSensemaking in unfamiliar domains can be challenging, demanding considerable user effort to compare different options with respect to various criteria. Prior research and our formative study found that people would benefit from reading an overview of an information space upfront, including the criteria others previously found useful. However, existing sensemaking tools struggle with the “cold-start” problem — it not only requires significant input from previous users to generate and share these overviews, but such overviews may also turn out to be biased and incomplete. In this work, we introduce a novel system, Selenite, which leverages Large Language Models (LLMs) as reasoning machines and knowledge retrievers to automatically produce a comprehensive overview of options and criteria to jumpstart users’ sensemaking processes. Subsequently, Selenite also adapts as people use it, helping users find, read, and navigate unfamiliar information in a systematic yet personalized manner. Through three studies, we found that Selenite produced accurate and high-quality overviews reliably, significantly accelerated users’ information processing, and effectively improved their overall comprehension and sensemaking experience. Michael Xieyang Liu, Sherry Tongshuang Wu, Tianying Chen 0001, Franklin Mingzhe Li, Aniket Kittur, Brad A. Myers |
CHI | 2 |
| 2024 | What Is Wrong with My Model? Identifying Systematic Problems with Semantic Data SlicingabstractMachine learning models make mistakes, yet sometimes it is difficult to identify the systematic problems behind the mistakes. Practitioners engage in various activities, including error analysis, testing, auditing, and red-teaming, to form hypotheses of what can go (or has gone) wrong with their models. To validate these hypotheses, practitioners employ data slicing to identify relevant examples. However, traditional data slicing is limited by available features and programmatic slicing functions. In this work, we propose SemSlicer, a framework that supports semantic data slicing, which identifies a semantically coherent slice, without the need for existing features. SemSlicer uses Large Language Models to annotate datasets and generate slices from any user-defined slicing criteria. We show that SemSlicer generates accurate slices with low cost, allows flexible trade-offs between different design dimensions, reliably identifies under-performing data slices, and helps practitioners identify useful data slices that reflect systematic problems. Chenyang Yang 0002, Yining Hong, Grace A. Lewis, Sherry Tongshuang Wu, Christian Kästner |
ASE | 4 |
| 2024 | Large Language Models Help Humans Verify Truthfulness - Except When They Are Convincingly WrongabstractChenglei Si, Navita Goyal, Tongshuang Wu, Chen Zhao, Shi Feng, Hal Daumé Iii, Jordan Boyd-Graber. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Chenglei Si, Navita Goyal, Sherry Tongshuang Wu, Chen Zhao 0013, Shi Feng 0005, Hal Daumé III, Jordan L. Boyd-Graber |
NAACL-HLT | 3 |
| 2024 | Large Language Models Enable Few-Shot ClusteringabstractAbstract Unlike traditional unsupervised clustering, semi-supervised clustering allows users to provide meaningful structure to the data, which helps the clustering algorithm to match the user’s intent. Existing approaches to semi-supervised clustering require a significant amount of feedback from an expert to improve the clusters. In this paper, we ask whether a large language model (LLM) can amplify an expert’s guidance to enable query-efficient, few-shot semi-supervised text clustering. We show that LLMs are surprisingly effective at improving clustering. We explore three stages where LLMs can be incorporated into clustering: before clustering (improving input features), during clustering (by providing constraints to the clusterer), and after clustering (using LLMs post-correction). We find that incorporating LLMs in the first two stages routinely provides significant improvements in cluster quality, and that LLMs enable a user to make trade-offs between cost and accuracy to produce desired clusters. We release our code and LLM prompts for the public to use.1 Vijay Viswanathan 0002, Kiril Gashteovski, Carolin Lawrence, Sherry Tongshuang Wu, Graham Neubig |
Trans. Assoc. Comput. Linguistics | 4 |
| 2024 | Do LLMs Exhibit Human-like Response Biases? A Case Study in Survey DesignabstractAbstract One widely cited barrier to the adoption of LLMs as proxies for humans in subjective tasks is their sensitivity to prompt wording—but interestingly, humans also display sensitivities to instruction changes in the form of response biases. We investigate the extent to which LLMs reflect human response biases, if at all. We look to survey design, where human response biases caused by changes in the wordings of “prompts” have been extensively explored in social psychology literature. Drawing from these works, we design a dataset and framework to evaluate whether LLMs exhibit human-like response biases in survey questionnaires. Our comprehensive evaluation of nine models shows that popular open and commercial LLMs generally fail to reflect human-like behavior, particularly in models that have undergone RLHF. Furthermore, even if a model shows a significant change in the same direction as humans, we find that they are sensitive to perturbations that do not elicit significant changes in humans. These results highlight the pitfalls of using LLMs as human proxies, and underscore the need for finer-grained characterizations of model behavior.1 Lindia Tjuatja, Valerie Chen, Sherry Tongshuang Wu, Ameet Talwalkwar, Graham Neubig |
Trans. Assoc. Comput. Linguistics | 3 |
| 2023 | DataFinder: Scientific Dataset Recommendation from Natural Language DescriptionsabstractModern machine learning relies on datasets to develop and validate research ideas.Given the growth of publicly available data, finding the right dataset to use is increasingly difficult.Any research question imposes explicit and implicit constraints on how well a given dataset will enable researchers to answer this question, such as dataset size, modality, and domain.We operationalize the task of recommending datasets given a short natural language description of a research idea, to help people find relevant datasets for their needs.Dataset recommendation poses unique challenges as an information retrieval problem; datasets are hard to directly index for search and there are no corpora readily available for this task.To facilitate this task, we build the DataFinder Dataset which consists of a larger automatically-constructed training set (17.5K queries) and a smaller expertannotated evaluation set (392 queries).Using this data, we compare various information retrieval algorithms on our test set and present a superior bi-encoder retriever for text-based dataset recommendation.This system, trained on the DataFinder Dataset, finds more relevant search results than existing third-party dataset search engines.To encourage progress on dataset recommendation, we release our dataset and models to the public.1 Vijay Viswanathan 0002, Luyu Gao, Sherry Tongshuang Wu, Pengfei Liu 0003, Graham Neubig |
ACL (1) | 3 |
| 2023 | BiasX: "Thinking Slow" in Toxic Content Moderation with Explanations of Implied Social BiasesabstractToxicity annotators and content moderators often default to mental shortcuts when making decisions.This can lead to subtle toxicity being missed, and seemingly toxic but harmless content being over-detected.We introduce BIASX, a framework that assists content moderators with free-text explanations of statements' implied social biases, and explore its effectiveness through a large-scale user study.We show that participants indeed benefit substantially from explanations for correctly moderating subtly (non-)toxic content.The quality of explanations is critical: imperfect machine-generated explanations (+2.4% on hard toxic examples) help less compared to expert-written human explanations (+7.2%).Our results showcase the promise of using free-text explanations to encourage more thoughtful toxicity moderation. 1 Yiming Zhang 0022, Sravani Nanduri, Sherry Tongshuang Wu, Maarten Sap |
EMNLP | 4 |
| 2023 | ScatterShot: Interactive In-context Example Curation for Text TransformationabstractThe in-context learning capabilities of LLMs like GPT-3 allow annotators to customize an LLM to their specific tasks with a small number of examples. However, users tend to include only the most obvious patterns when crafting examples, resulting in underspecified in-context functions that fall short on unseen cases. Further, it is hard to know when “enough” examples have been included even for known patterns. In this work, we present ScatterShot, an interactive system for building high-quality demonstration sets for in-context learning. ScatterShot iteratively slices unlabeled data into task-specific patterns, samples informative inputs from underexplored or not-yet-saturated slices in an active learning manner, and helps users label more efficiently with the help of an LLM and the current example set. In simulation studies on two text perturbation scenarios, ScatterShot sampling improves the resulting few-shot functions by 4-5 percentage points over random sampling, with less variance as more examples are added. In a user study, ScatterShot greatly helps users in covering different patterns in the input space and labeling in-context examples more efficiently, resulting in better in-context learning and less user effort. Sherry Tongshuang Wu, Hua Shen 0005, Daniel S. Weld, Jeffrey Heer, Marco Túlio Ribeiro |
IUI | 1 |
| 2023 | Synergi: A Mixed-Initiative System for Scholarly Synthesis and SensemakingabstractEfficiently reviewing scholarly literature and synthesizing prior art are crucial for scientific progress. Yet, the growing scale of publications and the burden of knowledge make synthesis of research threads more challenging than ever.While significant research has been devoted to helping scholars interact with individual papers, building research threads scattered across multiple papers remains a challenge.Most top-down synthesis (and LLMs) make it difficult to personalize and iterate on the output, while bottom-up synthesis is costly in time and effort.Here, we explore a new design space of mixed-initiative workflows.In doing so we develop a novel computational pipeline, Synergi, that ties together user input of relevant seed threads with citation graphs and LLMs, to expand and structure them, respectively.Synergiallows scholars to start with an entire threads-and-subthreads structure generated from papers relevant to their interests, and to iterate and customize on it as they wish. In our evaluation, we find that Synergi helps scholars efficiently make sense of relevant threads, broaden their perspectives, and increases their curiosity. We discuss future design implications for thread-based, mixed-initiative scholarly synthesis support tools. Hyeonsu B. Kang, Sherry Tongshuang Wu, Joseph Chee Chang, Aniket Kittur |
UIST | 2 |
| 2023 | Bridging the Gap: A Survey on Integrating (Human) Feedback for Natural Language GenerationabstractAbstract Natural language generation has witnessed significant advancements due to the training of large language models on vast internet-scale datasets. Despite these advancements, there exists a critical challenge: These models can inadvertently generate content that is toxic, inaccurate, and unhelpful, and existing automatic evaluation metrics often fall short of identifying these shortcomings. As models become more capable, human feedback is an invaluable signal for evaluating and improving models. This survey aims to provide an overview of recent research that has leveraged human feedback to improve natural language generation. First, we introduce a taxonomy distilled from existing research to categorize and organize the varied forms of feedback. Next, we discuss how feedback can be described by its format and objective, and cover the two approaches proposed to use feedback (either for training or decoding): directly using feedback or training feedback models. We also discuss existing datasets for human-feedback data collection, and concerns surrounding feedback collection. Finally, we provide an overview of the nascent field of AI feedback, which uses large language models to make judgments based on a set of principles and minimize the need for human intervention. We also release a website of this survey at feedback-gap-survey.info. Patrick Fernandes, Aman Madaan, Emmy Liu, António Farinhas, Pedro Henrique Martins, Amanda Bertsch, José Guilherme Camargo de Souza, Shuyan Zhou, Sherry Tongshuang Wu, Graham Neubig, André F. T. Martins |
Trans. Assoc. Comput. Linguistics | 9 |
| 2023 | Towards Natural Language-Based Visualization AuthoringabstractA key challenge to visualization authoring is the process of getting familiar with the complex user interfaces of authoring tools. Natural Language Interface (NLI) presents promising benefits due to its learnability and usability. However, supporting NLIs for authoring tools requires expertise in natural language processing, while existing NLIs are mostly designed for visual analytic workflow. In this paper, we propose an authoring-oriented NLI pipeline by introducing a structured representation of users' visualization editing intents, called editing actions, based on a formative study and an extensive survey on visualization construction tools. The editing actions are executable, and thus decouple natural language interpretation and visualization applications as an intermediate layer. We implement a deep learning-based NL interpreter to translate NL utterances into editing actions. The interpreter is reusable and extensible across authoring tools. The authoring tools only need to map the editing actions into tool-specific operations. To illustrate the usages of the NL interpreter, we implement an Excel chart editor and a proof-of-concept authoring tool, VisTalk. We conduct a user study with VisTalk to understand the usage patterns of NL-based authoring systems. Finally, we discuss observations on how users author charts with natural language, as well as implications for future research. Yun Wang 0012, Zhitao Hou, Leixian Shen, Sherry Tongshuang Wu, Dongmei Zhang 0001 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2022 | Tailor: Generating and Perturbing Text with Semantic ControlsabstractControlled text perturbation is useful for evaluating and improving model generalizability.However, current techniques rely on training a model for every target perturbation, which is expensive and hard to generalize.We present Tailor, a semantically-controlled text generation system.Tailor builds on a pretrained seq2seq model and produces textual outputs conditioned on control codes derived from semantic representations.We craft a set of operations to modify the control codes, which in turn steer generation towards targeted attributes.These operations can be further composed into higher-level ones, allowing for flexible perturbation strategies.We demonstrate the effectiveness of these perturbations in multiple applications.First, we use Tailor to automatically create high-quality contrast sets for four distinct natural language processing (NLP) tasks.These contrast sets contain fewer spurious artifacts and are complementary to manually annotated ones in their lexical diversity.Second, we show that Tailor perturbations can improve model generalization through data augmentation.Perturbing just ∼2% of training data leads to a 5.8-point gain on an NLI challenge set measuring reliance on syntactic heuristics. Alexis Ross, Sherry Tongshuang Wu, Hao Peng 0009, Matthew E. Peters, Matt Gardner 0001 |
ACL (1) | 2 |
| 2022 | Fantastic Questions and Where to Find Them: FairytaleQA - An Authentic Dataset for Narrative ComprehensionabstractYing Xu, Dakuo Wang, Mo Yu, Daniel Ritchie, Bingsheng Yao, Tongshuang Wu, Zheng Zhang, Toby Li, Nora Bradford, Branda Sun, Tran Hoang, Yisi Sang, Yufang Hou, Xiaojuan Ma, Diyi Yang, Nanyun Peng, Zhou Yu, Mark Warschauer. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Dakuo Wang, Mo Yu, Daniel Ritchie 0002, Bingsheng Yao, Sherry Tongshuang Wu, Zheng Zhang 0043, Toby Jia-Jun Li, Nora Bradford, Branda Sun, Tran Bao Hoang, Yisi Sang, Yufang Hou 0001, Xiaojuan Ma, Diyi Yang, Nanyun Peng 0001, Zhou Yu 0005, Mark Warschauer |
ACL (1) | 6 |
| 2022 | It is AI's Turn to Ask Humans a Question: Question-Answer Pair Generation for Children's Story BooksabstractExisting question answering (QA) techniques are created mainly to answer questions asked by humans.But in educational applications, teachers often need to decide what questions they should ask, in order to help students to improve their narrative understanding capabilities.We design an automated question-answer generation (QAG) system for this education scenario: given a story book at the kindergarten to eighth-grade level as input, our system can automatically generate QA pairs that are capable of testing a variety of dimensions of a student's comprehension skills.Our proposed QAG model architecture is demonstrated using a new expert-annotated FairytaleQA dataset, which has 278 child-friendly storybooks with 10,580 QA pairs.Automatic and human evaluations show that our model outperforms stateof-the-art QAG baseline systems.On top of our QAG system, we also start to build an interactive story-telling application for the future real-world deployment in this educational scenario. Bingsheng Yao, Dakuo Wang, Sherry Tongshuang Wu, Zheng Zhang 0043, Toby Jia-Jun Li, Mo Yu |
ACL (1) | 3 |
| 2022 | Pretty Princess vs. Successful Leader: Gender Roles in Greeting Card MessagesabstractPeople write personalized greeting cards on various occasions. While prior work has studied gender roles in greeting card messages, systematic analysis at scale and tools for raising the awareness of gender stereotyping remain under-investigated. To this end, we collect a large greeting card message corpus covering three different occasions (birthday, Valentine’s Day and wedding) from three sources (exemplars from greeting message websites, real-life greetings from social media and language model generated ones). We uncover a wide range of gender stereotypes in this corpus via topic modeling, odds ratio and Word Embedding Association Test (WEAT). We further conduct a survey to understand people’s perception of gender roles in messages from this corpus and if gender stereotyping is a concern. The results show that people want to be aware of gender roles in the messages, but remain unconcerned unless the perceived gender roles conflict with the recipient’s true personality. In response, we developed GreetA, an interactive visualization and writing assistant tool to visualize fine-grained topics in greeting card messages drafted by the users and the associated gender perception scores, but without suggesting text changes as an intervention. Jiao Sun, Sherry Tongshuang Wu, Yue Jiang 0002, Ronil Awalegaonkar, Xi Victoria Lin, Diyi Yang |
CHI | 2 |
| 2022 | AI Chains: Transparent and Controllable Human-AI Interaction by Chaining Large Language Model PromptsabstractAlthough large language models (LLMs) have demonstrated impressive potential on simple tasks, their breadth of scope, lack of transparency, and insufficient controllability can make them less effective when assisting humans on more complex tasks. In response, we introduce the concept of Chaining LLM steps together, where the output of one step becomes the input for the next, thus aggregating the gains per step. We first define a set of LLM primitive operations useful for Chain construction, then present an interactive system where users can modify these Chains, along with their intermediate results, in a modular way. In a 20-person user study, we found that Chaining not only improved the quality of task outcomes, but also significantly enhanced system transparency, controllability, and sense of collaboration. Additionally, we saw that users developed new ways of interacting with LLMs through Chains: they leveraged sub-tasks to calibrate model expectations, compared and contrasted alternative strategies by observing parallel downstream effects, and debugged unexpected model outputs by “unit-testing” sub-components of a Chain. In two case studies, we further explore how LLM Chains may be used in future applications. Sherry Tongshuang Wu, Michael Terry, Carrie J. Cai |
CHI | 1 |
| 2022 | StoryBuddy: A Human-AI Collaborative Chatbot for Parent-Child Interactive Storytelling with Flexible Parental InvolvementabstractDespite its benefits for children’s skill development and parent-child bonding, many parents do not often engage in interactive storytelling by having story-related dialogues with their child due to limited availability or challenges in coming up with appropriate questions. While recent advances made AI generation of questions from stories possible, the fully-automated approach excludes parent involvement, disregards educational goals, and underoptimizes for child engagement. Informed by need-finding interviews and participatory design (PD) results, we developed StoryBuddy, an AI-enabled system for parents to create interactive storytelling experiences. StoryBuddy’s design highlighted the need for accommodating dynamic user needs between the desire for parent involvement and parent-child bonding and the goal of minimizing parent intervention when busy. The PD revealed varied assessment and educational goals of parents, which StoryBuddy addressed by supporting configuring question types and tracking child progress. A user study validated StoryBuddy’s usability and suggested design insights for future parent-AI collaboration systems. Zheng Zhang 0043, Bingsheng Yao, Daniel Ritchie 0002, Sherry Tongshuang Wu, Mo Yu, Dakuo Wang, Toby Jia-Jun Li |
CHI | 6 |
| 2022 | DeHumor: Visual Analytics for Decomposing HumorabstractDespite being a critical communication skill, grasping humor is challenging-a successful use of humor requires a mixture of both engaging content build-up and an appropriate vocal delivery (e.g., pause). Prior studies on computational humor emphasize the textual and audio features immediately next to the punchline, yet overlooking longer-term context setup. Moreover, the theories are usually too abstract for understanding each concrete humor snippet. To fill in the gap, we develop DeHumor, a visual analytical system for analyzing humorous behaviors in public speaking. To intuitively reveal the building blocks of each concrete example, DeHumor decomposes each humorous video into multimodal features and provides inline annotations of them on the video script. In particular, to better capture the build-ups, we introduce content repetition as a complement to features introduced in theories of computational humor and visualize them in a context linking graph. To help users locate the punchlines that have the desired features to learn, we summarize the content (with keywords) and humor feature statistics on an augmented time matrix. With case studies on stand-up comedy shows and TED talks, we show that DeHumor is able to highlight various building blocks of humor examples. In addition, expert interviews with communication coaches and humor researchers demonstrate the effectiveness of DeHumor for multimodal humor analysis of speech content and vocal delivery. Xingbo Wang 0001, Yao Ming, Sherry Tongshuang Wu, Haipeng Zeng, Yong Wang 0021, Huamin Qu |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2021 | Polyjuice: Generating Counterfactuals for Explaining, Evaluating, and Improving ModelsabstractTongshuang Wu, Marco Tulio Ribeiro, Jeffrey Heer, Daniel Weld. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Sherry Tongshuang Wu, Marco Túlio Ribeiro, Jeffrey Heer, Daniel S. Weld |
ACL/IJCNLP (1) | 1 |
| 2021 | Does the Whole Exceed its Parts? The Effect of AI Explanations on Complementary Team PerformanceabstractMany researchers motivate explainable AI with studies showing that human-AI team performance on decision-making tasks improves when the AI explains its recommendations. However, prior studies observed improvements from explanations only when the AI, alone, outperformed both the human and the best team. Can explanations help lead to complementary performance, where team accuracy is higher than either the human or the AI working solo? We conduct mixed-method user studies on three datasets, where an AI with accuracy comparable to humans helps participants solve a task (explaining itself in some conditions). While we observed complementary improvements from AI augmentation, they were not increased by explanations. Rather, explanations increased the chance that humans will accept the AI’s recommendation, regardless of its correctness. Our result poses new challenges for human-centered AI: Can we develop explanatory approaches that encourage appropriate trust in AI, and therefore help generate (or improve) complementary performance? Gagan Bansal, Sherry Tongshuang Wu, Joyce Zhou, Raymond Fok, Besmira Nushi, Ece Kamar, Marco Túlio Ribeiro, Daniel S. Weld |
CHI | 2 |
| 2021 | Beyond Accuracy: Behavioral Testing of NLP Models with Checklist (Extended Abstract)abstractAlthough measuring held-out accuracy has been the primary approach to evaluate generalization, it often overestimates the performance of NLP models, while alternative approaches for evaluating models either focus on individual tasks or on specific behaviors. Inspired by principles of behavioral testing in software engineering, we introduce CheckList, a task-agnostic methodology for testing NLP models. CheckList includes a matrix of general linguistic capabilities and test types that facilitate comprehensive test ideation, as well as a software tool to generate a large and diverse number of test cases quickly. We illustrate the utility of CheckList with tests for three tasks, identifying critical failures in both commercial and state-of-art models. In a user study, a team responsible for a commercial sentiment analysis model found new and actionable bugs in an extensively tested model. In another user study, NLP practitioners with CheckList created twice as many tests, and found almost three times as many bugs as users without it. Marco Túlio Ribeiro, Sherry Tongshuang Wu, Carlos Guestrin, Sameer Singh 0001 |
IJCAI | 2 |
| 2020 | Beyond Accuracy: Behavioral Testing of NLP Models with CheckListabstractAlthough measuring held-out accuracy has been the primary approach to evaluate generalization, it often overestimates the performance of NLP models, while alternative approaches for evaluating models either focus on individual tasks or on specific behaviors.Inspired by principles of behavioral testing in software engineering, we introduce CheckList, a taskagnostic methodology for testing NLP models.CheckList includes a matrix of general linguistic capabilities and test types that facilitate comprehensive test ideation, as well as a software tool to generate a large and diverse number of test cases quickly.We illustrate the utility of CheckList with tests for three tasks, identifying critical failures in both commercial and state-of-art models.In a user study, a team responsible for a commercial sentiment analysis model found new and actionable bugs in an extensively tested model.In another user study, NLP practitioners with CheckList created twice as many tests, and found almost three times as many bugs as users without it. Marco Túlio Ribeiro, Sherry Tongshuang Wu, Carlos Guestrin, Sameer Singh 0001 |
ACL | 2 |
| 2020 | Interactive Attention Model Explorer for Natural Language Processing Tasks with Unbalanced Data SizesabstractConventional attention visualization tools compromise either the readability or the information conveyed when documents are lengthy, especially when these documents have imbalanced sizes. Our work strives toward a more intuitive visualization for a subset of Natural Language Processing tasks, where attention is mapped between documents with imbalanced sizes. We extend the flow map visualization to enhance the readability of the attention-augmented documents. Through interaction, our design enables semantic filtering that helps users prioritize important tokens and meaningful matching for an in-depth exploration. Case studies and informal user studies in machine comprehension prove that our visualization effectively helps users gain initial understandings about what their models are "paying attention to." We discuss how the work can be extended to other domains, as well as being plugged into more end-to-end systems for model error analysis. Zhihang Dong, Sherry Tongshuang Wu, Sicheng Song |
PacificVis | 2 |
| 2020 | No Explainability without Accountability: An Empirical Study of Explanations and Feedback in Interactive MLabstractAutomatically generated explanations of how machine learning (ML) models reason can help users understand and accept them. However, explanations can have unintended consequences: promoting over-reliance or undermining trust. This paper investigates how explanations shape users' perceptions of ML models with or without the ability to provide feedback to them: (1) does revealing model flaws increase users' desire to "fix" them; (2) does providing explanations cause users to believe - wrongly - that models are introspective, and will thus improve over time. Through two controlled experiments - varying model quality - we show how the combination of explanations and user feedback impacted perceptions, such as frustration and expectations of model improvement. Explanations without opportunity for feedback were frustrating with a lower quality model, while interactions between explanation and feedback for the higher quality model suggest that detailed feedback should not be requested without explanation. Users expected model correction, regardless of whether they provided feedback or received explanations. Alison Smith-Renner, Ron Fan, Melissa Birchfield, Sherry Tongshuang Wu, Jordan L. Boyd-Graber, Daniel S. Weld, Leah Findlater |
CHI | 4 |
| 2020 | Tempura: Query Analysis with Structural TemplatesabstractAnalyzing queries from search engines and intelligent assistants is difficult. A key challenge is organizing queries into interpretable, context-preserving, representative, and flexible groups. We present structural templates, abstract queries that replace tokens with their linguistic feature forms, as a query grouping method. The templates allow analysts to create query groups with structural similarity at different granularities. We introduce Tempura, an interactive tool that lets analysts explore a query dataset with structural templates. Tempura summarizes a query dataset by selecting a representative subset of templates to show the query distribution. The tool also helps analysts navigate the template space by suggesting related templates likely to yield further explorations. Our user study shows that Tempura helps analysts examine the distribution of a query dataset, find labeling errors, and discover model error patterns and outliers. Sherry Tongshuang Wu, Kanit Wongsuphasawat, Donghao Ren, Kayur Patel, Christopher DuBois |
CHI | 1 |
| 2019 | Errudite: Scalable, Reproducible, and Testable Error AnalysisabstractThough error analysis is crucial to understanding and improving NLP models, the common practice of manual, subjective categorization of a small sample of errors can yield biased and incomplete conclusions.This paper codifies model and task agnostic principles for informative error analysis, and presents Errudite, an interactive tool for better supporting this process.First, error groups should be precisely defined for reproducibility; Errudite supports this with an expressive domainspecific language.Second, to avoid spurious conclusions, a large set of instances should be analyzed, including both positive and negative examples; Errudite enables systematic grouping of relevant instances with filtering queries.Third, hypotheses about the cause of errors should be explicitly tested; Errudite supports this via automated counterfactual rewriting.We validate our approach with a user study, finding that Errudite (1) enables users to perform high quality and reproducible error analyses with less effort, (2) reveals substantial ambiguities in prior published error analyses practices, and (3) enhances the error analysis experience by allowing users to test and revise prior beliefs. Sherry Tongshuang Wu, Marco Túlio Ribeiro, Jeffrey Heer, Daniel S. Weld |
ACL (1) | 1 |
| 2019 | Interactive Context-Aware Anomaly Detection Guided by User FeedbackabstractAutomatic anomaly detection techniques have been extensively used to support decision making in abnormal situations. However, existing approaches are limited in their capacity of effectively identifying anomalies due to the complexity of the real-world environment, the uncertainty of the data input, and the unavailability of ground truth. In this paper, we propose an interactive context-aware anomaly detection algorithm framework that incorporates human judgment in searching for anomalous regions within a large geographic environment. In specific, our framework, 1) estimates a focal region and detect anomalous situations in real time, through which the user can observe and analyze suspicious entities, 2) leverages user feedback to refine results and guide further analysis, and 3) tolerates potential fault feedback provided by the users and resignal dubious anomalous points. Based on the framework, we propose two algorithm implementations, respectively, employ Bayes’ theorem and metric learning. We demonstrate the effectiveness of the proposed framework and corresponding implementations through two controlled user studies and a case study with a domain expert. Yang Shi 0007, Maoran Xu, Rongwen Zhao, Sherry Tongshuang Wu, Nan Cao 0001 |
IEEE Trans. Hum. Mach. Syst. | 5 |
| 2019 | Local Decision Pitfalls in Interactive Machine Learning: An Investigation into Feature Selection in Sentiment AnalysisabstractTools for Interactive Machine Learning (IML) enable end users to update models in a “rapid, focused, and incremental”—yet local—manner. In this work, we study the question of local decision making in an IML context around feature selection for a sentiment classification task. Specifically, we characterize the utility of interactive feature selection through a combination of human-subjects experiments and computational simulations. We find that, in expectation, interactive modification fails to improve model performance and may hamper generalization due to overfitting. We examine how these trends are affected by the dataset, learning algorithm, and the training set size. Across these factors we observe consistent generalization issues. Our results suggest that rapid iterations with IML systems can be dangerous if they encourage local actions divorced from global context, degrading overall model performance. We conclude by discussing the implications of our feature selection results to the broader area of IML systems and research. Sherry Tongshuang Wu, Daniel S. Weld, Jeffrey Heer |
ACM Trans. Comput. Hum. Interact. | 1 |
| 2017 | NameClarifier: A Visual Analytics System for Author Name DisambiguationabstractIn this paper, we present a novel visual analytics system called NameClarifier to interactively disambiguate author names in publications by keeping humans in the loop. Specifically, NameClarifier quantifies and visualizes the similarities between ambiguous names and those that have been confirmed in digital libraries. The similarities are calculated using three key factors, namely, co-authorships, publication venues, and temporal information. Our system estimates all possible allocations, and then provides visual cues to users to help them validate every ambiguous case. By looping users in the disambiguation process, our system can achieve more reliable results than general data mining models for highly ambiguous cases. In addition, once an ambiguous case is resolved, the result is instantly added back to our system and serves as additional cues for all the remaining unidentified names. In this way, we open up the black box in traditional disambiguation processes, and help intuitively and comprehensively explain why the corresponding classifications should hold. We conducted two use cases and an expert review to demonstrate the effectiveness of NameClarifier. Qiaomu Shen, Sherry Tongshuang Wu, Huamin Qu, Weiwei Cui 0001 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2016 | STAC: Enhancing stacked graphs for time series analysisabstractStacked graphs have been widely used to represent multiple time series simultaneously to show the changes of individual values and their aggregation over time. However, when the number of time series becomes very large, the layers representing time series with small values take up only very small proportions in the stacked graph, making them hard to trace. As a result, it is challenging for analysts to detect the correlation of individual layers and their aggregation, and find trend similarities and differences between layers solely with stacked graphs. In this paper, we study the correlations of individual layers, and their aggregation in time series data presented with stacked graphs, focusing on the local regions within any given time intervals. Specifically, we present STAC, an interactive visual analytics system, to help analysts gain insights into the correlations in stacked graphs. While preserving the original stacked shape, we further link a stacked graph with auxiliary views to facilitate the in-depth analysis of correlations in time series data. A case study based on a real-world dataset demonstrates the effectiveness of our system in gaining insights into time series data analysis and facilitating various analytical tasks. Yun Wang 0012, Sherry Tongshuang Wu, Chen Zhu-Tian, Qiong Luo 0001, Huamin Qu |
PacificVis | 2 |
| 2016 | NetworkSeer: Visual analysis for social network in MOOCsabstractThe rising trend of MOOCs has attracted wide ranging research interests. Among all the existing studies related to MOOCs, most of them focus on individuals' study behaviors and evaluations (e.g., analysis on click streams for video-watching behavior exploration, etc.) for course design purposes. However, in addition to traditional course materials, MOOCs also provide interactive user forums to encourage students to seek help from peers, which endows the courses with social network formation and interaction. Thus, we present NetworkSeer to help evaluate why MOOC students use forums, and what they do. NetworkSeer visualizes interactions in the forum, including where, when the interactions happen, and why. It also enables filtering out un-targeted groups. A case study is conducted to demonstrate its usefulness. Sherry Tongshuang Wu, Yuqing Duan, Xinzhi Fan, Huamin Qu |
PacificVis | 1 |
| 2016 | PieceStack: Toward Better Understanding of Stacked GraphsabstractStacked graphs have been widely adopted in various fields, because they are capable of hierarchically visualizing a set of temporal sequences as well as their aggregation. However, because of visual illusion issues, connections between overly-detailed individual layers and overly-generalized aggregation are intercepted. Consequently, information in this area has yet to be fully excavated. Thus, we present PieceStack in this paper, to reveal the relevance of stacked graphs in understanding intrinsic details of their displayed shapes. This new visual analytic design interprets the ways through which aggregations are generated with individual layers by interactively splitting and re-constructing the stacked graphs. A clustering algorithm is designed to partition stacked graphs into sub-aggregated pieces based on trend similarities of layers. We then visualize the pieces with augmented encoding to help analysts decompose and explore the graphs with respect to their interests. Case studies and a user study are conducted to demonstrate the usefulness of our technique in understanding the formation of stacked graphs. Sherry Tongshuang Wu, Yingcai Wu, Conglei Shi, Huamin Qu, Weiwei Cui 0001 |
IEEE Trans. Vis. Comput. Graph. | 1 |