VLDB 2026 Research / reviewers in the wild / expert
Sarah Wiegreffe
dblp:215/4324
· DBLP profile ↗
14ranked-venue papers
5as first author
11since 2021 · last 2026
0009-0009-5210-6716ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 5 first-author · 11 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Arguments that Alter Minds: LLM Rationales Sway Human (and LLM) Notions of PlausibilityabstractWe investigate the degree to which human (and LLM) plausibility judgments of multiplechoice commonsense benchmark answers are subject to influence by (im)plausibility arguments for or against an answer, in particular, using rationales generated by LLMs.We collect 3, 000 plausibility judgments from humans and another 13, 600 judgments from LLMs.Overall, we observe increases and decreases in mean human plausibility ratings in the presence of LLM-generated PRO and CON rationales, respectively, suggesting that, on the whole, human judges find these rationales convincing.Experiments with LLMs reveal similar patterns of influence.Our findings demonstrate a novel use of LLMs for studying aspects of human cognition, while also raising practical concerns that, even in domains where humans are "experts" (i.e., common sense), LLMs have the potential to exert considerable influence on people's beliefs. 1 Shramay Palta, Peter Rankel, Sarah Wiegreffe, Rachel Rudinger |
ACL (1) | 3 |
| 2025 | On Linear Representations and Pretraining Data Frequency in Language ModelsabstractPretraining data has a direct impact on the behaviors and quality of language models (LMs), but we only understand the most basic principles of this relationship. While most work focuses on pretraining data's effect on downstream task behavior, we investigate its relationship to LM representations. Previous work has discovered that, in language models, some concepts are encoded "linearly" in the representations, but what factors cause these representations to form (or not)? We study the connection between pretraining data frequency and models' linear representations of factual relations (e.g., mapping France to Paris in a capital prediction task). We find evidence that the formation of linear representations is strongly connected to pretraining term frequencies; specifically for subject-relation-object fact triplets, both subject-object co-occurrence frequency and in-context learning accuracy for the relation are highly correlated with linear representations. This is the case across all phases of pretraining, i.e., it is not affected by the model's underlying capability. In OLMo-7B and GPT-J (6B), we discover that a linear representation consistently (but not exclusively) forms when the subjects and objects within a relation co-occur at least 1k and 2k times, respectively, regardless of when these occurrences happen during pretraining (and around 4k times for OLMo-1B). Finally, we train a regression model on measurements of linear representation quality in fully-trained LMs that can predict how often a term was seen in pretraining. Our model achieves low error even on inputs from a different model with a different pretraining dataset, providing a new method for estimating properties of the otherwise-unknown training data of closed-data models. We conclude that the strength of linear representations in LMs contains signal about the models' pretraining corpora that may provide new avenues for controlling and improving model behavior: particularly, manipulating the models' training data to meet specific frequency thresholds. We release our code to support future work. Jack Merullo, Noah A. Smith, Sarah Wiegreffe, Yanai Elazar |
ICLR | 3 |
| 2025 | Answer, Assemble, Ace: Understanding How LMs Answer Multiple Choice QuestionsabstractMultiple-choice question answering (MCQA) is a key competence of performant transformer language models that is tested by mainstream benchmarks. However, recent evidence shows that models can have quite a range of performance, particularly when the task format is diversified slightly (such as by shuffling answer choice order). In this work we ask: how do successful models perform formatted MCQA? We employ vocabulary projection and activation patching methods to localize key hidden states that encode relevant information for predicting the correct answer. We find that prediction of a specific answer symbol is causally attributed to a few middle layers, and specifically their multi-head self-attention mechanisms. We show that subsequent layers increase the probability of the predicted answer symbol in vocabulary space, and that this probability increase is associated with a sparse set of attention heads with unique roles. We additionally uncover differences in how different models adjust to alternative symbols. Finally, we demonstrate that a synthetic task can disentangle sources of model error to pinpoint when a model has learned formatted MCQA, and show that logit differences between answer choice tokens continue to grow over the course of training. Sarah Wiegreffe, Oyvind Tafjord, Yonatan Belinkov, Hannaneh Hajishirzi, Ashish Sabharwal |
ICLR | 1 |
| 2025 | MIB: A Mechanistic Interpretability BenchmarkabstractHow can we know whether new mechanistic interpretability methods achieve real improvements?
In pursuit of lasting evaluation standards, we propose MIB, a Mechanistic Interpretability Benchmark, with two tracks spanning four tasks and five models. MIB favors methods that precisely and concisely recover relevant causal pathways or causal variables in neural language models. The circuit localization track compares methods that locate the model components---and connections between them---most important for performing a task (e.g., attribution patching or information flow routes). The causal variable track compares methods that featurize a hidden vector, e.g., sparse autoencoders (SAE) or distributed alignment search (DAS), and align those features to a task-relevant causal variable. Using MIB, we find that attribution and mask optimization methods perform best on circuit localization. For causal variable localization, we find that the supervised DAS method performs best, while SAEs features are not better than neurons, i.e., non-featurized hidden vectors. These findings illustrate that MIB enables meaningful comparisons, and increases our confidence that there has been real progress in the field. Aaron Mueller, Atticus Geiger, Sarah Wiegreffe, Dana Arad, Iván Arcuschin, Adam Belfki, Yik Siu Chan, Jaden Fiotto-Kaufman, Tal Haklay, Michael Hanna 0001, Rohan Gupta, Yaniv Nikankin, Hadas Orgad, Nikhil Prakash, Anja Reusch, Aruna Sankaranarayanan, Shun Shao, Alessandro Stolfo, Martin Tutek, Amir Zur, David Bau, Yonatan Belinkov |
ICML | 3 |
| 2024 | The Unreasonable Effectiveness of Easy Training Data for Hard TasksabstractHow can we train models to perform well on hard test data when hard training data is by definition difficult to label correctly?This question has been termed the scalable oversight problem and has drawn increasing attention as language models have continually improved.In this paper, we present the surprising conclusion that current pretrained language models often generalize relatively well from easy to hard data, even performing as well as oracle models finetuned on hard data.We demonstrate this kind of easy-to-hard generalization using simple finetuning methods like in-context learning, linear classifier heads, and QLoRA for seven different measures of datapoint hardness, including six empirically diverse human hardness measures (like grade level) and one model-based measure (loss-based).Furthermore, we show that even if one cares most about model performance on hard data, it can be better to collect easy data rather than hard data for finetuning, since hard data is generally noisier and costlier to collect.Our experiments use open models up to 70b in size and four publicly available question-answering datasets with questions ranging in difficulty from 3rd grade science questions to college level STEM questions and general-knowledge trivia.We conclude that easy-to-hard generalization in LMs is surprisingly strong for the tasks studied.1Test Input LM Generated Answer Q: John hires a driving service to get him to work each day.His work is 30 miles away and he has to go there and back each day.He goes to work 5 days a week for 50 weeks a year.He gets charged $2 per mile driven and he also gives his driver a $150 bonus per month How much does he pay a year for driving?A: John goes to work 5 days a week for 50 weeks a year.John goes to work 5 x 50 = <<5*50=250>>250 times a year.John pays $2 per mile driven.John pays $2 x 30 x 2 = <<2*30*2=120>>120 for each trip.John pays $120 x 250 = <<120*250=30000>>30000 for each year.John pays $150 bonus per month.John pays $150 x 12 = <<150*12=1800>>1800 for each year.John pays $30000 + $1800 = <<30000+1800=31800>>31800 for each year. So the answer is 31800Correct on hard problem given only easy data in prompt High-school supervision improves college performance Correct! Peter Hase, Mohit Bansal, Peter Clark, Sarah Wiegreffe |
ACL (1) | 4 |
| 2024 | The Art of Saying No: Contextual Noncompliance in Language ModelsabstractChat-based language models are designed to be helpful, yet they should not comply with every user request. While most existing work primarily focuses on refusal of ``unsafe'' queries, we posit that the scope of noncompliance should be broadened. We introduce a comprehensive taxonomy of contextual noncompliance describing when and how models should not comply with user requests. Our taxonomy spans a wide range of categories including incomplete, unsupported, indeterminate, and humanizing requests (in addition to unsafe requests). To test noncompliance capabilities of language models, we use this taxonomy to develop a new evaluation suite of 1000 noncompliance prompts. We find that most existing models show significantly high compliance rates in certain previously understudied categories with models like GPT-4 incorrectly complying with as many as 30\% of requests.To address these gaps, we explore different training strategies using a synthetically-generated training set of requests and expected noncompliant responses. Our experiments demonstrate that while direct finetuning of instruction-tuned models can lead to both over-refusal and a decline in general capabilities, using parameter efficient methods like low rank adapters helps to strike a good balance between appropriate noncompliance and other capabilities. Faeze Brahman, Sachin Kumar 0009, Vidhisha Balachandran, Pradeep Dasigi, Valentina Pyatkin, Abhilasha Ravichander, Sarah Wiegreffe, Nouha Dziri, Khyathi Raghavi Chandu, Jack Hessel, Yulia Tsvetkov, Noah A. Smith, Yejin Choi 0001, Hannaneh Hajishirzi |
NeurIPS | 7 |
| 2023 | Editing Common Sense in TransformersabstractEditing model parameters directly in Transformers makes updating open-source transformer-based models possible without re-training (Meng et al., 2023).However, these editing methods have only been evaluated on statements about encyclopedic knowledge with a single correct answer.Commonsense knowledge with multiple correct answers, e.g., an apple can be green or red but not transparent, has not been studied but is as essential for enhancing transformers' reliability and usefulness.In this paper, we investigate whether commonsense judgments are causally associated with localized, editable parameters in Transformers, and we provide an affirmative answer.We find that directly applying the MEMIT editing algorithm results in sub-par performance, and propose to improve it for the commonsense domain by varying edit tokens and improving the layer selection strategy, i.e., MEMIT CSK .GPT-2 Large and XL models edited using MEMIT CSK outperform best-fine-tuned baselines by 10.97% and 10.73% F1 scores on PEP3k and 20Q datasets.In addition, we propose a novel evaluation dataset, PROBE SET, that contains unaffected and affected neighborhoods, affected paraphrases, and affected reasoning challenges.MEMIT CSK performs well across the metrics while fine-tuning baselines show significant trade-offs between unaffected and affected metrics.These results suggest a compelling future direction for incorporating feedback about common sense into Transformers through direct model editing. 1 * Co-first and last authors.Lorraine's work done at AI2. 1 Code and datasets for all experiments are available at https://github.com/anshitag/memit_csk Anshita Gupta, Debanjan Mondal, Akshay Krishna Sheshadri, Wenlong Zhao 0001, Xiang Li 0069, Sarah Wiegreffe, Niket Tandon |
EMNLP | 6 |
| 2023 | Increasing Probability Mass on Answer Choices Does Not Always Improve AccuracyabstractWhen pretrained language models (LMs) are applied to discriminative tasks such as multiplechoice questions, they place probability mass on vocabulary tokens that aren't among the given answer choices.Spreading probability mass across multiple surface forms with identical meaning (such as "bath" and "bathtub") is thought to cause an underestimation of a model's true performance, referred to as the "surface form competition" (SFC) hypothesis.This has motivated the introduction of various probability normalization methods.However, many core questions remain unanswered.How do we measure SFC? Are there direct ways of reducing it, and does doing so improve task performance?We propose a mathematical formalism for SFC which allows us to quantify and bound its impact for the first time.We identify a simple method for reducing it-namely, increasing probability mass on the given answer choices by a) including them in the prompt and b) using in-context learning with even just one example.We show this method eliminates the impact of SFC in the majority of instances.Our experiments on three diverse datasets and six LMs reveal several additional surprising findings.For example, both normalization and prompting methods for reducing SFC can be ineffective or even detrimental to task performance for some LMs.We conclude with practical insights for effectively prompting LMs for multiple-choice tasks. 1 * Work done at AI2. 1 Code available at https://github.com/allenai/ revisiting_surface_form_competition. Sarah Wiegreffe, Matthew Finlayson, Oyvind Tafjord, Peter Clark, Ashish Sabharwal |
EMNLP | 1 |
| 2023 | Self-Refine: Iterative Refinement with Self-FeedbackabstractLike humans, large language models (LLMs) do not always generate the best output on their first try. Motivated by how humans refine their written text, we introduce Self-Refine, an approach for improving initial outputs from LLMs through iterative feedback and refinement. The main idea is to generate an initial output using an LLMs; then, the same LLMs provides *feedback* for its output and uses it to *refine* itself, iteratively. Self-Refine does not require any supervised training data, additional training, or reinforcement learning, and instead uses a single LLM as the generator, refiner and the feedback provider. We evaluate Self-Refine across 7 diverse tasks, ranging from dialog response generation to mathematical reasoning, using state-of-the-art (GPT-3.5, ChatGPT, and GPT-4) LLMs. Across all evaluated tasks, outputs generated with Self-Refine are preferred by humans and automatic metrics over those generated with the same LLM using conventional one-step generation, improving by $\sim$20\% absolute on average in task performance. Our work demonstrates that even state-of-the-art LLMs like GPT-4 can be further improved at test-time using our simple, standalone approach. Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon 0002, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang 0002, Bodhisattwa Prasad Majumder, Katherine Hermann, Sean Welleck, Amir Yazdanbakhsh, Peter Clark |
NeurIPS | 6 |
| 2022 | Reframing Human-AI Collaboration for Generating Free-Text ExplanationsabstractSarah Wiegreffe, Jack Hessel, Swabha Swayamdipta, Mark Riedl, Yejin Choi. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Sarah Wiegreffe, Jack Hessel, Swabha Swayamdipta, Mark O. Riedl, Yejin Choi 0001 |
NAACL-HLT | 1 |
| 2021 | Measuring Association Between Labels and Free-Text RationalesabstractIn interpretable NLP, we require faithful rationales that reflect the model's decision-making process for an explained instance.While prior work focuses on extractive rationales (a subset of the input words), we investigate their lessstudied counterpart: free-text natural language rationales.We demonstrate that pipelines, models for faithful rationalization on informationextraction style tasks, do not work as well on "reasoning" tasks requiring free-text rationales.We turn to models that jointly predict and rationalize, a class of widely used high-performance models for free-text rationalization.We investigate the extent to which the labels and rationales predicted by these models are associated, a necessary property of faithful explanation.Via two tests, robustness equivalence and feature importance agreement, we find that stateof-the-art T5-based joint models exhibit desirable properties for explaining commonsense question-answering and natural language inference, indicating their potential for producing faithful free-text rationales. 1 Sarah Wiegreffe, Ana Marasovic, Noah A. Smith |
EMNLP (1) | 1 |
| 2020 | Learning to Faithfully Rationalize by ConstructionabstractIn many settings it is important for one to be able to understand why a model made a particular prediction. In NLP this often entails extracting snippets of an input text ‘responsible for’ corresponding model output; when such a snippet comprises tokens that indeed informed the model’s prediction, it is a faithful explanation. In some settings, faithfulness may be critical to ensure transparency. Lei et al. (2016) proposed a model to produce faithful rationales for neural text classification by defining independent snippet extraction and prediction modules. However, the discrete selection over input tokens performed by this method complicates training, leading to high variance and requiring careful hyperparameter tuning. We propose a simpler variant of this approach that provides faithful explanations by construction. In our scheme, named FRESH, arbitrary feature importance scores (e.g., gradients from a trained model) are used to induce binary labels over token inputs, which an extractor can be trained to predict. An independent classifier module is then trained exclusively on snippets provided by the extractor; these snippets thus constitute faithful explanations, even if the classifier is arbitrarily complex. In both automatic and manual evaluations we find that variants of this simple framework yield predictive performance superior to ‘end-to-end’ approaches, while being more general and easier to train. Code is available at https://github.com/successar/FRESH. Sarah Wiegreffe, Yuval Pinter, Byron C. Wallace |
ACL | 2 |
| 2019 | Attention is not not ExplanationabstractSarah Wiegreffe, Yuval Pinter. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Sarah Wiegreffe, Yuval Pinter |
EMNLP/IJCNLP (1) | 1 |
| 2018 | Explainable Prediction of Medical Codes from Clinical TextabstractJames Mullenbach, Sarah Wiegreffe, Jon Duke, Jimeng Sun, Jacob Eisenstein. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. James Mullenbach, Sarah Wiegreffe, Jon Duke, Jimeng Sun 0001, Jacob Eisenstein |
NAACL-HLT | 2 |