VLDB 2026 Research / reviewers in the wild / expert
Christopher Potts
dblp:13/2617
· DBLP profile ↗
82ranked-venue papers
1as first author
54since 2021 · last 2025
0000-0002-7978-6055ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 77 · 1 first-author · 52 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 5 since 2021Databases, data management, data science and information retrieval · 7 · 3 since 2021Human-computer interaction and ubiquitous computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Exploring the mechanisms that enable multimodal reasoning about data visualizations in vision-language models
Alexa R. Tartaglini, Christopher Potts, Judith E. Fan |
CogSci | 2 |
| 2025 | ColBERT-Serve: Efficient Multi-stage Memory-Mapped Scoring
Kaili Huang, Thejas Venkatesh, Uma Dingankar, Antonio Mallia, Daniel Campos, Christopher Potts, Matei Zaharia, Kwabena Boahen 0001, Omar Khattab, Saarthak Sarup, Keshav Santhanam |
ECIR (4) | 7 |
| 2025 | Causal Interventions Reveal Shared Structure Across English Filler-Gap ConstructionsabstractLanguage Models (LMs) have emerged as powerful sources of evidence for linguists seeking to develop theories of syntax.In this paper, we argue that causal interpretability methods, applied to LMs, can greatly enhance the value of such evidence by helping us characterize the abstract mechanisms that LMs learn to use.Our empirical focus is a set of English filler-gap dependency constructions (e.g., questions, relative clauses).Linguistic theories largely agree that these constructions share many properties.Using experiments based in Distributed Interchange Interventions, we show that LMs converge on similar abstract analyses of these constructions.These analyses also reveal previously overlooked factorsrelating to frequency, filler type, and surrounding context -that could motivate changes to standard linguistic theory.Overall, these results suggest that mechanistic, internal analyses of LMs can push linguistic theory forward.https://github.com/SashaBoguraev/ causal-filler-gap Sasha Boguraev, Christopher Potts, Kyle Mahowald |
EMNLP | 2 |
| 2025 | MrT5: Dynamic Token Merging for Efficient Byte-level Language ModelsabstractModels that rely on subword tokenization have significant drawbacks, such as sensitivity to character-level noise like spelling errors and inconsistent compression rates across different languages and scripts. While character- or byte-level models like ByT5 attempt to address these concerns, they have not gained widespread adoption—processing raw byte streams without tokenization results in significantly longer sequence lengths, making training and inference inefficient. This work introduces MrT5 (MergeT5), a more efficient variant of ByT5 that integrates a token deletion mechanism in its encoder to dynamically shorten the input sequence length. After processing through a fixed number of encoder layers, a learned delete gate determines which tokens are to be removed and which are to be retained for subsequent layers. MrT5 effectively "merges" critical information from deleted tokens into a more compact sequence, leveraging contextual information from the remaining tokens. In continued pre-training experiments, we find that MrT5 can achieve significant gains in inference runtime with minimal effect on performance, as measured by bits-per-byte. Additionally, with multilingual training, MrT5 adapts to the orthographic characteristics of each language, learning language-specific compression rates. Furthermore, MrT5 shows comparable accuracy to ByT5 on downstream evaluations such as XNLI, TyDi QA, and character-level tasks while reducing sequence lengths by up to 75%. Our approach presents a solution to the practical limitations of existing byte-level models. Julie Kallini, Shikhar Murty, Christopher D. Manning, Christopher Potts, Róbert Csordás |
ICLR | 4 |
| 2025 | HyperDAS: Towards Automating Mechanistic Interpretability with HypernetworksabstractMechanistic interpretability has made great strides in identifying neural network features (e.g., directions in hidden activation space) that mediate concepts (e.g., *the birth year of a Nobel laureate*) and enable predictable manipulation. Distributed alignment search (DAS) leverages supervision from counterfactual data to learn concept features within hidden states, but DAS assumes we can afford to conduct a brute force search over potential feature locations. To address this, we present HyperDAS, a transformer-based hypernetwork architecture that (1) automatically locates the token-positions of the residual stream that a concept is realized in and (2) learns features of those residual stream vectors for the concept. In experiments with Llama3-8B, HyperDAS achieves state-of-the-art performance on the RAVEL benchmark for disentangling concepts in hidden states. In addition, we review the design decisions we made to mitigate the concern that HyperDAS (like all powerful interpretabilty methods) might inject new information into the target model rather than faithfully interpreting it. Jiuding Sun, Jing Huang 0014, Sidharth Baskaran, Karel D'Oosterlinck, Christopher Potts, Michael Sklar, Atticus Geiger |
ICLR | 5 |
| 2025 | Improving Pretraining Data Using Perplexity CorrelationsabstractQuality pretraining data is often seen as the key to high-performance language models. However, progress in understanding pretraining data has been slow due to the costly pretraining runs required for data selection experiments. We present a framework that avoids these costs and selects high-quality pretraining data without any LLM training of our own. Our work is based on a simple observation: LLM losses on many pretraining texts are correlated with downstream benchmark performance, and selecting high-correlation documents is an effective pretraining data selection method. We build a new statistical framework for data selection centered around estimates of perplexity-benchmark correlations and perform data selection using a sample of 90 LLMs taken from the Open LLM Leaderboard on texts from tens of thousands of web domains. In controlled pretraining experiments at the 160M parameter scale on 8 benchmarks, our approach outperforms DSIR on every benchmark, while matching the best data selector found in DataComp-LM, a hand-engineered bigram classifier. We have now also updated this paper to include results from preregistered experiments with new pretraining data on an aggregation of 22 benchmarks up to the 1.4B scale, showing increasing improvements of our method over others with more scale. A pip package with full documentation can be found here: https://github.com/TristanThrush/perplexity-correlations. Tristan Thrush, Christopher Potts, Tatsunori B. Hashimoto |
ICLR | 2 |
| 2025 | Internal Causal Mechanisms Robustly Predict Language Model Out-of-Distribution BehaviorsabstractInterpretability research now offers a variety of techniques for identifying abstract internal mechanisms in neural networks. Can such techniques be used to predict how models will behave on out-of-distribution examples? In this work, we provide a positive answer to this question. Through a diverse set of language modeling tasks—including symbol manipulation, knowledge retrieval, and instruction following—we show that the most robust features for correctness prediction are those that play a distinctive causal role in the model’s behavior. Specifically, we propose two methods that leverage causal mechanisms to predict the correctness of model outputs: counterfactual simulation (checking whether key causal variables are realized) and value probing (using the values of those variables to make predictions). Both achieve high AUC-ROC in distribution and outperform methods that rely on causal-agnostic features in out-of-distribution settings, where predicting model behaviors is more crucial. Our work thus highlights a novel and significant application for internal causal analysis of language models. Jing Huang 0014, Junyi Tao, Thomas Icard, Diyi Yang, Christopher Potts |
ICML | 5 |
| 2025 | AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse AutoencodersabstractFine-grained steering of language model outputs is essential for safety and reliability. Prompting and finetuning are widely used to achieve these goals, but interpretability researchers have proposed a variety of representation-based techniques as well, including sparse autoencoders (SAEs), linear artificial tomography, supervised steering vectors, linear probes, and representation finetuning. At present, there is no benchmark for making direct comparisons between these proposals. Therefore, we introduce AxBench, a large-scale benchmark for steering and concept detection, and report experiments on Gemma-2-2B and 9B. For steering, we find that prompting outperforms all existing methods, followed by finetuning. For concept detection, representation-based methods such as difference-in-means, perform the best. On both evaluations, SAEs are not competitive. We introduce a novel weakly-supervised representational method (Rank-1 Representation Finetuning; ReFT-r1), which is competitive on both tasks while providing the interpretability advantages that prompting lacks. Along with AxBench, we train and publicly release SAE-scale feature dictionaries for ReFT-r1 and DiffMean. Zhengxuan Wu, Aryaman Arora, Atticus Geiger, Zheng Wang 0078, Jing Huang 0014, Daniel Jurafsky, Christopher D. Manning, Christopher Potts |
ICML | 8 |
| 2025 | Do Language Models Use Their Depth Efficiently?abstractModern LLMs are increasingly deep, and depth correlates with performance, albeit with diminishing returns. However, do these models use their depth efficiently? Do they compose more features to create higher-order computations that are impossible in shallow models, or do they merely spread the same kinds of computation out over more layers? To address these questions, we analyze the residual stream of the Llama 3.1, Qwen 3, and OLMo 2 family of models. We find: First, comparing the output of the sublayers to the residual stream reveals that layers in the second half contribute much less than those in the first half, with a clear phase transition between the two halves. Second, skipping layers in the second half has a much smaller effect on future computations and output predictions. Third, for multihop tasks, we are unable to find evidence that models are using increased depth to compose subresults in examples involving many hops. Fourth, we seek to directly address whether deeper models are using their additional layers to perform new kinds of computation. To do this, we train linear maps from the residual stream of a shallow model to a deeper one. We find that layers with the same relative depth map best to each other, suggesting that the larger model simply spreads the same computations out over its many layers. All this evidence suggests that deeper models are not using their depth to learn new kinds of computation, but only using the greater depth to perform more fine-grained adjustments to the residual. This may help explain why increasing scale leads to diminishing returns for stacked Transformer architectures. Róbert Csordás, Christopher D. Manning, Christopher Potts |
NeurIPS | 3 |
| 2025 | Blackbox Model Provenance via Palimpsestic Membership InferenceabstractSuppose Alice trains an open-weight language model and Bob uses a blackbox derivative of Alice’s model to produce text. Can Alice prove that Bob is using her model, either by querying Bob’s derivative model (query setting) or from the text alone ( observational setting)? We formulate this question as an independence testing problem—in which the null hypothesis is that Bob’s model or text is independent of Alice’s randomized training run—and investigate it through the lens of palimpsestic memorization in language models: models are more likely to memorize data seen later in training, so we can test whether Bob is using Alice’s model using test statistics that capture correlation between Bob’s model or text and the ordering of training examples in Alice’s training run. If Alice has randomly shuffled her training data, then any significant correlation amounts to exactly quantifiable statistical evidence against the null hypothesis, regardless of the composition of Alice’s training data. In the query setting, we directly estimate (via prompting) the likelihood Bob’s model gives to Alice’s training examples and their training order; we correlate the likelihoods of over 40 fine-tunes of various Pythia and OLMo base models ranging from 1B to 12B parameters with the base model’s training data order, achieving a p-value on the order of at most $1 \times 10^{-8}$ in all but six cases. In the observational setting, we try two approaches based on estimating 1) the likelihood of Bob’s text overlapping with spans of Alice’s training examples and 2) the likelihood of Bob’s text with respect to different versions of Alice’s model we obtain by repeating the last phase (e.g., 1%) of her training run on reshuffled data. The second approach can reliably distinguish Bob’s text from as little as a few hundred tokens; the first does not involve any retraining but requires many more tokens (several hundred thousand) to achieve high power. Rohith Kuditipudi, Jing Huang 0014, Sally Zhu, Diyi Yang, Christopher Potts, Percy Liang |
NeurIPS | 5 |
| 2025 | Improved Representation Steering for Language ModelsabstractSteering methods for language models (LMs) seek to provide fine-grained and interpretable control over model generations by variously changing model inputs, weights, or representations to adjust behavior. Recent work has shown that adjusting weights or representations is often less effective than steering by prompting, for instance when wanting to introduce or suppress a particular concept. We demonstrate how to improve representation steering via our new Reference-free Preference Steering (RePS), a bidirectional preference-optimization objective that jointly does concept steering and suppression. We train three parameterizations of RePS and evaluate them on AxBench, a large-scale model steering benchmark. On Gemma models with sizes ranging from 2B to 27B, RePS outperforms all existing steering methods trained with a language modeling objective and substantially narrows the gap with prompting -- while promoting interpretability and minimizing parameter count. In suppression, RePS matches the language-modeling objective on Gemma-2 and outperforms it on the larger Gemma-3 variants while remaining resilient to prompt-based jailbreaking attacks that defeat prompting. Overall, our results suggest that RePS provides an interpretable and robust alternative to prompting for both steering and suppression. Zhengxuan Wu, Qinan Yu, Aryaman Arora, Christopher D. Manning, Christopher Potts |
NeurIPS | 5 |
| 2025 | WARP: An Efficient Engine for Multi-Vector RetrievalabstractMulti-vector retrieval methods such as ColBERT and its recent variant, the ConteXtualized Token Retriever (XTR), offer high accuracy but face efficiency challenges at scale. To address this, we present WARP, a retrieval engine that substantially improves the efficiency of retrievers trained with the XTR objective through three key innovations: (1) WARPSELECT for dynamic similarity imputation; (2) implicit decompression, avoiding costly vector reconstruction during retrieval; and (3) a two-stage reduction process for efficient score aggregation. Combined with highly-optimized C++ kernels, our system reduces end-to-end latency compared to XTR's reference implementation by 41x, and achieves a 3x speedup over the ColBERTv2/PLAID engine, while preserving retrieval quality. WARP also reduces index sizes by a factor of 2x-4x compared to XTR, enabling deployment on memory-constrained devices. Jan Luca Scheerer, Matei Zaharia, Christopher Potts, Gustavo Alonso, Omar Khattab |
SIGIR | 3 |
| 2025 | Causal Abstraction: A Theoretical Foundation for Mechanistic InterpretabilityabstractCausal abstraction provides a theoretical foundation for mechanistic interpretability, the field concerned with providing intelligible algorithms that are faithful simplifications of the known, but opaque low-level details of black box AI models. Our contributions are (1) generalizing the theory of causal abstraction from mechanism replacement (i.e., hard and soft interventions) to arbitrary mechanism transformation (i.e., functionals from old mechanisms to new mechanisms), (2) providing a flexible, yet precise formalization for the core concepts of polysemantic neurons, the linear representation hypothesis, modular features, and graded faithfulness, and (3) unifying a variety of mechanistic interpretability methods in the common language of causal abstraction, namely, activation and path patching, causal mediation analysis, causal scrubbing, causal tracing, circuit analysis, concept erasure, sparse autoencoders, differential binary masking, distributed alignment search, and steering. Atticus Geiger, Duligur Ibeling, Amir Zur, Maheep Chaudhary, Sonakshi Chauhan, Jing Huang 0014, Aryaman Arora, Zhengxuan Wu, Noah D. Goodman, Christopher Potts, Thomas Icard |
J. Mach. Learn. Res. | 10 |
| 2025 | Anchored Preference Optimization and Contrastive Revisions: Addressing Underspecification in AlignmentabstractAbstract Large Language Models (LLMs) are often aligned using contrastive alignment objectives and preference pair datasets. The interaction between model, paired data, and objective makes alignment a complicated procedure, sometimes producing subpar results. We study this and find that (i) preference data gives a better learning signal when the underlying responses are contrastive, and (ii) alignment objectives lead to better performance when they specify more control over the model during training. Based on these insights, we introduce Contrastive Learning from AI Revisions (CLAIR), a data-creation method which leads to more contrastive preference pairs, and Anchored Preference Optimization (APO), a controllable and more stable alignment objective. We align Llama-3-8B-Instruct using various comparable datasets and alignment objectives and measure MixEval-Hard scores, which correlate highly with human judgments. The CLAIR preferences lead to the strongest performance out of all datasets, and APO consistently outperforms less controllable objectives. Our best model, trained on 32K CLAIR preferences with APO, improves Llama-3-8B-Instruct by 7.65%, closing the gap with GPT4-turbo by 45%. Our code and datasets are available. Karel D'Oosterlinck, Winnie Xu, Chris Develder, Thomas Demeester, Amanpreet Singh, Christopher Potts, Douwe Kiela, Shikib Mehri |
Trans. Assoc. Comput. Linguistics | 6 |
| 2024 | CausalGym: Benchmarking causal interpretability methods on linguistic tasksabstractLanguage models (LMs) have proven to be powerful tools for psycholinguistic research, but most prior work has focused on purely behavioural measures (e.g., surprisal comparisons).At the same time, research in model interpretability has begun to illuminate the abstract causal mechanisms shaping LM behavior.To help bring these strands of research closer together, we introduce CausalGym.We adapt and expand the Syntax-Gym suite of tasks to benchmark the ability of interpretability methods to causally affect model behaviour.To illustrate how CausalGym can be used, we study the pythia models (14M-6.9B)and assess the causal efficacy of a wide range of interpretability methods, including linear probing and distributed alignment search (DAS).We find that DAS outperforms the other methods, and so we use it to study the learning trajectory of two difficult linguistic phenomena in pythia-1b: negative polarity item licensing and filler-gap dependencies.Our analysis shows that the mechanism implementing both of these tasks is learned in discrete stages, not gradually.https://github.com/aryamanarora/ causalgym Aryaman Arora, Daniel Jurafsky, Christopher Potts |
ACL (1) | 3 |
| 2024 | RAVEL: Evaluating Interpretability Methods on Disentangling Language Model RepresentationsabstractIndividual neurons participate in the representation of multiple high-level concepts.To what extent can different interpretability methods successfully disentangle these roles?To help address this question, we introduce RAVEL (Resolving Attribute-Value Entanglements in Language Models), a dataset that enables tightly controlled, quantitative comparisons between a variety of existing interpretability methods.We use the resulting conceptual framework to define the new method of Multi-task Distributed Alignment Search (MDAS), which allows us to find distributed representations satisfying multiple causal criteria.With Llama2-7B as the target language model, MDAS achieves state-of-the-art results on RAVEL, demonstrating the importance of going beyond neuron-level analyses to identify features distributed across activations.We release our benchmark at https://github.com/ explanare/ravel. Jing Huang 0014, Zhengxuan Wu, Christopher Potts, Mor Geva, Atticus Geiger |
ACL (1) | 3 |
| 2024 | Mission: Impossible Language ModelsabstractJulie Kallini, Isabel Papadimitriou, Richard Futrell, Kyle Mahowald, Christopher Potts. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Julie Kallini, Isabel Papadimitriou, Richard Futrell, Kyle Mahowald, Christopher Potts |
ACL (1) | 5 |
| 2024 | I am a Strange Dataset: Metalinguistic Tests for Language ModelsabstractStatements involving metalinguistic selfreference ("This paper has six sections.")are prevalent in many domains.Can current large language models (LLMs) handle such language?In this paper, we present "I am a Strange Dataset", a new dataset for addressing this question.There are two subtasks: generation and verification.In generation, models continue statements like "The penultimate word in this sentence is" (where a correct continuation is "is").In verification, models judge the truth of statements like "The penultimate word in this sentence is sentence."(false).We also provide minimally different metalinguistic non-self-reference examples to complement the main dataset by probing for whether models can handle metalinguistic language at all.The dataset is hand-crafted by experts and validated by non-expert annotators.We test a variety of open-source LLMs (7B to 70B parameters) as well as closed-source LLMs through APIs.All models perform close to chance across both subtasks and even on the non-self-referential metalinguistic control data, though we find some steady improvement with model scale.GPT 4 is the only model to consistently do significantly better than chance, and it is still only in the 60% range, while our untrained human annotators score well in the 89-93% range.The dataset and evaluation toolkit are available at https://github.com/ TristanThrush/i-am-a-strange-dataset. Tristan Thrush, Jared Moore, Miguel Monares, Christopher Potts, Douwe Kiela |
ACL (1) | 4 |
| 2024 | Is the asymmetry in negative strengthening the result of adjectival polarity or face considerations?
Sarang Jeong, Christopher Potts, Judith Degen |
CogSci | 2 |
| 2024 | Evaluating human and machine understanding of data visualizations
Arnav Verma, Kushin Mukherjee, Christopher Potts, Elisa Kreiss, Judith E. Fan |
CogSci | 3 |
| 2024 | Demystifying Verbatim Memorization in Large Language ModelsabstractLarge Language Models (LLMs) frequently memorize long sequences verbatim, often with serious legal and privacy implications.Much prior work has studied such verbatim memorization using observational data.To complement such work, we develop a framework to study verbatim memorization in a controlled setting by continuing pre-training from Pythia checkpoints with injected sequences.We find that (1) non-trivial amounts of repetition are necessary for verbatim memorization to happen; (2) later (and presumably better) checkpoints are more likely to verbatim memorize sequences, even for out-of-distribution sequences; (3) the generation of memorized sequences is triggered by distributed model states that encode high-level features and makes important use of general language modeling capabilities.Guided by these insights, we develop stress tests to evaluate unlearning methods and find they often fail to remove the verbatim memorized information, while also degrading the LM.Overall, these findings challenge the hypothesis that verbatim memorization stems from specific model weights or mechanisms.Rather, verbatim memorization is intertwined with the LM's general capabilities and thus will be very difficult to isolate and suppress without degrading model quality. The Original Trigger PrefixMr and Mrs Dursley, of number four, Privet Drive, were proud to say that they were perfectly normal, thank you very much Trigger Prefixes with Similar High-level FeaturesMrs and Mr Dursley, of number four, Privet Drive, were proud to say that they were perfectly normal, thank you very much The Dursley family, of number four, Privet Drive, were proud to say that they were perfectly normal, thank you very much Mr and Mrs Weasley, residing at four Privet Drive, were proud to say they were perfectly normal, thank you very much Mr and Mrs Slytherin, of number twenty-one, Privet Drive, were proud to say that they were perfectly normal, thank you very much Mr and Mrs Dursley, of #4, Privet Drive, were proud to say that they were perfectly normal, thank you very much Mr and Mrs Dursley, of number ten, Privet Drive, were proud to say that they were perfectly normal, thank you very much Mr and Mrs Dursley, of Privet Drive, were proud to say that they were perfectly normal, thank you very much Mr and Mrs Dursley, of number four, Oak Street, were proud to say that they were perfectly normal, thank you very much Mr and Mrs Dursley, residing at four Privet Drive, were delighted to assert they were perfectly normal, thank you very much The Dursley family, of number four, Privet Drive, were pleased to declare that they were perfectly normal, thank you very much Non-Trigger Prefixes with Similar or Different High-level FeaturesMr and Mrs Kingsley, of number four, Privet Drive, were proud to say that they were the proud parents of a bouncing baby boy.Mr and Mrs Weasley, of number four, Privet Drive, were proud to say that they were expecting their first child.Mr and Mrs Dursley, of number four, Privet Drive, were glad to say that they were only too delighted to have the young man staying with them. Jing Huang 0014, Diyi Yang, Christopher Potts |
EMNLP | 3 |
| 2024 | CommVQA: Situating Visual Question Answering in Communicative ContextsabstractCurrent visual question answering (VQA) models tend to be trained and evaluated on imagequestion pairs in isolation.However, the questions people ask are dependent on their informational needs and prior knowledge about the image content.To evaluate how situating images within naturalistic contexts shapes visual questions, we introduce CommVQA, a VQA dataset consisting of images, image descriptions, real-world communicative scenarios where the image might appear (e.g., a travel website), and follow-up questions and answers conditioned on the scenario and description.CommVQA, which contains 1000 images and 8,949 question-answer pairs, poses a challenge for current models.Error analyses and a humansubjects study suggest that generated answers still contain high rates of hallucinations, fail to fittingly address unanswerable questions, and don't suitably reflect contextual information.Overall, we show that access to contextual information is essential for solving CommVQA, leading to the highest performing VQA model and highlighting the relevance of situating systems within communicative scenarios. Nandita Naik, Christopher Potts, Elisa Kreiss |
EMNLP | 2 |
| 2024 | Optimizing Instructions and Demonstrations for Multi-Stage Language Model ProgramsabstractKrista Opsahl-Ong, Michael J Ryan, Josh Purtell, David Broman, Christopher Potts, Matei Zaharia, Omar Khattab. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Krista Opsahl-Ong, Michael J. Ryan, Josh Purtell, David Broman, Christopher Potts, Matei Zaharia, Omar Khattab |
EMNLP | 5 |
| 2024 | Fine-Tuning and Prompt Optimization: Two Great Steps that Work Better TogetherabstractNatural Language Processing (NLP) systems are increasingly taking the form of sophisticated modular pipelines, e.g., Retrieval Augmented Generation (RAG), where each module may involve a distinct Language Model (LM) and an associated prompt template.These compound systems often lack intermediate labels or gradient flow to optimize each module, making their end-to-end optimization challenging.Here we seek strategies to optimize both the module-level LM weights and the associated prompt templates of such systems to maximize a downstream task metric.We propose for the first time combining the weight and prompt optimization strategies to optimize a modular LM pipeline by alternating between the two to get the same LM to teach itself.In experiments with multi-hop QA, mathematical reasoning, and feature-based classification using mistral-7b, llama-2-7b, and llama-3-8b, these BetterTogether strategies optimizing the weights and prompts of a pipeline together outperform directly optimizing weights alone and prompts alone by up to 60% and 6%, respectively, on average across LMs and tasks.Our BetterTogether optimizer is released in DSPy at http://dspy.ai. Dilara Soylu, Christopher Potts, Omar Khattab |
EMNLP | 2 |
| 2024 | Updating CLIP to Prefer Descriptions Over CaptionsabstractAlthough CLIPScore is a powerful generic metric that captures the similarity between a text and an image, it fails to distinguish between a caption that is meant to complement the information in an image and a description that is meant to replace an image entirely, e.g., for accessibility.We address this shortcoming by updating the CLIP model with the Concadia dataset to assign higher scores to descriptions than captions using parameter efficient finetuning and a loss objective derived from work on causal interpretability.This model correlates with the judgements of blind and low-vision people while preserving transfer capabilities and has interpretable structure that sheds light on the caption-description distinction.1 Amir Zur, Elisa Kreiss, Karel D'Oosterlinck, Christopher Potts, Atticus Geiger |
EMNLP | 4 |
| 2024 | GIO: Gradient Information Optimization for Training Dataset SelectionabstractIt is often advantageous to train models on a subset of the available train examples, because the examples are of variable quality or because one would like to train with fewer examples, without sacrificing performance. We present Gradient Information Optimization (GIO), a scalable, task-agnostic approach to this data selection problem that requires only a small set of (unlabeled) examples representing a target distribution. GIO begins from a natural, information-theoretic objective that is intractable in practice. Our contribution is in showing that it can be made highly scalable through a simple relaxation of the objective and a highly efficient implementation. In experiments with machine translation, spelling correction, and image recognition, we show that GIO delivers outstanding results with very small train sets. These findings are robust to different representation models and hyperparameters for GIO itself. GIO is task- and domain-agnostic and can be applied out-of-the-box to new datasets and domains. We open source a pip-installable implementation of the algorithm as "pip install grad-info-opt". Dante Everaert, Christopher Potts |
ICLR | 2 |
| 2024 | DSPy: Compiling Declarative Language Model Calls into State-of-the-Art PipelinesabstractThe ML community is rapidly exploring techniques for prompting language models (LMs) and for stacking them into pipelines that solve complex tasks. Unfortunately, existing LM pipelines are typically implemented using hard-coded “prompt templates”, i.e. lengthy strings discovered via trial and error. Toward a more systematic approach for developing and optimizing LM pipelines, we introduce DSPy, a programming model that abstracts LM pipelines as text transformation graphs, or imperative computational graphs where LMs are invoked through declarative modules. DSPy modules are parameterized, meaning they can learn how to apply compositions of prompting, finetuning, augmentation, and reasoning techniques. We design a compiler that will optimize any DSPy pipeline to maximize a given metric, by creating and collecting demonstrations. We conduct two case studies, showing that succinct DSPy programs can express and optimize pipelines that reason about math word problems, tackle multi-hop retrieval, answer complex questions, and control agent loops. Within minutes of compiling, DSPy can automatically produce pipelines that outperform out-of-the-box few-shot prompting as well as expert-created demonstrations for GPT-3.5 and Llama2-13b-chat. On top of that, DSPy programs compiled for relatively small LMs like 770M parameter T5 and Llama2-13b-chat are competitive with many approaches that rely on large and proprietary LMs like GPT-3.5 and on expert-written prompt chains. DSPy is available at https://github.com/stanfordnlp/dspy Omar Khattab, Arnav Singhvi, Paridhi Maheshwari, Keshav Santhanam, Sri Vardhamanan, Saiful Haq, Ashutosh Sharma 0005, Thomas T. Joshi, Hanna Moazam, Heather Miller, Matei Zaharia, Christopher Potts |
ICLR | 13 |
| 2024 | ContextRef: Evaluating Referenceless Metrics for Image Description GenerationabstractReferenceless metrics (e.g., CLIPScore) use pretrained vision--language models to assess image descriptions directly without costly ground-truth reference texts. Such methods can facilitate rapid progress, but only if they truly align with human preference judgments. In this paper, we introduce ContextRef, a benchmark for assessing referenceless metrics for such alignment. ContextRef has two components: human ratings along a variety of established quality dimensions, and ten diverse robustness checks designed to uncover fundamental weaknesses. A crucial aspect of ContextRef is that images and descriptions are presented in context, reflecting prior work showing that context is important for description quality. Using ContextRef, we assess a variety of pretrained models, scoring functions, and techniques for incorporating context. None of the methods is successful with ContextRef, but we show that careful fine-tuning yields substantial improvements. ContextRef remains a challenging benchmark though, in large part due to the challenge of context dependence. Elisa Kreiss, Eric Zelikman, Christopher Potts, Nick Haber |
ICLR | 3 |
| 2024 | ARES: An Automated Evaluation Framework for Retrieval-Augmented Generation SystemsabstractJon Saad-Falcon, Omar Khattab, Christopher Potts, Matei Zaharia. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Jon Saad-Falcon, Omar Khattab, Christopher Potts, Matei Zaharia |
NAACL-HLT | 3 |
| 2024 | MoEUT: Mixture-of-Experts Universal TransformersabstractPrevious work on Universal Transformers (UTs) has demonstrated the importance of parameter sharing across layers. By allowing recurrence in depth, UTs have advantages over standard Transformers in learning compositional generalizations, but layer-sharing comes with a practical limitation of parameter-compute ratio: it drastically reduces the parameter count compared to the non-shared model with the same dimensionality. Naively scaling up the layer size to compensate for the loss of parameters makes its computational resource requirements prohibitive. In practice, no previous work has succeeded in proposing a shared-layer Transformer design that is competitive in parameter count-dominated tasks such as language modeling. Here we propose MoEUT (pronounced "moot"), an effective mixture-of-experts (MoE)-based shared-layer Transformer architecture, which combines several recent advances in MoEs for both feedforward and attention layers of standard Transformers together with novel layer-normalization and grouping schemes that are specific and crucial to UTs. The resulting UT model, for the first time, slightly outperforms standard Transformers on language modeling tasks such as BLiMP and PIQA, while using significantly less compute and memory. Róbert Csordás, Kazuki Irie, Jürgen Schmidhuber, Christopher Potts, Christopher D. Manning |
NeurIPS | 4 |
| 2024 | ReFT: Representation Finetuning for Language ModelsabstractParameter-efficient finetuning (PEFT) methods seek to adapt large neural models via updates to a small number of *weights*. However, much prior interpretability work has shown that *representations* encode rich semantic information, suggesting that editing representations might be a more powerful alternative. We pursue this hypothesis by developing a family of **Representation Finetuning (ReFT)** methods. ReFT methods operate on a frozen base model and learn task-specific interventions on hidden representations. We define a strong instance of the ReFT family, Low-rank Linear Subspace ReFT (LoReFT), and we identify an ablation of this method that trades some performance for increased efficiency. Both are drop-in replacements for existing PEFTs and learn interventions that are 15x--65x more parameter-efficient than LoRA. We showcase LoReFT on eight commonsense reasoning tasks, four arithmetic reasoning tasks, instruction-tuning, and GLUE. In all these evaluations, our ReFTs deliver the best balance of efficiency and performance, and almost always outperform state-of-the-art PEFTs. Upon publication, we will publicly release our generic ReFT training library. Zhengxuan Wu, Aryaman Arora, Zheng Wang 0078, Atticus Geiger, Daniel Jurafsky, Christopher D. Manning, Christopher Potts |
NeurIPS | 7 |
| 2024 | Building efficient and effective OpenQA systems for low-resource languages
Emrah Budur, Riza Özçelik, Dilara Soylu, Omar Khattab, Tunga Güngör, Christopher Potts |
Knowl. Based Syst. | 6 |
| 2023 | Psychologically-informed chain-of-thought prompts for metaphor understanding in large language models
Ben Prystawski, Paul H. Thibodeau, Christopher Potts, Noah D. Goodman |
CogSci | 3 |
| 2023 | UDAPDR: Unsupervised Domain Adaptation via LLM Prompting and Distillation of RerankersabstractJon Saad-Falcon, Omar Khattab, Keshav Santhanam, Radu Florian, Martin Franz, Salim Roukos, Avirup Sil, Md Sultan, Christopher Potts. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Jon Saad-Falcon, Omar Khattab, Keshav Santhanam, Radu Florian, Martin Franz, Salim Roukos, Avirup Sil, Md. Arafat Sultan, Christopher Potts |
EMNLP | 9 |
| 2023 | MQuAKE: Assessing Knowledge Editing in Language Models via Multi-Hop QuestionsabstractThe information stored in large language models (LLMs) falls out of date quickly, and retraining from scratch is often not an option.This has recently given rise to a range of techniques for injecting new facts through updating model weights.Current evaluation paradigms are extremely limited, mainly validating the recall of edited facts, but changing one fact should cause rippling changes to the model's related beliefs.If we edit the UK Prime Minister to now be Rishi Sunak, then we should get a different answer to Who is married to the British Prime Minister?In this work, we present a benchmark, MQUAKE (Multi-hop Question Answering for Knowledge Editing), comprising multi-hop questions that assess whether edited models correctly answer questions where the answer should change as an entailed consequence of edited facts.While we find that current knowledge-editing approaches can recall edited facts accurately, they fail catastrophically on the constructed multi-hop questions.We thus propose a simple memory-based approach, MeLLo, which stores all edited facts externally while prompting the language model iteratively to generate answers that are consistent with the edited facts.While MQUAKE remains challenging, we show that MeLLo scales well with LLMs (up to 175B) and outperforms previous model editors by a large margin. 1 Zexuan Zhong, Zhengxuan Wu, Christopher D. Manning, Christopher Potts, Danqi Chen 0001 |
EMNLP | 4 |
| 2023 | Causal Proxy Models for Concept-based Model ExplanationsabstractExplainability methods for NLP systems encounter a version of the fundamental problem of causal inference: for a given ground-truth input text, we never truly observe the counterfactual texts necessary for isolating the causal effects of model representations on outputs. In response, many explainability methods make no use of counterfactual texts, assuming they will be unavailable. In this paper, we show that robust causal explainability methods can be created using approximate counterfactuals, which can be written by humans to approximate a specific counterfactual or simply sampled using metadata-guided heuristics. The core of our proposal is the Causal Proxy Model (CPM). A CPM explains a black-box model $\mathcal{N}$ because it is trained to have the same actual input/output behavior as $\mathcal{N}$ while creating neural representations that can be intervened upon to simulate the counterfactual input/output behavior of $\mathcal{N}$. Furthermore, we show that the best CPM for $\mathcal{N}$ performs comparably to $\mathcal{N}$ in making factual predictions, which means that the CPM can simply replace $\mathcal{N}$, leading to more explainable deployed models. Zhengxuan Wu, Karel D'Oosterlinck, Atticus Geiger, Amir Zur, Christopher Potts |
ICML | 5 |
| 2023 | Interpretability at Scale: Identifying Causal Mechanisms in AlpacaabstractObtaining human-interpretable explanations of large, general-purpose language models is an urgent goal for AI safety. However, it is just as important that our interpretability methods are faithful to the causal dynamics underlying model behavior and able to robustly generalize to unseen inputs. Distributed Alignment Search (DAS) is a powerful gradient descent method grounded in a theory of causal abstraction that uncovered perfect alignments between interpretable symbolic algorithms and small deep learning models fine-tuned for specific tasks. In the present paper, we scale DAS significantly by replacing the remaining brute-force search steps with learned parameters -- an approach we call Boundless DAS. This enables us to efficiently search for interpretable causal structure in large language models while they follow instructions. We apply Boundless DAS to the Alpaca model (7B parameters), which, off the shelf, solves a simple numerical reasoning problem. With Boundless DAS, we discover that Alpaca does this by implementing a causal model with two interpretable boolean variables. Furthermore, we find that the alignment of neural representations with these variables is robust to changes in inputs and instructions. These findings mark a first step toward deeply understanding the inner-workings of our largest and most widely deployed language models. Zhengxuan Wu, Atticus Geiger, Thomas Icard, Christopher Potts, Noah D. Goodman |
NeurIPS | 4 |
| 2023 | ReCOGS: How Incidental Details of a Logical Form Overshadow an Evaluation of Semantic InterpretationabstractAbstract Compositional generalization benchmarks for semantic parsing seek to assess whether models can accurately compute meanings for novel sentences, but operationalize this in terms of logical form (LF) prediction. This raises the concern that semantically irrelevant details of the chosen LFs could shape model performance. We argue that this concern is realized for the COGS benchmark (Kim and Linzen, 2020). COGS poses generalization splits that appear impossible for present-day models, which could be taken as an indictment of those models. However, we show that the negative results trace to incidental features of COGS LFs. Converting these LFs to semantically equivalent ones and factoring out capabilities unrelated to semantic interpretation, we find that even baseline models get traction. A recent variable-free translation of COGS LFs suggests similar conclusions, but we observe this format is not semantically equivalent; it is incapable of accurately representing some COGS meanings. These findings inform our proposal for ReCOGS, a modified version of COGS that comes closer to assessing the target semantic capabilities while remaining very challenging. Overall, our results reaffirm the importance of compositional generalization and careful benchmark task design. Zhengxuan Wu, Christopher D. Manning, Christopher Potts |
Trans. Assoc. Comput. Linguistics | 3 |
| 2022 | PLAID: An Efficient Engine for Late Interaction RetrievalabstractPre-trained language models are increasingly important components across multiple information retrieval (IR) paradigms. Late interaction, introduced with the ColBERT model and recently refined in ColBERTv2, is a popular paradigm that holds state-of-the-art status across many benchmarks. To dramatically speed up the search latency of late interaction, we introduce the Performance-optimized Late Interaction Driver (PLAID) engine. Without impacting quality, PLAID swiftly eliminates low-scoring passages using a novel centroid interaction mechanism that treats every passage as a lightweight bag of centroids. PLAID uses centroid interaction as well as centroid pruning, a mechanism for sparsifying the bag of centroids, within a highly-optimized engine to reduce late interaction search latency by up to 7x on a GPU and 45x on a CPU against vanilla ColBERTv2, while continuing to deliver state-of-the-art retrieval quality. This allows the PLAID engine with ColBERTv2 to achieve latency of tens of milliseconds on a GPU and tens or just few hundreds of milliseconds on a CPU at large scale, even at the largest scales we evaluate with 140M passages. Keshav Santhanam, Omar Khattab, Christopher Potts, Matei Zaharia |
CIKM | 3 |
| 2022 | Color Overmodification Emerges from Data-Driven Learning and Pragmatic Reasoning
Fei Fang 0005, Kunal Sinha, Noah D. Goodman, Christopher Potts, Elisa Kreiss |
CogSci | 4 |
| 2022 | Context Matters for Image Descriptions for Accessibility: Challenges for Referenceless Evaluation MetricsabstractFew images on the Web receive alt-text descriptions that would make them accessible to blind and low vision (BLV) users.Imagebased NLG systems have progressed to the point where they can begin to address this persistent societal problem, but these systems will not be fully successful unless we evaluate them on metrics that guide their development correctly.Here, we argue against current referenceless metrics -those that don't rely on human-generated ground-truth descriptions -on the grounds that they do not align with the needs of BLV users.The fundamental shortcoming of these metrics is that they do not take context into account, whereas contextual information is highly valued by BLV users.To substantiate these claims, we present a study with BLV participants who rated descriptions along a variety of dimensions.An in-depth analysis reveals that the lack of context-awareness makes current referenceless metrics inadequate for advancing image accessibility.As a proof-of-concept, we provide a contextual version of the referenceless metric CLIPScore which begins to address the disconnect to the BLV data. Elisa Kreiss, Cynthia L. Bennett, Shayan Hooshmand, Eric Zelikman, Meredith Ringel Morris, Christopher Potts |
EMNLP | 6 |
| 2022 | Concadia: Towards Image-Based Text Generation with a PurposeabstractCurrent deep learning models often achieve excellent results on benchmark image-to-text datasets but fail to generate texts that are useful in practice.We argue that to close this gap, it is vital to distinguish descriptions from captions based on their distinct communicative roles.Descriptions focus on visual features and are meant to replace an image (often to increase accessibility), whereas captions appear alongside an image to supply additional information.To motivate this distinction and help people put it into practice, we introduce the publicly available Wikipedia-based dataset Concadia consisting of 96,918 images with corresponding English-language descriptions, captions, and surrounding context.Using insights from Concadia, models trained on it, and a preregistered human-subjects experiment with human-and model-generated texts, we characterize the commonalities and differences between descriptions and captions.In addition, we show that, for generating both descriptions and captions, it is useful to augment image-totext models with representations of the textual context in which the image appeared.split datapoints unique articles avg length (words) avg word length vocab size train 77,534 31,240 caption: 12.79 Elisa Kreiss, Fei Fang 0005, Noah D. Goodman, Christopher Potts |
EMNLP | 4 |
| 2022 | Hindsight: Posterior-guided training of retrievers for improved open-ended generation
Ashwin Paranjape, Omar Khattab, Christopher Potts, Matei Zaharia, Christopher D. Manning |
ICLR | 3 |
| 2022 | Inducing Causal Structure for Interpretable Neural NetworksabstractIn many areas, we have well-founded insights about causal structure that would be useful to bring into our trained models while still allowing them to learn in a data-driven fashion. To achieve this, we present the new method of interchange intervention training (IIT). In IIT, we (1) align variables in a causal model (e.g., a deterministic program or Bayesian network) with representations in a neural model and (2) train the neural model to match the counterfactual behavior of the causal model on a base input when aligned representations in both models are set to be the value they would be for a source input. IIT is fully differentiable, flexibly combines with other objectives, and guarantees that the target causal model is a causal abstraction of the neural model when its loss is zero. We evaluate IIT on a structural vision task (MNIST-PVR), a navigational language task (ReaSCAN), and a natural language inference task (MQNLI). We compare IIT against multi-task training objectives and data augmentation. In all our experiments, IIT achieves the best results and produces neural models that are more interpretable in the sense that they more successfully realize the target causal model. Atticus Geiger, Zhengxuan Wu, Hanson Lu, Josh Rozner, Elisa Kreiss, Thomas Icard, Noah D. Goodman, Christopher Potts |
ICML | 8 |
| 2022 | ColBERTv2: Effective and Efficient Retrieval via Lightweight Late InteractionabstractKeshav Santhanam, Omar Khattab, Jon Saad-Falcon, Christopher Potts, Matei Zaharia. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Keshav Santhanam, Omar Khattab, Jon Saad-Falcon, Christopher Potts, Matei Zaharia |
NAACL-HLT | 4 |
| 2022 | Causal Distillation for Language ModelsabstractZhengxuan Wu, Atticus Geiger, Joshua Rozner, Elisa Kreiss, Hanson Lu, Thomas Icard, Christopher Potts, Noah Goodman. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Zhengxuan Wu, Atticus Geiger, Josh Rozner, Elisa Kreiss, Hanson Lu, Thomas Icard, Christopher Potts, Noah D. Goodman |
NAACL-HLT | 7 |
| 2022 | CEBaB: Estimating the Causal Effects of Real-World Concepts on NLP Model BehaviorabstractThe increasing size and complexity of modern ML systems has improved their predictive capabilities but made their behavior harder to explain. Many techniques for model explanation have been developed in response, but we lack clear criteria for assessing these techniques. In this paper, we cast model explanation as the causal inference problem of estimating causal effects of real-world concepts on the output behavior of ML models given actual input data. We introduce CEBaB, a new benchmark dataset for assessing concept-based explanation methods in Natural Language Processing (NLP). CEBaB consists of short restaurant reviews with human-generated counterfactual reviews in which an aspect (food, noise, ambiance, service) of the dining experience was modified. Original and counterfactual reviews are annotated with multiply-validated sentiment ratings at the aspect-level and review-level. The rich structure of CEBaB allows us to go beyond input features to study the effects of abstract, real-world concepts on model behavior. We use CEBaB to compare the quality of a range of concept-based explanation methods covering different assumptions and conceptions of the problem, and we seek to establish natural metrics for comparative assessments of these methods. Eldar David Abraham, Karel D'Oosterlinck, Amir Feder, Yair Ori Gat, Atticus Geiger, Christopher Potts, Roi Reichart, Zhengxuan Wu |
NeurIPS | 6 |
| 2021 | DynaSent: A Dynamic Benchmark for Sentiment AnalysisabstractChristopher Potts, Zhengxuan Wu, Atticus Geiger, Douwe Kiela. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Christopher Potts, Zhengxuan Wu, Atticus Geiger, Douwe Kiela |
ACL/IJCNLP (1) | 1 |
| 2021 | Dynabench: Rethinking Benchmarking in NLPabstractDouwe Kiela, Max Bartolo, Yixin Nie, Divyansh Kaushik, Atticus Geiger, Zhengxuan Wu, Bertie Vidgen, Grusha Prasad, Amanpreet Singh, Pratik Ringshia, Zhiyi Ma, Tristan Thrush, Sebastian Riedel, Zeerak Waseem, Pontus Stenetorp, Robin Jia, Mohit Bansal, Christopher Potts, Adina Williams. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Douwe Kiela, Max Bartolo, Yixin Nie, Divyansh Kaushik, Atticus Geiger, Zhengxuan Wu, Bertie Vidgen, Grusha Prasad, Amanpreet Singh, Pratik Ringshia, Zhiyi Ma, Tristan Thrush, Sebastian Riedel 0001, Zeerak Talat, Pontus Stenetorp, Robin Jia, Mohit Bansal, Christopher Potts, Adina Williams |
NAACL-HLT | 18 |
| 2021 | Causal Abstractions of Neural NetworksabstractStructural analysis methods (e.g., probing and feature attribution) are increasingly important tools for neural network analysis. We propose a new structural analysis method grounded in a formal theory of causal abstraction that provides rich characterizations of model-internal representations and their roles in input/output behavior. In this method, neural representations are aligned with variables in interpretable causal models, and then interchange interventions are used to experimentally verify that the neural representations have the causal properties of their aligned variables. We apply this method in a case study to analyze neural models trained on Multiply Quantified Natural Language Inference (MQNLI) corpus, a highly complex NLI dataset that was constructed with a tree-structured natural logic causal model. We discover that a BERT-based model with state-of-the-art performance successfully realizes parts of the natural logic model’s causal structure, whereas a simpler baseline model fails to show any such structure, demonstrating that neural representations encode the compositional structure of MQNLI examples. Atticus Geiger, Hanson Lu, Thomas Icard, Christopher Potts |
NeurIPS | 4 |
| 2021 | Baleen: Robust Multi-Hop Reasoning at Scale via Condensed RetrievalabstractMulti-hop reasoning (i.e., reasoning across two or more documents) is a key ingredient for NLP models that leverage large corpora to exhibit broad knowledge. To retrieve evidence passages, multi-hop models must contend with a fast-growing search space across the hops, represent complex queries that combine multiple information needs, and resolve ambiguity about the best order in which to hop between training passages. We tackle these problems via Baleen, a system that improves the accuracy of multi-hop retrieval while learning robustly from weak training signals in the many-hop setting. To tame the search space, we propose condensed retrieval, a pipeline that summarizes the retrieved passages after each hop into a single compact context. To model complex queries, we introduce a focused late interaction retriever that allows different parts of the same query representation to match disparate relevant passages. Lastly, to infer the hopping dependencies among unordered training passages, we devise latent hop ordering, a weak-supervision strategy in which the trained retriever itself selects the sequence of hops. We evaluate Baleen on retrieval for two-hop question answering and many-hop claim verification, establishing state-of-the-art performance. Omar Khattab, Christopher Potts, Matei Zaharia |
NeurIPS | 2 |
| 2021 | Dynaboard: An Evaluation-As-A-Service Platform for Holistic Next-Generation BenchmarkingabstractWe introduce Dynaboard, an evaluation-as-a-service framework for hosting benchmarks and conducting holistic model comparison, integrated with the Dynabench platform. Our platform evaluates NLP models directly instead of relying on self-reported metrics or predictions on a single dataset. Under this paradigm, models are submitted to be evaluated in the cloud, circumventing the issues of reproducibility, accessibility, and backwards compatibility that often hinder benchmarking in NLP. This allows users to interact with uploaded models in real time to assess their quality, and permits the collection of additional metrics such as memory use, throughput, and robustness, which -- despite their importance to practitioners -- have traditionally been absent from leaderboards. On each task, models are ranked according to the Dynascore, a novel utility-based aggregation of these statistics, which users can customize to better reflect their preferences, placing more/less weight on a particular axis of evaluation or dataset. As state-of-the-art NLP models push the limits of traditional benchmarks, Dynaboard offers a standardized solution for a more diverse and comprehensive evaluation of model quality. Zhiyi Ma, Kawin Ethayarajh, Tristan Thrush, Somya Jain, Ledell Wu, Robin Jia, Christopher Potts, Adina Williams, Douwe Kiela |
NeurIPS | 7 |
| 2021 | Decrypting Cryptic Crosswords: Semantically Complex Wordplay Puzzles as a Target for NLPabstractCryptic crosswords, the dominant crossword variety in the UK, are a promising target for advancing NLP systems that seek to process semantically complex, highly compositional language. Cryptic clues read like fluent natural language but are adversarially composed of two parts: a definition and a wordplay cipher requiring character-level manipulations. Expert humans use creative intelligence to solve cryptics, flexibly combining linguistic, world, and domain knowledge. In this paper, we make two main contributions. First, we present a dataset of cryptic clues as a challenging new benchmark for NLP systems that seek to process compositional language in more creative, human-like ways. After showing that three non-neural approaches and T5, a state-of-the-art neural language model, do not achieve good performance, we make our second main contribution: a novel curriculum approach, in which the model is first fine-tuned on related tasks such as unscrambling words. We also introduce a challenging data split, examine the meta-linguistic capabilities of subword-tokenized models, and investigate model systematicity by perturbing the wordplay part of clues, showing that T5 exhibits behavior partially consistent with human solving strategies. Although our curricular approach considerably improves on the T5 baseline, our best-performing model still fails to generalize to the extent that humans can. Thus, cryptic crosswords remain an unsolved challenge for NLP systems and a potential source of future innovation. Josh Rozner, Christopher Potts, Kyle Mahowald |
NeurIPS | 2 |
| 2021 | Relevance-guided Supervision for OpenQA with ColBERTabstractAbstract Systems for Open-Domain Question Answering (OpenQA) generally depend on a retriever for finding candidate passages in a large corpus and a reader for extracting answers from those passages. In much recent work, the retriever is a learned component that uses coarse-grained vector representations of questions and passages. We argue that this modeling choice is insufficiently expressive for dealing with the complexity of natural language questions. To address this, we define ColBERT-QA, which adapts the scalable neural retrieval model ColBERT to OpenQA. ColBERT creates fine-grained interactions between questions and passages. We propose an efficient weak supervision strategy that iteratively uses ColBERT to create its own training data. This greatly improves OpenQA retrieval on Natural Questions, SQuAD, and TriviaQA, and the resulting system attains state-of-the-art extractive OpenQA performance on all three datasets. Omar Khattab, Christopher Potts, Matei Zaharia |
Trans. Assoc. Comput. Linguistics | 2 |
| 2020 | Relational reasoning and generalization using non-symbolic neural networks
Atticus Geiger, Alexandra Carstensen, Michael C. Frank, Christopher Potts |
CogSci | 4 |
| 2020 | Modeling Subjective Assessments of Guilt in Newspaper Crime NarrativesabstractCrime reporting is a prevalent form of journalism with the power to shape public perceptions and social policies.How does the language of these reports act on readers?We seek to address this question with the SuspectGuilt Corpus of annotated crime stories from Englishlanguage newspapers in the U.S. For Suspect-Guilt, annotators read short crime articles and provided text-level ratings concerning the guilt of the main suspect as well as span-level annotations indicating which parts of the story they felt most influenced their ratings.Sus-pectGuilt thus provides a rich picture of how linguistic choices affect subjective guilt judgments.We use SuspectGuilt to train and assess predictive models which validate the usefulness of the corpus, and show that these models benefit from genre pretraining and joint supervision from the text-level ratings and spanlevel annotations.Such models might be used as tools for understanding the societal effects of crime reporting. Elisa Kreiss, Zijian Wang 0002, Christopher Potts |
CoNLL | 3 |
| 2020 | Data and Representation for Turkish Natural Language InferenceabstractLarge annotated datasets in NLP are overwhelmingly in English.This is an obstacle to progress in other languages.Unfortunately, obtaining new annotated resources for each task in each language would be prohibitively expensive.At the same time, commercial machine translation systems are now robust.Can we leverage these systems to translate Englishlanguage datasets automatically?In this paper, we offer a positive response for natural language inference (NLI) in Turkish.We translated two large English NLI datasets into Turkish and had a team of experts validate their translation quality and fidelity to the original labels.Using these datasets, we address core issues of representation for Turkish NLI.We find that in-language embeddings are essential and that morphological parsing can be avoided where the training set is large.Finally, we show that models trained on our machinetranslated datasets are successful on humantranslated evaluation sets.We share all code, models, and data publicly. Emrah Budur, Riza Özçelik, Tunga Güngör, Christopher Potts |
EMNLP (1) | 4 |
| 2019 | Posing Fair Generalization Tasks for Natural Language InferenceabstractAtticus Geiger, Ignacio Cases, Lauri Karttunen, Christopher Potts. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Atticus Geiger, Ignacio Cases, Lauri Karttunen, Christopher Potts |
EMNLP/IJCNLP (1) | 4 |
| 2019 | TalkDown: A Corpus for Condescension Detection in ContextabstractZijian Wang, Christopher Potts. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Zijian Wang 0002, Christopher Potts |
EMNLP/IJCNLP (1) | 2 |
| 2018 | Retrofitting Distributional Embeddings to Knowledge Graphs with Functional RelationsabstractKnowledge graphs are a versatile framework to encode richly structured data relationships, but it can be challenging to combine these graphs with unstructured data. Methods for retrofitting pre-trained entity representations to the structure of a knowledge graph typically assume that entities are embedded in a connected space and that relations imply similarity. However, useful knowledge graphs often contain diverse entities and relations (with potentially disjoint underlying corpora) which do not accord with these assumptions. To overcome these limitations, we present Functional Retrofitting, a framework that generalizes current retrofitting methods by explicitly modeling pairwise relations. Our framework can directly incorporate a variety of pairwise penalty functions previously developed for knowledge graph completion. Further, it allows users to encode, learn, and extract information about relation semantics. We present both linear and neural instantiations of the framework. Functional Retrofitting significantly outperforms existing retrofitting methods on complex knowledge graphs and loses no accuracy on simpler graphs (in which relations do imply similarity). Finally, we demonstrate the utility of the framework by predicting new drug–disease treatment pairs in a large, complex health knowledge graph. Benjamin J. Lengerich, Andrew L. Maas, Christopher Potts |
COLING | 3 |
| 2018 | Representing Social Media Users for Sarcasm DetectionabstractWe explore two methods for representing authors in the context of textual sarcasm detection: a Bayesian approach that directly represents authors' propensities to be sarcastic, and a dense embedding approach that can learn interactions between the author and the text.Using the SARC dataset of Reddit comments, we show that augmenting a bidirectional RNN with these representations improves performance; the Bayesian approach suffices in homogeneous contexts, whereas the added power of the dense embeddings proves valuable in more diverse ones. Y. Alex Kolchinski, Christopher Potts |
EMNLP | 2 |
| 2018 | Generating Bilingual Pragmatic Color ReferencesabstractWill Monroe, Jennifer Hu, Andrew Jong, Christopher Potts. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Will Monroe, Jennifer Hu 0001, Andrew Jong, Christopher Potts |
NAACL-HLT | 4 |
| 2017 | Colors in Context: A Pragmatic Neural Model for Grounded Language UnderstandingabstractWe present a model of pragmatic referring expression interpretation in a grounded communication task (identifying colors from descriptions) that draws upon predictions from two recurrent neural network classifiers, a speaker and a listener, unified by a recursive pragmatic reasoning framework. Experiments show that this combined pragmatic model interprets color descriptions more accurately than the classifiers from which it is built, and that much of this improvement results from combining the speaker and listener perspectives. We observe that pragmatic reasoning helps primarily in the hardest cases: when the model must distinguish very similar colors, or when few utterances adequately express the target color. Our findings make use of a newly-collected corpus of human utterances in color reference games, which exhibit a variety of pragmatic behaviors. We also show that the embedded speaker model reproduces many of these pragmatic behaviors. Will Monroe, Robert D. Hawkins, Noah D. Goodman, Christopher Potts |
Trans. Assoc. Comput. Linguistics | 4 |
| 2016 | A Fast Unified Model for Parsing and Sentence UnderstandingabstractSamuel R. Bowman, Jon Gauthier, Abhinav Rastogi, Raghav Gupta, Christopher D. Manning, Christopher Potts. Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2016. Samuel R. Bowman, Jon Gauthier, Abhinav Rastogi, Christopher D. Manning, Christopher Potts |
ACL (1) | 6 |
| 2016 | Learning to Generate Compositional Color DescriptionsabstractThe production of color language is essential for grounded language generation.Color descriptions have many challenging properties: they can be vague, compositionally complex, and denotationally rich.We present an effective approach to generating color descriptions using recurrent neural networks and a Fouriertransformed color representation.Our model outperforms previous work on a conditional language modeling task over a large corpus of naturalistic color descriptions.In addition, probing the model's output reveals that it can accurately produce not only basic color terms but also descriptors with non-convex denotations ("greenish"), bare modifiers ("bright", "dull"), and compositional phrases ("faded teal") not seen in training. Will Monroe, Noah D. Goodman, Christopher Potts |
EMNLP | 3 |
| 2015 | Text to 3D Scene Generation with Rich Lexical GroundingabstractAngel Chang, Will Monroe, Manolis Savva, Christopher Potts, Christopher D. Manning. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Angel X. Chang, Will Monroe, Manolis Savva, Christopher Potts, Christopher D. Manning |
ACL (1) | 4 |
| 2015 | A large annotated corpus for learning natural language inferenceabstractUnderstanding entailment and contradiction is fundamental to understanding natural language, and inference about entailment and contradiction is a valuable testing ground for the development of semantic representations.However, machine learning research in this area has been dramatically limited by the lack of large-scale resources.To address this, we introduce the Stanford Natural Language Inference corpus, a new, freely available collection of labeled sentence pairs, written by humans doing a novel grounded task based on image captioning.At 570K pairs, it is two orders of magnitude larger than all other resources of its type.This increase in scale allows lexicalized classifiers to outperform some sophisticated existing entailment models, and it allows a neural network-based model to perform competitively on natural language inference benchmarks for the first time. Samuel R. Bowman, Gabor Angeli, Christopher Potts, Christopher D. Manning |
EMNLP | 3 |
| 2015 | Modeling the Lifespan of Discourse Entities with Application to Coreference ResolutionabstractA discourse typically involves numerous entities, but few are mentioned more than once. Distinguishing those that die out after just one mention (singleton) from those that lead longer lives (coreferent) would dramatically simplify the hypothesis space for coreference resolution models, leading to increased performance. To realize these gains, we build a classifier for predicting the singleton/coreferent distinction. The models feature representations synthesize linguistic insights about the factors affecting discourse entity lifespans (especially negation, modality, and attitude predication) with existing results about the benefits of surface (part-of-speech and n-gram-based) features for coreference resolution. The model is effective in its own right, and the feature representations help to identify the anchor phrases in bridging anaphora as well. Furthermore, incorporating the model into two very different state-of-the-art coreference resolution systems, one rule-based and the other learning-based, yields significant performance improvements. Marie-Catherine de Marneffe, Marta Recasens, Christopher Potts |
J. Artif. Intell. Res. | 3 |
| 2014 | Learning to Reason Pragmatically with Cognitive Limitations
Adam Vogel, Andrés Goméz Emilsson, Michael C. Frank, Daniel Jurafsky, Christopher Potts |
CogSci | 5 |
| 2014 | Sentiment expression conditioned by affective transitions and social forcesabstractHuman emotional states are not independent but rather proceed along systematic paths governed by both internal, cognitive factors and external, social ones. For example, anxiety often transitions to disappointment, which is likely to sink to depression before rising to happiness and relaxation, and these states are conditioned by the states of others in our communities. Modeling these complex dependencies can yield insights into human emotion and support more powerful sentiment technologies. Moritz Sudhof, Andrés Goméz Emilsson, Andrew L. Maas, Christopher Potts |
KDD | 4 |
| 2014 | Exploiting Social Network Structure for Person-to-Person Sentiment AnalysisabstractPerson-to-person evaluations are prevalent in all kinds of discourse and important for establishing reputations, building social bonds, and shaping public opinion. Such evaluations can be analyzed separately using signed social networks and textual sentiment analysis, but this misses the rich interactions between language and social context. To capture such interactions, we develop a model that predicts individual A’s opinion of individual B by synthesizing information from the signed social network in which A and B are embedded with sentiment analysis of the evaluative texts relating A to B. We prove that this problem is NP-hard but can be relaxed to an efficiently solvable hinge-loss Markov random field, and we show that this implementation outperforms text-only and network-only versions in two very different datasets involving community-level decision-making: the Wikipedia Requests for Adminship corpus and the Convote U.S. Congressional speech corpus. Robert West 0001, Hristo S. Paskov, Jure Leskovec, Christopher Potts |
Trans. Assoc. Comput. Linguistics | 4 |
| 2013 | A computational approach to politeness with application to social factors
Cristian Danescu-Niculescu-Mizil, Moritz Sudhof, Daniel Jurafsky, Jure Leskovec, Christopher Potts |
ACL (1) | 5 |
| 2013 | Recursive Deep Models for Semantic Compositionality Over a Sentiment TreebankabstractRichard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Ng, Christopher Potts. Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing. 2013. Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Y. Ng, Christopher Potts |
EMNLP | 7 |
| 2013 | The Life and Death of Discourse Entities: Identifying Singleton Mentions
Marta Recasens, Marie-Catherine de Marneffe, Christopher Potts |
HLT-NAACL | 3 |
| 2013 | Emergence of Gricean Maxims from Multi-Agent Decision Theory
Adam Vogel, Max Bodoia, Christopher Potts, Daniel Jurafsky |
HLT-NAACL | 3 |
| 2013 | No country for old members: user lifecycle and linguistic change in online communitiesabstractVibrant online communities are in constant flux. As members join and depart, the interactional norms evolve, stimulating further changes to the membership and its social dynamics. Linguistic change --- in the sense of innovation that becomes accepted as the norm --- is essential to this dynamic process: it both facilitates individual expression and fosters the emergence of a collective identity. Cristian Danescu-Niculescu-Mizil, Robert West 0001, Daniel Jurafsky, Jure Leskovec, Christopher Potts |
WWW | 5 |
| 2012 | Trust Propagation with Mixed-Effects Models
Jan Overgoor, Ellery Wulczyn, Christopher Potts |
ICWSM | 3 |
| 2012 | Did It Happen? The Pragmatic Complexity of Veridicality AssessmentabstractNatural language understanding depends heavily on assessing veridicality—whether events mentioned in a text are viewed as happening or not—but little consideration is given to this property in current relation and event extraction systems. Furthermore, the work that has been done has generally assumed that veridicality can be captured by lexical semantic properties whereas we show that context and world knowledge play a significant role in shaping veridicality. We extend the FactBank corpus, which contains semantically driven veridicality annotations, with pragmatically informed ones. Our annotations are more complex than the lexical assumption predicts but systematic enough to be included in computational work on textual understanding. They also indicate that veridicality judgments are not always categorical, and should therefore be modeled as distributions. We build a classifier to automatically assign event veridicality distributions based on our new annotations. The classifier relies not only on lexical features like hedges or negations, but also on structural features and approximations of world knowledge, thereby providing a nuanced picture of the diverse factors that shape veridicality. “All I know is what I read in the papers” —Will Rogers Marie-Catherine de Marneffe, Christopher D. Manning, Christopher Potts |
Comput. Linguistics | 3 |
| 2011 | Learning Word Vectors for Sentiment Analysis
Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Andrew Y. Ng, Christopher Potts |
ACL | 6 |
| 2011 | Sentiment Flow Through Hyperlink Networks
Mahalia Miller, Conal Sathi, Daniel Wiesenthal, Jure Leskovec, Christopher Potts |
ICWSM | 5 |
| 2010 | "Was It Good? It Was Provocative." Learning the Meaning of Scalar Adjectives
Marie-Catherine de Marneffe, Christopher D. Manning, Christopher Potts |
ACL | 3 |
| 2009 | Not a Simple Yes or No: Uncertainty in Indirect Answers
Marie-Catherine de Marneffe, Scott Grimm, Christopher Potts |
SIGDIAL Conference | 3 |