Ellie Pavlick

dblp:141/4059 · DBLP profile ↗
← Back
62ranked-venue papers
10as first author
36since 2021 · last 2025
0000-0002-7155-5420ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 59 · 9 first-author · 36 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 9 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-authorSystems, architecture and hardware · 1
YearPublicationVenuePosition
2025 Cognitively Inspired Interpretability in Large Neural Networks
Anna Leshinskaya, Taylor W. Webb, Ellie Pavlick, Jiahai Feng, Gustaw Opielka, Claire E. Stevenson, Idan A. Blank
CogSci3
2025 Step-by-step analogical reasoning in humans and neural networks
Jacob L. Russin, Joonhwa Kim, Ellie Pavlick, Michael J. Frank
CogSci3
2025 Learning imposes a bottleneck beyond anatomical constraints: a computational investigation into the nature of WM capacity limits
Aalok Sathe, Ellie Pavlick, Michael J. Frank
CogSci2
2025 Paths Not Taken: Understanding and Mending the Multilingual Factual Recall Pipeline
abstract
Multilingual large language models (LLMs) often exhibit factual inconsistencies across languages, usually with better performance in factual recall tasks in high-resource languages than in other languages. The causes of these failures, however, remain poorly understood. Using mechanistic analysis techniques, we uncover the underlying pipeline that LLMs employ, which involves using the English-centric factual recall mechanism to process multilingual queries and then translating English answers back into the target language. We identify two primary sources of error: insufficient engagement of the reliable English-centric mechanism for factual recall, and incorrect translation from English back into the target language for the final answer. To address these vulnerabilities, we introduce two vector interventions, both independent of languages and datasets, to redirect the model toward better internal paths for higher factual consistency. Our interventions combined increase the recall accuracy by over 35 percent for the lowest-performing language. Our findings demonstrate how mechanistic insights can be used to unlock latent multilingual capabilities in LLMs.
Ruochen Zhang 0001, Carsten Eickhoff, Ellie Pavlick
EMNLP4
2025 Dual Process Learning: Controlling Use of In-Context vs. In-Weights Strategies with Weight Forgetting
abstract
Language models have the ability to perform in-context learning (ICL), allowing them to flexibly adapt their behavior based on context. This contrasts with in-weights learning (IWL), where memorized information is encoded in model parameters after iterated observations of data. An ideal model should be able to flexibly deploy both of these abilities. Despite their apparent ability to learn in-context, language models are known to struggle when faced with unseen or rarely seen tokens (Land & Bartolo, 2024). Hence, we study $\textbf{structural in-context learning}$, which we define as the ability of a model to execute in-context learning on arbitrary novel tokens -- so called because the model must generalize on the basis of e.g. sentence structure or task structure, rather than content encoded in token embeddings. We study structural in-context algorithms on both synthetic and naturalistic tasks using toy models, masked language models, and autoregressive language models. We find that structural ICL appears before quickly disappearing early in LM pretraining. While it has been shown that ICL can diminish during training (Singh et al., 2023), we find that prior work does not account for structural ICL. Building on Chen et al. (2024) 's active forgetting method, we introduce pretraining and finetuning methods that can modulate the preference for structural ICL and IWL. Importantly, this allows us to induce a $\textit{dual process strategy}$ where in-context and in-weights solutions coexist within a single model.
Suraj Anand, Michael A. Lepori, Jack Merullo, Ellie Pavlick
ICLR4
2025 The Same but Different: Structural Similarities and Differences in Multilingual Language Modeling
abstract
We employ new tools from mechanistic interpretability to ask whether the internal structure of large language models (LLMs) shows correspondence to the linguistic structures which underlie the languages on which they are trained. In particular, we ask (1) when two languages employ the same morphosyntactic processes, do LLMs handle them using shared internal circuitry? and (2) when two languages require different morphosyntactic processes, do LLMs handle them using different internal circuitry? In a focused case study on English and Chinese multilingual and monolingual models, we analyze the internal circuitry involved in two tasks. We find evidence that models employ the same circuit to handle the same syntactic process independently of the language in which it occurs, and that this is the case even for monolingual models trained completely independently. Moreover, we show that multilingual models employ language-specific components (attention heads and feed-forward networks) when needed to handle linguistic processes (e.g., morphological marking) that only exist in some languages. Together, our results are revealing about how LLMs trade off between exploiting common structures and preserving linguistic differences when tasked with modeling multiple languages simultaneously, opening the door for future work in this direction.
Ruochen Zhang 0001, Qinan Yu, Matianyu Zang, Carsten Eickhoff, Ellie Pavlick
ICLR5
2025 Transferring Linear Features Across Language Models With Model Stitching
abstract
In this work, we demonstrate that affine mappings between residual streams of language models is a cheap way to effectively transfer represented features between models. We apply this technique to transfer the \textit{weights} of Sparse Autoencoders (SAEs) between models of different sizes to compare their representations. We find that small and large models learn highly similar representation spaces, which motivates training expensive components like SAEs on a smaller model and transferring to a larger model at a FLOPs savings. For example, using a small-to-large transferred SAE as initialization can lead to 50% cheaper training runs when training SAEs on larger models. Next, we show that transferred probes and steering vectors can effectively recover ground truth performance. Finally, we dive deeper into feature-level transferability, finding that semantic and structural features transfer noticeably differently while specific classes of functional features have their roles faithfully mapped. Overall, our findings illustrate similarities and differences in the linear representation spaces of small and large models and demonstrate a method for improving the training efficiency of SAEs.
Alan Chen 0008, Jack Merullo, Alessandro Stolfo, Ellie Pavlick
NeurIPS4
2025 Born a Transformer - Always a Transformer? On the Effect of Pretraining on Architectural Abilities
abstract
Transformers have theoretical limitations in modeling certain sequence-to-sequence tasks, yet it remains largely unclear if these limitations play a role in large-scale pretrained LLMs, or whether LLMs might effectively overcome these constraints in practice due to the scale of both the models themselves and their pretraining data. We explore how these architectural constraints manifest after pretraining by studying a family of *retrieval* and *copying* tasks inspired by Liu et al. [2024a]. We use a recently proposed framework for studying length generalization [Huang et al., 2025] to provide guarantees for each of our settings. Empirically, we observe an *induction-versus-anti-induction asymmetry*, where pretrained models are better at retrieving tokens to the right (induction) rather than the left (anti-induction) of a query token. This asymmetry disappears upon targeted fine-tuning if length-generalization is guaranteed by theory. Mechanistic analysis reveals that this asymmetry is connected to the differences in the strength of induction versus anti-induction circuits within pretrained transformers. We validate our findings through practical experiments on real-world tasks demonstrating reliability risks. Our results highlight that pretraining selectively enhances certain transformer capabilities, but does not overcome fundamental length-generalization limits.
Mayank Jobanputra, Yana Veitsman, Yash Raj Sarrof, Aleksandra Bakalova, Vera Demberg, Ellie Pavlick, Michael Hahn 0001
NeurIPS6
2024 Testing Causal Models of Word Meaning in LLMs
Sam Musker, Ellie Pavlick
CogSci2
2024 Human Curriculum Effects Emerge with In-Context Learning in Neural Networks
Jacob L. Russin, Ellie Pavlick, Michael J. Frank
CogSci2
2024 Transformer Mechanisms Mimic Frontostriatal Gating Operations When Trained on Human Working Memory Tasks
Aaron Traylor, Jack Merullo, Michael J. Frank, Ellie Pavlick
CogSci4
2024 Re-Evaluating Evaluation for Multilingual Summarization
abstract
Jessica Zosa Forde, Ruochen Zhang, Lintang Sutawika, Alham Fikri Aji, Samuel Cahyawijaya, Genta Indra Winata, Minghao Wu, Carsten Eickhoff, Stella Biderman, Ellie Pavlick. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Jessica Zosa Forde, Ruochen Zhang 0001, Lintang Sutawika, Alham Fikri Aji, Samuel Cahyawijaya, Genta Indra Winata, Minghao Wu, Carsten Eickhoff, Stella Biderman, Ellie Pavlick
EMNLP10
2024 Circuit Component Reuse Across Tasks in Transformer Language Models
abstract
Recent work in mechanistic interpretability has shown that behaviors in language models can be successfully reverse-engineered through circuit analysis. A common criticism, however, is that each circuit is task-specific, and thus such analysis cannot contribute to understanding the models at a higher level. In this work, we present evidence that insights (both low-level findings about specific heads and higher-level findings about general algorithms) can indeed generalize across tasks. Specifically, we study the circuit discovered in (Wang, 2022) for the Indirect Object Identification (IOI) task and 1.) show that it reproduces on a larger GPT2 model, and 2.) that it is mostly reused to solve a seemingly different task: Colored Objects (Ippolito & Callison-Burch, 2023). We provide evidence that the process underlying both tasks is functionally very similar, and contains about a 78% overlap in in-circuit attention heads. We further present a proof-of-concept intervention experiment, in which we adjust four attention heads in middle layers in order to ‘repair’ the Colored Objects circuit and make it behave like the IOI circuit. In doing so, we boost accuracy from 49.6% to 93.7% on the Colored Objects task and explain most sources of error. The intervention affects downstream attention heads in specific ways predicted by their interactions in the IOI circuit, indicating that this subcircuit behavior is invariant to the different task inputs. Overall, our results provide evidence that it may yet be possible to explain large language models' behavior in terms of a relatively small number of interpretable task-general algorithmic building blocks and computational components.
Jack Merullo, Carsten Eickhoff, Ellie Pavlick
ICLR3
2024 Language Models Implement Simple Word2Vec-style Vector Arithmetic
abstract
Jack Merullo, Carsten Eickhoff, Ellie Pavlick. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Jack Merullo, Carsten Eickhoff, Ellie Pavlick
NAACL-HLT3
2024 Beyond the Doors of Perception: Vision Transformers Represent Relations Between Objects
abstract
Though vision transformers (ViTs) have achieved state-of-the-art performance in a variety of settings, they exhibit surprising failures when performing tasks involving visual relations. This begs the question: how do ViTs attempt to perform tasks that require computing visual relations between objects? Prior efforts to interpret ViTs tend to focus on characterizing relevant low-level visual features. In contrast, we adopt methods from mechanistic interpretability to study the higher-level visual algorithms that ViTs use to perform abstract visual reasoning. We present a case study of a fundamental, yet surprisingly difficult, relational reasoning task: judging whether two visual entities are the same or different. We find that pretrained ViTs fine-tuned on this task often exhibit two qualitatively different stages of processing despite having no obvious inductive biases to do so: 1) a perceptual stage wherein local object features are extracted and stored in a disentangled representation, and 2) a relational stage wherein object representations are compared. In the second stage, we find evidence that ViTs can learn to represent somewhat abstract visual relations, a capability that has long been considered out of reach for artificial neural networks. Finally, we demonstrate that failures at either stage can prevent a model from learning a generalizable solution to our fairly simple tasks. By understanding ViTs in terms of discrete processing stages, one can more precisely diagnose and rectify shortcomings of existing and future models.
Michael A. Lepori, Alexa R. Tartaglini, Wai Keen Vong, Thomas Serre, Brenden M. Lake, Ellie Pavlick
NeurIPS6
2024 Talking Heads: Understanding Inter-Layer Communication in Transformer Language Models
abstract
Although it is known that transformer language models (LMs) pass features from early layers to later layers, it is not well understood how this information is represented and routed by the model. We analyze a mechanism used in two LMs to selectively inhibit items in a context in one task, and find that it underlies a commonly used abstraction across many context-retrieval behaviors. Specifically, we find that models write into low-rank subspaces of the residual stream to represent features which are then read out by later layers, forming low-rank *communication channels* (Elhage et al., 2021) between layers. A particular 3D subspace in model activations in GPT-2 can be traversed to positionally index items in lists, and we show that this mechanism can explain an otherwise arbitrary-seeming sensitivity of the model to the order of items in the prompt. That is, the model has trouble copying the correct information from context when many items ``crowd" this limited space. By decomposing attention heads with the Singular Value Decomposition (SVD), we find that previously described interactions between heads separated by one or more layers can be predicted via analysis of their weight matrices alone. We show that it is possible to manipulate the internal model representations as well as edit model weights based on the mechanism we discover in order to significantly improve performance on our synthetic Laundry List task, which requires recall from a list, often improving task accuracy by over 20\%. Our analysis reveals a surprisingly intricate interpretable structure learned from language model pretraining, and helps us understand why sophisticated LMs sometimes fail in simple domains, facilitating future analysis of more complex behaviors.
Jack Merullo, Carsten Eickhoff, Ellie Pavlick
NeurIPS3
2023 Analyzing Modular Approaches for Visual Question Decomposition
abstract
Modular neural networks without additional training have recently been shown to surpass end-to-end neural networks on challenging vision-language tasks.The latest such methods simultaneously introduce LLM-based code generation to build programs and a number of skill-specific, task-oriented modules to execute them.In this paper, we focus on ViperGPT and ask where its additional performance comes from and how much is due to the (state-of-art, end-to-end) BLIP-2 model it subsumes vs. additional symbolic components.To do so, we conduct a controlled study (comparing end-toend, modular, and prompting-based methods across several VQA benchmarks).We find that ViperGPT's reported gains over BLIP-2 can be attributed to its selection of task-specific modules, and when we run ViperGPT using a more task-agnostic selection of modules, these gains go away.ViperGPT retains much of its performance if we make prominent alterations to its selection of modules: e.g.removing or retaining only BLIP-2.We also compare ViperGPT against a prompting-based decomposition strategy and find that, on some benchmarks, modular approaches significantly benefit by representing subtasks with natural language, instead of code.Our code is fully available at https://github.com/brown-palm/ visual-question-decomposition.
Apoorv Khandelwal 0001, Ellie Pavlick, Chen Sun 0002
EMNLP2
2023 Characterizing Mechanisms for Factual Recall in Language Models
abstract
Language Models (LMs) often must integrate facts they memorized in pretraining with new information that appears in a given context.These two sources can disagree, causing competition within the model, and it is unclear how an LM will resolve the conflict.On a dataset that queries for knowledge of world capitals, we investigate both distributional and mechanistic determinants of LM behavior in such situations.Specifically, we measure the proportion of the time an LM will use a counterfactual prefix (e.g., "The capital of Poland is London") to overwrite what it learned in pretraining ("Warsaw").On Pythia and GPT2, the training frequency of both the query country ("Poland") and the in-context city ("London") highly affect the models' likelihood of using the counterfactual.We then use head attribution to identify individual attention heads that either promote the memorized answer or the in-context answer in the logits.By scaling up or down the value vector of these heads, we can control the likelihood of using the in-context answer on new data.This method can increase the rate of generating the in-context answer to 88% of the time simply by scaling a single head at runtime.Our work contributes to a body of evidence showing that we can often localize model behaviors to specific components and provides a proof of concept for how future methods might control model behavior dynamically at runtime.
Qinan Yu, Jack Merullo, Ellie Pavlick
EMNLP3
2023 Emergence of Abstract State Representations in Embodied Sequence Modeling
abstract
Decision making via sequence modeling aims to mimic the success of language models, where actions taken by an embodied agent are modeled as tokens to predict.Despite their promising performance, it remains unclear if embodied sequence modeling leads to the emergence of internal representations that represent the environmental state information.A model that lacks abstract state representations would be liable to make decisions based on surface statistics which fail to generalize.We take the BabyAI environment, a grid world in which language-conditioned navigation tasks are performed, and build a sequence modeling Transformer, which takes a language instruction, a sequence of actions, and environmental observations as its inputs.In order to investigate the emergence of abstract state representations, we design a "blindfolded" navigation task, where only the initial environmental layout, the language instruction, and the action sequence to complete the task are available for training.Our probing results show that intermediate environmental layouts can be reasonably reconstructed from the internal activations of a trained model, and that language instructions play a role in the reconstruction accuracy.Our results suggest that many key features of state representations can emerge via embodied sequence modeling, supporting an optimistic outlook for applications of sequence modeling objectives to more complex embodied decision-making domains.1
Tian Yun 0001, Zilai Zeng, Kunal Handa, Ashish V. Thapliyal, Bo Pang 0001, Ellie Pavlick, Chen Sun 0002
EMNLP6
2023 Linearly Mapping from Image to Text Space
Jack Merullo, Louis Castricato, Carsten Eickhoff, Ellie Pavlick
ICLR4
2023 Can Neural Networks Learn Implicit Logic from Physical Reasoning?
Aaron Traylor, Roman Feiman, Ellie Pavlick
ICLR3
2023 Break It Down: Evidence for Structural Compositionality in Neural Networks
abstract
Though modern neural networks have achieved impressive performance in both vision and language tasks, we know little about the functions that they implement. One possibility is that neural networks implicitly break down complex tasks into subroutines, implement modular solutions to these subroutines, and compose them into an overall solution to a task --- a property we term structural compositionality. Another possibility is that they may simply learn to match new inputs to learned templates, eliding task decomposition entirely. Here, we leverage model pruning techniques to investigate this question in both vision and language across a variety of architectures, tasks, and pretraining regimens. Our results demonstrate that models oftentimes implement solutions to subroutines via modular subnetworks, which can be ablated while maintaining the functionality of other subnetworks. This suggests that neural networks may be able to learn compositionality, obviating the need for specialized symbolic mechanisms.
Michael A. Lepori, Thomas Serre, Ellie Pavlick
NeurIPS3
2022 Mapping Language Models to Grounded Conceptual Spaces
Roma Patel, Ellie Pavlick
ICLR2
2022 The MultiBERTs: BERT Reproductions for Robustness Analysis
Thibault Sellam, Steve Yadlowsky, Ian Tenney, Jason Wei, Naomi Saphra, Alexander D'Amour, Tal Linzen, Jasmijn Bastings, Iulia Turc, Jacob Eisenstein, Dipanjan Das 0001, Ellie Pavlick
ICLR12
2022 Do Trajectories Encode Verb Meaning?
abstract
Distributional models learn representations of words from text, but are criticized for their lack of grounding, or the linking of text to the nonlinguistic world.Grounded language models have had success in learning to connect concrete categories like nouns and adjectives to the world via images and videos, but can struggle to isolate the meaning of the verbs themselves from the context in which they typically occur.In this paper, we investigate the extent to which trajectories (i.e. the position and rotation of objects over time) naturally encode verb semantics.We build a procedurally generated agent-object-interaction dataset, obtain human annotations for the verbs that occur in this data, and compare several methods for representation learning given the trajectories.We find that trajectories correlate as-is with some verbs (e.g., fall), and that additional abstraction via self-supervised pretraining can further capture nuanced differences in verb meaning (e.g., roll vs. slide).
Dylan Ebert, Chen Sun 0002, Ellie Pavlick
NAACL-HLT3
2022 Do Prompt-Based Models Really Understand the Meaning of Their Prompts?
abstract
Recently, a boom of papers has shown extraordinary progress in zero-shot and few-shot learning with various prompt-based models.It is commonly argued that prompts help models to learn faster in the same way that humans learn faster when provided with task instructions expressed in natural language.In this study, we experiment with over 30 prompt templates manually written for natural language inference (NLI).We find that models can learn just as fast with many prompts that are intentionally irrelevant or even pathologically misleading as they do with instructively "good" prompts.Further, such patterns hold even for models as large as 175 billion parameters (Brown et al., 2020) as well as the recently proposed instruction-tuned models which are trained on hundreds of prompts (Sanh et al., 2021).That is, instruction-tuned models often produce good predictions with irrelevant and misleading prompts even at zero shots.In sum, notwithstanding prompt-based models' impressive improvement, we find evidence of serious limitations that question the degree to which such improvement is derived from models understanding task instructions in ways analogous to humans' use of task instructions.* Unabridged version available on arXiv.Code, interactive figures, and statistical test results available at https://github. com/awebson/prompt_semantics arbitrary dimensions of a one-hot vector.In contrast, suppose a human is given a prompt such as: Given that " Given that " Given that " Given that " Given that " Given that " Given that " Given that " Given that " Given that " Given that " Given that " Given that " Given that " Given that " Given that " Given that "no weapons of mass destruction found in Iraq yet.", is it definitely correct that " ", is it definitely correct that " ", is it definitely correct that " ", is it definitely correct that " ", is it definitely correct that " ", is it definitely correct that " ", is it definitely correct that " ", is it definitely correct that " ", is it definitely correct that " ", is it definitely correct that " ", is it definitely correct that " ", is it definitely correct that " ", is it definitely correct that " ", is it definitely correct that " ", is it definitely correct that " ", is it definitely correct that " ", is it definitely correct that "weapons of mass destruction found in Iraq."? "? "? "? "?"? "? "? "? "? "? "? "? "? "? "? "? 1 Then it would be no surprise that they are able to perform the task more accurately and without needing many examples to figure out what the task is.
Albert Webson, Ellie Pavlick
NAACL-HLT2
2022 Evaluation beyond Task Performance: Analyzing Concepts in AlphaZero in Hex
abstract
AlphaZero, an approach to reinforcement learning that couples neural networks and Monte Carlo tree search (MCTS), has produced state-of-the-art strategies for traditional board games like chess, Go, shogi, and Hex. While researchers and game commentators have suggested that AlphaZero uses concepts that humans consider important, it is unclear how these concepts are captured in the network. We investigate AlphaZero's internal representations in the game of Hex using two evaluation techniques from natural language processing (NLP): model probing and behavioral tests. In doing so, we introduce several new evaluation tools to the RL community, and illustrate how evaluations other than task performance can be used to provide a more complete picture of a model's strengths and weaknesses. Our analyses in the game of Hex reveal interesting patterns and generate some testable hypotheses about how such models learn in general. For example, we find that the MCTS discovers concepts before the neural network learns to encode them. We also find that concepts related to short-term end-game planning are best encoded in the final layers of the model, whereas concepts related to long-term planning are encoded in the middle layers of the model.
Charles Lovering, Jessica Zosa Forde, George Dimitri Konidaris, Ellie Pavlick, Michael L. Littman
NeurIPS4
2022 Unit Testing for Concepts in Neural Networks
abstract
Abstract Many complex problems are naturally understood in terms of symbolic concepts. For example, our concept of “cat” is related to our concepts of “ears” and “whiskers” in a non-arbitrary way. Fodor (1998) proposes one theory of concepts, which emphasizes symbolic representations related via constituency structures. Whether neural networks are consistent with such a theory is open for debate. We propose unit tests for evaluating whether a system’s behavior is consistent with several key aspects of Fodor’s criteria. Using a simple visual concept learning task, we evaluate several modern neural architectures against this specification. We find that models succeed on tests of groundedness, modularity, and reusability of concepts, but that important questions about causality remain open. Resolving these will require new methods for analyzing models’ internal states.
Charles Lovering, Ellie Pavlick
Trans. Assoc. Comput. Linguistics2
2021 Which Linguist Invented the Lightbulb? Presupposition Verification for Question-Answering
abstract
Najoung Kim, Ellie Pavlick, Burcu Karagol Ayan, Deepak Ramachandran. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Najoung Kim, Ellie Pavlick, Burcu Karagol Ayan, Deepak Ramachandran
ACL/IJCNLP (1)2
2021 The Anatomy of Discourse: Linguistic Predictors of Narrative and Argument Quality
Sheridan Feucht, Babak Hemmatian, Rachel Avram, Alexander Wey, Kate Spitalnic, Muskaan Garg, Carsten Eickhoff, Ellie Pavlick, Björn Sandstede, Steven A. Sloman
CogSci8
2021 Can computers tell a story? Discourse Structure in Computer-generated Text and Humans
Alexander Wey, Babak Hemmatian, Rachel Avram, Sheridan Feucht, Kate Spitalnic, Muskaan Garg, Carsten Eickhoff, Ellie Pavlick, Björn Sandstede, Steven A. Sloman
CogSci8
2021 Can Language Models Encode Perceptual Structure Without Grounding? A Case Study in Color
abstract
Pretrained language models have been shown to encode relational information, such as the relations between entities or concepts in knowledge-bases -(Paris, Capital, France).However, simple relations of this type can often be recovered heuristically and the extent to which models implicitly reflect topological structure that is grounded in world, such as perceptual structure, is unknown.To explore this question, we conduct a thorough case study on color.Namely, we employ a dataset of monolexemic color terms and color chips represented in CIELAB, a color space with a perceptually meaningful distance metric.Using two methods of evaluating the structural alignment of colors in this space with textderived color term representations, we find significant correspondence.Analyzing the differences in alignment across the color spectrum, we find that warmer colors are, on average, better aligned to the perceptual color space than cooler ones, suggesting an intriguing connection to findings from recent work on efficient communication in color naming.Further analysis suggests that differences in alignment are, in part, mediated by collocationality and differences in syntactic usage, posing questions as to the relationship between color perception and usage and context.
Mostafa Abdou, Artur Kulmizev, Daniel Hershcovich, Stella Frank, Ellie Pavlick, Anders Søgaard
CoNLL5
2021 "Was it "stated" or was it "claimed"?: How linguistic bias affects generative language models
abstract
People use language in subtle and nuanced ways to convey their beliefs.For instance, saying claimed instead of said casts doubt on the truthfulness of the underlying proposition, thus representing the author's opinion on the matter.Several works have identified classes of words that induce such framing effects.In this paper, we test whether generative language models are sensitive to these linguistic cues.In particular, we test whether prompts that contain linguistic markers of author bias (e.g., hedges, implicatives, subjective intensifiers, assertives) influence the distribution of the generated text.Although these framing effects are subtle and stylistic, we find qualitative and quantitative evidence that they lead to measurable style and topic differences in the generated text, leading to language that is more polarised (both positively and negatively) and, anecdotally, appears more skewed towards controversial entities and events.
Roma Patel, Ellie Pavlick
EMNLP (1)2
2021 Frequency Effects on Syntactic Rule Learning in Transformers
abstract
Pre-trained language models perform well on a variety of linguistic tasks that require symbolic reasoning, raising the question of whether such models implicitly represent abstract symbols and rules.We investigate this question using the case study of BERT's performance on English subject-verb agreement.Unlike prior work, we train multiple instances of BERT from scratch, allowing us to perform a series of controlled interventions at pre-training time.We show that BERT often generalizes well to subject-verb pairs that never occurred in training, suggesting a degree of rule-governed behavior.We also find, however, that performance is heavily influenced by word frequency, with experiments showing that both the absolute frequency of a verb form, as well as the frequency relative to the alternate inflection, are causally implicated in the predictions BERT makes at inference time.Closer analysis of these frequency effects reveals that BERT's behavior is consistent with a system that correctly applies the SVA rule in general but struggles to overcome strong training priors and to estimate agreement features (singular vs. plural) on infrequent lexical items.
Jason Wei, Dan Garrette, Tal Linzen, Ellie Pavlick
EMNLP (1)4
2021 Predicting Inductive Biases of Pre-Trained Models
Charles Lovering, Rohan Jha, Tal Linzen, Ellie Pavlick
ICLR4
2021 Spatial Language Understanding for Object Search in Partially Observed City-scale Environments
abstract
Humans use spatial language to naturally describe object locations and their relations. Interpreting spatial language not only adds a perceptual modality for robots, but also reduces the barrier of interfacing with humans. Previous work primarily considers spatial language as goal specification for instruction following tasks in fully observable domains, often paired with reference paths for reward-based learning. However, spatial language is inherently subjective and potentially ambiguous or misleading. Hence, in this paper, we consider spatial language as a form of stochastic observation. We propose SLOOP (Spatial Language Object-Oriented POMDP), a new framework for partially observable decision making with a probabilistic observation model for spatial language. We apply SLOOP to object search in city-scale environments. To interpret ambiguous, context-dependent prepositions (e.g. front), we design a simple convolutional neural network that predicts the language provider’s latent frame of reference (FoR) given the environment context. Search strategies are computed via an online POMDP planner based on Monte Carlo Tree Search. Evaluation based on crowdsourced language data, collected over areas of five cities in OpenStreetMap, shows that our approach achieves faster search and higher success rate compared to baselines, with a wider margin as the spatial language becomes more complex. Finally, we demonstrate the proposed method in AirSim, a realistic simulator where a drone is tasked to find cars in a neighborhood environment.
Kaiyu Zheng, Deniz Bayazit, Rebecca Mathew, Ellie Pavlick, Stefanie Tellex
RO-MAN4
2020 Do "Undocumented Workers" == "Illegal Aliens"? Differentiating Denotation and Connotation in Vector Spaces
abstract
In politics, neologisms are frequently invented for partisan objectives.For example, "undocumented workers" and "illegal aliens" refer to the same group of people (i.e., they have the same denotation), but they carry clearly different connotations.Examples like these have traditionally posed a challenge to referencebased semantic theories and led to increasing acceptance of alternative theories (e.g., Two-Factor Semantics) among philosophers and cognitive scientists.In NLP, however, popular pretrained models encode both denotation and connotation as one entangled representation.In this study, we propose an adversarial neural network that decomposes a pretrained representation as independent denotation and connotation representations.For intrinsic interpretability, we show that words with the same denotation but different connotations (e.g., "immigrants" vs. "aliens", "estate tax" vs. "death tax") move closer to each other in denotation space while moving further apart in connotation space.For extrinsic application, we train an information retrieval system with our disentangled representations and show that the denotation vectors improve the viewpoint diversity of document rankings.
Albert Webson, Zhizhong Chen, Carsten Eickhoff, Ellie Pavlick
EMNLP (1)4
2020 Grounding Language to Landmarks in Arbitrary Outdoor Environments
abstract
Robots operating in outdoor, urban environments need the ability to follow complex natural language commands which refer to never-before-seen landmarks. Existing approaches to this problem are limited because they require training a language model for the landmarks of a particular environment before a robot can understand commands referring to those landmarks. To generalize to new environments outside of the training set, we present a framework that parses references to landmarks, then assesses semantic similarities between the referring expression and landmarks in a predefined semantic map of the world, and ultimately translates natural language commands to motion plans for a drone. This framework allows the robot to ground natural language phrases to landmarks in a map when both the referring expressions to landmarks and the landmarks themselves have not been seen during training. We test our framework with a 14-person user evaluation demonstrating an end-to-end accuracy of 76.19% in an unseen environment. Subjective measures show that users find our system to have high performance and low workload. These results demonstrate our approach enables untrained users to control a robot in large unseen outdoor environments with unconstrained natural language.
Matthew Berg, Deniz Bayazit, Rebecca Mathew, Ariel Rotter-Aboyoun, Ellie Pavlick, Stefanie Tellex
ICRA5
2020 Sochiatrist: Signals of Affect in Messaging Data
abstract
Messaging is a common mode of communication, with conversations written informally between individuals. Interpreting emotional affect from messaging data can lead to a powerful form of reflection or act as a support for clinical therapy. Existing analysis techniques for social media commonly use LIWC and VADER for automated sentiment estimation. We correlate LIWC, VADER, and ratings from human reviewers with affect scores from 25 participants. We explore differences in how and when each technique is successful. Results show that human review does better than VADER, the best automated technique, when humans are judging positive affect ($r_s=0.45$ correlation when confident, $r_s=0.30$ overall). Surprisingly, human reviewers only do slightly better than VADER when judging negative affect ($r_s=0.38$ correlation when confident, $r_s=0.29$ overall). Compared to prior literature, VADER correlates more closely with PANAS scores for private messaging than public social media. Our results indicate that while any technique that serves as a proxy for PANAS scores has moderate correlation at best, there are some areas to improve the automated techniques by better considering context and timing in conversations.
Talie Massachi, Grant Fong, Varun Mathur, Sachin R. Pendse, Gabriela Hoefer, Jessica J. Fu, Nikita Ramoji, Nicole Nugent, Megan Ranney, Daniel P. Dickstein, Michael F. Armey, Ellie Pavlick, Jeff Huang 0002
Proc. ACM Hum. Comput. Interact.13
2019 Right for the Wrong Reasons: Diagnosing Syntactic Heuristics in Natural Language Inference
abstract
A machine learning system can score well on a given test set by relying on heuristics that are effective for frequent example types but break down in more challenging cases.We study this issue within natural language inference (NLI), the task of determining whether one sentence entails another.We hypothesize that statistical NLI models may adopt three fallible syntactic heuristics: the lexical overlap heuristic, the subsequence heuristic, and the constituent heuristic.To determine whether models have adopted these heuristics, we introduce a controlled evaluation set called HANS (Heuristic Analysis for NLI Systems), which contains many examples where the heuristics fail.We find that models trained on MNLI, including BERT, a state-of-the-art model, perform very poorly on HANS, suggesting that they have indeed adopted these heuristics.We conclude that there is substantial room for improvement in NLI systems, and that the HANS dataset can motivate and measure progress in this area.
Tom McCoy 0001, Ellie Pavlick, Tal Linzen
ACL (1)2
2019 BERT Rediscovers the Classical NLP Pipeline
abstract
Pre-trained text encoders have rapidly advanced the state of the art on many NLP tasks.We focus on one such model, BERT, and aim to quantify where linguistic information is captured within the network.We find that the model represents the steps of the traditional NLP pipeline in an interpretable and localizable way, and that the regions responsible for each step appear in the expected sequence: POS tagging, parsing, NER, semantic roles, then coreference.Qualitative analysis reveals that the model can and often does adjust this pipeline dynamically, revising lowerlevel decisions on the basis of disambiguating information from higher-level representations.
Ian Tenney, Dipanjan Das 0001, Ellie Pavlick
ACL (1)3
2019 Can You Tell Me How to Get Past Sesame Street? Sentence-Level Pretraining Beyond Language Modeling
abstract
Alex Wang, Jan Hula, Patrick Xia, Raghavendra Pappagari, R. Thomas McCoy, Roma Patel, Najoung Kim, Ian Tenney, Yinghui Huang, Katherin Yu, Shuning Jin, Berlin Chen, Benjamin Van Durme, Edouard Grave, Ellie Pavlick, Samuel R. Bowman. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019.
Jan Hula, Patrick Xia 0002, Raghavendra Pappagari, Tom McCoy 0001, Roma Patel, Najoung Kim, Ian Tenney, Katherin Yu, Shuning Jin, Berlin Chen, Benjamin Van Durme, Edouard Grave, Ellie Pavlick, Samuel R. Bowman
ACL (1)15
2019 Emergent Compositionality in Signaling Games
Nicholas Tomlin, Ellie Pavlick
CogSci2
2019 How well do NLI models capture verb veridicality?
abstract
Alexis Ross, Ellie Pavlick. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Alexis Ross, Ellie Pavlick
EMNLP/IJCNLP (1)2
2019 What do you learn from context? Probing for sentence structure in contextualized word representations
Ian Tenney, Patrick Xia 0002, Berlin Chen, Adam Poliak, Tom McCoy 0001, Najoung Kim, Benjamin Van Durme, Samuel R. Bowman, Dipanjan Das 0001, Ellie Pavlick
ICLR (Poster)11
2019 Inherent Disagreements in Human Textual Inferences
abstract
We analyze human’s disagreements about the validity of natural language inferences. We show that, very often, disagreements are not dismissible as annotation “noise”, but rather persist as we collect more ratings and as we vary the amount of context provided to raters. We further show that the type of uncertainty captured by current state-of-the-art models for natural language inference is not reflective of the type of uncertainty present in human disagreements. We discuss implications of our results in relation to the recognizing textual entailment (RTE)/natural language inference (NLI) task. We argue for a refined evaluation objective that requires models to explicitly capture the full distribution of plausible human judgments.
Ellie Pavlick, Tom Kwiatkowski
Trans. Assoc. Comput. Linguistics1
2018 Learning Scalar Adjective Intensity from Paraphrases
abstract
Adjectives like warm, hot, and scalding all describe temperature but differ in intensity.Understanding these differences between adjectives is a necessary part of reasoning about natural language.We propose a new paraphrasebased method to automatically learn the relative intensity relation that holds between a pair of scalar adjectives.Our approach analyzes over 36k adjectival pairs from the Paraphrase Database under the assumption that, for example, paraphrase pair really hot ↔ scalding suggests that hot < scalding.We show that combining this paraphrase evidence with existing, complementary pattern-and lexicon-based approaches improves the quality of systems for automatically ordering sets of scalar adjectives and inferring the polarity of indirect answers to yes/no questions.
Anne Cocos, Veronica Wharton, Ellie Pavlick, Marianna Apidianaki, Chris Callison-Burch
EMNLP3
2018 WikiAtomicEdits: A Multilingual Corpus of Wikipedia Edits for Modeling Language and Discourse
abstract
We release a corpus of 43 million atomic edits across 8 languages.These edits are mined from Wikipedia edit history and consist of instances in which a human editor has inserted a single contiguous phrase into, or deleted a single contiguous phrase from, an existing sentence.We use the collected data to show that the language generated during editing differs from the language that we observe in standard corpora, and that models trained on edits encode different aspects of semantics and discourse than models trained on raw, unstructured text.We release the full corpus as a resource to aid ongoing research in semantics, discourse, and representation learning.
Manaal Faruqui, Ellie Pavlick, Ian Tenney, Dipanjan Das 0001
EMNLP2
2018 Collecting Diverse Natural Language Inference Problems for Sentence Representation Evaluation
abstract
Adam Poliak, Aparajita Haldar, Rachel Rudinger, J. Edward Hu, Ellie Pavlick, Aaron Steven White, Benjamin Van Durme. Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing. 2018.
Adam Poliak, Aparajita Haldar, Rachel Rudinger, Edward J. Hu, Ellie Pavlick, Aaron Steven White, Benjamin Van Durme
EMNLP5
2017 Identifying 1950s American Jazz Musicians: Fine-Grained IsA Extraction via Modifier Composition
abstract
We present a method for populating fine-grained classes (e.g., "1950s American jazz musicians") with instances (e.g., Charles Mingus).While stateof-the-art methods tend to treat class labels as single lexical units, the proposed method considers each of the individual modifiers in the class label relative to the head.An evaluation on the task of reconstructing Wikipedia category pages demonstrates a >10 point increase in AUC, over a strong baseline relying on widely-used Hearst patterns.
Ellie Pavlick, Marius Pasca
ACL (1)1
2016 Most "babies" are "little" and most "problems" are "huge": Compositional Entailment in Adjective-Nouns
abstract
We examine adjective-noun (AN) composition in the task of recognizing textual entailment (RTE).We analyze behavior of ANs in large corpora and show that, despite conventional wisdom, adjectives do not always restrict the denotation of the nouns they modify.We use natural logic to characterize the variety of entailment relations that can result from AN composition.Predicting these relations depends on context and on commonsense knowledge, making AN composition especially challenging for current RTE systems.We demonstrate the inability of current stateof-the-art systems to handle AN composition in a simplified RTE task which involves the insertion of only a single word.
Ellie Pavlick, Chris Callison-Burch
ACL (1)1
2016 Tense Manages to Predict Implicative Behavior in Verbs
abstract
Implicative verbs (e.g.manage) entail their complement clauses, while non-implicative verbs (e.g.want) do not.For example, while managing to solve the problem entails solving the problem, no such inference follows from wanting to solve the problem.Differentiating between implicative and non-implicative verbs is therefore an essential component of natural language understanding, relevant to applications such as textual entailment and summarization.We present a simple method for predicting implicativeness which exploits known constraints on the tense of implicative verbs and their complements.We show that this yields an effective, data-driven way of capturing this nuanced property in verbs.(0.14) UFJ wants to merge with Mitsubishi, a combination that'd surpass Citigroup as the world's biggest bank.⇒ The merger of Japanese Banks creates the world's biggest bank.(0.55)After graduating, Gallager chose to accept a full scholarship to play football for Temple University.⇒ Gallager attended Temple University.(0.68) Wilkins was allowed to leave in 1987 to join French outfit Paris Saint-Germain.⇒ Wilkins departed Milan in 1987.
Ellie Pavlick, Chris Callison-Burch
EMNLP1
2016 The Gun Violence Database: A new task and data set for NLP
abstract
We argue that NLP researchers are especially well-positioned to contribute to the national discussion about gun violence.Reasoning about the causes and outcomes of gun violence is typically dominated by politics and emotion, and data-driven research on the topic is stymied by a shortage of data and a lack of federal funding.However, data abounds in the form of unstructured text from news articles across the country.This is an ideal application of NLP technologies, such as relation extraction, coreference resolution, and event detection.We introduce a new and growing dataset, the Gun Violence Database, in order to facilitate the adaptation of current NLP technologies to the domain of gun violence, thus enabling better social science research on this important and under-resourced problem.
Ellie Pavlick, Heng Ji 0001, Xiaoman Pan, Chris Callison-Burch
EMNLP1
2016 An Empirical Analysis of Formality in Online Communication
abstract
This paper presents an empirical study of linguistic formality. We perform an analysis of humans’ perceptions of formality in four different genres. These findings are used to develop a statistical model for predicting formality, which is evaluated under different feature settings and genres. We apply our model to an investigation of formality in online discussion forums, and present findings consistent with theories of formality and linguistic coordination.
Ellie Pavlick, Joel R. Tetreault
Trans. Assoc. Comput. Linguistics1
2016 Optimizing Statistical Machine Translation for Text Simplification
abstract
Most recent sentence simplification systems use basic machine translation models to learn lexical and syntactic paraphrases from a manually simplified parallel corpus. These methods are limited by the quality and quantity of manually simplified corpora, which are expensive to build. In this paper, we conduct an in-depth adaptation of statistical machine translation to perform text simplification, taking advantage of large-scale paraphrases learned from bilingual texts and a small amount of manual simplifications with multiple references. Our work is the first to design automatic metrics that are effective for tuning and evaluating simplification systems, which will facilitate iterative development for this task.
Wei Xu 0004, Courtney Napoles, Ellie Pavlick, Quanze Chen, Chris Callison-Burch
Trans. Assoc. Comput. Linguistics3
2015 Adding Semantics to Data-Driven Paraphrasing
abstract
Ellie Pavlick, Johan Bos, Malvina Nissim, Charley Beller, Benjamin Van Durme, Chris Callison-Burch. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
Ellie Pavlick, Johan Bos, Malvina Nissim, Charley Beller, Benjamin Van Durme, Chris Callison-Burch
ACL (1)1
2015 Extracting Structured Information via Automatic + Human Computation
abstract
We present a system for extracting structured information from unstructured text using a combination of information retrieval, natural language processing, machine learning, and crowdsourcing. We test our pipeline by building a structured database of gun violence incidents in the United States. The results of our pilot study demonstrate that the proposed methodology is a viable way of collecting large-scale, up-to-date data for public health, public policy, and social science research.
Ellie Pavlick, Chris Callison-Burch
HCOMP1
2015 Crowdsourcing for NLP
abstract
Crowdsourced applications to scientific problems is a hot research area, with over 10,000 publications in the past five years. Platforms such as Amazons Mechanical Turk and CrowdFlower provide researchers with easy access to large numbers of workers. The crowds vast supply of inexpensive, intelligent labor allows people to attack problems that were previously impractical and gives potential for detailed scientific inquiry of social, psychological, economic, and linguistic phenomena via massive sample sizes of human annotated data. We introduce crowdsourcing and describe how it is being used in both industry and academia. Crowdsourcing is valuable to computational linguists both (a) as a source of labeled training data for use in machine learning and (b) as a means of collecting computational social science data that link language use to underlying beliefs and behavior. We present case studies for both categories: (a) collecting labeled data for use in natural language processing tasks such as word sense disambiguation and machine translation and (b) collecting experimental data in the context of psychology; e.g. finding how word use varies with age, sex, personality, health, and happiness. We will also cover tools and techniques for crowdsourcing. Effectively collecting crowdsourced data requires careful attention to the collection process, through selection of appropriately qualified workers, giving clear instructions that are understandable to non-?experts, and performing quality control on the results to eliminate spammers who complete tasks randomly or carelessly in order to collect the small financial reward. We will introduce different crowdsourcing platforms, review privacy and institutional review board issues, and provide rules of thumb for cost and time estimates. Crowdsourced data also has a particular structure that raises issues in statistical analysis; we describe some of the key methods to address these issues. No prior exposure to the area is required.
Chris Callison-Burch, Lyle H. Ungar, Ellie Pavlick
HLT-NAACL3
2015 Inducing Lexical Style Properties for Paraphrase and Genre Differentiation
abstract
We present an intuitive and effective method for inducing style scores on words and phrases. We exploit signal in a phrase’s rate of occurrence across stylistically contrasting corpora, making our method simple to implement and efficient to scale. We show strong results both intrinsically, by correlation with human judgements, and extrinsically, in applications to genre analysis and paraphrasing.
Ellie Pavlick, Ani Nenkova
HLT-NAACL1
2014 Are Two Heads Better than One? Crowdsourced Translation via a Two-Step Collaboration of Non-Professional Translators and Editors
abstract
Crowdsourcing is a viable mechanism for creating training data for machine translation.It provides a low cost, fast turnaround way of processing large volumes of data.However, when compared to professional translation, naive collection of translations from non-professionals yields low-quality results.Careful quality control is necessary for crowdsourcing to work well.In this paper, we examine the challenges of a two-step collaboration process with translation and post-editing by non-professionals.We develop graphbased ranking models that automatically select the best output from multiple redundant versions of translations and edits, and improves translation quality closer to professionals.
Mingkun Gao, Ellie Pavlick, Chris Callison-Burch
ACL (1)3
2014 Poetry of the Crowd: A Human Computation Algorithm to Convert Prose into Rhyming Verse
abstract
Poetry composition is a very complex task that requires a poet to satisfy multiple constraints concurrently. We believe that the task can be augmented by combining the creative abilities of humans with computational algorithms that efficiently constrain and permute available choices. We present a hybrid method for generating poetry from prose that combines crowdsourcing with natural language processing (NLP) machinery. We test the ability of crowd workers to accomplish the technically challenging and creative task of composing poems.
Quanze Chen, Chenyang Lei, Wei Xu 0004, Ellie Pavlick, Chris Callison-Burch
HCOMP4
2014 The Language Demographics of Amazon Mechanical Turk
abstract
We present a large scale study of the languages spoken by bilingual workers on Mechanical Turk (MTurk). We establish a methodology for determining the language skills of anonymous crowd workers that is more robust than simple surveying. We validate workers’ self-reported language skill claims by measuring their ability to correctly translate words, and by geolocating workers to see if they reside in countries where the languages are likely to be spoken. Rather than posting a one-off survey, we posted paid tasks consisting of 1,000 assignments to translate a total of 10,000 words in each of 100 languages. Our study ran for several months, and was highly visible on the MTurk crowdsourcing platform, increasing the chances that bilingual workers would complete it. Our study was useful both to create bilingual dictionaries and to act as census of the bilingual speakers on MTurk. We use this data to recommend languages with the largest speaker populations as good candidates for other researchers who want to develop crowdsourced, multilingual technologies. To further demonstrate the value of creating data via crowdsourcing, we hire workers to create bilingual parallel corpora in six Indian languages, and use them to train statistical machine translation systems.
Ellie Pavlick, Matt Post, Ann Irvine, Dmitry Kachaev, Chris Callison-Burch
Trans. Assoc. Comput. Linguistics1