Naomi Saphra

dblp:131/6883 · DBLP profile ↗
← Back
19ranked-venue papers
2as first author
15since 2021 · last 2025
0000-0001-8607-3515ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 18 · 2 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2025 Using Shapley interactions to understand how models use structure
abstract
Divyansh Singhvi, Diganta Misra, Andrej Erkelens, Raghav Jain, Isabel Papadimitriou, Naomi Saphra. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Divyansh Singhvi, Diganta Misra, Andrej Erkelens, Raghav Jain, Isabel Papadimitriou, Naomi Saphra
ACL (1)6
2025 Data Drives Unstable Hierarchical Generalization in LMs
abstract
Early in training, LMs can behave like n-gram models, but eventually, they often learn treebased syntactic rules and generalize hierarchically out of distribution (OOD).We study this shift using controlled grammar-learning tasks: question formation and tense inflection.We find a model learns to generalize hierarchically if its training data is complex-in particular, if it includes center-embedded clauses, a special syntactic structure.Under this definition, complex data drives hierarchical rules, while less complex data encourages shortcut learning in the form of n-gram-like linear rules.Furthermore, we find that a model uses rules to generalize, whether hierarchical or linear, if its training data is diverse-in particular, if it includes many distinct syntax trees in the training set.Under this definition, diverse data promotes stable rule learning, whereas less diverse data promotes memorization of individual syntactic sequences.Finally, intermediate diversity and intermediate complexity form an unstable regime, which is characterized by oscillatory learning dynamics and inconsistent behaviors across random seeds.These results highlight the central role of training data in shaping generalization and explain why competing strategies can lead to unstable outcomes.
Naomi Saphra, David Alvarez-Melis
EMNLP2
2025 Recite, Reconstruct, Recollect: Memorization in LMs as a Multifaceted Phenomenon
abstract
Memorization in language models is typically treated as a homogenous phenomenon, neglecting the specifics of the memorized data. We instead model memorization as the effect of a set of complex factors that describe each sample and relate it to the model and corpus. To build intuition around these factors, we break memorization down into a taxonomy: recitation of highly duplicated sequences, reconstruction of inherently predictable sequences, and recollection of sequences that are neither. We demonstrate the usefulness of our taxonomy by using it to construct a predictive model for memorization. By analyzing dependencies and inspecting the weights of the predictive model, we find that different factors have different influences on the likelihood of memorization depending on the taxonomic category.
USVSN Sai Prashanth, Alvin Deng, Kyle O'Brien, Jyothir S. V, Mohammad Aflah Khan, Jaydeep Borkar, Christopher A. Choquette-Choo, Jacob Ray Fuehne, Stella Biderman, Tracy Ke, Katherine Lee, Naomi Saphra
ICLR12
2025 PolyPythias: Stability and Outliers across Fifty Language Model Pre-Training Runs
abstract
The stability of language model pre-training and its effects on downstream performance are still understudied. Prior work shows that the training process can yield significantly different results in response to slight variations in initial conditions, e.g., the random seed. Crucially, the research community still lacks sufficient resources and tools to systematically investigate pre-training stability, particularly for decoder-only language models. We introduce the PolyPythias, a set of 45 new training runs for the Pythia model suite: 9 new seeds across 5 model sizes, from 14M to 410M parameters, resulting in about 7k new checkpoints that we release. Using these new 45 training runs, in addition to the 5 already available, we study the effects of different initial conditions determined by the seed---i.e., parameters' initialisation and data order---on (i) downstream performance, (ii) learned linguistic representations, and (iii) emergence of training phases. In addition to common scaling behaviours, our analyses generally reveal highly consistent training dynamics across both model sizes and initial conditions. Further, the new seeds for each model allow us to identify outlier training runs and delineate their characteristics. Our findings show the potential of using these methods to predict training stability.
Oskar van der Wal, Pietro Lesci, Max Müller-Eberstein, Naomi Saphra, Hailey Schoelkopf, Willem H. Zuidema, Stella Biderman
ICLR4
2024 Attribute Diversity Determines the Systematicity Gap in VQA
abstract
Although modern neural networks often generalize to new combinations of familiar concepts, the conditions that enable such compositionality have long been an open question.In this work, we study the systematicity gap in visual question answering: the performance difference between reasoning on previously seen and unseen combinations of object attributes.To test, we introduce a novel diagnostic dataset, CLEVR-HOPE.We find that the systematicity gap is not reduced by increasing the quantity of training data, but is reduced by increasing the diversity of training data.In particular, our experiments suggest that the more distinct attribute type combinations are seen during training, the more systematic we can expect the resulting model to be.We release our data and code at https://github.com/ ikb-a/systematicity-gap-in-vqa.
Ian Berlot-Attwell, Kumar Krishna Agrawal, Annabelle Michael Carrell, Naomi Saphra
EMNLP5
2024 ChatGPT Doesn't Trust Chargers Fans: Guardrail Sensitivity in Context
abstract
While the biases of language models in production are extensively documented, the biases of their guardrails have been neglected.This paper studies how contextual information about the user influences the likelihood of an LLM to refuse to execute a request.By generating user biographies that offer ideological and demographic information, we find a number of biases in guardrail sensitivity on GPT-3.5.Younger, female, and Asian-American personas are more likely to trigger a refusal guardrail when requesting censored or illegal information.Guardrails are also sycophantic, refusing to comply with requests for a political position the user is likely to disagree with.We find that certain identity groups and seemingly innocuous information, e.g., sports fandom, can elicit changes in guardrail sensitivity similar to direct statements of political ideology.For each demographic category and even for American football team fandom, we find that ChatGPT appears to infer a likely political ideology and modify guardrail behavior accordingly.
Victoria R. Li, Naomi Saphra
EMNLP3
2024 Fast Forwarding Low-Rank Training
abstract
Parameter efficient finetuning methods like lowrank adaptation (LoRA) aim to reduce the computational costs of finetuning pretrained Language Models (LMs).Enabled by these lowrank settings, we propose an even more efficient optimization strategy: Fast Forward, a simple and effective approach to accelerate large segments of training.In a Fast Forward stage, we repeat the most recent optimizer step until the loss stops improving on a tiny validation set.By alternating between regular optimization steps and Fast Forward stages, Fast Forward provides up to an 87% reduction in FLOPs and up to an 81% reduction in train time over standard SGD with Adam.We validate Fast Forward by finetuning various models on different tasks and demonstrate that it speeds up training without compromising model performance.Additionally, we analyze when and how to apply Fast Forward.
Adir Rahamim, Naomi Saphra, Sara Kangaslahti, Yonatan Belinkov
EMNLP2
2024 Sudden Drops in the Loss: Syntax Acquisition, Phase Transitions, and Simplicity Bias in MLMs
abstract
Most interpretability research in NLP focuses on understanding the behavior and features of a fully trained model. However, certain insights into model behavior may only be accessible by observing the trajectory of the training process. We present a case study of syntax acquisition in masked language models (MLMs) that demonstrates how analyzing the evolution of interpretable artifacts throughout training deepens our understanding of emergent behavior. In particular, we study Syntactic Attention Structure (SAS), a naturally emerging property of MLMs wherein specific Transformer heads tend to focus on specific syntactic relations. We identify a brief window in pretraining when models abruptly acquire SAS, concurrent with a steep drop in loss. This breakthrough precipitates the subsequent acquisition of linguistic capabilities. We then examine the causal role of SAS by manipulating SAS during training, and demonstrate that SAS is necessary for the development of grammatical capabilities. We further find that SAS competes with other beneficial traits during training, and that briefly suppressing SAS improves model quality. These findings offer an interpretation of a real-world example of both simplicity bias and breakthrough training dynamics.
Angelica Chen, Ravid Shwartz-Ziv, Kyunghyun Cho, Matthew L. Leavitt, Naomi Saphra
ICLR5
2024 TRAM: Bridging Trust Regions and Sharpness Aware Minimization
abstract
Sharpness-aware minimization (SAM) reports improving domain generalization by reducing the loss surface curvature in the parameter space. However, generalization during _fine-tuning_ is often more dependent on the transferability of _representations_ in the function space. Trust-region methods (TR) target this goal by regularizing representation curvature to reduce catastrophic forgetting of pre-trained task-agnostic information while adopting task-specific skills. We consider unifying these strategies for low curvature in both parameter space and function space to improve out-of-domain (OOD) generalization. We propose **Trust Region Aware Minimization** (TRAM), a SAM algorithm fine-tuning for low parameter sharpness and smooth, informative representations preserving pre-trained structure. TRAM uses a trust region bound to inform the SAM adversarial neighborhood, introducing an awareness of function curvature within optimization for flatter minima. We empirically validate TRAM in vision (cross-dataset adaptation) and text (OOD language modeling, zero-shot cross-lingual transfer) tasks where robust domain transfer and representation generality are critical. TRAM outperforms SAM- and TR-based optimization across all tasks, notably surpassing competing methods for hard transfer between _anticorrelated_ domains. TRAM establishes a novel standard in fine-tuning for domain-generalizable models with minimal additional computation over previous sharpness-aware methods.
Tom Sherborne, Naomi Saphra, Pradeep Dasigi, Hao Peng 0009
ICLR2
2024 First Tragedy, then Parse: History Repeats Itself in the New Era of Large Language Models
abstract
Naomi Saphra, Eve Fleisig, Kyunghyun Cho, Adam Lopez. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Naomi Saphra, Eve Fleisig, Kyunghyun Cho, Adam Lopez
NAACL-HLT1
2024 Transcendence: Generative Models Can Outperform The Experts That Train Them
abstract
Generative models are trained with the simple objective of imitating the conditional probability distribution induced by the data they are trained on. Therefore, when trained on data generated by humans, we may not expect the artificial model to outperform the humans on their original objectives. In this work, we study the phenomenon of *transcendence*: when a generative model achieves capabilities that surpass the abilities of the experts generating its data. We demonstrate transcendence by training an autoregressive transformer to play chess from game transcripts, and show that the trained model can sometimes achieve better performance than all players in the dataset. We theoretically prove that transcendence is enabled by low-temperature sampling, and rigorously assess this experimentally. Finally, we discuss other sources of transcendence, laying the groundwork for future investigation of this phenomenon in a broader setting.
Edwin Zhang, Vincent Zhu, Naomi Saphra, Anat Kleiman, Benjamin L. Edelman, Milind Tambe, Sham M. Kakade, Eran Malach
NeurIPS3
2023 Linear Connectivity Reveals Generalization Strategies
Jeevesh Juneja, Rachit Bansal, Kyunghyun Cho, João Sedoc, Naomi Saphra
ICLR5
2022 Benchmarking Compositionality with Formal Languages
abstract
Recombining known primitive concepts into larger novel combinations is a quintessentially human cognitive capability. Whether large neural models in NLP acquire this ability while learning from data is an open question. In this paper, we look at this problem from the perspective of formal languages. We use deterministic finite-state transducers to make an unbounded number of datasets with controllable properties governing compositionality. By randomly sampling over many transducers, we explore which of their properties (number of states, alphabet size, number of transitions etc.) contribute to learnability of a compositional relation by a neural network. In general, we find that the models either learn the relations completely or not at all. The key is transition coverage, setting a soft learnability limit at 400 examples per transition.
Josef Valvoda, Naomi Saphra, Jonathan Rawski, Adina Williams, Ryan Cotterell
COLING2
2022 The MultiBERTs: BERT Reproductions for Robustness Analysis
Thibault Sellam, Steve Yadlowsky, Ian Tenney, Jason Wei, Naomi Saphra, Alexander D'Amour, Tal Linzen, Jasmijn Bastings, Iulia Turc, Jacob Eisenstein, Dipanjan Das 0001, Ellie Pavlick
ICLR5
2021 A Non-Linear Structural Probe
abstract
Probes are models devised to investigate the encoding of knowledge—e.g. syntactic structure—in contextual representations. Probes are often designed for simplicity, which has led to restrictions on probe design that may not allow for the full exploitation of the structure of encoded information; one such restriction is linearity. We examine the case of a structural probe (Hewitt and Manning, 2019), which aims to investigate the encoding of syntactic structure in contextual representations through learning only linear transformations. By observing that the structural probe learns a metric, we are able to kernelize it and develop a novel non-linear variant with an identical number of parameters. We test on 6 languages and find that the radial-basis function (RBF) kernel, in conjunction with regularization, achieves a statistically significant improvement over the baseline in all languages—implying that at least part of the syntactic knowledge is encoded non-linearly. We conclude by discussing how the RBF kernel resembles BERT’s self-attention layers and speculate that this resemblance leads to the RBF-based probe’s stronger performance.
Jennifer C. White, Tiago Pimentel, Naomi Saphra, Ryan Cotterell
NAACL-HLT3
2020 Understanding Privacy-Related Questions on Stack Overflow
abstract
We analyse Stack Overflow (SO) to understand challenges and confusions developers face while dealing with privacy-related topics. We apply topic modelling techniques to 1,733 privacy-related questions to identify topics and then qualitatively analyse a random sample of 315 privacy-related questions. Identified topics include privacy policies, privacy concerns, access control, and version changes. Results show that developers do ask SO for support on privacy-related issues. We also find that platforms such as Apple and Google are defining privacy requirements for developers by specifying what "sensitive" information is and what types of information developers need to communicate to users (e.g. privacy policies). We also examine the accepted answers in our sample and find that 28% of them link to official documentation and more than half are answered by SO users without references to any external resources.
Mohammad Tahaei, Kami Vaniea, Naomi Saphra
CHI3
2020 Pareto Probing: Trading Off Accuracy for Complexity
abstract
The question of how to probe contextual word representations for linguistic structure in a way that is both principled and useful has seen significant attention recently in the NLP literature.In our contribution to this discussion, we argue for a probe metric that reflects the fundamental trade-off between probe complexity and performance: the Pareto hypervolume.To measure complexity, we present a number of parametric and non-parametric metrics.Our experiments using Pareto hypervolume as an evaluation metric show that probes often do not conform to our expectations-e.g., why should the non-contextual fastText representations encode more morpho-syntactic information than the contextual BERT representations?These results suggest that common, simplistic probing tasks, such as part-of-speech labeling and dependency arc labeling, are inadequate to evaluate the linguistic structure encoded in contextual word representations.This leads us to propose full dependency parsing as a probing task.In support of our suggestion that harder probing tasks are necessary, our experiments with dependency parsing reveal a wide gap in syntactic knowledge between contextual and non-contextual representations.Our code can be found at https://github. com/rycolab/pareto-probing.
Tiago Pimentel, Naomi Saphra, Adina Williams, Ryan Cotterell
EMNLP (1)2
2015 AMRICA: an AMR Inspector for Cross-language Alignments
abstract
Meaning Representation (AMR), an annotation scheme for natural language semantics, has drawn attention for its simplicity and representational power.Because AMR annotations are not designed for human readability, we present AMRICA, a visual aid for exploration of AMR annotations.AMRICA can visualize an AMR or the difference between two AMRs to help users diagnose interannotator disagreement or errors from an AMR parser.AMRICA can also automatically align and visualize the AMRs of a sentence and its translation in a parallel text.We believe AMRICA will simplify and streamline exploratory research on cross-lingual AMR corpora.
Naomi Saphra, Adam Lopez
HLT-NAACL1
2014 Understanding Objects in Detail with Fine-Grained Attributes
abstract
We study the problem of understanding objects in detail, intended as recognizing a wide array of fine-grained object attributes. To this end, we introduce a dataset of 7, 413 airplanes annotated in detail with parts and their attributes, leveraging images donated by airplane spotters and crowd-sourcing both the design and collection of the detailed annotations. We provide a number of insights that should help researchers interested in designing fine-grained datasets for other basic level categories. We show that the collected data can be used to study the relation between part detection and attribute prediction by diagnosing the performance of classifiers that pool information from different parts of an object. We note that the prediction of certain attributes can benefit substantially from accurate part detection. We also show that, differently from previous results in object detection, employing a large number of part templates can improve detection accuracy at the expenses of detection speed. We finally propose a coarse-to-fine approach to speed up detection through a hierarchical cascade algorithm.
Andrea Vedaldi, Siddharth Mahendran, Stavros Tsogkas, Subhransu Maji, Ross B. Girshick, Juho Kannala, Esa Rahtu, Iasonas Kokkinos, Matthew B. Blaschko, David J. Weiss, Ben Taskar, Karen Simonyan, Naomi Saphra, Sammy Mohamed
CVPR13