Adina Williams

dblp:199/2104 · DBLP profile ↗
← Back
41ranked-venue papers
5as first author
27since 2021 · last 2025
0000-0001-5281-3343ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 41 · 5 first-author · 27 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
YearPublicationVenuePosition
2025 On the Role of Speech Data in Reducing Toxicity Detection Bias
abstract
Samuel Bell, Mariano Coria Meglioli, Megan Richards, Eduardo Sánchez, Christophe Ropers, Skyler Wang, Adina Williams, Levent Sagun, Marta R. Costa-jussà. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Samuel J. Bell, Mariano Coria Meglioli, Megan Richards, Eduardo Sánchez, Christophe Ropers, Skyler Wang, Adina Williams, Levent Sagun, Marta R. Costa-jussà
NAACL (Long Papers)7
2025 Improving Model Evaluation using SMART Filtering of Benchmark Datasets
abstract
Vipul Gupta, Candace Ross, David Pantoja, Rebecca J. Passonneau, Megan Ung, Adina Williams. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Candace Ross, David Pantoja, Rebecca J. Passonneau, Megan Ung, Adina Williams
NAACL (Long Papers)6
2025 Do different prompting methods yield a common task representation in language models?
abstract
Demonstrations and instructions are two primary approaches for prompting language models to perform in-context learning (ICL) tasks. Do identical tasks elicited in different ways result in similar representations of the task? An improved understanding of task representation mechanisms would offer interpretability insights and may aid in steering models. We study this through function vectors (FVs), recently proposed as a mechanism to extract few-shot ICL task representations. We generalize FVs to alternative task presentations, focusing on short textual instruction prompts, and successfully extract instruction function vectors that promote zero-shot task accuracy. We find evidence that demonstration- and instruction-based function vectors leverage different model components, and offer several controls to dissociate their contributions to task performance. Our results suggest that different task prompting forms do not induce a common task representation through FVs but elicit different, partly overlapping mechanisms. Our findings offer principled support to the practice of combining instructions and task demonstrations, imply challenges in universally monitoring task inference across presentation forms, and encourage further examinations of LLM task inference mechanisms.
Guy Davidson, Todd M. Gureckis, Brenden M. Lake, Adina Williams
NeurIPS4
2025 What's in Common? Multimodal Models Hallucinate When Reasoning Across Scenes
abstract
Multimodal language models possess a remarkable ability to handle an open-vocabulary worth of objects. Yet the best models still suffer from hallucinations when reasoning about scenes in the real world, revealing a gap between their seemingly strong performance on existing perception benchmarks that are saturating and their reasoning in the real world. To address this gap, we build a novel benchmark of in-the-wild scenes that we call Common-O Bench with more than 10.5k examples using exclusively new images not found in web training data to avoid contamination, Common-O goes beyond just perception, inspired by cognitive tests for humans, to probe reasoning across scenes by asking ``what’s in common?''. We evaluate leading multimodal language models, including models specifically trained to reason. We find that perceiving objects in single images is easy for most models, yet reasoning across scenes is very challenging even for the best models, including reasoning models. Despite saturating many leaderboards focusing on perception, the best performing model only achieves 35\% on Common-O Bench---and on Common-O Complex, consisting of more complex scenes, the best model achieves only 1\%. Curiously, we find models are more prone to hallucinate when similar objects are present in the scene, suggesting models may be relying on object co-occurrence seen during training. Among the models we evaluated, we found scale can provide modest improvements while models explicitly trained with multi-image inputs show bigger improvements, suggesting scaled multi-image training may offer promise. We make our benchmark publicly available to spur research into the challenge of hallucination when reasoning across scenes.
Candace Ross, Florian Bordes, Adina Williams, Polina Kirichenko, Mark Ibrahim
NeurIPS3
2024 Insights from the first BabyLM Challenge: Training sample-efficient language models on a developmentally plausible corpus
Alex Warstadt, Aaron Mueller, Leshem Choshen, Ethan Wilcox, Chengxu Zhuang, Adina Williams, Ryan Cotterell, Tal Linzen
CogSci6
2024 Compositional learning of functions in humans and machines
Yanli Zhou, Brenden M. Lake, Adina Williams
CogSci3
2024 EmphAssess : a Prosodic Benchmark on Assessing Emphasis Transfer in Speech-to-Speech Models
abstract
We introduce EmphAssess, a prosodic benchmark designed to evaluate the capability of speech-to-speech models to encode and reproduce prosodic emphasis.We apply this to two tasks: speech resynthesis and speech-to-speech translation.In both cases, the benchmark evaluates the ability of the model to encode emphasis in the speech input and accurately reproduce it in the output, potentially across a change of speaker and language.As part of the evaluation pipeline, we introduce EmphaClass, a new model that classifies emphasis at the frame or word level.
Maureen de Seyssel, Antony D'Avirro, Adina Williams, Emmanuel Dupoux
EMNLP3
2024 The PRISM Alignment Dataset: What Participatory, Representative and Individualised Human Feedback Reveals About the Subjective and Multicultural Alignment of Large Language Models
abstract
Human feedback is central to the alignment of Large Language Models (LLMs). However, open questions remain about the methods (how), domains (where), people (who) and objectives (to what end) of feedback processes. To navigate these questions, we introduce PRISM, a new dataset which maps the sociodemographics and stated preferences of 1,500 diverse participants from 75 countries, to their contextual preferences and fine-grained feedback in 8,011 live conversations with 21 LLMs. With PRISM, we contribute (i) wider geographic and demographic participation in feedback; (ii) census-representative samples for two countries (UK, US); and (iii) individualised ratings that link to detailed participant profiles, permitting personalisation and attribution of sample artefacts. We target subjective and multicultural perspectives on value-laden and controversial issues, where we expect interpersonal and cross-cultural disagreement. We use PRISM in three case studies to demonstrate the need for careful consideration of which humans provide alignment data.
Hannah Kirk, Alexander Whitefield, Paul Röttger, Andrew M. Bean 0001, Aikaterini Margatina, Rafael Mosquera, Juan Ciro, Max Bartolo, Adina Williams, He He 0001, Bertie Vidgen, Scott A. Hale
NeurIPS9
2024 The Factorization Curse: Which Tokens You Predict Underlie the Reversal Curse and More
abstract
Today's best language models still struggle with "hallucinations", factually incorrect generations, which impede their ability to reliably retrieve information seen during training. The *reversal curse*, where models cannot recall information when probed in a different order than was encountered during training, exemplifies limitations in information retrieval. To better understand these limitations, we reframe the reversal curse as a *factorization curse* --- a failure of models to learn the same joint distribution under different factorizations. We more closely simulate finetuning workflows which train pretrained models on specialized knowledge by introducing *WikiReversal*, a realistic testbed based on Wikipedia knowledge graphs. Through a series of controlled experiments with increasing levels of realism, including non-reciprocal relations, we find that reliable information retrieval is an inherent failure of the next-token prediction objective used in popular large language models. Moreover, we demonstrate reliable information retrieval cannot be solved with scale, reversed tokens, or even naive bidirectional-attention training. Consequently, various approaches to finetuning on specialized data would necessarily provide mixed results on downstream tasks, unless the model has already seen the right sequence of tokens. Across five tasks of varying levels of complexity, our results uncover a promising path forward: factorization-agnostic objectives can significantly mitigate the reversal curse and hint at improved knowledge storage and planning capabilities.
Ouail Kitouni, Niklas Nolte, Adina Williams, Michael G. Rabbat, Diane Bouchacourt, Mark Ibrahim
NeurIPS3
2023 A Latent-Variable Model for Intrinsic Probing
abstract
The success of pre-trained contextualized representations has prompted researchers to analyze them for the presence of linguistic information. Indeed, it is natural to assume that these pre-trained representations do encode some level of linguistic knowledge as they have brought about large empirical improvements on a wide variety of NLP tasks, which suggests they are learning true linguistic generalization. In this work, we focus on intrinsic probing, an analysis technique where the goal is not only to identify whether a representation encodes a linguistic attribute but also to pinpoint where this attribute is encoded. We propose a novel latent-variable formulation for constructing intrinsic probes and derive a tractable variational approximation to the log-likelihood. Our results show that our model is versatile and yields tighter mutual information estimates than two intrinsic probes previously proposed in the literature. Finally, we find empirical evidence that pre-trained representations develop a cross-lingually entangled notion of morphosyntax.
Karolina Stanczak, Lucas Torroba Hennigen, Adina Williams, Ryan Cotterell, Isabelle Augenstein
AAAI3
2023 Language model acceptability judgements are not always robust to context
abstract
Koustuv Sinha, Jon Gauthier, Aaron Mueller, Kanishka Misra, Keren Fuentes, Roger Levy, Adina Williams. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Koustuv Sinha, Jon Gauthier, Aaron Mueller, Kanishka Misra, Keren Fuentes, Roger Levy, Adina Williams
ACL (1)7
2023 The Validity of Evaluation Results: Assessing Concurrence Across Compositionality Benchmarks
abstract
NLP models have progressed drastically in recent years, according to numerous datasets proposed to evaluate performance.Questions remain, however, about how particular dataset design choices may impact the conclusions we draw about model capabilities.In this work, we investigate this question in the domain of compositional generalization.We examine the performance of six modeling approaches across 4 datasets, split according to 8 compositional splitting strategies, ranking models by 18 compositional generalization splits in total.Our results show that: i) the datasets, although all designed to evaluate compositional generalization, rank modeling approaches differently; ii) datasets generated by humans align better with each other than they with synthetic datasets, or than synthetic datasets among themselves; iii) generally, whether datasets are sampled from the same source is more predictive of the resulting model ranking than whether they maintain the same interpretation of compositionality; and iv) which lexical items are used in the data can strongly impact conclusions.Overall, our results demonstrate that much work remains to be done when it comes to assessing whether popular evaluation datasets measure what they intend to measure, and suggests that elucidating more rigorous standards for establishing the validity of evaluation sets could benefit the field. 1
Kaiser Sun, Adina Williams, Dieuwke Hupkes
CoNLL2
2023 ROBBIE: Robust Bias Evaluation of Large Generative Language Models
abstract
David Esiobu, Xiaoqing Tan, Saghar Hosseini, Megan Ung, Yuchen Zhang, Jude Fernandes, Jane Dwivedi-Yu, Eleonora Presani, Adina Williams, Eric Smith. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023.
David Esiobu, Xiaoqing Ellen Tan, Saghar Hosseini, Megan Ung, Jude Fernandes, Jane Dwivedi-Yu, Eleonora Presani, Adina Williams, Eric Michael Smith
EMNLP9
2023 DataPerf: Benchmarks for Data-Centric AI Development
abstract
Machine learning research has long focused on models rather than datasets, and prominent datasets are used for common ML tasks without regard to the breadth, difficulty, and faithfulness of the underlying problems. Neglecting the fundamental importance of data has given rise to inaccuracy, bias, and fragility in real-world applications, and research is hindered by saturation across existing dataset benchmarks. In response, we present DataPerf, a community-led benchmark suite for evaluating ML datasets and data-centric algorithms. We aim to foster innovation in data-centric AI through competition, comparability, and reproducibility. We enable the ML community to iterate on datasets, instead of just architectures, and we provide an open, online platform with multiple rounds of challenges to support this iterative development. The first iteration of DataPerf contains five benchmarks covering a wide spectrum of data-centric techniques, tasks, and modalities in vision, speech, acquisition, debugging, and diffusion prompting, and we support hosting new contributed benchmarks from the community. The benchmarks, online evaluation platform, and baseline implementations are open source, and the MLCommons Association will maintain DataPerf to ensure long-term benefits to academia and industry.
Mark Mazumder, Colby R. Banbury, Xiaozhe Yao, Bojan Karlas, William Gaviria Rojas, Sudnya Frederick Diamos, Gregory Frederick Diamos, Lynn He, Alicia Parrish, Hannah Kirk, Jessica Quaye, Charvi Rastogi, Douwe Kiela, David Jurado, David Kanter, Rafael Mosquera, Will Cukierski, Juan Ciro, Lora Aroyo, Bilge Acun, Lingjiao Chen, Mehul Raje, Max Bartolo, Sabri Eyuboglu, Amirata Ghorbani, Emmett D. Goodman, Addison Howard, Oana Inel, Tariq Kane, Christine R. Kirkpatrick, D. Sculley, Tzu-Sheng Kuo, Jonas Mueller 0001, Tristan Thrush, Joaquin Vanschoren, Margaret Warren, Adina Williams, Serena Yeung-Levy, Newsha Ardalani, Praveen K. Paritosh, Ce Zhang 0001, James Zou 0001, Carole-Jean Wu, Cody Coleman, Andrew Y. Ng, Peter Mattson, Vijay Janapa Reddi
NeurIPS37
2022 Investigating Failures of Automatic Translationin the Case of Unambiguous Gender
abstract
Transformer-based models are the modern work horses for neural machine translation (NMT), reaching state of the art across several benchmarks.Despite their impressive accuracy, we observe a systemic and rudimentary class of errors made by current state-of-the-art NMT models with regards to translating from a language that doesn't mark gender on nouns into others that do.We find that even when the surrounding context provides unambiguous evidence of the appropriate grammatical gender marking, no tested model was able to accurately gender occupation nouns systematically.We release an evaluation scheme and dataset for measuring the ability of NMT models to translate gender morphology correctly in unambiguous contexts across syntactically diverse sentences.Our dataset translates from an English source into 20 languages from several different language families.With the availability of this dataset, our hope is that the NMT community can iterate on solutions for this class of especially egregious errors. Source/Target LabelSrc: My sister is a carpenter 4 .Correct Tgt: Mi hermana es carpenteria(f) 4 .Src: That nurse 1 is a funny man .Wrong Tgt: Esa enfermera(f) 1 es un tipo gracioso .Src: The engineer 1 is her emotional mother .Inconclusive Tgt: La ingeniería(?) 1 es su madre emocional .
Adithya Renduchintala, Adina Williams
ACL (1)2
2022 Evaluating locality in NMT models
Itay Itzhak, Koustuv Sinha, Brenden M. Lake, Adina Williams, Dieuwke Hupkes
CogSci4
2022 Benchmarking Compositionality with Formal Languages
abstract
Recombining known primitive concepts into larger novel combinations is a quintessentially human cognitive capability. Whether large neural models in NLP acquire this ability while learning from data is an open question. In this paper, we look at this problem from the perspective of formal languages. We use deterministic finite-state transducers to make an unbounded number of datasets with controllable properties governing compositionality. By randomly sampling over many transducers, we explore which of their properties (number of states, alphabet size, number of transitions etc.) contribute to learnability of a compositional relation by a neural network. In general, we find that the models either learn the relations completely or not at all. The key is transition coverage, setting a soft learnability limit at 400 examples per transition.
Josef Valvoda, Naomi Saphra, Jonathan Rawski, Adina Williams, Ryan Cotterell
COLING4
2022 Winoground: Probing Vision and Language Models for Visio-Linguistic Compositionality
abstract
We present a novel task and dataset for evaluating the ability of vision and language models to conduct visio-linguistic compositional reasoning, which we call Winoground. Given two images and two captions, the goal is to match them correctly-but crucially, both captions contain a completely identical set of words, only in a different order. The dataset was carefully hand-curated by expert annotators and is labeled with a rich set offine-grained tags to assist in analyzing model performance. We probe a diverse range of state-of-the-art vision and language models and find that, surprisingly, none of them do much better than chance. Evidently, these models are not as skilled at visio-linguistic compositional reasoning as we might have hoped. We perform an extensive analysis to obtain insights into how future work might try to mitigate these models' shortcomings. We aim for Winoground to serve as a useful evaluation set for advancing the state of the art and driving further progress in the field. The dataset is available at https://huggingface.co/datasets/facebook/winoground.
Tristan Thrush, Ryan Jiang, Max Bartolo, Amanpreet Singh, Adina Williams, Douwe Kiela, Candace Ross
CVPR5
2022 Perturbation Augmentation for Fairer NLP
abstract
Unwanted and often harmful social biases are becoming ever more salient in NLP research, affecting both models and datasets.In this work, we ask whether training on demographically perturbed data leads to fairer language models.We collect a large dataset of human annotated text perturbations and train a neural perturbation model, which we show outperforms heuristic alternatives.We find that (i) language models (LMs) pre-trained on demographically perturbed corpora are typically more fair, and (ii) LMs finetuned on perturbed GLUE datasets exhibit less demographic bias on downstream tasks, and (iii) fairness improvements do not come at the expense of performance on downstream tasks.Lastly, we discuss outstanding questions about how best to evaluate the (un)fairness of large language models.We hope that this exploration of neural demographic perturbation will help drive more improvement towards fairer NLP.
Rebecca Qian, Candace Ross, Jude Fernandes, Eric Michael Smith, Douwe Kiela, Adina Williams
EMNLP6
2022 "I'm sorry to hear that": Finding New Biases in Language Models with a Holistic Descriptor Dataset
abstract
As language models grow in popularity, it becomes increasingly important to clearly measure all possible markers of demographic identity in order to avoid perpetuating existing societal harms.Many datasets for measuring bias currently exist, but they are restricted in their coverage of demographic axes and are commonly used with preset bias tests that presuppose which types of biases models can exhibit.In this work, we present a new, more inclusive bias measurement dataset, HOLIS-TICBIAS, which includes nearly 600 descriptor terms across 13 different demographic axes.HOLISTICBIAS was assembled in a participatory process including experts and community members with lived experience of these terms.These descriptors combine with a set of bias measurement templates to produce over 450,000 unique sentence prompts, which we use to explore, identify, and reduce novel forms of bias in several generative models.We demonstrate that HOLISTICBIAS is effective at measuring previously undetectable biases in token likelihoods from language models, as well as in an offensiveness classifier.We will invite additions and amendments to the dataset, which we hope will serve as a basis for more easy-to-use and standardized methods for evaluating bias in NLP models.
Eric Michael Smith, Melissa Hall, Melanie Kambadur, Eleonora Presani, Adina Williams
EMNLP5
2022 On the Machine Learning of Ethical Judgments from Natural Language
abstract
Zeerak Talat, Hagen Blix, Josef Valvoda, Maya Indira Ganesh, Ryan Cotterell, Adina Williams. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Zeerak Talat, Hagen Blix, Josef Valvoda, Maya Indira Ganesh, Ryan Cotterell, Adina Williams
NAACL-HLT6
2021 UnNatural Language Inference
abstract
Koustuv Sinha, Prasanna Parthasarathi, Joelle Pineau, Adina Williams. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Koustuv Sinha, Prasanna Parthasarathi, Joelle Pineau, Adina Williams
ACL/IJCNLP (1)4
2021 Generalising to German Plural Noun Classes, from the Perspective of a Recurrent Neural Network
abstract
Inflectional morphology has since long been a useful testing ground for broader questions about generalisation in language and the viability of neural network models as cognitive models of language.Here, in line with that tradition, we explore how recurrent neural networks acquire the complex German plural system and reflect upon how their strategy compares to human generalisation and rule-based models of this system.We perform analyses including behavioural experiments, diagnostic classification, representation analysis and causal interventions, suggesting that the models rely on features that are also key predictors in rule-based models of German plurals.However, the models also display shortcut learning, which is crucial to overcome in search of more cognitively plausible generalisation behaviour.
Verna Dankers, Anna Langedijk, Kate McCurdy, Adina Williams, Dieuwke Hupkes
CoNLL4
2021 Masked Language Modeling and the Distributional Hypothesis: Order Word Matters Pre-training for Little
abstract
A possible explanation for the impressive performance of masked language model (MLM) pre-training is that such models have learned to represent the syntactic structures prevalent in classical NLP pipelines.In this paper, we propose a different explanation: MLMs succeed on downstream tasks mostly due to their ability to model higher-order word cooccurrence statistics.To demonstrate this, we pre-train MLMs on sentences with randomly shuffled word order, and we show that these models still achieve high accuracy after finetuning on many downstream tasks -including tasks specifically designed to be challenging for models that ignore word order.Our models also perform surprisingly well according to some parametric syntactic probes, indicating possible deficiencies in how we test representations for syntactic information.Overall, our results show that purely distributional information largely explains the success of pretraining, and they underscore the importance of curating challenging evaluation datasets that require deeper linguistic knowledge.
Koustuv Sinha, Robin Jia, Dieuwke Hupkes, Joelle Pineau, Adina Williams, Douwe Kiela
EMNLP (1)5
2021 Dynabench: Rethinking Benchmarking in NLP
abstract
Douwe Kiela, Max Bartolo, Yixin Nie, Divyansh Kaushik, Atticus Geiger, Zhengxuan Wu, Bertie Vidgen, Grusha Prasad, Amanpreet Singh, Pratik Ringshia, Zhiyi Ma, Tristan Thrush, Sebastian Riedel, Zeerak Waseem, Pontus Stenetorp, Robin Jia, Mohit Bansal, Christopher Potts, Adina Williams. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Douwe Kiela, Max Bartolo, Yixin Nie, Divyansh Kaushik, Atticus Geiger, Zhengxuan Wu, Bertie Vidgen, Grusha Prasad, Amanpreet Singh, Pratik Ringshia, Zhiyi Ma, Tristan Thrush, Sebastian Riedel 0001, Zeerak Talat, Pontus Stenetorp, Robin Jia, Mohit Bansal, Christopher Potts, Adina Williams
NAACL-HLT19
2021 Dynaboard: An Evaluation-As-A-Service Platform for Holistic Next-Generation Benchmarking
abstract
We introduce Dynaboard, an evaluation-as-a-service framework for hosting benchmarks and conducting holistic model comparison, integrated with the Dynabench platform. Our platform evaluates NLP models directly instead of relying on self-reported metrics or predictions on a single dataset. Under this paradigm, models are submitted to be evaluated in the cloud, circumventing the issues of reproducibility, accessibility, and backwards compatibility that often hinder benchmarking in NLP. This allows users to interact with uploaded models in real time to assess their quality, and permits the collection of additional metrics such as memory use, throughput, and robustness, which -- despite their importance to practitioners -- have traditionally been absent from leaderboards. On each task, models are ranked according to the Dynascore, a novel utility-based aggregation of these statistics, which users can customize to better reflect their preferences, placing more/less weight on a particular axis of evaluation or dataset. As state-of-the-art NLP models push the limits of traditional benchmarks, Dynaboard offers a standardized solution for a more diverse and comprehensive evaluation of model quality.
Zhiyi Ma, Kawin Ethayarajh, Tristan Thrush, Somya Jain, Ledell Wu, Robin Jia, Christopher Potts, Adina Williams, Douwe Kiela
NeurIPS8
2021 On the Relationships Between the Grammatical Genders of Inanimate Nouns and Their Co-Occurring Adjectives and Verbs
abstract
Abstract We use large-scale corpora in six different gendered languages, along with tools from NLP and information theory, to test whether there is a relationship between the grammatical genders of inanimate nouns and the adjectives used to describe those nouns. For all six languages, we find that there is a statistically significant relationship. We also find that there are statistically significant relationships between the grammatical genders of inanimate nouns and the verbs that take those nouns as direct objects, as indirect objects, and as subjects. We defer deeper investigation of these relationships for future work.
Adina Williams, Lawrence Wolf-Sonkin, Damián E. Blasi, Hanna M. Wallach, Ryan Cotterell
Trans. Assoc. Comput. Linguistics1
2020 Are Natural Language Inference Models IMPPRESsive? Learning IMPlicature and PRESupposition
abstract
Natural language inference (NLI) is an increasingly important task for natural language understanding, which requires one to infer whether a sentence entails another.However, the ability of NLI models to make pragmatic inferences remains understudied.We create an IMPlicature and PRESupposition diagnostic dataset (IMPPRES), consisting of >25k semiautomatically generated sentence pairs illustrating well-studied pragmatic inference types.We use IMPPRES to evaluate whether BERT, InferSent, and BOW NLI models trained on MultiNLI (Williams et al., 2018) learn to make pragmatic inferences.Although MultiNLI appears to contain very few pairs illustrating these inference types, we find that BERT learns to draw pragmatic inferences.It reliably treats scalar implicatures triggered by "some" as entailments.For some presupposition triggers like only, BERT reliably recognizes the presupposition as an entailment, even when the trigger is embedded under an entailment canceling operator like negation.BOW and InferSent show weaker evidence of pragmatic reasoning.We conclude that NLI training encourages models to learn some, but not all, pragmatic inferences.Type Example Trigger Jo's cat yawned.Presupposition Jo has a cat.Negated Trigger Jo's cat didn't yawn.Modal Trigger It's possible that Jo's cat yawned.Interrog.Trigger Did Jo's cat yawn?Cond.Trigger If Jo's cat yawned, it's OK.Negated Prsp.Jo doesn't have a cat.Neutral Prsp.Amy has a cat.
Paloma Jeretic, Alex Warstadt, Suvrat Bhooshan, Adina Williams
ACL4
2020 A Tale of a Probe and a Parser
abstract
Measuring what linguistic information is encoded in neural models of language has become popular in NLP.Researchers approach this enterprise by training "probes"supervised models designed to extract linguistic structure from another model's output.One such probe is the structural probe (Hewitt and Manning, 2019), designed to quantify the extent to which syntactic information is encoded in contextualised word representations.The structural probe has a novel design, unattested in the parsing literature, the precise benefit of which is not immediately obvious.To explore whether syntactic probes would do better to make use of existing techniques, we compare the structural probe to a more traditional parser with an identical lightweight parameterisation.The parser outperforms structural probe on UUAS in seven of nine analysed languages, often by a substantial amount (e.g. by 11.1 points in English).Under a second less common metric, however, there is the opposite trend-the structural probe outperforms the parser.This begs the question: which metric should we prefer?
Rowan Hall Maudslay, Josef Valvoda, Tiago Pimentel, Adina Williams, Ryan Cotterell
ACL4
2020 Adversarial NLI: A New Benchmark for Natural Language Understanding
abstract
We introduce a new large-scale NLI benchmark dataset, collected via an iterative, adversarial human-and-model-in-the-loop procedure.We show that training models on this new dataset leads to state-of-the-art performance on a variety of popular NLI benchmarks, while posing a more difficult challenge with its new test set.Our analysis sheds light on the shortcomings of current state-of-theart models, and shows that non-expert annotators are successful at finding their weaknesses.The data collection method can be applied in a never-ending learning scenario, becoming a moving target for NLU, rather than a static benchmark that will quickly saturate.
Yixin Nie, Adina Williams, Emily Dinan, Mohit Bansal, Jason Weston, Douwe Kiela
ACL2
2020 Information-Theoretic Probing for Linguistic Structure
abstract
The success of neural networks on a diverse set of NLP tasks has led researchers to question how much these networks actually "know" about natural language.Probes are a natural way of assessing this.When probing, a researcher chooses a linguistic task and trains a supervised model to predict annotations in that linguistic task from the network's learned representations.If the probe does well, the researcher may conclude that the representations encode knowledge related to the task.A commonly held belief is that using simpler models as probes is better; the logic is that simpler models will identify linguistic structure, but not learn the task itself.We propose an information-theoretic operationalization of probing as estimating mutual information that contradicts this received wisdom: one should always select the highest performing probe one can, even if it is more complex, since it will result in a tighter estimate, and thus reveal more of the linguistic information inherent in the representation.The experimental portion of our paper focuses on empirically estimating the mutual information between a linguistic property and BERT, comparing these estimates to several baselines.We evaluate on a set of ten typologically diverse languages often underrepresented in NLP research-plus Englishtotalling eleven languages.Our implementation is available in https://github.com/ rycolab/info-theoretic-probing.
Tiago Pimentel, Josef Valvoda, Rowan Hall Maudslay, Ran Zmigrod, Adina Williams, Ryan Cotterell
ACL5
2020 Predicting Declension Class from Form and Meaning
abstract
The noun lexica of many natural languages are divided into several declension classes with characteristic morphological properties.Class membership is far from deterministic, but the phonological form of a noun and its meaning can often provide imperfect clues.Here, we investigate the strength of those clues.More specifically, we operationalize "strength" as measuring how much information, in bits, we can glean about declension class from knowing the form and meaning of nouns.We know that form and meaning are often also indicative of grammatical gender-which, as we quantitatively verify, can itself share information with declension class-so we also control for gender.We find for two Indo-European languages (Czech and German) that form and meaning share a significant amount of information with class (and contribute additional information beyond gender).The three-way interaction between class, form, and meaning (given gender) is also significant.Our study is important for two reasons: First, we introduce a new method that provides additional quantitative support for a classic linguistic finding that form and meaning are relevant for the classification of nouns into declensions.Second, we show not only that individual declension classes vary in the strength of their clues within a language, but also that the variations between classes vary across languages.The code is publicly available at https://github.com/ rycolab/declension-mi.
Adina Williams, Tiago Pimentel, Hagen Blix, Arya McCarthy, Eleanor Chodroff, Ryan Cotterell
ACL1
2020 Queens are Powerful too: Mitigating Gender Bias in Dialogue Generation
abstract
Models often easily learn biases present in the training data, and their predictions directly reflect this bias.We analyze gender bias in dialogue data, and examine how this bias is actually amplified in subsequent generative chit-chat dialogue models.We measure gender bias in six existing dialogue datasets, and focus on the most biased one, the multiplayer text-based fantasy adventure dataset LIGHT (Urbanek et al., 2019), as a testbed for our bias mitigation techniques.The LIGHT dataset is highly imbalanced with respect to gender, containing predominantly male characters, likely because it is entirely collected by crowdworkers and reflects common biases that exist in fantasy or medieval settings.We consider three techniques to mitigate gender bias: counterfactual data augmentation, targeted data collection, and bias controlled training.We show that our proposed techniques mitigate gender bias in LIGHT by balancing the genderedness of generated dialogue utterances and are particularly effective in combination.We quantify performance using various evaluation methods-such as quantity of gendered words, a dialogue safety classifier, and human studies-all of which show that our models generate less gendered, but equally engaging chit-chat responses.
Emily Dinan, Angela Fan, Adina Williams, Jack Urbanek, Douwe Kiela, Jason Weston
EMNLP (1)3
2020 Multi-Dimensional Gender Bias Classification
abstract
Machine learning models are trained to find patterns in data.NLP models can inadvertently learn socially undesirable patterns when training on gender biased text.In this work, we propose a novel, general framework that decomposes gender bias in text along several pragmatic and semantic dimensions: bias from the gender of the person being spoken about, bias from the gender of the person being spoken to, and bias from the gender of the speaker.Using this fine-grained framework, we automatically annotate eight large scale datasets with gender information.In addition, we collect a new, crowdsourced evaluation benchmark.Distinguishing between gender bias along multiple dimensions enables us to train better and more fine-grained gender bias classifiers.We show our classifiers are valuable for a variety of applications, like controlling for gender bias in generative models, detecting gender bias in arbitrary text, and classifying text as offensive based on its genderedness.
Emily Dinan, Angela Fan, Ledell Wu, Jason Weston, Douwe Kiela, Adina Williams
EMNLP (1)6
2020 Intrinsic Probing through Dimension Selection
abstract
Most modern NLP systems make use of pretrained contextual representations that attain astonishingly high performance on a variety of tasks.Such high performance should not be possible unless some form of linguistic structure inheres in these representations, and a wealth of research has sprung up on probing for it.In this paper, we draw a distinction between intrinsic probing, which examines how linguistic information is structured within a representation, and the extrinsic probing popular in prior work, which only argues for the presence of such information by showing that it can be successfully extracted.To enable intrinsic probing, we propose a novel framework based on a decomposable multivariate Gaussian probe that allows us to determine whether the linguistic information in word embeddings is dispersed or focal.We then probe fastText and BERT for various morphosyntactic attributes across 36 languages.We find that most attributes are reliably encoded by only a few neurons, with fastText concentrating its linguistic structure more than BERT. 1
Lucas Torroba Hennigen, Adina Williams, Ryan Cotterell
EMNLP (1)2
2020 Measuring the Similarity of Grammatical Gender Systems by Comparing Partitions
abstract
A grammatical gender system divides a lexicon into a small number of relatively fixed grammatical categories.How similar are these gender systems across languages?To quantify the similarity, we define gender systems extensionally, thereby reducing the problem of comparisons between languages' gender systems to cluster evaluation.We borrow a rich inventory of statistical tools for cluster evaluation from the field of community detection (Driver and Kroeber, 1932;Cattell, 1945), that enable us to craft novel information-theoretic metrics for measuring similarity between gender systems.We first validate our metrics, then use them to measure gender system similarity in 20 languages.Finally, we ask whether our gender system similarities alone are sufficient to reconstruct historical relationships between languages.Towards this end, we make phylogenetic predictions on the popular, but thorny, problem from historical linguistics of inducing a phylogenetic tree over extant Indo-European languages.Languages on the same branch of our phylogenetic tree are notably similar, whereas languages from separate branches are no more similar than chance.
Arya McCarthy, Adina Williams, Shijia Liu, David Yarowsky, Ryan Cotterell
EMNLP (1)2
2020 Pareto Probing: Trading Off Accuracy for Complexity
abstract
The question of how to probe contextual word representations for linguistic structure in a way that is both principled and useful has seen significant attention recently in the NLP literature.In our contribution to this discussion, we argue for a probe metric that reflects the fundamental trade-off between probe complexity and performance: the Pareto hypervolume.To measure complexity, we present a number of parametric and non-parametric metrics.Our experiments using Pareto hypervolume as an evaluation metric show that probes often do not conform to our expectations-e.g., why should the non-contextual fastText representations encode more morpho-syntactic information than the contextual BERT representations?These results suggest that common, simplistic probing tasks, such as part-of-speech labeling and dependency arc labeling, are inadequate to evaluate the linguistic structure encoded in contextual word representations.This leads us to propose full dependency parsing as a probing task.In support of our suggestion that harder probing tasks are necessary, our experiments with dependency parsing reveal a wide gap in syntactic knowledge between contextual and non-contextual representations.Our code can be found at https://github. com/rycolab/pareto-probing.
Tiago Pimentel, Naomi Saphra, Adina Williams, Ryan Cotterell
EMNLP (1)3
2019 Quantifying the Semantic Core of Gender Systems
abstract
Adina Williams, Damian Blasi, Lawrence Wolf-Sonkin, Hanna Wallach, Ryan Cotterell. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Adina Williams, Damián E. Blasi, Lawrence Wolf-Sonkin, Hanna M. Wallach, Ryan Cotterell
EMNLP/IJCNLP (1)1
2018 XNLI: Evaluating Cross-lingual Sentence Representations
abstract
Alexis Conneau, Ruty Rinott, Guillaume Lample, Adina Williams, Samuel Bowman, Holger Schwenk, Veselin Stoyanov. Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing. 2018.
Alexis Conneau, Ruty Rinott, Guillaume Lample, Adina Williams, Samuel R. Bowman, Holger Schwenk, Veselin Stoyanov
EMNLP4
2018 A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference
abstract
Adina Williams, Nikita Nangia, Samuel Bowman. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Adina Williams, Nikita Nangia, Samuel R. Bowman
NAACL-HLT1
2018 Do latent tree learning models identify meaningful structure in sentences?
abstract
Recent work on the problem of latent tree learning has made it possible to train neural networks that learn to both parse a sentence and use the resulting parse to interpret the sentence, all without exposure to ground-truth parse trees at training time. Surprisingly, these models often perform better at sentence understanding tasks than models that use parse trees from conventional parsers. This paper aims to investigate what these latent tree learning models learn. We replicate two such models in a shared codebase and find that (i) only one of these models outperforms conventional tree-structured models on sentence classification, (ii) its parsing strategies are not especially consistent across random restarts, (iii) the parses it produces tend to be shallower than standard Penn Treebank (PTB) parses, and (iv) they do not resemble those of PTB or any other semantic or syntactic formalism that the authors are aware of.
Adina Williams, Andrew Drozdov, Samuel R. Bowman
Trans. Assoc. Comput. Linguistics1