Samuel R. Bowman

dblp:116/0502 · also Sam Bowman 0001 · DBLP profile ↗
← Back
51ranked-venue papers
8as first author
20since 2021 · last 2025
0000-0001-6737-4603ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 50 · 8 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 Language Models Learn to Mislead Humans via RLHF
abstract
Language models (LMs) can produce errors that are hard to detect for humans, especially when the task is complex. RLHF, the most popular post-training method, may exacerbate this problem: to achieve higher rewards, LMs might get better at convincing humans that they are right even when they are wrong. We study this phenomenon under a standard RLHF pipeline, calling it ``U-Sophistry'' since it is \textbf{U}nintended by model developers. Specifically, we ask time-constrained (e.g., 3-10 minutes) human subjects to evaluate the correctness of model outputs and calculate humans' accuracy against gold labels. On a question-answering task (QuALITY) and programming task (APPS), RLHF makes LMs better at convincing our subjects but not at completing the task correctly. RLHF also makes the model harder to evaluate: our subjects' false positive rate increases by 24.1% on QuALITY and 18.3% on APPS. Finally, we show that probing, a state-of-the-art approach for detecting \textbf{I}ntended Sophistry (e.g.~backdoored LMs), does not generalize to U-Sophistry. Our results highlight an important failure mode of RLHF and call for more research in assisting humans to align them.
Jiaxin Wen, Ruiqi Zhong, Akbir Khan, Ethan Perez, Jacob Steinhardt, Minlie Huang, Samuel R. Bowman, He He 0001, Shi Feng 0005
ICLR7
2024 Towards Understanding Sycophancy in Language Models
abstract
Reinforcement learning from human feedback (RLHF) is a popular technique for training high-quality AI assistants. However, RLHF may also encourage model responses that match user beliefs over truthful responses, a behavior known as sycophancy. We investigate the prevalence of sycophancy in RLHF-trained models and whether human preference judgments are responsible. We first demonstrate that five state-of-the-art AI assistants consistently exhibit sycophancy behavior across four varied free-form text-generation tasks. To understand if human preferences drive this broadly observed behavior of RLHF models, we analyze existing human preference data. We find that when a response matches a user's views, it is more likely to be preferred. Moreover, both humans and preference models (PMs) prefer convincingly-written sycophantic responses over correct ones a non-negligible fraction of the time. Optimizing model outputs against PMs also sometimes sacrifices truthfulness in favor of sycophancy. Overall, our results indicate that sycophancy is a general behavior of RLHF models, likely driven in part by human preference judgments favoring sycophantic responses.
Mrinank Sharma, Meg Tong, Tomasz Korbak, David Duvenaud, Amanda Askell, Samuel R. Bowman, Esin Durmus, Zac Hatfield-Dodds, Scott R. Johnston, Shauna Kravec, Timothy Maxwell, Sam McCandlish, Kamal Ndousse, Oliver Rausch, Nicholas Schiefer, Miranda Zhang, Ethan Perez
ICLR6
2024 Debating with More Persuasive LLMs Leads to More Truthful Answers
abstract
Common methods for aligning large language models (LLMs) with desired behaviour heavily rely on human-labelled data. However, as models grow increasingly sophisticated, they will surpass human expertise, and the role of human evaluation will evolve into non-experts overseeing experts. In anticipation of this, we ask: can weaker models assess the correctness of stronger models? We investigate this question in an analogous setting, where stronger models (experts) possess the necessary information to answer questions and weaker models (non-experts) lack this information. The method we evaluate is debate, where two LLM experts each argue for a different answer, and a non-expert selects the answer. We find that debate consistently helps both non-expert models and humans answer questions, achieving 76% and 88% accuracy respectively (naive baselines obtain 48% and 60%). Furthermore, optimising expert debaters for persuasiveness in an unsupervised manner improves non-expert ability to identify the truth in debates. Our results provide encouraging empirical evidence for the viability of aligning models with debate in the absence of ground truth.
Akbir Khan, Dan Valentine, Laura Ruis, Kshitij Sachan, Ansh Radhakrishnan, Edward Grefenstette, Samuel R. Bowman, Tim Rocktäschel, Ethan Perez
ICML8
2024 Many-shot Jailbreaking
abstract
We investigate a family of simple long-context attacks on large language models: prompting with hundreds of demonstrations of undesirable behavior. This attack is newly feasible with the larger context windows recently deployed by language model providers like Google DeepMind, OpenAI and Anthropic. We find that in diverse, realistic circumstances, the effectiveness of this attack follows a power law, up to hundreds of shots. We demonstrate the success of this attack on the most widely used state-of-the-art closed-weight models, and across various tasks. Our results suggest very long contexts present a rich new attack surface for LLMs.
Cem Anil, Esin Durmus, Nina Panickssery, Mrinank Sharma, Joe Benton, Sandipan Kundu, Joshua Batson, Meg Tong, Jesse Mu, Daniel Ford, Francesco Mosconi, Rajashree Agrawal, Rylan Schaeffer, Naomi Bashkansky, Samuel Svenningsen, Mike Lambert, Ansh Radhakrishnan, Carson Denison, Evan Hubinger, Yuntao Bai, Trenton Bricken, Timothy Maxwell, Nicholas Schiefer, James Sully, Alex Tamkin, Tamera Lanham, Karina Nguyen, Tomek Korbak, Jared Kaplan, Deep Ganguli, Samuel R. Bowman, Ethan Perez, Roger B. Grosse, David Duvenaud
NeurIPS31
2024 LLM Evaluators Recognize and Favor Their Own Generations
abstract
Self-evaluation using large language models (LLMs) has proven valuable not only in benchmarking but also methods like reward modeling, constitutional AI, and self-refinement. But new biases are introduced due to the same LLM acting as both the evaluator and the evaluatee. One such bias is self-preference, where an LLM evaluator scores its own outputs higher than others’ while human annotators consider them of equal quality. But do LLMs actually recognize their own outputs when they give those texts higher scores, or is it just a coincidence? In this paper, we investigate if self-recognition capability contributes to self-preference. We discover that, out of the box, LLMs such as GPT-4 and Llama 2 have non-trivial accuracy at distinguishing themselves from other LLMs and humans. By finetuning LLMs, we discover a linear correlation between self-recognition capability and the strength of self-preference bias; using controlled experiments, we show that the causal explanation resists straightforward confounders. We discuss how self-recognition can interfere with unbiased evaluations and AI safety more generally.
Arjun Panickssery, Samuel R. Bowman, Shi Feng 0005
NeurIPS2
2023 Instruction Induction: From Few Examples to Natural Language Task Descriptions
abstract
Large language models are able to perform a task by conditioning on a few input-output demonstrations -a paradigm known as incontext learning.We show that language models can explicitly infer an underlying task from a few demonstrations by prompting them to generate a natural language instruction that fits the examples.To explore this ability, we introduce the instruction induction challenge, compile a dataset consisting of 24 tasks, and define a novel evaluation metric based on executing the generated instruction.We discover that, to a large extent, the ability to generate instructions does indeed emerge when using a model that is both large enough and aligned to follow instructions; InstructGPT achieves 65.7% of human performance in our execution-based metric, while the original GPT-3 model reaches only 9.8% of human performance.This surprising result suggests that instruction induction might be a viable learning paradigm in and of itself, where instead of fitting a set of latent continuous parameters to the data, one searches for the best description in the natural language hypothesis space. 1
Or Honovich, Uri Shaham 0002, Samuel R. Bowman, Omer Levy
ACL (1)3
2023 (QA)²: Question Answering with Questionable Assumptions
abstract
Naturally occurring information-seeking questions often contain questionable assumptions-assumptions that are false or unverifiable.Questions containing questionable assumptions are challenging because they require a distinct answer strategy that deviates from typical answers for information-seeking questions.For instance, the question When did Marie Curie discover Uranium?cannot be answered as a typical when question without addressing the false assumption Marie Curie discovered Uranium.In this work, we propose (QA) 2 (Question Answering with Questionable Assumptions), an open-domain evaluation dataset consisting of naturally occurring search engine queries that may or may not contain questionable assumptions.To be successful on (QA) 2 , systems must be able to detect questionable assumptions and also be able to produce adequate responses for both typical information-seeking questions and ones with questionable assumptions.Through human rater acceptability on end-to-end QA with (QA) 2 , we find that current models do struggle with handling questionable assumptions, leaving substantial headroom for progress.* Equal contribution, corresponding authors ∆ Work partly done at NYU before joining BU. δ Work done at NYU before joining Amazon. 1 We use the term questionable assumptions instead of presupposition failure to capture failures of both true presup-
Najoung Kim, Phu Mon Htut, Samuel R. Bowman, Jackson Petty
ACL (1)3
2023 What Do NLP Researchers Believe? Results of the NLP Community Metasurvey
abstract
Julian Michael, Ari Holtzman, Alicia Parrish, Aaron Mueller, Alex Wang, Angelica Chen, Divyam Madaan, Nikita Nangia, Richard Yuanzhe Pang, Jason Phang, Samuel R. Bowman. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Julian Michael, Ari Holtzman, Alicia Parrish, Aaron Mueller, Angelica Chen, Divyam Madaan, Nikita Nangia, Richard Yuanzhe Pang, Jason Phang, Samuel R. Bowman
ACL (1)11
2023 Pretraining Language Models with Human Preferences
abstract
Language models (LMs) are pretrained to imitate text from large and diverse datasets that contain content that would violate human preferences if generated by an LM: falsehoods, offensive comments, personally identifiable information, low-quality or buggy code, among others. Here, we explore alternative objectives for pretraining LMs in a way that also guides them to generate text aligned with human preferences. We benchmark five objectives for pretraining with human feedback across three tasks and study how they affect the alignment and capabilities of pretrained LMs. We find a Pareto-optimal and simple approach among those we explored: conditional training, or learning distribution over tokens conditional on their human preference scores. Conditional training reduces the rate of undesirable content by up to an order of magnitude, both when generating without a prompt and with an adversarially-chosen prompt. Moreover, conditional training maintains the downstream task performance of standard LM pretraining, both before and after task-specific finetuning. Pretraining with human feedback results in much better preference satisfaction than standard LM pretraining followed by finetuning with feedback, i.e., learning and then unlearning undesirable behavior. Our results suggest that we should move beyond imitation learning when pretraining LMs and incorporate human preferences from the start of training.
Tomasz Korbak, Kejian Shi, Angelica Chen, Rasika Bhalerao, Christopher L. Buckley, Jason Phang, Samuel R. Bowman, Ethan Perez
ICML7
2023 Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting
abstract
Large Language Models (LLMs) can achieve strong performance on many tasks by producing step-by-step reasoning before giving a final output, often referred to as chain-of-thought reasoning (CoT). It is tempting to interpret these CoT explanations as the LLM's process for solving a task. This level of transparency into LLMs' predictions would yield significant safety benefits. However, we find that CoT explanations can systematically misrepresent the true reason for a model's prediction. We demonstrate that CoT explanations can be heavily influenced by adding biasing features to model inputs—e.g., by reordering the multiple-choice options in a few-shot prompt to make the answer always "(A)"—which models systematically fail to mention in their explanations. When we bias models toward incorrect answers, they frequently generate CoT explanations rationalizing those answers. This causes accuracy to drop by as much as 36% on a suite of 13 tasks from BIG-Bench Hard, when testing with GPT-3.5 from OpenAI and Claude 1.0 from Anthropic. On a social-bias task, model explanations justify giving answers in line with stereotypes without mentioning the influence of these social biases. Our findings indicate that CoT explanations can be plausible yet misleading, which risks increasing our trust in LLMs without guaranteeing their safety. Building more transparent and explainable systems will require either improving CoT faithfulness through targeted efforts or abandoning CoT in favor of alternative methods.
Miles Turpin, Julian Michael, Ethan Perez, Samuel R. Bowman
NeurIPS4
2022 The Dangers of Underclaiming: Reasons for Caution When Reporting How NLP Systems Fail
abstract
Researchers in NLP often frame and discuss research results in ways that serve to deemphasize the field's successes, often in response to the field's widespread hype.Though wellmeaning, this has yielded many misleading or false claims about the limits of our best technology.This is a problem, and it may be more serious than it looks: It harms our credibility in ways that can make it harder to mitigate present-day harms, like those involving biased systems for content moderation or resume screening.It also limits our ability to prepare for the potentially enormous impacts of more distant future advances.This paper urges researchers to be careful about these claims and suggests some research directions and communication strategies that will make it easier to avoid or rebut them. Model Year SQuAD AS AOS
Samuel R. Bowman
ACL (1)1
2022 What Makes Reading Comprehension Questions Difficult?
abstract
For a natural language understanding benchmark to be useful in research, it has to consist of examples that are diverse and difficult enough to discriminate among current and near-future state-of-the-art systems.However, we do not yet know how best to select text sources to collect a variety of challenging examples.In this study, we crowdsource multiple-choice reading comprehension questions for passages taken from seven qualitatively distinct sources, analyzing what attributes of passages contribute to the difficulty and question types of the collected examples.To our surprise, we find that passage source, length, and readability measures do not significantly affect question difficulty.Through our manual annotation of seven reasoning types, we observe several trends between passage sources and reasoning types, e.g., logical reasoning is more often required in questions written for technical passages.These results suggest that when creating a new benchmark dataset, selecting a diverse set of passages can help ensure a diverse range of question types, but that passage difficulty need not be a priority.
Saku Sugawara, Nikita Nangia, Alex Warstadt, Samuel R. Bowman
ACL (1)4
2022 SocioProbe: What, When, and Where Language Models Learn about Sociodemographics
abstract
Pre-trained language models (PLMs) have outperformed other NLP models on a wide range of tasks.Opting for a more thorough understanding of their capabilities and inner workings, researchers have established the extend to which they capture lower-level knowledge like grammaticality, and mid-level semantic knowledge like factual understanding.However, there is still little understanding of their knowledge of higher-level aspects of language.In particular, despite the importance of sociodemographic aspects in shaping our language, the questions of whether, where, and how PLMs encode these aspects, e.g., gender or age, is still unexplored.We address this research gap by probing the sociodemographic knowledge of different single-GPU PLMs on multiple English data sets via traditional classifier probing and information-theoretic minimum description length probing.Our results show that PLMs do encode these sociodemographics, and that this knowledge is sometimes spread across the layers of some of the tested PLMs.We further conduct a multilingual analysis and investigate the effect of supplementary training to further explore to what extent, where, and with what amount of pre-training data the knowledge is encoded.Our overall results indicate that sociodemographic knowledge is still a major challenge for NLP.PLMs require large amounts of pre-training data to acquire the knowledge and models that excel in general language understanding do not seem to own more knowledge about these aspects.
Anne Lauscher, Federico Bianchi 0001, Samuel R. Bowman, Dirk Hovy
EMNLP3
2022 SQuALITY: Building a Long-Document Summarization Dataset the Hard Way
abstract
Summarization datasets are often assembled either by scraping naturally occurring publicdomain summaries-which are nearly always in difcult-to-work-with technical domainsor by using approximate heuristics to extract them from everyday text-which frequently yields unfaithful summaries.In this work, we turn to a slower but more straightforward approach to developing summarization benchmark data: We hire highly-qualied contractors to read stories and write original summaries from scratch.To amortize reading time, we collect ve summaries per document, with the rst giving an overview and the subsequent four addressing specic questions.We use this protocol to collect SQuAL-ITY, a dataset of question-focused summaries built on the same public-domain short stories as the multiple-choice dataset QuALITY (Pang et al., 2021b).Experiments with stateof-the-art summarization systems show that our dataset is challenging and that existing automatic evaluation metrics are weak indicators of quality.
Richard Yuanzhe Pang, Angelica Chen, Jason Phang, Samuel R. Bowman
EMNLP5
2022 QuALITY: Question Answering with Long Input Texts, Yes!
abstract
Richard Yuanzhe Pang, Alicia Parrish, Nitish Joshi, Nikita Nangia, Jason Phang, Angelica Chen, Vishakh Padmakumar, Johnny Ma, Jana Thompson, He He, Samuel Bowman. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Richard Yuanzhe Pang, Alicia Parrish, Nitish Joshi, Nikita Nangia, Jason Phang, Angelica Chen, Vishakh Padmakumar, Johnny Ma, Jana Thompson, He He 0001, Samuel R. Bowman
NAACL-HLT11
2021 What Ingredients Make for an Effective Crowdsourcing Protocol for Difficult NLU Data Collection Tasks?
abstract
Nikita Nangia, Saku Sugawara, Harsh Trivedi, Alex Warstadt, Clara Vania, Samuel R. Bowman. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Nikita Nangia, Saku Sugawara, Harsh Trivedi, Alex Warstadt, Clara Vania, Samuel R. Bowman
ACL/IJCNLP (1)6
2021 Comparing Test Sets with Item Response Theory
abstract
Clara Vania, Phu Mon Htut, William Huang, Dhara Mungra, Richard Yuanzhe Pang, Jason Phang, Haokun Liu, Kyunghyun Cho, Samuel R. Bowman. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Clara Vania, Phu Mon Htut, William Huang, Dhara A. Mungra, Richard Yuanzhe Pang, Jason Phang, Haokun Liu, Kyunghyun Cho, Samuel R. Bowman
ACL/IJCNLP (1)9
2021 When Do You Need Billions of Words of Pretraining Data?
abstract
Yian Zhang, Alex Warstadt, Xiaocheng Li, Samuel R. Bowman. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Yian Zhang, Alex Warstadt, Xiaocheng Li, Samuel R. Bowman
ACL/IJCNLP (1)4
2021 NOPE: A Corpus of Naturally-Occurring Presuppositions in English
abstract
Alicia Parrish, Sebastian Schuster, Alex Warstadt, Omar Agha, Soo-Hwan Lee, Zhuoye Zhao, Samuel R. Bowman, Tal Linzen. Proceedings of the 25th Conference on Computational Natural Language Learning. 2021.
Alicia Parrish, Sebastian Schuster 0001, Alex Warstadt, Omar Agha, Soo-Hwan Lee, Zhuoye Zhao, Samuel R. Bowman, Tal Linzen
CoNLL7
2021 What Will it Take to Fix Benchmarking in Natural Language Understanding?
abstract
Evaluation for many natural language understanding (NLU) tasks is broken: Unreliable and biased systems score so highly on standard benchmarks that there is little room for researchers who develop better systems to demonstrate their improvements.The recent trend to abandon IID benchmarks in favor of adversarially-constructed, out-of-distribution test sets ensures that current models will perform poorly, but ultimately only obscures the abilities that we want our benchmarks to measure.In this position paper, we lay out four criteria that we argue NLU benchmarks should meet.We argue most current benchmarks fail at these criteria, and that adversarial data collection does not meaningfully address the causes of these failures.Instead, restoring a healthy evaluation ecosystem will require significant progress in the design of benchmark datasets, the reliability with which they are annotated, their size, and the ways they handle social bias.
Samuel R. Bowman, George E. Dahl
NAACL-HLT1
2020 Learning to Learn Morphological Inflection for Resource-Poor Languages
abstract
We propose to cast the task of morphological inflection—mapping a lemma to an indicated inflected form—for resource-poor languages as a meta-learning problem. Treating each language as a separate task, we use data from high-resource source languages to learn a set of model parameters that can serve as a strong initialization point for fine-tuning on a resource-poor target language. Experiments with two model architectures on 29 target languages from 3 families show that our suggested approach outperforms all baselines. In particular, it obtains a 31.7% higher absolute accuracy than a previously proposed cross-lingual transfer model and outperforms the previous state of the art by 1.7% absolute accuracy on average over languages.
Katharina Kann, Samuel R. Bowman, Kyunghyun Cho
AAAI2
2020 Intermediate-Task Transfer Learning with Pretrained Language Models: When and Why Does It Work?
abstract
Yada Pruksachatkun, Jason Phang, Haokun Liu, Phu Mon Htut, Xiaoyi Zhang, Richard Yuanzhe Pang, Clara Vania, Katharina Kann, Samuel R. Bowman. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020.
Yada Pruksachatkun, Jason Phang, Haokun Liu, Phu Mon Htut, Richard Yuanzhe Pang, Clara Vania, Katharina Kann, Samuel R. Bowman
ACL9
2020 Can neural networks acquire a structural bias from raw linguistic data?
Alex Warstadt, Samuel R. Bowman
CogSci2
2020 New Protocols and Negative Results for Textual Entailment Data Collection
abstract
Natural language inference (NLI) data has proven useful in benchmarking and, especially, as pretraining data for tasks requiring language understanding.However, the crowdsourcing protocol that was used to collect this data has known issues and was not explicitly optimized for either of these purposes, so it is likely far from ideal.We propose four alternative protocols, each aimed at improving either the ease with which annotators can produce sound training examples or the quality and diversity of those examples.Using these alternatives and a fifth baseline protocol, we collect and compare five new 8.5k-example training sets.In evaluations focused on transfer learning applications, our results are solidly negative, with models trained on our baseline dataset yielding good transfer performance to downstream tasks, but none of our four new methods (nor the recent ANLI) showing any improvements over that baseline.In a small silver lining, we observe that all four new protocols, especially those where annotators edit pre-filled text boxes, reduce previously observed issues with annotation artifacts. * Work done while visiting Google.Base
Samuel R. Bowman, Jennimaria Palomaki, Livio B. Soares, Emily Pitler
EMNLP (1)1
2020 Precise Task Formalization Matters in Winograd Schema Evaluations
abstract
Performance on the Winograd Schema Challenge (WSC), a respected English commonsense reasoning benchmark, recently rocketed from chance accuracy to 89% on the Super-GLUE leaderboard, with relatively little corroborating evidence of a correspondingly large improvement in reasoning ability.We hypothesize that much of this improvement comes from recent changes in task formalizationthe combination of input specification, loss function, and reuse of pretrained parametersby users of the dataset, rather than improvements in the pretrained model's reasoning ability.We perform an ablation on two Winograd Schema datasets that interpolates between the formalizations used before and after this surge, and find (i) framing the task as multiple choice improves performance by 2-6 points and (ii) several additional techniques, including the reuse of a pretrained language modeling head, can mitigate the model's extreme sensitivity to hyperparameters.We urge future benchmark creators to impose additional structure to minimize the impact of formalization decisions on reported results.
Haokun Liu, William Huang, Dhara A. Mungra, Samuel R. Bowman
EMNLP (1)4
2020 CrowS-Pairs: A Challenge Dataset for Measuring Social Biases in Masked Language Models
abstract
Warning: This paper contains explicit statements of offensive stereotypes and may be upsetting.Pretrained language models, especially masked language models (MLMs) have seen success across many NLP tasks.However, there is ample evidence that they use the cultural biases that are undoubtedly present in the corpora they are trained on, implicitly creating harm with biased representations.To measure some forms of social bias in language models against protected demographic groups in the US, we introduce the Crowdsourced Stereotype Pairs benchmark (CrowS-Pairs).CrowS-Pairs has 1508 examples that cover stereotypes dealing with nine types of bias, like race, religion, and age.In CrowS-Pairs a model is presented with two sentences: one that is more stereotyping and another that is less stereotyping.The data focuses on stereotypes about historically disadvantaged groups and contrasts them with advantaged groups.We find that all three of the widelyused MLMs we evaluate substantially favor sentences that express stereotypes in every category in CrowS-Pairs.As work on building less biased models advances, this dataset can be used as a benchmark to evaluate progress.
Nikita Nangia, Clara Vania, Rasika Bhalerao, Samuel R. Bowman
EMNLP (1)4
2020 Learning Which Features Matter: RoBERTa Acquires a Preference for Linguistic Generalizations (Eventually)
abstract
One reason pretraining on self-supervised linguistic tasks is effective is that it teaches models features that are helpful for language understanding.However, we want pretrained models to learn not only to represent linguistic features, but also to use those features preferentially during fine-turning.With this goal in mind, we introduce a new English-language diagnostic set called MSGS (the Mixed Signals Generalization Set), which consists of 20 ambiguous binary classification tasks that we use to test whether a pretrained model prefers linguistic or surface generalizations during finetuning.We pretrain RoBERTa models from scratch on quantities of data ranging from 1M to 1B words and compare their performance on MSGS to the publicly available RoBERTa BASE .We find that models can learn to represent linguistic features with little pretraining data, but require far more data to learn to prefer linguistic generalizations over surface ones.Eventually, with about 30B words of pretraining data, RoBERTa BASE does demonstrate a linguistic bias with some regularity.We conclude that while self-supervised pretraining is an effective way to learn helpful inductive biases, there is likely room to improve the rate at which models learn which features matter. Feature type Feature description Positive example Negative example SurfaceAbsolute position Is the first token of S "the"?The cat chased a mouse.A cat chased a mouse.Length Is S longer than n (e.g., 3) words?The cat chased a mouse.The cat meowed.Lexical content Does S contain "the"?That cat chased the mouse.That cat chased a mouse.Relative position Does "the" precede "a"?The cat chased a mouse.A cat chased the mouse.Orthography Does S appear in title case?The Cat Chased a Mouse.The cat chased a mouse. LinguisticMorphology Does S have an irregular past verb?The cats slept.The cats meow.Syn.category Does S have an adjective?Lincoln was tall.Lincoln was president.Syn.construction Is S the control construction?Sue is eager to sleep.Sue is likely to sleep.Syn.position Is the main verb in "ing" form?Cats who eat mice are purring.Cats who are eating mice purr.
Alex Warstadt, Yian Zhang, Xiaocheng Li, Haokun Liu, Samuel R. Bowman
EMNLP (1)5
2020 BLiMP: The Benchmark of Linguistic Minimal Pairs for English
abstract
We introduce The Benchmark of Linguistic Minimal Pairs (BLiMP),1 a challenge set for evaluating the linguistic knowledge of language models (LMs) on major grammatical phenomena in English. BLiMP consists of 67 individual datasets, each containing 1,000 minimal pairs—that is, pairs of minimally different sentences that contrast in grammatical acceptability and isolate specific phenomenon in syntax, morphology, or semantics. We generate the data according to linguist-crafted grammar templates, and human aggregate agreement with the labels is 96.4%. We evaluate n-gram, LSTM, and Transformer (GPT-2 and Transformer-XL) LMs by observing whether they assign a higher probability to the acceptable sentence in each minimal pair. We find that state-of-the-art models identify morphological contrasts related to agreement reliably, but they struggle with some subtle semantic and syntactic phenomena, such as negative polarity items and extraction islands.
Alex Warstadt, Alicia Parrish, Haokun Liu, Anhad Mohananey, Wei Peng 0013, Sheng-Fu Wang, Samuel R. Bowman
Trans. Assoc. Comput. Linguistics7
2020 Erratum: "BLiMP: The Benchmark of Linguistic Minimal Pairs for English"
abstract
We correct wrongly reported results on BLiMP.
Alex Warstadt, Alicia Parrish, Haokun Liu, Anhad Mohananey, Wei Peng 0013, Sheng-Fu Wang, Samuel R. Bowman
Trans. Assoc. Comput. Linguistics7
2019 Human vs. Muppet: A Conservative Estimate of Human Performance on the GLUE Benchmark
abstract
The GLUE benchmark (Wang et al., 2019b) is a suite of language understanding tasks which has seen dramatic progress in the past year, with average performance moving from 70.0 at launch to 83.9, state of the art at the time of writing (May 24, 2019).Here, we measure human performance on the benchmark, in order to learn whether significant headroom remains for further progress.We provide a conservative estimate of human performance on the benchmark through crowdsourcing: Our annotators are non-experts who must learn each task from a brief set of instructions and 20 examples.In spite of limited training, these annotators robustly outperform the state of the art on six of the nine GLUE tasks and achieve an average score of 87.1.Given the fast pace of progress however, the headroom we observe is quite limited.To reproduce the datapoor setting that our annotators must learn in, we also train the BERT model (Devlin et al., 2019) in limited-data regimes, and conclude that low-resource sentence classification remains a challenge for modern neural network approaches to text understanding.How do you prepare for a job interview?How do I prepare for my first job interview? 1 Table 6: Another ten randomly sampled examples from QQP's development set.Pairs of sentences with a label of 1 are marked as paraphrases in QQP.
Nikita Nangia, Samuel R. Bowman
ACL (1)2
2019 Can You Tell Me How to Get Past Sesame Street? Sentence-Level Pretraining Beyond Language Modeling
abstract
Alex Wang, Jan Hula, Patrick Xia, Raghavendra Pappagari, R. Thomas McCoy, Roma Patel, Najoung Kim, Ian Tenney, Yinghui Huang, Katherin Yu, Shuning Jin, Berlin Chen, Benjamin Van Durme, Edouard Grave, Ellie Pavlick, Samuel R. Bowman. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019.
Jan Hula, Patrick Xia 0002, Raghavendra Pappagari, Tom McCoy 0001, Roma Patel, Najoung Kim, Ian Tenney, Katherin Yu, Shuning Jin, Berlin Chen, Benjamin Van Durme, Edouard Grave, Ellie Pavlick, Samuel R. Bowman
ACL (1)16
2019 Towards Realistic Practices In Low-Resource Natural Language Processing: The Development Set
abstract
Katharina Kann, Kyunghyun Cho, Samuel R. Bowman. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Katharina Kann, Kyunghyun Cho, Samuel R. Bowman
EMNLP/IJCNLP (1)3
2019 Investigating BERT's Knowledge of Language: Five Analysis Methods with NPIs
abstract
Alex Warstadt, Yu Cao, Ioana Grosu, Wei Peng, Hagen Blix, Yining Nie, Anna Alsop, Shikha Bordia, Haokun Liu, Alicia Parrish, Sheng-Fu Wang, Jason Phang, Anhad Mohananey, Phu Mon Htut, Paloma Jeretic, Samuel R. Bowman. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Alex Warstadt, Ioana Grosu, Wei Peng 0013, Hagen Blix, Yining Nie, Anna Alsop, Shikha Bordia, Haokun Liu, Alicia Parrish, Sheng-Fu Wang, Jason Phang, Anhad Mohananey, Phu Mon Htut, Paloma Jeretic, Samuel R. Bowman
EMNLP/IJCNLP (1)16
2019 What do you learn from context? Probing for sentence structure in contextualized word representations
Ian Tenney, Patrick Xia 0002, Berlin Chen, Adam Poliak, Tom McCoy 0001, Najoung Kim, Benjamin Van Durme, Samuel R. Bowman, Dipanjan Das 0001, Ellie Pavlick
ICLR (Poster)9
2019 GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding
Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, Samuel R. Bowman
ICLR (Poster)6
2019 Can Unconditional Language Models Recover Arbitrary Sentences?
abstract
Neural network-based generative language models like ELMo and BERT can work effectively as general purpose sentence encoders in text classification without further fine-tuning. Is it possible to adapt them in a similar way for use as general-purpose decoders? For this to be possible, it would need to be the case that for any target sentence of interest, there is some continuous representation that can be passed to the language model to cause it to reproduce that sentence. We set aside the difficult problem of designing an encoder that can produce such representations and, instead, ask directly whether such representations exist at all. To do this, we introduce a pair of effective, complementary methods for feeding representations into pretrained unconditional language models and a corresponding set of methods to map sentences into and out of this representation space, the reparametrized sentence space. We then investigate the conditions under which a language model can be made to generate a sentence through the identification of a point in such a space and find that it is possible to recover arbitrary sentences nearly perfectly with language models and representations of moderate size.
Nishant Subramani, Samuel R. Bowman, Kyunghyun Cho
NeurIPS2
2019 SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems
abstract
In the last year, new models and methods for pretraining and transfer learning have driven striking performance improvements across a range of language understanding tasks. The GLUE benchmark, introduced a little over one year ago, offers a single-number metric that summarizes progress on a diverse set of such tasks, but performance on the benchmark has recently surpassed the level of non-expert humans, suggesting limited headroom for further research. In this paper we present SuperGLUE, a new benchmark styled after GLUE with a new set of more difficult language understanding tasks, a software toolkit, and a public leaderboard. SuperGLUE is available at https://super.gluebenchmark.com.
Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, Samuel R. Bowman
NeurIPS8
2019 Neural Network Acceptability Judgments
abstract
This paper investigates the ability of artificial neural networks to judge the grammatical acceptability of a sentence, with the goal of testing their linguistic competence. We introduce the Corpus of Linguistic Acceptability (CoLA), a set of 10,657 English sentences labeled as grammatical or ungrammatical from published linguistics literature. As baselines, we train several recurrent neural network models on acceptability classification, and find that our models outperform unsupervised models by Lau et al. (2016) on CoLA. Error-analysis on specific grammatical phenomena reveals that both Lau et al.’s models and ours learn systematic generalizations like subject-verb-object order. However, all models we test perform far below human level on a wide range of grammatical constructions.
Alex Warstadt, Amanpreet Singh, Samuel R. Bowman
Trans. Assoc. Comput. Linguistics3
2018 The Lifted Matrix-Space Model for Semantic Composition
abstract
Tree-structured neural network architectures for sentence encoding draw inspiration from the approach to semantic composition generally seen in formal linguistics, and have shown empirical improvements over comparable sequence models by doing so.Moreover, adding multiplicative interaction terms to the composition functions in these models can yield significant further improvements.However, existing compositional approaches that adopt such a powerful composition function scale poorly, with parameter counts exploding as model dimension or vocabulary size grows.We introduce the Lifted Matrix-Space model, which uses a global transformation to map vector word embeddings to matrices, which can then be composed via an operation based on matrix-matrix multiplication.Its composition function effectively transmits a larger number of activations across layers with relatively few model parameters.We evaluate our model on the Stanford NLI corpus, the Multi-Genre NLI corpus, and the Stanford Sentiment Treebank and find that it consistently outperforms TreeLSTM (Tai et al., 2015), the previous best known composition function for treestructured models.
Woojin Chung, Sheng-Fu Wang, Samuel R. Bowman
CoNLL3
2018 A Stable and Effective Learning Strategy for Trainable Greedy Decoding
abstract
Beam search is a widely used approximate search strategy for neural network decoders, and it generally outperforms simple greedy decoding on tasks like machine translation.However, this improvement comes at substantial computational cost.In this paper, we propose a flexible new method that allows us to reap nearly the full benefits of beam search with nearly no additional computational cost.The method revolves around a small neural network actor that is trained to observe and manipulate the hidden state of a previouslytrained decoder.To train this actor network, we introduce the use of a pseudo-parallel corpus built using the output of beam search on a base model, ranked by a target quality metric like BLEU.Our method is inspired by earlier work on this problem, but requires no reinforcement learning, and can be trained reliably on a range of models.Experiments on three parallel corpora and three architectures show that the method yields substantial improvements in translation quality and speed over each base system.
Yun Chen 0007, Victor O. K. Li, Kyunghyun Cho, Samuel R. Bowman
EMNLP4
2018 XNLI: Evaluating Cross-lingual Sentence Representations
abstract
Alexis Conneau, Ruty Rinott, Guillaume Lample, Adina Williams, Samuel Bowman, Holger Schwenk, Veselin Stoyanov. Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing. 2018.
Alexis Conneau, Ruty Rinott, Guillaume Lample, Adina Williams, Samuel R. Bowman, Holger Schwenk, Veselin Stoyanov
EMNLP5
2018 Grammar Induction with Neural Language Models: An Unusual Replication
abstract
A substantial thread of recent work on latent tree learning has attempted to develop neural network models with parse-valued latent variables and train them on non-parsing tasks, in the hope of having them discover interpretable tree structure.In a recent paper, Shen et al. (2018) introduce such a model and report nearstate-of-the-art results on the target task of language modeling, and the first strong latent tree learning result on constituency parsing.In an attempt to reproduce these results, we discover issues that make the original results hard to trust, including tuning and even training on what is effectively the test set.Here, we attempt to reproduce these results in a fair experiment and to extend them to two new datasets.We find that the results of this work are robust: All variants of the model under study outperform all latent tree learning baselines, and perform competitively with symbolic grammar induction systems.We find that this model represents the first empirical success for latent tree learning, and that neural network language modeling warrants further study as a setting for grammar induction.
Phu Mon Htut, Kyunghyun Cho, Samuel R. Bowman
EMNLP3
2018 A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference
abstract
Adina Williams, Nikita Nangia, Samuel Bowman. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Adina Williams, Nikita Nangia, Samuel R. Bowman
NAACL-HLT3
2018 Do latent tree learning models identify meaningful structure in sentences?
abstract
Recent work on the problem of latent tree learning has made it possible to train neural networks that learn to both parse a sentence and use the resulting parse to interpret the sentence, all without exposure to ground-truth parse trees at training time. Surprisingly, these models often perform better at sentence understanding tasks than models that use parse trees from conventional parsers. This paper aims to investigate what these latent tree learning models learn. We replicate two such models in a shared codebase and find that (i) only one of these models outperforms conventional tree-structured models on sentence classification, (ii) its parsing strategies are not especially consistent across random restarts, (iii) the parses it produces tend to be shallower than standard Penn Treebank (PTB) parses, and (iv) they do not resemble those of PTB or any other semantic or syntactic formalism that the authors are aware of.
Adina Williams, Andrew Drozdov, Samuel R. Bowman
Trans. Assoc. Comput. Linguistics3
2016 A Fast Unified Model for Parsing and Sentence Understanding
abstract
Samuel R. Bowman, Jon Gauthier, Abhinav Rastogi, Raghav Gupta, Christopher D. Manning, Christopher Potts. Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2016.
Samuel R. Bowman, Jon Gauthier, Abhinav Rastogi, Christopher D. Manning, Christopher Potts
ACL (1)1
2016 Generating Sentences from a Continuous Space
abstract
The standard recurrent neural network language model (rnnlm) generates sentences one word at a time and does not work from an explicit global sentence representation.In this work, we introduce and study an rnn-based variational autoencoder generative model that incorporates distributed latent representations of entire sentences.This factorization allows it to explicitly model holistic properties of sentences such as style, topic, and high-level syntactic features.Samples from the prior over these sentence representations remarkably produce diverse and well-formed sentences through simple deterministic decoding.By examining paths through this latent space, we are able to generate coherent novel sentences that interpolate between known sentences.We present techniques for solving the difficult learning problem presented by this model, demonstrate its effectiveness in imputing missing words, explore many interesting properties of the model's latent sentence space, and present negative results on the use of the model in language modeling.but now , as they parked out front and owen stepped out of the car , he could see True: that the transition was complete .RNNLM: it , " i said .VAE: through the driver 's door .you kill
Samuel R. Bowman, Luke Vilnis, Oriol Vinyals, Andrew M. Dai, Rafal Józefowicz, Samy Bengio
CoNLL1
2015 A large annotated corpus for learning natural language inference
abstract
Understanding entailment and contradiction is fundamental to understanding natural language, and inference about entailment and contradiction is a valuable testing ground for the development of semantic representations.However, machine learning research in this area has been dramatically limited by the lack of large-scale resources.To address this, we introduce the Stanford Natural Language Inference corpus, a new, freely available collection of labeled sentence pairs, written by humans doing a novel grounded task based on image captioning.At 570K pairs, it is two orders of magnitude larger than all other resources of its type.This increase in scale allows lexicalized classifiers to outperform some sophisticated existing entailment models, and it allows a neural network-based model to perform competitively on natural language inference benchmarks for the first time.
Samuel R. Bowman, Gabor Angeli, Christopher Potts, Christopher D. Manning
EMNLP1
2014 A Gold Standard Dependency Corpus for English
Natalia Silveira, Timothy Dozat, Marie-Catherine de Marneffe, Samuel R. Bowman, Miriam Connor, John Bauer, Christopher D. Manning
LREC4
2012 Automatic Animacy Classification
Samuel R. Bowman, Harshit Chopra
HLT-NAACL1
2011 Speech recognitionwith segmental conditional random fields: A summary of the JHU CLSP 2010 Summer Workshop
abstract
This paper summarizes the 2010 CLSP Summer Workshop on speech recognition at Johns Hopkins University. The key theme of the workshop was to improve on state-of-the-art speech recognition systems by using Segmental Conditional Random Fields (SCRFs) to integrate multiple types of information. This approach uses a state of-the-art baseline as a springboard from which to add a suite of novel features including ones derived from acoustic templates, deep neural net phoneme detections, duration models, modulation features, and whole word point-process models. The SCRF framework is able to appropriately weight these different information sources to produce significant gains on both die Broadcast News and Wall Street Journal tasks.
Geoffrey Zweig, Patrick Nguyen, Dirk Van Compernolle, Kris Demuynck, Les E. Atlas, Pascal Clark, Gregory Sell, Meihong Wang, Fei Sha, Hynek Hermansky, Damianos Karakos, Aren Jansen, Samuel Thomas 0001, Sivaram G. S. V. S., Samuel R. Bowman, Justine T. Kao
ICASSP15
2010 Modeling pronunciation variation with context-dependent articulatory feature decision trees
abstract
We consider the problem of predicting the surface pronunciations of a word in conversational speech, using a model of pronunciation variation based on articulatory features. We build context-dependent decision trees for both phone-based and feature-based models, and compare their perplexities on conversational data from the Switchboard Transcription Project. We find that a fully-factored model, with separate decision trees for each articulatory feature, does not perform well, but a feature-based model using a smaller number of “feature bundles” outperforms both the fully-factored model and a phonebased model. The articulatory feature-based decision trees are also much more robust to reductions in training data. We also analyze the usefulness of various context variables.
Samuel R. Bowman, Karen Livescu
INTERSPEECH1