EDBT 2026 Demo / reviewers in the wild / expert
Alex Warstadt
dblp:220/5281 · also Alexander Warstadt
· DBLP profile ↗
28ranked-venue papers
7as first author
21since 2021 · last 2026
0000-0002-5397-3151ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 28 · 7 first-author · 21 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | It's Not What You Say, It's How You Say It: Evaluating LLM Responses to Expressions of BeliefabstractUsers frequently express their beliefs to large language models (LLMs).In some situations, it is ideal for the LLM to accept this contextual information as true, while in others, it is ideal to stick to prior knowledge.Users' expressions of belief (EoBs) can take linguistically diverse forms-using presuppositions, evidential and certainty markers, or varied toneseach of which may have a different persuasiveness over the LLMs.We introduce a benchmark to systematically evaluate how different EoBs affect whether models follow context versus prior knowledge.We propose a typology grounded in four linguistically motivated dimensions: form, evidentiality, epistemic stance, and tone, spanning 19 fine-grained types.By pairing these EoBs with world knowledge facts, we generate controlled EoB-query pairs that isolate the effect of linguistic variation.We use our benchmark to evaluate 18 LLMs that differ in architecture (Llama3, Qwen3, Gemma3), scale (1B-30B parameters), and training stages (base vs instruct).We identify meaningful variations in response behavior across these axes: For example, bigger models and instruction models tend to be less context-following than smaller models and base models.We further identify specific EoBs that statistically significantly persuade LMs more consistently than others.These systematic patterns in how linguistic framing affects LLM context integration serve to evaluate model robustness and inform best practices for prompt engineering.We publicly release code and data used in this project. Kevin Du, Clara Kümpel, Michelle Wastl, Alex Warstadt |
ACL (1) | 4 |
| 2026 | Dual Alignment Between Language Model Layers and Human Sentence ProcessingabstractA recent study (Kuribayashi et al., 2025) has shown that human sentence processing behavior, typically measured on syntactically unchallenging constructions, can be effectively modeled using surprisal from early layers of large language models (LLMs).This raises the question of whether such advantages of internal layers extend to more syntactically challenging constructions, where surprisal has been reported to underestimate human cognitive effort.In this paper, we begin by exploring internal layers that better estimate human cognitive effort observed in syntactic ambiguity processing in English.Our experiments show that, in contrast to naturalistic reading, later layers better estimate such a cognitive effort, but still underestimate the human data.This dual alignment sheds light on different modes of sentence processing in humans and LMs: naturalistic reading employs a somewhat weak prediction akin to earlier layers of LMs, while syntactically challenging processing requires more fully-contextualized representations, better modeled by later layers of LMs.Motivated by these findings, we also explore several probability-update measures using shallow and deep layers of LMs, showing a complementary advantage to single-layer's surprisal in reading time modeling. https://github.com/kuribayashi4/ internal_surprisal_targeted_assessmentPhenomena Example MVRR D + : The girl fed the lamb remained relatively calm before the sunset in silence.D -: The girl who was fed the lamb remained relatively calm before the sunset in silence.NPS D + : The girl found the lamb remained relatively calm near the wooden fence.D -: The girl found that the lamb remained relatively calm near the wooden fence.NPZ D + : When the girl attacked the lamb remained relatively calm despite the sudden noise.D -: When the girl attacked, the lamb remained relatively calm despite the sudden noise.RC D + : The bus driver that the kids followed waited patiently at dawn.D -: The bus driver that followed the kids waited patiently at dawn.Attachment D + : Janet charmed the executive of the assistants who decides almost everything during long weekly meetings. Tatsuki Kuribayashi, Alex Warstadt, Yohei Oseki, Ethan Wilcox |
ACL (1) | 2 |
| 2026 | What Do Prosody and Text Convey? Characterizing How Meaningful Information is Distributed Across Multiple ChannelsabstractAditya Yadavalli, Tiago Pimentel, Tamar I Regev, Ethan Gotlieb Wilcox, Alex Warstadt. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Aditya Yadavalli, Tiago Pimentel, Tamar I. Regev, Ethan Wilcox, Alex Warstadt |
ACL (1) | 5 |
| 2026 | An Information-Theoretic Study of RLHF-Induced Uniformity in Large Language Model OutputsabstractReinforcement Learning with Human Feedback (RLHF) is an increasingly popular post-training procedure for Large Language Models (LLMs) to better align outputs with human preferences.Therefore, one might expect some sense of human-like audience design to be induced into LLMs.However, RLHF and other post-training alignment methods have many complex effects on the outputs of LLMs that can be difficult to study quantitatively.We apply an informationtheoretic lens to investigate the changes in the "naturalness" of language and the presence of audience design in LLMs before and after posttraining.The Uniform Information Density (UID) Hypothesis posits that humans optimize language production and comprehension across a noisy channel by transferring information at a more uniform rate.Accordingly, we analyze and compare how information is distributed within model-generated and human-generated text belonging to various domains to investigate the presence and form of audience design in LLMs.We find that pretrained and posttrained LLMs both show superhuman uniformity across various text domains, while RLHF encourages slightly more human-like, i.e., less uniform, outputs.However, other post-training approaches have a similar effect, suggesting that information uniformity is not a significant driver of human preferences. Nolan Chai, Alex Warstadt |
CoNLL | 3 |
| 2026 | Can Language Models Learn Typologically Implausible Languages?abstractAbstract Grammatical features across human languages exhibit intriguing correlations, often attributed to learning biases in humans. Language models (LMs) provide a scalable and naturalistic framework for studying artificial language learning—one not available in human research. We investigate how learnability varies across typologically plausible and implausible languages that closely follow the word order universals identified by linguistic typologists. Our study trains LMs on highly naturalistic counterfactual versions of English (head-initial) and Japanese (head-final). Compared to prior work, our datasets more precisely target the boundary between typological plausibility and implausibility. Our experiments show that LMs learn subtly implausible languages more slowly, though they eventually reach similar performance on some metrics regardless of typological plausibility. These findings suggest that LMs exhibit typologically aligned learning preferences and that certain typological patterns may emerge from general learning biases. https://github.com/sally-xu-42/Typological_Universals. Tianyang Xu 0002, Tatsuki Kuribayashi, Yohei Oseki, Ryan Cotterell, Alex Warstadt |
Trans. Assoc. Comput. Linguistics | 5 |
| 2025 | The time scale of redundancy between prosody and linguistic contextabstractTamar I Regev, Chiebuka Ohams, Shaylee Xie, Lukas Wolf, Evelina Fedorenko, Alex Warstadt, Ethan Wilcox, Tiago Pimentel. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Tamar I. Regev, Chiebuka Ohams, Shaylee Xie, Lukas Wolf, Evelina Fedorenko, Alex Warstadt, Ethan Wilcox, Tiago Pimentel |
ACL (1) | 6 |
| 2025 | The Harmonic Structure of Information ContoursabstractEleftheria Tsipidi, Samuel Kiegeland, Franz Nowak, Tianyang Xu, Ethan Wilcox, Alex Warstadt, Ryan Cotterell, Mario Giulianelli. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Eleftheria Tsipidi, Samuel Kiegeland, Franz Nowak, Tianyang Xu 0002, Ethan Wilcox, Alex Warstadt, Ryan Cotterell, Mario Giulianelli |
ACL (1) | 6 |
| 2025 | Using Information Theory to Characterize Prosodic Typology: The Case of Tone, Pitch-Accent and Stress-AccentabstractEthan Wilcox, Cui Ding, Giovanni Acampa, Tiago Pimentel, Alex Warstadt, Tamar I Regev. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Ethan Wilcox, Cui Ding, Giovanni Acampa, Tiago Pimentel, Alex Warstadt, Tamar I. Regev |
ACL (1) | 5 |
| 2025 | A Distributional Perspective on Word Learning in Neural Language ModelsabstractFilippo Ficarra, Ryan Cotterell, Alex Warstadt. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Filippo Ficarra, Ryan Cotterell, Alex Warstadt |
NAACL (Long Papers) | 3 |
| 2025 | Investigating Critical Period Effects in Language Acquisition through Neural Language ModelsabstractAbstract Humans appear to have a critical period (CP) for language acquisition: Second language (L2) acquisition becomes harder after early childhood, and ceasing exposure to a first language (L1) after this period (but not before) typically does not lead to substantial loss of L1 proficiency. It is unknown whether these CP effects result from innately determined brain maturation or as a stabilization of neural connections naturally induced by experience. In this study, we use language models (LMs) to test the extent to which these phenomena are peculiar to humans, or shared by a broader class of language learners. We vary the age of exposure by training LMs on language pairs in various experimental conditions, and find that LMs, which lack any direct analog to innate maturational stages, do not show CP effects when the age of exposure of L2 is delayed. Our results contradict the claim that CP effects are an inevitable result of statistical learning, and they are consistent with an innate mechanism for CP effects. We show that we can reverse-engineer the CP by introducing a regularizer partway through training to simulate a maturational decrease in plasticity. All in all, our results suggest that L1 learning on its own may not be enough to induce a CP, and additional engineering is necessary to make language models more cognitively plausible. Ionut Constantinescu, Tiago Pimentel, Ryan Cotterell, Alex Warstadt |
Trans. Assoc. Comput. Linguistics | 4 |
| 2024 | Insights from the first BabyLM Challenge: Training sample-efficient language models on a developmentally plausible corpus
Alex Warstadt, Aaron Mueller, Leshem Choshen, Ethan Wilcox, Chengxu Zhuang, Adina Williams, Ryan Cotterell, Tal Linzen |
CogSci | 1 |
| 2024 | Automatic Annotation of Grammaticality in Child-Caregiver ConversationsabstractThe acquisition of grammar has been a central question to adjudicate between theories of language acquisition. In order to conduct faster, more reproducible, and larger-scale corpus studies on grammaticality in child-caregiver conversations, tools for automatic annotation can offer an effective alternative to tedious manual annotation. We propose a coding scheme for context-dependent grammaticality in child-caregiver conversations and annotate more than 4,000 utterances from a large corpus of transcribed conversations. Based on these annotations, we train and evaluate a range of NLP models. Our results show that fine-tuned Transformer-based models perform best, achieving human inter-annotation agreement levels. As a first application and sanity check of this tool, we use the trained models to annotate a corpus almost two orders of magnitude larger than the manually annotated data and verify that children’s grammaticality shows a steady increase with age. This work contributes to the growing literature on applying state-of-the-art NLP methods to help study child language acquisition at scale. Mitja Nikolaus, Abhishek Agrawal, Petros Kaklamanis, Alex Warstadt, Abdellah Fourtassi |
LREC/COLING | 4 |
| 2024 | Surprise! Uniform Information Density Isn't the Whole Story: Predicting Surprisal Contours in Long-form DiscourseabstractThe Uniform Information Density (UID) hypothesis posits that speakers tend to distribute information evenly across linguistic units to achieve efficient communication.Of course, information rate in texts and discourses is not perfectly uniform.While these fluctuations can be viewed as theoretically uninteresting noise on top of a uniform target, another explanation is that UID is not the only functional pressure regulating information content in a language.Speakers may also seek to maintain interest, adhere to writing conventions, and build compelling arguments.In this paper, we propose one such functional pressure; namely that speakers modulate information rate based on location within a hierarchically-structured model of discourse.We term this the Structured Context Hypothesis and test it by predicting the surprisal contours of naturally occurring discourses extracted from large language models using predictors derived from discourse structure.We find that hierarchical predictors are significant predictors of a discourse's information contour and that deeply nested hierarchical predictors are more predictive than shallow ones.This work takes an initial step beyond UID to propose testable hypotheses for why the information rate fluctuates in predictable ways.https://github.com/rycolab/ surprisal-discourse Eleftheria Tsipidi, Franz Nowak, Ryan Cotterell, Ethan Wilcox, Mario Giulianelli, Alex Warstadt |
EMNLP | 6 |
| 2023 | Generalizing Backpropagation for Gradient-Based InterpretabilityabstractMany popular feature-attribution methods for interpreting deep neural networks rely on computing the gradients of a model's output with respect to its inputs.While these methods can indicate which input features may be important for the model's prediction, they reveal little about the inner workings of the model itself.In this paper, we observe that the gradient computation of a model is a special case of a more general formulation using semirings.This observation allows us to generalize the backpropagation algorithm to efficiently compute other interpretable statistics about the gradient graph of a neural network, such as the highest-weighted path and entropy.We implement this generalized algorithm, evaluate it on synthetic datasets to better understand the statistics it computes, and apply it to study BERT's behavior on the subject-verb number agreement task (SVA).With this method, we (a) validate that the amount of gradient flow through a component of a model reflects its importance to a prediction and (b) for SVA, identify which pathways of the self-attention mechanism are most important. Kevin Du, Lucas Torroba Hennigen, Niklas Stoehr, Alex Warstadt, Ryan Cotterell |
ACL (1) | 4 |
| 2023 | Quantifying the redundancy between prosody and textabstractProsody-the suprasegmental component of speech, including pitch, loudness, and tempocarries critical aspects of meaning.However, the relationship between the information conveyed by prosody vs. by the words themselves remains poorly understood.We use large language models (LLMs) to estimate how much information is redundant between prosody and the words themselves.Using a large spoken corpus of English audiobooks, we extract prosodic features aligned to individual words and test how well they can be predicted from LLM embeddings, compared to non-contextual word embeddings.We find a high degree of redundancy between the information carried by the words and prosodic information across several prosodic features, including intensity, duration, pauses, and pitch contours.Furthermore, a word's prosodic information is redundant with both the word itself and the context preceding as well as following it.Still, we observe that prosodic features can not be fully predicted from text, suggesting that prosody carries information above and beyond the words.Along with this paper, we release a general-purpose data processing pipeline for quantifying the relationship between linguistic information and extra-linguistic features.https://github.com/lu-wo/ quantifying-redundancy Lukas Wolf, Tiago Pimentel, Evelina Fedorenko, Ryan Cotterell, Alex Warstadt, Ethan Wilcox, Tamar I. Regev |
EMNLP | 5 |
| 2022 | What Makes Reading Comprehension Questions Difficult?abstractFor a natural language understanding benchmark to be useful in research, it has to consist of examples that are diverse and difficult enough to discriminate among current and near-future state-of-the-art systems.However, we do not yet know how best to select text sources to collect a variety of challenging examples.In this study, we crowdsource multiple-choice reading comprehension questions for passages taken from seven qualitatively distinct sources, analyzing what attributes of passages contribute to the difficulty and question types of the collected examples.To our surprise, we find that passage source, length, and readability measures do not significantly affect question difficulty.Through our manual annotation of seven reasoning types, we observe several trends between passage sources and reasoning types, e.g., logical reasoning is more often required in questions written for technical passages.These results suggest that when creating a new benchmark dataset, selecting a diverse set of passages can help ensure a diverse range of question types, but that passage difficulty need not be a priority. Saku Sugawara, Nikita Nangia, Alex Warstadt, Samuel R. Bowman |
ACL (1) | 3 |
| 2022 | Entailment Semantics Can Be Extracted from an Ideal Language ModelabstractLanguage models are often trained on text alone, without additional grounding.There is debate as to how much of natural language semantics can be inferred from such a procedure.We prove that entailment judgments between sentences can be extracted from an ideal language model that has perfectly learned its target distribution, assuming the training sentences are generated by Gricean agents, i.e., agents who follow fundamental principles of communication from the linguistic theory of pragmatics.We also show entailment judgments can be decoded from the predictions of a language model trained on such Gricean data.Our results reveal a pathway for understanding the semantic information encoded in unlabeled linguistic data and a potential framework for extracting semantics from language models. William Merrill, Alex Warstadt, Tal Linzen |
CoNLL | 2 |
| 2021 | What Ingredients Make for an Effective Crowdsourcing Protocol for Difficult NLU Data Collection Tasks?abstractNikita Nangia, Saku Sugawara, Harsh Trivedi, Alex Warstadt, Clara Vania, Samuel R. Bowman. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Nikita Nangia, Saku Sugawara, Harsh Trivedi, Alex Warstadt, Clara Vania, Samuel R. Bowman |
ACL/IJCNLP (1) | 4 |
| 2021 | When Do You Need Billions of Words of Pretraining Data?abstractYian Zhang, Alex Warstadt, Xiaocheng Li, Samuel R. Bowman. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Yian Zhang, Alex Warstadt, Xiaocheng Li, Samuel R. Bowman |
ACL/IJCNLP (1) | 2 |
| 2021 | NOPE: A Corpus of Naturally-Occurring Presuppositions in EnglishabstractAlicia Parrish, Sebastian Schuster, Alex Warstadt, Omar Agha, Soo-Hwan Lee, Zhuoye Zhao, Samuel R. Bowman, Tal Linzen. Proceedings of the 25th Conference on Computational Natural Language Learning. 2021. Alicia Parrish, Sebastian Schuster 0001, Alex Warstadt, Omar Agha, Soo-Hwan Lee, Zhuoye Zhao, Samuel R. Bowman, Tal Linzen |
CoNLL | 3 |
| 2021 | CLiMP: A Benchmark for Chinese Language Model EvaluationabstractLinguistically informed analyses of language models (LMs) contribute to the understanding and improvement of these models.Here, we introduce the corpus of Chinese linguistic minimal pairs (CLiMP), which can be used to investigate what knowledge Chinese LMs acquire.CLiMP consists of sets of 1,000 minimal pairs (MPs) for 16 syntactic contrasts in Mandarin, covering 9 major Mandarin linguistic phenomena.The MPs are semiautomatically generated, and human agreement with the labels in CLiMP is 95.8%.We evaluate 11 different LMs on CLiMP, covering n-grams, LSTMs, and Chinese BERT.We find that classifier-noun agreement and verb complement selection are the phenomena that models generally perform best at.However, models struggle the most with the bǎ construction, binding, and filler-gap dependencies.Overall, Chinese BERT achieves an 81.8% average accuracy, while the performances of LSTMs and 5-grams are only moderately above chance level. Beilei Xiang, Changbing Yang, Alex Warstadt, Katharina Kann |
EACL | 4 |
| 2020 | Are Natural Language Inference Models IMPPRESsive? Learning IMPlicature and PRESuppositionabstractNatural language inference (NLI) is an increasingly important task for natural language understanding, which requires one to infer whether a sentence entails another.However, the ability of NLI models to make pragmatic inferences remains understudied.We create an IMPlicature and PRESupposition diagnostic dataset (IMPPRES), consisting of >25k semiautomatically generated sentence pairs illustrating well-studied pragmatic inference types.We use IMPPRES to evaluate whether BERT, InferSent, and BOW NLI models trained on MultiNLI (Williams et al., 2018) learn to make pragmatic inferences.Although MultiNLI appears to contain very few pairs illustrating these inference types, we find that BERT learns to draw pragmatic inferences.It reliably treats scalar implicatures triggered by "some" as entailments.For some presupposition triggers like only, BERT reliably recognizes the presupposition as an entailment, even when the trigger is embedded under an entailment canceling operator like negation.BOW and InferSent show weaker evidence of pragmatic reasoning.We conclude that NLI training encourages models to learn some, but not all, pragmatic inferences.Type Example Trigger Jo's cat yawned.Presupposition Jo has a cat.Negated Trigger Jo's cat didn't yawn.Modal Trigger It's possible that Jo's cat yawned.Interrog.Trigger Did Jo's cat yawn?Cond.Trigger If Jo's cat yawned, it's OK.Negated Prsp.Jo doesn't have a cat.Neutral Prsp.Amy has a cat. Paloma Jeretic, Alex Warstadt, Suvrat Bhooshan, Adina Williams |
ACL | 2 |
| 2020 | Can neural networks acquire a structural bias from raw linguistic data?
Alex Warstadt, Samuel R. Bowman |
CogSci | 1 |
| 2020 | Learning Which Features Matter: RoBERTa Acquires a Preference for Linguistic Generalizations (Eventually)abstractOne reason pretraining on self-supervised linguistic tasks is effective is that it teaches models features that are helpful for language understanding.However, we want pretrained models to learn not only to represent linguistic features, but also to use those features preferentially during fine-turning.With this goal in mind, we introduce a new English-language diagnostic set called MSGS (the Mixed Signals Generalization Set), which consists of 20 ambiguous binary classification tasks that we use to test whether a pretrained model prefers linguistic or surface generalizations during finetuning.We pretrain RoBERTa models from scratch on quantities of data ranging from 1M to 1B words and compare their performance on MSGS to the publicly available RoBERTa BASE .We find that models can learn to represent linguistic features with little pretraining data, but require far more data to learn to prefer linguistic generalizations over surface ones.Eventually, with about 30B words of pretraining data, RoBERTa BASE does demonstrate a linguistic bias with some regularity.We conclude that while self-supervised pretraining is an effective way to learn helpful inductive biases, there is likely room to improve the rate at which models learn which features matter. Feature type Feature description Positive example Negative example SurfaceAbsolute position Is the first token of S "the"?The cat chased a mouse.A cat chased a mouse.Length Is S longer than n (e.g., 3) words?The cat chased a mouse.The cat meowed.Lexical content Does S contain "the"?That cat chased the mouse.That cat chased a mouse.Relative position Does "the" precede "a"?The cat chased a mouse.A cat chased the mouse.Orthography Does S appear in title case?The Cat Chased a Mouse.The cat chased a mouse. LinguisticMorphology Does S have an irregular past verb?The cats slept.The cats meow.Syn.category Does S have an adjective?Lincoln was tall.Lincoln was president.Syn.construction Is S the control construction?Sue is eager to sleep.Sue is likely to sleep.Syn.position Is the main verb in "ing" form?Cats who eat mice are purring.Cats who are eating mice purr. Alex Warstadt, Yian Zhang, Xiaocheng Li, Haokun Liu, Samuel R. Bowman |
EMNLP (1) | 1 |
| 2020 | BLiMP: The Benchmark of Linguistic Minimal Pairs for EnglishabstractWe introduce The Benchmark of Linguistic Minimal Pairs (BLiMP),1 a challenge set for evaluating the linguistic knowledge of language models (LMs) on major grammatical phenomena in English. BLiMP consists of 67 individual datasets, each containing 1,000 minimal pairs—that is, pairs of minimally different sentences that contrast in grammatical acceptability and isolate specific phenomenon in syntax, morphology, or semantics. We generate the data according to linguist-crafted grammar templates, and human aggregate agreement with the labels is 96.4%. We evaluate n-gram, LSTM, and Transformer (GPT-2 and Transformer-XL) LMs by observing whether they assign a higher probability to the acceptable sentence in each minimal pair. We find that state-of-the-art models identify morphological contrasts related to agreement reliably, but they struggle with some subtle semantic and syntactic phenomena, such as negative polarity items and extraction islands. Alex Warstadt, Alicia Parrish, Haokun Liu, Anhad Mohananey, Wei Peng 0013, Sheng-Fu Wang, Samuel R. Bowman |
Trans. Assoc. Comput. Linguistics | 1 |
| 2020 | Erratum: "BLiMP: The Benchmark of Linguistic Minimal Pairs for English"abstractWe correct wrongly reported results on BLiMP. Alex Warstadt, Alicia Parrish, Haokun Liu, Anhad Mohananey, Wei Peng 0013, Sheng-Fu Wang, Samuel R. Bowman |
Trans. Assoc. Comput. Linguistics | 1 |
| 2019 | Investigating BERT's Knowledge of Language: Five Analysis Methods with NPIsabstractAlex Warstadt, Yu Cao, Ioana Grosu, Wei Peng, Hagen Blix, Yining Nie, Anna Alsop, Shikha Bordia, Haokun Liu, Alicia Parrish, Sheng-Fu Wang, Jason Phang, Anhad Mohananey, Phu Mon Htut, Paloma Jeretic, Samuel R. Bowman. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Alex Warstadt, Ioana Grosu, Wei Peng 0013, Hagen Blix, Yining Nie, Anna Alsop, Shikha Bordia, Haokun Liu, Alicia Parrish, Sheng-Fu Wang, Jason Phang, Anhad Mohananey, Phu Mon Htut, Paloma Jeretic, Samuel R. Bowman |
EMNLP/IJCNLP (1) | 1 |
| 2019 | Neural Network Acceptability JudgmentsabstractThis paper investigates the ability of artificial neural networks to judge the grammatical acceptability of a sentence, with the goal of testing their linguistic competence. We introduce the Corpus of Linguistic Acceptability (CoLA), a set of 10,657 English sentences labeled as grammatical or ungrammatical from published linguistics literature. As baselines, we train several recurrent neural network models on acceptability classification, and find that our models outperform unsupervised models by Lau et al. (2016) on CoLA. Error-analysis on specific grammatical phenomena reveals that both Lau et al.’s models and ours learn systematic generalizations like subject-verb-object order. However, all models we test perform far below human level on a wide range of grammatical constructions. Alex Warstadt, Amanpreet Singh, Samuel R. Bowman |
Trans. Assoc. Comput. Linguistics | 1 |