Tal Linzen

dblp:169/3438 · DBLP profile ↗
← Back
56ranked-venue papers
6as first author
30since 2021 · last 2026
0000-0003-0435-6912ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 56 · 6 first-author · 30 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 3 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Language Models Struggle to Use Representations Learned In-Context
abstract
Though language models (LMs) have enabled great success across a wide variety of tasks, they still appear to fall short of one of the loftier goals of artificial intelligence research: creating an artificial system that can adapt its behavior to radically new contexts upon deployment (Shi et al., 2024).One important step towards this goal is to create systems that can induce rich representations of data that are seen in-context, and then flexibly deploy these representations to accomplish goals (Lampinen et al., 2024).Recently, Park et al. (2025a) demonstrated that current LMs are indeed capable of inducing such representation from context (i.e., in-context representation learning).The present study investigates whether LMs can use these representations to complete simple downstream tasks.We first assess whether open-weights LMs can use in-context representations for next-token prediction, and then probe models using a novel task, adaptive world modeling.In both tasks, we find evidence that open-weights LMs struggle to deploy representations of novel semantics that are defined in-context, even if they encode these semantics in their latent representations.Furthermore, we assess closed-source, state-of-the-art reasoning models on the adaptive world modeling task, and demonstrate that even the most performant LMs cannot reliably leverage novel patterns presented in-context.Overall, this work seeks to inspire novel methods for encouraging models to not only encode information presented in-context, but to do so in a manner that supports flexible deployment of this information.
Michael A. Lepori, Tal Linzen, Ann Yuan, Katja Filippova
ACL (1)2
2026 RELIC: Evaluating Complex Reasoning via the Recognition of Languages In-Context
abstract
Abstract Large language models (LLMs) are increasingly used to solve complex tasks where they must retrieve and compose many pieces of in-context information in long reasoning chains. For many real-world tasks it is hard to accurately gauge how model performance and strategy change as task complexity grows. To evaluate models’ complex reasoning capability in a scalable and verifiable way, we introduce RELIC (Recognition of Languages In-Context), a framework that evaluates an LLM’s ability to decide whether a given string belongs to the context-free language (CFL) generated by a grammar presented in-context. CFL recognition allows us to modulate the intrinsic complexity of the problem by varying grammar size and string length and translate this asymptotic complexity into predictions for ideal LLM performance. We find that even the most advanced reasoning models perform poorly on RELIC, not only failing to appropriately scale their inference compute to keep pace with task difficulty, but even reducing the number of reasoning tokens they use as task complexity increases. We find that these decreases in compute accompany changes in reasoning strategy, as models move from identifying and implementing algorithmic solutions to guessing. For models whose full completions go uninspected, this manifests as “quiet quitting” on hard tasks. Code: https://jpetty.org/relic
Jackson Petty, Michael Y. Hu, Shauli Ravfogel, William Merrill, Tal Linzen
Trans. Assoc. Comput. Linguistics6
2025 Between Circuits and Chomsky: Pre-pretraining on Formal Languages Imparts Linguistic Biases
abstract
Pretraining language models on formal language can improve their acquisition of natural language.Which features of the formal language impart an inductive bias that leads to effective transfer?Drawing on insights from linguistics and complexity theory, we hypothesize that effective transfer occurs when two conditions are met: the formal language should capture the dependency structures present in natural language, and it should remain within the computational limitations of the model architecture.We experiment with pre-pretraining (training on formal language before natural languages) on transformers and find that formal languages capturing hierarchical dependencies indeed enable language models to achieve lower loss on natural language and better linguistic generalization compared to other formal languages.We also find modest support for the hypothesis that the formal language should fall within the computational limitations of the architecture.Strikingly, pre-pretraining reduces loss more efficiently than training on a matched amount of natural language.For a 1B-parameter language model trained on roughly 1.6B tokens of natural language, pre-pretraining achieves the same loss and better linguistic generalization with a 33% smaller token budget.Finally, we also give mechanistic evidence of transfer from formal to natural language: attention heads acquired during pre-pretraining remain crucial for the model's performance on syntactic evaluations. 1
Michael Y. Hu, Jackson Petty, William Merrill, Tal Linzen
ACL (1)5
2025 Rapid Word Learning Through Meta In-Context Learning
abstract
Humans can quickly learn a new word from a few illustrative examples, and then systematically and flexibly use it in novel contexts.Yet the abilities of current language models for fewshot word learning, and methods for improving these abilities, are underexplored.In this study, we introduce a novel method, Meta-training for IN-context learNing Of Words (Minnow).This method trains language models to generate new examples of a word's usage given a few in-context examples, using a special placeholder token to represent the new word.This training is repeated on many new words to develop a general word-learning ability.We find that training models from scratch with Minnow on human-scale child-directed language enables strong few-shot word learning, comparable to a large language model (LLM) pretrained on orders of magnitude more data.Furthermore, through discriminative and generative evaluations, we demonstrate that finetuning pre-trained LLMs with Minnow improves their ability to discriminate between new words, identify syntactic categories of new words, and generate reasonable new usages and definitions for new words, based on one or a few in-context examples.These findings highlight the data efficiency of Minnow and its potential to improve language model performance in word learning tasks.
Guangyuan Jiang, Tal Linzen, Brenden M. Lake
EMNLP3
2025 Multilingual Prompting for Improving LLM Generation Diversity
abstract
Large Language Models (LLMs) are known to lack cultural representation and overall diversity in their generations, from expressing opinions to answering factual questions.To mitigate this problem, we propose multilingual prompting: a prompting method which generates several variations of a base prompt with added cultural and linguistic cues from several cultures, generates responses, and then combines the results.Building on evidence that LLMs have language-specific knowledge, multilingual prompting seeks to increase diversity by activating a broader range of cultural knowledge embedded in model training data.Through experiments across multiple models (GPT-4o, GPT-4o-mini, LLaMA 70B, and LLaMA 8B), we show that multilingual prompting consistently outperforms existing diversity-enhancing techniques such as hightemperature sampling, step-by-step recall, and persona prompting.Further analyses show that the benefits of multilingual prompting vary between high and low resource languages and across model sizes, and that aligning the prompting language with cultural cues reduces hallucination about culturally-specific information. Can you recommend some singersto follow?
Shidong Pan, Tal Linzen, Emily Black
EMNLP3
2025 What Goes Into a LM Acceptability Judgment? Rethinking the Impact of Frequency and Length
abstract
Lindia Tjuatja, Graham Neubig, Tal Linzen, Sophie Hao. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Lindia Tjuatja, Graham Neubig, Tal Linzen, Sophie Hao
NAACL (Long Papers)3
2025 Emergence of Linear Truth Encodings in Language Models
abstract
Recent probing studies reveal that large language models exhibit linear subspaces that separate true from false statements, yet the mechanism behind their emergence is unclear. We introduce a transparent, one-layer transformer toy model that reproduces such truth subspaces end-to-end and exposes one concrete route by which they can arise. We study one simple setting in which truth encoding can emerge: a data distribution where factual statements co-occur with other factual statements (and vice-versa), encouraging the model to learn this distinction in order to lower the LM loss on future tokens. We corroborate this pattern with experiments in pretrained language models. Finally, in the toy setting we observe a two-phase learning dynamic: networks first memorize individual factual associations in a few steps, then---over a longer horizon---learn to linearly separate true from false, which in turn lowers language-modeling loss. Together, these results provide both a mechanistic demonstration and an empirical motivation for how and why linear truth representations can emerge in language models.
Shauli Ravfogel, Gilad Yehudai, Tal Linzen, Joan Bruna, Alberto Bietti
NeurIPS3
2024 Insights from the first BabyLM Challenge: Training sample-efficient language models on a developmentally plausible corpus
Alex Warstadt, Aaron Mueller, Leshem Choshen, Ethan Wilcox, Chengxu Zhuang, Adina Williams, Ryan Cotterell, Tal Linzen
CogSci8
2024 A Systematic Comparison of Syllogistic Reasoning in Humans and Language Models
abstract
Tiwalayo Eisape, Michael Tessler, Ishita Dasgupta, Fei Sha, Sjoerd Steenkiste, Tal Linzen. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Tiwalayo Eisape, Michael Henry Tessler, Ishita Dasgupta 0001, Fei Sha, Sjoerd van Steenkiste, Tal Linzen
NAACL-HLT6
2024 In-context Learning Generalizes, But Not Always Robustly: The Case of Syntax
abstract
Aaron Mueller, Albert Webson, Jackson Petty, Tal Linzen. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Aaron Mueller, Albert Webson, Jackson Petty, Tal Linzen
NAACL-HLT4
2024 The Impact of Depth on Compositional Generalization in Transformer Language Models
abstract
Jackson Petty, Sjoerd Steenkiste, Ishita Dasgupta, Fei Sha, Dan Garrette, Tal Linzen. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Jackson Petty, Sjoerd van Steenkiste, Ishita Dasgupta 0001, Fei Sha, Dan Garrette, Tal Linzen
NAACL-HLT6
2024 Do Language Models' Words Refer?
abstract
Abstract What do language models (LMs) do with language? They can produce sequences of (mostly) coherent strings closely resembling English. But do those sentences mean something, or are LMs simply babbling in a convincing simulacrum of language use? We address one aspect of this broad question: whether LMs’ words can refer, that is, achieve “word-to-world” connections. There is prima facie reason to think they do not, since LMs do not interact with the world in the way that ordinary language users do. Drawing on the externalist tradition in philosophy of language, we argue that those appearances are misleading: Even if the inputs to LMs are simply strings of text, they are strings of text with natural histories, and that may suffice for LMs’ words to refer.
Matthew Mandelkern, Tal Linzen
Comput. Linguistics2
2023 How to Plant Trees in Language Models: Data and Architectural Effects on the Emergence of Syntactic Inductive Biases
abstract
Accurate syntactic representations are essential for robust generalization in natural language.Recent work has found that pre-training can teach language models to rely on hierarchical syntactic features-as opposed to incorrect linear features-when performing tasks after finetuning.We test what aspects of pre-training are important for endowing encoder-decoder Transformers with an inductive bias that favors hierarchical syntactic generalizations.We focus on architectural features (depth, width, and number of parameters), as well as the genre and size of the pre-training corpus, diagnosing inductive biases using two syntactic transformation tasks: question formation and passivization, both in English.We find that the number of parameters alone does not explain hierarchical generalization: model depth plays greater role than model width.We also find that pre-training on simpler language, such as child-directed speech, induces a hierarchical bias using an order-of-magnitude less data than pre-training on more typical datasets based on web text or Wikipedia; this suggests that in cognitively plausible language acquisition settings, neural language models may be more data-efficient than previously thought.
Aaron Mueller, Tal Linzen
ACL (1)2
2023 How poor is the stimulus? Evaluating hierarchical generalization in neural networks trained on child-directed speech
abstract
When acquiring syntax, children consistently choose hierarchical rules over competing nonhierarchical possibilities.Is this preference due to a learning bias for hierarchical structure, or due to more general biases that interact with hierarchical cues in children's linguistic input?We explore these possibilities by training LSTMs and Transformers-two types of neural networks without a hierarchical biason data similar in quantity and content to children's linguistic input: text from the CHILDES corpus.We then evaluate what these models have learned about English yes/no questions, a phenomenon for which hierarchical structure is crucial.We find that, though they perform well at capturing the surface statistics of childdirected speech (as measured by perplexity), both model types generalize in a way more consistent with an incorrect linear rule than the correct hierarchical rule.These results suggest that human-like generalization from text alone requires stronger biases than the general sequence-processing biases of standard neural network architectures.
Aditya Yedetore, Tal Linzen, Robert Frank 0001, Tom McCoy 0001
ACL (1)2
2023 SLOG: A Structural Generalization Benchmark for Semantic Parsing
abstract
The goal of compositional generalization benchmarks is to evaluate how well models generalize to new complex linguistic expressions.Existing benchmarks often focus on lexical generalization, the interpretation of novel lexical items in syntactic structures familiar from training.Structural generalization tasks, where a model needs to interpret syntactic structures that are themselves unfamiliar from training, are often underrepresented, resulting in overly optimistic perceptions of how well models can generalize.We introduce SLOG, a semantic parsing dataset that extends COGS (Kim and Linzen, 2020) with 17 structural generalization cases.In our experiments, the generalization accuracy of Transformer models, including pretrained ones, only reaches 40.6%, while a structure-aware parser only achieves 70.8%.These results are far from the near-perfect accuracy existing models achieve on COGS, demonstrating the role of SLOG in foregrounding the large discrepancy between models' lexical and structural generalization capacities.
Bingzhi Li, Lucia Donatelli, Alexander Koller, Tal Linzen, Yuekun Yao, Najoung Kim
EMNLP4
2023 How Much Do Language Models Copy From Their Training Data? Evaluating Linguistic Novelty in Text Generation Using RAVEN
abstract
Abstract Current language models can generate high-quality text. Are they simply copying text they have seen before, or have they learned generalizable linguistic abstractions? To tease apart these possibilities, we introduce RAVEN, a suite of analyses for assessing the novelty of generated text, focusing on sequential structure (n-grams) and syntactic structure. We apply these analyses to four neural language models trained on English (an LSTM, a Transformer, Transformer-XL, and GPT-2). For local structure—e.g., individual dependencies—text generated with a standard sampling scheme is substantially less novel than our baseline of human-generated text from each model’s test set. For larger-scale structure—e.g., overall sentence structure—model-generated text is as novel or even more novel than the human-generated baseline, but models still sometimes copy substantially, in some cases duplicating passages over 1,000 words long from the training set. We also perform extensive manual analysis, finding evidence that GPT-2 uses both compositional and analogical generalization mechanisms and showing that GPT-2’s novel text is usually well-formed morphologically and syntactically but has reasonably frequent semantic issues (e.g., being self-contradictory).
Tom McCoy 0001, Paul Smolensky, Tal Linzen, Jianfeng Gao 0001, Asli Celikyilmaz
Trans. Assoc. Comput. Linguistics3
2022 LSTMs Can Learn Basic Wh- and Relative Clause Dependencies in Norwegian
Anastasia Kobzeva, Suhas Arehalli, Tal Linzen, Dave Kush
CogSci3
2022 Syntactic Surprisal From Neural Models Predicts, But Underestimates, Human Processing Difficulty From Syntactic Ambiguities
abstract
Humans exhibit garden path effects: When reading sentences that are temporarily structurally ambiguous, they slow down when the structure is disambiguated in favor of the less preferred alternative.Surprisal theory (Hale, 2001;Levy, 2008), a prominent explanation of this finding, proposes that these slowdowns are due to the unpredictability of each of the words that occur in these sentences.Challenging this hypothesis, van Schijndel and Linzen (2021) find that estimates of the cost of word predictability derived from language models severely underestimate the magnitude of human garden path effects.In this work, we consider whether this underestimation is due to the fact that humans weight syntactic factors in their predictions more highly than language models do.We propose a method for estimating syntactic predictability from a language model, allowing us to weigh the cost of lexical and syntactic predictability independently.We find that treating syntactic predictability independently from lexical predictability indeed results in larger estimates of garden path.At the same time, even when syntactic predictability is independently weighted, surprisal still greatly underestimate the magnitude of human garden path effects.Our results support the hypothesis that predictability is not the only factor responsible for the processing cost associated with garden path sentences.
Suhas Arehalli, Brian Dillon, Tal Linzen
CoNLL3
2022 Characterizing Verbatim Short-Term Memory in Neural Language Models
abstract
When a language model is trained to predict natural language sequences, its prediction at each moment depends on a representation of prior context.What kind of information about the prior context can language models retrieve?We tested whether language models could retrieve the exact words that occurred previously in a text.In our paradigm, language models (transformers and an LSTM) processed English text in which a list of nouns occurred twice.We operationalized retrieval as the reduction in surprisal from the first to the second list.We found that the transformers retrieved both the identity and ordering of nouns from the first list.Further, the transformers' retrieval was markedly enhanced when they were trained on a larger corpus and with greater model depth.Lastly, their ability to index prior tokens was dependent on learned attention patterns.In contrast, the LSTM exhibited less precise retrieval, which was limited to list-initial tokens and to short intervening texts.The LSTM's retrieval was not sensitive to the order of nouns and it improved when the list was semantically coherent.We conclude that transformers implemented something akin to a working memory system that could flexibly retrieve individual token representations across arbitrary delays; conversely, the LSTM maintained a coarser and more rapidly-decaying semantic gist of prior tokens, weighted toward the earliest items.
Kristijan Armeni, Christopher J. Honey, Tal Linzen
CoNLL3
2022 Entailment Semantics Can Be Extracted from an Ideal Language Model
abstract
Language models are often trained on text alone, without additional grounding.There is debate as to how much of natural language semantics can be inferred from such a procedure.We prove that entailment judgments between sentences can be extracted from an ideal language model that has perfectly learned its target distribution, assuming the training sentences are generated by Gricean agents, i.e., agents who follow fundamental principles of communication from the linguistic theory of pragmatics.We also show entailment judgments can be decoded from the predictions of a language model trained on such Gricean data.Our results reveal a pathway for understanding the semantic information encoded in unlabeled linguistic data and a potential framework for extracting semantics from language models.
William Merrill, Alex Warstadt, Tal Linzen
CoNLL3
2022 Causal Analysis of Syntactic Agreement Neurons in Multilingual Language Models
abstract
Structural probing work has found evidence for latent syntactic information in pre-trained language models.However, much of this analysis has focused on monolingual models, and analyses of multilingual models have employed correlational methods that are confounded by the choice of probing tasks.In this study, we causally probe multilingual language models (XGLM and multilingual BERT) as well as monolingual BERT-based models across various languages; we do this by performing counterfactual perturbations on neuron activations and observing the effect on models' subjectverb agreement probabilities.We observe where in the model and to what extent syntactic agreement is encoded in each language.We find significant neuron overlap across languages in autoregressive multilingual language models, but not masked language models.We also find two distinct layer-wise effect patterns and two distinct sets of neurons used for syntactic agreement, depending on whether the subject and verb are separated by other tokens.Finally, we find that behavioral analyses of language models are likely underestimating how sensitive masked language models are to syntactic information.
Aaron Mueller, Tal Linzen
CoNLL3
2022 The MultiBERTs: BERT Reproductions for Robustness Analysis
Thibault Sellam, Steve Yadlowsky, Ian Tenney, Jason Wei, Naomi Saphra, Alexander D'Amour, Tal Linzen, Jasmijn Bastings, Iulia Turc, Jacob Eisenstein, Dipanjan Das 0001, Ellie Pavlick
ICLR7
2022 Improving Compositional Generalization with Latent Structure and Data Augmentation
abstract
Linlu Qiu, Peter Shaw, Panupong Pasupat, Pawel Nowak, Tal Linzen, Fei Sha, Kristina Toutanova. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Linlu Qiu, Peter Shaw 0004, Panupong Pasupat, Pawel Krzysztof Nowak, Tal Linzen, Fei Sha, Kristina Toutanova
NAACL-HLT5
2022 When a sentence does not introduce a discourse entity, Transformer-based models still sometimes refer to it
abstract
based models still sometimes refer to it
Sebastian Schuster 0001, Tal Linzen
NAACL-HLT2
2022 Evaluating Attribution in Dialogue Systems: The BEGIN Benchmark
abstract
Abstract Knowledge-grounded dialogue systems powered by large language models often generate responses that, while fluent, are not attributable to a relevant source of information. Progress towards models that do not exhibit this issue requires evaluation metrics that can quantify its prevalence. To this end, we introduce the Benchmark for Evaluation of Grounded INteraction (Begin), comprising 12k dialogue turns generated by neural dialogue systems trained on three knowledge-grounded dialogue corpora. We collect human annotations assessing the extent to which the models’ responses can be attributed to the given background information. We then use Begin to analyze eight evaluation metrics. We find that these metrics rely on spurious correlations, do not reliably distinguish attributable abstractive responses from unattributable ones, and perform substantially worse when the knowledge source is longer. Our findings underscore the need for more sophisticated and robust evaluation metrics for knowledge-grounded dialogue. We make Begin publicly available at https://github.com/google/BEGIN-dataset.
Nouha Dziri, Hannah Rashkin, Tal Linzen, David Reitter
Trans. Assoc. Comput. Linguistics3
2021 Causal Analysis of Syntactic Agreement Mechanisms in Neural Language Models
abstract
Matthew Finlayson, Aaron Mueller, Sebastian Gehrmann, Stuart Shieber, Tal Linzen, Yonatan Belinkov. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Matthew Finlayson, Aaron Mueller, Sebastian Gehrmann, Stuart M. Shieber, Tal Linzen, Yonatan Belinkov
ACL/IJCNLP (1)5
2021 NOPE: A Corpus of Naturally-Occurring Presuppositions in English
abstract
Alicia Parrish, Sebastian Schuster, Alex Warstadt, Omar Agha, Soo-Hwan Lee, Zhuoye Zhao, Samuel R. Bowman, Tal Linzen. Proceedings of the 25th Conference on Computational Natural Language Learning. 2021.
Alicia Parrish, Sebastian Schuster 0001, Alex Warstadt, Omar Agha, Soo-Hwan Lee, Zhuoye Zhao, Samuel R. Bowman, Tal Linzen
CoNLL8
2021 Counterfactual Interventions Reveal the Causal Effect of Relative Clause Representations on Agreement Prediction
abstract
When language models process syntactically complex sentences, do they use their representations of syntax in a manner that is consistent with the grammar of the language?We propose AlterRep, an intervention-based method to address this question.For any linguistic feature of a given sentence, AlterRep generates counterfactual representations by altering how the feature is encoded, while leaving intact all other aspects of the original representation.By measuring the change in a model's word prediction behavior when these counterfactual representations are substituted for the original ones, we can draw conclusions about the causal effect of the linguistic feature in question on the model's behavior.We apply this method to study how BERT models of different sizes process relative clauses (RCs).We find that BERT variants use RC boundary information during word prediction in a manner that is consistent with the rules of English grammar; this RC boundary information generalizes to a considerable extent across different RC types, suggesting that BERT represents RCs as an abstract linguistic category.
Shauli Ravfogel, Grusha Prasad, Tal Linzen, Yoav Goldberg
CoNLL3
2021 Frequency Effects on Syntactic Rule Learning in Transformers
abstract
Pre-trained language models perform well on a variety of linguistic tasks that require symbolic reasoning, raising the question of whether such models implicitly represent abstract symbols and rules.We investigate this question using the case study of BERT's performance on English subject-verb agreement.Unlike prior work, we train multiple instances of BERT from scratch, allowing us to perform a series of controlled interventions at pre-training time.We show that BERT often generalizes well to subject-verb pairs that never occurred in training, suggesting a degree of rule-governed behavior.We also find, however, that performance is heavily influenced by word frequency, with experiments showing that both the absolute frequency of a verb form, as well as the frequency relative to the alternate inflection, are causally implicated in the predictions BERT makes at inference time.Closer analysis of these frequency effects reveals that BERT's behavior is consistent with a system that correctly applies the SVA rule in general but struggles to overcome strong training priors and to estimate agreement features (singular vs. plural) on infrequent lexical items.
Jason Wei, Dan Garrette, Tal Linzen, Ellie Pavlick
EMNLP (1)3
2021 Predicting Inductive Biases of Pre-Trained Models
Charles Lovering, Rohan Jha, Tal Linzen, Ellie Pavlick
ICLR3
2020 Representations of Syntax [MASK] Useful: Effects of Constituency and Dependency Structure in Recursive LSTMs
abstract
Sequence-based neural networks show significant sensitivity to syntactic structure, but they still perform less well on syntactic tasks than tree-based networks.Such tree-based networks can be provided with a constituency parse, a dependency parse, or both.We evaluate which of these two representational schemes more effectively introduces biases for syntactic structure that increase performance on the subject-verb agreement prediction task.We find that a constituency-based network generalizes more robustly than a dependencybased one, and that combining the two types of structure does not yield further improvement.Finally, we show that the syntactic robustness of sequential models can be substantially improved by fine-tuning on a small amount of constructed data, suggesting that data augmentation is a viable alternative to explicit constituency structure for imparting the syntactic biases that sequential models are lacking.
Michael A. Lepori, Tal Linzen, Tom McCoy 0001
ACL2
2020 How Can We Accelerate Progress Towards Human-like Linguistic Generalization?
abstract
This position paper describes and critiques the Pretraining-Agnostic Identically Distributed (PAID) evaluation paradigm, which has become a central tool for measuring progress in natural language understanding.This paradigm consists of three stages: (1) pretraining of a word prediction model on a corpus of arbitrary size; (2) fine-tuning (transfer learning) on a training set representing a classification task; (3) evaluation on a test set drawn from the same distribution as that training set.This paradigm favors simple, low-bias architectures, which, first, can be scaled to process vast amounts of data, and second, can capture the fine-grained statistical properties of a particular data set, regardless of whether those properties are likely to generalize to examples of the task outside the data set.This contrasts with humans, who learn language from several orders of magnitude less data than the systems favored by this evaluation paradigm, and generalize to new tasks in a consistent way.We advocate for supplementing or replacing PAID with paradigms that reward architectures that generalize as quickly and robustly as humans.
Tal Linzen
ACL1
2020 Syntactic Data Augmentation Increases Robustness to Inference Heuristics
abstract
Pretrained neural models such as BERT, when fine-tuned to perform natural language inference (NLI), often show high accuracy on standard datasets, but display a surprising lack of sensitivity to word order on controlled challenge sets.We hypothesize that this issue is not primarily caused by the pretrained model's limitations, but rather by the paucity of crowdsourced NLI examples that might convey the importance of syntactic structure at the finetuning stage.We explore several methods to augment standard training sets with syntactically informative examples, generated by applying syntactic transformations to sentences from the MNLI corpus.The best-performing augmentation method, subject/object inversion, improved BERT's accuracy on controlled examples that diagnose sensitivity to word order from 0.28 to 0.73, without affecting performance on the MNLI test set.This improvement generalized beyond the particular construction used for data augmentation, suggesting that augmentation causes BERT to recruit abstract syntactic representations.
Junghyun Min, Tom McCoy 0001, Dipanjan Das 0001, Emily Pitler, Tal Linzen
ACL5
2020 Cross-Linguistic Syntactic Evaluation of Word Prediction Models
abstract
A range of studies have concluded that neural word prediction models can distinguish grammatical from ungrammatical sentences with high accuracy.However, these studies are based primarily on monolingual evidence from English.To investigate how these models' ability to learn syntax varies by language, we introduce CLAMS (Cross-Linguistic Assessment of Models on Syntax), a syntactic evaluation suite for monolingual and multilingual models.CLAMS includes subject-verb agreement challenge sets for English, French, German, Hebrew and Russian, generated from grammars we develop.We use CLAMS to evaluate LSTM language models as well as monolingual and multilingual BERT.Across languages, monolingual LSTMs achieved high accuracy on dependencies without attractors, and generally poor accuracy on agreement across object relative clauses.On other constructions, agreement accuracy was generally higher in languages with richer morphology.Multilingual models generally underperformed monolingual models.Multilingual BERT showed high syntactic accuracy on English, but noticeable deficiencies in other languages.
Aaron Mueller, Garrett Nicolai, Panayiota Petrou-Zeniou, Natalia Talmina, Tal Linzen
ACL5
2020 Neural Language Models Capture Some, But Not All Agreement Attraction Effects
Suhas Arehalli, Tal Linzen
CogSci2
2020 Universal linguistic inductive biases via meta-learning
Tom McCoy 0001, Erin Grant, Paul Smolensky, Thomas L. Griffiths 0001, Tal Linzen
CogSci5
2020 COGS: A Compositional Generalization Challenge Based on Semantic Interpretation
abstract
Natural language is characterized by compositionality: the meaning of a complex expression is constructed from the meanings of its constituent parts.To facilitate the evaluation of the compositional abilities of language processing architectures, we introduce COGS, a semantic parsing dataset based on a fragment of English.The evaluation portion of COGS contains multiple systematic gaps that can only be addressed by compositional generalization; these include new combinations of familiar syntactic structures, or new combinations of familiar words and familiar structures.In experiments with Transformers and LSTMs, we found that in-distribution accuracy on the COGS test set was near-perfect (96-99%), but generalization accuracy was substantially lower (16-35%) and showed high sensitivity to random seed (±6-8%).These findings indicate that contemporary standard NLP models are limited in their compositional generalization capacity, and position COGS as a good way to measure progress.
Najoung Kim, Tal Linzen
EMNLP (1)2
2020 Does Syntax Need to Grow on Trees? Sources of Hierarchical Inductive Bias in Sequence-to-Sequence Networks
abstract
Learners that are exposed to the same training data might generalize differently due to differing inductive biases. In neural network models, inductive biases could in theory arise from any aspect of the model architecture. We investigate which architectural factors affect the generalization behavior of neural sequence-to-sequence models trained on two syntactic tasks, English question formation and English tense reinflection. For both tasks, the training set is consistent with a generalization based on hierarchical structure and a generalization based on linear order. All architectural factors that we investigated qualitatively affected how models generalized, including factors with no clear connection to hierarchical structure. For example, LSTMs and GRUs displayed qualitatively different inductive biases. However, the only factor that consistently contributed a hierarchical bias across tasks was the use of a tree-structured model rather than a model with sequential recurrence, suggesting that human-like syntactic generalization requires architectural syntactic structure.
Tom McCoy 0001, Robert Frank 0001, Tal Linzen
Trans. Assoc. Comput. Linguistics3
2019 Right for the Wrong Reasons: Diagnosing Syntactic Heuristics in Natural Language Inference
abstract
A machine learning system can score well on a given test set by relying on heuristics that are effective for frequent example types but break down in more challenging cases.We study this issue within natural language inference (NLI), the task of determining whether one sentence entails another.We hypothesize that statistical NLI models may adopt three fallible syntactic heuristics: the lexical overlap heuristic, the subsequence heuristic, and the constituent heuristic.To determine whether models have adopted these heuristics, we introduce a controlled evaluation set called HANS (Heuristic Analysis for NLI Systems), which contains many examples where the heuristics fail.We find that models trained on MNLI, including BERT, a state-of-the-art model, perform very poorly on HANS, suggesting that they have indeed adopted these heuristics.We conclude that there is substantial room for improvement in NLI systems, and that the HANS dataset can motivate and measure progress in this area.
Tom McCoy 0001, Ellie Pavlick, Tal Linzen
ACL (1)3
2019 Human few-shot learning of compositional instructions
Brenden M. Lake, Tal Linzen, Marco Baroni
CogSci2
2019 How much harder are hard garden-path sentences than easy ones?
Grusha Prasad, Tal Linzen
CogSci2
2019 Using Priming to Uncover the Organization of Syntactic Representations in Neural Language Models
abstract
Neural language models (LMs) perform well on tasks that require sensitivity to syntactic structure.Drawing on the syntactic priming paradigm from psycholinguistics, we propose a novel technique to analyze the representations that enable such success.By establishing a gradient similarity metric between structures, this technique allows us to reconstruct the organization of the LMs' syntactic representational space.We use this technique to demonstrate that LSTM LMs' representations of different types of sentences with relative clauses are organized hierarchically in a linguistically interpretable manner, suggesting that the LMs track abstract properties of the sentence.
Grusha Prasad, Marten van Schijndel, Tal Linzen
CoNLL3
2019 Quantity doesn't buy quality syntax with neural language models
abstract
Marten van Schijndel, Aaron Mueller, Tal Linzen. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Marten van Schijndel, Aaron Mueller, Tal Linzen
EMNLP/IJCNLP (1)3
2019 RNNs implicitly implement tensor-product representations
Tom McCoy 0001, Tal Linzen, Ewan Dunbar, Paul Smolensky
ICLR (Poster)2
2019 Analyzing and interpreting neural networks for NLP: A report on the first BlackboxNLP workshop
abstract
Abstract The Empirical Methods in Natural Language Processing (EMNLP) 2018 workshop BlackboxNLP was dedicated to resources and techniques specifically developed for analyzing and understanding the inner-workings and representations acquired by neural models of language. Approaches included: systematic manipulation of input to neural networks and investigating the impact on their performance, testing whether interpretable knowledge can be decoded from intermediate representations acquired by neural networks, proposing modifications to neural network architectures to make their knowledge state or generated output more explainable, and examining the performance of networks on simplified or formal languages. Here we review a number of representative studies in each category.
Afra Alishahi, Grzegorz Chrupala, Tal Linzen
Nat. Lang. Eng.3
2018 Distinct patterns of syntactic agreement errors in recurrent networks and humans
Tal Linzen, Brian Leonard
CogSci1
2018 Revisiting the poverty of the stimulus: hierarchical generalization without a hierarchical bias in recurrent neural networks
Tom McCoy 0001, Robert Frank 0001, Tal Linzen
CogSci3
2018 Modeling garden path effects without explicit hierarchical syntax
Marten van Schijndel, Tal Linzen
CogSci2
2018 Targeted Syntactic Evaluation of Language Models
abstract
We present a dataset for evaluating the grammaticality of the predictions of a language model.We automatically construct a large number of minimally different pairs of English sentences, each consisting of a grammatical and an ungrammatical sentence.The sentence pairs represent different variations of structure-sensitive phenomena: subject-verb agreement, reflexive anaphora and negative polarity items.We expect a language model to assign a higher probability to the grammatical sentence than the ungrammatical one.In an experiment using this data set, an LSTM language model performed poorly on many of the constructions.Multi-task training with a syntactic objective (CCG supertagging) improved the LSTM's accuracy, but a large gap remained between its performance and the accuracy of human participants recruited online.This suggests that there is considerable room for improvement over LSTMs in capturing syntax in a language model.
Rebecca Marvin, Tal Linzen
EMNLP2
2018 A Neural Model of Adaptation in Reading
abstract
It has been argued that humans rapidly adapt their lexical and syntactic expectations to match the statistics of the current linguistic context.We provide further support to this claim by showing that the addition of a simple adaptation mechanism to a neural language model improves our predictions of human reading times compared to a non-adaptive model.We analyze the performance of the model on controlled materials from psycholinguistic experiments and show that it adapts not only to lexical items but also to abstract syntactic structures.
Marten van Schijndel, Tal Linzen
EMNLP2
2018 Colorless Green Recurrent Networks Dream Hierarchically
abstract
Kristina Gulordava, Piotr Bojanowski, Edouard Grave, Tal Linzen, Marco Baroni. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Kristina Gulordava, Piotr Bojanowski, Edouard Grave, Tal Linzen, Marco Baroni
NAACL-HLT4
2017 Prediction and uncertainty in an artificial language
Tal Linzen, Noam Siegelman, Louisa Bogaerts
CogSci1
2017 Exploring the Syntactic Abilities of RNNs with Multi-task Learning
abstract
Recent work has explored the syntactic abilities of RNNs using the subject-verb agreement task, which diagnoses sensitivity to sentence structure.RNNs performed this task well in common cases, but faltered in complex sentences (Linzen et al., 2016).We test whether these errors are due to inherent limitations of the architecture or to the relatively indirect supervision provided by most agreement dependencies in a corpus.We trained a single RNN to perform both the agreement task and an additional task, either CCG supertagging or language modeling.Multitask training led to significantly lower error rates, in particular on complex sentences, suggesting that RNNs have the ability to evolve more sophisticated syntactic representations than shown before.We also show that easily available agreement training data can improve performance on other syntactic tasks, in particular when only a limited amount of training data is available for those tasks.The multi-task paradigm can also be leveraged to inject grammatical knowledge into language models.
Émile Enguehard, Yoav Goldberg, Tal Linzen
CoNLL3
2016 Assessing the Ability of LSTMs to Learn Syntax-Sensitive Dependencies
abstract
The success of long short-term memory (LSTM) neural networks in language processing is typically attributed to their ability to capture long-distance statistical regularities. Linguistic regularities are often sensitive to syntactic structure; can such dependencies be captured by LSTMs, which do not have explicit structural representations? We begin addressing this question using number agreement in English subject-verb dependencies. We probe the architecture’s grammatical competence both using training objectives with an explicit grammatical target (number prediction, grammaticality judgments) and using language models. In the strongly supervised settings, the LSTM achieved very high overall accuracy (less than 1% errors), but errors increased when sequential and structural information conflicted. The frequency of such errors rose sharply in the language-modeling setting. We conclude that LSTMs can capture a non-trivial amount of grammatical structure given targeted supervision, but stronger architectures may be required to further reduce errors; furthermore, the language modeling signal is insufficient for capturing syntax-sensitive dependencies, and should be supplemented with more direct supervision if such dependencies need to be captured.
Tal Linzen, Emmanuel Dupoux, Yoav Goldberg
Trans. Assoc. Comput. Linguistics1
2015 A model of rapid phonotactic generalization
abstract
The phonotactics of a language describes the ways in which the sounds of the language combine to form possible morphemes and words.Humans can learn phonotactic patterns at the level of abstract classes, generalizing across sounds (e.g., "words can end in a voiced stop").Moreover, they rapidly acquire these generalizations, even before they acquire soundspecific patterns.We present a probabilistic model intended to capture this earlyabstraction phenomenon.The model represents both abstract and concrete generalizations in its hypothesis space from the outset of learning.This-combined with a parsimony bias in favor of compact descriptions of the input data-leads the model to favor rapid abstraction in a way similar to human learners.
Tal Linzen, Timothy J. O'Donnell
EMNLP1
2014 The timecourse of phonotactic learning
Tal Linzen, Gillian Gallagher
CogSci1