Roger Levy

dblp:23/90 · also Roger P. Levy · DBLP profile ↗
← Back
107ranked-venue papers
8as first author
45since 2021 · last 2026
0000-0002-4493-8864ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 105 · 8 first-author · 45 since 2021Applied, interdisciplinary, general and emerging computing · 52 · 1 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 Implicit Representations of Grammaticality in Language Models
abstract
Grammaticality and likelihood are distinct notions in human language.Pretrained language models (LMs), which are probabilistic models of language fitted to maximize corpus likelihood, generate grammatically well-formed text and discriminate well between grammatical and ungrammatical sentences in tightly controlled minimal pairs.However, their string probabilities do not sharply discriminate between grammatical and ungrammatical sentences overall.But do LMs implicitly acquire a grammaticality distinction distinct from string probability?We explore this question through studying internal representations of LMs, by training a linear probe on a dataset of grammatical and (synthetic) ungrammatical sentences obtained by applying perturbations to a naturalistic text corpus.We find that this simple grammaticality probe generalizes to human-curated grammaticality judgment benchmarks and outperforms LM probability-based grammaticality judgments.When applied to semantic plausibility benchmarks, in which both members of a minimal pair are grammatical and differ in only plausibility, the probe however performs worse than string probability.The English-trained probe also exhibits nontrivial cross-lingual generalization, outperforming string probabilities on grammaticality benchmarks in numerous other languages.Additionally, probe scores correlate only weakly with string probabilities.These results collectively suggest that LMs acquire to some extent an implicit grammaticality distinction within their hidden layers. 1
Yingshan Susan Wang, Linlu Qiu, Zhaofeng Wu, Roger Levy
ACL (1)4
2026 Readers make targeted regressions to plausible errors in reanalysis of "noisy-channel garden-path" sentences
abstract
A key question in psycholinguistics is how inferences about the meaning of linguistic input unfold incrementally a comprehender's mind.In this work, we study reading dynamics for "noisy-channel garden-path" sentences, which temporarily appear well-formed but feature late-appearing violations of expectation that can be resolved not by inferring an alternative syntactic structure, but by inferring the presence of an error.We find evidence for targeted regressions -eye movements towards regions that are promising loci of possible errors in light of later-arriving information, showing patterns consistent with the posterior inferences of a model of noisy-channel processing with reanalysis.We discuss the implications of these findings for theories of noisy-channel language comprehension and information-theoretic explanations of reading dynamics.
Thomas Hikaru Clark, Roger Levy, Edward Gibson
CoNLL2
2026 What Can String Probability Tell Us About Grammaticality?
abstract
Abstract What have language models (LMs) learned about grammar? This question remains hotly debated, with major ramifications for linguistic theory. However, since probability and grammaticality are distinct notions in linguistics, it is not obvious what string probabilities can reveal about an LM’s underlying grammatical knowledge. We present a theoretical analysis of the relationship between grammar, meaning, and string probability, based on simple assumptions about the generative process of corpus data. Our framework makes three predictions, which we validate empirically using 280K sentence pairs in English and Chinese: (1) correlation between the probability of strings within minimal pairs, i.e., string pairs with minimal semantic differences; (2) correlation between models’ and humans’ deltas within minimal pairs; and (3) poor separation in probability space between unpaired grammatical and ungrammatical strings. Our analyses give theoretical grounding for using probability to learn about LMs’ structural knowledge, and suggest directions for future work in LM grammatical evaluation.
Jennifer Hu 0001, Ethan Wilcox, Siyuan Song, Kyle Mahowald, Roger Levy
Trans. Assoc. Comput. Linguistics5
2025 A Model of Approximate and Incremental Noisy-Channel Language Processing
Thomas Hikaru Clark, Jacob Hoover Vigly, Edward Gibson, Roger Levy
CogSci4
2025 Surprisal and developmental sentence processing: exploring the role of language exposure through neural language models
Kuan-Jung Huang, Roger Levy, Yi Ting Huang
CogSci2
2025 Finding structure in logographic writing with library learning II: Grapheme, sound, and meaning systematicity
Guangyuan Jiang, Matthias Hofer 0002, Jiayuan Mao, Lionel Wong, Josh Tenenbaum, Roger Levy
CogSci6
2025 Efficient compression in locomotion verbs across languages
Thomas A. Langlois, Roger Levy, Nidhi Seethapathi, Noga Zaslavsky
CogSci2
2025 Resource-Rational Noisy-Channel Language Processing: Testing the Effect of Algorithmic Constraints on Inferences
abstract
Human language use is robust to errors: comprehenders can and do mentally correct utterances that are implausible or anomalous.How are humans able to solve these problems in real time, picking out alternatives from an unbounded space of options using limited cognitive resources?And can language models trained on next-word prediction for typical language be augmented to handle language anomalies in a human-like way?Using a language model as a prior and an error model to encode likelihoods, we use Sequential Monte Carlo with optional rejuvenation to perform incremental and approximate probabilistic inference over intended sentences and production errors.We demonstrate that the model captures previously established patterns in human sentence processing, and that a trade-off between human-like noisy-channel inferences and computational resources falls out of this model.From a psycholinguistic perspective, our results offer a candidate algorithmic model of rational inference in language processing.From an NLP perspective, our results showcase how to elicit human-like noisy-channel inference behavior from a relatively small LLM while controlling the amount of computation available during inference.Our model is implemented in the Gen.jl probabilistic programming language, and our code is available at https://github. com/thomashikaru/noisy_channel_model.
Thomas Hikaru Clark, Jacob Hoover Vigly, Edward Gibson, Roger Levy
EMNLP4
2025 On the Same Wavelength? Evaluating Pragmatic Reasoning in Language Models across Broad Concepts
abstract
Language use is shaped by pragmatics-i.e., reasoning about communicative goals and norms in context.As language models (LMs) are increasingly used as conversational agents, it becomes ever more important to understand their pragmatic reasoning abilities.We propose an evaluation framework derived from Wavelength, a popular communication game where a speaker and a listener communicate about a broad range of concepts in a granular manner.We study a range of LMs on both language comprehension and language production using direct and Chain-of-Thought (CoT) prompting, and further explore a Rational Speech Act (RSA) approach to incorporating Bayesian pragmatic reasoning into LM inference.We find that state-of-the-art LMs, but not smaller ones, achieve strong performance on language comprehension, obtaining similar-to-human accuracy and exhibiting high correlations with human judgments even without CoT prompting or RSA.On language production, CoT can outperform direct prompting, and using RSA provides significant improvements over both approaches.Our study helps identify the strengths and limitations in LMs' pragmatic reasoning abilities and demonstrates the potential for improving them with RSA, opening up future avenues for understanding conceptual representation, language understanding, and social reasoning in LMs and humans. 1 Left Concept (0) Target Value Right Concept (100) Human-written Clues Chosen Clue Human Mean Deep thought 10 Shallow thought Evolution, Solving complex problems, Chess, Einstein, Meditation, Quantum mechanics Solv.complex prob.
Linlu Qiu, Cedegao E. Zhang, Josh Tenenbaum, Roger Levy
EMNLP5
2025 Language Model Behavioral Phases are Consistent Across Architecture, Training Data, and Scale
abstract
We show that across architecture (Transformer vs. Mamba vs. RWKV), training dataset (OpenWebText vs. The Pile), and scale (14 million parameters to 12 billion parameters), autoregressive language models exhibit highly consistent patterns of change in their behavior over the course of pretraining. Based on our analysis of over 1,400 language model checkpoints on over 110,000 tokens of English, we find that up to 98% of the variance in language model behavior at the word level can be explained by three simple heuristics: the unigram probability (frequency) of a given word, the $n$-gram probability of the word, and the semantic similarity between the word and its context. Furthermore, we see consistent behavioral phases in all language models, with their predicted probabilities for words overfitting to those words' $n$-gram probabilities for increasing $n$ over the course of training. Taken together, these results suggest that learning in neural language models may follow a similar trajectory irrespective of model details.
James A. Michaelov, Roger Levy, Ben Bergen 0001
NeurIPS2
2025 Elements of World Knowledge (EWoK): A Cognition-Inspired Framework for Evaluating Basic World Knowledge in Language Models
Anna A. Ivanova, Aalok Sathe, Benjamin Lipkin, Unnathi U. Kumar, Setayesh Radkani, Thomas Hikaru Clark, Carina Kauf, Jennifer Hu 0001, R. T. Pramod, Gabriel Grand, Vivian C. Paulun, Maria Ryskina, Ekin Akyürek, Ethan Wilcox, Nafisa Rashid, Leshem Choshen, Roger Levy, Evelina Fedorenko, Josh Tenenbaum, Jacob Andreas
Trans. Assoc. Comput. Linguistics17
2024 Inferring errors and intended meanings with a generative model of language production in aphasia
Thomas Hikaru Clark, Edward Gibson, Roger Levy
CogSci3
2024 Finding structure in logographic writing with library learning
Guangyuan Jiang, Matthias Hofer 0002, Jiayuan Mao, Lionel Wong, Josh Tenenbaum, Roger Levy
CogSci6
2024 Multimodal Input Aids a Bayesian Model of Phonetic Learning
Sophia Zhi, Roger Levy, Stephan C. Meylan
CogSci2
2024 Bridging semantics and pragmatics in information-theoretic emergent communication
abstract
Human languages support both semantic categorization and local pragmatic interactions that require context-sensitive reasoning about meaning. While semantics and pragmatics are two fundamental aspects of language, they are typically studied independently and their co-evolution is largely under-explored. Here, we aim to bridge this gap by studying how a shared lexicon may emerge from local pragmatic interactions. To this end, we extend a recent information-theoretic framework for emergent communication in artificial agents, which integrates utility maximization, associated with pragmatics, with general communicative constraints that are believed to shape human semantic systems. Specifically, we show how to adapt this framework to train agents via unsupervised pragmatic interactions, and then evaluate their emergent lexical semantics. We test this approach in a rich visual domain of naturalistic images, and find that key human-like properties of the lexicon emerge when agents are guided by both context-specific utility and general communicative pressures, suggesting that both aspects are crucial for understanding how language may evolve in humans and in artificial agents.
Eleonora Gualdoni, Mycal Tucker, Roger Levy, Noga Zaslavsky
NeurIPS3
2023 Language model acceptability judgements are not always robust to context
abstract
Koustuv Sinha, Jon Gauthier, Aaron Mueller, Kanishka Misra, Keren Fuentes, Roger Levy, Adina Williams. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Koustuv Sinha, Jon Gauthier, Aaron Mueller, Kanishka Misra, Keren Fuentes, Roger Levy, Adina Williams
ACL (1)6
2023 Unsupervised Discontinuous Constituency Parsing with Mildly Context-Sensitive Grammars
abstract
We study grammar induction with mildly context-sensitive grammars for unsupervised discontinuous parsing.Using the probabilistic linear context-free rewriting system (LCFRS) formalism, our approach fixes the rule structure in advance and focuses on parameter learning with maximum likelihood.To reduce the computational complexity of both parsing and parameter estimation, we restrict the grammar formalism to binary LCFRS with fan-out two and further discard rules that require O(ℓ 6 ) time to parse, reducing inference to O(ℓ 5 ).We find that using a large number of nonterminals is beneficial and thus make use of tensor decomposition-based rank-space dynamic programming with an embedding-based parameterization of rule probabilities to scale up the number of nonterminals.Experiments on German and Dutch show that our approach is able to induce linguistically meaningful trees with continuous and discontinuous structures.
Roger Levy
ACL (1)2
2023 Simplicity and Informativeness in the Evolution of Combinatorial Structure
Matthias Hofer 0002, Simon Kirby, Roger Levy
CogSci3
2023 The neural dynamics of word recognition and integration
abstract
Listeners recognize and integrate words in rapid and noisy everyday speech by combining expectations about upcoming content with incremental sensory evidence.We present a computational model of word recognition which formalizes this perceptual process in Bayesian decision theory.We fit this model to explain scalp EEG signals recorded as subjects passively listened to a fictional story, revealing both the dynamics of the online auditory word recognition process and the neural correlates of the recognition and integration of words.
Jon Gauthier, Roger Levy
EMNLP2
2023 Prompting is not a substitute for probability measurements in large language models
abstract
Prompting is now a dominant method for evaluating the linguistic knowledge of large language models (LLMs).While other methods directly read out models' probability distributions over strings, prompting requires models to access this internal information by processing linguistic input, thereby implicitly testing a new type of emergent ability: metalinguistic judgment.In this study, we compare metalinguistic prompting and direct probability measurements as ways of measuring models' linguistic knowledge.Broadly, we find that LLMs' metalinguistic judgments are inferior to quantities directly derived from representations.Furthermore, consistency gets worse as the prompt query diverges from direct measurements of next-word probabilities.Our findings suggest that negative results relying on metalinguistic prompts cannot be taken as conclusive evidence that an LLM lacks a particular linguistic generalization.Our results also highlight the value that is lost with the move to closed APIs where access to probability distributions is limited.
Jennifer Hu 0001, Roger Levy
EMNLP2
2023 LINC: A Neurosymbolic Approach for Logical Reasoning by Combining Language Models with First-Order Logic Provers
abstract
Logical reasoning, i.e., deductively inferring the truth value of a conclusion from a set of premises, is an important task for artificial intelligence with wide potential impacts on science, mathematics, and society.While many prompting-based strategies have been proposed to enable Large Language Models (LLMs) to do such reasoning more effectively, they still appear unsatisfactory, often failing in subtle and unpredictable ways.In this work, we investigate the validity of instead reformulating such tasks as modular neurosymbolic programming, which we call LINC: Logical Inference via Neurosymbolic Computation.In LINC, the LLM acts as a semantic parser, translating premises and conclusions from natural language to expressions in first-order logic.These expressions are then offloaded to an external theorem prover, which symbolically performs deductive inference.Leveraging this approach, we observe significant performance gains on FOLIO and a balanced subset of ProofWriter for three different models in nearly all experimental conditions we evaluate.On ProofWriter, augmenting the comparatively small open-source StarCoder+ (15.5B parameters) with LINC even outperforms GPT-3.5 and GPT-4 with Chain-of-Thought (CoT) prompting by an absolute 38% and 10%, respectively.When used with GPT-4, LINC scores 26% higher than CoT on ProofWriter while performing comparatively on FOLIO.Further analysis reveals that although both methods on average succeed roughly equally often on this dataset, they exhibit distinct and complementary failure modes.We thus provide promising evidence for how logical reasoning over natural language can be tackled through jointly leveraging LLMs alongside symbolic provers.All corresponding code is publicly available.
Theo X. Olausson, Alex Gu, Benjamin Lipkin, Cedegao E. Zhang, Armando Solar-Lezama, Josh Tenenbaum, Roger Levy
EMNLP7
2023 Probing Self-supervised Speech Models for Phonetic and Phonemic Information: A Case Study in Aspiration
Kinan Martin, Jon Gauthier, Canaan Breiss, Roger Levy
INTERSPEECH4
2023 A Cross-Linguistic Pressure for Uniform Information Density in Word Order
abstract
Abstract While natural languages differ widely in both canonical word order and word order flexibility, their word orders still follow shared cross-linguistic statistical patterns, often attributed to functional pressures. In the effort to identify these pressures, prior work has compared real and counterfactual word orders. Yet one functional pressure has been overlooked in such investigations: The uniform information density (UID) hypothesis, which holds that information should be spread evenly throughout an utterance. Here, we ask whether a pressure for UID may have influenced word order patterns cross-linguistically. To this end, we use computational models to test whether real orders lead to greater information uniformity than counterfactual orders. In our empirical study of 10 typologically diverse languages, we find that: (i) among SVO languages, real word orders consistently have greater uniformity than reverse word orders, and (ii) only linguistically implausible counterfactual orders consistently exceed the uniformity of real orders. These findings are compatible with a pressure for information uniformity in the development and usage of natural languages.1
Thomas Hikaru Clark, Clara Meister, Tiago Pimentel, Michael Hahn 0001, Ryan Cotterell, Richard Futrell, Roger Levy
Trans. Assoc. Comput. Linguistics7
2023 Expectations over Unspoken Alternatives Predict Pragmatic Inferences
abstract
Abstract Scalar inferences (SI) are a signature example of how humans interpret language based on unspoken alternatives. While empirical studies have demonstrated that human SI rates are highly variable—both within instances of a single scale, and across different scales—there have been few proposals that quantitatively explain both cross- and within-scale variation. Furthermore, while it is generally assumed that SIs arise through reasoning about unspoken alternatives, it remains debated whether humans reason about alternatives as linguistic forms, or at the level of concepts. Here, we test a shared mechanism explaining SI rates within and across scales: context-driven expectations about the unspoken alternatives. Using neural language models to approximate human predictive distributions, we find that SI rates are captured by the expectedness of the strong scalemate as an alternative. Crucially, however, expectedness robustly predicts cross-scale variation only under a meaning-based view of alternatives. Our results suggest that pragmatic inferences arise from context-driven expectations over alternatives, and these expectations operate at the level of concepts.1
Jennifer Hu 0001, Roger Levy, Judith Degen, Sebastian Schuster 0001
Trans. Assoc. Comput. Linguistics2
2023 On the Effect of Anticipation on Reading Times
abstract
Abstract Over the past two decades, numerous studies have demonstrated how less-predictable (i.e., higher surprisal) words take more time to read. In general, these studies have implicitly assumed the reading process is purely responsive: Readers observe a new word and allocate time to process it as required. We argue that prior results are also compatible with a reading process that is at least partially anticipatory: Readers could make predictions about a future word and allocate time to process it based on their expectation. In this work, we operationalize this anticipation as a word’s contextual entropy. We assess the effect of anticipation on reading by comparing how well surprisal and contextual entropy predict reading times on four naturalistic reading datasets: two self-paced and two eye-tracking. Experimentally, across datasets and analyses, we find substantial evidence for effects of contextual entropy over surprisal on a word’s reading time (RT): In fact, entropy is sometimes better than surprisal in predicting a word’s RT. Spillover effects, however, are generally not captured by entropy, but only by surprisal. Further, we hypothesize four cognitive mechanisms through which contextual entropy could impact RTs—three of which we are able to design experiments to analyze. Overall, our results support a view of reading that is not just responsive, but also anticipatory.1
Tiago Pimentel, Clara Meister, Ethan Wilcox, Roger Levy, Ryan Cotterell
Trans. Assoc. Comput. Linguistics4
2023 Testing the Predictions of Surprisal Theory in 11 Languages
abstract
Abstract Surprisal theory posits that less-predictable words should take more time to process, with word predictability quantified as surprisal, i.e., negative log probability in context. While evidence supporting the predictions of surprisal theory has been replicated widely, much of it has focused on a very narrow slice of data: native English speakers reading English texts. Indeed, no comprehensive multilingual analysis exists. We address this gap in the current literature by investigating the relationship between surprisal and reading times in eleven different languages, distributed across five language families. Deriving estimates from language models trained on monolingual and multilingual corpora, we test three predictions associated with surprisal theory: (i) whether surprisal is predictive of reading times, (ii) whether expected surprisal, i.e., contextual entropy, is predictive of reading times, and (iii) whether the linking function between surprisal and reading times is linear. We find that all three predictions are borne out crosslinguistically. By focusing on a more diverse set of languages, we argue that these results offer the most robust link to date between information theory and incremental language processing across languages.
Ethan Wilcox, Tiago Pimentel, Clara Meister, Ryan Cotterell, Roger Levy
Trans. Assoc. Comput. Linguistics5
2022 Flexible Generation from Fragmentary Linguistic Input
abstract
The dominant paradigm for high-performance models in novel NLP tasks today is direct specialization for the task via training from scratch or fine-tuning large pre-trained models.But does direct specialization capture how humans approach novel language tasks?We hypothesize that human performance is better characterized by flexible inference through composition of basic computational motifs available to the human language user.To test this hypothesis, we formulate a set of novel fragmentary text completion tasks, and compare the behavior of three direct-specialization models against a new model we introduce, GibbsComplete, which composes two basic computational motifs central to contemporary models: masked and autoregressive word prediction.We conduct three types of evaluation: human judgments of completion quality, satisfaction of syntactic constraints imposed by the input fragment, and similarity to human behavior in the structural statistics of the completions.With no task-specific parameter tuning, GibbsComplete performs comparably to direct-specialization models in the first two evaluations, and outperforms all directspecialization models in the third evaluation.These results support our hypothesis that human behavior in novel language tasks and environments may be better characterized by flexible composition of basic computational motifs rather than by direct specialization. Generic representationEnd-to-end learning/adaptation Task-specific algorithm This model predicts the model predicts the next Show [blank] infill [SEP] how how to word next
Roger Levy
ACL (1)2
2022 The emergence of discrete and systematic communication in a continuous signal-meaning space
Alicia M. Chen, Matthias Hofer 0002, Moshe Poliak, Roger Levy, Noga Zaslavsky
CogSci4
2022 Evidence for Availability Effects on Speaker Choice in the Russian Comparative Alternation
Thomas Hikaru Clark, Ethan Wilcox, Edward Gibson, Roger Levy
CogSci4
2022 Rational Inference from Number Agreement Mismatch
Edward Gibson, Roger Levy
CogSci3
2022 Teasing apart models of pragmatics using optimal reference game design
Irene Zhou, Jennifer Hu 0001, Roger Levy, Noga Zaslavsky
CogSci3
2022 When Does Syntax Mediate Neural Language Model Performance? Evidence from Dropout Probes
abstract
Mycal Tucker, Tiwalayo Eisape, Peng Qian, Roger Levy, Julie Shah. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Mycal Tucker, Tiwalayo Eisape, Roger Levy, Julie A. Shah
NAACL-HLT4
2022 Trading off Utility, Informativeness, and Complexity in Emergent Communication
abstract
Emergent communication (EC) research often focuses on optimizing task-specific utility as a driver for communication. However, there is increasing evidence that human languages are shaped by task-general communicative constraints and evolve under pressure to optimize the Information Bottleneck (IB) tradeoff between the informativeness and complexity of the lexicon. Here, we integrate these two approaches by trading off utility, informativeness, and complexity in EC. To this end, we propose Vector-Quantized Variational Information Bottleneck (VQ-VIB), a method for training neural agents to encode inputs into discrete signals embedded in a continuous space. We evaluate our approach in multi-agent reinforcement learning settings and in color reference games and show that: (1) VQ-VIB agents can continuously adapt to changing communicative needs and, in the color domain, align with human languages; (2) the emergent VQ-VIB embedding spaces are semantically meaningful and perceptually grounded; and (3) encouraging informativeness leads to faster convergence rates and improved utility, both in VQ-VIB and in prior neural architectures for symbolic EC, with VQ-VIB achieving higher utility for any given complexity. This work offers a new framework for EC that is grounded in information-theoretic principles that are believed to characterize human language evolution and that may facilitate human-agent interaction.
Mycal Tucker, Roger Levy, Julie A. Shah, Noga Zaslavsky
NeurIPS2
2021 Structural Guidance for Transformer Language Models
abstract
Peng Qian, Tahira Naseem, Roger Levy, Ramón Fernandez Astudillo. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Tahira Naseem, Roger Levy, Ramón Fernandez Astudillo
ACL/IJCNLP (1)3
2021 A Targeted Assessment of Incremental Processing in Neural Language Models and Humans
abstract
Ethan Wilcox, Pranali Vani, Roger Levy. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Ethan Wilcox, Pranali Vani, Roger Levy
ACL/IJCNLP (1)3
2021 Learning Evolved Combinatorial Symbols with a Neuro-symbolic Generative Model
Matthias Hofer 0002, Tuan Anh Le 0001, Roger Levy, Josh Tenenbaum
CogSci3
2021 Eye Movement Traces of Linguistic Knowledge
Yevgeni Berzak, Roger Levy
CogSci2
2021 On Factors Influencing Typing Time: Insights from a Viral Online Typing Game
Roger Levy, Tiwalayo Eisape
CogSci2
2021 Competition from novel features drives scalar inferences in reference games
Jennifer Hu 0001, Noga Zaslavsky, Roger Levy
CogSci3
2021 Child-directed Listening: How Caregiver Inference Enables Children's Early Verbal Communication
Stephan C. Meylan, Ruthe Foushee, Elika Bergelson, Roger Levy
CogSci4
2021 Using the Interpolated Maze Task to Assess Incremental Processing in English Relative Clauses
Pranali Vani, Ethan Wilcox, Roger Levy
CogSci3
2021 Empirical Support for a Rate-Distortion Account of Pragmatic Reasoning
Irene Zhou, Jennifer Hu 0001, Roger Levy, Noga Zaslavsky
CogSci3
2021 Revisiting the Uniform Information Density Hypothesis
abstract
The uniform information density (UID) hypothesis posits a preference among language users for utterances structured such that information is distributed uniformly across a signal.While its implications on language production have been well explored, the hypothesis potentially makes predictions about language comprehension and linguistic acceptability as well.Further, it is unclear how uniformity in a linguistic signal-or lack thereof-should be measured, and over which linguistic unit, e.g., the sentence or language level, this uniformity should hold.Here we investigate these facets of the UID hypothesis using reading time and acceptability data.While our reading time results are generally consistent with previous work, they are also consistent with a weakly super-linear effect of surprisal, which would be compatible with UID's predictions.For acceptability judgments, we find clearer evidence that non-uniformity in information density is predictive of lower acceptability.We then explore multiple operationalizations of UID, motivated by different interpretations of the original hypothesis, and analyze the scope over which the pressure towards uniformity is exerted.The explanatory power of a subset of the proposed operationalizations suggests that the strongest trend may be a regression towards a mean surprisal across the language, rather than the phrase, sentence, or document-a finding that supports a typical interpretation of UID, namely that it is the byproduct of language users maximizing the use of a (hypothetical) communication channel. 1
Clara Meister, Tiago Pimentel, Patrick Haller 0001, Lena A. Jäger, Ryan Cotterell, Roger Levy
EMNLP (1)6
2021 Controlled Evaluation of Grammatical Knowledge in Mandarin Chinese Language Models
abstract
Prior work has shown that structural supervision helps English language models learn generalizations about syntactic phenomena such as subject-verb agreement.However, it remains unclear if such an inductive bias would also improve language models' ability to learn grammatical dependencies in typologically different languages.Here we investigate this question in Mandarin Chinese, which has a logographic, largely syllable-based writing system; different word order; and sparser morphology than English.We train LSTMs, Recurrent Neural Network Grammars, Transformer language models, and Transformerparameterized generative parsing models on two Mandarin Chinese datasets of different sizes.We evaluate the models' ability to learn different aspects of Mandarin grammar that assess syntactic and semantic relationships.We find suggestive evidence that structural supervision helps with representing syntactic state across intervening content and improves performance in low-data settings, suggesting that the benefits of hierarchical inductive biases in acquiring dependency relationships may extend beyond English.
Jennifer Hu 0001, Roger Levy
EMNLP (1)3
2021 Grammar-Based Grounded Lexicon Learning
abstract
We present Grammar-Based Grounded Language Learning (G2L2), a lexicalist approach toward learning a compositional and grounded meaning representation of language from grounded data, such as paired images and texts. At the core of G2L2 is a collection of lexicon entries, which map each word to a tuple of a syntactic type and a neuro-symbolic semantic program. For example, the word shiny has a syntactic type of adjective; its neuro-symbolic semantic program has the symbolic form $\lambda x.\textit{filter}(x, \textbf{SHINY})$, where the concept SHINY is associated with a neural network embedding, which will be used to classify shiny objects. Given an input sentence, G2L2 first looks up the lexicon entries associated with each token. It then derives the meaning of the sentence as an executable neuro-symbolic program by composing lexical meanings based on syntax. The recovered meaning programs can be executed on grounded inputs. To facilitate learning in an exponentially-growing compositional space, we introduce a joint parsing and expected execution algorithm, which does local marginalization over derivations to reduce the training time. We evaluate G2L2 on two domains: visual reasoning and language-driven navigation. Results show that G2L2 can generalize from small amounts of data to novel compositions of words.
Jiayuan Mao, Freda Shi, Jiajun Wu 0001, Roger Levy, Josh Tenenbaum
NeurIPS4
2020 STARC: Structured Annotations for Reading Comprehension
abstract
We present STARC (Structured Annotations for Reading Comprehension), a new annotation framework for assessing reading comprehension with multiple choice questions.Our framework introduces a principled structure for the answer choices and ties them to textual span annotations.The framework is implemented in OneStopQA, a new high-quality dataset for evaluation and analysis of reading comprehension in English.We use this dataset to demonstrate that STARC can be leveraged for a key new application for the development of SAT-like reading comprehension materials: automatic annotation quality probing via span ablation experiments.We further show that it enables in-depth analyses and comparisons between machine and human reading comprehension behavior, including error distributions and guessing ability.Our experiments also reveal that the standard multiple choice dataset in NLP, RACE (Lai et al., 2017), is limited in its ability to measure reading comprehension.47% of its questions can be guessed by machines without accessing the passage, and 18% are unanimously judged by humans as not having a unique correct answer.OneStopQA provides an alternative test set for reading comprehension which alleviates these shortcomings and has a substantially higher human ceiling performance.1
Yevgeni Berzak, Jonathan Malmaud, Roger Levy
ACL3
2020 A Systematic Assessment of Syntactic Generalization in Neural Language Models
abstract
While state-of-the-art neural network models continue to achieve lower perplexity scores on language modeling benchmarks, it remains unknown whether optimizing for broad-coverage predictive performance leads to human-like syntactic knowledge.Furthermore, existing work has not provided a clear picture about the model properties required to produce proper syntactic generalizations.We present a systematic evaluation of the syntactic knowledge of neural language models, testing 20 combinations of model types and data sizes on a set of 34 English-language syntactic test suites.We find substantial differences in syntactic generalization performance by model architecture, with sequential models underperforming other architectures.Factorially manipulating model architecture and training dataset size (1M-40M words), we find that variability in syntactic generalization performance is substantially greater by architecture than by dataset size for the corpora tested in our experiments.Our results also reveal a dissociation between perplexity and syntactic generalization performance.
Jennifer Hu 0001, Jon Gauthier, Ethan Wilcox, Roger Levy
ACL5
2020 Hierarchical Inferences Support Systematicity in the Lexicon
Matthias Hofer 0002, Tessa Verhoef, Roger Levy
CogSci3
2020 Jointly learning motion verbs and frame semantics from natural language and grounded scenes
Jon Gauthier, Jiayuan Mao, Tianmin Shu, Roger Levy, Josh Tenenbaum
CogSci4
2020 Children's Expressive and Receptive Knowledge of the English Regular Plural
Stephan C. Meylan, Roger Levy, Elika Bergelson
CogSci2
2020 Integrating Semantics Into Developmental Models of Morphology Learning
Abi Tenenbaum, Mika Braginsky, Roger Levy
CogSci3
2020 Informational goals, sentence structure, and comparison class inference
Michael Henry Tessler, Polina Tsvilodub, Jesse Snedeker, Roger Levy
CogSci4
2020 On the Predictive Power of Neural Language Models for Human Real-Time Comprehension Behavior
Ethan Wilcox, Jon Gauthier, Jennifer Hu 0001, Roger Levy
CogSci5
2020 Cloze Distillation: Improving Neural Language Models with Human Next-Word Prediction
abstract
Contemporary autoregressive language models (LMs) trained purely on corpus data have been shown to capture numerous features of human incremental processing.However, past work has also suggested dissociations between corpus probabilities and human next-word predictions.Here we evaluate several state-of-theart language models for their match to human next-word predictions and to reading time behavior from eye movements.We then propose a novel method for distilling the linguistic information implicit in human linguistic predictions into pre-trained LMs: Cloze Distillation.We apply this method to a baseline neural LM and show potential improvement in reading time prediction and generalization to held-out human cloze data.
Tiwalayo Eisape, Noga Zaslavsky, Roger Levy
CoNLL3
2020 Bridging Information-Seeking Human Gaze and Machine Reading Comprehension
abstract
In this work, we analyze how human gaze during reading comprehension is conditioned on the given reading comprehension question, and whether this signal can be beneficial for machine reading comprehension.To this end, we collect a new eye-tracking dataset with a large number of participants engaging in a multiple choice reading comprehension task.Our analysis of this data reveals increased fixation times over parts of the text that are most relevant for answering the question.Motivated by this finding, we propose making automated reading comprehension more human-like by mimicking human information-seeking reading behavior during reading comprehension.We demonstrate that this approach leads to performance gains on multiple choice question answering in English for a state-of-the-art reading comprehension model.
Jonathan Malmaud, Roger Levy, Yevgeni Berzak
CoNLL2
2020 Structural Supervision Improves Few-Shot Learning and Syntactic Generalization in Neural Language Models
abstract
Humans can learn structural properties about a word from minimal experience, and deploy their learned syntactic representations uniformly in different grammatical contexts. We assess the ability of modern neural language models to reproduce this behavior in English and evaluate the effect of structural supervision on learning outcomes. First, we assess few-shot learning capabilities by developing controlled experiments that probe models' syntactic nominal number and verbal argument structure generalizations for tokens seen as few as two times during training. Second, we assess invariance properties of learned representation: the ability of a model to transfer syntactic generalizations from a base context (e.g., a simple declarative active-voice sentence) to a transformed context (e.g., an interrogative sentence). We test four models trained on the same dataset: an n-gram baseline, an LSTM, and two LSTM-variants trained with explicit structural supervision (Dyer et al.,2016; Charniak et al., 2016). We find that in most cases, the neural models are able to induce the proper syntactic generalizations after minimal exposure, often from just two examples during training, and that the two structurally supervised models generalize more accurately than the LSTM model. All neural models are able to leverage information learned in base contexts to drive expectations in transformed contexts, indicating that they have learned some invariance properties of syntax.
Ethan Wilcox, Richard Futrell, Ryosuke Kohita, Roger Levy, Miguel Ballesteros
EMNLP (1)5
2019 Query-guided visual search
Junyi Chu, Jon Gauthier, Roger Levy, Josh Tenenbaum, Laura Schulz
CogSci3
2019 A rational model of syntactic bootstrapping
Jon Gauthier, Roger Levy, Josh Tenenbaum
CogSci2
2019 Iconicity and Structure in the Emergence of Combinatoriality
Matthias Hofer 0002, Roger Levy
CogSci2
2019 Inferring Structured Visual Concepts from Minimal Data
Luke B. Hewitt, Josh Tenenbaum, Roger Levy
CogSci4
2019 Incorporating Semantic Constraints into Algorithms for Unsupervised Learning of Morphology
Abi Tenenbaum, Roger Levy
CogSci2
2019 Incremental understanding of conjunctive generic sentences
Michael Henry Tessler, Karen Gu, Roger Levy
CogSci3
2019 What Syntactic Structures block Dependencies in RNN Language Models?
Ethan Wilcox, Roger Levy, Richard Futrell
CogSci2
2019 Availability-Based Production Predicts Speakers' Real-time Choices of Mandarin Classifiers
Meilin Zhan, Roger Levy
CogSci2
2019 Representation of Constituents in Neural Language Models: Coordination Phrase as a Case Study
abstract
Aixiu An, Peng Qian, Ethan Wilcox, Roger Levy. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Aixiu An, Ethan Wilcox, Roger Levy
EMNLP/IJCNLP (1)4
2019 Linking artificial and human neural representations of language
abstract
Jon Gauthier, Roger Levy. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Jon Gauthier, Roger Levy
EMNLP/IJCNLP (1)2
2018 Word learning and the acquisition of syntactic-semantic overhypotheses
Jon Gauthier, Roger Levy, Josh Tenenbaum
CogSci2
2018 Inductive Biases in the Evolution of Combinatorial Structure in Language
Matthias Hofer 0002, Josh Tenenbaum, Roger Levy
CogSci3
2018 Pragmatic Inference of Intended Referents from Binomial Word Order
Anna A. Ivanova, Roger Levy
CogSci2
2018 Communicative Efficiency, Uniform Information Density, and the Rational Speech Act Theory
Roger Levy
CogSci1
2018 Comparing Theories of Speaker Choice Using Classifier Production in Mandarin Chinese
Meilin Zhan, Roger Levy
CogSci2
2018 Comparing Models of Associative Meaning: An Empirical Investigation of Reference in Simple Language Games
abstract
Simple reference games (Wittgenstein, 1953) are of central theoretical and empirical importance in the study of situated language use.Although language provides rich, compositional truth-conditional semantics to facilitate reference, speakers and listeners may sometimes lack the overall lexical and cognitive resources to guarantee successful reference through these means alone.However, language also has rich associational structures that can serve as a further resource for achieving successful reference.Here we investigate this use of associational information in a setting where only associational information is available: a simplified version of the popular game Codenames.Using optimal experiment design techniques, we compare a range of models varying in the type of associative information deployed and in level of pragmatic sophistication against human behavior.In this setting we find that listeners' behavior reflects direct bigram collocational associations more strongly than word-embedding or semantic knowledge graph-based associations and that there is little evidence for pragmatically sophisticated behavior by either speakers or listeners of the type that might be predicted by recursive-reasoning models such as the Rational Speech Acts theory.These results shed light on the nature of the lexical resources that speakers and listeners can bring to bear in achieving reference through associative meaning alone.
Judy Hanwen Shen, Matthias Hofer 0002, Bjarke Felbo, Roger Levy
CoNLL4
2018 Assessing Language Proficiency from Eye Movements in Reading
abstract
Yevgeni Berzak, Boris Katz, Roger Levy. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Yevgeni Berzak, Boris Katz, Roger Levy
NAACL-HLT3
2018 Comparing Theories of Speaker Choice Using a Model of Classifier Production in Mandarin Chinese
abstract
Speakers often have more than one way to express the same meaning.What general principles govern speaker choice in the face of optionality when near semantically invariant alternation exists?Studies have shown that optional reduction in language is sensitive to contextual predictability, such that the more predictable a linguistic unit is, the more likely it is to get reduced.Yet it is unclear to what extent these cases of speaker choice are driven by audience design versus toward facilitating production.Here we argue that for a different optionality phenomenon, namely classifier choice in Mandarin Chinese, Uniform Information Density and at least one plausible variant of availability-based production make opposite predictions regarding the relationship between the predictability of the upcoming material and speaker choices.In a corpus analysis of Mandarin Chinese, we show that the distribution of speaker choices supports the availability-based production account, and not Uniform Information Density.
Meilin Zhan, Roger Levy
NAACL-HLT2
2017 Modeling Sources of Uncertainty in Spoken Word Learning
Matthias Hofer 0002, Roger Levy
CogSci2
2017 Noisy-context surprisal as a human sentence processing cost model
abstract
We use the noisy-channel theory of human sentence comprehension to develop an incremental processing cost model that unifies and extends key features of expectation-based and memory-based models.In this model, which we call noisy-context surprisal, the processing cost of a word is the surprisal of the word given a noisy representation of the preceding context.We show that this model accounts for an outstanding puzzle in sentence comprehension, language-dependent structural forgetting effects (Gibson and Thomas, 1999;Vasishth et al., 2010;Frank et al., 2016), which are previously not well modeled by either expectation-based or memory-based approaches.Additionally, we show that this model derives and generalizes locality effects (Gibson, 1998;Demberg and Keller, 2008), a signature prediction of memory-based models.We give corpusbased evidence for a key assumption in this derivation.
Richard Futrell, Roger Levy
EACL (1)2
2016 Finding Non-Arbitrary Form-Meaning Systematicity Using String-Metric Learning for Kernel Regression
abstract
Arbitrariness of the sign-the notion that the forms of words are unrelated to their meanings-is an underlying assumption of many linguistic theories.Two lines of research have recently challenged this assumption, but they produce differing characterizations of non-arbitrariness in language.Behavioral and corpus studies have confirmed the validity of localized form-meaning patterns manifested in limited subsets of the lexicon.Meanwhile, global (lexicon-wide) statistical analyses instead find diffuse form-meaning systematicity across the lexicon as a whole.We bridge the gap with an approach that can detect both local and global formmeaning systematicity in language.In the kernel regression formulation we introduce, form-meaning relationships can be used to predict words' distributional semantic vectors from their forms.Furthermore, we introduce a novel metric learning algorithm that can learn weighted edit distances that minimize kernel regression error.Our results suggest that the English lexicon exhibits far more global form-meaning systematicity than previously discovered, and that much of this systematicity is focused in localized formmeaning patterns.
E. Dario Gutiérrez, Roger Levy, Ben Bergen 0001
ACL (1)2
2016 Structure-sensitive Noise Inference: Comprehenders Expect Exchange Errors
Till Poppels, Roger Levy
CogSci2
2016 Bayesian Pronoun Interpretation in Mandarin Chinese
Meilin Zhan, Roger Levy, Andrew Kehler
CogSci2
2016 Data-driven learning of symbolic constraints for a log-linear model in a phonological setting
abstract
We propose a non-parametric Bayesian model for learning and weighting symbolically-defined constraints to populate a log-linear model. The model jointly infers a vector of binary constraint values for each candidate output and likely definitions for these constraints, combining observations of the output classes with a (potentially infinite) grammar over potential constraint definitions. We present results on a small morphophonological system, English regular plurals, as a test case. The inferred constraints, based on a grammar of articulatory features, perform as well as theoretically-defined constraints on both observed and novel forms of English regular plurals. The learned constraint values and definitions also closely resemble standard constraints defined within phonological theory.
Gabriel Doyle, Roger Levy
COLING2
2015 Modeling idiosyncratic preferences: How generative knowledge and expression frequency jointly determine language structure
Emily Morgan, Roger Levy
CogSci2
2014 Nonparametric Learning of Phonological Constraints in Optimality Theory
abstract
We present a method to jointly learn fea-tures and weights directly from distri-butional data in a log-linear framework. Specifically, we propose a non-parametric Bayesian model for learning phonologi-cal markedness constraints directly from the distribution of input-output mappings in an Optimality Theory (OT) setting. The model uses an Indian Buffet Process prior to learn the feature values used in the log-linear method, and is the first algorithm for learning phonological constraints with-out presupposing constraint structure. The model learns a system of constraints that explains observed data as well as the phonologically-grounded constraints of a standard analysis, with a violation struc-ture corresponding to the standard con-straints. These results suggest an alterna-tive data-driven source for constraints in-stead of a fully innate constraint set. 1
Gabriel Doyle, Klinton Bicknell, Roger Levy
ACL (1)3
2014 On the Role of Correlation and Abstraction in Cross-Modal Multimedia Retrieval
abstract
The problem of cross-modal retrieval from multimedia repositories is considered. This problem addresses the design of retrieval systems that support queries across content modalities, for example, using an image to search for texts. A mathematical formulation is proposed, equating the design of cross-modal retrieval systems to that of isomorphic feature spaces for different content modalities. Two hypotheses are then investigated regarding the fundamental attributes of these spaces. The first is that low-level cross-modal correlations should be accounted for. The second is that the space should enable semantic abstraction. Three new solutions to the cross-modal retrieval problem are then derived from these hypotheses: correlation matching (CM), an unsupervised method which models cross-modal correlations, semantic matching (SM), a supervised technique that relies on semantic representation, and semantic correlation matching (SCM), which combines both. An extensive evaluation of retrieval performance is conducted to test the validity of the hypotheses. All approaches are shown successful for text retrieval in response to image queries and vice versa. It is concluded that both hypotheses hold, in a complementary form, although evidence in favor of the abstraction hypothesis is stronger than that for correlation.
José Costa Pereira, Emanuele Coviello, Gabriel Doyle, Nikhil Rasiwasia, Gert R. G. Lanckriet, Roger Levy, Nuno Vasconcelos
IEEE Trans. Pattern Anal. Mach. Intell.6
2013 Evidence for cognitively controlled saccade targeting in reading
Klinton Bicknell, Emily C. Higgins, Roger Levy, Keith Rayner
CogSci3
2013 The Funny Thing About Incongruity: A Computational Model of Humor in Puns
Justine T. Kao, Roger Levy, Noah D. Goodman
CogSci2
2013 Modeling the Development of Determiner Productivity in Children's Early Speech
Stephan C. Meylan, Michael C. Frank, Roger Levy
CogSci3
2013 Combining multiple information types in Bayesian word segmentation
Gabriel Doyle, Roger Levy
HLT-NAACL2
2012 That's what she (could have) said: How alternative utterances affect language use
Leon Bergen, Noah D. Goodman, Roger Levy
CogSci3
2012 Verb omission errors: Evidence of rational processing of noisy language inputs
Leon Bergen, Roger Levy, Edward Gibson
CogSci2
2012 Word predictability and frequency effects in a rational model of reading
Klinton Bicknell, Roger Levy
CogSci2
2012 Can native-language perceptual bias facilitate learning words in a new language?
Bozena Pajak, Sarah C. Creel, Roger Levy
CogSci3
2011 Integrating surprisal and uncertain-input models in online sentence comprehension: formal techniques and empirical results
Roger Levy
ACL1
2011 Automated Whole Sentence Grammar Correction Using a Noisy Channel Model
Y. Albert Park, Roger Levy
ACL2
2011 Why readers regress to previous words: A statistical analysis
Klinton Bicknell, Roger Levy
CogSci2
2011 Phonological generalization from distributional evidence
Bozena Pajak, Roger Levy
CogSci2
2011 Cloze but no cigar: The complex relationship between cloze, corpus, and subjective probabilities in language processing
Nathaniel Smith, Roger Levy
CogSci2
2010 A Rational Model of Eye Movement Control in Reading
Klinton Bicknell, Roger Levy
ACL2
2010 A new approach to cross-modal multimedia retrieval
abstract
The problem of joint modeling the text and image components of multimedia documents is studied. The text component is represented as a sample from a hidden topic model, learned with latent Dirichlet allocation, and images are represented as bags of visual (SIFT) features. Two hypotheses are investigated: that 1) there is a benefit to explicitly modeling correlations between the two components, and 2) this modeling is more effective in feature spaces with higher levels of abstraction. Correlations between the two components are learned with canonical correlation analysis. Abstraction is achieved by representing text and images at a more general, semantic level. The two hypotheses are studied in the context of the task of cross-modal document retrieval. This includes retrieving the text that most closely matches a query image, or retrieving the images that most closely match a query text. It is shown that accounting for cross-modal correlations and semantic abstraction both improve retrieval accuracy. The cross-modal model is also shown to outperform state-of-the-art image retrieval systems on a unimodal retrieval task.
Nikhil Rasiwasia, José Costa Pereira, Emanuele Coviello, Gabriel Doyle, Gert R. G. Lanckriet, Roger Levy, Nuno Vasconcelos
ACM Multimedia6
2009 A model of local coherence effects in human sentence processing as consequences of updates from bottom-up prior to posterior beliefs
Klinton Bicknell, Roger Levy
HLT-NAACL2
2009 Minimal-length linearizations for mildly context-sensitive dependency trees
Y. Albert Park, Roger Levy
HLT-NAACL2
2008 A Noisy-Channel Model of Human Sentence Comprehension under Uncertain Input
Roger Levy
EMNLP1
2008 Modeling the effects of memory on human online sentence processing with particle filters
abstract
Language comprehension in humans is significantly constrained by memory, yet rapid, highly incremental, and capable of utilizing a wide range of contextual information to resolve ambiguity and form expectations about future input. In contrast, most of the leading psycholinguistic models and fielded algorithms for natural language parsing are non-incremental, have run time superlinear in input length, and/or enforce structural locality constraints on probabilistic dependencies between events. We present a new limited-memory model of sentence comprehension which involves an adaptation of the particle filter, a sequential Monte Carlo method, to the problem of incremental parsing. We show that this model can reproduce classic results in online sentence comprehension, and that it naturally provides the first rational account of an outstanding problem in psycholinguistics, in which the preferred alternative in a syntactic ambiguity seems to grow more attractive over time even in the absence of strong disambiguating information.
Roger Levy, Florencia Reali, Thomas L. Griffiths 0001
NIPS1
2006 Tregex and Tsurgeon: tools for querying and manipulating tree data structures
Roger Levy, Galen Andrew
LREC1
2006 Speakers optimize information density through syntactic reduction
abstract
If language users are rational, they might choose to structure their utterances so as to optimize communicative properties. In particular, information-theoretic and psycholinguistic considerations suggest that this may include maximizing the uniformity of information density in an utterance. We investigate this possibility in the context of syntactic reduction, where the speaker has the option of either marking a higher-order unit (a phrase) with an extra word, or leaving it unmarked. We demonstrate that speakers are more likely to reduce less information-dense phrases. In a second step, we combine a stochastic model of structured utterance production with a logistic-regression model of syntactic reduction to study which types of cues speakers employ when estimating the predictability of upcoming elements. We demonstrate that the trend toward predictability-sensitive syntactic reduction (Jaeger, 2006) is robust in the face of a wide variety of control variables, and present evidence that speakers use both surface and structural cues for predictability estimation.
Roger Levy, T. Florian Jaeger
NIPS1
2004 Deep Dependencies from Context-Free Statistical Parsers: Correcting the Surface Dependency Approximation
abstract
We present a linguistically-motivated algorithm for reconstructing nonlocal dependency in broad-coverage context-free parse trees derived from treebanks. We use an algorithm based on loglinear classifiers to augment and reshape context-free trees so as to reintroduce underlying nonlocal dependencies lost in the context-free approximation. We find that our algorithm compares favorably with prior work on English using an existing evaluation metric, and also introduce and argue for a new dependency-based evaluation metric. By this new evaluation metric our algorithm achieves 60% error reduction on gold-standard input trees and 5% error reduction on state-of-the-art machine-parsed input trees, when compared with the best previous work. We also present the first results on non-local dependency reconstruction for a language other than English, comparing performance on English and German. Our new evaluation metric quantitatively corroborates the intuition that in a language with freer word order, the surface dependencies in context-free parse trees are a poorer approximation to underlying dependency structure.
Roger Levy, Christopher D. Manning
ACL1
2003 Is it Harder to Parse Chinese, or the Chinese Treebank?
abstract
L¼ ¥ S " " h S 9{ | ¦S t 9 w{ ¥¬ w .
Roger Levy, Christopher D. Manning
ACL1
2003 A Generative Model for Semantic Role Labeling
Cynthia A. Thompson, Roger Levy, Christopher D. Manning
ECML2