VLDB 2026 Research / reviewers in the wild / expert
Ethan Wilcox
dblp:227/3505 · also Ethan Gotlieb Wilcox
· DBLP profile ↗
37ranked-venue papers
10as first author
29since 2021 · last 2026
0000-0002-5128-9890ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 37 · 10 first-author · 29 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 5 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Dual Alignment Between Language Model Layers and Human Sentence ProcessingabstractA recent study (Kuribayashi et al., 2025) has shown that human sentence processing behavior, typically measured on syntactically unchallenging constructions, can be effectively modeled using surprisal from early layers of large language models (LLMs).This raises the question of whether such advantages of internal layers extend to more syntactically challenging constructions, where surprisal has been reported to underestimate human cognitive effort.In this paper, we begin by exploring internal layers that better estimate human cognitive effort observed in syntactic ambiguity processing in English.Our experiments show that, in contrast to naturalistic reading, later layers better estimate such a cognitive effort, but still underestimate the human data.This dual alignment sheds light on different modes of sentence processing in humans and LMs: naturalistic reading employs a somewhat weak prediction akin to earlier layers of LMs, while syntactically challenging processing requires more fully-contextualized representations, better modeled by later layers of LMs.Motivated by these findings, we also explore several probability-update measures using shallow and deep layers of LMs, showing a complementary advantage to single-layer's surprisal in reading time modeling. https://github.com/kuribayashi4/ internal_surprisal_targeted_assessmentPhenomena Example MVRR D + : The girl fed the lamb remained relatively calm before the sunset in silence.D -: The girl who was fed the lamb remained relatively calm before the sunset in silence.NPS D + : The girl found the lamb remained relatively calm near the wooden fence.D -: The girl found that the lamb remained relatively calm near the wooden fence.NPZ D + : When the girl attacked the lamb remained relatively calm despite the sudden noise.D -: When the girl attacked, the lamb remained relatively calm despite the sudden noise.RC D + : The bus driver that the kids followed waited patiently at dawn.D -: The bus driver that followed the kids waited patiently at dawn.Attachment D + : Janet charmed the executive of the assistants who decides almost everything during long weekly meetings. Tatsuki Kuribayashi, Alex Warstadt, Yohei Oseki, Ethan Wilcox |
ACL (1) | 4 |
| 2026 | What Do Prosody and Text Convey? Characterizing How Meaningful Information is Distributed Across Multiple ChannelsabstractAditya Yadavalli, Tiago Pimentel, Tamar I Regev, Ethan Gotlieb Wilcox, Alex Warstadt. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Aditya Yadavalli, Tiago Pimentel, Tamar I. Regev, Ethan Wilcox, Alex Warstadt |
ACL (1) | 4 |
| 2026 | Function Words as Statistical Cues for Language LearningabstractWhat statistical properties might support learning abstract grammatical knowledge from linear input?We address this question by examining the statistical distribution of function words.Function words have been argued to aid acquisition through three distributional properties: high frequency, reliable syntactic association, and phrase-boundary alignment.We conduct a cross-linguistic corpus analysis of 186 languages, which confirms that all three properties are universal.Using counterfactual language modeling and ablation experiments on English, we show that preserving these properties facilitates acquisition in neural learners, with a Goldilocks effect: function words must be frequent enough to be reliable, yet diverse enough to remain informative to structural dependency.Probing analyses further reveal that different learning conditions produce systematically different reliance on function words. 1 Xiulin Yang, Heidi Getz, Ethan Wilcox |
ACL (1) | 3 |
| 2026 | Information-Theoretic Storage Cost in Sentence ComprehensionabstractReal-time sentence comprehension imposes a significant load on working memory, as comprehenders must maintain contextual information to anticipate future input.While measures of such load have played an important role in psycholinguistic theories, they have largely been formalized using symbolic grammars, which assign discrete, uniform costs to syntactic predictions.This study proposes a measure of processing storage cost based on an information-theoretic formalization, as the amount of information previous words carry about future context, under uncertainty.Unlike previous discrete, grammar-based metrics, this measure is continuous, probabilistic, theory-neutral, and can be estimated from pretrained neural language models.The validity of this approach is demonstrated through three analyses in English: our measure (i) recovers well-known processing asymmetries in center embeddings and relative clauses, (ii) correlates with a grammar-based storage cost in a syntactically-annotated corpus, and (iii) predicts reading-time variance in two large-scale naturalistic datasets over and above baseline models with traditional information-based predictors.Our code is available at https: //github.com/kohei-kaji/info-storage. Kohei Kajikawa, Shinnosuke Isono, Ethan Wilcox |
CoNLL | 3 |
| 2026 | Language Models Learn Constructional Semantics, Not To Mention Syntax: Investigating LM Understanding of Paired-Focus ConstructionsabstractGrasping the semantics of rare constructions (form-meaning pairings) has been shown to be a challenging problem that has currently only been solved by the largest LLMs.It remains an open question if open-source models have robust constructional understanding, and if so, what learning dynamics underlie the acquisition of this knowledge.Focusing on a set of rare PAIRED-FOCUS constructions in English (e.g."let alone", "much less"), we construct a novel dataset to test their meanings using both scalar adjectival semantics and general world knowledge.Testing a wide range of models differing in parameter count, architecture, and pretraining dataset size, we find that several modestly sized models are sensitive to both the forms and the meanings of PAIRED-FOCUS constructions, though models trained on human-scale data fail at all meaning evaluations.Turning to training dynamics for a set of open-checkpoint models, we find that PAIRED-FOCUS understanding emerges later in training than PAIRED-FOCUS syntactic knowledge, and that learning of PAIRED-FOCUS semantics is correlated with gains in some domains of world knowledge.Overall, our empirical results support the conclusion that modestly sized open-source models can grasp the rare PAIRED-FOCUS constructions, and demonstrate a connection between knowledge of PAIRED-FOCUS constructions and other meaning domains. Wesley Scivetti, Ethan Wilcox, Nathan Schneider 0001, Kanishka Misra, Leonie Weissweiler |
CoNLL | 2 |
| 2026 | What Can String Probability Tell Us About Grammaticality?abstractAbstract What have language models (LMs) learned about grammar? This question remains hotly debated, with major ramifications for linguistic theory. However, since probability and grammaticality are distinct notions in linguistics, it is not obvious what string probabilities can reveal about an LM’s underlying grammatical knowledge. We present a theoretical analysis of the relationship between grammar, meaning, and string probability, based on simple assumptions about the generative process of corpus data. Our framework makes three predictions, which we validate empirically using 280K sentence pairs in English and Chinese: (1) correlation between the probability of strings within minimal pairs, i.e., string pairs with minimal semantic differences; (2) correlation between models’ and humans’ deltas within minimal pairs; and (3) poor separation in probability space between unpaired grammatical and ungrammatical strings. Our analyses give theoretical grounding for using probability to learn about LMs’ structural knowledge, and suggest directions for future work in LM grammatical evaluation. Jennifer Hu 0001, Ethan Wilcox, Siyuan Song, Kyle Mahowald, Roger Levy |
Trans. Assoc. Comput. Linguistics | 2 |
| 2025 | Language Models Grow Less Humanlike beyond Phase TransitionabstractLMs' alignment with human reading behavior (i.e.psychometric predictive power; PPP) is known to improve during pretraining up to a tipping point, beyond which it either plateaus or degrades.Various factors, such as word frequency, recency bias in attention, and context size, have been theorized to affect PPP, yet there is no current account that explains why such a tipping point exists, and how it interacts with LMs' pretraining dynamics more generally.We hypothesize that the underlying factor is a pretraining phase transition, characterized by the rapid emergence of specialized attention heads.We conduct a series of correlational and causal experiments to show that such a phase transition is responsible for the tipping point in PPP.We then show that, rather than producing attention patterns that contribute to the degradation in PPP, phase transitions alter the subsequent learning dynamics of the model, such that further training keeps damaging PPP. Tatsuya Aoyama, Ethan Wilcox |
ACL (1) | 2 |
| 2025 | The time scale of redundancy between prosody and linguistic contextabstractTamar I Regev, Chiebuka Ohams, Shaylee Xie, Lukas Wolf, Evelina Fedorenko, Alex Warstadt, Ethan Wilcox, Tiago Pimentel. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Tamar I. Regev, Chiebuka Ohams, Shaylee Xie, Lukas Wolf, Evelina Fedorenko, Alex Warstadt, Ethan Wilcox, Tiago Pimentel |
ACL (1) | 7 |
| 2025 | The Harmonic Structure of Information ContoursabstractEleftheria Tsipidi, Samuel Kiegeland, Franz Nowak, Tianyang Xu, Ethan Wilcox, Alex Warstadt, Ryan Cotterell, Mario Giulianelli. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Eleftheria Tsipidi, Samuel Kiegeland, Franz Nowak, Tianyang Xu 0002, Ethan Wilcox, Alex Warstadt, Ryan Cotterell, Mario Giulianelli |
ACL (1) | 5 |
| 2025 | Using Information Theory to Characterize Prosodic Typology: The Case of Tone, Pitch-Accent and Stress-AccentabstractEthan Wilcox, Cui Ding, Giovanni Acampa, Tiago Pimentel, Alex Warstadt, Tamar I Regev. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Ethan Wilcox, Cui Ding, Giovanni Acampa, Tiago Pimentel, Alex Warstadt, Tamar I. Regev |
ACL (1) | 1 |
| 2025 | Anything Goes? A Crosslinguistic Study of (Im)possible Language Learning in LMsabstractDo language models (LMs) offer insights into human language learning?A common argument against this idea is that because their architecture and training paradigm are so vastly different from humans, LMs can learn arbitrary inputs as easily as natural languages.We test this claim by training LMs to model impossible and typologically unattested languages.Unlike previous work, which has focused exclusively on English, we conduct experiments on 12 languages from 4 language families with two newly constructed parallel corpora.Our results show that while GPT-2 small can largely distinguish attested languages from their impossible counterparts, it does not achieve perfect separation between all the attested languages and all the impossible ones.We further test whether GPT-2 small distinguishes typologically attested from unattested languages with different NP orders by manipulating word order based on Greenberg's Universal 20.We find that the model's perplexity scores do not distinguish attested vs. unattested word orders, while its performance on the generalization test does.These findings suggest that LMs exhibit some human-like inductive biases, though these biases are weaker than those found in human learners. Xiulin Yang, Tatsuya Aoyama, Yuekun Yao, Ethan Wilcox |
ACL (1) | 4 |
| 2025 | Modeling Bottom-up Information Quality during Language ProcessingabstractContemporary theories model language processing as integrating both top-down expectations and bottom-up inputs.One major prediction of such models is that the quality of the bottom-up inputs modulates ease of processing-noisy inputs should lead to difficult and effortful comprehension.We test this prediction in the domain of reading.First, we propose an information-theoretic operationalization for the "quality" of bottom-up information as the mutual information (MI) between visual information and word identity.We formalize this prediction in a mathematical model of reading as a Bayesian update.Second, we test our operationalization by comparing participants' reading times in conditions where words' information quality has been reduced, either by occluding their top or bottom half, with full words.We collect data in English and Chinese.We then use multimodal language models to estimate the mutual information between visual inputs and words.We use these data to estimate the specific effect of reduced information quality on reading times.Finally, we compare how information is distributed across visual forms.In English and Chinese, the upper half contains more information about word identity than the lower half.However, the asymmetry is more pronounced in English, a pattern which is reflected in the reading times. Cui Ding, Yanning Yin, Lena A. Jäger, Ethan Wilcox |
EMNLP | 4 |
| 2025 | Unpacking Let Alone: Human-Scale Models Generalize to a Rare Construction in Form but not MeaningabstractHumans have a remarkable ability to acquire and understand grammatical phenomena that are seen rarely, if ever, during childhood.Recent evidence suggests that language models with human-scale pretraining data may possess a similar ability by generalizing from frequent to rare constructions.However, it remains an open question how widespread this generalization ability is, and to what extent this knowledge extends to meanings of rare constructions, as opposed to just their forms.We fill this gap by testing human-scale transformer language models on their knowledge of both the form and meaning of the (rare and quirky) English LET-ALONE construction.To evaluate our LMs we construct a bespoke synthetic benchmark that targets syntactic and semantic properties of the construction.We find that human-scale LMs are sensitive to form, even when related constructions are filtered from the dataset.However, human-scale LMs do not make correct generalizations about LET-ALONE's meaning.These results point to an asymmetry in the current architectures' sample efficiency between language form and meaning, something which is not present in human language learners.1 Wesley Scivetti, Tatsuya Aoyama, Ethan Wilcox, Nathan Schneider 0001 |
EMNLP | 3 |
| 2025 | Elements of World Knowledge (EWoK): A Cognition-Inspired Framework for Evaluating Basic World Knowledge in Language Models
Anna A. Ivanova, Aalok Sathe, Benjamin Lipkin, Unnathi U. Kumar, Setayesh Radkani, Thomas Hikaru Clark, Carina Kauf, Jennifer Hu 0001, R. T. Pramod, Gabriel Grand, Vivian C. Paulun, Maria Ryskina, Ekin Akyürek, Ethan Wilcox, Nafisa Rashid, Leshem Choshen, Roger Levy, Evelina Fedorenko, Josh Tenenbaum, Jacob Andreas |
Trans. Assoc. Comput. Linguistics | 14 |
| 2024 | How does dependency type mediate gender agreement in Russian?
Cui Ding, Ethan Wilcox, Metehan Oguz, Zuzanna Fuchs |
CogSci | 2 |
| 2024 | Insights from the first BabyLM Challenge: Training sample-efficient language models on a developmentally plausible corpus
Alex Warstadt, Aaron Mueller, Leshem Choshen, Ethan Wilcox, Chengxu Zhuang, Adina Williams, Ryan Cotterell, Tal Linzen |
CogSci | 4 |
| 2024 | Reverse-Engineering the ReaderabstractNumerous previous studies have sought to determine to what extent language models, pretrained on natural language text, can serve as useful models of human cognition.In this paper, we are interested in the opposite question: whether we can directly optimize a language model to be a useful cognitive model by aligning it to human psychometric data.To achieve this, we introduce a novel alignment technique in which we fine-tune a language model to implicitly optimize the parameters of a linear regressor that directly predicts humans' reading times of in-context linguistic units, e.g., phonemes, morphemes, or words, using surprisal estimates derived from the language model.Using words as a test case, we evaluate our technique across multiple model sizes and datasets and find that it improves language models' psychometric predictive power.However, we find an inverse relationship between psychometric power and a model's performance on downstream NLP tasks as well as its perplexity on held-out test data.While this latter trend has been observed before (Oh et al., 2022;Shain et al., 2024), we are the first to induce it by manipulating a model's alignment to psychometric data. Samuel Kiegeland, Ethan Wilcox, Afra Amini, David R. Reich, Ryan Cotterell |
EMNLP | 2 |
| 2024 | On the Role of Context in Reading Time PredictionabstractWe present a new perspective on how readers integrate context during real-time language comprehension.Our proposals build on surprisal theory, which posits that the processing effort of a linguistic unit (e.g., a word) is an affine function of its in-context information content.We first observe that surprisal is only one out of many potential ways that a contextual predictor can be derived from a language model.Another one is the pointwise mutual information (PMI) between a unit and its context, which turns out to yield the same predictive power as surprisal when controlling for unigram frequency.Moreover, both PMI and surprisal are correlated with frequency.This means that neither PMI nor surprisal contains information about context alone.In response to this, we propose a technique where we project surprisal onto the orthogonal complement of frequency, yielding a new contextual predictor that is uncorrelated with frequency.Our experiments show that the proportion of variance in reading times explained by context is a lot smaller when context is represented by the orthogonalized predictor.From an interpretability standpoint, this indicates that previous studies may have overstated the role that context has in predicting reading times.https://github.com/rycolab/ context-reading-time Andreas Opedal, Eleanor Chodroff, Ryan Cotterell, Ethan Wilcox |
EMNLP | 4 |
| 2024 | Surprise! Uniform Information Density Isn't the Whole Story: Predicting Surprisal Contours in Long-form DiscourseabstractThe Uniform Information Density (UID) hypothesis posits that speakers tend to distribute information evenly across linguistic units to achieve efficient communication.Of course, information rate in texts and discourses is not perfectly uniform.While these fluctuations can be viewed as theoretically uninteresting noise on top of a uniform target, another explanation is that UID is not the only functional pressure regulating information content in a language.Speakers may also seek to maintain interest, adhere to writing conventions, and build compelling arguments.In this paper, we propose one such functional pressure; namely that speakers modulate information rate based on location within a hierarchically-structured model of discourse.We term this the Structured Context Hypothesis and test it by predicting the surprisal contours of naturally occurring discourses extracted from large language models using predictors derived from discourse structure.We find that hierarchical predictors are significant predictors of a discourse's information contour and that deeply nested hierarchical predictors are more predictive than shallow ones.This work takes an initial step beyond UID to propose testable hypotheses for why the information rate fluctuates in predictable ways.https://github.com/rycolab/ surprisal-discourse Eleftheria Tsipidi, Franz Nowak, Ryan Cotterell, Ethan Wilcox, Mario Giulianelli, Alex Warstadt |
EMNLP | 4 |
| 2023 | On the Efficacy of Sampling AdaptersabstractSampling is a common strategy for generating text from probabilistic models, yet standard ancestral sampling often results in text that is incoherent or ungrammatical.To alleviate this issue, various modifications to a model's sampling distribution, such as nucleus or top-k sampling, have been introduced and are now ubiquitously used in language generation systems.We propose a unified framework for understanding these techniques, which we term sampling adapters.Sampling adapters often lead to qualitatively better text, which raises the question: From a formal perspective, how are they changing the (sub)word-level distributions of language generation models?And why do these local changes lead to higher-quality text?We argue that the shift they enforce can be viewed as a trade-off between precision and recall: while the model loses its ability to produce certain strings, its precision rate on desirable text increases.While this trade-off is not reflected in standard metrics of distribution quality (such as perplexity), we find that several precision-emphasizing measures indeed indicate that sampling adapters can lead to probability distributions more aligned with the true distribution.Further, these measures correlate with higher sequence-level quality scores, specifically, MAUVE.https://github.com/rycolab/ sampling-adapters Clara Meister, Tiago Pimentel, Luca Malagutti, Ethan Wilcox, Ryan Cotterell |
ACL (1) | 4 |
| 2023 | Revisiting the Optimality of Word LengthsabstractZipf (1935) posited that wordforms are optimized to minimize utterances' communicative costs.Under the assumption that cost is given by an utterance's length, he supported this claim by showing that words' lengths are inversely correlated with their frequencies.Communicative cost, however, can be operationalized in different ways.Piantadosi et al. (2011) claim that cost should be measured as the distance between an utterance's information rate and channel capacity, which we dub the channel capacity hypothesis (CCH) here.Following this logic, they then proposed that a word's length should be proportional to the expected value of its surprisal (negative log-probability in context).In this work, we show that Piantadosi et al.'s derivation does not minimize CCH's cost, but rather a lower bound, which we term CCH ↓ .We propose a novel derivation, suggesting an improved way to minimize CCH's cost.Under this method, we find that a language's word lengths should instead be proportional to the surprisal's expectation plus its variance-tomean ratio.Experimentally, we compare these three communicative cost functions: Zipf's, CCH ↓ , and CCH.Across 13 languages and several experimental settings, we find that length is better predicted by frequency than either of the other hypotheses.In fact, when surprisal's expectation, or expectation plus variance-to-mean ratio, is estimated using better language models, it leads to worse word length predictions.We take these results as evidence that Zipf's longstanding hypothesis holds.https://github.com/tpimentelms/ optimality-of-word-lengths Tiago Pimentel, Clara Meister, Ethan Wilcox, Kyle Mahowald, Ryan Cotterell |
EMNLP | 3 |
| 2023 | Language Model Quality Correlates with Psychometric Predictive Power in Multiple LanguagesabstractSurprisal theory (Hale, 2001;Levy, 2008) posits that a word's reading time is proportional to its surprisal (i.e., to its negative log probability given the proceeding context).It has been empirically tested using surprisal estimates from language models (LMs).Under the premise that surprisal theory holds, we would expect that higher quality language models, whose predictions are more accurate, provide more powerful predictors of human reading behavior-a conjecture we dub the quality-power (QP) hypothesis.Unfortunately, empirical support for the QP hypothesis is mixed.Some studies in English have found correlations between LM quality and psychometric predictive power, but other studies using Japanese data, as well as using larger English LMs, find no such correlations.In this work, we conduct a systematic crosslinguistic assessment of the QP hypothesis.We train LMs from scratch on small-and medium-sized datasets from 13 languages (across five language families) and assess their ability to predict eye tracking data.We find correlations between LM quality and psychometric predictive power in eleven of these thirteen languages, suggesting that, within the range of model classes and sizes tested, better language models provide better predictors of human language processing behaviors.https://github.com/rycolab/ quality-power-hypothesis Ethan Wilcox, Clara Meister, Ryan Cotterell, Tiago Pimentel |
EMNLP | 1 |
| 2023 | Quantifying the redundancy between prosody and textabstractProsody-the suprasegmental component of speech, including pitch, loudness, and tempocarries critical aspects of meaning.However, the relationship between the information conveyed by prosody vs. by the words themselves remains poorly understood.We use large language models (LLMs) to estimate how much information is redundant between prosody and the words themselves.Using a large spoken corpus of English audiobooks, we extract prosodic features aligned to individual words and test how well they can be predicted from LLM embeddings, compared to non-contextual word embeddings.We find a high degree of redundancy between the information carried by the words and prosodic information across several prosodic features, including intensity, duration, pauses, and pitch contours.Furthermore, a word's prosodic information is redundant with both the word itself and the context preceding as well as following it.Still, we observe that prosodic features can not be fully predicted from text, suggesting that prosody carries information above and beyond the words.Along with this paper, we release a general-purpose data processing pipeline for quantifying the relationship between linguistic information and extra-linguistic features.https://github.com/lu-wo/ quantifying-redundancy Lukas Wolf, Tiago Pimentel, Evelina Fedorenko, Ryan Cotterell, Alex Warstadt, Ethan Wilcox, Tamar I. Regev |
EMNLP | 6 |
| 2023 | Controlled Text Generation with Natural Language InstructionsabstractLarge language models can be prompted to pro- duce fluent output for a wide range of tasks without being specifically trained to do so. Nevertheless, it is notoriously difficult to control their generation in such a way that it satisfies user-specified constraints. In this paper, we present InstructCTG, a simple controlled text generation framework that incorporates different constraints by verbalizing them as natural language instructions. We annotate natural texts through a combination of off-the-shelf NLP tools and simple heuristics with the linguistic and extra-linguistic constraints they satisfy. Then, we verbalize the constraints into natural language instructions to form weakly supervised training data, i.e., we prepend the natural language verbalizations of the constraints in front of their corresponding natural language sentences. Next, we fine-tune a pre-trained language model on the augmented corpus. Compared to existing methods, InstructCTG is more flexible in terms of the types of constraints it allows the practitioner to use. It also does not require any modification of the decoding procedure. Finally, InstructCTG allows the model to adapt to new constraints without re-training through the use of in-context learning. Wangchunshu Zhou, Yuchen Eleanor Jiang, Ethan Wilcox, Ryan Cotterell, Mrinmaya Sachan |
ICML | 3 |
| 2023 | On the Effect of Anticipation on Reading TimesabstractAbstract Over the past two decades, numerous studies have demonstrated how less-predictable (i.e., higher surprisal) words take more time to read. In general, these studies have implicitly assumed the reading process is purely responsive: Readers observe a new word and allocate time to process it as required. We argue that prior results are also compatible with a reading process that is at least partially anticipatory: Readers could make predictions about a future word and allocate time to process it based on their expectation. In this work, we operationalize this anticipation as a word’s contextual entropy. We assess the effect of anticipation on reading by comparing how well surprisal and contextual entropy predict reading times on four naturalistic reading datasets: two self-paced and two eye-tracking. Experimentally, across datasets and analyses, we find substantial evidence for effects of contextual entropy over surprisal on a word’s reading time (RT): In fact, entropy is sometimes better than surprisal in predicting a word’s RT. Spillover effects, however, are generally not captured by entropy, but only by surprisal. Further, we hypothesize four cognitive mechanisms through which contextual entropy could impact RTs—three of which we are able to design experiments to analyze. Overall, our results support a view of reading that is not just responsive, but also anticipatory.1 Tiago Pimentel, Clara Meister, Ethan Wilcox, Roger Levy, Ryan Cotterell |
Trans. Assoc. Comput. Linguistics | 3 |
| 2023 | Testing the Predictions of Surprisal Theory in 11 LanguagesabstractAbstract Surprisal theory posits that less-predictable words should take more time to process, with word predictability quantified as surprisal, i.e., negative log probability in context. While evidence supporting the predictions of surprisal theory has been replicated widely, much of it has focused on a very narrow slice of data: native English speakers reading English texts. Indeed, no comprehensive multilingual analysis exists. We address this gap in the current literature by investigating the relationship between surprisal and reading times in eleven different languages, distributed across five language families. Deriving estimates from language models trained on monolingual and multilingual corpora, we test three predictions associated with surprisal theory: (i) whether surprisal is predictive of reading times, (ii) whether expected surprisal, i.e., contextual entropy, is predictive of reading times, and (iii) whether the linking function between surprisal and reading times is linear. We find that all three predictions are borne out crosslinguistically. By focusing on a more diverse set of languages, we argue that these results offer the most robust link to date between information theory and incremental language processing across languages. Ethan Wilcox, Tiago Pimentel, Clara Meister, Ryan Cotterell, Roger Levy |
Trans. Assoc. Comput. Linguistics | 1 |
| 2022 | Evidence for Availability Effects on Speaker Choice in the Russian Comparative Alternation
Thomas Hikaru Clark, Ethan Wilcox, Edward Gibson, Roger Levy |
CogSci | 2 |
| 2021 | A Targeted Assessment of Incremental Processing in Neural Language Models and HumansabstractEthan Wilcox, Pranali Vani, Roger Levy. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Ethan Wilcox, Pranali Vani, Roger Levy |
ACL/IJCNLP (1) | 1 |
| 2021 | Using the Interpolated Maze Task to Assess Incremental Processing in English Relative Clauses
Pranali Vani, Ethan Wilcox, Roger Levy |
CogSci | 2 |
| 2020 | A Systematic Assessment of Syntactic Generalization in Neural Language ModelsabstractWhile state-of-the-art neural network models continue to achieve lower perplexity scores on language modeling benchmarks, it remains unknown whether optimizing for broad-coverage predictive performance leads to human-like syntactic knowledge.Furthermore, existing work has not provided a clear picture about the model properties required to produce proper syntactic generalizations.We present a systematic evaluation of the syntactic knowledge of neural language models, testing 20 combinations of model types and data sizes on a set of 34 English-language syntactic test suites.We find substantial differences in syntactic generalization performance by model architecture, with sequential models underperforming other architectures.Factorially manipulating model architecture and training dataset size (1M-40M words), we find that variability in syntactic generalization performance is substantially greater by architecture than by dataset size for the corpora tested in our experiments.Our results also reveal a dissociation between perplexity and syntactic generalization performance. Jennifer Hu 0001, Jon Gauthier, Ethan Wilcox, Roger Levy |
ACL | 4 |
| 2020 | On the Predictive Power of Neural Language Models for Human Real-Time Comprehension Behavior
Ethan Wilcox, Jon Gauthier, Jennifer Hu 0001, Roger Levy |
CogSci | 1 |
| 2020 | Structural Supervision Improves Few-Shot Learning and Syntactic Generalization in Neural Language ModelsabstractHumans can learn structural properties about a word from minimal experience, and deploy their learned syntactic representations uniformly in different grammatical contexts. We assess the ability of modern neural language models to reproduce this behavior in English and evaluate the effect of structural supervision on learning outcomes. First, we assess few-shot learning capabilities by developing controlled experiments that probe models' syntactic nominal number and verbal argument structure generalizations for tokens seen as few as two times during training. Second, we assess invariance properties of learned representation: the ability of a model to transfer syntactic generalizations from a base context (e.g., a simple declarative active-voice sentence) to a transformed context (e.g., an interrogative sentence). We test four models trained on the same dataset: an n-gram baseline, an LSTM, and two LSTM-variants trained with explicit structural supervision (Dyer et al.,2016; Charniak et al., 2016). We find that in most cases, the neural models are able to induce the proper syntactic generalizations after minimal exposure, often from just two examples during training, and that the two structurally supervised models generalize more accurately than the LSTM model. All neural models are able to leverage information learned in base contexts to drive expectations in transformed contexts, indicating that they have learned some invariance properties of syntax. Ethan Wilcox, Richard Futrell, Ryosuke Kohita, Roger Levy, Miguel Ballesteros |
EMNLP (1) | 1 |
| 2019 | Phonological Cues to Syntactic Structure in a Large-Scale Corpus
Ethan Wilcox |
CogSci | 1 |
| 2019 | Testing Gender Markedness of Nouns with Self a Paced Reading Study
Ethan Wilcox |
CogSci | 1 |
| 2019 | What Syntactic Structures block Dependencies in RNN Language Models?
Ethan Wilcox, Roger Levy, Richard Futrell |
CogSci | 1 |
| 2019 | The Role of Prior Beliefs in The Rational Speech Act Model of Pragmatics: Exhaustivity as a Case Study
Ethan Wilcox, Benjamin Spector |
CogSci | 1 |
| 2019 | Representation of Constituents in Neural Language Models: Coordination Phrase as a Case StudyabstractAixiu An, Peng Qian, Ethan Wilcox, Roger Levy. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Aixiu An, Ethan Wilcox, Roger Levy |
EMNLP/IJCNLP (1) | 3 |