EDBT 2026 Demo / reviewers in the wild / expert
Kyle Mahowald
dblp:38/11196
· DBLP profile ↗
22ranked-venue papers
1as first author
21since 2021 · last 2026
0000-0002-9786-8716ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 22 · 1 first-author · 21 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Mechanisms of Prompt-Induced Hallucination in Vision-Language ModelsabstractWilliam Rudman, Michal Golovanevsky, Dana Arad, Yonatan Belinkov, Carsten Eickhoff, Ritambhara Singh, Kyle Mahowald. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. William Rudman, Michal Golovanevsky, Dana Arad, Yonatan Belinkov, Carsten Eickhoff, Ritambhara Singh, Kyle Mahowald |
ACL (1) | 7 |
| 2026 | What Can String Probability Tell Us About Grammaticality?abstractAbstract What have language models (LMs) learned about grammar? This question remains hotly debated, with major ramifications for linguistic theory. However, since probability and grammaticality are distinct notions in linguistics, it is not obvious what string probabilities can reveal about an LM’s underlying grammatical knowledge. We present a theoretical analysis of the relationship between grammar, meaning, and string probability, based on simple assumptions about the generative process of corpus data. Our framework makes three predictions, which we validate empirically using 280K sentence pairs in English and Chinese: (1) correlation between the probability of strings within minimal pairs, i.e., string pairs with minimal semantic differences; (2) correlation between models’ and humans’ deltas within minimal pairs; and (3) poor separation in probability space between unpaired grammatical and ungrammatical strings. Our analyses give theoretical grounding for using probability to learn about LMs’ structural knowledge, and suggest directions for future work in LM grammatical evaluation. Jennifer Hu 0001, Ethan Wilcox, Siyuan Song, Kyle Mahowald, Roger Levy |
Trans. Assoc. Comput. Linguistics | 4 |
| 2025 | Reinforcement learning produces efficient case-marking systems
Sasha Boguraev, Katrin Erk, Kyle Mahowald, James W. Shearer, Stephen Wechsler |
CogSci | 3 |
| 2025 | Causal Interventions Reveal Shared Structure Across English Filler-Gap ConstructionsabstractLanguage Models (LMs) have emerged as powerful sources of evidence for linguists seeking to develop theories of syntax.In this paper, we argue that causal interpretability methods, applied to LMs, can greatly enhance the value of such evidence by helping us characterize the abstract mechanisms that LMs learn to use.Our empirical focus is a set of English filler-gap dependency constructions (e.g., questions, relative clauses).Linguistic theories largely agree that these constructions share many properties.Using experiments based in Distributed Interchange Interventions, we show that LMs converge on similar abstract analyses of these constructions.These analyses also reveal previously overlooked factorsrelating to frequency, filler type, and surrounding context -that could motivate changes to standard linguistic theory.Overall, these results suggest that mechanistic, internal analyses of LMs can push linguistic theory forward.https://github.com/SashaBoguraev/ causal-filler-gap Sasha Boguraev, Christopher Potts, Kyle Mahowald |
EMNLP | 3 |
| 2025 | Convergence and Divergence of Language Models under Different Random SeedsabstractIn this paper, we investigate the convergence of language models (LMs) trained under different random seeds, measuring convergence as the expected per-token Kullback-Leibler (KL) divergence across seeds.By comparing LM convergence as a function of model size and training checkpoint, we identify a four-phase convergence pattern: (i) an initial uniform phase, (ii) a sharp-convergence phase, (iii) a sharp-divergence phase, and (iv) a slowreconvergence phase.Further, we observe that larger models reconverge faster in later training stages, while smaller models never actually reconverge; these results suggest that a certain model size may be necessary to learn stable distributions.Restricting our analysis to specific token frequencies or part-of-speech (PoS) tags further reveals that convergence is uneven across linguistic categories: frequent tokens and function words converge faster and more reliably than their counterparts (infrequent tokens and content words).Overall, our findings highlight factors that influence the stability of the learned distributions in model training. Finlay Fehlauer, Kyle Mahowald, Tiago Pimentel |
EMNLP | 2 |
| 2025 | Constructions are Revealed in Word DistributionsabstractConstruction grammar posits that constructions, or form-meaning pairings, are acquired through experience with language (the distributional learning hypothesis).But how much information about constructions does this distribution actually contain?Corpus-based analyses provide some answers, but text alone cannot answer counterfactual questions about what caused a particular word to occur.This requires computable models of the distribution over strings-namely, pretrained language models (PLMs).Here, we treat a RoBERTa model as a proxy for this distribution and hypothesize that constructions will be revealed within it as patterns of statistical affinity.We support this hypothesis experimentally: many constructions are robustly distinguished, including (i) hard cases where semantically distinct constructions are superficially similar, as well as (ii) schematic constructions, whose "slots" can be filled by abstract word classes.Despite this success, we also provide qualitative evidence that statistical affinity alone may be insufficient to identify all constructions from text.Thus, statistical affinity is likely an important, but partial, signal available to learners. 1 Josh Rozner, Leonie Weissweiler, Kyle Mahowald, Cory Shain |
EMNLP | 3 |
| 2025 | To CoT or not to CoT? Chain-of-thought helps mainly on math and symbolic reasoningabstractChain-of-thought (CoT) via prompting is the de facto method for eliciting reasoning capabilities from large language models (LLMs). But for what kinds of tasks is this extra "thinking" really helpful? To analyze this, we conducted a quantitative meta-analysis covering over 100 papers using CoT and ran our own evaluations of 20 datasets across 14 models. Our results show that CoT gives strong performance benefits primarily on tasks involving math or logic, with much smaller gains on other types of tasks. On MMLU, directly generating the answer without CoT leads to almost identical accuracy as CoT unless the question or model's response contains an equals sign, indicating symbolic operations and reasoning. Following this finding, we analyze the behavior of CoT on these problems by separating planning and execution and comparing against tool-augmented LLMs. Much of CoT's gain comes from improving symbolic execution, but it underperforms relative to using a symbolic solver. Our results indicate that CoT can be applied selectively, maintaining performance while saving inference costs. Furthermore, they suggest a need to move beyond prompt-based CoT to new paradigms that better leverage intermediate computation across the whole range of LLM applications. Zayne Sprague, Fangcong Yin, Juan Diego Rodriguez, Dongwei Jiang, Manya Wadhwa, Prasann Singhal, Xi Ye 0003, Kyle Mahowald, Greg Durrett |
ICLR | 9 |
| 2024 | Mission: Impossible Language ModelsabstractJulie Kallini, Isabel Papadimitriou, Richard Futrell, Kyle Mahowald, Christopher Potts. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Julie Kallini, Isabel Papadimitriou, Richard Futrell, Kyle Mahowald, Christopher Potts |
ACL (1) | 4 |
| 2024 | Participle-Prepended Nominals Have Lower Entropy Than Nominals Appended After the Participle
Kristie Denlinger, Stephen Wechsler, Kyle Mahowald |
CogSci | 3 |
| 2024 | Experimental Contexts Can Facilitate Robust Semantic Property Inference in Language Models, but InconsistentlyabstractRecent zero-shot evaluations have highlighted important limitations in the abilities of language models (LMs) to perform meaning extraction.However, it is now well known that LMs can demonstrate radical improvements in the presence of experimental contexts such as in-context examples and instructions.How well does this translate to previously studied meaning-sensitive tasks?We present a casestudy on the extent to which experimental contexts can improve LMs' robustness in performing property inheritance-predicting semantic properties of novel concepts, a task that they have been previously shown to fail on.Upon carefully controlling the nature of the in-context examples and the instructions, our work reveals that they can indeed lead to nontrivial property inheritance behavior in LMs.However, this ability is inconsistent: with a minimal reformulation of the task, some LMs were found to pick up on shallow, non-semantic heuristics from their inputs, suggesting that the computational principles of semantic property inference are yet to be mastered by LMs. Kanishka Misra, Allyson Ettinger, Kyle Mahowald |
EMNLP | 3 |
| 2024 | Language Models Learn Rare Phenomena from Less Rare Phenomena: The Case of the Missing AANNsabstractLanguage models learn rare syntactic phenomena, but the extent to which this is attributable to generalization vs. memorization is a major open question.To that end, we iteratively trained transformer language models on systematically manipulated corpora which were human-scale in size, and then evaluated their learning of a rare grammatical phenomenon: the English Article+Adjective+Numeral+Noun (AANN) construction ("a beautiful five days").We compared how well this construction was learned on the default corpus relative to a counterfactual corpus in which AANN sentences were removed.We found that AANNs were still learned better than systematically perturbed variants of the construction.Using additional counterfactual corpora, we suggest that this learning occurs through generalization from related constructions (e.g., "a few days").An additional experiment showed that this learning is enhanced when there is more variability in the input.Taken together, our results provide an existence proof that LMs can learn rare grammatical phenomena by generalization from less rare phenomena. Kanishka Misra, Kyle Mahowald |
EMNLP | 2 |
| 2023 | A Method for Studying Semantic Construal in Grammatical Constructions with Interpretable Contextual Embedding SpacesabstractWe study semantic construal in grammatical constructions using large language models.First, we project contextual word embeddings into three interpretable semantic spaces, each defined by a different set of psycholinguistic feature norms.We validate these interpretable spaces and then use them to automatically derive semantic characterizations of lexical items in two grammatical constructions: nouns in subject or object position within the same sentence, and the AANN construction (e.g., 'a beautiful three days').We show that a word in subject position is interpreted as more agentive than the very same word in object position, and that the nouns in the AANN construction are interpreted as more measurement-like than when in the canonical alternation.Our method can probe the distributional meaning of syntactic constructions at a templatic level, abstracted away from specific lexemes. Gabriella Chronis, Kyle Mahowald, Katrin Erk |
ACL (1) | 2 |
| 2023 | A Discerning Several Thousand Judgments: GPT-3 Rates the Article + Adjective + Numeral + Noun ConstructionabstractKnowledge of syntax includes knowledge of rare, idiosyncratic constructions.LLMs must overcome frequency biases in order to master such constructions.In this study, I prompt GPT-3 to give acceptability judgments on the English-language Article + Adjective + Numeral + Noun construction (e.g., "a lovely five days").I validate the prompt using the CoLA corpus of acceptability judgments and then zero in on the AANN construction.I compare GPT-3's judgments to crowdsourced human judgments on a subset of sentences.GPT-3's judgments are broadly similar to human judgments and generally align with proposed constraints in the literature but, in some cases, GPT-3's judgments and human judgments diverge from the literature and from each other. Kyle Mahowald |
EACL | 1 |
| 2023 | Revisiting the Optimality of Word LengthsabstractZipf (1935) posited that wordforms are optimized to minimize utterances' communicative costs.Under the assumption that cost is given by an utterance's length, he supported this claim by showing that words' lengths are inversely correlated with their frequencies.Communicative cost, however, can be operationalized in different ways.Piantadosi et al. (2011) claim that cost should be measured as the distance between an utterance's information rate and channel capacity, which we dub the channel capacity hypothesis (CCH) here.Following this logic, they then proposed that a word's length should be proportional to the expected value of its surprisal (negative log-probability in context).In this work, we show that Piantadosi et al.'s derivation does not minimize CCH's cost, but rather a lower bound, which we term CCH ↓ .We propose a novel derivation, suggesting an improved way to minimize CCH's cost.Under this method, we find that a language's word lengths should instead be proportional to the surprisal's expectation plus its variance-tomean ratio.Experimentally, we compare these three communicative cost functions: Zipf's, CCH ↓ , and CCH.Across 13 languages and several experimental settings, we find that length is better predicted by frequency than either of the other hypotheses.In fact, when surprisal's expectation, or expectation plus variance-to-mean ratio, is estimated using better language models, it leads to worse word length predictions.We take these results as evidence that Zipf's longstanding hypothesis holds.https://github.com/tpimentelms/ optimality-of-word-lengths Tiago Pimentel, Clara Meister, Ethan Wilcox, Kyle Mahowald, Ryan Cotterell |
EMNLP | 4 |
| 2023 | Elaborative Simplification as Implicit Questions Under DiscussionabstractAutomated text simplification, a technique useful for making text more accessible to people such as children and emergent bilinguals, is often thought of as a monolingual translation task from complex to simplified text.This view fails to account for elaborative simplification, where new information is added into the simplified text.This paper proposes to view elaborative simplification through the lens of the Question Under Discussion (QUD) framework, providing a robust way to investigate what writers elaborate upon, how they elaborate, and how elaborations fit into the discourse context by viewing elaborations as explicit answers to implicit questions.We introduce ELABQUD, consisting of 1.3K elaborations accompanied with implicit QUDs, to study these phenomena.We show that explicitly modeling QUD (via question generation) not only provides essential understanding of elaborative simplification and how the elaborations connect with the rest of the discourse, but also substantially improves the quality of elaboration generation. Yating Wu 0002, William Sheffield, Kyle Mahowald, Junyi Jessy Li |
EMNLP | 3 |
| 2022 | Why is Winoground Hard? Investigating Failures in Visuolinguistic CompositionalityabstractRecent visuolinguistic pre-trained models show promising progress on various end tasks such as image retrieval and video captioning.Yet, they fail miserably on the recently proposed Winoground dataset (Thrush et al., 2022), which challenges models to match paired images and English captions, with items constructed to overlap lexically but differ in meaning (e.g., "there is a mug in some grass" vs. "there is some grass in a mug").By annotating the dataset using new fine-grained tags, we show that solving the Winoground task requires not just compositional language understanding, but a host of other abilities like commonsense reasoning or locating small, out-of-focus objects in low-resolution images.In this paper, we identify the dataset's main challenges through a suite of experiments on related tasks (probing task, image retrieval task), data augmentation, and manual inspection of the dataset.Our analysis suggests that a main challenge in visuolinguistic models may lie in fusing visual and textual representations, rather than in compositional language understanding.We release our annotation and code at https://github. com/ajd12342/why-winoground-hard. Anuj Diwan, Layne Berry, Eunsol Choi, David F. Harwath, Kyle Mahowald |
EMNLP | 5 |
| 2022 | What do tokens know about their characters and how do they know it?abstractPre-trained language models (PLMs) that use subword tokenization schemes can succeed at a variety of language tasks that require characterlevel information, despite lacking explicit access to the character composition of tokens.Here, studying a range of models (e.g., GPT-J, BERT, RoBERTa, GloVe), we probe what word pieces encode about character-level information by training classifiers to predict the presence or absence of a particular alphabetical character in a token, based on its embedding (e.g., probing whether the model embedding for "cat" encodes that it contains the character "a").We find that these models robustly encode character-level information and, in general, larger models perform better at the task.We show that these results generalize to characters from non-Latin alphabets (Arabic, Devanagari, and Cyrillic).Then, through a series of experiments and analyses, we investigate the mechanisms through which PLMs acquire English-language character information during training and argue that this knowledge is acquired through multiple phenomena, including a systematic relationship between particular characters and particular parts of speech, as well as natural variability in the tokenization of related strings. Ayush Kaushal, Kyle Mahowald |
NAACL-HLT | 2 |
| 2021 | Deep Subjecthood: Higher-Order Grammatical Features in Multilingual BERTabstractWe investigate how Multilingual BERT (mBERT) encodes grammar by examining how the high-order grammatical feature of morphosyntactic alignment (how different languages define what counts as a "subject") is manifested across the embedding spaces of different languages.To understand if and how morphosyntactic alignment affects contextual embedding spaces, we train classifiers to recover the subjecthood of mBERT embeddings in transitive sentences (which do not contain overt information about morphosyntactic alignment) and then evaluate them zero-shot on intransitive sentences (where subjecthood classification depends on alignment), within and across languages.We find that the resulting classifier distributions reflect the morphosyntactic alignment of their training languages.Our results demonstrate that mBERT representations are influenced by high-level grammatical features that are not manifested in any one input sentence, and that this is robust across languages.Further examining the characteristics that our classifiers rely on, we find that features such as passive voice, animacy and case strongly correlate with classification decisions, suggesting that mBERT does not encode subjecthood purely syntactically, but that subjecthood embedding is continuous and dependent on semantic and discourse factors, as is proposed in much of the functional linguistics literature.Together, these results provide insight into how grammatical features manifest in contextual embedding spaces, at a level of abstraction not covered by previous work.1 Isabel Papadimitriou, Ethan A. Chi, Richard Futrell, Kyle Mahowald |
EACL | 4 |
| 2021 | A Massively Multilingual Analysis of Cross-linguality in Shared Embedding SpaceabstractIn cross-lingual language models, representations for many different languages live in the same space.Here, we investigate the linguistic and non-linguistic factors affecting sentencelevel alignment in cross-lingual pretrained language models for 101 languages and 5,050 language pairs.Using BERT-based LaBSE and BiLSTM-based LASER as our models, and the Bible as our corpus, we compute a taskbased measure of cross-lingual alignment in the form of bitext retrieval performance, as well as four intrinsic measures of vector space alignment and isomorphism.We then examine a range of linguistic, quasi-linguistic, and training-related features as potential predictors of these alignment metrics.The results of our analyses show that word order agreement and agreement in morphological complexity are two of the strongest linguistic predictors of cross-linguality.We also note in-family training data as a stronger predictor than languagespecific training data across the board.We verify some of our linguistic findings by looking at the effect of morphological segmentation on English-Inuktitut alignment, in addition to examining the effect of word order agreement on isomorphism for 66 zero-shot language pairs from a different corpus.We make the data and code for our experiments publicly available.1 William Yang Wang, Kyle Mahowald |
EMNLP (1) | 3 |
| 2021 | How (Non-)Optimal is the Lexicon?abstractTiago Pimentel, Irene Nikkarinen, Kyle Mahowald, Ryan Cotterell, Damián Blasi. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Tiago Pimentel, Irene Nikkarinen, Kyle Mahowald, Ryan Cotterell, Damián E. Blasi |
NAACL-HLT | 3 |
| 2021 | Decrypting Cryptic Crosswords: Semantically Complex Wordplay Puzzles as a Target for NLPabstractCryptic crosswords, the dominant crossword variety in the UK, are a promising target for advancing NLP systems that seek to process semantically complex, highly compositional language. Cryptic clues read like fluent natural language but are adversarially composed of two parts: a definition and a wordplay cipher requiring character-level manipulations. Expert humans use creative intelligence to solve cryptics, flexibly combining linguistic, world, and domain knowledge. In this paper, we make two main contributions. First, we present a dataset of cryptic clues as a challenging new benchmark for NLP systems that seek to process compositional language in more creative, human-like ways. After showing that three non-neural approaches and T5, a state-of-the-art neural language model, do not achieve good performance, we make our second main contribution: a novel curriculum approach, in which the model is first fine-tuned on related tasks such as unscrambling words. We also introduce a challenging data split, examine the meta-linguistic capabilities of subword-tokenized models, and investigate model systematicity by perturbing the wordplay part of clues, showing that T5 exhibits behavior partially consistent with human solving strategies. Although our curricular approach considerably improves on the T5 baseline, our best-performing model still fails to generalize to the extent that humans can. Thus, cryptic crosswords remain an unsolved challenge for NLP systems and a potential source of future innovation. Josh Rozner, Christopher Potts, Kyle Mahowald |
NeurIPS | 3 |
| 2020 | With Little Power Comes Great ResponsibilityabstractDespite its importance to experimental design, statistical power (the probability that, given a real effect, an experiment will reject the null hypothesis) has largely been ignored by the NLP community.Underpowered experiments make it more difficult to discern the difference between statistical noise and meaningful model improvements, and increase the chances of exaggerated findings.By metaanalyzing a set of existing NLP papers and datasets, we characterize typical power for a variety of settings and conclude that underpowered experiments are common in the NLP literature.In particular, for several tasks in the popular GLUE benchmark, small test sets mean that most attempted comparisons to state of the art models will not be adequately powered.Similarly, based on reasonable assumptions, we find that the most typical experimental design for human rating studies will be underpowered to detect small model differences, of the sort that are frequently studied.For machine translation, we find that typical test sets of 2000 sentences have approximately 75% power to detect differences of 1 BLEU point.To improve the situation going forward, we give an overview of best practices for power analysis in NLP and release a series of notebooks to assist with future power analyses.1 Dallas Card, Peter Henderson 0002, Urvashi Khandelwal, Robin Jia, Kyle Mahowald, Daniel Jurafsky |
EMNLP (1) | 5 |