Timothy J. O'Donnell

dblp:89/3188 · also Timothy John O'Donnell, Timothy O'Donnell 0001 · DBLP profile ↗
← Back
30ranked-venue papers
2as first author
15since 2021 · last 2026
0000-0002-5711-977XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 29 · 2 first-author · 15 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 2 first-author · 4 since 2021Systems, architecture and hardware · 1Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2026 LexiPhon: A Collection of Phonetically Transcribed Lexicons from Wikipedia
Amanda Doucette, Timothy J. O'Donnell, Morgan Sonderegger
LREC2
2025 Information Locality as an Inductive Bias for Neural Language Models
abstract
Taiga Someya, Anej Svete, Brian DuSell, Timothy J. O’Donnell, Mario Giulianelli, Ryan Cotterell. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Taiga Someya, Anej Svete, Brian DuSell, Timothy J. O'Donnell, Mario Giulianelli, Ryan Cotterell
ACL (1)4
2025 When unpredictable does not mean difficult to process
Jacob Hoover Vigly, Morgan Sonderegger, Timothy J. O'Donnell
CogSci4
2025 Modeling Open-World Cognition as On-Demand Synthesis of Probabilistic Models
Lionel Wong, Katie Collins, Lance Ying, Cedegao E. Zhang, Adrian Weller, Tobias Gerstenberg, Timothy J. O'Donnell, Alexander K. Lew, Jacob Andreas, Tyler Brooke-Wilson, Josh Tenenbaum
CogSci7
2025 Syntactic and Semantic Control of Large Language Models via Sequential Monte Carlo
abstract
A wide range of LM applications require generating text that conforms to syntactic or semantic constraints. Imposing such constraints can be naturally framed as _probabilistic conditioning_, but exact generation from the resulting distribution—which can differ substantially from the LM’s base distribution—is generally intractable. In this work, we develop an architecture for controlled LM generation based on sequential Monte Carlo (SMC). Our SMC framework allows us to flexibly incorporate domain- and problem-specific constraints at inference time, and efficiently reallocate computational resources in light of new information during the course of generation. By comparing to a number of alternatives and ablations on four challenging domains---Python code generation for data science, text-to-SQL, goal inference, and molecule synthesis—we demonstrate that, with little overhead, our approach allows small open-source language models to outperform models over 8$\times$ larger, as well as closed-source, fine-tuned ones. In support of the probabilistic perspective, we show that these performance improvements are driven by better approximation to the posterior distribution. [Our system](https://github.com/probcomp/genlm-control) builds on the framework of Lew et al. (2023) and integrates with its _language model probabilistic programming language_, giving users a simple, programmable way to apply SMC to a broad variety of controlled generation problems.
João Loula, Benjamin LeBrun, Benjamin Lipkin, Clemente Pasti, Gabriel Grand, Tianyu Liu 0004, Yahya Emara, Marjorie Freedman, Jason Eisner, Ryan Cotterell, Vikash Mansinghka 0001, Alexander K. Lew, Tim Vieira, Timothy J. O'Donnell
ICLR15
2025 Language Models over Canonical Byte-Pair Encodings
abstract
Modern language models represent probability distributions over character strings as distributions over (shorter) token strings derived via a deterministic tokenizer, such as byte-pair encoding. While this approach is highly effective at scaling up language models to large corpora, its current incarnations have a concerning property: the model assigns nonzero probability mass to an exponential number of noncanonical token encodings of each character string—these are token strings that decode to valid character strings but are impossible under the deterministic tokenizer (i.e., they will never be seen in any training corpus, no matter how large). This misallocation is both erroneous, as noncanonical strings never appear in training data, and wasteful, diverting probability mass away from plausible outputs. These are avoidable mistakes! In this work, we propose methods to enforce canonicality in token-level language models, ensuring that only canonical token strings are assigned positive probability. We present two approaches: (1) canonicality by conditioning, leveraging test-time inference strategies without additional training, and (2) canonicality by construction, a model parameterization that guarantees canonical outputs but requires training. We demonstrate that fixing canonicality mistakes improves the likelihood of held-out data for several models and corpora.
Tim Vieira, Tianyu Liu 0004, Clemente Pasti, Yahya Emara, Brian DuSell, Benjamin LeBrun, Mario Giulianelli, Juan Luis Gastaldi, Timothy J. O'Donnell, Ryan Cotterell
ICML9
2025 From Language Models over Tokens to Language Models over Characters
abstract
Modern language models are internally—and mathematically—distributions over token strings rather than character strings, posing numerous challenges for programmers building user applications on top of them. For example, if a prompt is specified as a character string, it must be tokenized before passing it to the token-level language model. Thus, the tokenizer and consequent processing are very sensitive to the specification of the prompt (e.g., whether the prompt ends with a space or not). This paper presents algorithms for converting token-level language models to character-level ones. We present both exact and approximate algorithms. In the empirical portion of the paper, we benchmark the practical runtime and approximation quality. Across four publicly available language models, we find that—even with a small computation budget—our method is able to accurately approximate the character-level distribution at reasonably fast speeds, and that a significant improvement in the language model’s compression rate (bits/byte) is achieved.
Tim Vieira, Benjamin LeBrun, Mario Giulianelli, Juan Luis Gastaldi, Brian DuSell, John Terilla, Timothy J. O'Donnell, Ryan Cotterell
ICML7
2022 Compositional Generalization in Dependency Parsing
abstract
Compositionality-the ability to combine familiar units like words into novel phrases and sentences-has been the focus of intense interest in artificial intelligence in recent years.To test compositional generalization in semantic parsing, Keysers et al. (2020) introduced Compositional Freebase Queries (CFQ).This dataset maximizes the similarity between the test and train distributions over primitive units, like words, while maximizing the compound divergence-the dissimilarity between test and train distributions over larger structures, like phrases.Dependency parsing, however, lacks a compositional generalization benchmark.In this work, we introduce a gold-standard set of dependency parses for CFQ, and use this to analyze the behavior of a state-of-the art dependency parser (Qi et al., 2020) on the CFQ dataset.We find that increasing compound divergence degrades dependency parsing performance, although not as dramatically as semantic parsing performance.Additionally, we find the performance of the dependency parser does not uniformly degrade relative to compound divergence, and the parser performs differently on different splits with the same compound divergence.We explore a number of hypotheses for what causes the non-uniform degradation in dependency parsing performance, and identify a number of syntactic structures that drive the dependency parser's lower performance on the most challenging splits.
Emily Goodwin, Siva Reddy, Timothy J. O'Donnell, Dzmitry Bahdanau
ACL (1)3
2022 Characterizing Idioms: Conventionality and Contingency
abstract
Idioms are unlike most phrases in two important ways.First, words in an idiom have non-canonical meanings.Second, the noncanonical meanings of words in an idiom are contingent on the presence of other words in the idiom.Linguistic theories differ on whether these properties depend on one another, as well as whether special theoretical machinery is needed to accommodate idioms.We define two measures that correspond to the properties above, and we implement them using BERT (Devlin et al., 2019) and XLNet (Yang et al., 2019).We show that English idioms fall at the expected intersection of the two dimensions, but that the dimensions themselves are not correlated.Our results suggest that special machinery to handle idioms may not be warranted.
Michaela Socolof, Jackie Chi Kit Cheung, Michael Wagner 0019, Timothy J. O'Donnell
ACL (1)4
2022 Measuring Morphological Fusion Using Partial Information Decomposition
abstract
Morphological systems across languages vary when it comes to the relation between form and meaning. In some languages, a single meaning feature corresponds to a single morpheme, whereas in other languages, multiple meaning features are bundled together into one morpheme. The two types of languages have been called agglutinative and fusional, respectively, but this distinction does not capture the graded nature of the phenomenon. We provide a mathematically precise way of characterizing morphological systems using partial information decomposition, a framework for decomposing mutual information into three components: unique, redundant, and synergistic information. We show that highly fusional languages are characterized by high levels of synergy.
Michaela Socolof, Jacob Hoover Vigly, Richard Futrell, Alessandro Sordoni, Timothy J. O'Donnell
COLING5
2022 Evaluating Distributional Distortion in Neural Language Modeling
Benjamin LeBrun, Alessandro Sordoni, Timothy J. O'Donnell
ICLR3
2021 The Learnability of Goal-directedness in Jazz Music
Daniel Harasim, Timothy J. O'Donnell, Martin Rohrmeier
CogSci2
2021 Processing differences among irregular inflection classes
Maya C. Watt, Mika Braginsky, Timothy J. O'Donnell
CogSci3
2021 Linguistic Dependencies and Statistical Dependence
abstract
Are pairs of words that tend to occur together also likely to stand in a linguistic dependency?This empirical question is motivated by a long history of literature in cognitive science, psycholinguistics, and NLP.In this work we contribute an extensive analysis of the relationship between linguistic dependencies and statistical dependence between words.Improving on previous work, we introduce the use of large pretrained language models to compute contextualized estimates of the pointwise mutual information between words (CPMI).For multiple models and languages, we extract dependency trees which maximize CPMI, and compare to gold standard linguistic dependencies.Overall, we find that CPMI dependencies achieve an unlabelled undirected attachment score of at most ≈ 0.5.While far above chance, and consistently above a non-contextualized PMI baseline, this score is generally comparable to a simple baseline formed by connecting adjacent words.We analyze which kinds of linguistic dependencies are best captured in CPMI dependencies, and also find marked differences between the estimates of the large pretrained language models, illustrating how their different training schemes affect the type of dependencies they capture.
Jacob Hoover Vigly, Wenyu Du, Alessandro Sordoni, Timothy J. O'Donnell
EMNLP (1)4
2021 Systematic Generalization with Edge Transformers
abstract
Recent research suggests that systematic generalization in natural language understanding remains a challenge for state-of-the-art neural models such as Transformers and Graph Neural Networks. To tackle this challenge, we propose Edge Transformer, a new model that combines inspiration from Transformers and rule-based symbolic AI. The first key idea in Edge Transformers is to associate vector states with every edge, that is, with every pair of input nodes---as opposed to just every node, as it is done in the Transformer model. The second major innovation is a triangular attention mechanism that updates edge representations in a way that is inspired by unification from logic programming. We evaluate Edge Transformer on compositional generalization benchmarks in relational reasoning, semantic parsing, and dependency parsing. In all three settings, the Edge Transformer outperforms Relation-aware, Universal and classical Transformer baselines.
Leon Bergen, Timothy J. O'Donnell, Dzmitry Bahdanau
NeurIPS2
2020 Exploiting Syntactic Structure for Better Language Modeling: A Syntactic Distance Approach
abstract
It is commonly believed that knowledge of syntactic structure should improve language modeling.However, effectively and computationally efficiently incorporating syntactic structure into neural language models has been a challenging topic.In this paper, we make use of a multi-task objective, i.e., the models simultaneously predict words as well as ground truth parse trees in a form called "syntactic distances", where information between these two separate objectives shares the same intermediate representation.Experimental results on the Penn Treebank and Chinese Treebank datasets show that when ground truth parse trees are provided as additional training signals, the model is able to achieve lower perplexity and induce trees with better quality.
Wenyu Du, Zhouhan Lin, Yikang Shen, Timothy J. O'Donnell, Yoshua Bengio, Yue Zhang 0004
ACL4
2020 Probing Linguistic Systematicity
abstract
Recently, there has been much interest in the question of whether deep natural language understanding models exhibit systematicitygeneralizing such that units like words make consistent contributions to the meaning of the sentences in which they appear.There is accumulating evidence that neural models often generalize non-systematically.We examined the notion of systematicity from a linguistic perspective, defining a set of probes and a set of metrics to measure systematic behaviour.We also identified ways in which network architectures can generalize non-systematically, and discuss why such forms of generalization may be unsatisfying.As a case study, we performed a series of experiments in the setting of natural language inference (NLI), demonstrating that some NLU systems achieve high overall performance despite being non-systematic.
Emily Goodwin, Koustuv Sinha, Timothy J. O'Donnell
ACL3
2020 Storage and Computation of Multimorphemic Words in Turkish
Rabia Ergin, Emily Morgan, Timothy J. O'Donnell
CogSci3
2019 Morphological Irregularity Correlates with Frequency
abstract
We present a study of morphological irregularity. Following recent work, we define an information-theoretic measure of irregularity based on the predictability of forms in a language. Using a neural transduction model, we estimate this quantity for the forms in 28 languages. We first present several validatory and exploratory analyses of irregularity. We then show that our analyses provide evidence for a correlation between irregularity and frequency: higher frequency items are more likely to be irregular and irregular items are more likely be highly frequent. To our knowledge, this result is the first of its breadth and confirms longstanding proposals from the linguistics literature. The correlation is more robust when aggregated at the level of whole paradigms—providing support for models of linguistic structure in which inflected forms are unified by abstract underlying stems or lexemes.
Ryan Cotterell, Timothy J. O'Donnell
ACL (1)3
2019 Evaluating systematicity in neural networks with natural language inference
Emily Goodwin, Koustuv Sinha, Timothy J. O'Donnell
CogSci3
2019 On Robustness: An Undervalued Dimension of Human Rationality
Ardavan Salehi Nobandegani, Kevin da Silva Castanheira, Timothy J. O'Donnell, Thomas R. Shultz
CogSci3
2017 Evaluating Hierarchies of Verb Argument Structure with Hierarchical Clustering
abstract
Verbs can only be used with a few specific arrangements of their arguments (syntactic frames). Most theorists note that verbs can be organized into a hierarchy of verb classes based on the frames they admit. Here we show that such a hierarchy is objectively well-supported by the patterns of verbs and frames in English, since a systematic hierarchical clustering algorithm converges on the same structure as the handcrafted taxonomy of VerbNet, a broad-coverage verb lexicon. We also show that the hierarchies capture meaningful psychological dimensions of generalization by predicting novel verb coercions by human participants. We discuss limitations of a simple hierarchical representation and suggest similar approaches for identifying the representations underpinning verb argument structure.
Jesse Mu, Joshua K. Hartshorne, Timothy J. O'Donnell
EMNLP3
2017 A Generative Model of Phonotactics
abstract
We present a probabilistic model of phonotactics, the set of well-formed phoneme sequences in a language. Unlike most computational models of phonotactics (Hayes and Wilson, 2008; Goldsmith and Riggle, 2012), we take a fully generative approach, modeling a process where forms are built up out of subparts by phonologically-informed structure building operations. We learn an inventory of subparts by applying stochastic memoization (Johnson et al., 2007; Goodman et al., 2008) to a generative process for phonemes structured as an and-or graph, based on concepts of feature hierarchy from generative phonology (Clements, 1985; Dresher, 2009). Subparts are combined in a way that allows tier-based feature interactions. We evaluate our models’ ability to capture phonotactic distributions in the lexicons of 14 languages drawn from the WOLEX corpus (Graff, 2012). Our full model robustly assigns higher probabilities to held-out forms than a sophisticated N-gram model for all languages. We also present novel analyses that probe model behavior in more detail.
Richard Futrell, Adam Albright, Peter Graff, Timothy J. O'Donnell
Trans. Assoc. Comput. Linguistics4
2016 Unsupervised learning of VerbNet argument structure
Jesse Mu, Timothy J. O'Donnell, Joshua K. Hartshorne
CogSci2
2015 A model of rapid phonotactic generalization
abstract
The phonotactics of a language describes the ways in which the sounds of the language combine to form possible morphemes and words.Humans can learn phonotactic patterns at the level of abstract classes, generalizing across sounds (e.g., "words can end in a voiced stop").Moreover, they rapidly acquire these generalizations, even before they acquire soundspecific patterns.We present a probabilistic model intended to capture this earlyabstraction phenomenon.The model represents both abstract and concrete generalizations in its hypothesis space from the outset of learning.This-combined with a parsimony bias in favor of compact descriptions of the input data-leads the model to favor rapid abstraction in a way similar to human learners.
Tal Linzen, Timothy J. O'Donnell
EMNLP2
2015 Unsupervised Lexicon Discovery from Acoustic Input
abstract
We present a model of unsupervised phonological lexicon discovery—the problem of simultaneously learning phoneme-like and word-like units from acoustic input. Our model builds on earlier models of unsupervised phone-like unit discovery from acoustic data (Lee and Glass, 2012), and unsupervised symbolic lexicon discovery using the Adaptor Grammar framework (Johnson et al., 2006), integrating these earlier approaches using a probabilistic model of phonological variation. We show that the model is competitive with state-of-the-art spoken term discovery systems, and present analyses exploring the model’s behavior and the kinds of linguistic structures it learns.
Chia-ying Lee, Timothy J. O'Donnell, James R. Glass
Trans. Assoc. Comput. Linguistics2
2011 Productivity and Reuse in Language
Timothy J. O'Donnell, Jesse Snedeker, Josh Tenenbaum, Noah D. Goodman
CogSci1
2011 Productivity and Reuse in Language: a Developmental Study
Timothy J. O'Donnell, Jesse Snedeker, Josh Tenenbaum, Noah D. Goodman
CogSci1
2011 Storage and computation in syntax: Evidence from relative clause priming
Melissa Troyer, Timothy J. O'Donnell, Evelina Fedorenko, Edward Gibson
CogSci2
1996 Analysis of the Early Workload on the Cornell Theory Center IBM SP2
abstract
Parallel computers have matured to the point where they are capable of running a significant production workload. Characterizing this workload, however, is far more complicated than for the single-processor case. Besides the varying number of processors that may be invoked, the nodes themselves may provide differing computational resources (memory size, for example). In addition, the batch schedulers may introduce further categories of service which must be considered in the analysis.The Cornell Theory Center (CTC) put a 512-node IBM SP2 system into production in early 1995. Extended traces of batch jobs began to be collected in mid-1995 when the usage base became sufficiently large. This paper offers an analysis of this early batch workload.
Steven Hotovy, David J. Schneider 0001, Timothy J. O'Donnell
SIGMETRICS3