VLDB 2026 Research / reviewers in the wild / expert
John Terilla
dblp:179/8683
· DBLP profile ↗
3ranked-venue papers
0as first author
3since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Language models and text generation · 76% Learning theory · 19% Efficient and distributed learning · 6% |
Topics — the 5 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
tokenization |
1.7 | 2 | 2025 | From Language Models over Tokens to Language Models over Characters · ICML 2025 The Foundations of Tokenization: Statistical and Computational Concerns · ICLR 2025 |
Natural language and speech › Language models and text generation › language modeling
character-level language modeling |
0.9 | 1 | 2025 | From Language Models over Tokens to Language Models over Characters · ICML 2025 |
Machine learning › Learning theory › statistical estimation
statistical consistency |
0.9 | 1 | 2025 | The Foundations of Tokenization: Statistical and Computational Concerns · ICLR 2025 |
Natural language and speech › Language models and text generation › language modeling
statistical language modeling |
0.9 | 1 | 2025 | From Language Models over Tokens to Language Models over Characters · ICML 2025 |
Machine learning › Efficient and distributed learning
compression |
0.3 | 1 | 2025 | From Language Models over Tokens to Language Models over Characters · ICML 2025 |
Methods — techniques the papers use, named apart from their topics
stochastic map · 0.9exact and approximate marginalization · 0.9category theory · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | The Foundations of Tokenization: Statistical and Computational ConcernsabstractTokenization — the practice of converting strings of characters from an alphabet into sequences of tokens over a vocabulary — is a critical step in the NLP pipeline. The use of token representations is widely credited with increased model performance but is also the source of many undesirable behaviors, such as spurious ambiguity or inconsistency. Despite its recognized importance as a standard representation method in NLP, the theoretical underpinnings of tokenization are not yet fully understood. In particular, the impact of tokenization on language model estimation has been investigated primarily through empirical means. The present paper contributes to addressing this theoretical gap by proposing a unified formal framework for representing and analyzing tokenizer models. Based on the category of stochastic maps, this framework enables us to establish general conditions for a principled use of tokenizers and, most importantly, the necessary and sufficient conditions for a tokenizer model to preserve the consistency of statistical estimators. In addition, we discuss statistical and computational concerns crucial for designing and implementing tokenizer models, such as inconsistency, ambiguity, finiteness, and sequentiality. The framework and results advanced in this paper contribute to building robust theoretical foundations for representations in neural language modeling that can inform future theoretical and empirical research. Juan Luis Gastaldi, John Terilla, Luca Malagutti, Brian DuSell, Tim Vieira, Ryan Cotterell |
ICLR | 2 |
| 2025 | From Language Models over Tokens to Language Models over CharactersabstractModern language models are internally—and mathematically—distributions over token strings rather than character strings, posing numerous challenges for programmers building user applications on top of them. For example, if a prompt is specified as a character string, it must be tokenized before passing it to the token-level language model. Thus, the tokenizer and consequent processing are very sensitive to the specification of the prompt (e.g., whether the prompt ends with a space or not). This paper presents algorithms for converting token-level language models to character-level ones. We present both exact and approximate algorithms. In the empirical portion of the paper, we benchmark the practical runtime and approximation quality. Across four publicly available language models, we find that—even with a small computation budget—our method is able to accurately approximate the character-level distribution at reasonably fast speeds, and that a significant improvement in the language model’s compression rate (bits/byte) is achieved. Tim Vieira, Benjamin LeBrun, Mario Giulianelli, Juan Luis Gastaldi, Brian DuSell, John Terilla, Timothy J. O'Donnell, Ryan Cotterell |
ICML | 6 |
| 2021 | Tensor Networks for Probabilistic Sequence ModelingabstractTensor networks are a powerful modeling framework developed for computational many-body physics, which have only recently been applied within machine learning. In this work we utilize a uniform matrix product state (u-MPS) model for probabilistic modeling of sequence data. We first show that u-MPS enable sequence-level parallelism, with length-n sequences able to be evaluated in depth O(log n). We then introduce a novel generative algorithm giving trained u-MPS the ability to efficiently sample from a wide variety of conditional distributions, each one defined by a regular expression. Special cases of this algorithm correspond to autoregressive and fill-in-the-blank sampling, but more complex regular expressions permit the generation of richly structured data in a manner that has no direct analogue in neural generative models. Experiments on sequence modeling with synthetic and real text data show u-MPS outperforming a variety of baselines and effectively generalizing their predictions in the presence of limited data. Jacob Miller 0003, Guillaume Rabusseau, John Terilla |
AISTATS | 3 |