VLDB 2026 Research / reviewers in the wild / expert
Benjamin LeBrun
dblp:317/0720
· DBLP profile ↗
4ranked-venue papers
1as first author
4since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Language models and text generation · 78% Probabilistic and Bayesian machine learning · 19% Efficient and distributed learning · 3% |
Topics — the 9 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation › language modeling
statistical language modeling |
1.7 | 2 | 2025 | From Language Models over Tokens to Language Models over Characters · ICML 2025 Language Models over Canonical Byte-Pair Encodings · ICML 2025 |
Natural language and speech › Language models and text generation
tokenization |
1.7 | 2 | 2025 | From Language Models over Tokens to Language Models over Characters · ICML 2025 Language Models over Canonical Byte-Pair Encodings · ICML 2025 |
Natural language and speech › Language models and text generation › tokenization › subword tokenization
byte-pair encoding |
0.9 | 1 | 2025 | Language Models over Canonical Byte-Pair Encodings · ICML 2025 |
Natural language and speech › Language models and text generation › language modeling
character-level language modeling |
0.9 | 1 | 2025 | From Language Models over Tokens to Language Models over Characters · ICML 2025 |
Machine learning › Probabilistic and Bayesian machine learning
probabilistic inference |
0.9 | 1 | 2025 | Syntactic and Semantic Control of Large Language Models via Sequential Monte Carlo · ICLR 2025 |
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods
sequential monte carlo |
0.9 | 1 | 2025 | Syntactic and Semantic Control of Large Language Models via Sequential Monte Carlo · ICLR 2025 |
Natural language and speech › Language models and text generation
text generation |
0.9 | 1 | 2025 | Syntactic and Semantic Control of Large Language Models via Sequential Monte Carlo · ICLR 2025 |
Natural language and speech › Language models and text generation
neural language model |
0.6 | 1 | 2022 | Evaluating Distributional Distortion in Neural Language Modeling · ICLR 2022 |
Machine learning › Efficient and distributed learning
compression |
0.3 | 1 | 2025 | From Language Models over Tokens to Language Models over Characters · ICML 2025 |
Methods — techniques the papers use, named apart from their topics
sequential monte carlo · 0.9probabilistic conditioning · 0.9model reparameterization · 0.9exact and approximate marginalization · 0.9constrained inference · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Syntactic and Semantic Control of Large Language Models via Sequential Monte CarloabstractA wide range of LM applications require generating text that conforms to syntactic or semantic constraints. Imposing such constraints can be naturally framed as _probabilistic conditioning_, but exact generation from the resulting distribution—which can differ substantially from the LM’s base distribution—is generally intractable. In this work,
we develop an architecture for controlled LM generation based on sequential Monte Carlo (SMC). Our SMC framework allows us to flexibly incorporate domain- and problem-specific constraints at inference time, and efficiently reallocate computational resources in light of new information during the course of generation. By comparing to a number of alternatives and ablations on four challenging domains---Python code generation for data science, text-to-SQL, goal inference, and molecule synthesis—we demonstrate that, with little overhead, our approach allows small open-source language models to outperform models over 8$\times$ larger, as well as closed-source, fine-tuned ones.
In support of the probabilistic perspective, we show that these performance improvements are driven by better approximation to the posterior distribution.
[Our system](https://github.com/probcomp/genlm-control) builds on the framework of Lew et al. (2023) and integrates with its _language model probabilistic programming language_, giving users a simple, programmable way to apply SMC to a broad variety of controlled generation problems. João Loula, Benjamin LeBrun, Benjamin Lipkin, Clemente Pasti, Gabriel Grand, Tianyu Liu 0004, Yahya Emara, Marjorie Freedman, Jason Eisner, Ryan Cotterell, Vikash Mansinghka 0001, Alexander K. Lew, Tim Vieira, Timothy J. O'Donnell |
ICLR | 2 |
| 2025 | Language Models over Canonical Byte-Pair EncodingsabstractModern language models represent probability distributions over character strings as distributions over (shorter) token strings derived via a deterministic tokenizer, such as byte-pair encoding. While this approach is highly effective at scaling up language models to large corpora, its current incarnations have a concerning property: the model assigns nonzero probability mass to an exponential number of noncanonical token encodings of each character string—these are token strings that decode to valid character strings but are impossible under the deterministic tokenizer (i.e., they will never be seen in any training corpus, no matter how large). This misallocation is both erroneous, as noncanonical strings never appear in training data, and wasteful, diverting probability mass away from plausible outputs. These are avoidable mistakes! In this work, we propose methods to enforce canonicality in token-level language models, ensuring that only canonical token strings are assigned positive probability. We present two approaches: (1) canonicality by conditioning, leveraging test-time inference strategies without additional training, and (2) canonicality by construction, a model parameterization that guarantees canonical outputs but requires training. We demonstrate that fixing canonicality mistakes improves the likelihood of held-out data for several models and corpora. Tim Vieira, Tianyu Liu 0004, Clemente Pasti, Yahya Emara, Brian DuSell, Benjamin LeBrun, Mario Giulianelli, Juan Luis Gastaldi, Timothy J. O'Donnell, Ryan Cotterell |
ICML | 6 |
| 2025 | From Language Models over Tokens to Language Models over CharactersabstractModern language models are internally—and mathematically—distributions over token strings rather than character strings, posing numerous challenges for programmers building user applications on top of them. For example, if a prompt is specified as a character string, it must be tokenized before passing it to the token-level language model. Thus, the tokenizer and consequent processing are very sensitive to the specification of the prompt (e.g., whether the prompt ends with a space or not). This paper presents algorithms for converting token-level language models to character-level ones. We present both exact and approximate algorithms. In the empirical portion of the paper, we benchmark the practical runtime and approximation quality. Across four publicly available language models, we find that—even with a small computation budget—our method is able to accurately approximate the character-level distribution at reasonably fast speeds, and that a significant improvement in the language model’s compression rate (bits/byte) is achieved. Tim Vieira, Benjamin LeBrun, Mario Giulianelli, Juan Luis Gastaldi, Brian DuSell, John Terilla, Timothy J. O'Donnell, Ryan Cotterell |
ICML | 2 |
| 2022 | Evaluating Distributional Distortion in Neural Language Modeling
Benjamin LeBrun, Alessandro Sordoni, Timothy J. O'Donnell |
ICLR | 1 |