Benjamin LeBrun

dblp:317/0720 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
4since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Language models and text generation · 78% Probabilistic and Bayesian machine learning · 19% Efficient and distributed learning · 3%

Topics — the 9 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation › language modeling
statistical language modeling
1.722025
From Language Models over Tokens to Language Models over Characters · ICML 2025
Language Models over Canonical Byte-Pair Encodings · ICML 2025
Natural language and speech › Language models and text generation
tokenization
1.722025
From Language Models over Tokens to Language Models over Characters · ICML 2025
Language Models over Canonical Byte-Pair Encodings · ICML 2025
Natural language and speech › Language models and text generation › tokenization › subword tokenization
byte-pair encoding
0.912025
Language Models over Canonical Byte-Pair Encodings · ICML 2025
Natural language and speech › Language models and text generation › language modeling
character-level language modeling
0.912025
From Language Models over Tokens to Language Models over Characters · ICML 2025
Machine learning › Probabilistic and Bayesian machine learning
probabilistic inference
0.912025
Syntactic and Semantic Control of Large Language Models via Sequential Monte Carlo · ICLR 2025
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods
sequential monte carlo
0.912025
Syntactic and Semantic Control of Large Language Models via Sequential Monte Carlo · ICLR 2025
Natural language and speech › Language models and text generation
text generation
0.912025
Syntactic and Semantic Control of Large Language Models via Sequential Monte Carlo · ICLR 2025
Natural language and speech › Language models and text generation
neural language model
0.612022
Evaluating Distributional Distortion in Neural Language Modeling · ICLR 2022
Machine learning › Efficient and distributed learning
compression
0.312025
From Language Models over Tokens to Language Models over Characters · ICML 2025

Methods — techniques the papers use, named apart from their topics

sequential monte carlo · 0.9probabilistic conditioning · 0.9model reparameterization · 0.9exact and approximate marginalization · 0.9constrained inference · 0.9
YearPublicationVenuePosition
2025 Syntactic and Semantic Control of Large Language Models via Sequential Monte Carlo
abstract
A wide range of LM applications require generating text that conforms to syntactic or semantic constraints. Imposing such constraints can be naturally framed as _probabilistic conditioning_, but exact generation from the resulting distribution—which can differ substantially from the LM’s base distribution—is generally intractable. In this work, we develop an architecture for controlled LM generation based on sequential Monte Carlo (SMC). Our SMC framework allows us to flexibly incorporate domain- and problem-specific constraints at inference time, and efficiently reallocate computational resources in light of new information during the course of generation. By comparing to a number of alternatives and ablations on four challenging domains---Python code generation for data science, text-to-SQL, goal inference, and molecule synthesis—we demonstrate that, with little overhead, our approach allows small open-source language models to outperform models over 8$\times$ larger, as well as closed-source, fine-tuned ones. In support of the probabilistic perspective, we show that these performance improvements are driven by better approximation to the posterior distribution. [Our system](https://github.com/probcomp/genlm-control) builds on the framework of Lew et al. (2023) and integrates with its _language model probabilistic programming language_, giving users a simple, programmable way to apply SMC to a broad variety of controlled generation problems.
João Loula, Benjamin LeBrun, Benjamin Lipkin, Clemente Pasti, Gabriel Grand, Tianyu Liu 0004, Yahya Emara, Marjorie Freedman, Jason Eisner, Ryan Cotterell, Vikash Mansinghka 0001, Alexander K. Lew, Tim Vieira, Timothy J. O'Donnell
ICLR2
2025 Language Models over Canonical Byte-Pair Encodings
abstract
Modern language models represent probability distributions over character strings as distributions over (shorter) token strings derived via a deterministic tokenizer, such as byte-pair encoding. While this approach is highly effective at scaling up language models to large corpora, its current incarnations have a concerning property: the model assigns nonzero probability mass to an exponential number of noncanonical token encodings of each character string—these are token strings that decode to valid character strings but are impossible under the deterministic tokenizer (i.e., they will never be seen in any training corpus, no matter how large). This misallocation is both erroneous, as noncanonical strings never appear in training data, and wasteful, diverting probability mass away from plausible outputs. These are avoidable mistakes! In this work, we propose methods to enforce canonicality in token-level language models, ensuring that only canonical token strings are assigned positive probability. We present two approaches: (1) canonicality by conditioning, leveraging test-time inference strategies without additional training, and (2) canonicality by construction, a model parameterization that guarantees canonical outputs but requires training. We demonstrate that fixing canonicality mistakes improves the likelihood of held-out data for several models and corpora.
Tim Vieira, Tianyu Liu 0004, Clemente Pasti, Yahya Emara, Brian DuSell, Benjamin LeBrun, Mario Giulianelli, Juan Luis Gastaldi, Timothy J. O'Donnell, Ryan Cotterell
ICML6
2025 From Language Models over Tokens to Language Models over Characters
abstract
Modern language models are internally—and mathematically—distributions over token strings rather than character strings, posing numerous challenges for programmers building user applications on top of them. For example, if a prompt is specified as a character string, it must be tokenized before passing it to the token-level language model. Thus, the tokenizer and consequent processing are very sensitive to the specification of the prompt (e.g., whether the prompt ends with a space or not). This paper presents algorithms for converting token-level language models to character-level ones. We present both exact and approximate algorithms. In the empirical portion of the paper, we benchmark the practical runtime and approximation quality. Across four publicly available language models, we find that—even with a small computation budget—our method is able to accurately approximate the character-level distribution at reasonably fast speeds, and that a significant improvement in the language model’s compression rate (bits/byte) is achieved.
Tim Vieira, Benjamin LeBrun, Mario Giulianelli, Juan Luis Gastaldi, Brian DuSell, John Terilla, Timothy J. O'Donnell, Ryan Cotterell
ICML2
2022 Evaluating Distributional Distortion in Neural Language Modeling
Benjamin LeBrun, Alessandro Sordoni, Timothy J. O'Donnell
ICLR1