Thomas Bauwens

dblp:384/2142 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
4since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Language models and text generation · 62% Representation and self-supervised learning · 18% Trustworthy machine learning · 10%

Topics — the 10 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
language modeling
0.912025
Confounding Factors in Relating Model Performance to Morphology · EMNLP 2025
Machine learning › Trustworthy machine learning › robustness › distribution shift
robustness to distribution shift
0.912025
GRaMPa: Subword Regularisation by Skewing Uniform Segmentation Distributions with an Efficient Path-counting Markov Model · ACL (1) 2025
Natural language and speech › Language models and text generation › tokenization › subword tokenization
subword regularization
0.912025
GRaMPa: Subword Regularisation by Skewing Uniform Segmentation Distributions with an Efficient Path-counting Markov Model · ACL (1) 2025
Natural language and speech › Language models and text generation › tokenization
subword tokenization
0.912025
GRaMPa: Subword Regularisation by Skewing Uniform Segmentation Distributions with an Efficient Path-counting Markov Model · ACL (1) 2025
Natural language and speech › Language models and text generation
tokenization
0.912025
Confounding Factors in Relating Model Performance to Morphology · EMNLP 2025
Natural language and speech › Language models and text generation › tokenization
tokenization robustness
0.912025
GRaMPa: Subword Regularisation by Skewing Uniform Segmentation Distributions with an Efficient Path-counting Markov Model · ACL (1) 2025
Machine learning › Representation and self-supervised learning › representation analysis
layer-wise representation analysis
0.812024
Pixology: Probing the Linguistic and Visual Capabilities of Pixel-based Language Models · EMNLP 2024
Natural language and speech › Information extraction and text analysis
linguistic probing
0.812024
Pixology: Probing the Linguistic and Visual Capabilities of Pixel-based Language Models · EMNLP 2024
Natural language and speech › Language models and text generation › multimodal language model
pixel-based language models
0.812024
Pixology: Probing the Linguistic and Visual Capabilities of Pixel-based Language Models · EMNLP 2024
Machine learning › Representation and self-supervised learning
probing
0.812024
Pixology: Probing the Linguistic and Visual Capabilities of Pixel-based Language Models · EMNLP 2024

Methods — techniques the papers use, named apart from their topics

uniform sampling · 0.9token bigram metrics · 0.9path counting · 0.9markov model · 0.9vision transformer analysis · 0.8probing tasks · 0.8
YearPublicationVenuePosition
2025 GRaMPa: Subword Regularisation by Skewing Uniform Segmentation Distributions with an Efficient Path-counting Markov Model
abstract
Stochastically sampling word segmentations from a subword tokeniser, also called subword regularisation, is a known way to increase robustness of language models to out-of-distribution inputs, such as text containing spelling errors. Recent work has observed that usual augmentations that make popular deterministic subword tokenisers stochastic still cause only a handful of all possible segmentations to be sampled. It has been proposed to uniformly sample across these instead, through rejection sampling of paths in an unweighted segmentation graph. In this paper, we argue that uniformly random segmentation in turn skews the distributions of certain segmentational properties (e.g. token lengths and amount of tokens produced) away from uniformity, which still ends up hiding meaningfully diverse tokenisations. We propose an alternative uniform sampler using the same segmentation graph, but weighted by counting the paths through it. Our sampling algorithm, GRaMPa, provides hyperparameters allowing sampled tokenisations to skew towards fewer, longer tokens. Furthermore, GRaMPa is single-pass, guaranteeing significantly better computational complexity than previous approaches relying on rejection sampling. We show experimentally that language models trained with GRaMPa outperform existing regularising tokenisers in a data-scarce setting on token-level tasks such as dependency parsing, especially with spelling errors present.
Thomas Bauwens, David Kaczér, Miryam de Lhoneux
ACL (1)1
2025 Confounding Factors in Relating Model Performance to Morphology
abstract
The extent to which individual language characteristics influence tokenization and language modeling is an open question.Differences in morphological systems have been suggested as both unimportant and crucial to consider (Cotterell et al., 2018; Gerz et al., 2018a; Park et al., 2021, inter alia).We argue this conflicting evidence is due to confounding factors in experimental setups, making it hard to compare results and draw conclusions.We identify such factors in analyses trying to answer the question of whether, and how, morphology relates to language modeling.Next, we re-assess three hypotheses by Arnett and Bergen (2025) for why modeling agglutinative languages results in higher perplexities than fusional languages: they look at morphological alignment of tokenization, tokenization efficiency, and dataset size.We show that each conclusion includes confounding factors and suggest methodological improvements.Finally, we introduce token bigram metrics as an intrinsic way to predict the difficulty of causal language modeling, and find that they are gradient proxies for morphological complexity that do not require expert annotation.Ultimately, we outline necessities to reliably answer whether, and how, morphology relates to language modeling.
Wessel Poelman, Thomas Bauwens, Miryam de Lhoneux
EMNLP2
2024 Pixology: Probing the Linguistic and Visual Capabilities of Pixel-based Language Models
abstract
Pixel-based language models have emerged as a compelling alternative to subword-based language modelling, particularly because they can represent virtually any script.PIXEL, a canonical example of such a model, is a vision transformer that has been pre-trained on rendered text.While PIXEL has shown promising cross-script transfer abilities and robustness to orthographic perturbations, it falls short of outperforming monolingual subword counterparts like BERT in most other contexts.This discrepancy raises questions about the amount of linguistic knowledge learnt by these models and whether their performance in language tasks stems more from their visual capabilities than their linguistic ones.To explore this, we probe PIXEL using a variety of linguistic and visual tasks to assess its position on the vision-to-language spectrum.Our findings reveal a substantial gap between the model's visual and linguistic understanding.The lower layers of PIXEL predominantly capture superficial visual features, whereas the higher layers gradually learn more syntactic and semantic abstractions.Additionally, we examine variants of PIXEL trained with different text rendering strategies, discovering that introducing certain orthographic constraints at the input level can facilitate earlier learning of surface-level features.With this study, we hope to provide insights that aid the further development of pixelbased language models. 1
Kushal Tatariya, Vladimir Araujo, Thomas Bauwens, Miryam de Lhoneux
EMNLP3
2024 BPE-knockout: Pruning Pre-existing BPE Tokenisers with Backwards-compatible Morphological Semi-supervision
abstract
Thomas Bauwens, Pieter Delobelle. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Thomas Bauwens, Pieter Delobelle
NAACL-HLT1