EDBT 2026 Demo / reviewers in the wild / expert
Miryam de Lhoneux
dblp:163/1873
· DBLP profile ↗
18ranked-venue papers
2as first author
13since 2021 · last 2026
0000-0001-8844-2126ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 2 first-author · 13 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | How Good is Your Wikipedia? Auditing Data Quality for Low-resource and Multilingual NLPabstractKushal Tatariya, Artur Kulmizev, Wessel Poelman, Esther Ploeger, Marcel Bollmann, Johannes Bjerva, Jiaming Luo, Heather Lent, Miryam de Lhoneux. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Kushal Tatariya, Artur Kulmizev, Wessel Poelman, Esther Ploeger, Marcel Bollmann, Johannes Bjerva, Jiaming Luo, Heather C. Lent, Miryam de Lhoneux |
ACL (1) | 9 |
| 2025 | GRaMPa: Subword Regularisation by Skewing Uniform Segmentation Distributions with an Efficient Path-counting Markov ModelabstractStochastically sampling word segmentations from a subword tokeniser, also called subword regularisation, is a known way to increase robustness of language models to out-of-distribution inputs, such as text containing spelling errors. Recent work has observed that usual augmentations that make popular deterministic subword tokenisers stochastic still cause only a handful of all possible segmentations to be sampled. It has been proposed to uniformly sample across these instead, through rejection sampling of paths in an unweighted segmentation graph. In this paper, we argue that uniformly random segmentation in turn skews the distributions of certain segmentational properties (e.g. token lengths and amount of tokens produced) away from uniformity, which still ends up hiding meaningfully diverse tokenisations. We propose an alternative uniform sampler using the same segmentation graph, but weighted by counting the paths through it. Our sampling algorithm, GRaMPa, provides hyperparameters allowing sampled tokenisations to skew towards fewer, longer tokens. Furthermore, GRaMPa is single-pass, guaranteeing significantly better computational complexity than previous approaches relying on rejection sampling. We show experimentally that language models trained with GRaMPa outperform existing regularising tokenisers in a data-scarce setting on token-level tasks such as dependency parsing, especially with spelling errors present. Thomas Bauwens, David Kaczér, Miryam de Lhoneux |
ACL (1) | 3 |
| 2025 | Confounding Factors in Relating Model Performance to MorphologyabstractThe extent to which individual language characteristics influence tokenization and language modeling is an open question.Differences in morphological systems have been suggested as both unimportant and crucial to consider (Cotterell et al., 2018; Gerz et al., 2018a; Park et al., 2021, inter alia).We argue this conflicting evidence is due to confounding factors in experimental setups, making it hard to compare results and draw conclusions.We identify such factors in analyses trying to answer the question of whether, and how, morphology relates to language modeling.Next, we re-assess three hypotheses by Arnett and Bergen (2025) for why modeling agglutinative languages results in higher perplexities than fusional languages: they look at morphological alignment of tokenization, tokenization efficiency, and dataset size.We show that each conclusion includes confounding factors and suggest methodological improvements.Finally, we introduce token bigram metrics as an intrinsic way to predict the difficulty of causal language modeling, and find that they are gradient proxies for morphological complexity that do not require expert annotation.Ultimately, we outline necessities to reliably answer whether, and how, morphology relates to language modeling. Wessel Poelman, Thomas Bauwens, Miryam de Lhoneux |
EMNLP | 3 |
| 2024 | What is "Typological Diversity" in NLP?abstractThe NLP research community has devoted increased attention to languages beyond English, resulting in considerable improvements for multilingual NLP.However, these improvements only apply to a small subset of the world's languages.An increasing number of papers aspires to enhance generalizable multilingual performance across languages.To this end, linguistic typology is commonly used to motivate language selection, on the basis that a broad typological sample ought to imply generalization across a broad range of languages.These selections are often described as being 'typologically diverse'.In this meta-analysis, we systematically investigate NLP research that includes claims regarding typological diversity.We find there are no set definitions or criteria for such claims.We introduce metrics to approximate the diversity of resulting language samples along several axes and find that the results vary considerably across papers.Crucially, we show that skewed language selection can lead to overestimated multilingual performance.We recommend future work to include an operationalization of typological diversity that empirically justifies the diversity of language samples.To help facilitate this, we release the code for our diversity measures.1 * Equal contribution. 1 Our code and data are publicly available: https://github.com/WPoelman/typ-div Esther Ploeger, Wessel Poelman, Miryam de Lhoneux, Johannes Bjerva |
EMNLP | 3 |
| 2024 | Pixology: Probing the Linguistic and Visual Capabilities of Pixel-based Language ModelsabstractPixel-based language models have emerged as a compelling alternative to subword-based language modelling, particularly because they can represent virtually any script.PIXEL, a canonical example of such a model, is a vision transformer that has been pre-trained on rendered text.While PIXEL has shown promising cross-script transfer abilities and robustness to orthographic perturbations, it falls short of outperforming monolingual subword counterparts like BERT in most other contexts.This discrepancy raises questions about the amount of linguistic knowledge learnt by these models and whether their performance in language tasks stems more from their visual capabilities than their linguistic ones.To explore this, we probe PIXEL using a variety of linguistic and visual tasks to assess its position on the vision-to-language spectrum.Our findings reveal a substantial gap between the model's visual and linguistic understanding.The lower layers of PIXEL predominantly capture superficial visual features, whereas the higher layers gradually learn more syntactic and semantic abstractions.Additionally, we examine variants of PIXEL trained with different text rendering strategies, discovering that introducing certain orthographic constraints at the input level can facilitate earlier learning of surface-level features.With this study, we hope to provide insights that aid the further development of pixelbased language models. 1 Kushal Tatariya, Vladimir Araujo, Thomas Bauwens, Miryam de Lhoneux |
EMNLP | 4 |
| 2024 | CreoleVal: Multilingual Multitask Benchmarks for CreolesabstractAbstract Creoles represent an under-explored and marginalized group of languages, with few available resources for NLP research. While the genealogical ties between Creoles and a number of highly resourced languages imply a significant potential for transfer learning, this potential is hampered due to this lack of annotated data. In this work we present CreoleVal, a collection of benchmark datasets spanning 8 different NLP tasks, covering up to 28 Creole languages; it is an aggregate of novel development datasets for reading comprehension relation classification, and machine translation for Creoles, in addition to a practical gateway to a handful of preexisting benchmarks. For each benchmark, we conduct baseline experiments in a zero-shot setting in order to further ascertain the capabilities and limitations of transfer learning for Creoles. Ultimately, we see CreoleVal as an opportunity to empower research on Creoles in NLP and computational linguistics, and in general, a step towards more equitable language technology around the globe. Heather C. Lent, Kushal Tatariya, Raj Dabre, Yiyi Chen 0002, Marcell Fekete, Esther Ploeger, Li Zhou 0010, Ruth-Ann Armstrong, Abee Eijansantos, Catriona Malau, Hans Erik Heje, Ernests Lavrinovics, Diptesh Kanojia, Paul Belony, Marcel Bollmann, Loïc Grobol, Miryam de Lhoneux, Daniel Hershcovich, Michel DeGraff, Anders Søgaard, Johannes Bjerva |
Trans. Assoc. Comput. Linguistics | 17 |
| 2023 | A Two-Sided Discussion of Preregistration of NLP ResearchabstractVan Miltenburg et al. (2021) suggest NLP research should adopt preregistration to prevent fishing expeditions and to promote publication of negative results.At face value, this is a very reasonable suggestion, seemingly solving many methodological problems with NLP research.We discuss pros and cons-some old, some new: a) Preregistration is challenged by the practice of retrieving hypotheses after the results are known; b) preregistration may bias NLP toward confirmatory research; c) preregistration must allow for reclassification of research as exploratory; d) preregistration may increase publication bias; e) preregistration may increase flag-planting; f) preregistration may increase p-hacking; and finally, g) preregistration may make us less risk tolerant.We cast our discussion as a dialogue, presenting both sides of the debate. Anders Søgaard, Daniel Hershcovich, Miryam de Lhoneux |
EACL | 3 |
| 2023 | Language Modelling with Pixels
Phillip Rust, Jonas F. Lotz, Emanuele Bugliarello, Elizabeth Salesky, Miryam de Lhoneux, Desmond Elliott |
ICLR | 5 |
| 2022 | Challenges and Strategies in Cross-Cultural NLPabstractDaniel Hershcovich, Stella Frank, Heather Lent, Miryam de Lhoneux, Mostafa Abdou, Stephanie Brandl, Emanuele Bugliarello, Laura Cabello Piqueras, Ilias Chalkidis, Ruixiang Cui, Constanza Fierro, Katerina Margatina, Phillip Rust, Anders Søgaard. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Daniel Hershcovich, Stella Frank, Heather C. Lent, Miryam de Lhoneux, Mostafa Abdou, Stephanie Brandl, Emanuele Bugliarello, Laura Cabello Piqueras, Ilias Chalkidis, Ruixiang Cui, Constanza Fierro, Aikaterini Margatina, Phillip Rust, Anders Søgaard |
ACL (1) | 4 |
| 2022 | Finding Structural Knowledge in Multimodal-BERTabstractIn this work, we investigate the knowledge learned in the embeddings of multimodal-BERT models.More specifically, we probe their capabilities of storing the grammatical structure of linguistic data and the structure learned over objects in visual data.To reach that goal, we first make the inherent structure of language and visuals explicit by a dependency parse of the sentences that describe the image and by the dependencies between the object regions in the image, respectively.We call this explicit visual structure the scene tree, that is based on the dependency tree of the language description.Extensive probing experiments show that the multimodal-BERT models do not encode these scene trees.Code available at https://github. com/VSJMilewski/multimodal-probes. Victor Milewski, Miryam de Lhoneux, Marie-Francine Moens |
ACL (1) | 2 |
| 2022 | What a Creole Wants, What a Creole NeedsabstractIn recent years, the natural language processing (NLP) community has given increased attention to the disparity of efforts directed towards high-resource languages over low-resource ones. Efforts to remedy this delta often begin with translations of existing English datasets into other languages. However, this approach ignores that different language communities have different needs. We consider a group of low-resource languages, creole languages. Creoles are both largely absent from the NLP literature, and also often ignored by society at large due to stigma, despite these languages having sizable and vibrant communities. We demonstrate, through conversations with creole experts and surveys of creole-speaking communities, how the things needed from language technology can change dramatically from one language to another, even when the languages are considered to be very similar to each other, as with creoles. We discuss the prominent themes arising from these conversations, and ultimately demonstrate that useful language technology cannot be built without involving the relevant community. Heather C. Lent, Kelechi Ogueji, Miryam de Lhoneux, Orevaoghene Ahia, Anders Søgaard |
LREC | 3 |
| 2021 | A Multilingual Benchmark for Probing Negation-Awareness with Minimal PairsabstractMareike Hartmann, Miryam de Lhoneux, Daniel Hershcovich, Yova Kementchedjhieva, Lukas Nielsen, Chen Qiu, Anders Søgaard. Proceedings of the 25th Conference on Computational Natural Language Learning. 2021. Mareike Hartmann, Miryam de Lhoneux, Daniel Hershcovich, Yova Kementchedjhieva, Lukas Nielsen, Chen Qiu 0005, Anders Søgaard |
CoNLL | 2 |
| 2021 | On Language Models for CreolesabstractCreole languages such as Nigerian Pidgin English and Haitian Creole are under-resourced and largely ignored in the NLP literature.Creoles typically result from the fusion of a foreign language with multiple local languages, and what grammatical and lexical features are transferred to the creole is a complex process (Sessarego, 2020).While creoles are generally stable, the prominence of some features may be much stronger with certain demographics or in some linguistic situations (Winford, 1999;Patrick, 1999).This paper makes several contributions: We collect existing corpora and release models for Haitian Creole, Nigerian Pidgin English, and Singaporean Colloquial English.We evaluate these models on intrinsic and extrinsic tasks.Motivated by the above literature, we compare standard language models with distributionally robust ones and find that, somewhat surprisingly, the standard language models are superior to the distributionally robust ones.We investigate whether this is an effect of overparameterization or relative distributional stability, and find that the difference persists in the absence of over-parameterization, and that drift is limited, confirming the relative stability of creole languages. Heather C. Lent, Emanuele Bugliarello, Miryam de Lhoneux, Chen Qiu 0005, Anders Søgaard |
CoNLL | 3 |
| 2020 | Comparison by Conversion: Reverse-Engineering UCCA from Syntax and Lexical SemanticsabstractBuilding robust natural language understanding systems will require a clear characterization of whether and how various linguistic meaning representations complement each other.To perform a systematic comparative analysis, we evaluate the mapping between meaning representations from different frameworks using two complementary methods: (i) a rule-based converter, and (ii) a supervised delexicalized parser that parses to one framework using only information from the other as features.We apply these methods to convert the STREUSLE corpus (with syntactic and lexical semantic annotations) to UCCA (a graph-structured full-sentence meaning representation).Both methods yield surprisingly accurate target representations, close to fully supervised UCCA parser quality-indicating that UCCA annotations are partially redundant with STREUSLE annotations.Despite this substantial convergence between frameworks, we find several important areas of divergence. Daniel Hershcovich, Nathan Schneider 0001, Dotan Dvir, Jakob Prange, Miryam de Lhoneux, Omri Abend |
COLING | 5 |
| 2020 | What Should/Do/Can LSTMs Learn When Parsing Auxiliary Verb Constructions?abstractThere is a growing interest in investigating what neural NLP models learn about language. A prominent open question is the question of whether or not it is necessary to model hierarchical structure. We present a linguistic investigation of a neural parser adding insights to this question. We look at transitivity and agreement information of auxiliary verb constructions (AVCs) in comparison to finite main verbs (FMVs). This comparison is motivated by theoretical work in dependency grammar and in particular the work of Tesnière ( 1959 ), where AVCs and FMVs are both instances of a nucleus, the basic unit of syntax. An AVC is a dissociated nucleus; it consists of at least two words, and an FMV is its non-dissociated counterpart, consisting of exactly one word. We suggest that the representation of AVCs and FMVs should capture similar information. We use diagnostic classifiers to probe agreement and transitivity information in vectors learned by a transition-based neural parser in four typologically different languages. We find that the parser learns different information about AVCs and FMVs if only sequential models (BiLSTMs) are used in the architecture but similar information when a recursive layer is used. We find explanations for why this is the case by looking closely at how information is learned in the network and looking at what happens with different dependency representations of AVCs. We conclude that there may be benefits to using a recursive layer in dependency parsing and that we have not yet found the best way to integrate it in our parsers. Miryam de Lhoneux, Sara Stymne, Joakim Nivre |
Comput. Linguistics | 1 |
| 2019 | Deep Contextualized Word Embeddings in Transition-Based and Graph-Based Dependency Parsing - A Tale of Two Parsers RevisitedabstractArtur Kulmizev, Miryam de Lhoneux, Johannes Gontrum, Elena Fano, Joakim Nivre. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Artur Kulmizev, Miryam de Lhoneux, Johannes Gontrum, Elena Fano, Joakim Nivre |
EMNLP/IJCNLP (1) | 2 |
| 2018 | Parameter sharing between dependency parsers for related languagesabstractPrevious work has suggested that parameter sharing between transition-based neural dependency parsers for related languages can lead to better performance, but there is no consensus on what parameters to share.We present an evaluation of 27 different parameter sharing strategies across 10 languages, representing five pairs of related languages, each pair from a different language family.We find that sharing transition classifier parameters always helps, whereas the usefulness of sharing word and/or character LSTM parameters varies.Based on this result, we propose an architecture where the transition classifier is shared, and the sharing of word and character parameters is controlled by a parameter that can be tuned on validation data.This model is linguistically motivated and obtains significant improvements over a mono-lingually trained baseline.We also find that sharing transition classifier parameters helps when training a parser on unrelated language pairs, but we find that, in the case of unrelated languages, sharing too many parameters does not help. Miryam de Lhoneux, Johannes Bjerva, Isabelle Augenstein, Anders Søgaard |
EMNLP | 1 |
| 2018 | An Investigation of the Interactions Between Pre-Trained Word Embeddings, Character Models and POS Tags in Dependency ParsingabstractWe provide a comprehensive analysis of the interactions between pre-trained word embeddings, character models and POS tags in a transition-based dependency parser.While previous studies have shown POS information to be less important in the presence of character models, we show that in fact there are complex interactions between all three techniques.In isolation each produces large improvements over a baseline system using randomly initialised word embeddings only, but combining them quickly leads to diminishing returns.We categorise words by frequency, POS tag and language in order to systematically investigate how each of the techniques affects parsing quality.For many word categories, applying any two of the three techniques is almost as good as the full combined system.Character models tend to be more important for low-frequency open-class words, especially in morphologically rich languages, while POS tags can help disambiguate highfrequency function words.We also show that large character embedding sizes help even for languages with small character sets, especially in morphologically rich languages. Aaron Smith, Miryam de Lhoneux, Sara Stymne, Joakim Nivre |
EMNLP | 2 |