Jaap Jumelet

dblp:225/7711 · DBLP profile ↗
← Back
11ranked-venue papers
2as first author
10since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 2 first-author · 10 since 2021
YearPublicationVenuePosition
2026 Vocabulary Shapes Cross-Lingual Variation of Word-Order Learnability in Language Models
abstract
Why do some languages like Czech permit free word order, while others like English do not?We address this question by pretraining transformer language models on a spectrum of synthetic word-order variants of natural languages.We observe that greater word-order irregularity consistently raises model surprisal, indicating reduced learnability.Sentence reversal, however, affects learnability only weakly.A coarse distinction of free-(e.g., Czech and Finnish) and fixed-word-order languages (e.g., English and French) does not explain crosslingual variation.Instead, the structure of the word and subword vocabulary strongly predicts the model surprisal.Overall, vocabulary structure emerges as a key driver of computational word-order learnability across languages.
Jonas Mayer Martins, Jaap Jumelet, Viola Priesemann, Lisa Beinborn
ACL (1)2
2026 CAIT: A Syntactic Parsing Toolkit for Child-Adult InTeractions
abstract
Francesca Padovani, Xiulin Yang, Bastian Bunzeck, Jaap Jumelet, Yevgen Matusevych, Nathan Schneider, Arianna Bisazza. Proceedings of the 30th Conference on Computational Natural Language Learning. 2026.
Francesca Padovani, Xiulin Yang, Bastian Bunzeck, Jaap Jumelet, Yevgen Matusevych, Nathan Schneider 0001, Arianna Bisazza
CoNLL4
2026 MultiBLiMP 1.0: A Massively Multilingual Benchmark of Linguistic Minimal Pairs
abstract
Abstract We introduce MultiBLiMP 1.0, a massively multilingual benchmark of linguistic minimal pairs, covering 101 languages and 2 types of subject-verb agreement, containing more than 128,000 minimal pairs. Our minimal pairs are created using a fully automated pipeline, leveraging the large-scale linguistic resources of Universal Dependencies and UniMorph. MultiBLiMP 1.0 evaluates abilities of LLMs at an unprecedented multilingual scale, and highlights the shortcomings of the current state-of-the-art in modelling low-resource languages.1
Jaap Jumelet, Leonie Weissweiler, Joakim Nivre, Arianna Bisazza
Trans. Assoc. Comput. Linguistics1
2025 TurBLiMP: A Turkish Benchmark of Linguistic Minimal Pairs
abstract
We introduce TurBLiMP, the first Turkish benchmark of linguistic minimal pairs, designed to evaluate the linguistic abilities of monolingual and multilingual language models (LMs).Covering 16 linguistic phenomena with 1000 minimal pairs each, TurBLiMP fills an important gap in linguistic evaluation resources for Turkish.In designing the benchmark, we give extra attention to two properties of Turkish that remain understudied in current syntactic evaluations of LMs, namely word order flexibility and subordination through morphological processes.Our experiments on a wide range of LMs and a newly collected set of human acceptability judgments reveal that even cutting-edge Large LMs still struggle with grammatical phenomena that are not challenging for humans, and may also exhibit different sensitivities to word order and morphological complexity compared to humans.
Ezgi Basar, Francesca Padovani, Jaap Jumelet, Arianna Bisazza
EMNLP3
2025 Child-Directed Language Does Not Consistently Boost Syntax Learning in Language Models
abstract
Seminal work by Huebner et al. (2021) showed that language models (LMs) trained on English Child-Directed Language (CDL) can reach similar syntactic abilities as LMs trained on much larger amounts of adult-directed written text, suggesting that CDL could provide more effective LM training material than the commonly used internet-crawled data.However, the generalizability of these results across languages, model types, and evaluation settings remains unclear.We test this by comparing models trained on CDL vs. Wikipedia across two LM objectives (masked and causal), three languages (English, French, German), and three syntactic minimal-pair benchmarks.Our results on these benchmarks show inconsistent benefits of CDL, which in most cases is outperformed by Wikipedia models.We then identify various shortcomings in previous benchmarks, and introduce a novel testing methodology, FIT-CLAMS, which uses a frequency-controlled design to enable balanced comparisons across training corpora.Through minimal pair evaluations and regression analysis we show that training on CDL does not yield stronger generalizations for acquiring syntax and highlight the importance of controlling for frequency effects when evaluating syntactic ability. 1
Francesca Padovani, Jaap Jumelet, Yevgen Matusevych, Arianna Bisazza
EMNLP2
2024 Interpretability of Language Models via Task Spaces
abstract
The usual way to interpret language models (LMs) is to test their performance on different benchmarks and subsequently infer their internal processes.In this paper, we present an alternative approach, concentrating on the quality of LM processing, with a focus on their language abilities.To this end, we construct 'linguistic task spaces' -representations of an LM's language conceptualisation -that shed light on the connections LMs draw between language phenomena.Task spaces are based on the interactions of the learning signals from different linguistic phenomena, which we assess via a method we call 'similarity probing'.To disentangle the learning signals of linguistic phenomena, we further introduce a method called 'fine-tuning via gradient differentials' (FTGD).We apply our methods to language models of three different scales and find that larger models generalise better to overarching general concepts for linguistic tasks, making better use of their shared structure.Further, the distributedness of linguistic processing increases with pre-training through increased parameter sharing between related linguistic tasks.The overall generalisation patterns are mostly stable throughout training and not marked by incisive stages, potentially explaining the lack of successful curriculum strategies for LMs.
Lucas Weber, Jaap Jumelet, Elia Bruni, Dieuwke Hupkes
ACL (1)2
2024 Filtered Corpus Training (FiCT) Shows that Language Models Can Generalize from Indirect Evidence
abstract
Abstract This paper introduces Filtered Corpus Training, a method that trains language models (LMs) on corpora with certain linguistic constructions filtered out from the training data, and uses it to measure the ability of LMs to perform linguistic generalization on the basis of indirect evidence. We apply the method to both LSTM and Transformer LMs (of roughly comparable size), developing filtered corpora that target a wide range of linguistic phenomena. Our results show that while transformers are better qua LMs (as measured by perplexity), both models perform equally and surprisingly well on linguistic generalization measures, suggesting that they are capable of generalizing from indirect evidence.
Abhinav Patil, Jaap Jumelet, Yu Ying Chiu, Andy Lapastora, Peter Shen, Lexie Wang, Clevis Willrich, Shane Steinert-Threlkeld
Trans. Assoc. Comput. Linguistics2
2023 Attribution and Alignment: Effects of Local Context Repetition on Utterance Production and Comprehension in Dialogue
abstract
Language models are often used as the backbone of modern dialogue systems.These models are pre-trained on large amounts of written fluent language.Repetition is typically penalised when evaluating language model generations.However, it is a key component of dialogue.Humans use local and partner specific repetitions; these are preferred by human users and lead to more successful communication in dialogue.In this study, we evaluate (a) whether language models produce humanlike levels of repetition in dialogue, and (b) what are the processing mechanisms related to lexical re-use they use during comprehension.We believe that such joint analysis of model production and comprehension behaviour can inform the development of cognitively inspired dialogue generation systems.
Aron Molnar, Jaap Jumelet, Mario Giulianelli, Arabella Sinclair
CoNLL2
2022 Structural Persistence in Language Models: Priming as a Window into Abstract Language Representations
abstract
Abstract We investigate the extent to which modern neural language models are susceptible to structural priming, the phenomenon whereby the structure of a sentence makes the same structure more probable in a follow-up sentence. We explore how priming can be used to study the potential of these models to learn abstract structural information, which is a prerequisite for good performance on tasks that require natural language understanding skills. We introduce a novel metric and release Prime-LM, a large corpus where we control for various linguistic factors that interact with priming strength. We find that Transformer models indeed show evidence of structural priming, but also that the generalizations they learned are to some extent modulated by semantic information. Our experiments also show that the representations acquired by the models may not only encode abstract sequential structure but involve certain level of hierarchical syntactic information. More generally, our study shows that the priming paradigm is a useful, additional tool for gaining insights into the capacities of language models and opens the door to future priming-based investigations that probe the model’s internal states.1
Arabella Sinclair, Jaap Jumelet, Willem H. Zuidema, Raquel Fernández
Trans. Assoc. Comput. Linguistics2
2021 Language Modelling as a Multi-Task Problem
abstract
In this paper, we propose to study language modelling as a multi-task problem, bringing together three strands of research: multitask learning, linguistics, and interpretability.Based on hypotheses derived from linguistic theory, we investigate whether language models adhere to learning principles of multi-task learning during training.To showcase the idea, we analyse the generalisation behaviour of language models as they learn the linguistic concept of Negative Polarity Items (NPIs).Our experiments demonstrate that a multi-task setting naturally emerges within the objective of the more general task of language modelling.We argue that this insight is valuable for multitask learning, linguistics and interpretability research and can lead to exciting new findings in all three domains.
Lucas Weber, Jaap Jumelet, Elia Bruni, Dieuwke Hupkes
EACL2
2019 Analysing Neural Language Models: Contextual Decomposition Reveals Default Reasoning in Number and Gender Assignment
abstract
Extensive research has recently shown that recurrent neural language models are able to process a wide range of grammatical phenomena.How these models are able to perform these remarkable feats so well, however, is still an open question.To gain more insight into what information LSTMs base their decisions on, we propose a generalisation of Contextual Decomposition (GCD).In particular, this setup enables us to accurately distil which part of a prediction stems from semantic heuristics, which part truly emanates from syntactic cues and which part arise from the model biases themselves instead.We investigate this technique on tasks pertaining to syntactic agreement and co-reference resolution and discover that the model strongly relies on a default reasoning effect to perform these tasks.
Jaap Jumelet, Willem H. Zuidema, Dieuwke Hupkes
CoNLL1