Dieuwke Hupkes

dblp:184/8838 · DBLP profile ↗
← Back
22ranked-venue papers
5as first author
14since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 22 · 5 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Lost in Inference: Rediscovering the Role of Natural Language Inference for Large Language Models
abstract
Lovish Madaan, David Esiobu, Pontus Stenetorp, Barbara Plank, Dieuwke Hupkes. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Lovish Madaan, David Esiobu, Pontus Stenetorp, Barbara Plank, Dieuwke Hupkes
NAACL (Long Papers)5
2024 Interpretability of Language Models via Task Spaces
abstract
The usual way to interpret language models (LMs) is to test their performance on different benchmarks and subsequently infer their internal processes.In this paper, we present an alternative approach, concentrating on the quality of LM processing, with a focus on their language abilities.To this end, we construct 'linguistic task spaces' -representations of an LM's language conceptualisation -that shed light on the connections LMs draw between language phenomena.Task spaces are based on the interactions of the learning signals from different linguistic phenomena, which we assess via a method we call 'similarity probing'.To disentangle the learning signals of linguistic phenomena, we further introduce a method called 'fine-tuning via gradient differentials' (FTGD).We apply our methods to language models of three different scales and find that larger models generalise better to overarching general concepts for linguistic tasks, making better use of their shared structure.Further, the distributedness of linguistic processing increases with pre-training through increased parameter sharing between related linguistic tasks.The overall generalisation patterns are mostly stable throughout training and not marked by incisive stages, potentially explaining the lack of successful curriculum strategies for LMs.
Lucas Weber, Jaap Jumelet, Elia Bruni, Dieuwke Hupkes
ACL (1)4
2024 From Form(s) to Meaning: Probing the Semantic Depths of Language Models Using Multisense Consistency
abstract
Abstract The staggering pace with which the capabilities of large language models (LLMs) are increasing, as measured by a range of commonly used natural language understanding (NLU) benchmarks, raises many questions regarding what “understanding” means for a language model and how it compares to human understanding. This is especially true since many LLMs are exclusively trained on text, casting doubt on whether their stellar benchmark performances are reflective of a true understanding of the problems represented by these benchmarks, or whether LLMs simply excel at uttering textual forms that correlate with what someone who understands the problem would say. In this philosophically inspired work, we aim to create some separation between form and meaning, with a series of tests that leverage the idea that world understanding should be consistent across presentational modes—inspired by Fregean senses—of the same meaning. Specifically, we focus on consistency across languages as well as paraphrases. Taking GPT-3.5 as our object of study, we evaluate multisense consistency across five different languages and various tasks. We start the evaluation in a controlled setting, asking the model for simple facts, and then proceed with an evaluation on four popular NLU benchmarks. We find that the model’s multisense consistency is lacking and run several follow-up analyses to verify that this lack of consistency is due to a sense-dependent task understanding. We conclude that, in this aspect, the understanding of LLMs is still quite far from being consistent and human-like, and deliberate on how this impacts their utility in the context of learning about human language and understanding.
Xenia Ohmer, Elia Bruni, Dieuwke Hupkes
Comput. Linguistics3
2023 The Validity of Evaluation Results: Assessing Concurrence Across Compositionality Benchmarks
abstract
NLP models have progressed drastically in recent years, according to numerous datasets proposed to evaluate performance.Questions remain, however, about how particular dataset design choices may impact the conclusions we draw about model capabilities.In this work, we investigate this question in the domain of compositional generalization.We examine the performance of six modeling approaches across 4 datasets, split according to 8 compositional splitting strategies, ranking models by 18 compositional generalization splits in total.Our results show that: i) the datasets, although all designed to evaluate compositional generalization, rank modeling approaches differently; ii) datasets generated by humans align better with each other than they with synthetic datasets, or than synthetic datasets among themselves; iii) generally, whether datasets are sampled from the same source is more predictive of the resulting model ranking than whether they maintain the same interpretation of compositionality; and iv) which lexical items are used in the data can strongly impact conclusions.Overall, our results demonstrate that much work remains to be done when it comes to assessing whether popular evaluation datasets measure what they intend to measure, and suggests that elucidating more rigorous standards for establishing the validity of evaluation sets could benefit the field. 1
Kaiser Sun, Adina Williams, Dieuwke Hupkes
CoNLL3
2023 Mind the instructions: a holistic evaluation of consistency and interactions in prompt-based learning
abstract
Finding the best way of adapting pre-trained language models to a task is a big challenge in current NLP.Just like the previous generation of task-tuned models (TT), models that are adapted to tasks via in-context-learning (ICL) are robust in some setups but not in others.Here, we present a detailed analysis of which design choices cause instabilities and inconsistencies in LLM predictions.First, we show how spurious correlations between input distributions and labels -a known issue in TT models -form only a minor problem for prompted models.Then, we engage in a systematic, holistic evaluation of different factors that have been found to influence predictions in a prompting setup.We test all possible combinations of a range of factors on both vanilla and instructiontuned (IT) LLMs of different scale and statistically analyse the results to show which factors are the most influential, interactive or stable.Our results show which factors can be used without precautions and which should be avoided or handled with care in most settings.
Lucas Weber, Elia Bruni, Dieuwke Hupkes
CoNLL3
2023 Memorisation Cartography: Mapping out the Memorisation-Generalisation Continuum in Neural Machine Translation
abstract
When training a neural network, it will quickly memorise some source-target mappings from your dataset but never learn some others.Yet, memorisation is not easily expressed as a binary feature that is good or bad: individual datapoints lie on a memorisation-generalisation continuum.What determines a datapoint's position on that spectrum, and how does that spectrum influence neural models' performance?We address these two questions for neural machine translation (NMT) models.We use the counterfactual memorisation metric to (1) build a resource that places 5M NMT datapoints on a memorisation-generalisation map, (2) illustrate how the datapoints' surface-level characteristics and a models' per-datum training signals are predictive of memorisation in NMT, (3) and describe the influence that subsets of that map have on NMT systems' performance.1 * Work partially conducted during an internship at FAIR. 1 Click here to interactively explore the NMT memorisation maps in our demo.
Verna Dankers, Ivan Titov 0001, Dieuwke Hupkes
EMNLP3
2023 Neural Agents Struggle to Take Turns in Bidirectional Emergent Communication
Valentin Taillandier, Dieuwke Hupkes, Benoît Sagot, Emmanuel Dupoux, Paul Michel
ICLR2
2022 The Paradox of the Compositionality of Natural Language: A Neural Machine Translation Case Study
abstract
Obtaining human-like performance in NLP is often argued to require compositional generalisation.Whether neural networks exhibit this ability is usually studied by training models on highly compositional synthetic data.However, compositionality in natural language is much more complex than the rigid, arithmeticlike version such data adheres to, and artificial compositionality tests thus do not allow us to determine how neural models deal with more realistic forms of compositionality.In this work, we re-instantiate three compositionality tests from the literature and reformulate them for neural machine translation (NMT).Our results highlight that: i) unfavourably, models trained on more data are more compositional; ii) models are sometimes less compositional than expected, but sometimes more, exemplifying that different levels of compositionality are required, and models are not always able to modulate between them correctly; iii) some of the non-compositional behaviours are mistakes, whereas others reflect the natural variation in data.Apart from an empirical study, our work is a call to action: we should rethink the evaluation of compositionality in neural networks and develop benchmarks using real data to evaluate compositionality on natural language, where composing meaning is not as straightforward as doing the math. 1
Verna Dankers, Elia Bruni, Dieuwke Hupkes
ACL (1)3
2022 Evaluating locality in NMT models
Itay Itzhak, Koustuv Sinha, Brenden M. Lake, Adina Williams, Dieuwke Hupkes
CogSci5
2022 Can Transformers Process Recursive Nested Constructions, Like Humans?
abstract
Recursive processing is considered a hallmark of human linguistic abilities. A recent study evaluated recursive processing in recurrent neural language models (RNN-LMs) and showed that such models perform below chance level on embedded dependencies within nested constructions – a prototypical example of recursion in natural language. Here, we study if state-of-the-art Transformer LMs do any better. We test eight different Transformer LMs on two different types of nested constructions, which differ in whether the embedded (inner) dependency is short or long range. We find that Transformers achieve near-perfect performance on short-range embedded dependencies, significantly better than previous results reported for RNN-LMs and humans. However, on long-range embedded dependencies, Transformers’ performance sharply drops below chance level. Remarkably, the addition of only three words to the embedded dependency caused Transformers to fall from near-perfect to below-chance performance. Taken together, our results reveal how brittle syntactic processing is in Transformers, compared to humans.
Yair Lakretz, Theo Desbordes, Dieuwke Hupkes, Stanislas Dehaene
COLING3
2021 Generalising to German Plural Noun Classes, from the Perspective of a Recurrent Neural Network
abstract
Inflectional morphology has since long been a useful testing ground for broader questions about generalisation in language and the viability of neural network models as cognitive models of language.Here, in line with that tradition, we explore how recurrent neural networks acquire the complex German plural system and reflect upon how their strategy compares to human generalisation and rule-based models of this system.We perform analyses including behavioural experiments, diagnostic classification, representation analysis and causal interventions, suggesting that the models rely on features that are also key predictors in rule-based models of German plurals.However, the models also display shortcut learning, which is crucial to overcome in search of more cognitively plausible generalisation behaviour.
Verna Dankers, Anna Langedijk, Kate McCurdy, Adina Williams, Dieuwke Hupkes
CoNLL5
2021 Co-evolution of language and agents in referential games
abstract
Referential games offer a grounded learning environment for neural agents which accounts for the fact that language is functionally used to communicate.However, they do not take into account a second constraint considered to be fundamental for the shape of human language: that it must be learnable by new language learners.Cogswell et al. (2019) introduced cultural transmission within referential games through a changing population of agents to constrain the emerging language to be learnable.However, the resulting languages remain inherently biased by the agents' underlying capabilities.In this work, we introduce Language Transmission Simulator to model both cultural and architectural evolution in a population of agents.As our core contribution, we empirically show that the optimal situation is to take into account also the learning biases of the language learners and thus let language and agents coevolve.When we allow the agent population to evolve through architectural evolution, we achieve across the board improvements on all considered metrics and surpass the gains made with cultural transmission.These results stress the importance of studying the underlying agent architecture and pave the way to investigate the co-evolution of language and agent in language emergence studies.
Gautier Dagan, Dieuwke Hupkes, Elia Bruni
EACL2
2021 Language Modelling as a Multi-Task Problem
abstract
In this paper, we propose to study language modelling as a multi-task problem, bringing together three strands of research: multitask learning, linguistics, and interpretability.Based on hypotheses derived from linguistic theory, we investigate whether language models adhere to learning principles of multi-task learning during training.To showcase the idea, we analyse the generalisation behaviour of language models as they learn the linguistic concept of Negative Polarity Items (NPIs).Our experiments demonstrate that a multi-task setting naturally emerges within the objective of the more general task of language modelling.We argue that this insight is valuable for multitask learning, linguistics and interpretability research and can lead to exciting new findings in all three domains.
Lucas Weber, Jaap Jumelet, Elia Bruni, Dieuwke Hupkes
EACL4
2021 Masked Language Modeling and the Distributional Hypothesis: Order Word Matters Pre-training for Little
abstract
A possible explanation for the impressive performance of masked language model (MLM) pre-training is that such models have learned to represent the syntactic structures prevalent in classical NLP pipelines.In this paper, we propose a different explanation: MLMs succeed on downstream tasks mostly due to their ability to model higher-order word cooccurrence statistics.To demonstrate this, we pre-train MLMs on sentences with randomly shuffled word order, and we show that these models still achieve high accuracy after finetuning on many downstream tasks -including tasks specifically designed to be challenging for models that ignore word order.Our models also perform surprisingly well according to some parametric syntactic probes, indicating possible deficiencies in how we test representations for syntactic information.Overall, our results show that purely distributional information largely explains the success of pretraining, and they underscore the importance of curating challenging evaluation datasets that require deeper linguistic knowledge.
Koustuv Sinha, Robin Jia, Dieuwke Hupkes, Joelle Pineau, Adina Williams, Douwe Kiela
EMNLP (1)3
2020 Location Attention for Extrapolation to Longer Sequences
abstract
Neural networks are surprisingly good at interpolating and perform remarkably well when the training set examples resemble those in the test set.However, they are often unable to extrapolate patterns beyond the seen data, even when the abstractions required for such patterns are simple.In this paper, we first review the notion of extrapolation, why it is important, and how one could hope to tackle it.We then focus on a specific type of extrapolation, which is especially useful for natural language processing: generalization to sequences longer than those seen during training.We hypothesize that models with a separate contentand location-based attention are more likely to extrapolate than those with common attention mechanisms.We empirically support our claim for recurrent seq2seq models with our proposed attention on variants of the Lookup Table task.This sheds light on some striking failures of neural models for sequences and on possible methods to approaching such issues.
Yann Dubois, Gautier Dagan, Dieuwke Hupkes, Elia Bruni
ACL3
2020 The Grammar of Emergent Languages
abstract
In this paper, we consider the syntactic properties of languages emerged in referential games, using unsupervised grammar induction (UGI) techniques originally designed to analyse natural language.We show that the considered UGI techniques are appropriate to analyse emergent languages and we then study if the languages that emerge in a typical referential game setup exhibit syntactic structure, and to what extent this depends on the maximum message length and number of symbols that the agents are allowed to use.Our experiments demonstrate that a certain message length and vocabulary size are required for structure to emerge, but they also illustrate that more sophisticated game scenarios are required to obtain syntactic properties more akin to those observed in human language.We argue that UGI techniques should be part of the standard toolkit for analysing emergent languages and release a comprehensive library to facilitate such analysis for future researchers.
Oskar van der Wal, Silvan de Boer, Elia Bruni, Dieuwke Hupkes
EMNLP (1)4
2020 Compositionality Decomposed: How do Neural Networks Generalise? (Extended Abstract)
abstract
Despite a multitude of empirical studies, little consensus exists on whether neural networks are able to generalise compositionally. As a response to this controversy, we present a set of tests that provide a bridge between, on the one hand, the vast amount of linguistic and philosophical theory about compositionality of language and, on the other, the successful neural models of language. We collect different interpretations of compositionality and translate them into five theoretically grounded tests for models that are formulated on a task-independent level. To demonstrate the usefulness of this evaluation paradigm, we instantiate these five tests on a highly compositional data set which we dub PCFG SET, apply the resulting tests to three popular sequence-to-sequence models and provide an in-depth analysis of the results.
Dieuwke Hupkes, Verna Dankers, Mathijs Mul, Elia Bruni
IJCAI1
2020 Compositionality Decomposed: How do Neural Networks Generalise?
abstract
Despite a multitude of empirical studies, little consensus exists on whether neural networks are able to generalise compositionally, a controversy that, in part, stems from a lack of agreement about what it means for a neural model to be compositional. As a response to this controversy, we present a set of tests that provide a bridge between, on the one hand, the vast amount of linguistic and philosophical theory about compositionality of language and, on the other, the successful neural models of language. We collect different interpretations of compositionality and translate them into five theoretically grounded tests for models that are formulated on a task-independent level. In particular, we provide tests to investigate (i) if models systematically recombine known parts and rules (ii) if models can extend their predictions beyond the length they have seen in the training data (iii) if models’ composition operations are local or global (iv) if models’ predictions are robust to synonym substitutions and (v) if models favour rules or exceptions during training. To demonstrate the usefulness of this evaluation paradigm, we instantiate these five tests on a highly compositional data set which we dub PCFG SET and apply the resulting tests to three popular sequence-to-sequence models: a recurrent, a convolution-based and a transformer model. We provide an in-depth analysis of the results, which uncover the strengths and weaknesses of these three architectures and point to potential areas of improvement.
Dieuwke Hupkes, Verna Dankers, Mathijs Mul, Elia Bruni
J. Artif. Intell. Res.1
2019 Analysing Neural Language Models: Contextual Decomposition Reveals Default Reasoning in Number and Gender Assignment
abstract
Extensive research has recently shown that recurrent neural language models are able to process a wide range of grammatical phenomena.How these models are able to perform these remarkable feats so well, however, is still an open question.To gain more insight into what information LSTMs base their decisions on, we propose a generalisation of Contextual Decomposition (GCD).In particular, this setup enables us to accurately distil which part of a prediction stems from semantic heuristics, which part truly emanates from syntactic cues and which part arise from the model biases themselves instead.We investigate this technique on tasks pertaining to syntactic agreement and co-reference resolution and discover that the model strongly relies on a default reasoning effect to perform these tasks.
Jaap Jumelet, Willem H. Zuidema, Dieuwke Hupkes
CoNLL3
2018 Visualisation and 'Diagnostic Classifiers' Reveal how Recurrent and Recursive Neural Networks Process Hierarchical Structure (Extended Abstract)
abstract
In this paper, we investigate how recurrent neural networks can learn and process languages with hierarchical, compositional semantics. To this end, we define the artificial task of processing nested arithmetic expressions, and study whether different types of neural networks can learn to compute their meaning. We find that simple recurrent networks cannot find a generalising solution to this task, but gated recurrent neural networks perform surprisingly well: networks learn to predict the outcome of the arithmetic expressions with high accuracy, although performance deteriorates somewhat with increasing length. We test multiple hypotheses on the information that is encoded and processed by the networks using a method called diagnostic classification. In this method, simple neural classifiers are used to test sequences of predictions about features of the hidden state representations at each time step. Our results indicate that the networks follow a strategy similar to our hypothesised ‘cumulative strategy’, which explains the high accuracy of the network on novel expressions, the generalisation to longer expressions than seen in training, and the mild deterioration with increasing length. This, in turn, shows that diagnostic classifiers can be a useful technique for opening up the black box of neural networks.
Dieuwke Hupkes, Willem H. Zuidema
IJCAI1
2018 Visualisation and 'Diagnostic Classifiers' Reveal How Recurrent and Recursive Neural Networks Process Hierarchical Structure
abstract
We investigate how neural networks can learn and process languages with hierarchical, compositional semantics. To this end, we define the artificial task of processing nested arithmetic expressions, and study whether different types of neural networks can learn to compute their meaning. We find that recursive neural networks can implement a generalising solution to this problem, and we visualise this solution by breaking it up in three steps: project, sum and squash. As a next step, we investigate recurrent neural networks, and show that a gated recurrent unit, that processes its input incrementally, also performs very well on this task: the network learns to predict the outcome of the arithmetic expressions with high accuracy, although performance deteriorates somewhat with increasing length. To develop an understanding of what the recurrent network encodes, visualisation techniques alone do not suffice. Therefore, we develop an approach where we formulate and test multiple hypotheses on the information encoded and processed by the network. For each hypothesis, we derive predictions about features of the hidden state representations at each time step, and train 'diagnostic classifiers' to test those predictions. Our results indicate that the networks follow a strategy similar to our hypothesised 'cumulative strategy', which explains the high accuracy of the network on novel expressions, the generalisation to longer expressions than seen in training, and the mild deterioration with increasing length. This in turn shows that diagnostic classifiers can be a useful technique for opening up the black box of neural networks. We argue that diagnostic classification, unlike most visualisation techniques, does scale up from small networks in a toy domain, to larger and deeper recurrent networks dealing with real-life data, and may therefore contribute to a better understanding of the internal dynamics of current state-of-the-art models in natural language processing.
Dieuwke Hupkes, Sara Veldhoen, Willem H. Zuidema
J. Artif. Intell. Res.1
2016 POS-tagging of Historical Dutch
Dieuwke Hupkes, Rens Bod
LREC1