VLDB 2026 Research / reviewers in the wild / expert
Yevgen Matusevych
dblp:176/0029
· DBLP profile ↗
21ranked-venue papers
9as first author
12since 2021 · last 2026
0000-0003-2910-4025ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 9 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 7 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CAIT: A Syntactic Parsing Toolkit for Child-Adult InTeractionsabstractFrancesca Padovani, Xiulin Yang, Bastian Bunzeck, Jaap Jumelet, Yevgen Matusevych, Nathan Schneider, Arianna Bisazza. Proceedings of the 30th Conference on Computational Natural Language Learning. 2026. Francesca Padovani, Xiulin Yang, Bastian Bunzeck, Jaap Jumelet, Yevgen Matusevych, Nathan Schneider 0001, Arianna Bisazza |
CoNLL | 5 |
| 2025 | Child-Directed Language Does Not Consistently Boost Syntax Learning in Language ModelsabstractSeminal work by Huebner et al. (2021) showed that language models (LMs) trained on English Child-Directed Language (CDL) can reach similar syntactic abilities as LMs trained on much larger amounts of adult-directed written text, suggesting that CDL could provide more effective LM training material than the commonly used internet-crawled data.However, the generalizability of these results across languages, model types, and evaluation settings remains unclear.We test this by comparing models trained on CDL vs. Wikipedia across two LM objectives (masked and causal), three languages (English, French, German), and three syntactic minimal-pair benchmarks.Our results on these benchmarks show inconsistent benefits of CDL, which in most cases is outperformed by Wikipedia models.We then identify various shortcomings in previous benchmarks, and introduce a novel testing methodology, FIT-CLAMS, which uses a frequency-controlled design to enable balanced comparisons across training corpora.Through minimal pair evaluations and regression analysis we show that training on CDL does not yield stronger generalizations for acquiring syntax and highlight the importance of controlling for frequency effects when evaluating syntactic ability. 1 Francesca Padovani, Jaap Jumelet, Yevgen Matusevych, Arianna Bisazza |
EMNLP | 3 |
| 2025 | The mutual exclusivity bias of bilingual visually grounded speech modelsabstractMutual exclusivity (ME) is a strategy where a novel word is associated with a novel object rather than a familiar one, facilitating language learning in children. Recent work has found an ME bias in a visually grounded speech (VGS) model trained on English speech with paired images. But ME has also been studied in bilingual children, who may employ it less due to cross-lingual ambiguity. We explore this pattern computationally using bilingual VGS models trained on combinations of English, French, and Dutch. We find that bilingual models generally exhibit a weaker ME bias than monolingual models, though exceptions exist. Analyses show that the combined visual embeddings of bilingual models have a smaller variance for familiar data, partly explaining the increase in confusion between novel and familiar concepts. We also provide new insights into why the ME bias exists in VGS models in the first place. Code and data: https://github.com/danoneata/me-vgs. Dan Oneata, Leanne Nortje, Yevgen Matusevych, Herman Kamper |
INTERSPEECH | 3 |
| 2024 | Cognitive Plausibility in Natural Language ProcessingabstractRecent successes in Natural Language Processing (NLP) give rise to more and more computational models aimed at generating and understanding human language. Traditionally, researchers evaluate such models by looking at how accurate they are at a given task, but the focus in the field is slowly shifting to evaluating other characteristics, such as models’ fairness, interpretability, efficiency, and so forth. An often overlooked aspect is a model’s cognitive plausibility: that is, how “human-like” a model is. Cognitive plausibility is a multifaceted concept grounded in the field of cognitive science, which often uses computational models to study various aspects of human cognition, including language production and comprehension. The present book is a contribution to narrowing the wide gap between state-of-the-art NLP architectures and human cognition.In the quickly changing world of NLP, it is not easy to summarize recent advances, nonetheless, this book’s comprehensive selection of examples from various studies offers a solid overview of the current research on models’ cognitively plausibility. The examples are combined with discussions of theoretical and methodological issues at the interface of NLP and cognitive science. Moreover, each of the main content chapters ends with an overview of relevant ethical issues, a necessary consideration in the era of powerful language models. The book will be compelling reading to NLP researchers interested in human cognition and model interpretability, as well as to cognitive scientists and psycholinguists willing to better understand computational modeling approaches in the language domain. Gradual presentation of the material, with two introductory chapters (see below), also makes the book generally accessible to students and those with only basic knowledge of NLP.The book consists of seven chapters. Chapter 1, “Introduction,” explains why cognitive plausibility is important to consider in NLP: In addition to obvious advantages of cognitively plausible models for cognitive science, the behavior of such models is more intuitive to interpret for human speakers. Exploring this link between cognitive plausibility and interpretability is one of the book’s goals, and it is repeatedly—and very successfully—exploited throughout all the chapters. Another important message in the introduction is that cognitive plausibility is a graded concept that involves multiple dimensions, and this book focuses on three of them: the similarity between the decisions made by models and humans, between the representational structures they use, and between their procedural strategies.The five chapters with the main content can be divided into two large sections: the more introductory ones (2–3) and more in-depth ones (4–6). Specifically, Chapter 2, “Foundations of Language Modeling,” provides an overview of basic concepts in language modeling, such as conditional probability of a sentence, model perplexity, recurrent neural networks, pretrained language models, and so on. Particular emphasis is made on modeling choices that link language models to human language processing: For example, a model’s objective function and architecture can determine whether the model processes words sequentially or not, while input units that models are trained on may or may not reflect the underlying linguistic structure. Chapter 3, “Cognitive Signals of Language Processing,” introduces the types of data that can be used for evaluating cognitive plausibility of computational models. Here, the reader is made aware that the focus of the book is on language comprehension rather than production, which some may find slightly disappointing given that many state-of-the-art generative language models are celebrated largely for their ability to produce human-like language. On the brighter side, even in comprehension there is plenty of relevant data sets, which capture both human speakers’ behavioral responses and their brain activity patterns. Moreover, one can combine different types of data for an even more comprehensive model evaluation. Importantly, many data sets collected from human speakers need to be preprocessed or otherwise adapted to NLP settings, and this chapter also offers an overview of common techniques in this area.The following three chapters are more in-depth and discuss the three dimensions of cognitive plausibility mentioned above. Chapter 4, “Behavioral Patterns,” largely focuses on model interpretability methods through the cognitive lens. First of all, one can analyze a model’s input and output: properties of the input data, model’s behavior on well-defined subpopulations of data and its performance on specific instances depending on their difficulty. Second, there is a variety of tests or even out-of-the-box test suites for targeted evaluation of NLP models. Many of them focus on models’ linguistic abilities, while others (e.g., occlusion or perturbation tests) are designed to stress-test models’ robustness on intricately designed examples. Here, the authors call for using finer-grained model evaluation and for developing multilingual models that are not optimized for English data, as is often the case in NLP. Another promising approach to narrowing the gap between models’ and human speakers’ behavior, according to the authors, consists in designing cognitively plausible curricula for model training, but sadly, in the 2023 BabyLM challenge curriculum learning methods only resulted in modest improvements (Warstadt et al. 2023), and the jury is still out on this subject.While the previous chapter focuses on models’ inputs and outputs, Chapter 5, “Representational Structure,” discusses interpretability of models’ internal representations. Neural representations are commonly expressed as vectors in a high-dimensional space, and measuring similarity between them is central to research in this area. From the cognitive perspective, however, similarity—even between words, let alone longer units—is a complicated concept: units can be judged similar for a variety of reasons, which also highly depend on the context. Moving beyond a single representational space, one can also measure the similarity of different spaces, including the ones derived from human speakers’ behavioral or brain responses. In many cases, it is crucial to have a concrete hypothesis about a model’s representations and test it in a targeted way—for example, using probing classifiers, a common method for finding out whether a specific feature is encoded in a model’s representation space. Another fruitful research direction is mapping models’ representations to brain responses: For example, can a probing classifier learn to predict brain activation patterns?The final dimension of cognitive plausibility is discussed in Chapter 6, “Procedural Strategies.” Strategies that a model adopts are determined by its architecture. For the time being, transformers are by far the most common architecture in NLP, and a large part of this chapter discusses the mechanisms of attention and self-attention used in transformers. Somewhat sadly from this book’s perspective, multi-head attention can hardly be considered cognitively plausible, but nevertheless, studying attention weights can still help us understand relative importance of input units (e.g., words), and the model’s importance values can be compared to human data, such as gaze patterns during reading. Overall, this chapter mostly focuses on sentence processing tasks, since they provide a fruitful ground for studying various procedural strategies and effects: incremental processing (cf. Chapter 2), priming, hierarchical processing, and so on. Concrete proposals to improve the cognitive plausibility of models’ algorithms include the use of multi-task and transfer learning setups, through the explicit integration of human problem-solving strategies into models’ training process.The final Chapter 7, “Towards Cognitively More Plausible Models,” is a brief recap of the book. Although the authors conclude they could not find a silver bullet to make a model cognitively plausible on all the three dimensions considered, they nevertheless successfully propose concrete methods for designing more cognitively plausible models. Among others, these proposals include taking into account instance difficulty in model evaluation, developing more context-aware tools for representation analysis, integrating information from multiple modalities into the models, and adopting a truly multilingual perspective on model design.One key takeaway from this book is that cognitive plausibility is a complex concept—not only because there are several dimensions to it, but also because data collected from human speakers is often less straightforward than what’s dictated by existing NLP models’ objective functions. The authors provide plenty of examples of this complexity throughout the book: Human speakers can disagree in their linguistic annotations, their conceptual representations tend to be fluid, their similarity judgments are nuanced and graded. One possible way forward for NLP is to embrace this uncertainty of human language behavior, an avenue that the field is only starting to explore (e.g., Baan et al. 2023; Liu et al. 2023). Yevgen Matusevych |
Comput. Linguistics | 1 |
| 2024 | Visually Grounded Speech Models Have a Mutual Exclusivity BiasabstractAbstract When children learn new words, they employ constraints such as the mutual exclusivity (ME) bias: A novel word is mapped to a novel object rather than a familiar one. This bias has been studied computationally, but only in models that use discrete word representations as input, ignoring the high variability of spoken words. We investigate the ME bias in the context of visually grounded speech models that learn from natural images and continuous speech audio. Concretely, we train a model on familiar words and test its ME bias by asking it to select between a novel and a familiar object when queried with a novel word. To simulate prior acoustic and visual knowledge, we experiment with several initialization strategies using pretrained speech and vision networks. Our findings reveal the ME bias across the different initialization approaches, with a stronger bias in models with more prior (in particular, visual) knowledge. Additional tests confirm the robustness of our results, even when different loss functions are considered. Based on detailed analyses to piece out the model’s representation space, we attribute the ME bias to how familiar and novel classes are distinctly separated in the resulting space. Leanne Nortje, Dan Oneata, Yevgen Matusevych, Herman Kamper |
Trans. Assoc. Comput. Linguistics | 3 |
| 2022 | Bilingual Sentence Processing: when Models Meet Experiments
Stefan L. Frank, Xavier Hinaut, Edith Kaan, Yung Han Khoe, Irene Elisabeth Winther, Yevgen Matusevych |
CogSci | 7 |
| 2022 | Trees neural those: RNNs can learn the hierarchical structure of noun phrases
Yevgen Matusevych, Jennifer Culbertson |
CogSci | 1 |
| 2022 | Modeling Sentence Processing Effects in Bilingual Speakers: A Comparison of Neural Architectures
Rasmus Roslund, Yevgen Matusevych |
CogSci | 2 |
| 2021 | Cumulative frequency can explain cognate facilitation in language models
Irene Elisabeth Winther, Yevgen Matusevych, Martin J. Pickering |
CogSci | 2 |
| 2021 | A phonetic model of non-native spoken word processingabstractYevgen Matusevych, Herman Kamper, Thomas Schatz, Naomi Feldman, Sharon Goldwater. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021. Yevgen Matusevych, Herman Kamper, Thomas Schatz, Naomi Feldman, Sharon Goldwater |
EACL | 1 |
| 2021 | Acoustic Word Embeddings for Zero-Resource Languages Using Self-Supervised Contrastive Learning and Multilingual AdaptationabstractAcoustic word embeddings (AWEs) are fixed-dimensional representations of variable-length speech segments. For zero-resource languages where labelled data is not available, one AWE approach is to use unsupervised autoencoder-based re-current models. Another recent approach is to use multilingual transfer: a supervised AWE model is trained on several well-resourced languages and then applied to an unseen zero-resource language. We consider how a recent contrastive learning loss can be used in both the purely unsupervised and multilingual transfer settings. Firstly, we show that terms from an unsupervised term discovery system can be used for contrastive self-supervision, resulting in improvements over previous unsupervised monolingual AWE models. Secondly, we consider how multilingual AWE models can be adapted to a specific zero-resource language using discovered terms. We find that self-supervised contrastive adaptation outperforms adapted multilingual correspondence autoencoder and Siamese AWE models, giving the best overall results in a word discrimination task on six zero-resource languages. Christiaan Jacobs, Yevgen Matusevych, Herman Kamper |
SLT | 2 |
| 2021 | Improved Acoustic Word Embeddings for Zero-Resource Languages Using Multilingual TransferabstractAcoustic word embeddings are fixed-dimensional representations of variable-length speech segments. Such embeddings can form the basis for speech search, indexing and discovery systems when conventional speech recognition is not possible. In zero-resource settings where unlabelled speech is the only available resource, we need a method that gives robust embeddings on an arbitrary language. Here we explore multilingual transfer: we train a single supervised embedding model on labelled data from multiple well-resourced languages and then apply it to unseen zero-resource languages. We consider three multilingual recurrent neural network (RNN) models: a classifier trained on the joint vocabularies of all training languages; a Siamese RNN trained to discriminate between same and different words from multiple languages; and a correspondence autoencoder (CAE) RNN trained to reconstruct word pairs. In a word discrimination task on six target languages, all of these models outperform state-of-the-art unsupervised models trained on the zero-resource languages themselves, giving relative improvements of more than 30% in average precision. When using only a few training languages, the multilingual CAE-RNN performs better, but with more training languages the other multilingual models perform similarly. Using more training languages is generally beneficial, but improvements are marginal on some languages. We present probing experiments which show that the CAE-RNN encodes more phonetic, word duration, language identity and speaker information than the other multilingual models. Herman Kamper, Yevgen Matusevych, Sharon Goldwater |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2020 | Input matters in the modeling of early phonetic learning
Ruolan Li, Thomas Schatz, Yevgen Matusevych, Sharon Goldwater, Naomi Feldman |
CogSci | 3 |
| 2020 | Evaluating computational models of infant phonetic learning across languages
Yevgen Matusevych, Thomas Schatz, Herman Kamper, Naomi Feldman, Sharon Goldwater |
CogSci | 1 |
| 2020 | Multilingual Acoustic Word Embedding Models for Processing Zero-resource LanguagesabstractAcoustic word embeddings are fixed-dimensional representations of variable-length speech segments. In settings where unlabelled speech is the only available resource, such embeddings can be used in "zero-resource" speech search, indexing and discovery systems. Here we propose to train a single supervised embedding model on labelled data from multiple well-resourced languages and then apply it to unseen zero-resource languages. For this transfer learning approach, we consider two multilingual recurrent neural network models: a discriminative classifier trained on the joint vocabularies of all training languages, and a correspondence autoencoder trained to reconstruct word pairs. We test these using a word discrimination task on six target zero-resource languages. When trained on seven well-resourced languages, both models perform similarly and outperform unsupervised models trained on the zero-resource languages. With just a single training language, the second model works better, but performance depends more on the particular training-testing language pair. Herman Kamper, Yevgen Matusevych, Sharon Goldwater |
ICASSP | 2 |
| 2019 | Are we there yet? Encoder-decoder neural networks as cognitive models of English past tense inflectionabstractThe cognitive mechanisms needed to account for the English past tense have long been a subject of debate in linguistics and cognitive science.Neural network models were proposed early on, but were shown to have clear flaws.Recently, however, Kirov and Cotterell (2018) showed that modern encoder-decoder (ED) models overcome many of these flaws.They also presented evidence that ED models demonstrate humanlike performance in a nonce-word task.Here, we look more closely at the behaviour of their model in this task.We find that (1) the model exhibits instability across multiple simulations in terms of its correlation with human data, and (2) even when results are aggregated across simulations (treating each simulation as an individual human participant), the fit to the human data is not strong-worse than an older rule-based model.These findings hold up through several alternative training regimes and evaluation measures.Although other neural architectures might do better, we conclude that there is still insufficient evidence to claim that neural nets are a good cognitive model for this task. Maria Corkery, Yevgen Matusevych, Sharon Goldwater |
ACL (1) | 2 |
| 2018 | Crosslinguistic transfer as category adjustment: Modeling conceptual color shift in bilingualism
Yevgen Matusevych, Barend Beekhuizen, Suzanne Stevenson |
CogSci | 1 |
| 2018 | Analyzing and modeling free word associations
Yevgen Matusevych, Suzanne Stevenson |
CogSci | 1 |
| 2015 | Distributional determinants of learning argument structure constructions in first and second language
Yevgen Matusevych, Afra Alishahi, Ad Backus |
CogSci | 1 |
| 2014 | Isolating second language learning factors in a computational study of bilingual construction acquisition
Yevgen Matusevych, Afra Alishahi, Ad Backus |
CogSci | 1 |
| 2013 | Automatic generation of naturalistic child-adult interaction data
Yevgen Matusevych, Afra Alishahi, Paul Vogt |
CogSci | 1 |