Leonie Weissweiler

dblp:212/0281 · DBLP profile ↗
← Back
11ranked-venue papers
4as first author
11since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 4 first-author · 11 since 2021
YearPublicationVenuePosition
2026 Language Models Learn Constructional Semantics, Not To Mention Syntax: Investigating LM Understanding of Paired-Focus Constructions
abstract
Grasping the semantics of rare constructions (form-meaning pairings) has been shown to be a challenging problem that has currently only been solved by the largest LLMs.It remains an open question if open-source models have robust constructional understanding, and if so, what learning dynamics underlie the acquisition of this knowledge.Focusing on a set of rare PAIRED-FOCUS constructions in English (e.g."let alone", "much less"), we construct a novel dataset to test their meanings using both scalar adjectival semantics and general world knowledge.Testing a wide range of models differing in parameter count, architecture, and pretraining dataset size, we find that several modestly sized models are sensitive to both the forms and the meanings of PAIRED-FOCUS constructions, though models trained on human-scale data fail at all meaning evaluations.Turning to training dynamics for a set of open-checkpoint models, we find that PAIRED-FOCUS understanding emerges later in training than PAIRED-FOCUS syntactic knowledge, and that learning of PAIRED-FOCUS semantics is correlated with gains in some domains of world knowledge.Overall, our empirical results support the conclusion that modestly sized open-source models can grasp the rare PAIRED-FOCUS constructions, and demonstrate a connection between knowledge of PAIRED-FOCUS constructions and other meaning domains.
Wesley Scivetti, Ethan Wilcox, Nathan Schneider 0001, Kanishka Misra, Leonie Weissweiler
CoNLL5
2026 MultiBLiMP 1.0: A Massively Multilingual Benchmark of Linguistic Minimal Pairs
abstract
Abstract We introduce MultiBLiMP 1.0, a massively multilingual benchmark of linguistic minimal pairs, covering 101 languages and 2 types of subject-verb agreement, containing more than 128,000 minimal pairs. Our minimal pairs are created using a fully automated pipeline, leveraging the large-scale linguistic resources of Universal Dependencies and UniMorph. MultiBLiMP 1.0 evaluates abilities of LLMs at an unprecedented multilingual scale, and highlights the shortcomings of the current state-of-the-art in modelling low-resource languages.1
Jaap Jumelet, Leonie Weissweiler, Joakim Nivre, Arianna Bisazza
Trans. Assoc. Comput. Linguistics2
2025 Constructions are Revealed in Word Distributions
abstract
Construction grammar posits that constructions, or form-meaning pairings, are acquired through experience with language (the distributional learning hypothesis).But how much information about constructions does this distribution actually contain?Corpus-based analyses provide some answers, but text alone cannot answer counterfactual questions about what caused a particular word to occur.This requires computable models of the distribution over strings-namely, pretrained language models (PLMs).Here, we treat a RoBERTa model as a proxy for this distribution and hypothesize that constructions will be revealed within it as patterns of statistical affinity.We support this hypothesis experimentally: many constructions are robustly distinguished, including (i) hard cases where semantically distinct constructions are superficially similar, as well as (ii) schematic constructions, whose "slots" can be filled by abstract word classes.Despite this success, we also provide qualitative evidence that statistical affinity alone may be insufficient to identify all constructions from text.Thus, statistical affinity is likely an important, but partial, signal available to learners. 1
Josh Rozner, Leonie Weissweiler, Kyle Mahowald, Cory Shain
EMNLP2
2025 BabyLM's First Constructions: Causal interventions provide a signal of learning
abstract
Construction grammar posits that language learners acquire constructions (form-meaning pairings) from the statistics of their environment.Recent work supports this hypothesis by showing sensitivity to constructions in pretrained language models (PLMs), including one recent study (Rozner et al., 2025) demonstrating that constructions shape RoBERTa's output distribution.However, models under study have generally been trained on developmentally implausible amounts of data, casting doubt on their relevance to human language learning.Here we use Rozner et al.'s methods to evaluate construction learning in masked language models from the 2024 BabyLM Challenge.Our results show that even when trained on developmentally plausible quantities of data, models learn diverse constructions, even hard cases that are superficially indistinguishable.We further find correlational evidence that constructional performance may be functionally relevant: models that better represent constructions perform better on the BabyLM benchmarks. 1
Josh Rozner, Leonie Weissweiler, Cory Shain
EMNLP2
2024 Verbing Weirds Language (Models): Evaluation of English Zero-Derivation in Five LLMs
abstract
Lexical-syntactic flexibility, in the form of conversion (or zero-derivation) is a hallmark of English morphology. In conversion, a word with one part of speech is placed in a non-prototypical context, where it is coerced to behave as if it had a different part of speech. However, while this process affects a large part of the English lexicon, little work has been done to establish the degree to which language models capture this type of generalization. This paper reports the first study on the behavior of large language models with reference to conversion. We design a task for testing lexical-syntactic flexibility—the degree to which models can generalize over words in a construction with a non-prototypical part of speech. This task is situated within a natural language inference paradigm. We test the abilities of five language models—two proprietary models (GPT-3.5 and GPT-4), three open source model (Mistral 7B, Falcon 40B, and Llama 2 70B). We find that GPT-4 performs best on the task, followed by GPT-3.5, but that the open source language models are also able to perform it and that the 7-billion parameter Mistral displays as little difference between its baseline performance on the natural language inference task and the non-prototypical syntactic category task, as the massive GPT-4.
David R. Mortensen, Valentina Izrailevitch, Yunze Xiao, Hinrich Schütze, Leonie Weissweiler
LREC/COLING5
2024 UCxn: Typologically-Informed Annotation of Constructions Atop Universal Dependencies
abstract
The Universal Dependencies (UD) project has created an invaluable collection of treebanks with contributions in over 140 languages. However, the UD annotations do not tell the full story. Grammatical constructions that convey meaning through a particular combination of several morphosyntactic elements—for example, interrogative sentences with special markers and/or word orders—are not labeled holistically. We argue for (i) augmenting UD annotations with a ‘UCxn’ annotation layer for such meaning-bearing grammatical constructions, and (ii) approaching this in a typologically informed way so that morphosyntactic strategies can be compared across languages. As a case study, we consider five construction families in ten languages, identifying instances of each construction in UD treebanks through the use of morphosyntactic patterns. In addition to findings regarding these particular constructions, our study yields important insights on methodology for describing and identifying constructions in language-general and language-particular ways, and lays the foundation for future constructional enrichment of UD treebanks.
Leonie Weissweiler, Nina Böbel, Kirian Guiller, Santiago Herrera, Wesley Scivetti, Arthur Lorenzi Almeida, Nurit Melnik, Archna Bhatia, Hinrich Schütze, Lori S. Levin, Amir Zeldes, Joakim Nivre, William Croft 0001, Nathan Schneider 0001
LREC/COLING1
2024 Constructions Are So Difficult That Even Large Language Models Get Them Right for the Wrong Reasons
abstract
In this paper, we make a contribution that can be understood from two perspectives: from an NLP perspective, we introduce a small challenge dataset for NLI with large lexical overlap, which minimises the possibility of models discerning entailment solely based on token distinctions, and show that GPT-4 and Llama 2 fail it with strong bias. We then create further challenging sub-tasks in an effort to explain this failure. From a Computational Linguistics perspective, we identify a group of constructions with three classes of adjectives which cannot be distinguished by surface features. This enables us to probe for LLM’s understanding of these constructions in various ways, and we find that they fail in a variety of ways to distinguish between them, suggesting that they don’t adequately represent their meaning or capture the lexical properties of phrasal heads.
Shijia Zhou, Leonie Weissweiler, Taiqi He, Hinrich Schütze, David R. Mortensen, Lori S. Levin
LREC/COLING2
2023 A Crosslingual Investigation of Conceptualization in 1335 Languages
abstract
Yihong Liu, Haotian Ye, Leonie Weissweiler, Philipp Wicke, Renhao Pei, Robert Zangenfeind, Hinrich Schütze. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Yihong Liu 0001, Haotian Ye, Leonie Weissweiler, Philipp Wicke, Renhao Pei, Robert Zangenfeind, Hinrich Schütze
ACL (1)3
2023 Counting the Bugs in ChatGPT's Wugs: A Multilingual Investigation into the Morphological Capabilities of a Large Language Model
abstract
Leonie Weissweiler, Valentin Hofmann, Anjali Kantharuban, Anna Cai, Ritam Dutt, Amey Hengle, Anubha Kabra, Atharva Kulkarni, Abhishek Vijayakumar, Haofei Yu, Hinrich Schuetze, Kemal Oflazer, David Mortensen. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023.
Leonie Weissweiler, Valentin Hofmann, Anjali Kantharuban, Anna Cai, Ritam Dutt, Amey Hengle, Anubha Kabra, Atharva Kulkarni, Abhishek Vijayakumar, Haofei Yu, Hinrich Schütze, Kemal Oflazer, David R. Mortensen
EMNLP1
2022 CaMEL: Case Marker Extraction without Labels
abstract
We introduce CaMEL (Case Marker Extraction without Labels), a novel and challenging task in computational morphology that is especially relevant for low-resource languages.We propose a first model for CaMEL that uses a massively multilingual corpus to extract case markers in 83 languages based only on a noun phrase chunker and an alignment system.To evaluate CaMEL, we automatically construct a silver standard from UniMorph.The case markers extracted by our model can be used to detect and visualise similarities and differences between the case systems of different languages as well as to annotate fine-grained deep cases in languages in which they are not overtly marked.
Leonie Weissweiler, Valentin Hofmann, Masoud Jalili Sabet, Hinrich Schütze
ACL (1)1
2022 The better your Syntax, the better your Semantics? Probing Pretrained Language Models for the English Comparative Correlative
abstract
Construction Grammar (CxG) is a paradigm from cognitive linguistics emphasising the connection between syntax and semantics.Rather than rules that operate on lexical items, it posits constructions as the central building blocks of language, i.e., linguistic units of different granularity that combine syntax and semantics.As a first step towards assessing the compatibility of CxG with the syntactic and semantic knowledge demonstrated by state-ofthe-art pretrained language models (PLMs), we present an investigation of their capability to classify and understand one of the most commonly studied constructions, the English comparative correlative (CC).We conduct experiments examining the classification accuracy of a syntactic probe on the one hand and the models' behaviour in a semantic application task on the other, with BERT, RoBERTa, and DeBERTa as the example PLMs.Our results show that all three investigated PLMs are able to recognise the structure of the CC but fail to use its meaning.While human-like performance of PLMs on many NLP tasks has been alleged, this indicates that PLMs still suffer from substantial shortcomings in central domains of linguistic knowledge.
Leonie Weissweiler, Valentin Hofmann, Abdullatif Köksal, Hinrich Schütze
EMNLP1