VLDB 2026 Research / reviewers in the wild / expert
Dominik Schlechtweg
dblp:184/8971
· DBLP profile ↗
17ranked-venue papers
5as first author
10since 2021 · last 2026
0000-0002-0685-2576ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 5 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | APODICTUS: Automatic Processing of DICTionary Update candidateS
Felix Blessing, Johannes S. Sax, Julian Kaufmann, Wei Zhao 0033, Nikolay Arefyev, Dominik Schlechtweg |
LREC | 6 |
| 2026 | Insights from Transfer Learning Experiments with Word-in-Context and Word Sense Disambiguation Models
Alp Mujko, Dominik Schlechtweg |
LREC | 2 |
| 2025 | Can Large Language Models Compete with Specialized Models in Lexical Semantic Change Detection?abstractIn this paper, we present a comprehensive comparison between specialized Lexical Semantic Change Detection (LSCD) models and Large Language Models (LLMs) for the LSCD task. In addition to comparing models, we also investigate the role of automatic prompt selection for improving LLM performance. We evaluate three approaches: Average Pairwise Distance (APD), Word-in-Context (WiC), and Word Sense Induction (WSI). Using Spearman correlation as the evaluation metric, we assess the performance of Mixtral, Llama 3.1, Llama 3.3, and specialized LSCD models across English and Spanish datasets. Our results show that by using prompt optimization and LLMs, we achieve state-of-the-art performance for the English dataset and outperform specialized LSCD models at the annotation level in the same dataset. For Spanish, specialized models outperform LLMs across all three approaches—WiC, APD, and WSI—indicating that specialized LSCD models are still more effective for semantic change detection in Spanish. Frank D. Zamora-Reina, Felipe Bravo-Marquez, Dominik Schlechtweg, Nikolay Arefyev |
ECAI | 3 |
| 2024 | Enriching Word Usage Graphs with Cluster DefinitionsabstractWe present a dataset of word usage graphs (WUGs), where the existing WUGs for multiple languages are enriched with cluster labels functioning as sense definitions. They are generated from scratch by fine-tuned encoder-decoder language models. The conducted human evaluation has shown that these definitions match the existing clusters in WUGs better than the definitions chosen from WordNet by two baseline systems. At the same time, the method is straightforward to use and easy to extend to new languages. The resulting enriched datasets can be extremely helpful for moving on to explainable semantic change modeling. Andrey Kutuzov, Mariia Fedorova, Dominik Schlechtweg, Nikolay Arefyev |
LREC/COLING | 3 |
| 2024 | TRoTR: A Framework for Evaluating the Re-contextualization of Text ReuseabstractCurrent approaches for detecting text reuse do not focus on recontextualization, i.e., how the new context(s) of a reused text differs from its original context(s).In this paper, we propose a novel framework called TRoTR that relies on the notion of topic relatedness for evaluating the diachronic change of context in which text is reused.TRoTR includes two NLP tasks: TRiC and TRaC.TRiC is designed to evaluate the topic relatedness between a pair of recontextualizations. TRaC is designed to evaluate the overall topic variation within a set of recontextualizations.We also provide a curated TRoTR benchmark of biblical text reuse, human-annotated with topic relatedness.The benchmark exhibits an inter-annotator agreement of .811.We evaluate multiple, established SBERT models on the TRoTR tasks and find that they exhibit greater sensitivity to textual similarity than topic relatedness.Our experiments show that fine-tuning these models can mitigate such a kind of sensitivity. Francesco Periti, Pierluigi Cassotti, Stefano Montanelli, Nina Tahmasebi, Dominik Schlechtweg |
EMNLP | 5 |
| 2024 | More DWUGs: Extending and Evaluating Word Usage Graph Datasets in Multiple LanguagesabstractDominik Schlechtweg, Pierluigi Cassotti, Bill Noble, David Alfter, Sabine Schulte Im Walde, Nina Tahmasebi. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Dominik Schlechtweg, Pierluigi Cassotti, Bill Noble, David Alfter, Sabine Schulte im Walde, Nina Tahmasebi |
EMNLP | 1 |
| 2022 | DiaWUG: A Dataset for Diatopic Lexical Semantic Variation in SpanishabstractWe provide a novel dataset – DiaWUG – with judgements on diatopic lexical semantic variation for six Spanish variants in Europe and Latin America. In contrast to most previous meaning-based resources and studies on semantic diatopic variation, we collect annotations on semantic relatedness for Spanish target words in their contexts from both a semasiological perspective (i.e., exploring the meanings of a word given its form, thus including polysemy) and an onomasiological perspective (i.e., exploring identical meanings of words with different forms, thus including synonymy). In addition, our novel dataset exploits and extends the existing framework DURel for annotating word senses in context (Erk et al., 2013; Schlechtweg et al., 2018) and the framework-embedded Word Usage Graphs (WUGs) – which up to now have mainly be used for semasiological tasks and resources – in order to distinguish, visualize and interpret lexical semantic variation of contextualized words in Spanish from these two perspectives, i.e., semasiological and onomasiological language variation. Gioia Baldissin, Dominik Schlechtweg, Sabine Schulte im Walde |
LREC | 2 |
| 2021 | Lexical Semantic Change DiscoveryabstractSinan Kurtyigit, Maike Park, Dominik Schlechtweg, Jonas Kuhn, Sabine Schulte im Walde. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Sinan Kurtyigit, Maike Park, Dominik Schlechtweg, Jonas Kuhn, Sabine Schulte im Walde |
ACL/IJCNLP (1) | 3 |
| 2021 | Effects of Pre- and Post-Processing on type-based Embeddings in Lexical Semantic Change DetectionabstractLexical semantic change detection is a new and innovative research field.The optimal fine-tuning of models including pre-and postprocessing is largely unclear.We optimize existing models by (i) pre-training on large corpora and refining on diachronic target corpora tackling the notorious small data problem, and (ii) applying post-processing transformations that have been shown to improve performance on synchronic tasks.Our results provide a guide for the application and optimization of lexical semantic change detection models across various learning scenarios. Jens Kaiser, Sinan Kurtyigit, Serge Kotchourko, Dominik Schlechtweg |
EACL | 4 |
| 2021 | DWUG: A large Resource of Diachronic Word Usage Graphs in Four LanguagesabstractWord meaning is notoriously difficult to capture, both synchronically and diachronically.In this paper, we describe the creation of the largest resource of graded contextualized, diachronic word meaning annotation in four different languages, based on 100,000 human semantic proximity judgments.We describe in detail the multi-round incremental annotation process, the choice for a clustering algorithm to group usages into senses, and possible -diachronic and synchronic -uses for this dataset. Dominik Schlechtweg, Nina Tahmasebi, Simon Hengchen, Haim Dubossarsky, Barbara McGillivray |
EMNLP (1) | 1 |
| 2020 | Predicting Degrees of Technicality in Automatic Terminology ExtractionabstractWhile automatic term extraction is a wellresearched area, computational approaches to distinguish between degrees of technicality are still understudied.We semi-automatically create a German gold standard of technicality across four domains, and illustrate the impact of a web-crawled general-language corpus on predicting technicality.When defining a classification approach that combines general-language and domain-specific word embeddings, we go beyond previous work and align vector spaces to gain comparative embeddings.We suggest two novel models to exploit general-vs.domain-specific comparisons: a simple neural network model with pre-computed comparative-embedding information as input, and a multi-channel model computing the comparison internally.Both models outperform previous approaches, with the multi-channel model performing best. Anna Hätty, Dominik Schlechtweg, Michael Dorna, Sabine Schulte im Walde |
ACL | 2 |
| 2020 | CCOHA: Clean Corpus of Historical American EnglishabstractModelling language change is an increasingly important area of interest within the fields of sociolinguistics and historical linguistics. In recent years, there has been a growing number of publications whose main concern is studying changes that have occurred within the past centuries. The Corpus of Historical American English (COHA) is one of the most commonly used large corpora in diachronic studies in English. This paper describes methods applied to the downloadable version of the COHA corpus in order to overcome its main limitations, such as inconsistent lemmas and malformed tokens, without compromising its qualitative and distributional properties. The resulting corpus CCOHA contains a larger number of cleaned word tokens which can offer better insights into language change and allow for a larger variety of tasks to be performed. Reem Alatrash, Dominik Schlechtweg, Jonas Kuhn, Sabine Schulte im Walde |
LREC | 2 |
| 2019 | Time-Out: Temporal Referencing for Robust Modeling of Lexical Semantic ChangeabstractState-of-the-art models of lexical semantic change detection suffer from noise stemming from vector space alignment.We have empirically tested the Temporal Referencing method for lexical semantic change and show that, by avoiding alignment, it is less affected by this noise.We show that, trained on a diachronic corpus, the skip-gram with negative sampling architecture with temporal referencing outperforms alignment models on a synthetic task as well as a manual testset.We introduce a principled way to simulate lexical semantic change and systematically control for possible biases. Haim Dubossarsky, Simon Hengchen, Nina Tahmasebi, Dominik Schlechtweg |
ACL (1) | 4 |
| 2019 | A Wind of Change: Detecting and Evaluating Lexical Semantic Change across Times and DomainsabstractWe perform an interdisciplinary large-scale evaluation for detecting lexical semantic divergences in a diachronic and in a synchronic task: semantic sense changes across time, and semantic sense changes across domains.Our work addresses the superficialness and lack of comparison in assessing models of diachronic lexical change, by bringing together and extending benchmark models on a common state-of-the-art evaluation task.In addition, we demonstrate that the same evaluation task and modelling approaches can successfully be utilised for the synchronic detection of domain-specific sense divergences in the field of term extraction. Dominik Schlechtweg, Anna Hätty, Marco Del Tredici, Sabine Schulte im Walde |
ACL (1) | 1 |
| 2017 | German in Flux: Detecting Metaphoric Change via Word EntropyabstractThis paper explores the informationtheoretic measure entropy to detect metaphoric change, transferring ideas from hypernym detection to research on language change.We also build the first diachronic test set for German as a standard for metaphoric change annotation.Our model shows high performance, is unsupervised, language-independent and generalizable to other processes of semantic change. Dominik Schlechtweg, Stefanie Eckmann, Enrico Santus, Sabine Schulte im Walde, Daniel Hole |
CoNLL | 1 |
| 2017 | Hypernyms under Siege: Linguistically-motivated Artillery for Hypernymy DetectionabstractThe fundamental role of hypernymy in NLP has motivated the development of many methods for the automatic identification of this relation, most of which rely on word distribution.We investigate an extensive number of such unsupervised measures, using several distributional semantic models that differ by context type and feature weighting.We analyze the performance of the different methods based on their linguistic motivation.Comparison to the state-of-the-art supervised methods shows that while supervised methods generally outperform the unsupervised ones, the former are sensitive to the distribution of training instances, hurting their reliability.Being based on general linguistic hypotheses and independent from training data, unsupervised measures are more robust, and therefore are still useful artillery for hypernymy detection. Vered Shwartz, Enrico Santus, Dominik Schlechtweg |
EACL (1) | 3 |
| 2016 | Exploitation of Co-reference in Distributional Semantics
Dominik Schlechtweg |
LREC | 1 |