VLDB 2026 Research / reviewers in the wild / expert
Bolette S. Pedersen
dblp:15/1759 · also Bolette Sandford Pedersen
· DBLP profile ↗
33ranked-venue papers
15as first author
7since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 33 · 15 first-author · 7 since 2021Databases, data management, data science and information retrieval · 9 · 4 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Dying or Departing? Euphemism Detection for Death Discourse in Historical TextsabstractEuphemisms are a linguistic device used to soften discussions of sensitive or uncomfortable topics, with death being a prominent example. In this paper, we present a study on the detection of death-related euphemisms in historical literary texts from a corpus containing Danish and Norwegian novels from the late 19th century. We introduce an annotated dataset of euphemistic and literal references to death, including both common and rare euphemisms, ranging from well-established terms to more culturally nuanced expressions. We evaluate the performances of state-of-the-art pre-trained language models fine-tuned for euphemism detection. Our findings show that fixed, literal expressions of death became less frequent over time, while metaphorical euphemisms grew in prevalence. Additionally, euphemistic language was more common in historical novels, whereas contemporary novels tended to refer to death more literally, reflecting the rise of secularism. These results shed light on the shifting discourse on death during a period when the concept of death as final became prominent. Ali Allaith, Alexander Conroy, Jens Bjerring-Hansen, Bolette S. Pedersen, Carsten Levisen, Daniel Hershcovich |
COLING | 4 |
| 2024 | Towards a Danish Semantic Reasoning Benchmark - Compiled from Lexical-Semantic Resources for Assessing Selected Language Understanding Capabilities of Large Language ModelsabstractWe present the first version of a semantic reasoning benchmark for Danish compiled semi-automatically from a number of human-curated lexical-semantic resources, which function as our gold standard. Taken together, the datasets constitute a benchmark for assessing selected language understanding capacities of large language models (LLMs) for Danish. This first version comprises 25 datasets across 6 different tasks and include 3,800 test instances. Although still somewhat limited in size, we go beyond comparative evaluation datasets for Danish by including both negative and contrastive examples as well as low-frequent vocabulary; aspects which tend to challenge current LLMs when based substantially on language transfer. The datasets focus on features such as semantic inference and entailment, similarity, relatedness, and ability to disambiguate words in context. We use ChatGPT to assess to which degree our datasets challenge the ceiling performance of state-of-the-art LLMs, average performance being relatively high with an average accuracy of 0.6 on ChatGPT 3.5 turbo and 0.8 on ChatGPT 4.0. Bolette S. Pedersen, Nathalie Carmen Hau Sørensen, Sussi Olsen, Sanni Nimb, Simon Gray |
LREC/COLING | 1 |
| 2023 | Reusing the Danish WordNet for a New Central Word Register for Danish - a Project ReportabstractIn this paper we report on a new Danish lexical initiative, the Central Word Register for Danish, (COR), which aims at providing an open-source, well curated and large-coverage lexicon for AI purposes.The semantic part of the lexicon (COR-S) relies to a large extent on the lexical-semantic information provided in the Danish wordnet, DanNet.However, we have taken the opportunity to evaluate and curate the wordnet information while compiling the new resource.Some information types have been simplified and more systematically curated.This is the case for the hyponymy relations, the ontological typing, and the sense inventory, i.e. the treatment of polysemy, including systematic polysemy. Bolette S. Pedersen, Sanni Nimb, Nathalie Carmen Hau Sørensen, Sussi Olsen, Ida Flørke, Thomas Troelsgård |
GWC | 1 |
| 2023 | How do We Treat Systematic Polysemy in Wordnets and Similar Resources? - Using Human Intuition and Contextualized Embeddings as GuidanceabstractSystematic polysemy is a well-known linguistic phenomenon where a group of lemmas follow the same polysemy pattern.However, when compiling a lexical resource like a wordnet, a problem arises regarding when to underspecify the two (or more) meanings by one (complex) sense and when to systematically split into separate senses.In this work, we present an extensive analysis of the systematic polysemy patterns in Danish, and in our preliminary study, we examine a subset of these with experiments on human intuition and contextual embeddings.The aim of this preparatory work is to enable future guidelines for each polysemy type.In the future, we hope to expand this approach and thereby hopefully obtain a sense inventory which is distributionally verified and thereby more suitable for NLP. Nathalie Carmen Hau Sørensen, Sanni Nimb, Bolette S. Pedersen |
GWC | 3 |
| 2022 | A Thesaurus-based Sentiment Lexicon for Danish: The Danish Sentiment LexiconabstractThis paper describes how a newly published Danish sentiment lexicon with a high lexical coverage was compiled by use of lexicographic methods and based on the links between groups of words listed in semantic order in a thesaurus and the corresponding word sense descriptions in a comprehensive monolingual dictionary. The overall idea was to identify negative and positive sections in a thesaurus, extract the words from these sections and combine them with the dictionary information via the links. The annotation task of the dataset included several steps, and was based on the comparison of synonyms and near synonyms within a semantic field. In the cases where one of the words were included in the smaller Danish sentiment lexicon AFINN, its value there was used as inspiration and expanded to the synonyms when appropriate. In order to obtain a more practical lexicon with overall polarity values at lemma level, all the senses of the lemma were afterwards compared, taking into consideration dictionary information such as usage, style and frequency. The final lexicon contains 13,859 Danish polarity lemmas and includes morphological information. It is freely available at https://github.com/dsldk/danish-sentiment-lexicon (licence CC-BY-SA 4.0 International). Sanni Nimb, Sussi Olsen, Bolette S. Pedersen, Thomas Troelsgård |
LREC | 3 |
| 2022 | Compiling a Suitable Level of Sense Granularity in a Lexicon for AI Purposes: The Open Source COR LexiconabstractWe present The Central Word Register for Danish (COR), which is an open source lexicon project for general AI purposes funded and initiated by the Danish Agency for Digitisation as part of an AI initiative embarked by the Danish Government in 2020. We focus here on the lexical semantic part of the project (COR-S) and describe how we – based on the existing fine-grained sense inventory from Den Danske Ordbog (DDO) – compile a more AI suitable sense granularity level of the vocabulary. A three-step methodology is applied: We establish a set of linguistic principles for defining core senses in COR-S and from there, we generate a hand-crafted gold standard of 6,000 lemmas depicting how to come from the fine-grained DDO sense to the COR inventory. Finally, we experiment with a number of language models in order to automatize the sense reduction of the rest of the lexicon. The models comprise a ruled-based model that applies our linguistic principles in terms of features, a word2vec model using cosine similarity to measure the sense proximity, and finally a deep neural BERT model fine-tuned on our annotations. The rule-based approach shows best results, in particular on adjectives, however, when focusing on the average polysemous vocabulary, the BERT model shows promising results too. Bolette S. Pedersen, Nathalie Carmen Hau Sørensen, Sanni Nimb, Ida Flørke, Sussi Olsen, Thomas Troelsgård |
LREC | 1 |
| 2021 | DanNet2: Extending the coverage of adjectives in DanNet based on thesaurus data (project presentation)abstractThe paper describes work in progress in the DanNet2 project financed by the Carlsberg Foundation.The project aim is to extend the original Danish wordnet, DanNet, in several ways.Main focus is on extension of the coverage and description of the adjectives, a part of speech that was rather sparsely described in the original wordnet.We describe the methodology and initial work of semiautomatically transferring adjectives from the Danish Thesaurus to the wordnet with the aim of easily enlarging the coverage from 3,000 to approx.13,000 adjectival synsets.Transfer is performed by manually encoding all missing adjectival subsection headwords from the thesaurus and thereafter employing a semiautomatic procedure where adjectives from the same subsection are transferred to the wordnet as either 1) near synonyms to the section's headword, 2) hyponyms to the section's headword, or 3) as members of the same synset as the headword.We also discuss how to deal with the problem of multiple representations of the same sense in the thesaurus, and present other types of information from the thesaurus that we plan to integrate, such as thematic and sentiment information. Sanni Nimb, Bolette S. Pedersen, Sussi Olsen |
GWC | 2 |
| 2020 | A Multilingual Evaluation Dataset for Monolingual Word Sense AlignmentabstractAligning senses across resources and languages is a challenging task with beneficial applications in the field of natural language processing and electronic lexicography. In this paper, we describe our efforts in manually aligning monolingual dictionaries. The alignment is carried out at sense-level for various resources in 15 languages. Moreover, senses are annotated with possible semantic relationships such as broadness, narrowness, relatedness, and equivalence. In comparison to previous datasets for this task, this dataset covers a wide range of languages and resources and focuses on the more challenging task of linking general-purpose language. We believe that our data will pave the way for further advances in alignment and evaluation of word senses by creating new solutions, particularly those notoriously requiring data such as neural networks. Our resources are publicly available at https://github.com/elexis-eu/MWSA. Sina Ahmadi, John P. McCrae, Sanni Nimb, Anas Fahad Khan, Monica Monachini, Bolette S. Pedersen, Thierry Declerck, Tanja Wissik, Andrea Bellandi, Irene Pisani, Thomas Troelsgård, Sussi Olsen, Simon Krek, Veronika Lipp, Tamás Váradi, László Simon, András Gyorffy, Carole Tiberius, Tanneke Schoonheim, Yifat Ben Moshe, Maya Rudich, Raya Abu Ahmad, Dorielle Lonke, Kira Kovalenko, Margit Langemets, Jelena Kallas, Oksana Dereza, Theodorus Fransen, David Cillessen, David Lindemann, Mikel Alonso, Ana Salgado, José-Luis Sancho-Gómez, Rafael-J. Ureña-Ruiz, Jordi Porta-Zamorano, Kiril Ivanov Simov, Petya Osenova, Zara Kancheva, Ivaylo Radev, Ranka Stankovic, Andrej Perdih, Dejan Gabrovsek |
LREC | 6 |
| 2020 | World Class Language Technology - Developing a Language Technology Strategy for DanishabstractAlthough Denmark is one of the most digitized countries in Europe, no coordinated efforts have been made in recent years to support the Danish language with regard to language technology and artificial intelligence. In March 2019, however, the Danish government adopted a new, ambitious strategy for LT and artificial intelligence. In this paper, we describe the process behind the development of the language-related parts of the strategy: A Danish Language Technology Committee was constituted and a comprehensive series of workshops were organized in which users, suppliers, developers, and researchers gave their valuable input based on their experiences. We describe how, based on this experience, the focus areas and recommendations for the LT strategy were established, and which steps are currently taken in order to put the strategy into practice. Sabine Kirchmeier, Bolette S. Pedersen, Sanni Nimb, Philip Diderichsen, Peter Juel Henrichsen |
LREC | 2 |
| 2020 | The European Language Technology Landscape in 2020: Language-Centric and Human-Centric AI for Cross-Cultural Communication in Multilingual EuropeabstractMultilingualism is a cultural cornerstone of Europe and firmly anchored in the European treaties including full language equality. However, language barriers impacting business, cross-lingual and cross-cultural communication are still omnipresent. Language Technologies (LTs) are a powerful means to break down these barriers. While the last decade has seen various initiatives that created a multitude of approaches and technologies tailored to Europe’s specific needs, there is still an immense level of fragmentation. At the same time, AI has become an increasingly important concept in the European Information and Communication Technology area. For a few years now, AI – including many opportunities, synergies but also misconceptions – has been overshadowing every other topic. We present an overview of the European LT landscape, describing funding programmes, activities, actions and challenges in the different countries with regard to LT, including the current state of play in industry and the LT market. We present a brief overview of the main LT-related activities on the EU level in the last ten years and develop strategic guidance with regard to four key dimensions. Georg Rehm, Katrin Marheinecke, Stefanie Hegele, Stelios Piperidis, Kalina Bontcheva, Jan Hajic 0001, Khalid Choukri, Andrejs Vasiljevs, Gerhard Backfried, Christoph Prinz, José Manuél Gómez-Pérez, Luc Meertens, Paul Lukowicz, Josef van Genabith, Andrea Lösch, Philipp Slusallek, Morten Irgens, Patrick Gatellier, Joachim Köhler, Laure Le Bars, Dimitra Anastasiou, Albina Auksoriute, Núria Bel, António Branco, Gerhard Budin, Walter Daelemans, Koenraad De Smedt, Radovan Garabík, Maria Gavrilidou, Dagmar Gromann, Svetla Koeva, Simon Krek, Cvetana Krstev, Krister Lindén, Bernardo Magnini, Jan Odijk, Maciej Ogrodniczuk, Eiríkur Rögnvaldsson, Mike Rosner, Bolette S. Pedersen, Inguna Skadina, Marko Tadic, Dan Tufis, Tamás Váradi, Kadri Vider, Andy Way, François Yvon |
LREC | 40 |
| 2020 | Towards a Gold Standard for Evaluating Danish Word EmbeddingsabstractThis paper presents the process of compiling a model-agnostic similarity goal standard for evaluating Danish word embeddings based on human judgments made by 42 native speakers of Danish. Word embeddings resemble semantic similarity solely by distribution (meaning that word vectors do not reflect relatedness as differing from similarity), and we argue that this generalization poses a problem in most intrinsic evaluation scenarios. In order to be able to evaluate on both dimensions, our human-generated dataset is therefore designed to reflect the distinction between relatedness and similarity. The goal standard is applied for evaluating the “goodness” of six existing word embedding models for Danish, and it is discussed how a relatively low correlation can be explained by the fact that semantic similarity is substantially more challenging to model than relatedness, and that there seems to be a need for future human judgments to measure similarity in full context and along more than a single spectrum. Nina Schneidermann, Rasmus Hvingelby, Bolette S. Pedersen |
LREC | 3 |
| 2019 | Merging DanNet with Princeton WordnetabstractIn this paper we describe the merge of the Danish wordnet, DanNet, with Princeton Wordnet applying a two-step approach.We first link from the English Princeton core to Danish (5,000 base concepts) and then proceed to linking the rest of the Danish vocabulary to English, thus going from Danish to English.Since the Danish wordnet is built bottom-up from Danish lexica and corpora, all taxonomies are monolingually based and thus not necessarily directly compatible with the coverage and structure of the Princeton WordNet.This fact proves to pose some challenges to the linking procedure since a considerable number of the links cannot be realised via the preferred crosslanguage synonym link which implies a more or less precise correlation between the two concepts.Instead, a subpart of the links are realised through near synonym or hyponymy links to compensate for the fact that no precise translation can be found in the target resource.The tool WordnetLoom is currently used for manual linking but procedures for a more automatic procedure in future is discussed.We conclude that the two resources actually differ from each other quite more than expected, both vocabulary-and structure-wise. Bolette S. Pedersen, Sanni Nimb, Ida Rørmann Olsen, Sussi Olsen |
GWC | 1 |
| 2018 | A Danish FrameNet Lexicon and an Annotated Corpus Used for Training and Evaluating a Semantic Frame Classifier
Bolette S. Pedersen, Sanni Nimb, Anders Søgaard, Mareike Hartmann, Sussi Olsen |
LREC | 1 |
| 2018 | Towards a principled approach to sense clustering - a case study of wordnet and dictionary senses in DanishabstractOur aim is to develop principled methods for sense clustering which can make existing lexical resources practically useful in NLPnot too fine-grained to be operational and yet finegrained enough to be worth the trouble.Where traditional dictionaries have a highly structured sense inventory typically describing the vocabulary by means of main-and subsenses, wordnets are generally fine-grained and unstructured.We present a series of clustering and annotation experiments with 10 of the most polysemous nouns in Danish.We combine the structured information of a traditional Danish dictionary with the ontological types found in the Danish wordnet, DanNet.This constellation enables us to automatically cluster senses in a principled way and improve inter-annotator agreement and wsd performance. Bolette S. Pedersen, Manex Agirrezabal, Sanni Nimb, Ida Rørmann Olsen, Sussi Olsen |
GWC | 1 |
| 2018 | ELEXIS - a European infrastructure fostering cooperation and information exchange among lexicographical research communitiesabstractThe paper describes objectives, concept and methodology for ELEXIS, a European infrastructure fostering cooperation and information exchange among lexicographical research communities.The infrastructure is a newly granted project under the Horizon 2020 INFRAIA call, with the topic Integrating Activities for Starting Communities.The project is planned to start in January 2018. Bolette S. Pedersen, John P. McCrae, Carole Tiberius, Simon Krek |
GWC | 1 |
| 2016 | The SemDaX Corpus ― Sense Annotations with Scalable Sense Inventories
Bolette S. Pedersen, Anna Braasch, Anders Johannsen, Héctor Martínez Alonso, Sanni Nimb, Sussi Olsen, Anders Søgaard, Nicolai Hartvig Sørensen |
LREC | 1 |
| 2016 | An empirically grounded expansion of the supersense inventoryabstractIn this article we present an expansion of the supersense inventory.All new supersenses are extensions of members of the current inventory, which we postulate by identifying semantically coherent groups of synsets.We cover the expansion of the already-established supernsense inventory for nouns and verbs, the addition of coarse supersenses for adjectives in absence of a canonical supersense inventory, and supersenses for verbal satellites.We evaluate the viability of the new senses examining the annotation agreement, frequency and co-ocurrence patterns. Héctor Martínez Alonso, Anders Johannsen, Sanni Nimb, Sussi Olsen, Bolette S. Pedersen |
GWC | 5 |
| 2014 | Encompassing a spectrum of LT users in the CLARIN-DK Infrastructure
Lina Henriksen, Dorte Haltrup Hansen, Bente Maegaard, Bolette S. Pedersen, Claus Povlsen |
LREC | 4 |
| 2014 | The Strategic Impact of META-NET on the Regional, National and International Level
Georg Rehm, Hans Uszkoreit, Sophia Ananiadou, Núria Bel, Audroné Bieleviciené, Lars Borin, António Branco, Gerhard Budin, Nicoletta Calzolari, Walter Daelemans, Radovan Garabík, Marko Grobelnik, Carmen García-Mateo, Josef van Genabith, Jan Hajic 0001, Inma Hernáez Rioja, John Judge, Svetla Koeva, Simon Krek, Cvetana Krstev, Krister Lindén, Bernardo Magnini, Joseph Mariani, John McNaught, Maite Melero, Monica Monachini, Asunción Moreno, Jan Odijk, Maciej Ogrodniczuk, Piotr Pezik, Stelios Piperidis, Adam Przepiórkowski, Eiríkur Rögnvaldsson, Mike Rosner, Bolette S. Pedersen, Inguna Skadina, Koenraad De Smedt, Marko Tadic, Paul Thompson 0002, Dan Tufis, Tamás Váradi, Andrejs Vasiljevs, Kadri Vider, Jolanta Zabarskaite |
LREC | 35 |
| 2014 | CLARA: A New Generation of Researchers in Common Language Resources and Their Applications
Koenraad De Smedt, Erhard W. Hinrichs, Detmar Meurers, Inguna Skadina, Bolette S. Pedersen, Costanza Navarretta, Núria Bel, Krister Lindén, Markéta Lopatková, Jan Hajic 0001, Gisle Andersen, Przemyslaw Lenkiewicz |
LREC | 5 |
| 2012 | A voting scheme to detect semantic underspecification
Héctor Martínez Alonso, Núria Bel, Bolette S. Pedersen |
LREC | 3 |
| 2012 | Towards a richer wordnet representation of properties
Sanni Nimb, Bolette S. Pedersen |
LREC | 2 |
| 2012 | Creation of an Open Shared Language Resource Repository in the Nordic and Baltic Countries
Andrejs Vasiljevs, Markus Forsberg, Tatiana Gornostay, Dorte Haltrup Hansen, Kristín Jóhannsdóttir, Gunn Inger Lyse, Krister Lindén, Lene Offersgaard, Sussi Olsen, Bolette S. Pedersen, Eiríkur Rögnvaldsson, Inguna Skadina, Koenraad De Smedt, Ville Oksanen, Roberts Rozis |
LREC | 10 |
| 2010 | Merging Specialist Taxonomies and Folk Taxonomies in Wordnets - A case Study of Plants, Animals and Foods in the Danish Wordnet
Bolette S. Pedersen, Sanni Nimb, Anna Braasch |
LREC | 1 |
| 2008 | Merging a Syntactic Resource with a WordNet: a Feasibility Study of a Merge between STO and DanNet
Bolette S. Pedersen, Anna Braasch, Lina Henriksen, Sussi Olsen, Claus Povlsen |
LREC | 1 |
| 2007 | Using shallow linguistic analysis to improve search on Danish compoundsabstractIn this paper we focus on a specific search-related query expansion topic, namely search on Danish compounds and expansion to some of their synonymous phrases. Compounds constitute a specific issue in search, in particular in languages where they are written in one word, as is the case for Danish and the other Scandinavian languages. For such languages, expansion of the query compound into separate lemmas is a way of finding the often frequent alternative synonymous phrases in which the content of a compound can also be expressed. However, it is crucial to note that the number of irrelevant hits is generally very high when using this expansion strategy. The aim of this paper is therefore to examine how we can obtain better search results on split compounds, partly by looking at the internal structure of the original compound, partly by analyzing the context in which the split compound occurs. In this context, we pursue two hypotheses: (1) that some categories of compounds are more likely to have synonymous ‘split’ counterparts than others; and (2) that search results where both the search words (obtained by splitting the compound) occur in the same noun phrase, are more likely to contain a synonymous phrase to the original compound query. The search results from 410 enhanced compound queries are used as a test bed for our experiments. On these search results, we perform a shallow linguistic analysis and introduce a new, linguistically based threshold for retrieved hits. The results obtained by using this strategy demonstrate that compound splitting combined with a shallow linguistic analysis focusing on the argument structure of the compound head as well as on the recognition of NPs, can improve search by substantially bringing down the number of irrelevant hits. Bolette S. Pedersen |
Nat. Lang. Eng. | 1 |
| 2006 | Query Expansion on Compounds
Bolette S. Pedersen |
LREC | 1 |
| 2004 | Human Language Technology Elements in a Knowledge Organisation System - The VID Project
Costanza Navarretta, Bolette S. Pedersen, Dorte Haltrup Hansen |
LREC | 2 |
| 2004 | Content-based text querying with ontological descriptors
Troels Andreasen, Per Anker Jensen, Jørgen Fischer Nilsson, Patrizia Paggio, Bolette S. Pedersen, Hanne Erdman Thomsen |
Data Knowl. Eng. | 5 |
| 2002 | Semantic Lexical Resources Applied to Content-based Querying - the OntoQuery Project
Bolette S. Pedersen, Patrizia Paggio |
LREC | 1 |
| 2002 | Ontological Extraction of Content for Text Querying
Troels Andreasen, Per Anker Jensen, Jørgen Fischer Nilsson, Patrizia Paggio, Bolette S. Pedersen, Hanne Erdman Thomsen |
NLDB | 5 |
| 2000 | Semantic Encoding of Danish Verbs in SIMPLE - Adapting a Verb Framed Model to a Satellite-framed Language
Bolette S. Pedersen, Sanni Nimb |
LREC | 1 |
| 1999 | Systematic Verb Polysemy in MT: A Study of Danish Motion Verbs with Comparisons with Spanish
Bolette S. Pedersen |
Mach. Transl. | 1 |