Sussi Olsen

dblp:56/8155 · DBLP profile ↗
← Back
26ranked-venue papers
1as first author
9since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 26 · 1 first-author · 9 since 2021Databases, data management, data science and information retrieval · 7 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 DAMETA: An LLM Benchmark for Danish Metaphor Interpretation with Systematically Varied Distractors
Nina Schneidermann, Sanni Nimb, Nathalie Carmen Hau Norman, Sussi Olsen, Bolette Pedersen
LREC4
2026 A Parallel Cross-Lingual Benchmark for Multimodal Idiomaticity Understanding
abstract
Potentially idiomatic expressions (PIEs) carry meanings inherently tied to the everyday experience of a given language community. As such, they constitute an interesting challenge for assessing the linguistic (and to some extent cultural) capabilities of NLP systems. In this paper, we present XMPIE, a parallel multilingual and multimodal dataset of potentially idiomatic expressions. The dataset, containing 34 languages and over ten thousand items, allows comparative analyses of idiomatic patterns among language-specific realisations and preferences in order to gather insights about shared cultural aspects. This parallel dataset allows evaluation of language model performance for a given PIE in different languages and whether idiomatic understanding in one language can be transferred to another. Moreover, the dataset supports the study of PIEs across textual and visual modalities, to measure to what extent PIE understanding in one modality transfers or implies in understanding in another modality (text vs. image). The data was created by language experts, with both textual and visual components crafted under multilingual guidelines, and each PIE is accompanied by five images representing a spectrum from idiomatic to literal meanings, including semantically related and random distractors. The result is a high-quality benchmark for evaluating multilingual and multimodal idiomatic language understanding.
Dilara Torunoglu-Selamet, Dogukan Arslan, Rodrigo Wilkens, Wei He 0017, Doruk Eryigit, Thomas Pickard, Adriana S. Pagano, Aline Villavicencio, Gülsen Eryigit, Ágnes Abuczki, Aida Cardoso, Alesia Lazarenka, Dina Almassova, Amália Mendes, Anna Kanellopoulou, Antoni Brosa-Rodríguez, Baiba Valkovska, Beata Wojtowicz, Bolette Pedersen, Carlos Manuel Hidalgo-Ternero, Chaya Liebeskind, Danka Jokic, Diego Alves, Eleni Triantafyllidi, Erik Velldal, Fred Philippy, Giedre Valunaite Oleskeviciene, Ieva Rizgeliene, Inguna Skadina, Irina Lobzhanidze, Isabell Stinessen Haugen, Jauza Akbar Krito, Jelena M. Markovic, Johanna Monti, Josue Alejandro Sauca, Kaja Dobrovoljc, Kingsley O. Ugwuanyi, Laura Rituma, Lilja Øvrelid, Maha Tufail Agro, Manzura Abjalova, Maria Chatzigrigoriou, María del Mar Sánchez Ramos, Marija Pendevska, Masoumeh Seyyedrezaei, Mehrnoush Shamsfard, Momina Ahsan, Muhammad Ahsan Riaz Khan, Nathalie Carmen Hau Norman, Nilay Erdem Ayyildiz, Nina Hosseini-Kivanani, Noémi Ligeti-Nagy, Numaan Naeem, Olha Kanishcheva, Olha Yatsyshyna, Daniil Orel, Petra Giommarelli, Petya Osenova, Radovan Garabík, Regina E. Semou, Rozane Rebechi, Salsabila Zahirah Pranida, Samia Touileb, Sanni Nimb, Sarvinoz Sharipova, Shahar Golan, Shaoxiong Ji, Sopuruchi Christian Aboh, Srdjan Sucur, Stella Markantonatou, Sussi Olsen, Vahideh Tajalli, Veronika Lipp, Voula Giouli, Yelda Yesildal Eraydin, Zahra Saaberi, Zhuohan Xie
LREC72
2024 Towards a Danish Semantic Reasoning Benchmark - Compiled from Lexical-Semantic Resources for Assessing Selected Language Understanding Capabilities of Large Language Models
abstract
We present the first version of a semantic reasoning benchmark for Danish compiled semi-automatically from a number of human-curated lexical-semantic resources, which function as our gold standard. Taken together, the datasets constitute a benchmark for assessing selected language understanding capacities of large language models (LLMs) for Danish. This first version comprises 25 datasets across 6 different tasks and include 3,800 test instances. Although still somewhat limited in size, we go beyond comparative evaluation datasets for Danish by including both negative and contrastive examples as well as low-frequent vocabulary; aspects which tend to challenge current LLMs when based substantially on language transfer. The datasets focus on features such as semantic inference and entailment, similarity, relatedness, and ability to disambiguate words in context. We use ChatGPT to assess to which degree our datasets challenge the ceiling performance of state-of-the-art LLMs, average performance being relatively high with an average accuracy of 0.6 on ChatGPT 3.5 turbo and 0.8 on ChatGPT 4.0.
Bolette S. Pedersen, Nathalie Carmen Hau Sørensen, Sussi Olsen, Sanni Nimb, Simon Gray
LREC/COLING3
2023 A uniform RDF-based Representation of the Interlinking of Wordnets and Sign Language Data
Thierry Declerck, Sam Bigeard, Dorians Callus, Benjamin Matthews, Sussi Olsen, Loran Ripard Xuereb
LDK5
2023 Towards an RDF Representation of the Infrastructure consisting in using Wordnets as a conceptual Interlingua between multilingual Sign Language Datasets
abstract
We present ongoing work dealing with a Linked Data compliant representation of infrastructures using wordnets for connecting multilingual Sign Language data sets.We build for this on already existing RDF and OntoLex representations of Open Multilingual Wordnet (OMW) data sets and work done by the European EAS-IER research project on the use of the CSV files of OMW for linking glosses and basic semantic information associated with Sign Language data sets in two languages: German and Greek.In this context, we started the transformation into RDF of a Danish data set, which links Danish Sign Language data and the wordnet for Danish, DanNet.The final objective of our work is to include Sign Language data sets (and their conceptual cross-linking via wordnets) in the Linguistic Linked Open Data cloud.
Thierry Declerck, Thomas Troelsgård, Sussi Olsen
GWC3
2023 Reusing the Danish WordNet for a New Central Word Register for Danish - a Project Report
abstract
In this paper we report on a new Danish lexical initiative, the Central Word Register for Danish, (COR), which aims at providing an open-source, well curated and large-coverage lexicon for AI purposes.The semantic part of the lexicon (COR-S) relies to a large extent on the lexical-semantic information provided in the Danish wordnet, DanNet.However, we have taken the opportunity to evaluate and curate the wordnet information while compiling the new resource.Some information types have been simplified and more systematically curated.This is the case for the hyponymy relations, the ontological typing, and the sense inventory, i.e. the treatment of polysemy, including systematic polysemy.
Bolette S. Pedersen, Sanni Nimb, Nathalie Carmen Hau Sørensen, Sussi Olsen, Ida Flørke, Thomas Troelsgård
GWC4
2022 A Thesaurus-based Sentiment Lexicon for Danish: The Danish Sentiment Lexicon
abstract
This paper describes how a newly published Danish sentiment lexicon with a high lexical coverage was compiled by use of lexicographic methods and based on the links between groups of words listed in semantic order in a thesaurus and the corresponding word sense descriptions in a comprehensive monolingual dictionary. The overall idea was to identify negative and positive sections in a thesaurus, extract the words from these sections and combine them with the dictionary information via the links. The annotation task of the dataset included several steps, and was based on the comparison of synonyms and near synonyms within a semantic field. In the cases where one of the words were included in the smaller Danish sentiment lexicon AFINN, its value there was used as inspiration and expanded to the synonyms when appropriate. In order to obtain a more practical lexicon with overall polarity values at lemma level, all the senses of the lemma were afterwards compared, taking into consideration dictionary information such as usage, style and frequency. The final lexicon contains 13,859 Danish polarity lemmas and includes morphological information. It is freely available at https://github.com/dsldk/danish-sentiment-lexicon (licence CC-BY-SA 4.0 International).
Sanni Nimb, Sussi Olsen, Bolette S. Pedersen, Thomas Troelsgård
LREC2
2022 Compiling a Suitable Level of Sense Granularity in a Lexicon for AI Purposes: The Open Source COR Lexicon
abstract
We present The Central Word Register for Danish (COR), which is an open source lexicon project for general AI purposes funded and initiated by the Danish Agency for Digitisation as part of an AI initiative embarked by the Danish Government in 2020. We focus here on the lexical semantic part of the project (COR-S) and describe how we – based on the existing fine-grained sense inventory from Den Danske Ordbog (DDO) – compile a more AI suitable sense granularity level of the vocabulary. A three-step methodology is applied: We establish a set of linguistic principles for defining core senses in COR-S and from there, we generate a hand-crafted gold standard of 6,000 lemmas depicting how to come from the fine-grained DDO sense to the COR inventory. Finally, we experiment with a number of language models in order to automatize the sense reduction of the rest of the lexicon. The models comprise a ruled-based model that applies our linguistic principles in terms of features, a word2vec model using cosine similarity to measure the sense proximity, and finally a deep neural BERT model fine-tuned on our annotations. The rule-based approach shows best results, in particular on adjectives, however, when focusing on the average polysemous vocabulary, the BERT model shows promising results too.
Bolette S. Pedersen, Nathalie Carmen Hau Sørensen, Sanni Nimb, Ida Flørke, Sussi Olsen, Thomas Troelsgård
LREC5
2021 DanNet2: Extending the coverage of adjectives in DanNet based on thesaurus data (project presentation)
abstract
The paper describes work in progress in the DanNet2 project financed by the Carlsberg Foundation.The project aim is to extend the original Danish wordnet, DanNet, in several ways.Main focus is on extension of the coverage and description of the adjectives, a part of speech that was rather sparsely described in the original wordnet.We describe the methodology and initial work of semiautomatically transferring adjectives from the Danish Thesaurus to the wordnet with the aim of easily enlarging the coverage from 3,000 to approx.13,000 adjectival synsets.Transfer is performed by manually encoding all missing adjectival subsection headwords from the thesaurus and thereafter employing a semiautomatic procedure where adjectives from the same subsection are transferred to the wordnet as either 1) near synonyms to the section's headword, 2) hyponyms to the section's headword, or 3) as members of the same synset as the headword.We also discuss how to deal with the problem of multiple representations of the same sense in the thesaurus, and present other types of information from the thesaurus that we plan to integrate, such as thematic and sentiment information.
Sanni Nimb, Bolette S. Pedersen, Sussi Olsen
GWC3
2020 A Multilingual Evaluation Dataset for Monolingual Word Sense Alignment
abstract
Aligning senses across resources and languages is a challenging task with beneficial applications in the field of natural language processing and electronic lexicography. In this paper, we describe our efforts in manually aligning monolingual dictionaries. The alignment is carried out at sense-level for various resources in 15 languages. Moreover, senses are annotated with possible semantic relationships such as broadness, narrowness, relatedness, and equivalence. In comparison to previous datasets for this task, this dataset covers a wide range of languages and resources and focuses on the more challenging task of linking general-purpose language. We believe that our data will pave the way for further advances in alignment and evaluation of word senses by creating new solutions, particularly those notoriously requiring data such as neural networks. Our resources are publicly available at https://github.com/elexis-eu/MWSA.
Sina Ahmadi, John P. McCrae, Sanni Nimb, Anas Fahad Khan, Monica Monachini, Bolette S. Pedersen, Thierry Declerck, Tanja Wissik, Andrea Bellandi, Irene Pisani, Thomas Troelsgård, Sussi Olsen, Simon Krek, Veronika Lipp, Tamás Váradi, László Simon, András Gyorffy, Carole Tiberius, Tanneke Schoonheim, Yifat Ben Moshe, Maya Rudich, Raya Abu Ahmad, Dorielle Lonke, Kira Kovalenko, Margit Langemets, Jelena Kallas, Oksana Dereza, Theodorus Fransen, David Cillessen, David Lindemann, Mikel Alonso, Ana Salgado, José-Luis Sancho-Gómez, Rafael-J. Ureña-Ruiz, Jordi Porta-Zamorano, Kiril Ivanov Simov, Petya Osenova, Zara Kancheva, Ivaylo Radev, Ranka Stankovic, Andrej Perdih, Dejan Gabrovsek
LREC12
2019 Merging DanNet with Princeton Wordnet
abstract
In this paper we describe the merge of the Danish wordnet, DanNet, with Princeton Wordnet applying a two-step approach.We first link from the English Princeton core to Danish (5,000 base concepts) and then proceed to linking the rest of the Danish vocabulary to English, thus going from Danish to English.Since the Danish wordnet is built bottom-up from Danish lexica and corpora, all taxonomies are monolingually based and thus not necessarily directly compatible with the coverage and structure of the Princeton WordNet.This fact proves to pose some challenges to the linking procedure since a considerable number of the links cannot be realised via the preferred crosslanguage synonym link which implies a more or less precise correlation between the two concepts.Instead, a subpart of the links are realised through near synonym or hyponymy links to compensate for the fact that no precise translation can be found in the target resource.The tool WordnetLoom is currently used for manual linking but procedures for a more automatic procedure in future is discussed.We conclude that the two resources actually differ from each other quite more than expected, both vocabulary-and structure-wise.
Bolette S. Pedersen, Sanni Nimb, Ida Rørmann Olsen, Sussi Olsen
GWC4
2018 A Danish FrameNet Lexicon and an Annotated Corpus Used for Training and Evaluating a Semantic Frame Classifier
Bolette S. Pedersen, Sanni Nimb, Anders Søgaard, Mareike Hartmann, Sussi Olsen
LREC5
2018 Towards a principled approach to sense clustering - a case study of wordnet and dictionary senses in Danish
abstract
Our aim is to develop principled methods for sense clustering which can make existing lexical resources practically useful in NLPnot too fine-grained to be operational and yet finegrained enough to be worth the trouble.Where traditional dictionaries have a highly structured sense inventory typically describing the vocabulary by means of main-and subsenses, wordnets are generally fine-grained and unstructured.We present a series of clustering and annotation experiments with 10 of the most polysemous nouns in Danish.We combine the structured information of a traditional Danish dictionary with the ontological types found in the Danish wordnet, DanNet.This constellation enables us to automatically cluster senses in a principled way and improve inter-annotator agreement and wsd performance.
Bolette S. Pedersen, Manex Agirrezabal, Sanni Nimb, Ida Rørmann Olsen, Sussi Olsen
GWC5
2016 Providing a Catalogue of Language Resources for Commercial Users
Bente Maegaard, Lina Henriksen, Andrew Joscelyne, Vesna Lusicky, Margaretha Mazura, Sussi Olsen, Claus Povlsen, Philippe Wacker
LREC6
2016 The SemDaX Corpus ― Sense Annotations with Scalable Sense Inventories
Bolette S. Pedersen, Anna Braasch, Anders Johannsen, Héctor Martínez Alonso, Sanni Nimb, Sussi Olsen, Anders Søgaard, Nicolai Hartvig Sørensen
LREC6
2016 An empirically grounded expansion of the supersense inventory
abstract
In this article we present an expansion of the supersense inventory.All new supersenses are extensions of members of the current inventory, which we postulate by identifying semantically coherent groups of synsets.We cover the expansion of the already-established supernsense inventory for nouns and verbs, the addition of coarse supersenses for adjectives in absence of a canonical supersense inventory, and supersenses for verbal satellites.We evaluate the viability of the new senses examining the annotation agreement, frequency and co-ocurrence patterns.
Héctor Martínez Alonso, Anders Johannsen, Sanni Nimb, Sussi Olsen, Bolette S. Pedersen
GWC4
2014 Using TEI, CMDI and ISOcat in CLARIN-DK
Dorte Haltrup Hansen, Lene Offersgaard, Sussi Olsen
LREC3
2012 A Distributed Resource Repository for Cloud-Based Machine Translation
Jörg Tiedemann, Dorte Haltrup Hansen, Lene Offersgaard, Sussi Olsen, Matthias Zumpe
LREC4
2012 Creation of an Open Shared Language Resource Repository in the Nordic and Baltic Countries
Andrejs Vasiljevs, Markus Forsberg, Tatiana Gornostay, Dorte Haltrup Hansen, Kristín Jóhannsdóttir, Gunn Inger Lyse, Krister Lindén, Lene Offersgaard, Sussi Olsen, Bolette S. Pedersen, Eiríkur Rögnvaldsson, Inguna Skadina, Koenraad De Smedt, Ville Oksanen, Roberts Rozis
LREC9
2010 Quality Indicators of LSP Texts - Selection and Measurements Measuring the Terminological Usefulness of Documents for an LSP Corpus
Jakob Halskov, Dorte Haltrup Hansen, Anna Braasch, Sussi Olsen
LREC4
2008 Annotating Abstract Pronominal Anaphora in the DAD Project
Costanza Navarretta, Sussi Olsen
LREC2
2008 Merging a Syntactic Resource with a WordNet: a Feasibility Study of a Merge between STO and DanNet
Bolette S. Pedersen, Anna Braasch, Lina Henriksen, Sussi Olsen, Claus Povlsen
LREC4
2004 STO: A Danish Lexicon Resource - Ready for Applications
Anna Braasch, Sussi Olsen
LREC2
2002 Lemma selection in domain specific computational lexica - some specific problems
Sussi Olsen
LREC1
2000 Towards a Strategy for a Representation of Collocations - Extending the Danish PAROLE-lexicon
Anna Braasch, Sussi Olsen
LREC2
1998 A large scale lexicon for danish in the information society
Anna Braasch, A. B. Christensen, Sussi Olsen, B. S. Perdersen
LREC3