EDBT 2026 Demo / reviewers in the wild / expert
Lilja Øvrelid
dblp:03/6797
· DBLP profile ↗
35ranked-venue papers
7as first author
15since 2021 · last 2026
0000-0002-8880-8404ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 34 · 7 first-author · 14 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Context Is (Almost) Everything: Llama-3 on Structured Output and AMR Parsing
Maja Buljan, Stephan Oepen, Lilja Øvrelid |
LREC | 3 |
| 2026 | MUC-4 Revisited: Document-level Event Analysis beyond Span-based Arguments
Helene Olsen, Erik Velldal, Lilja Øvrelid |
LREC | 3 |
| 2026 | Entity-Level Sentiment Analysis with Sentence Relevance Detection
Egil Rønningstad, Roman Klinger, Lilja Øvrelid, Erik Velldal |
LREC | 3 |
| 2026 | A Parallel Cross-Lingual Benchmark for Multimodal Idiomaticity UnderstandingabstractPotentially idiomatic expressions (PIEs) carry meanings inherently tied to the everyday experience of a given language community. As such, they constitute an interesting challenge for assessing the linguistic (and to some extent cultural) capabilities of NLP systems. In this paper, we present XMPIE, a parallel multilingual and multimodal dataset of potentially idiomatic expressions. The dataset, containing 34 languages and over ten thousand items, allows comparative analyses of idiomatic patterns among language-specific realisations and preferences in order to gather insights about shared cultural aspects. This parallel dataset allows evaluation of language model performance for a given PIE in different languages and whether idiomatic understanding in one language can be transferred to another. Moreover, the dataset supports the study of PIEs across textual and visual modalities, to measure to what extent PIE understanding in one modality transfers or implies in understanding in another modality (text vs. image). The data was created by language experts, with both textual and visual components crafted under multilingual guidelines, and each PIE is accompanied by five images representing a spectrum from idiomatic to literal meanings, including semantically related and random distractors. The result is a high-quality benchmark for evaluating multilingual and multimodal idiomatic language understanding. Dilara Torunoglu-Selamet, Dogukan Arslan, Rodrigo Wilkens, Wei He 0017, Doruk Eryigit, Thomas Pickard, Adriana S. Pagano, Aline Villavicencio, Gülsen Eryigit, Ágnes Abuczki, Aida Cardoso, Alesia Lazarenka, Dina Almassova, Amália Mendes, Anna Kanellopoulou, Antoni Brosa-Rodríguez, Baiba Valkovska, Beata Wojtowicz, Bolette Pedersen, Carlos Manuel Hidalgo-Ternero, Chaya Liebeskind, Danka Jokic, Diego Alves, Eleni Triantafyllidi, Erik Velldal, Fred Philippy, Giedre Valunaite Oleskeviciene, Ieva Rizgeliene, Inguna Skadina, Irina Lobzhanidze, Isabell Stinessen Haugen, Jauza Akbar Krito, Jelena M. Markovic, Johanna Monti, Josue Alejandro Sauca, Kaja Dobrovoljc, Kingsley O. Ugwuanyi, Laura Rituma, Lilja Øvrelid, Maha Tufail Agro, Manzura Abjalova, Maria Chatzigrigoriou, María del Mar Sánchez Ramos, Marija Pendevska, Masoumeh Seyyedrezaei, Mehrnoush Shamsfard, Momina Ahsan, Muhammad Ahsan Riaz Khan, Nathalie Carmen Hau Norman, Nilay Erdem Ayyildiz, Nina Hosseini-Kivanani, Noémi Ligeti-Nagy, Numaan Naeem, Olha Kanishcheva, Olha Yatsyshyna, Daniil Orel, Petra Giommarelli, Petya Osenova, Radovan Garabík, Regina E. Semou, Rozane Rebechi, Salsabila Zahirah Pranida, Samia Touileb, Sanni Nimb, Sarvinoz Sharipova, Shahar Golan, Shaoxiong Ji, Sopuruchi Christian Aboh, Srdjan Sucur, Stella Markantonatou, Sussi Olsen, Vahideh Tajalli, Veronika Lipp, Voula Giouli, Yelda Yesildal Eraydin, Zahra Saaberi, Zhuohan Xie |
LREC | 39 |
| 2025 | ELOQUENT CLEF Shared Tasks for Evaluation of Generative Language Model Quality, 2025 Edition
Jussi Karlgren, Ekaterina Artemova, Ondrej Bojar, Vladislav Mikhailov, Magnus Sahlgren, Erik Velldal, Lilja Øvrelid |
ECIR (5) | 7 |
| 2024 | EDEN: A Dataset for Event Detection in Norwegian NewsabstractWe present EDEN, the first Norwegian dataset annotated with event information at the sentence level, adapting the widely used ACE event schema to Norwegian. The paper describes the manual annotation of Norwegian text as well as transcribed speech in the news domain, together with inter-annotator agreement and discussions of relevant dataset statistics. We also present preliminary modeling results using a graph-based event parser. The resulting dataset will be freely available for download and use. Samia Touileb, Jeanett Murstad, Petter Mæhlum, Lubos Steskal, Lilja Charlotte Storset, Huiling You, Lilja Øvrelid |
LREC/COLING | 7 |
| 2024 | PashtoEmo: Enhancing Text-Based Emotion Analysis in the Pashto Language Through Dataset CreationabstractThis paper presents the comprehensive PashtoEmo dataset for emotion analysis in the Pashto language. PashtoEmo contains 8016 text instances from $$\mathbb {X}$$ social media platform covering several topics, including politics, women’s rights, social justice, culture, sports, and education. The texts are annotated with six emotional categories and an additional ‘Other’ class. We tested PashtoEmo for emotion analysis with several machine learning, deep learning, and transformer models. Even the best model, XLM-RoBERTa-large, achieved only a 76.85% F1 score and a 78.11% accuracy. This indicates that the data set must be extended with adequate resources in Pashto language text-based applications. Mohammad Arif Payenda, Abdul Razaq Vahidi, Mohammad Ali Hussiny, Andreas Prinz 0001, Lilja Øvrelid |
NLDB (2) | 5 |
| 2023 | Measuring Normative and Descriptive Biases in Language Models Using Census DataabstractWe investigate in this paper how distributions of occupations with respect to gender is reflected in pre-trained language models.Such distributions are not always aligned to normative ideals, nor do they necessarily reflect a descriptive assessment of reality.In this paper, we introduce an approach for measuring to what degree pre-trained language models are aligned to normative and descriptive occupational distributions.To this end, we use official demographic information about gender-occupation distributions provided by the national statistics agencies of France, Norway, United Kingdom, and the United States.We manually generate template-based sentences combining gendered pronouns and nouns with occupations, and subsequently probe a selection of ten language models covering the English, French, and Norwegian languages.The scoring system we introduce in this work is language independent, and can be used on any combination of template-based sentences, occupations, and languages.The approach could also be extended to other dimensions of national census data and other demographic variables. Samia Touileb, Lilja Øvrelid, Erik Velldal |
EACL | 2 |
| 2022 | Entity-Level Sentiment Analysis (ELSA): An Exploratory Task SurveyabstractThis paper explores the task of identifying the overall sentiment expressed towards volitional entities (persons and organizations) in a document - what we refer to as Entity-Level Sentiment Analysis (ELSA). While identifying sentiment conveyed towards an entity is well researched for shorter texts like tweets, we find little to no research on this specific task for longer texts with multiple mentions and opinions towards the same entity. This lack of research would be understandable if ELSA can be derived from existing tasks and models. To assess this, we annotate a set of professional reviews for their overall sentiment towards each volitional entity in the text. We sample from data already annotated for document-level, sentence-level, and target-level sentiment in a multi-domain review corpus, and our results indicate that there is no single proxy task that provides this overall sentiment we seek for the entities at a satisfactory level of performance. We present a suite of experiments aiming to assess the contribution towards ELSA provided by document-, sentence-, and target-level sentiment analysis, and provide a discussion of their shortcomings. We show that sentiment in our dataset is expressed not only with an entity mention as target, but also towards targets with a sentiment-relevant relation to a volitional entity. In our data, these relations extend beyond anaphoric coreference resolution, and our findings call for further research of the topic. Finally, we also present a survey of previous relevant work. Egil Rønningstad, Erik Velldal, Lilja Øvrelid |
COLING | 3 |
| 2022 | Bootstrapping Text Anonymization Models with Distant SupervisionabstractWe propose a novel method to bootstrap text anonymization models based on distant supervision. Instead of requiring manually labeled training data, the approach relies on a knowledge graph expressing the background information assumed to be publicly available about various individuals. This knowledge graph is employed to automatically annotate text documents including personal data about a subset of those individuals. More precisely, the method determines which text spans ought to be masked in order to guarantee k-anonymity, assuming an adversary with access to both the text documents and the background information expressed in the knowledge graph. The resulting collection of labeled documents is then used as training data to fine-tune a pre-trained language model for text anonymization. We illustrate this approach using a knowledge graph extracted from Wikidata and short biographical texts from Wikipedia. Evaluation results with a RoBERTa-based model and a manually annotated collection of 553 summaries showcase the potential of the approach, but also unveil a number of issues that may arise if the knowledge graph is noisy or incomplete. The results also illustrate that, contrary to most sequence labeling problems, the text anonymization task may admit several alternative solutions. Anthi Papadopoulou, Pierre Lison, Lilja Øvrelid, Ildikó Pilán |
LREC | 3 |
| 2022 | The Text Anonymization Benchmark (TAB): A Dedicated Corpus and Evaluation Framework for Text AnonymizationabstractAbstract We present a novel benchmark and associated evaluation metrics for assessing the performance of text anonymization methods. Text anonymization, defined as the task of editing a text document to prevent the disclosure of personal information, currently suffers from a shortage of privacy-oriented annotated text resources, making it difficult to properly evaluate the level of privacy protection offered by various anonymization methods. This paper presents TAB (Text Anonymization Benchmark), a new, open-source annotated corpus developed to address this shortage. The corpus comprises 1,268 English-language court cases from the European Court of Human Rights (ECHR) enriched with comprehensive annotations about the personal information appearing in each document, including their semantic category, identifier type, confidential attributes, and co-reference relations. Compared with previous work, the TAB corpus is designed to go beyond traditional de-identification (which is limited to the detection of predefined semantic categories), and explicitly marks which text spans ought to be masked in order to conceal the identity of the person to be protected. Along with presenting the corpus and its annotation layers, we also propose a set of evaluation metrics that are specifically tailored toward measuring the performance of text anonymization, both in terms of privacy protection and utility preservation. We illustrate the use of the benchmark and the proposed metrics by assessing the empirical performance of several baseline text anonymization models. The full corpus along with its privacy-oriented annotation guidelines, evaluation scripts, and baseline models are available on: https://github.com/NorskRegnesentral/text-anonymization-benchmark. Ildikó Pilán, Pierre Lison, Lilja Øvrelid, Anthi Papadopoulou, David Sánchez 0001, Montserrat Batet |
Comput. Linguistics | 3 |
| 2021 | Structured Sentiment Analysis as Dependency Graph ParsingabstractJeremy Barnes, Robin Kurtz, Stephan Oepen, Lilja Øvrelid, Erik Velldal. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Jeremy Barnes 0001, Robin Kurtz, Stephan Oepen, Lilja Øvrelid, Erik Velldal |
ACL/IJCNLP (1) | 4 |
| 2021 | Anonymisation Models for Text Data: State of the art, Challenges and Future DirectionsabstractPierre Lison, Ildikó Pilán, David Sanchez, Montserrat Batet, Lilja Øvrelid. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Pierre Lison, Ildikó Pilán, David Sánchez 0001, Montserrat Batet, Lilja Øvrelid |
ACL/IJCNLP (1) | 5 |
| 2021 | If you've got it, flaunt it: Making the most of fine-grained sentiment annotationsabstractFine-grained sentiment analysis attempts to extract sentiment holders, targets and polar expressions and resolve the relationship between them, but progress has been hampered by the difficulty of annotation.Targeted sentiment analysis, on the other hand, is a more narrow task, focusing on extracting sentiment targets and classifying their polarity.In this paper, we explore whether incorporating holder and expression information can improve target extraction and classification and perform experiments on eight English datasets.We conclude that jointly predicting target and polarity BIO labels improves target extraction, and that augmenting the input text with gold expressions generally improves targeted polarity classification.This highlights the potential importance of annotating expressions for fine-grained sentiment datasets.At the same time, our results show that performance of current models for predicting polar expressions is poor, hampering the benefit of this information in practice. Jeremy Barnes 0001, Lilja Øvrelid, Erik Velldal |
EACL | 2 |
| 2021 | Improving sentiment analysis with multi-task learning of negationabstractAbstract Sentiment analysis is directly affected by compositional phenomena in language that act on the prior polarity of the words and phrases found in the text.Negationis the most prevalent of these phenomena, and in order to correctly predict sentiment, a classifier must be able to identify negation and disentangle the effect that its scope has on the final polarity of a text. This paper proposes a multi-task approach to explicitly incorporate information about negation in sentiment analysis, which we show outperforms learning negation implicitly in an end-to-end manner. We describe our approach, a cascading and hierarchical neural architecture with selective sharing of Long Short-term Memory layers, and show that explicitly training the model with negation as an auxiliary task helps improve the main task of sentiment analysis. The effect is demonstrated across several different standard English-language data sets for both tasks, and we analyze several aspects of our system related to its performance, varying types and amounts of input data and different multi-task setups. Jeremy Barnes 0001, Erik Velldal, Lilja Øvrelid |
Nat. Lang. Eng. | 3 |
| 2020 | A Tale of Three Parsers: Towards Diagnostic Evaluation for Meaning Representation ParsingabstractWe discuss methodological choices in contrastive and diagnostic evaluation in meaning representation parsing, i.e. mapping from natural language utterances to graph-based encodings of its semantic structure. Drawing inspiration from earlier work in syntactic dependency parsing, we transfer and refine several quantitative diagnosis techniques for use in the context of the 2019 shared task on Meaning Representation Parsing (MRP). As in parsing proper, moving evaluation from simple rooted trees to general graphs brings along its own range of challenges. Specifically, we seek to begin to shed light on relative strenghts and weaknesses in different broad families of parsing techniques. In addition to these theoretical reflections, we conduct a pilot experiment on a selection of top-performing MRP systems and one of the five meaning representation frameworks in the shared task. Empirical results suggest that the proposed methodology can be meaningfully applied to parsing into graph-structured target representations, uncovering hitherto unknown properties of the different systems that can inform future development and cross-fertilization across approaches. Maja Buljan, Joakim Nivre, Stephan Oepen, Lilja Øvrelid |
LREC | 4 |
| 2020 | NorNE: Annotating Named Entities for NorwegianabstractThis paper presents NorNE, a manually annotated corpus of named entities which extends the annotation of the existing Norwegian Dependency Treebank. Comprising both of the official standards of written Norwegian (Bokmål and Nynorsk), the corpus contains around 600,000 tokens and annotates a rich set of entity types including persons, organizations, locations, geo-political entities, products, and events, in addition to a class corresponding to nominals derived from names. We here present details on the annotation effort, guidelines, inter-annotator agreement and an experimental analysis of the corpus using a neural sequence labeling architecture. Fredrik Jørgensen, Tobias Aasmoe, Anne-Stine Ruud Husevåg, Lilja Øvrelid, Erik Velldal |
LREC | 4 |
| 2020 | A Fine-grained Sentiment Dataset for NorwegianabstractWe here introduce NoReC_fine, a dataset for fine-grained sentiment analysis in Norwegian, annotated with respect to polar expressions, targets and holders of opinion. The underlying texts are taken from a corpus of professionally authored reviews from multiple news-sources and across a wide variety of domains, including literature, games, music, products, movies and more. We here present a detailed description of this annotation effort. We provide an overview of the developed annotation guidelines, illustrated with examples and present an analysis of inter-annotator agreement. We also report the first experimental results on the dataset, intended as a preliminary benchmark for further experiments. Lilja Øvrelid, Petter Mæhlum, Jeremy Barnes 0001, Erik Velldal |
LREC | 1 |
| 2019 | THREAT: A Large Annotated Corpus for Detection of Violent ThreatsabstractUnderstanding, detecting, moderating and in extreme cases deleting hateful comments in online discussions and social media are well-known challenges. In this paper we present a dataset consisting of a total of around 30 000 sentences from around 10 000 YouTube comments. Each sentence is manually annotated as either being a violent threat or not. Violent threats is the most extreme form of hateful communication and is of particular importance from an online radicalization and national security perspective. This is the first publicly available dataset with such an annotation. The dataset can further be useful to develop automatic moderation tools or may even be useful from a social science perspective for analyzing the characteristics of online threats and how hateful discussions evolve. Hugo Hammer, Michael Riegler 0001, Lilja Øvrelid, Erik Velldal |
CBMI | 3 |
| 2018 | Diachronic word embeddings and semantic shifts: a surveyabstractRecent years have witnessed a surge of publications aimed at tracing temporal changes in lexical semantics using distributional methods, particularly prediction-based word embedding models. However, this vein of research lacks the cohesion, common terminology and shared practices of more established areas of natural language processing. In this paper, we survey the current state of academic research related to diachronic word embeddings and semantic shifts detection. We start with discussing the notion of semantic shifts, and then continue with an overview of the existing methods for tracing such time-related shifts with word embedding models. We propose several axes along which these methods can be compared, and outline the main challenges before this emerging subfield of NLP, as well as prospects and possible applications. Andrey Kutuzov, Lilja Øvrelid, Terrence Szymanski, Erik Velldal |
COLING | 2 |
| 2018 | Evaluation of Domain-specific Word Embeddings using Knowledge Resources
Farhad Nooralahzadeh, Lilja Øvrelid, Jan Tore Lønning |
LREC | 2 |
| 2018 | The LIA Treebank of Spoken Norwegian Dialects
Lilja Øvrelid, Andre Kåsen, Kristin Hagen, Anders Nøklestad, Per Erik Solberg, Janne Bondi Johannessen |
LREC | 1 |
| 2018 | NoReC: The Norwegian Review Corpus
Erik Velldal, Lilja Øvrelid, Eivind Alexander Bergem, Cathrine Stadsnes, Samia Touileb, Fredrik Jørgensen |
LREC | 2 |
| 2017 | Temporal dynamics of semantic relations in word embeddings: an application to predicting armed conflict participantsabstractThis paper deals with using word embedding models to trace the temporal dynamics of semantic relations between pairs of words.The set-up is similar to the well-known analogies task, but expanded with a time dimension.To this end, we apply incremental updating of the models with new training texts, including incremental vocabulary expansion, coupled with learned transformation matrices that let us map between members of the relation.The proposed approach is evaluated on the task of predicting insurgent armed groups based on geographical locations.The gold standard data for the time span 1994-2010 is extracted from the UCDP Armed Conflicts dataset.The results show that the method is feasible and outperforms the baselines, but also that important work still remains to be done. Andrey Kutuzov, Erik Velldal, Lilja Øvrelid |
EMNLP | 3 |
| 2016 | Redefining part-of-speech classes with distributional semantic modelsabstractThis paper studies how word embeddings trained on the British National Corpus interact with part of speech boundaries.Our work targets the Universal PoS tag set, which is currently actively being used for annotation of a range of languages.We experiment with training classifiers for predicting PoS tags for words based on their embeddings.The results show that the information about PoS affiliation contained in the distributional vectors allows us to discover groups of words with distributional patterns that differ from other words of the same part of speech.This data often reveals hidden inconsistencies of the annotation process or guidelines.At the same time, it supports the notion of 'soft' or 'graded' part of speech affiliations.Finally, we show that information about PoS is distributed among dozens of vector components, not limited to only one or two features. Andrey Kutuzov, Erik Velldal, Lilja Øvrelid |
CoNLL | 3 |
| 2016 | Universal Dependencies for Norwegian
Lilja Øvrelid, Petter Hohle |
LREC | 1 |
| 2014 | The Norwegian Dependency Treebank
Per Erik Solberg, Arne Skjærholt, Lilja Øvrelid, Kristin Hagen, Janne Bondi Johannessen |
LREC | 3 |
| 2012 | The WeSearch Corpus, Treebank, and Treecache - A Comprehensive Sample of User-Generated Content
Jonathon Read, Dan Flickinger, Rebecca Dridan, Stephan Oepen, Lilja Øvrelid |
LREC | 5 |
| 2012 | Speculation and Negation: Rules, Rankers, and the Role of SyntaxabstractThis article explores a combination of deep and shallow approaches to the problem of resolving the scope of speculation and negation within a sentence, specifically in the domain of biomedical research literature. The first part of the article focuses on speculation. After first showing how speculation cues can be accurately identified using a very simple classifier informed only by local lexical context, we go on to explore two different syntactic approaches to resolving the in-sentence scopes of these cues. Whereas one uses manually crafted rules operating over dependency structures, the other automatically learns a discriminative ranking function over nodes in constituent trees. We provide an in-depth error analysis and discussion of various linguistic properties characterizing the problem, and show that although both approaches perform well in isolation, even better results can be obtained by combining them, yielding the best published results to date on the CoNLL-2010 Shared Task data. The last part of the article describes how our speculation system is ported to also resolve the scope of negation. With only modest modifications to the initial design, the system obtains state-of-the-art results on this task also. Erik Velldal, Lilja Øvrelid, Jonathon Read, Stephan Oepen |
Comput. Linguistics | 2 |
| 2010 | Syntactic Scope Resolution in Uncertainty Analysis
Lilja Øvrelid, Erik Velldal, Stephan Oepen |
COLING | 1 |
| 2010 | Towards a Large Parallel Corpus of Cleft Constructions
Gerlof Bouma, Lilja Øvrelid, Jonas Kuhn |
LREC | 2 |
| 2010 | Training Parsers on Partial Trees: A Cross-language Comparison
Kathrin Spreyer, Lilja Øvrelid, Jonas Kuhn |
LREC | 2 |
| 2009 | Empirical Evaluations of Animacy Annotation
Lilja Øvrelid |
EACL | 1 |
| 2008 | Linguistic features in data-driven dependency parsing
Lilja Øvrelid |
CoNLL | 1 |
| 2006 | Towards Robust Animacy Classification Using Morphosyntactic Distributional Features
Lilja Øvrelid |
EACL | 1 |