Erik Velldal

dblp:04/8218 · DBLP profile ↗
← Back
24ranked-venue papers
4as first author
9since 2021 · last 2026
0009-0008-6479-4512ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 23 · 4 first-author · 8 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Information extraction and text analysis · 88% Learning paradigms · 12% Language models and text generation · 1%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Computational social science and digital humanities · 100%

Topics — the 7 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Information extraction and text analysis › syntactic parsing › dependency parsing
graph-based dependency parsing
0.512021
Structured Sentiment Analysis as Dependency Graph Parsing · ACL/IJCNLP (1) 2021
Natural language and speech › Information extraction and text analysis
sentiment analysis
0.512021
Structured Sentiment Analysis as Dependency Graph Parsing · ACL/IJCNLP (1) 2021
Natural language and speech › Information extraction and text analysis › sentiment analysis
structured sentiment analysis
0.512021
Structured Sentiment Analysis as Dependency Graph Parsing · ACL/IJCNLP (1) 2021
Machine learning › Learning paradigms
multi-task learning
0.312018
Transfer and Multi-Task Learning for Noun-Noun Compound Interpretation · EMNLP 2018
Natural language and speech › Information extraction and text analysis › lexical semantics › multiword expression
noun compound interpretation
0.312018
Transfer and Multi-Task Learning for Noun-Noun Compound Interpretation · EMNLP 2018
Natural language and speech › Information extraction and text analysis › text classification
semantic classification
0.312018
Transfer and Multi-Task Learning for Noun-Noun Compound Interpretation · EMNLP 2018
Natural language and speech › Information extraction and text analysis
relation extraction
0.312017
Temporal dynamics of semantic relations in word embeddings: an application to predicting armed conflict participants · EMNLP 2017

Methods — techniques the papers use, named apart from their topics

word embeddings · 0.6transformation matrices · 0.6incremental learning · 0.6graph-based parsing · 0.5dependency parsing · 0.5transfer learning · 0.3neural classification · 0.3multi-task learning · 0.3statistical ranking · 0.1
YearPublicationVenuePosition
2026 MUC-4 Revisited: Document-level Event Analysis beyond Span-based Arguments
Helene Olsen, Erik Velldal, Lilja Øvrelid
LREC2
2026 Entity-Level Sentiment Analysis with Sentence Relevance Detection
Egil Rønningstad, Roman Klinger, Lilja Øvrelid, Erik Velldal
LREC4
2026 A Parallel Cross-Lingual Benchmark for Multimodal Idiomaticity Understanding
abstract
Potentially idiomatic expressions (PIEs) carry meanings inherently tied to the everyday experience of a given language community. As such, they constitute an interesting challenge for assessing the linguistic (and to some extent cultural) capabilities of NLP systems. In this paper, we present XMPIE, a parallel multilingual and multimodal dataset of potentially idiomatic expressions. The dataset, containing 34 languages and over ten thousand items, allows comparative analyses of idiomatic patterns among language-specific realisations and preferences in order to gather insights about shared cultural aspects. This parallel dataset allows evaluation of language model performance for a given PIE in different languages and whether idiomatic understanding in one language can be transferred to another. Moreover, the dataset supports the study of PIEs across textual and visual modalities, to measure to what extent PIE understanding in one modality transfers or implies in understanding in another modality (text vs. image). The data was created by language experts, with both textual and visual components crafted under multilingual guidelines, and each PIE is accompanied by five images representing a spectrum from idiomatic to literal meanings, including semantically related and random distractors. The result is a high-quality benchmark for evaluating multilingual and multimodal idiomatic language understanding.
Dilara Torunoglu-Selamet, Dogukan Arslan, Rodrigo Wilkens, Wei He 0017, Doruk Eryigit, Thomas Pickard, Adriana S. Pagano, Aline Villavicencio, Gülsen Eryigit, Ágnes Abuczki, Aida Cardoso, Alesia Lazarenka, Dina Almassova, Amália Mendes, Anna Kanellopoulou, Antoni Brosa-Rodríguez, Baiba Valkovska, Beata Wojtowicz, Bolette Pedersen, Carlos Manuel Hidalgo-Ternero, Chaya Liebeskind, Danka Jokic, Diego Alves, Eleni Triantafyllidi, Erik Velldal, Fred Philippy, Giedre Valunaite Oleskeviciene, Ieva Rizgeliene, Inguna Skadina, Irina Lobzhanidze, Isabell Stinessen Haugen, Jauza Akbar Krito, Jelena M. Markovic, Johanna Monti, Josue Alejandro Sauca, Kaja Dobrovoljc, Kingsley O. Ugwuanyi, Laura Rituma, Lilja Øvrelid, Maha Tufail Agro, Manzura Abjalova, Maria Chatzigrigoriou, María del Mar Sánchez Ramos, Marija Pendevska, Masoumeh Seyyedrezaei, Mehrnoush Shamsfard, Momina Ahsan, Muhammad Ahsan Riaz Khan, Nathalie Carmen Hau Norman, Nilay Erdem Ayyildiz, Nina Hosseini-Kivanani, Noémi Ligeti-Nagy, Numaan Naeem, Olha Kanishcheva, Olha Yatsyshyna, Daniil Orel, Petra Giommarelli, Petya Osenova, Radovan Garabík, Regina E. Semou, Rozane Rebechi, Salsabila Zahirah Pranida, Samia Touileb, Sanni Nimb, Sarvinoz Sharipova, Shahar Golan, Shaoxiong Ji, Sopuruchi Christian Aboh, Srdjan Sucur, Stella Markantonatou, Sussi Olsen, Vahideh Tajalli, Veronika Lipp, Voula Giouli, Yelda Yesildal Eraydin, Zahra Saaberi, Zhuohan Xie
LREC25
2025 ELOQUENT CLEF Shared Tasks for Evaluation of Generative Language Model Quality, 2025 Edition
Jussi Karlgren, Ekaterina Artemova, Ondrej Bojar, Vladislav Mikhailov, Magnus Sahlgren, Erik Velldal, Lilja Øvrelid
ECIR (5)6
2023 Measuring Normative and Descriptive Biases in Language Models Using Census Data
abstract
We investigate in this paper how distributions of occupations with respect to gender is reflected in pre-trained language models.Such distributions are not always aligned to normative ideals, nor do they necessarily reflect a descriptive assessment of reality.In this paper, we introduce an approach for measuring to what degree pre-trained language models are aligned to normative and descriptive occupational distributions.To this end, we use official demographic information about gender-occupation distributions provided by the national statistics agencies of France, Norway, United Kingdom, and the United States.We manually generate template-based sentences combining gendered pronouns and nouns with occupations, and subsequently probe a selection of ten language models covering the English, French, and Norwegian languages.The scoring system we introduce in this work is language independent, and can be used on any combination of template-based sentences, occupations, and languages.The approach could also be extended to other dimensions of national census data and other demographic variables.
Samia Touileb, Lilja Øvrelid, Erik Velldal
EACL3
2022 Entity-Level Sentiment Analysis (ELSA): An Exploratory Task Survey
abstract
This paper explores the task of identifying the overall sentiment expressed towards volitional entities (persons and organizations) in a document - what we refer to as Entity-Level Sentiment Analysis (ELSA). While identifying sentiment conveyed towards an entity is well researched for shorter texts like tweets, we find little to no research on this specific task for longer texts with multiple mentions and opinions towards the same entity. This lack of research would be understandable if ELSA can be derived from existing tasks and models. To assess this, we annotate a set of professional reviews for their overall sentiment towards each volitional entity in the text. We sample from data already annotated for document-level, sentence-level, and target-level sentiment in a multi-domain review corpus, and our results indicate that there is no single proxy task that provides this overall sentiment we seek for the entities at a satisfactory level of performance. We present a suite of experiments aiming to assess the contribution towards ELSA provided by document-, sentence-, and target-level sentiment analysis, and provide a discussion of their shortcomings. We show that sentiment in our dataset is expressed not only with an entity mention as target, but also towards targets with a sentiment-relevant relation to a volitional entity. In our data, these relations extend beyond anaphoric coreference resolution, and our findings call for further research of the topic. Finally, we also present a survey of previous relevant work.
Egil Rønningstad, Erik Velldal, Lilja Øvrelid
COLING2
2021 Structured Sentiment Analysis as Dependency Graph Parsing
abstract
Jeremy Barnes, Robin Kurtz, Stephan Oepen, Lilja Øvrelid, Erik Velldal. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Jeremy Barnes 0001, Robin Kurtz, Stephan Oepen, Lilja Øvrelid, Erik Velldal
ACL/IJCNLP (1)5
2021 If you've got it, flaunt it: Making the most of fine-grained sentiment annotations
abstract
Fine-grained sentiment analysis attempts to extract sentiment holders, targets and polar expressions and resolve the relationship between them, but progress has been hampered by the difficulty of annotation.Targeted sentiment analysis, on the other hand, is a more narrow task, focusing on extracting sentiment targets and classifying their polarity.In this paper, we explore whether incorporating holder and expression information can improve target extraction and classification and perform experiments on eight English datasets.We conclude that jointly predicting target and polarity BIO labels improves target extraction, and that augmenting the input text with gold expressions generally improves targeted polarity classification.This highlights the potential importance of annotating expressions for fine-grained sentiment datasets.At the same time, our results show that performance of current models for predicting polar expressions is poor, hampering the benefit of this information in practice.
Jeremy Barnes 0001, Lilja Øvrelid, Erik Velldal
EACL3
2021 Improving sentiment analysis with multi-task learning of negation
abstract
Abstract Sentiment analysis is directly affected by compositional phenomena in language that act on the prior polarity of the words and phrases found in the text.Negationis the most prevalent of these phenomena, and in order to correctly predict sentiment, a classifier must be able to identify negation and disentangle the effect that its scope has on the final polarity of a text. This paper proposes a multi-task approach to explicitly incorporate information about negation in sentiment analysis, which we show outperforms learning negation implicitly in an end-to-end manner. We describe our approach, a cascading and hierarchical neural architecture with selective sharing of Long Short-term Memory layers, and show that explicitly training the model with negation as an auxiliary task helps improve the main task of sentiment analysis. The effect is demonstrated across several different standard English-language data sets for both tasks, and we analyze several aspects of our system related to its performance, varying types and amounts of input data and different multi-task setups.
Jeremy Barnes 0001, Erik Velldal, Lilja Øvrelid
Nat. Lang. Eng.2
2020 NorNE: Annotating Named Entities for Norwegian
abstract
This paper presents NorNE, a manually annotated corpus of named entities which extends the annotation of the existing Norwegian Dependency Treebank. Comprising both of the official standards of written Norwegian (Bokmål and Nynorsk), the corpus contains around 600,000 tokens and annotates a rich set of entity types including persons, organizations, locations, geo-political entities, products, and events, in addition to a class corresponding to nominals derived from names. We here present details on the annotation effort, guidelines, inter-annotator agreement and an experimental analysis of the corpus using a neural sequence labeling architecture.
Fredrik Jørgensen, Tobias Aasmoe, Anne-Stine Ruud Husevåg, Lilja Øvrelid, Erik Velldal
LREC5
2020 A Fine-grained Sentiment Dataset for Norwegian
abstract
We here introduce NoReC_fine, a dataset for fine-grained sentiment analysis in Norwegian, annotated with respect to polar expressions, targets and holders of opinion. The underlying texts are taken from a corpus of professionally authored reviews from multiple news-sources and across a wide variety of domains, including literature, games, music, products, movies and more. We here present a detailed description of this annotation effort. We provide an overview of the developed annotation guidelines, illustrated with examples and present an analysis of inter-annotator agreement. We also report the first experimental results on the dataset, intended as a preliminary benchmark for further experiments.
Lilja Øvrelid, Petter Mæhlum, Jeremy Barnes 0001, Erik Velldal
LREC4
2019 THREAT: A Large Annotated Corpus for Detection of Violent Threats
abstract
Understanding, detecting, moderating and in extreme cases deleting hateful comments in online discussions and social media are well-known challenges. In this paper we present a dataset consisting of a total of around 30 000 sentences from around 10 000 YouTube comments. Each sentence is manually annotated as either being a violent threat or not. Violent threats is the most extreme form of hateful communication and is of particular importance from an online radicalization and national security perspective. This is the first publicly available dataset with such an annotation. The dataset can further be useful to develop automatic moderation tools or may even be useful from a social science perspective for analyzing the characteristics of online threats and how hateful discussions evolve.
Hugo Hammer, Michael Riegler 0001, Lilja Øvrelid, Erik Velldal
CBMI4
2018 Diachronic word embeddings and semantic shifts: a survey
abstract
Recent years have witnessed a surge of publications aimed at tracing temporal changes in lexical semantics using distributional methods, particularly prediction-based word embedding models. However, this vein of research lacks the cohesion, common terminology and shared practices of more established areas of natural language processing. In this paper, we survey the current state of academic research related to diachronic word embeddings and semantic shifts detection. We start with discussing the notion of semantic shifts, and then continue with an overview of the existing methods for tracing such time-related shifts with word embedding models. We propose several axes along which these methods can be compared, and outline the main challenges before this emerging subfield of NLP, as well as prospects and possible applications.
Andrey Kutuzov, Lilja Øvrelid, Terrence Szymanski, Erik Velldal
COLING4
2018 Transfer and Multi-Task Learning for Noun-Noun Compound Interpretation
abstract
In this paper, we empirically evaluate the utility of transfer and multi-task learning on a challenging semantic classification task: semantic interpretation of noun-noun compounds.Through a comprehensive series of experiments and in-depth error analysis, we show that transfer learning via parameter initialization and multi-task learning via parameter sharing can help a neural classification model generalize over a highly skewed distribution of relations.Further, we demonstrate how dual annotation with two distinct sets of relations over the same set of compounds can be exploited to improve the overall accuracy of a neural classifier and its F 1 scores on the less frequent, but more difficult relations.
Murhaf Fares, Stephan Oepen, Erik Velldal
EMNLP3
2018 NoReC: The Norwegian Review Corpus
Erik Velldal, Lilja Øvrelid, Eivind Alexander Bergem, Cathrine Stadsnes, Samia Touileb, Fredrik Jørgensen
LREC1
2017 Temporal dynamics of semantic relations in word embeddings: an application to predicting armed conflict participants
abstract
This paper deals with using word embedding models to trace the temporal dynamics of semantic relations between pairs of words.The set-up is similar to the well-known analogies task, but expanded with a time dimension.To this end, we apply incremental updating of the models with new training texts, including incremental vocabulary expansion, coupled with learned transformation matrices that let us map between members of the relation.The proposed approach is evaluated on the task of predicting insurgent armed groups based on geographical locations.The gold standard data for the time span 1994-2010 is extracted from the UCDP Armed Conflicts dataset.The results show that the method is feasible and outperforms the baselines, but also that important work still remains to be done.
Andrey Kutuzov, Erik Velldal, Lilja Øvrelid
EMNLP2
2016 Redefining part-of-speech classes with distributional semantic models
abstract
This paper studies how word embeddings trained on the British National Corpus interact with part of speech boundaries.Our work targets the Universal PoS tag set, which is currently actively being used for annotation of a range of languages.We experiment with training classifiers for predicting PoS tags for words based on their embeddings.The results show that the information about PoS affiliation contained in the distributional vectors allows us to discover groups of words with distributional patterns that differ from other words of the same part of speech.This data often reveals hidden inconsistencies of the annotation process or guidelines.At the same time, it supports the notion of 'soft' or 'graded' part of speech affiliations.Finally, we show that information about PoS is distributed among dozens of vector components, not limited to only one or two features.
Andrey Kutuzov, Erik Velldal, Lilja Øvrelid
CoNLL2
2016 A Corpus of Clinical Practice Guidelines Annotated with the Importance of Recommendations
Jonathon Read, Erik Velldal, Marc Cavazza, Gersende Georg
LREC2
2014 Off-Road LAF: Encoding and Processing Annotations in NLP Workflows
Emanuele Lapponi, Erik Velldal, Stephan Oepen, Rune Lain Knudsen
LREC2
2012 Speculation and Negation: Rules, Rankers, and the Role of Syntax
abstract
This article explores a combination of deep and shallow approaches to the problem of resolving the scope of speculation and negation within a sentence, specifically in the domain of biomedical research literature. The first part of the article focuses on speculation. After first showing how speculation cues can be accurately identified using a very simple classifier informed only by local lexical context, we go on to explore two different syntactic approaches to resolving the in-sentence scopes of these cues. Whereas one uses manually crafted rules operating over dependency structures, the other automatically learns a discriminative ranking function over nodes in constituent trees. We provide an in-depth error analysis and discussion of various linguistic properties characterizing the problem, and show that although both approaches perform well in isolation, even better results can be obtained by combining them, yielding the best published results to date on the CoNLL-2010 Shared Task data. The last part of the article describes how our speculation system is ported to also resolve the scope of negation. With only modest modifications to the initial design, the system obtains state-of-the-art results on this task also.
Erik Velldal, Lilja Øvrelid, Jonathon Read, Stephan Oepen
Comput. Linguistics1
2011 Deep open-source machine translation
Francis Bond, Stephan Oepen, Eric Nichols, Dan Flickinger, Erik Velldal, Petter Haugereid
Mach. Transl.5
2010 Syntactic Scope Resolution in Uncertainty Analysis
Lilja Øvrelid, Erik Velldal, Stephan Oepen
COLING2
2006 Statistical Ranking in Tactical Generation
Erik Velldal, Stephan Oepen
EMNLP1
2005 Maximum Entropy Models for Realization Ranking
abstract
In this paper we describe and evaluate different statistical models for the task of realization ranking, i.e. the problem of discriminating between competing surface realizations generated for a given input semantics. Three models are trained and tested; an n-gram language model, a discriminative maximum entropy model using structural features, and a combination of these two. Our realization component forms part of a larger, hybrid MT system.
Erik Velldal, Stephan Oepen
MTSummit1