Els Lefever

dblp:70/46 · DBLP profile ↗
← Back
31ranked-venue papers
9as first author
12since 2021 · last 2026
0000-0002-7755-0591ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 28 · 8 first-author · 11 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Counter-Hypothesis Generation: Towards Evaluating How LLMs Reason about Alternatives
abstract
Reasoning about alternatives is a fundamental component of human cognition and argumentation, yet it remains unclear whether large language models (LLMs) can coherently generate and assess them. This paper introduces Counter-Hypothesis Generation (CHG), a novel task for evaluating how LLMs construct plausible hypotheses when contextual information changes. Inspired by open-domain commonsense reasoning, where models infer and compare multiple explanations, CHG bridges commonsense and counterfactual reasoning by requiring models to generate hypotheses that remain logically consistent with modified premises. We present a test set annotated by a human expert and complemented with counter-hypotheses generated by OpenAI-o3 and DeepSeek-r1. Experimental results reveal that even advanced reasoning models exhibit notable limitations in counter-hypothesis generation.
Marzieh Abdolmaleki, Aaron Maladry, Véronique Hoste, Els Lefever
LREC4
2026 LoveHate: Stance Detection and Generation for Multiple Topics in User-generated Comments in Russian and English
abstract
This paper introduces LoveHate, a new multi-topic corpus of user-generated arguments in Russian, collected from the historical data of the debate platform lovehate.ru. The dataset contains nearly 19,000 posts spanning socially and politically relevant topics, each mapped to binary pro and con stances. We test multiple approaches to stance detection and stance generation across Russian and English data, including translated variants, using both classifier-based (Roberta, RuRoberta) and instruction-tuned generative (Llama, Qwen) models. Results demonstrate that language-specific pretraining yields the strongest performance for stance classification (F1 = 0.892 with RuRoberta), while multilingual generative models – when fine-tuned on sufficient data – can effectively generate stance in Russian without explicit Russian pretraining. Cross-domain experiments show that English datasets generalise better across corpora, whereas Russian data capture language- and culture-specific argumentation but are less effective for generalizable models. Generating topics remains a more challenging task for both Russian and English data. The dataset and accompanying results contribute to multilingual stance research and provide a valuable new resource for argument mining in Russian.
Natalia Evgrafova, Véronique Hoste, Els Lefever
LREC3
2026 TextLens & LeTTuce: Automated Corpus Annotation and Multilingual Tagging as a Service
abstract
We present TextLens, a web-based platform for automated linguistic annotation designed to lower technical barriers for researchers in digital humanities, linguistics and translation studies. Hosted by the Dutch Language Institute (INT), TextLens allows users to upload and annotate corpora in a variety of formats (.txt, .tsv, CoNLL-U, FoLiA, TEI, and NAF) using state-of-the-art NLP tools, without the need for local installation or computational resources. The platform supports multilingual data processing and provides a persistent dashboard for managing, monitoring and sharing annotation projects. Alongside this service, we introduce the LeTTuce-PoS Dataset, a new multilingual, manually annotated dataset for part-of-speech tagging in English, French, Dutch and German, covering multiple genres and offering a valuable resource to the research community. This paper also reports benchmark results for different PoS taggers (LeTs Preprocess, LeTTuce, spaCy and Stanza) on the dataset. Together, TextLens and the LeTTuce-PoS Dataset provide an accessible, scalable platform for high-quality annotation and a robust multilingual dataset that support comparable and reproducible research in multilingual contexts.
Cynthia Van Hee, Jonas Doumen, Vincent Prins, Pranaydeep Singh, Vincent Vandeghinste, Els Lefever
LREC6
2026 Towards Reliable Evaluation of Emotional Text Generation in LLMs: Human vs. Automatic Metrics
Sadegh Jafari, Els Lefever, Véronique Hoste
LREC2
2026 CorpusClues: Scalable Unsupervised Similarity Search for Historical Texts Using MinHash-LSH
abstract
CorpusClues is a prototype web-based platform for large-scale, unsupervised clustering of textual data, designed to address the specific challenges of historical corpora. It leverages the well-established computational techniques of MinHash and Locality-Sensitive Hashing (LSH) at the character level in order to detect structural similarities between texts even when exact patterns diverge. This approach makes CorpusClues robust to orthographic variation, such as historical spelling differences, while remaining fast and language-agnostic, capable of processing large and heterogeneous corpora without relying on language-specific models or preprocessing. Researchers can explore resulting clusters through interactive visualizations and exportable data, gaining access to patterns that would otherwise require the slow and uncertain process of manual collation. Evaluation against labeled gold standards shows that the system consistently produces high-quality clustering, accurately reconstructing relationships between texts despite substantial orthographic variation. By combining computational efficiency with user-friendly design, CorpusClues provides an accessible yet rigorous means of uncovering formulaicity and textual transmission at scale, opening new possibilities for the study of historical textual traditions.
Paulien Lemay, Klaas Bentein, Els Lefever
LREC3
2026 Exploring the Transfer of Irony Explanation Generation from English to Dutch
abstract
Explanation generation has gained increasing attention in the field of NLP because it makes the output of classification models more intuitively understandable for humans. This is particularly relevant for complex semantic tasks such as irony detection, where there may not be any explicit linguistic markers. Generative models have shown great potential for irony explanation in earlier work, but most studies have been limited to English. Since this is the highest-resourced language, these capabilities may not be available in languages other than English. To address this gap, this paper analyses the performance of generative models for explanation generation in Dutch, a lower-resourced but closely related language to English. Our work shows that larger proprietary models, like GPT-4, can generate meaningful explanations based on relevant world knowledge, whereas smaller open-source models still struggle to perform this task. Besides quality evaluation, we also analyse the limitations of these models, showing that GPT models struggle most with verbosity and that both open-source and proprietary models exhibit circular reasoning ("this text is ironic because the person expresses this in an ironic way”). Finally, open-source models struggle in particular for Dutch because they fail to produce the relevant world knowledge that is required to understand the irony. All models and data used for the experiments is available at iRONNIE on Hugging Face.
Aaron Maladry, Els Lefever, Cynthia Van Hee, Véronique Hoste
LREC2
2025 Evaluating Transformers for OCR Post-Correction in Early Modern Dutch Theatre
abstract
This paper explores the effectiveness of two types of transformer models — large generative models and sequence-to-sequence models — for automatically post-correcting Optical Character Recognition (OCR) output in early modern Dutch plays. To address the need for optimally aligned data, we create a parallel dataset based on the OCRed and ground truth versions from the EmDComF corpus using state-of-the-art alignment techniques. By combining character-based and semantic methods, we design and release a qualitative OCR-to-gold parallel dataset, selecting the alignment with the lowest Character Error Rate (CER) for all alignment pairs. We then fine-tune and evaluate five generative models and four sequence-to-sequence models on the OCR post-correction dataset. Results show that sequence-to-sequence models generally outperform generative models in this task, correcting more OCR errors and overgenerating and undergenerating less, with mBART as the best performing system.
Florian Debaene, Aaron Maladry, Els Lefever, Véronique Hoste
COLING3
2025 LDW: Label Divergence Weighting for Multimodal Sentiment Analysis
abstract
Multimodal sentiment analysis (MSA) traditionally assumes a unified emotional signal across modalities such as text, audio, and video. However, recent findings suggest that each modality may convey distinct affective perspectives. Motivated by perspectivist theories from cognitive science and natural language processing, this paper introduces Label Divergence Weighting (LDW), a modality-weighting strategy that dynamically adjusts trust in each modality based on its alignment with the overall sentiment label. The LDW framework leverages training-time supervision from the divergence between unimodal and multimodal sentiment annotations to learn modality reliability, and applies this learning to unseen data without requiring unimodal labels at inference time. Integrated into a multitask variant of the Tensor Fusion Network (MTFN), the proposed LDW-MTFN model achieves state-of-the-art results on both the acted Chinese dataset CH-SIMS and the authentic English dataset UniC. Extensive experiments and ablation studies demonstrate the robustness and generalizability of LDW across datasets with different cultural, linguistic, and environmental characteristics.
Quanqi Du, Loic De Langhe, Els Lefever, Véronique Hoste
ACM Multimedia3
2024 Human and System Perspectives on the Expression of Irony: An Analysis of Likelihood Labels and Rationales
abstract
In this paper, we examine the recognition of irony by both humans and automatic systems. We achieve this by enhancing the annotations of an English benchmark data set for irony detection. This enhancement involves a layer of human-annotated irony likelihood using a 7-point Likert scale that combines binary annotation with a confidence measure. Additionally, the annotators indicated the trigger words that led them to perceive the text as ironic, which leveraged necessary theoretical insights into the definition of irony and its various forms. By comparing these trigger word spans across annotators, we determine the extent to which humans agree on the source of irony in a text. Finally, we compare the human-annotated spans with sub-token importance attributions for fine-tuned transformers using Layer Integrated Gradients, a state-of-the-art interpretability metric. Our results indicate that our model achieves better performance on tweets that were annotated with high confidence and high agreement. Although automatic systems can identify trigger words with relative success, they still attribute a significant amount of their importance to the wrong tokens.
Aaron Maladry, Alessandra Teresa Cignarella, Els Lefever, Cynthia Van Hee, Véronique Hoste
LREC/COLING3
2024 At the Crossroad of Cuneiform and NLP: Challenges for Fine-grained Part-of-speech Tagging
abstract
The study of ancient Middle Eastern cultures is dominated by the vast number of cuneiform texts. Multiple languages and language families were expressed in cuneiform. The most dominant language written in cuneiform is the Semitic Akkadian, which is the focus of this paper. We are specifically focusing on letters written in the dialect used in modern-day Baghdad and south towards the Persian Gulf during the Old Babylonian period (c. 2000-1600 B.C.E.). The Akkadian language was rediscovered in the 19th century and is now being scrutinised by Natural Language Processing (NLP) methods. However, existing Akkadian text publications are not always suitable for digital editions. We therefore risk applying NLP methods onto renderings of Akkadian unfit for the purpose. In this paper we want to investigate the input material and try to initiate a discussion about best-practices in the crossroad where NLP meets cuneiform studies. Specifically, we want to question the use of pre-trained embeddings, sentence segmentation and the type of cuneiform input used to fine-tune language models for the task of fine-grained part-of-speech tagging. We examine the issues by theoretical and practical approaches in a way that we hope spurs discussions that are relevant for automatic processing of other ancient languages.
Gustav Ryberg Smidt, Els Lefever, Katrien de Graef
LREC/COLING2
2024 Lemmatisation of Medieval Greek: Against the Limits of Transformer's Capabilities?
abstract
This paper presents preliminary experiments for the lemmatisation of unedited, Byzantine Greek epigrams. This type of Greek is quite different from its classical ancestor, mostly because of its orthographic inconsistencies. Existing lemmatisation algorithms display an accuracy drop of around 30pp when tested on these Byzantine book epigrams. We conducted seven different lemmatisation experiments, which were either transformer-based or based on neural edit-trees. The best performing lemmatiser was a hybrid method combining transformer-based embeddings with a dictionary look-up. We compare our results with existing lemmatisers, and provide a detailed error analysis revealing why unedited, Byzantine Greek is so challenging for lemmatisation.
Colin Swaelens, Pranaydeep Singh, Ilse De Vos, Els Lefever
LREC/COLING4
2022 When the Student Becomes the Master: Learning Better and Smaller Monolingual Models from mBERT
abstract
In this research, we present pilot experiments to distil monolingual models from a jointly trained model for 102 languages (mBERT). We demonstrate that it is possible for the target language to outperform the original model, even with a basic distillation setup. We evaluate our methodology for 6 languages with varying amounts of resources and language families.
Pranaydeep Singh, Els Lefever
COLING2
2020 Identifying Cognates in English-Dutch and French-Dutch by means of Orthographic Information and Cross-lingual Word Embeddings
abstract
This paper investigates the validity of combining more traditional orthographic information with cross-lingual word embeddings to identify cognate pairs in English-Dutch and French-Dutch. In a first step, lists of potential cognate pairs in English-Dutch and French-Dutch are manually labelled. The resulting gold standard is used to train and evaluate a multi-layer perceptron that can distinguish cognates from non-cognates. Fifteen orthographic features capture string similarities between source and target words, while the cosine similarity between their word embeddings represents the semantic relation between these words. By adding domain-specific information to pretrained fastText embeddings, we are able to obtain good embeddings for words that did not yet have a pretrained embedding (e.g. Dutch compound nouns). These embeddings are then aligned in a cross-lingual vector space by exploiting their structural similarity (cf. adversarial learning). Our results indicate that although the classifier already achieves good results on the basis of orthographic information, the performance further improves by including semantic information in the form of cross-lingual word embeddings.
Els Lefever, Sofie Labat, Pranaydeep Singh
LREC1
2020 Uncovering the language of wine experts
abstract
Abstract Talking about odors and flavors is difficult for most people, yet experts appear to be able to convey critical information about wines in their reviews. This seems to be a contradiction, and wine expert descriptions are frequently received with criticism. Here, we propose a method for probing the language of wine reviews, and thus offer a means to enhance current vocabularies, and as a by-product question the general assumption that wine reviews are gibberish. By means of two different quantitative analyses—support vector machines for classification and Termhood analysis—on a corpus of online wine reviews, we tested whether wine reviews are written in a consistent manner, and thus may be considered informative; and whether reviews feature domain-specific language. First, a classification paradigm was trained on wine reviews from one set of authors for which the color, grape variety, and origin of a wine were known, and subsequently tested on data from a new author. This analysis revealed that, regardless of individual differences in vocabulary preferences, color and grape variety were predicted with high accuracy. Second, using Termhood as a measure of how words are used in wine reviews in a domain-specific manner compared to other genres in English, a list of 146 wine-specific terms was uncovered. These words were compared to existing lists of wine vocabulary that are currently used to train experts. Some overlap was observed, but there were also gaps revealed in the extant lists, suggesting these lists could be improved by our automatic analysis.
Ilja Croijmans, Iris Hendrickx, Els Lefever, Asifa Majid, Antal van den Bosch
Nat. Lang. Eng.3
2018 Smart Computer-Aided Translation Environment (SCATE): Highlights
abstract
We present the highlights of the now finished 4-year SCATE project. It was completed in February 2018 and funded by the Flemish Government IWT-SBO, project No. 130041.1
Vincent Vandeghinste, Tom Vanallemeersch, Bram Bulté, Liesbeth Augustinus, Frank Van Eynde, Joris Pelemans, Lyan Verwimp, Patrick Wambacq, Geert Heyman, Marie-Francine Moens, Iulianna Van der Lek-Ciudin, Frieda Steurs, Ayla Rigouts Terryn, Els Lefever, Arda Tezcan, Lieve Macken, Sven Coppers, Jens Brulmans, Jan Van den Bergh 0001, Kris Luyten, Karin Coninx
EAMT14
2018 MIsA: Multilingual "IsA" Extraction from Corpora
Stefano Faralli 0001, Els Lefever, Simone Paolo Ponzetto
LREC2
2018 Discovering the Language of Wine Reviews: A Text Mining Account
Els Lefever, Iris Hendrickx, Ilja Croijmans, Antal van den Bosch, Asifa Majid
LREC1
2018 A Gold Standard for Multilingual Automatic Term Extraction from Comparable Corpora: Term Structure and Translation Equivalents
Ayla Rigouts Terryn, Véronique Hoste, Els Lefever
LREC3
2018 We Usually Don't Like Going to the Dentist: Using Common Sense to Detect Irony on Twitter
abstract
Although common sense and connotative knowledge come naturally to most people, computers still struggle to perform well on tasks for which such extratextual information is required. Automatic approaches to sentiment analysis and irony detection have revealed that the lack of such world knowledge undermines classification performance. In this article, we therefore address the challenge of modeling implicit or prototypical sentiment in the framework of automatic irony detection. Starting from manually annotated connoted situation phrases (e.g., “flight delays,” “sitting the whole day at the doctor’s office”), we defined the implicit sentiment held towards such situations automatically by using both a lexico-semantic knowledge base and a data-driven method. We further investigate how such implicit sentiment information affects irony detection by assessing a state-of-the-art irony classifier before and after it is informed with implicit sentiment information.
Cynthia Van Hee, Els Lefever, Véronique Hoste
Comput. Linguistics2
2016 Monday mornings are my fave : ) #not Exploring the Automatic Recognition of Irony in English tweets
abstract
Recognising and understanding irony is crucial for the improvement natural language processing tasks including sentiment analysis. In this study, we describe the construction of an English Twitter corpus and its annotation for irony based on a newly developed fine-grained annotation scheme. We also explore the feasibility of automatic irony recognition by exploiting a varied set of features including lexical, syntactic, sentiment and semantic (Word2Vec) information. Experiments on a held-out test set show that our irony classifier benefits from this combined information, yielding an F1-score of 67.66%. When explicit hashtag information like #irony is included in the data, the system even obtains an F1-score of 92.77%. A qualitative analysis of the output reveals that recognising irony that results from a polarity clash appears to be (much) more feasible than recognising other forms of ironic utterances (e.g., descriptions of situational irony).
Cynthia Van Hee, Els Lefever, Véronique Hoste
COLING2
2016 Exploring the Realization of Irony in Twitter Data
Cynthia Van Hee, Els Lefever, Véronique Hoste
LREC2
2016 A Classification-based Approach to Economic Event Detection in Dutch News Text
Els Lefever, Véronique Hoste
LREC1
2014 Does size matter? A comparison of a large web corpus and a smaller focused corpus for medical term extraction
Klaar Vanopstal, Els Lefever, Véronique Hoste
AMIA2
2014 Evaluation of Automatic Hypernym Extraction from Technical Corpora in English and Dutch
Els Lefever, Marjan Van de Kauter, Véronique Hoste
LREC1
2013 Five Languages Are Better Than One: An Attempt to Bypass the Data Acquisition Bottleneck for WSD
Els Lefever, Véronique Hoste, Martine De Cock
CICLing (1)1
2012 Discovering Missing Wikipedia Inter-language Links by means of Cross-lingual Word Sense Disambiguation
Els Lefever, Véronique Hoste, Martine De Cock
LREC1
2010 Construction of a Benchmark Data Set for Cross-lingual Word Sense Disambiguation
Els Lefever, Véronique Hoste
LREC1
2010 Clustering web people search results using fuzzy ants
Els Lefever, Timur Fayruzov, Véronique Hoste, Martine De Cock
Inf. Sci.1
2009 Language-Independent Bilingual Terminology Extraction from a Multilingual Parallel Corpus
Els Lefever, Lieve Macken, Véronique Hoste
EACL1
2008 Linguistically-Based Sub-Sentential Alignment for Terminology Extraction from a Bilingual Automotive Corpus
Lieve Macken, Els Lefever, Véronique Hoste
COLING2
2008 Learning-based Detection of Scientific Terms in Patient Information
Véronique Hoste, Els Lefever, Klaar Vanopstal, Isabelle Delaere
LREC2