VLDB 2026 Research / reviewers in the wild / expert
Luis Espinosa Anke
dblp:140/3490
· DBLP profile ↗
43ranked-venue papers
8as first author
22since 2021 · last 2025
0000-0001-6830-9176ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 39 · 8 first-author · 19 since 2021Databases, data management, data science and information retrieval · 7 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | GEAR: A Simple GENERATE, EMBED, AVERAGE AND RANK Approach for Unsupervised Reverse DictionaryabstractReverse Dictionary (RD) is the task of obtaining the most relevant word or set of words given a textual description or dictionary definition. Effective RD methods have applications in accessibility, translation or writing support systems. Moreover, in NLP research we find RD to be used to benchmark text encoders at various granularities, as it often requires word, definition and sentence embeddings. In this paper, we propose a simple approach to RD that leverages LLMs in combination with embedding models. Despite its simplicity, this approach outperforms supervised baselines in well studied RD datasets, while also showing less overfitting. We also conduct a number of experiments on different dictionaries and analyze how different styles, registers and target audiences impact the quality of RD systems. We conclude that, on average, untuned embeddings alone fare way below an LLM-only baseline (although they are competitive in highly technical dictionaries), but are crucial for boosting performance in combined methods. Fatemah Almeman, Luis Espinosa Anke |
COLING | 2 |
| 2025 | Automatic Extraction of Metaphoric Analogies from Literary Texts: Task Formulation, Dataset Construction, and EvaluationabstractExtracting metaphors and analogies from free text requires high-level reasoning abilities such as abstraction and language understanding. Our study focuses on the extraction of the concepts forming metaphoric analogies in literary texts. To this end, we construct a novel dataset in this domain with the help of domain experts. We compare the out-of-the-box ability of recent large language models (LLMs) to structure metaphoric mappings from fragments of texts containing rather explicit proportional analogies. The models are further evaluated on the generation of implicit elements of the analogy, which are indirectly suggested in the texts and inferred by human readers. The competitive results obtained by LLMs in our experiments are encouraging and open up new avenues such as automatically extracting analogies and metaphors from text instead of investing resources in domain experts to manually label data. Joanne Boisson, Zara Siddique, Hsuvas Borkakoty, Dimosthenis Antypas, Luis Espinosa Anke, José Camacho-Collados |
COLING | 5 |
| 2025 | TACTICAL: A Framework for Building Wikipedia-Derived Timelines of Atomic ChangesabstractThe well-known temporal misalignment in large language models (LLMs) emerges when they fail to recall temporal information. This is due to their training process, which happens without any explicit temporal grounding. To mitigate this issue, multiple approaches have been proposed, including fine-tuning on up-to-date data, retrieval augmented generation – where an LLM is directed to a recent dataset – or modifying an LLM’s knowledge via knowledge editing. Regardless of the method, however, the question of building datasets that accurately and faithfully reflect changes to events or entities remains open. Doing this in free text form and not only as triplets is desirable because LLMs benefit downstream from more context and can capture more nuanced relationships and cascading knowledge updates. Resources like Wikipedia can be leveraged for this thanks to their revision histories, which are expressed in free text and are both less biased and more comprehensive than knowledge graphs like Wikidata. In this paper, we propose TACTICAL, a methodology for creating timelines of Wikipedia entities and events, represented as revision pairs extracted from a wikititle’s timeline, and are categorized according to the atomicity of the changes affecting such entities or events. Our results suggest that LLMs struggle to recall event and entity timelines, even if they have seen them during pretraining. TACTICAL, on the other hand, proves to be an effective method for building temporally grounded datasets that are, in turn, effective tools for activating LLMs’ temporal knowledge. Hsuvas Borkakoty, Luis Espinosa Anke |
ECAI | 2 |
| 2025 | Grouping Entities with Shared Properties using Multi-Facet Prompting and Property EmbeddingsabstractMethods for learning taxonomies from data have been widely studied.We study a specific version of this task, called commonality identification, where only the set of entities is given and we need to find meaningful ways to group those entities.While LLMs should intuitively excel at this task, it is difficult to directly use such models in large domains.In this paper, we instead use LLMs to describe the different properties that are satisfied by each of the entities individually.We then use pretrained embeddings to cluster these properties, and finally group entities that have properties which belong to the same cluster.To achieve good results, it is paramount that the properties predicted by the LLM are sufficiently diverse.We find that this diversity can be improved by prompting the LLM to structure the predicted properties into different facets of knowledge.1 Amit Gajbhiye, Thomas Bailleux, Zied Bouraoui, Luis Espinosa Anke, Steven Schockaert |
EMNLP | 4 |
| 2024 | WordNet under Scrutiny: Dictionary Examples in the Era of Large Language ModelsabstractDictionary definitions play a prominent role in a wide range of NLP tasks, for instance by providing additional context about the meaning of rare and emerging terms. Many dictionaries also provide examples to illustrate the prototypical usage of words, which brings further opportunities for training or enriching NLP models. The intrinsic qualities of dictionaries, and related lexical resources such as glossaries and encyclopedias, are however still not well-understood. While there has been significant work on developing best practices, such guidance has been aimed at traditional usages of dictionaries (e.g. supporting language learners), and it is currently unclear how different quality aspects affect the NLP systems that rely on them. To address this issue, we compare WordNet, the most commonly used lexical resource in NLP, with a variety of dictionaries, as well as with examples that were generated by ChatGPT. Our analysis involves human judgments as well as automatic metrics. We furthermore study the quality of word embeddings derived from dictionary examples, as a proxy for downstream performance. We find that WordNet’s examples lead to lower-quality embeddings than those from the Oxford dictionary. Surprisingly, however, the ChatGPT generated examples were found to be most effective overall. Fatemah Almeman, Steven Schockaert, Luis Espinosa Anke |
LREC/COLING | 3 |
| 2024 | AMenDeD: Modelling Concepts by Aligning Mentions, Definitions and Decontextualised EmbeddingsabstractContextualised Language Models (LM) improve on traditional word embeddings by encoding the meaning of words in context. However, such models have also made it possible to learn high-quality decontextualised concept embeddings. Three main strategies for learning such embeddings have thus far been considered: (i) fine-tuning the LM to directly predict concept embeddings from the name of the concept itself, (ii) averaging contextualised representations of mentions of the concept in a corpus, and (iii) encoding definitions of the concept. As these strategies have complementary strengths and weaknesses, we propose to learn a unified embedding space in which all three types of representations can be integrated. We show that this allows us to outperform existing approaches in tasks such as ontology completion, which heavily depends on access to high-quality concept embeddings. We furthermore find that mentions and definitions are well-aligned in the resulting space, enabling tasks such as target sense verification, even without the need for any fine-tuning. Amit Gajbhiye, Zied Bouraoui, Luis Espinosa Anke, Steven Schockaert |
LREC/COLING | 3 |
| 2024 | LongEval: Longitudinal Evaluation of Model Performance at CLEF 2024
Rabab Alkhalifa, Hsuvas Borkakoty, Romain Deveaud, Alaa El-Ebshihy, Luis Espinosa Anke, Tobias Fink, Gabriela González Sáez, Petra Galuscáková, Lorraine Goeuriot, David Iommi, Maria Liakata, Harish Tayyar Madabushi, Pablo Medina-Alias, Philippe Mulhem, Florina Piroi, Martin Popel, Christophe Servan, Arkaitz Zubiaga |
ECIR (6) | 5 |
| 2024 | Who is better at math, Jenny or Jingzhen? Uncovering Stereotypes in Large Language ModelsabstractLarge language models (LLMs) have been shown to propagate and amplify harmful stereotypes, particularly those that disproportionately affect marginalised communities.To understand the effect of these stereotypes more comprehensively, we introduce GlobalBias, a dataset of 876k sentences incorporating 40 distinct gender-by-ethnicity groups alongside descriptors typically used in bias literature, which enables us to study a broad set of stereotypes from around the world.We use GlobalBias to directly probe a suite of LMs via perplexity, which we use as a proxy to determine how certain stereotypes are represented in the model's internal representations.Following this, we generate character profiles based on given names and evaluate the prevalence of stereotypes in model outputs.We find that the demographic groups associated with various stereotypes remain consistent across model likelihoods and model outputs.Furthermore, larger models consistently display higher levels of stereotypical outputs, even when explicitly instructed not to. Zara Siddique, Liam D. Turner, Luis Espinosa Anke |
EMNLP | 3 |
| 2023 | LongEval: Longitudinal Evaluation of Model Performance at CLEF 2023
Rabab Alkhalifa, Iman Munire Bilal, Hsuvas Borkakoty, José Camacho-Collados, Romain Deveaud, Alaa El-Ebshihy, Luis Espinosa Anke, Gabriela González Sáez, Petra Galuscáková, Lorraine Goeuriot, Elena Kochkina, Maria Liakata, Daniel Loureiro, Harish Tayyar Madabushi, Philippe Mulhem, Florina Piroi, Martin Popel, Christophe Servan, Arkaitz Zubiaga |
ECIR (3) | 7 |
| 2023 | Construction Artifacts in Metaphor Identification DatasetsabstractMetaphor identification aims at understanding whether a given expression is used figuratively in context.However, in this paper we show how existing metaphor identification datasets can be gamed by fully ignoring the potential metaphorical expression or the context in which it occurs.We test this hypothesis in a variety of datasets and settings, and show that metaphor identification systems based on language models without complete information can be competitive with those using the full context.This is due to the construction procedures to build such datasets, which introduce unwanted biases for positive and negative classes.Finally, we test the same hypothesis on datasets that are carefully sampled from natural corpora and where this bias is not present, making these datasets more challenging and reliable. Accuracy PrecisionRecall F1 Dataset Maj Def PME Mask Def PME Mask Def PME Mask Def PME Mask Psy CARD_N 50.0 87.5 44.5 85.9 90.5 42.8 86.7 83.9 33.9 85.1 87.0 35.4 85.6 CARD_V 50.0 83.9 42.9 82.9 87.2 40.8 83.1 80.8 36.9 83.4 83.3 38.3 83.0 JANK 66.7 85.3 51.1 84.2 81.2 45.1 82.2 75.0 28.4 68.5 76.5 45.1 74.2 TroFi 57.6 88.5 71.3 84.6 91.2 73.7 88.1 88.7 78.2 84.8 89.9 75.8 86.4 TSV_AN 50.3 89.3 59.4 79.2 91.3 62.2 80.2 86.9 52.0 78.0 89.1 55.8 78.9 GUT 53.6 98.3 65.4 95.4 98.9 66.8 96.0 97.9 70.6 95.3 98.4 68.6 95.7 NLP MOH 78.8 78.3 73.2 73.8 48.8 32.7 33.2 35.8 22.1 20.0 40.7 45.1 23.9 LLC 58.7 86.5 80.5 74.1 83.1 76.8 69.2 84.5 75.8 68.1 83.8 76.2 68.4 CHAK 66.7 69.7 74.4 66.7 76.0 80.2 66.7 79.8 81.8 100.0 77.8 81.0 80.0 NEU 56.0 76.0 56.0 72.0 81.3 58.8 79.9 79.8 66.0 74.5 77.6 60.4 72.9 DUNN 66.7 71.7 63.3 71.7 78.7 66.3 73.3 77.4 96.4 93.5 76.8 77.4 81.0 IDIX 51.4 93.8 85.3 84.6 93.3 83.1 84.4 93.7 86.8 83.0 93.5 84.9 83.7 PVC 65.1 85.9 76.9 80.8 89.0 79.8 83.2 89.4 86.6 88.6 89.2 82.8 85.7 VNC 78.5 96.0 93.5 87.0 97.1 95.0 89.8 97.9 96.8 94.1 97.5 95.9 91.9 PIE SE2013_ALL 59.5 91.5 83.2 85.3 93.5 84.5 86.3 92.3 87.9 89.3 92.8 86.2 87.8 SE2013_LEX 50.6 92.6 61.9 89.3 93.4 60.4 89.0 91.9 73.7 90.2 92.6 65.8 89.6 MAD 52.0 94.6 87.0 76.0 93.1 82.2 75.4 95.9 93.1 74.6 94.4 87.3 74.9 PIE 52.6 94.9 94.4 86.2 93.3 92.5 83.1 96.1 95.8 89.1 94.7 94.1 86.0 MAGPIE 74.7 96.1 93.3 86.7 97.4 95.4 88.6 97.3 95.5 94.3 97.4 95.5 91.4 VUAC_DO 50.4 75.5 77.4 63.0 74.1 77.1 62.1 79.1 78.4 68.0 76.5 77.7 64.9 VUAC VUAC_ST1 71.6 86.2 75.8 77.2 77.3 57.5 61.4 73.1 56.9 53.9 75.1 57.2 57.4 VUAC_ST2 84.3 92.3 86.5 84.7 76.7 56.4 51.4 73.0 61.1 39.9 74.8 58.6 45.0 VUAC_BO 50.5 85.4 63.3 76.0 84.6 63.4 75.4 86.9 64.6 77.8 85.8 64.0 76.6Table 4: Majority class accuracy (Maj) accuracy is shown in the first result column.Accuracy, precision, recall and F1 results for the metaphor class, averaged over 5 cross-validation folds, for the Default (Def), only PME (PME), and Masked settings on the random splits of metaphor identification datasets appear in the following columns. Accuracy Precision Recall F1 Dataset Maj Def PME MaskDef PME Mask Def PME Mask Def PME Mask CARD_N 50.0 89.7 50.0 86.7 90.1 50.0 88.6 89.5 59.4 84.4 89.7 54.0 86.4 Psy CARD_V 50.0 86.1 50.0 82.1 87.6 50.0 85.4 84.3 54.3 77.9 85.6 48.7 81.2 JANK 66.7 84.2 51.4 83.3 78.6 26.7 78.8 74.2 45.8 70.0 75.9 33.3 73.6 TroFi 57.6 82.3 62.9 78.5 85.1 63.4 81.2 85.0 84.1 82.9 84.7 72.2 81.6 TSV_AN 50.3 87.0 63.5 77.8 88.0 64.6 76.9 85.2 59.9 79.2 86.5 61.2 77.9 TSV_AN_L2 50.3 87.6 59.5 78.2 91.7 61.5 79.6 83.0 53.4 77.0 87.1 56.8 78.0 GUT 53.6 95.3 52.1 94.2 96.0 47.6 93.9 94.8 55.0 94.9 95.3 45.6 94.3 NLP GUT_L2 53.6 97.6 66.0 94.8 97.8 68.1 95.8 97.7 69.6 94.5 97.7 68.5 95.1 MOH 78.8 79.0 74.4 74.2 56.9 34.9 34.4 34.4 22.2 45.1 40.7 24.5 45.1 LLC 58.7 85.1 79.1 72.8 83.2 73.7 67.5 80.0 76.7 66.0 81.6 75.2 66.7 CHAK 66.7 64.3 73.2 63.7 74.6 81.0 65.6 71.1 77.9 95.3 72.2 79.3 77.6 NEU 56.0 76.0 50.0 77.0 84.4 43.0 82.2 73.6 80.0 81.0 73.8 54.5 79.1 DUNN 66.7 66.7 56.7 70.0 72.7 64.8 74.7 82.5 75.0 87.5 76.2 65.9 79.8 IDIX 51.4 75.5 63.2 74.1 76.5 66.3 74.5 77.4 62.5 74.9 75.1 62.5 73.5 PVC_V 65.1 69.5 59.2 67.3 73.4 66.8 68.9 78.9 76.3 80.0 75.3 67.4 72.8 VNC 78.5 84.3 73.4 81.8 90.4 85.6 86.4 90.5 80.6 91.2 89.5 81.7 88.5 PIE SE2013_ALL 59.5 79.2 49.2 79.1 81.5 57.8 78.7 82.3 50.6 85.2 79.9 48.7 80.4 SE2013_LEX 50.6 81.3 47.1 78.7 80.2 41.0 79.8 86.0 50.9 78.2 81.8 42.1 78.1 MAD 52.0 78.1 71.1 69.3 78.5 65.9 68.0 75.0 79.8 66.1 76.0 72.0 66.9 PIE 52.6 87.2 87.6 74.7 84.0 82.4 71.4 90.3 94.1 79.1 86.8 87.8 74.9 MAGPIE 73.7 90.2 84.9 83.4 94.4 88.4 88.4 92.2 91.5 89.2 93.3 89.9 88.8 VUAC_DO 57.6 74.5 75.0 59.6 78.5 78.7 64.7 76.7 77.5 65.7 77.6 78.0 65.2 VUAC VUAC_ST1 68.5 77.0 67.0 73.2 62.7 47.8 58.5 66.8 49.9 51.6 64.7 48.8 54.8 VUAC_ST2 82.6 88.3 83.3 84.2 68.8 53.2 56.1 60.0 36.3 41.7 64.1 43.2 47.9 VUAC_BO 53.4 82.2 65.5 73.0 80.7 64.9 71.5 87.7 77.0 82.1 84.1 77.0 76.5 Joanne Boisson, Luis Espinosa Anke, José Camacho-Collados |
EMNLP | 2 |
| 2023 | What do Deck Chairs and Sun Hats Have in Common? Uncovering Shared Properties in Large Concept VocabulariesabstractConcepts play a central role in many applications.This includes settings where concepts have to be modelled in the absence of sentence context.Previous work has therefore focused on distilling decontextualised concept embeddings from language models.But concepts can be modelled from different perspectives, whereas concept embeddings typically mostly capture taxonomic structure.To address this issue, we propose a strategy for identifying what different concepts, from a potentially large concept vocabulary, have in common with others.We then represent concepts in terms of the properties they share with the other concepts.To demonstrate the practical usefulness of this way of modelling concepts, we consider the task of ultra-fine entity typing, which is a challenging multi-label classification problem.We show that by augmenting the label set with shared properties, we can improve the performance of the state-of-the-art models for this task. 1 Amit Gajbhiye, Zied Bouraoui, Na Li 0018, Usashi Chatterjee, Luis Espinosa Anke, Steven Schockaert |
EMNLP | 5 |
| 2023 | Meemi: A simple method for post-processing and integrating cross-lingual word embeddingsabstractAbstract Word embeddings have become a standard resource in the toolset of any Natural Language Processing practitioner. While monolingual word embeddings encode information about words in the context of a particular language, cross-lingual embeddings define a multilingual space where word embeddings from two or more languages are integrated together. Current state-of-the-art approaches learn these embeddings by aligning two disjoint monolingual vector spaces through an orthogonal transformation which preserves the structure of the monolingual counterparts. In this work, we propose to apply an additional transformation after this initial alignment step, which aims to bring the vector representations of a given word and its translations closer to their average. Since this additional transformation is non-orthogonal, it also affects the structure of the monolingual spaces. We show that our approach both improves the integration of the monolingual spaces and the quality of the monolingual spaces themselves. Furthermore, because our transformation can be applied to an arbitrary number of languages, we are able to effectively obtain a truly multilingual space. The resulting (monolingual and multilingual) spaces show consistent gains over the current state-of-the-art in standard intrinsic tasks, namely dictionary induction and word similarity, as well as in extrinsic tasks such as cross-lingual hypernym discovery and cross-lingual natural language inference. Yerai Doval, José Camacho-Collados, Luis Espinosa Anke, Steven Schockaert |
Nat. Lang. Eng. | 3 |
| 2022 | Self-Supervised Intermediate Fine-Tuning of Biomedical Language Models for Interpreting Patient Case DescriptionsabstractInterpreting patient case descriptions has emerged as a challenging problem for biomedical NLP, where the aim is typically to predict diagnoses, to recommended treatments, or to answer questions about cases more generally. Previous work has found that biomedical language models often lack the knowledge that is needed for such tasks. In this paper, we aim to improve their performance through a self-supervised intermediate fine-tuning strategy based on PubMed abstracts. Our solution builds on the observation that many of these abstracts are case reports, and thus essentially patient case descriptions. As a general strategy, we propose to fine-tune biomedical language models on the task of predicting masked medical concepts from such abstracts. We find that the success of this strategy crucially depends on the selection of the medical concepts to be masked. By ensuring that these concepts are sufficiently salient, we can substantially boost the performance of biomedical language models, achieving state-of-the-art results on two benchmarks. Israa Alghanmi, Luis Espinosa Anke, Steven Schockaert |
COLING | 2 |
| 2022 | Modelling Commonsense Properties Using Pre-Trained Bi-EncodersabstractGrasping the commonsense properties of everyday concepts is an important prerequisite to language understanding. While contextualised language models are reportedly capable of predicting such commonsense properties with human-level accuracy, we argue that such results have been inflated because of the high similarity between training and test concepts. This means that models which capture concept similarity can perform well, even if they do not capture any knowledge of the commonsense properties themselves. In settings where there is no overlap between the properties that are considered during training and testing, we find that the empirical performance of standard language models drops dramatically. To address this, we study the possibility of fine-tuning language models to explicitly model concepts and their properties. In particular, we train separate concept and property encoders on two types of readily available data: extracted hyponym-hypernym pairs and generic sentences. Our experimental results show that the resulting encoders allow us to predict commonsense properties with much higher accuracy than is possible by directly fine-tuning language models. We also present experimental results for the related task of unsupervised hypernym discovery. Amit Gajbhiye, Luis Espinosa Anke, Steven Schockaert |
COLING | 2 |
| 2022 | TempoWiC: An Evaluation Benchmark for Detecting Meaning Shift in Social MediaabstractLanguage evolves over time, and word meaning changes accordingly. This is especially true in social media, since its dynamic nature leads to faster semantic shifts, making it challenging for NLP models to deal with new content and trends. However, the number of datasets and models that specifically address the dynamic nature of these social platforms is scarce. To bridge this gap, we present TempoWiC, a new benchmark especially aimed at accelerating research in social media-based meaning shift. Our results show that TempoWiC is a challenging benchmark, even for recently-released language models specialized in social media. Daniel Loureiro, Aminette D'Souza, Areej Nasser Muhajab, Isabella A. White, Gabriel Wong, Luis Espinosa Anke, Leonardo Neves, Francesco Barbieri, José Camacho-Collados |
COLING | 6 |
| 2022 | XLM-T: Multilingual Language Models in Twitter for Sentiment Analysis and BeyondabstractLanguage models are ubiquitous in current NLP, and their multilingual capacity has recently attracted considerable attention. However, current analyses have almost exclusively focused on (multilingual variants of) standard benchmarks, and have relied on clean pre-training and task-specific corpora as multilingual signals. In this paper, we introduce XLM-T, a model to train and evaluate multilingual language models in Twitter. In this paper we provide: (1) a new strong multilingual baseline consisting of an XLM-R (Conneau et al. 2020) model pre-trained on millions of tweets in over thirty languages, alongside starter code to subsequently fine-tune on a target task; and (2) a set of unified sentiment analysis Twitter datasets in eight different languages and a XLM-T model trained on this dataset. Francesco Barbieri, Luis Espinosa Anke, José Camacho-Collados |
LREC | 2 |
| 2022 | Pre-Training Language Models for Identifying Patronizing and Condescending Language: An AnalysisabstractPatronizing and Condescending Language (PCL) is a subtle but harmful type of discourse, yet the task of recognizing PCL remains under-studied by the NLP community. Recognizing PCL is challenging because of its subtle nature, because available datasets are limited in size, and because this task often relies on some form of commonsense knowledge. In this paper, we study to what extent PCL detection models can be improved by pre-training them on other, more established NLP tasks. We find that performance gains are indeed possible in this way, in particular when pre-training on tasks focusing on sentiment, harmful language and commonsense morality. In contrast, for tasks focusing on political speech and social justice, no or only very small improvements were witnessed. These findings improve our understanding of the nature of PCL. Carla Pérez-Almendros, Luis Espinosa Anke, Steven Schockaert |
LREC | 2 |
| 2022 | Sentence Selection Strategies for Distilling Word Embeddings from BERTabstractMany applications crucially rely on the availability of high-quality word vectors. To learn such representations, several strategies based on language models have been proposed in recent years. While effective, these methods typically rely on a large number of contextualised vectors for each word, which makes them impractical. In this paper, we investigate whether similar results can be obtained when only a few contextualised representations of each word can be used. To this end, we analyse a range of strategies for selecting the most informative sentences. Our results show that with a careful selection strategy, high-quality word vectors can be learned from as few as 5 to 10 sentences. Zied Bouraoui, Luis Espinosa Anke, Steven Schockaert |
LREC | 3 |
| 2022 | Interpreting Patient Descriptions using Distantly Supervised Similar Case RetrievalabstractBiomedical natural language processing often involves the interpretation of patient descriptions, for instance for diagnosis or for recommending treatments. Current methods, based on biomedical language models, have been found to struggle with such tasks. Moreover, retrieval augmented strategies have only had limited success, as it is rare to find sentences which express the exact type of knowledge that is needed for interpreting a given patient description. For this reason, rather than attempting to retrieve explicit medical knowledge, we instead propose to rely on a nearest neighbour strategy. First, we retrieve text passages that are similar to the given patient description, and are thus likely to describe patients in similar situations, while also mentioning some hypothesis (e.g.\ a possible diagnosis of the patient). We then judge the likelihood of the hypothesis based on the similarity of the retrieved passages. Identifying similar cases is challenging, however, as descriptions of similar patients may superficially look rather different, among others because they often contain an abundance of irrelevant details. To address this challenge, we propose a strategy that relies on a distantly supervised cross-encoder. Despite its conceptual simplicity, we find this strategy to be effective in practice. Israa Alghanmi, Luis Espinosa Anke, Steven Schockaert |
SIGIR | 2 |
| 2021 | BERT is to NLP what AlexNet is to CV: Can Pre-Trained Language Models Identify Analogies?abstractAsahi Ushio, Luis Espinosa Anke, Steven Schockaert, Jose Camacho-Collados. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Asahi Ushio, Luis Espinosa Anke, Steven Schockaert, José Camacho-Collados |
ACL/IJCNLP (1) | 2 |
| 2021 | Evaluating language models for the retrieval and categorization of lexical collocationsabstractComunicació presentada a: EACL 2021 celebrat del 19 a 23 d'abril de 2021 en línia. Luis Espinosa Anke, Joan Codina, Leo Wanner |
EACL | 1 |
| 2021 | Modelling General Properties of Nouns by Selectively Averaging Contextualised EmbeddingsabstractWhile the success of pre-trained language models has largely eliminated the need for high-quality static word vectors in many NLP applications, static word vectors continue to play an important role in tasks where word meaning needs to be modelled in the absence of linguistic context. In this paper, we explore how the contextualised embeddings predicted by BERT can be used to produce high-quality word vectors for such domains, in particular related to knowledge base completion, where our focus is on capturing the semantic properties of nouns. We find that a simple strategy of averaging the contextualised embeddings of masked word mentions leads to vectors that outperform the static word vectors learned by BERT, as well as those from standard word embedding models, in property induction tasks. We notice in particular that masking target words is critical to achieve this strong performance, as the resulting vectors focus less on idiosyncratic properties and more on general semantic properties. Inspired by this view, we propose a filtering strategy which is aimed at removing the most idiosyncratic mention vectors, allowing us to obtain further performance gains in property induction. Na Li 0018, Zied Bouraoui, José Camacho-Collados, Luis Espinosa Anke, Qing Gu 0001, Steven Schockaert |
IJCAI | 4 |
| 2020 | Modelling Semantic Categories Using Conceptual NeighborhoodabstractWhile many methods for learning vector space embeddings have been proposed in the field of Natural Language Processing, these methods typically do not distinguish between categories and individuals. Intuitively, if individuals are represented as vectors, we can think of categories as (soft) regions in the embedding space. Unfortunately, meaningful regions can be difficult to estimate, especially since we often have few examples of individuals that belong to a given category. To address this issue, we rely on the fact that different categories are often highly interdependent. In particular, categories often have conceptual neighbors, which are disjoint from but closely related to the given category (e.g. fruit and vegetable). Our hypothesis is that more accurate category representations can be learned by relying on the assumption that the regions representing such conceptual neighbors should be adjacent in the embedding space. We propose a simple method for identifying conceptual neighbors and then show that incorporating these conceptual neighbors indeed leads to more accurate region based representations. Zied Bouraoui, José Camacho-Collados, Luis Espinosa Anke, Steven Schockaert |
AAAI | 3 |
| 2020 | Don't Patronize Me! An Annotated Dataset with Patronizing and Condescending Language towards Vulnerable CommunitiesabstractIn this paper, we introduce a new annotated dataset which is aimed at supporting the development of NLP models to identify and categorize language that is patronizing or condescending towards vulnerable communities (e.g. refugees, homeless people, poor families). While the prevalence of such language in the general media has long been shown to have harmful effects, it differs from other types of harmful language, in that it is generally used unconsciously and with good intentions. We furthermore believe that the often subtle nature of patronizing and condescending language (PCL) presents an interesting technical challenge for the NLP community. Our analysis of the proposed dataset shows that identifying PCL is hard for standard NLP models, with language models such as BERT achieving the best results. Carla Pérez-Almendros, Luis Espinosa Anke, Steven Schockaert |
COLING | 2 |
| 2020 | Capturing Word Order in Averaging Based Sentence EmbeddingsabstractOne of the most remarkable findings in the literature on sentence embeddings has been that simple word vector averaging can compete with state-of-the-art models in many tasks. While counter-intuitive, a convincing explanation has been provided by Arora et al., who showed that the bag-of-words representation of a sentence can be recovered from its word vector average with almost perfect accuracy. Beyond word vector averaging, however, most sentence embedding models are essentially black boxes: while there is abundant empirical evidence about their strengths and weaknesses, it is not clear why and how different embedding strategies are able to capture particular properties of sentences. In this paper, we focus in particular on how sentence embedding models are able to capture word order. For instance, it seems intuitively puzzling that simple LSTM autoencoders are able to learn sentence vectors from which the original sentence can be reconstructed almost perfectly. With the aim of elucidating this phenomenon, we show that to capture word order, it is in fact sufficient to supplement standard word vector averages with averages of bigram and trigram vectors. To this end, we first study the problem of reconstructing bags-of-bigrams, focusing in particular on how suitable bigram vectors should be encoded. We then show that LSTMs are capable, in principle, of learning our proposed sentence embeddings. Empirically, we find that our embeddings outperform those learned by LSTM autoencoders on the task of sentence reconstruction, while needing almost no training data. Jae Hee Lee 0001, José Camacho-Collados, Luis Espinosa Anke, Steven Schockaert |
ECAI | 3 |
| 2020 | Learning Cross-Lingual Word Embeddings from Twitter via Distant Supervision
José Camacho-Collados, Yerai Doval, Eugenio Martínez-Cámara, Luis Espinosa Anke, Francesco Barbieri, Steven Schockaert |
ICWSM | 4 |
| 2020 | On the Robustness of Unsupervised and Semi-supervised Cross-lingual Word Embedding LearningabstractCross-lingual word embeddings are vector representations of words in different languages where words with similar meaning are represented by similar vectors, regardless of the language. Recent developments which construct these embeddings by aligning monolingual spaces have shown that accurate alignments can be obtained with little or no supervision, which usually comes in the form of bilingual dictionaries. However, the focus has been on a particular controlled scenario for evaluation, and there is no strong evidence on how current state-of-the-art systems would fare with noisy text or for language pairs with major linguistic differences. In this paper we present an extensive evaluation over multiple cross-lingual embedding models, analyzing their strengths and limitations with respect to different variables such as target language, training corpora and amount of supervision. Our conclusions put in doubt the view that high-quality cross-lingual embeddings can always be learned without much supervision. Yerai Doval, José Camacho-Collados, Luis Espinosa Anke, Steven Schockaert |
LREC | 3 |
| 2019 | Collocation Classification with Unsupervised Relation VectorsabstractLexical relation classification is the task of predicting whether a certain relation holds between a given pair of words.In this paper, we explore to which extent the current distributional landscape based on word embeddings provides a suitable basis for classification of collocations, i.e., pairs of words between which idiosyncratic lexical relations hold.First, we introduce a novel dataset with collocations categorized according to lexical functions.Second, we conduct experiments on a subset of this benchmark, comparing it in particular to the well known DiffVec dataset.In these experiments, in addition to simple word vector arithmetic operations, we also investigate the role of unsupervised relation vectors as a complementary input.While these relation vectors indeed help, we also show that lexical function classification poses a greater challenge than the syntactic and semantic relations that are typically used for benchmarks in the literature. Luis Espinosa Anke, Steven Schockaert, Leo Wanner |
ACL (1) | 1 |
| 2019 | Relational Word EmbeddingsabstractWhile word embeddings have been shown to implicitly encode various forms of attributional knowledge, the extent to which they capture relational information is far more limited.In previous work, this limitation has been addressed by incorporating relational knowledge from external knowledge bases when learning the word embedding.Such strategies may not be optimal, however, as they are limited by the coverage of available resources and conflate similarity with other forms of relatedness.As an alternative, in this paper we propose to encode relational knowledge in a separate word embedding, which is aimed to be complementary to a given standard word embedding.This relational word embedding is still learned from co-occurrence statistics, and can thus be used even when no external knowledge base is available.Our analysis shows that relational word vectors do indeed capture information that is complementary to what is encoded in standard word embeddings. José Camacho-Collados, Luis Espinosa Anke, Steven Schockaert |
ACL (1) | 2 |
| 2019 | A Latent Variable Model for Learning Distributional Relation VectorsabstractRecently a number of unsupervised approaches have been proposed for learning vectors that capture the relationship between two words. Inspired by word embedding models, these approaches rely on co-occurrence statistics that are obtained from sentences in which the two target words appear. However, the number of such sentences is often quite small, and most of the words that occur in them are not relevant for characterizing the considered relationship. As a result, standard co-occurrence statistics typically lead to noisy relation vectors. To address this issue, we propose a latent variable model that aims to explicitly determine what words from the given sentences best characterize the relationship between the two target words. Relation vectors then correspond to the parameters of a simple unigram language model which is estimated from these words. José Camacho-Collados, Luis Espinosa Anke, Shoaib Jameel, Steven Schockaert |
IJCAI | 2 |
| 2018 | SeVeN: Augmenting Word Embeddings with Unsupervised Relation VectorsabstractWe present SeVeN (Semantic Vector Networks), a hybrid resource that encodes relationships between words in the form of a graph. Different from traditional semantic networks, these relations are represented as vectors in a continuous vector space. We propose a simple pipeline for learning such relation vectors, which is based on word vector averaging in combination with an ad hoc autoencoder. We show that by explicitly encoding relational information in a dedicated vector space we can capture aspects of word meaning that are complementary to what is captured by word embeddings. For example, by examining clusters of relation vectors, we observe that relational similarities can be identified at a more abstract level than with traditional word vector differences. Finally, we test the effectiveness of semantic vector networks in two tasks: measuring word similarity and neural text categorization. SeVeN is available at bitbucket.org/luisespinosa/seven. Luis Espinosa Anke, Steven Schockaert |
COLING | 1 |
| 2018 | Interpretable Emoji Prediction via Label-Wise Attention LSTMsabstractHuman language has evolved towards newer forms of communication such as social media, where emojis (i.e., ideograms bearing a visual meaning) play a key role.While there is an increasing body of work aimed at the computational modeling of emoji semantics, there is currently little understanding about what makes a computational model represent or predict a given emoji in a certain way.In this paper we propose a label-wise attention mechanism with which we attempt to better understand the nuances underlying emoji prediction.In addition to advantages in terms of interpretability, we show that our proposed architecture improves over standard baselines in emoji prediction, and does particularly well when predicting infrequent emojis. Francesco Barbieri, Luis Espinosa Anke, José Camacho-Collados, Steven Schockaert, Horacio Saggion |
EMNLP | 2 |
| 2018 | Improving Cross-Lingual Word Embeddings by Meeting in the MiddleabstractCross-lingual word embeddings are becoming increasingly important in multilingual NLP.Recently, it has been shown that these embeddings can be effectively learned by aligning two disjoint monolingual vector spaces through linear transformations, using no more than a small bilingual dictionary as supervision.In this work, we propose to apply an additional transformation after the initial alignment step, which moves cross-lingual synonyms towards a middle point between them.By applying this transformation our aim is to obtain a better cross-lingual integration of the vector spaces.In addition, and perhaps surprisingly, the monolingual spaces also improve by this transformation.This is in contrast to the original alignment, which is typically learned such that the structure of the monolingual spaces is preserved.Our experiments confirm that the resulting cross-lingual embeddings outperform state-of-the-art models in both monolingual and cross-lingual evaluation tasks. Yerai Doval, José Camacho-Collados, Luis Espinosa Anke, Steven Schockaert |
EMNLP | 3 |
| 2016 | ExTaSem! Extending, Taxonomizing and Semantifying Domain TerminologiesabstractWe introduce ExTaSem!, a novel approach for the automatic learning of lexical taxonomies from domain terminologies. First, we exploit a very large semantic network to collect housands of in-domain textual definitions. Second, we extract (hyponym, hypernym) pairs from each definition with a CRF-based algorithm trained on manually-validated data. Finally, we introduce a graph induction procedure which constructs a full-fledged taxonomy where each edge is weighted according to its domain pertinence. ExTaSem! achieves state-of-the-art results in the following taxonomy evaluation experiments: (1) Hypernym discovery, (2) Reconstructing gold standard taxonomies, and (3) Taxonomy quality according to structural measures. We release weighted taxonomies for six domains for the use and scrutiny of the community. Luis Espinosa Anke, Horacio Saggion, Francesco Ronzano, Roberto Navigli |
AAAI | 1 |
| 2016 | Extending WordNet with Fine-Grained Collocational Information via Supervised Distributional LearningabstractWordNet is probably the best known lexical resource in Natural Language Processing. While it is widely regarded as a high quality repository of concepts and semantic relations, updating and extending it manually is costly. One important type of relation which could potentially add enormous value to WordNet is the inclusion of collocational information, which is paramount in tasks such as Machine Translation, Natural Language Generation and Second Language Learning. In this paper, we present ColWordNet (CWN), an extended WordNet version with fine-grained collocational information, automatically introduced thanks to a method exploiting linear relations between analogous sense-level embeddings spaces. We perform both intrinsic and extrinsic evaluations, and release CWN for the use and scrutiny of the community. Luis Espinosa Anke, José Camacho-Collados, Sara Rodríguez-Fernández, Horacio Saggion, Leo Wanner |
COLING | 1 |
| 2016 | Supervised Distributional Hypernym Discovery via Domain AdaptationabstractComunicació presentada a la Conference on Empirical Methods in Natural Language Processing celebrada els dies 1 a 5 de novembre de 2016 a Austin, Texas. Luis Espinosa Anke, José Camacho-Collados, Claudio Delli Bovi, Horacio Saggion |
EMNLP | 1 |
| 2016 | ELMD: An Automatically Generated Entity Linking Gold Standard Dataset in the Music Domain
Sergio Oramas, Luis Espinosa Anke, Mohamed Sordo, Horacio Saggion, Xavier Serra |
LREC | 2 |
| 2016 | Example-based Acquisition of Fine-grained Collocation Resources
Sara Rodríguez-Fernández, Roberto Carlini, Luis Espinosa Anke, Leo Wanner |
LREC | 3 |
| 2016 | Information extraction for knowledge base construction in the music domain
Sergio Oramas, Luis Espinosa Anke, Mohamed Sordo, Horacio Saggion, Xavier Serra |
Data Knowl. Eng. | 2 |
| 2015 | Hypernym Extraction: Combining Machine-Learning and Dependency Grammar
Luis Espinosa Anke, Francesco Ronzano, Horacio Saggion |
CICLing (1) | 1 |
| 2015 | Knowledge Base Unification via Sense Embeddings and DisambiguationabstractPaper presented at The 2015 Conference on Empirical Methods in Natural Language; 2015 Sept 17-21; Lisbon, Portugal. Claudio Delli Bovi, Luis Espinosa Anke, Roberto Navigli |
EMNLP | 2 |
| 2015 | Extracting Relations from Unstructured Text Sources for Music Recommendation
Mohamed Sordo, Sergio Oramas, Luis Espinosa Anke |
NLDB | 3 |
| 2014 | Applying Dependency Relations to Definition Extraction
Luis Espinosa Anke, Horacio Saggion |
NLDB | 1 |