José Camacho-Collados

dblp:165/0790 · DBLP profile ↗
← Back
60ranked-venue papers
8as first author
34since 2021 · last 2026
0000-0003-1618-7239ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 58 · 7 first-author · 33 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 False Friends or Cognates? A Cross-lingual Semantic Ambiguity Evaluation for Galician, Portuguese and Spanish
abstract
The linguistic proximity between Galician, Portuguese, and Spanish results in a lexical overlap that often conceals semantic interference.This is particularly evident in false friends, posing a challenge for NLP systems.In this work, we assess whether state-of-the-art language models can identify and process false friends among these languages.We introduce six cross-lingual datasets -created manually or using semi-automatic methods, with all instances being carefully verified-covering cognates and false friends.We evaluate a broad range of encoder and decoder models of varying sizes via zero-shot and few-shot settings.Our results highlight the challenging nature of the task, but also show the clear progress made by LLMs in recent years, particularly those of a larger size, with smaller language models struggling on the task.Notably, unlike other tasks where language distance poses additional challenges, we find that linguistic proximity itself introduces errors: closely related language pairs tend to perform worse, reflecting the challenge of semantic discrimination due to lexical overlap.
Marta Vázquez Abuín, José Camacho-Collados, Marcos García 0001
ACL (1)2
2026 Understanding LLM Performance Degradation in Multi-Instance Processing: The Roles of Instance Count and Context Length
abstract
Users often rely on Large Language Models (LLMs) for processing multiple documents or performing analysis over a number of instances.For example, analysing the overall sentiment of a number of movie reviews requires an LLM to process the sentiment of each review individually in order to provide a final aggregated answer.While LLM performance on such individual tasks is generally high, there has been little research on how LLMs perform when dealing with multi-instance inputs.In this paper, we perform a comprehensive evaluation of the multi-instance processing (MIP) ability of LLMs for tasks in which they excel individually.The results show that all LLMs follow a pattern of slight performance degradation for small numbers of instances (≈20-100), followed by a performance collapse on larger instance counts.Crucially, our analysis shows that while context length is associated with this degradation, the number of instances has a stronger effect on the final results.This finding suggests that when optimising LLM performance for MIP, attention should be paid to both context length and, in particular, instance count. 1
Jingxuan Chen, Mohammad Taher Pilehvar, José Camacho-Collados
ACL (1)3
2026 Report-based Recommendations for Policy Making and Agency Operations: Dataset and LLM Evaluation
Aleksandra Edwards, Thomas Edwards, José Camacho-Collados, Alun D. Preece
LREC3
2026 Enhancing and Evaluating Tabular Models on the Fly via Synthetic Question-Answer Generation
Jorge Osés Grijalba, Eugenio Martínez-Cámara, Luis Alfonso Ureñ López, José Camacho-Collados
LREC4
2025 Analysing Zero-Shot Readability-Controlled Sentence Simplification
abstract
Readability-controlled text simplification (RCTS) rewrites texts to lower readability levels while preserving their meaning. RCTS models often depend on parallel corpora with readability annotations on both source and target sides. Such datasets are scarce and difficult to curate, especially at the sentence level. To reduce reliance on parallel data, we explore using instruction-tuned large language models for zero-shot RCTS. Through automatic and manual evaluations, we examine: (1) how different types of contextual information affect a model’s ability to generate sentences with the desired readability, and (2) the trade-off between achieving target readability and preserving meaning. Results show that all tested models struggle to simplify sentences (especially to the lowest levels) due to models’ limitations and characteristics of the source sentences that impede adequate rewriting. Our experiments also highlight the need for better automatic evaluation metrics tailored to RCTS, as standard ones often misinterpret common simplification operations, and inaccurately assess readability and meaning preservation.
Abdullah Barayan, José Camacho-Collados, Fernando Alva-Manchego
COLING2
2025 Automatic Extraction of Metaphoric Analogies from Literary Texts: Task Formulation, Dataset Construction, and Evaluation
abstract
Extracting metaphors and analogies from free text requires high-level reasoning abilities such as abstraction and language understanding. Our study focuses on the extraction of the concepts forming metaphoric analogies in literary texts. To this end, we construct a novel dataset in this domain with the help of domain experts. We compare the out-of-the-box ability of recent large language models (LLMs) to structure metaphoric mappings from fragments of texts containing rather explicit proportional analogies. The models are further evaluated on the generation of implicit elements of the analogy, which are indirectly suggested in the texts and inferred by human readers. The competitive results obtained by LLMs in our experiments are encouraging and open up new avenues such as automatically extracting analogies and metaphors from text instead of investing resources in domain experts to manually label data.
Joanne Boisson, Zara Siddique, Hsuvas Borkakoty, Dimosthenis Antypas, Luis Espinosa Anke, José Camacho-Collados
COLING6
2025 Morables: A Benchmark for Assessing Abstract Moral Reasoning in LLMs with Fables
abstract
As LLMs excel on standard reading comprehension benchmarks, attention is shifting toward evaluating their capacity for complex abstract reasoning and inference.Literaturebased benchmarks, with their rich narrative and moral depth, provide a compelling framework for evaluating such deeper comprehension skills.Here, we present MORABLES, a humanverified benchmark built from fables and short stories drawn from historical literature.The main task is structured as multiple-choice questions targeting moral inference, with carefully crafted distractors that challenge models to go beyond shallow, extractive question answering.To further stress-test model robustness, we introduce adversarial variants designed to surface LLM vulnerabilities and shortcuts due to issues such as data contamination.Our findings show that, while larger models outperform smaller ones, they remain susceptible to adversarial manipulation and often rely on superficial patterns rather than true moral reasoning.This brittleness results in significant self-contradiction, with the best models refuting their own answers in roughly 20% of cases depending on the framing of the moral choice.Interestingly, reasoning-enhanced models fail to bridge this gap, suggesting that scale -not reasoning ability -is the primary driver of performance.
Matteo Marcuzzo, Alessandro Zangari, Andrea Albarelli, José Camacho-Collados, Mohammad Taher Pilehvar
EMNLP4
2025 Pun Unintended: LLMs and the Illusion of Humor Understanding
abstract
Puns are a form of humorous wordplay that exploits polysemy and phonetic similarity.While LLMs have shown promise in detecting puns, we show in this paper that their understanding often remains shallow, lacking the nuanced grasp typical of human interpretation.By systematically analyzing and reformulating existing pun benchmarks, we demonstrate how subtle changes in puns are sufficient to mislead LLMs.Our contributions include comprehensive and nuanced pun detection benchmarks, human evaluation across recent LLMs, and an analysis of the robustness challenges these models face in processing puns.
Alessandro Zangari, Matteo Marcuzzo, Andrea Albarelli, Mohammad Taher Pilehvar, José Camacho-Collados
EMNLP5
2025 RelBERT: Embedding relations with language models
abstract
Many applications need access to background knowledge about how different concepts and entities are related. Although Large Language Models (LLM) can address this need to some extent, LLMs are inefficient and difficult to control. As an alternative, we propose to extract relation embeddings from relatively small language models. In particular, we show that masked language models such as RoBERTa can be straightforwardly fine-tuned for this purpose, using only a small amount of training data. The resulting model, which we call RelBERT, captures relational similarity in a surprisingly fine-grained way, allowing us to set a new state-of-the-art in analogy benchmarks. Crucially, RelBERT is capable of modelling relations that go well beyond what the model has seen during training. For instance, we obtained strong results on relations between named entities with a model that was only trained on lexical relations between concepts, and we observed that RelBERT can recognise morphological analogies despite not being trained on such examples. Overall, we find that RelBERT significantly outperforms strategies based on prompting language models that are several orders of magnitude larger, including recent GPT-based models and open source models. 1
Asahi Ushio, José Camacho-Collados, Steven Schockaert
Artif. Intell.2
2024 Language Models for Text Classification: Is In-Context Learning Enough?
abstract
Recent foundational language models have shown state-of-the-art performance in many NLP tasks in zero- and few-shot settings. An advantage of these models over more standard approaches based on fine-tuning is the ability to understand instructions written in natural language (prompts), which helps them generalise better to different tasks and domains without the need for specific training data. This makes them suitable for addressing text classification problems for domains with limited amounts of annotated instances. However, existing research is limited in scale and lacks understanding of how text generation models combined with prompting techniques compare to more established methods for text classification such as fine-tuning masked language models. In this paper, we address this research gap by performing a large-scale evaluation study for 16 text classification datasets covering binary, multiclass, and multilabel problems. In particular, we compare zero- and few-shot approaches of large language models to fine-tuning smaller language models. We also analyse the results by prompt, classification type, domain, and number of labels. In general, the results show how fine-tuning smaller and more efficient language models can still outperform few-shot approaches of larger language models, which have room for improvement when it comes to text classification.
Aleksandra Edwards, José Camacho-Collados
LREC/COLING2
2024 Question Answering over Tabular Data with DataBench: A Large-Scale Empirical Evaluation of LLMs
abstract
Large Language Models (LLMs) are showing emerging abilities, and one of the latest recognized ones deals with their ability to reason and answer questions from tabular data. Although there are some available datasets to assess question answering systems on tabular data, they are not large and diverse enough to properly assess the capabilities of LLMs. To this end, we propose DataBench, a benchmark composed of 65 real-world datasets over several domains, including 20 human-generated questions per dataset, totaling 1300 questions and answers overall. Using this benchmark, we perform a large-scale empirical comparison of several open and closed source models, including both code-generating and in-context learning models. The results highlight the current gap between open-source and closed-source models, with all types of model having room for improvement even in simple boolean questions or involving a single column.
Jorge Osés Grijalba, Luis Alfonso Ureña López, Eugenio Martínez-Cámara, José Camacho-Collados
LREC/COLING4
2024 Do Large Language Models Understand Mansplaining? Well, Actually
abstract
Gender bias has been widely studied by the NLP community. However, other more subtle variations of it, such as mansplaining, have yet received little attention. Mansplaining is a discriminatory behaviour that consists of a condescending treatment or discourse towards women. In this paper, we introduce and analyze Well, actually..., a corpus of 886 mansplaining stories experienced by women. We analyze the corpus in terms of features such as offensiveness, sentiment or misogyny, among others. We also explore to what extent Large Language Models (LLMs) can understand and identify mansplaining and other gender-related microaggressions. Specifically, we experiment with ChatGPT-3.5-Turbo and LLaMA-2 (13b and 70b), with both targeted and open questions. Our findings suggest that, although they can identify mansplaining to some extent, LLMs still struggle to point out this attitude and will even reproduce some of the social patterns behind mansplaining situations, for instance by praising men for giving unsolicited advice to women.
Carla Pérez-Almendros, José Camacho-Collados
LREC/COLING2
2024 TweetTER: A Benchmark for Target Entity Retrieval on Twitter without Knowledge Bases
abstract
Entity linking is a well-established task in NLP consisting of associating entity mentions with entries in a knowledge base. Current models have demonstrated competitive performance in standard text settings. However, when it comes to noisy domains such as social media, certain challenges still persist. Typically, to evaluate entity linking on existing benchmarks, a comprehensive knowledge base is necessary and models are expected to possess an understanding of all the entities contained within the knowledge base. However, in practical scenarios where the objective is to retrieve sentences specifically related to a particular entity, strict adherence to a complete understanding of all entities in the knowledge base may not be necessary. To address this gap, we introduce TweetTER (Tweet Target Entity Retrieval), a novel benchmark that aims to bridge the challenges in entity linking. The distinguishing feature of this benchmark is its approach of re-framing entity linking as a binary entity retrieval task. This enables the evaluation of language models’ performance without relying on a conventional knowledge base, providing a more practical and versatile evaluation framework for assessing the effectiveness of language models in entity retrieval tasks.
Kiamehr Rezaee, José Camacho-Collados, Mohammad Taher Pilehvar
LREC/COLING2
2024 A RelEntLess Benchmark for Modelling Graded Relations between Named Entities
abstract
Relations such as "is influenced by", "is known for" or "is a competitor of" are inherently graded: we can rank entity pairs based on how well they satisfy these relations, but it is hard to draw a line between those pairs that satisfy them and those that do not.Such graded relations play a central role in many applications, yet they are typically not covered by existing Knowledge Graphs.In this paper, we consider the possibility of using Large Language Models (LLMs) to fill this gap.To this end, we introduce a new benchmark, in which entity pairs have to be ranked according to how much they satisfy a given graded relation.The task is formulated as a few-shot ranking problem, where models only have access to a description of the relation and five prototypical instances.We use the proposed benchmark to evaluate state-of-the-art relation embedding strategies as well as several publicly available LLMs and closed conversational models such as GPT-4.We find that smaller language models struggle to outperform a naive baseline.Overall, the best results are obtained with the 11B parameter Flan-T5 model and the 13B parameter OPT model, where further increasing the model size does not seem to be beneficial.For all models, a clear gap with human performance remains.
Asahi Ushio, José Camacho-Collados, Steven Schockaert
EACL (1)2
2024 Multilingual Topic Classification in X: Dataset and Analysis
abstract
In the dynamic realm of social media, diverse topics are discussed daily, transcending linguistic boundaries.However, the complexities of understanding and categorising this content across various languages remain an important challenge with traditional techniques like topic modelling often struggling to accommodate this multilingual diversity.In this paper, we introduce X-Topic, a multilingual dataset featuring content in four distinct languages (English, Spanish, Japanese, and Greek), crafted for the purpose of tweet topic classification.Our dataset includes a wide range of topics, tailored for social media content, making it a valuable resource for scientists and professionals working on cross-linguistic analysis, the development of robust multilingual models, and computational scientists studying online dialogue.Finally, we leverage X-Topic to perform a comprehensive cross-linguistic and multilingual analysis, and compare the capabilities of current general-and domain-specific language models.
Dimosthenis Antypas, Asahi Ushio, Francesco Barbieri, José Camacho-Collados
EMNLP4
2024 Evaluating Short-Term Temporal Fluctuations of Social Biases in Social Media Data and Masked Language Models
abstract
Social biases such as gender or racial biases have been reported in language models (LMs), including Masked Language Models (MLMs).Given that MLMs are continuously trained with increasing amounts of additional data collected over time, an important yet unanswered question is how the social biases encoded with MLMs vary over time.In particular, the number of social media users continues to grow at an exponential rate, and it is a valid concern for the MLMs trained specifically on social media data whether their social biases (if any) would also amplify over time.To empirically analyse this problem, we use a series of MLMs pretrained on chronologically ordered temporal snapshots of corpora.Our analysis reveals that, although social biases are present in all MLMs, most types of social bias remain relatively stable over time (with a few exceptions).To further understand the mechanisms that influence social biases in MLMs, we analyse the temporal corpora used to train the MLMs.Our findings show that some demographic groups, such as male, obtain higher preference over the other, such as female on the training corpora constantly.1
Yi Zhou 0019, Danushka Bollegala, José Camacho-Collados
EMNLP3
2024 Exploring Cross-Cultural Differences in English Hate Speech Annotations: From Dataset Construction to Analysis
abstract
Nayeon Lee, Chani Jung, Junho Myung, Jiho Jin, Jose Camacho-Collados, Juho Kim, Alice Oh. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Nayeon Lee, Chani Jung, Junho Myung, Jiho Jin, José Camacho-Collados, Juho Kim 0001, Alice Oh
NAACL-HLT5
2024 BLEnD: A Benchmark for LLMs on Everyday Knowledge in Diverse Cultures and Languages
abstract
Large language models (LLMs) often lack culture-specific everyday knowledge, especially across diverse regions and non-English languages. Existing benchmarks for evaluating LLMs' cultural sensitivities are usually limited to a single language or online sources like Wikipedia, which may not reflect the daily habits, customs, and lifestyles of different regions. That is, information about the food people eat for their birthday celebrations, spices they typically use, musical instruments youngsters play or the sports they practice in school is not always explicitly written online. To address this issue, we introduce BLEnD, a hand-crafted benchmark designed to evaluate LLMs' everyday knowledge across diverse cultures and languages. The benchmark comprises 52.6k question-answer pairs from 16 countries/regions, in 13 different languages, including low-resource ones such as Amharic, Assamese, Azerbaijani, Hausa, and Sundanese. We evaluate LLMs in two formats: short-answer questions, and multiple-choice questions. We show that LLMs perform better in cultures that are more present online, with a maximum 57.34% difference in GPT-4, the best-performing model, in the short-answer format.Furthermore, we find that LLMs perform better in their local languages for mid-to-high-resource languages. Interestingly, for languages deemed to be low-resource, LLMs provide better answers in English. We make our dataset publicly available at: https://github.com/nlee0212/BLEnD.
Junho Myung, Nayeon Lee, Yi Zhou 0019, Jiho Jin, Rifki Afina Putri, Dimosthenis Antypas, Hsuvas Borkakoty, Eunsu Kim, Carla Pérez-Almendros, Abinew Ali Ayele, Víctor Gutiérrez-Basulto, Yazmín Ibáñez-García, Hwaran Lee, Shamsuddeen Hassan Muhammad, Ki-Woong Park, Anar Rzayev, Nina White, Seid Muhie Yimam, Mohammad Taher Pilehvar, Nedjma Ousidhoum, José Camacho-Collados, Alice Oh
NeurIPS21
2024 Federated Learning for Exploiting Annotators' Disagreements in Natural Language Processing
abstract
Abstract The annotation of ambiguous or subjective NLP tasks is usually addressed by various annotators. In most datasets, these annotations are aggregated into a single ground truth. However, this omits divergent opinions of annotators, hence missing individual perspectives. We propose FLEAD (Federated Learning for Exploiting Annotators’ Disagreements), a methodology built upon federated learning to independently learn from the opinions of all the annotators, thereby leveraging all their underlying information without relying on a single ground truth. We conduct an extensive experimental study and analysis in diverse text classification tasks to show the contribution of our approach with respect to mainstream approaches based on majority voting and other recent methodologies that also learn from annotator disagreements.
Nuria Rodríguez Barroso, Eugenio Martínez-Cámara, José Camacho-Collados, María Victoria Luzón, Francisco Herrera
Trans. Assoc. Comput. Linguistics3
2023 LongEval: Longitudinal Evaluation of Model Performance at CLEF 2023
Rabab Alkhalifa, Iman Munire Bilal, Hsuvas Borkakoty, José Camacho-Collados, Romain Deveaud, Alaa El-Ebshihy, Luis Espinosa Anke, Gabriela González Sáez, Petra Galuscáková, Lorraine Goeuriot, Elena Kochkina, Maria Liakata, Daniel Loureiro, Harish Tayyar Madabushi, Philippe Mulhem, Florina Piroi, Martin Popel, Christophe Servan, Arkaitz Zubiaga
ECIR (3)4
2023 Construction Artifacts in Metaphor Identification Datasets
abstract
Metaphor identification aims at understanding whether a given expression is used figuratively in context.However, in this paper we show how existing metaphor identification datasets can be gamed by fully ignoring the potential metaphorical expression or the context in which it occurs.We test this hypothesis in a variety of datasets and settings, and show that metaphor identification systems based on language models without complete information can be competitive with those using the full context.This is due to the construction procedures to build such datasets, which introduce unwanted biases for positive and negative classes.Finally, we test the same hypothesis on datasets that are carefully sampled from natural corpora and where this bias is not present, making these datasets more challenging and reliable. Accuracy PrecisionRecall F1 Dataset Maj Def PME Mask Def PME Mask Def PME Mask Def PME Mask Psy CARD_N 50.0 87.5 44.5 85.9 90.5 42.8 86.7 83.9 33.9 85.1 87.0 35.4 85.6 CARD_V 50.0 83.9 42.9 82.9 87.2 40.8 83.1 80.8 36.9 83.4 83.3 38.3 83.0 JANK 66.7 85.3 51.1 84.2 81.2 45.1 82.2 75.0 28.4 68.5 76.5 45.1 74.2 TroFi 57.6 88.5 71.3 84.6 91.2 73.7 88.1 88.7 78.2 84.8 89.9 75.8 86.4 TSV_AN 50.3 89.3 59.4 79.2 91.3 62.2 80.2 86.9 52.0 78.0 89.1 55.8 78.9 GUT 53.6 98.3 65.4 95.4 98.9 66.8 96.0 97.9 70.6 95.3 98.4 68.6 95.7 NLP MOH 78.8 78.3 73.2 73.8 48.8 32.7 33.2 35.8 22.1 20.0 40.7 45.1 23.9 LLC 58.7 86.5 80.5 74.1 83.1 76.8 69.2 84.5 75.8 68.1 83.8 76.2 68.4 CHAK 66.7 69.7 74.4 66.7 76.0 80.2 66.7 79.8 81.8 100.0 77.8 81.0 80.0 NEU 56.0 76.0 56.0 72.0 81.3 58.8 79.9 79.8 66.0 74.5 77.6 60.4 72.9 DUNN 66.7 71.7 63.3 71.7 78.7 66.3 73.3 77.4 96.4 93.5 76.8 77.4 81.0 IDIX 51.4 93.8 85.3 84.6 93.3 83.1 84.4 93.7 86.8 83.0 93.5 84.9 83.7 PVC 65.1 85.9 76.9 80.8 89.0 79.8 83.2 89.4 86.6 88.6 89.2 82.8 85.7 VNC 78.5 96.0 93.5 87.0 97.1 95.0 89.8 97.9 96.8 94.1 97.5 95.9 91.9 PIE SE2013_ALL 59.5 91.5 83.2 85.3 93.5 84.5 86.3 92.3 87.9 89.3 92.8 86.2 87.8 SE2013_LEX 50.6 92.6 61.9 89.3 93.4 60.4 89.0 91.9 73.7 90.2 92.6 65.8 89.6 MAD 52.0 94.6 87.0 76.0 93.1 82.2 75.4 95.9 93.1 74.6 94.4 87.3 74.9 PIE 52.6 94.9 94.4 86.2 93.3 92.5 83.1 96.1 95.8 89.1 94.7 94.1 86.0 MAGPIE 74.7 96.1 93.3 86.7 97.4 95.4 88.6 97.3 95.5 94.3 97.4 95.5 91.4 VUAC_DO 50.4 75.5 77.4 63.0 74.1 77.1 62.1 79.1 78.4 68.0 76.5 77.7 64.9 VUAC VUAC_ST1 71.6 86.2 75.8 77.2 77.3 57.5 61.4 73.1 56.9 53.9 75.1 57.2 57.4 VUAC_ST2 84.3 92.3 86.5 84.7 76.7 56.4 51.4 73.0 61.1 39.9 74.8 58.6 45.0 VUAC_BO 50.5 85.4 63.3 76.0 84.6 63.4 75.4 86.9 64.6 77.8 85.8 64.0 76.6Table 4: Majority class accuracy (Maj) accuracy is shown in the first result column.Accuracy, precision, recall and F1 results for the metaphor class, averaged over 5 cross-validation folds, for the Default (Def), only PME (PME), and Masked settings on the random splits of metaphor identification datasets appear in the following columns. Accuracy Precision Recall F1 Dataset Maj Def PME MaskDef PME Mask Def PME Mask Def PME Mask CARD_N 50.0 89.7 50.0 86.7 90.1 50.0 88.6 89.5 59.4 84.4 89.7 54.0 86.4 Psy CARD_V 50.0 86.1 50.0 82.1 87.6 50.0 85.4 84.3 54.3 77.9 85.6 48.7 81.2 JANK 66.7 84.2 51.4 83.3 78.6 26.7 78.8 74.2 45.8 70.0 75.9 33.3 73.6 TroFi 57.6 82.3 62.9 78.5 85.1 63.4 81.2 85.0 84.1 82.9 84.7 72.2 81.6 TSV_AN 50.3 87.0 63.5 77.8 88.0 64.6 76.9 85.2 59.9 79.2 86.5 61.2 77.9 TSV_AN_L2 50.3 87.6 59.5 78.2 91.7 61.5 79.6 83.0 53.4 77.0 87.1 56.8 78.0 GUT 53.6 95.3 52.1 94.2 96.0 47.6 93.9 94.8 55.0 94.9 95.3 45.6 94.3 NLP GUT_L2 53.6 97.6 66.0 94.8 97.8 68.1 95.8 97.7 69.6 94.5 97.7 68.5 95.1 MOH 78.8 79.0 74.4 74.2 56.9 34.9 34.4 34.4 22.2 45.1 40.7 24.5 45.1 LLC 58.7 85.1 79.1 72.8 83.2 73.7 67.5 80.0 76.7 66.0 81.6 75.2 66.7 CHAK 66.7 64.3 73.2 63.7 74.6 81.0 65.6 71.1 77.9 95.3 72.2 79.3 77.6 NEU 56.0 76.0 50.0 77.0 84.4 43.0 82.2 73.6 80.0 81.0 73.8 54.5 79.1 DUNN 66.7 66.7 56.7 70.0 72.7 64.8 74.7 82.5 75.0 87.5 76.2 65.9 79.8 IDIX 51.4 75.5 63.2 74.1 76.5 66.3 74.5 77.4 62.5 74.9 75.1 62.5 73.5 PVC_V 65.1 69.5 59.2 67.3 73.4 66.8 68.9 78.9 76.3 80.0 75.3 67.4 72.8 VNC 78.5 84.3 73.4 81.8 90.4 85.6 86.4 90.5 80.6 91.2 89.5 81.7 88.5 PIE SE2013_ALL 59.5 79.2 49.2 79.1 81.5 57.8 78.7 82.3 50.6 85.2 79.9 48.7 80.4 SE2013_LEX 50.6 81.3 47.1 78.7 80.2 41.0 79.8 86.0 50.9 78.2 81.8 42.1 78.1 MAD 52.0 78.1 71.1 69.3 78.5 65.9 68.0 75.0 79.8 66.1 76.0 72.0 66.9 PIE 52.6 87.2 87.6 74.7 84.0 82.4 71.4 90.3 94.1 79.1 86.8 87.8 74.9 MAGPIE 73.7 90.2 84.9 83.4 94.4 88.4 88.4 92.2 91.5 89.2 93.3 89.9 88.8 VUAC_DO 57.6 74.5 75.0 59.6 78.5 78.7 64.7 76.7 77.5 65.7 77.6 78.0 65.2 VUAC VUAC_ST1 68.5 77.0 67.0 73.2 62.7 47.8 58.5 66.8 49.9 51.6 64.7 48.8 54.8 VUAC_ST2 82.6 88.3 83.3 84.2 68.8 53.2 56.1 60.0 36.3 41.7 64.1 43.2 47.9 VUAC_BO 53.4 82.2 65.5 73.0 80.7 64.9 71.5 87.7 77.0 82.1 84.1 77.0 76.5
Joanne Boisson, Luis Espinosa Anke, José Camacho-Collados
EMNLP3
2023 A Predictive Factor Analysis of Social Biases and Task-Performance in Pretrained Masked Language Models
abstract
Various types of social biases have been reported with pretrained Masked Language Models (MLMs) in prior work.However, multiple underlying factors are associated with an MLM such as its model size, size of the training data, training objectives, the domain from which pretraining data is sampled, tokenization, and languages present in the pretrained corpora, to name a few.It remains unclear as to which of those factors influence social biases that are learned by MLMs.To study the relationship between model factors and the social biases learned by an MLM, as well as the downstream task performance of the model, we conduct a comprehensive study over 39 pretrained MLMs covering different model sizes, training objectives, tokenization methods, training data domains and languages.Our results shed light on important factors often neglected in prior literature, such as tokenization or model objectives.
Yi Zhou 0019, José Camacho-Collados, Danushka Bollegala
EMNLP2
2023 Meemi: A simple method for post-processing and integrating cross-lingual word embeddings
abstract
Abstract Word embeddings have become a standard resource in the toolset of any Natural Language Processing practitioner. While monolingual word embeddings encode information about words in the context of a particular language, cross-lingual embeddings define a multilingual space where word embeddings from two or more languages are integrated together. Current state-of-the-art approaches learn these embeddings by aligning two disjoint monolingual vector spaces through an orthogonal transformation which preserves the structure of the monolingual counterparts. In this work, we propose to apply an additional transformation after this initial alignment step, which aims to bring the vector representations of a given word and its translations closer to their average. Since this additional transformation is non-orthogonal, it also affects the structure of the monolingual spaces. We show that our approach both improves the integration of the monolingual spaces and the quality of the monolingual spaces themselves. Furthermore, because our transformation can be applied to an arbitrary number of languages, we are able to effectively obtain a truly multilingual space. The resulting (monolingual and multilingual) spaces show consistent gains over the current state-of-the-art in standard intrinsic tasks, namely dictionary induction and word similarity, as well as in extrinsic tasks such as cross-lingual hypernym discovery and cross-lingual natural language inference.
Yerai Doval, José Camacho-Collados, Luis Espinosa Anke, Steven Schockaert
Nat. Lang. Eng.2
2022 Twitter Topic Classification
abstract
Social media platforms host discussions about a wide variety of topics that arise everyday. Making sense of all the content and organising it into categories is an arduous task. A common way to deal with this issue is relying on topic modeling, but topics discovered using this technique are difficult to interpret and can differ from corpus to corpus. In this paper, we present a new task based on tweet topic classification and release two associated datasets. Given a wide range of topics covering the most important discussion points in social media, we provide training and testing data from recent time periods that can be used to evaluate tweet classification models. Moreover, we perform a quantitative evaluation and analysis of current general- and domain-specific language models on the task, which provide more insights on the challenges and nature of the task.
Dimosthenis Antypas, Asahi Ushio, José Camacho-Collados, Vítor Silva 0003, Leonardo Neves, Francesco Barbieri
COLING3
2022 TempoWiC: An Evaluation Benchmark for Detecting Meaning Shift in Social Media
abstract
Language evolves over time, and word meaning changes accordingly. This is especially true in social media, since its dynamic nature leads to faster semantic shifts, making it challenging for NLP models to deal with new content and trends. However, the number of datasets and models that specifically address the dynamic nature of these social platforms is scarce. To bridge this gap, we present TempoWiC, a new benchmark especially aimed at accelerating research in social media-based meaning shift. Our results show that TempoWiC is a challenging benchmark, even for recently-released language models specialized in social media.
Daniel Loureiro, Aminette D'Souza, Areej Nasser Muhajab, Isabella A. White, Gabriel Wong, Luis Espinosa Anke, Leonardo Neves, Francesco Barbieri, José Camacho-Collados
COLING9
2022 Generative Language Models for Paragraph-Level Question Generation
abstract
Powerful generative models have led to recent progress in question generation (QG).However, it is difficult to measure advances in QG research since there are no standardized resources that allow a uniform comparison among approaches.In this paper, we introduce QG-Bench, a multilingual and multidomain benchmark for QG that unifies existing question answering datasets by converting them to a standard QG setting.It includes generalpurpose datasets such as SQuAD (Rajpurkar et al., 2016) for English, datasets from ten domains and two styles, as well as datasets in eight different languages.Using QG-Bench as a reference, we perform an extensive analysis of the capabilities of language models for the task.First, we propose robust QG baselines based on fine-tuning generative language models.Then, we complement automatic evaluation based on standard metrics with an extensive manual evaluation, which in turn sheds light on the difficulty of evaluating QG models.Finally, we analyse both the domain adaptability of these models as well as the effectiveness of multilingual models in languages other than English.QG-Bench is released along with the fine-tuned models presented in the paper, 1 which are also available as a demo. 2
Asahi Ushio, Fernando Alva-Manchego, José Camacho-Collados
EMNLP3
2022 XLM-T: Multilingual Language Models in Twitter for Sentiment Analysis and Beyond
abstract
Language models are ubiquitous in current NLP, and their multilingual capacity has recently attracted considerable attention. However, current analyses have almost exclusively focused on (multilingual variants of) standard benchmarks, and have relied on clean pre-training and task-specific corpora as multilingual signals. In this paper, we introduce XLM-T, a model to train and evaluate multilingual language models in Twitter. In this paper we provide: (1) a new strong multilingual baseline consisting of an XLM-R (Conneau et al. 2020) model pre-trained on millions of tweets in over thirty languages, alongside starter code to subsequently fine-tune on a target task; and (2) a set of unified sentiment analysis Twitter datasets in eight different languages and a XLM-T model trained on this dataset.
Francesco Barbieri, Luis Espinosa Anke, José Camacho-Collados
LREC3
2022 LMMS reloaded: Transformer-based sense embeddings for disambiguation and beyond
abstract
Distributional semantics based on neural approaches is a cornerstone of Natural Language Processing, with surprising connections to human meaning representation as well. Recent Transformer-based Language Models have proven capable of producing contextual word representations that reliably convey sense-specific information, simply as a product of self-supervision. Prior work has shown that these contextual representations can be used to accurately represent large sense inventories as sense embeddings, to the extent that a distance-based solution to Word Sense Disambiguation (WSD) tasks outperforms models trained specifically for the task. Still, there remains much to understand on how to use these Neural Language Models (NLMs) to produce sense embeddings that can better harness each NLM's meaning representation abilities. In this work we introduce a more principled approach to leverage information from all layers of NLMs, informed by a probing analysis on 14 NLM variants. We also emphasize the versatility of these sense embeddings in contrast to task-specific models, applying them on several sense-related tasks, besides WSD, while demonstrating improved performance using our proposed approach over prior work focused on sense embeddings. Finally, we discuss unexpected findings regarding layer and model performance variations, and potential applications for downstream tasks.
Daniel Loureiro, Alípio Mário Jorge, José Camacho-Collados
Artif. Intell.3
2021 BERT is to NLP what AlexNet is to CV: Can Pre-Trained Language Models Identify Analogies?
abstract
Asahi Ushio, Luis Espinosa Anke, Steven Schockaert, Jose Camacho-Collados. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Asahi Ushio, Luis Espinosa Anke, Steven Schockaert, José Camacho-Collados
ACL/IJCNLP (1)4
2021 WiC-TSV: An Evaluation Benchmark for Target Sense Verification of Words in Context
abstract
Anna Breit, Artem Revenko, Kiamehr Rezaee, Mohammad Taher Pilehvar, Jose Camacho-Collados. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021.
Anna Breit, Artem Revenko, Kiamehr Rezaee, Mohammad Taher Pilehvar, José Camacho-Collados
EACL5
2021 Distilling Relation Embeddings from Pretrained Language Models
abstract
Pre-trained language models have been found to capture a surprisingly rich amount of lexical knowledge, ranging from commonsense properties of everyday concepts to detailed factual knowledge about named entities.Among others, this makes it possible to distill high-quality word vectors from pre-trained language models.However, it is currently unclear to what extent it is possible to distill relation embeddings, i.e. vectors that characterize the relationship between two words.Such relation embeddings are appealing because they can, in principle, encode relational knowledge in a more finegrained way than is possible with knowledge graphs.To obtain relation embeddings from a pre-trained language model, we encode word pairs using a (manually or automatically generated) prompt, and we fine-tune the language model such that relationally similar word pairs yield similar output vectors.We find that the resulting relation embeddings are highly competitive on analogy (unsupervised) and relation classification (supervised) benchmarks, even without any task-specific fine-tuning. 1
Asahi Ushio, José Camacho-Collados, Steven Schockaert
EMNLP (1)2
2021 Back to the Basics: A Quantitative Analysis of Statistical and Graph-Based Term Weighting Schemes for Keyword Extraction
abstract
Term weighting schemes are widely used in Natural Language Processing and Information Retrieval.In particular, term weighting is the basis for keyword extraction.However, there are relatively few evaluation studies that shed light about the strengths and shortcomings of each weighting scheme.In fact, in most cases researchers and practitioners resort to the wellknown tf-idf as default, despite the existence of other suitable alternatives, including graphbased models.In this paper, we perform an exhaustive and large-scale empirical comparison of both statistical and graph-based term weighting methods in the context of keyword extraction.Our analysis reveals some interesting findings such as the advantages of the lessknown lexical specificity with respect to tf-idf, or the qualitative differences between statistical and graph-based methods.Finally, based on our findings we discuss and devise some suggestions for practitioners. 1
Asahi Ushio, Federico Liberatore, José Camacho-Collados
EMNLP (1)3
2021 Modelling General Properties of Nouns by Selectively Averaging Contextualised Embeddings
abstract
While the success of pre-trained language models has largely eliminated the need for high-quality static word vectors in many NLP applications, static word vectors continue to play an important role in tasks where word meaning needs to be modelled in the absence of linguistic context. In this paper, we explore how the contextualised embeddings predicted by BERT can be used to produce high-quality word vectors for such domains, in particular related to knowledge base completion, where our focus is on capturing the semantic properties of nouns. We find that a simple strategy of averaging the contextualised embeddings of masked word mentions leads to vectors that outperform the static word vectors learned by BERT, as well as those from standard word embedding models, in property induction tasks. We notice in particular that masking target words is critical to achieve this strong performance, as the resulting vectors focus less on idiosyncratic properties and more on general semantic properties. Inspired by this view, we propose a filtering strategy which is aimed at removing the most idiosyncratic mention vectors, allowing us to obtain further performance gains in property induction.
Na Li 0018, Zied Bouraoui, José Camacho-Collados, Luis Espinosa Anke, Qing Gu 0001, Steven Schockaert
IJCAI3
2021 Analysis and Evaluation of Language Models for Word Sense Disambiguation
abstract
Abstract Transformer-based language models have taken many fields in NLP by storm. BERT and its derivatives dominate most of the existing evaluation benchmarks, including those for Word Sense Disambiguation (WSD), thanks to their ability in capturing context-sensitive semantic nuances. However, there is still little knowledge about their capabilities and potential limitations in encoding and recovering word senses. In this article, we provide an in-depth quantitative and qualitative analysis of the celebrated BERT model with respect to lexical ambiguity. One of the main conclusions of our analysis is that BERT can accurately capture high-level sense distinctions, even when a limited number of examples is available for each word sense. Our analysis also reveals that in some cases language models come close to solving coarse-grained noun disambiguation under ideal conditions in terms of availability of training data and computing resources. However, this scenario rarely occurs in real-world settings and, hence, many practical challenges remain even in the coarse-grained setting.We also perform an in-depth comparison of the two main language model based WSD strategies, namely, fine-tuning and feature extraction, finding that the latter approach is more robust with respect to sense bias and it can better exploit limited available training data. In fact, the simple feature extraction strategy of averaging contextualized embeddings proves robust even using only three training sentences per word sense, with minimal improvements obtained by increasing the size of this training data.
Daniel Loureiro, Kiamehr Rezaee, Mohammad Taher Pilehvar, José Camacho-Collados
Comput. Linguistics4
2020 Modelling Semantic Categories Using Conceptual Neighborhood
abstract
While many methods for learning vector space embeddings have been proposed in the field of Natural Language Processing, these methods typically do not distinguish between categories and individuals. Intuitively, if individuals are represented as vectors, we can think of categories as (soft) regions in the embedding space. Unfortunately, meaningful regions can be difficult to estimate, especially since we often have few examples of individuals that belong to a given category. To address this issue, we rely on the fact that different categories are often highly interdependent. In particular, categories often have conceptual neighbors, which are disjoint from but closely related to the given category (e.g. fruit and vegetable). Our hypothesis is that more accurate category representations can be learned by relying on the assumption that the regions representing such conceptual neighbors should be adjacent in the embedding space. We propose a simple method for identifying conceptual neighbors and then show that incorporating these conceptual neighbors indeed leads to more accurate region based representations.
Zied Bouraoui, José Camacho-Collados, Luis Espinosa Anke, Steven Schockaert
AAAI2
2020 Inducing Relational Knowledge from BERT
abstract
One of the most remarkable properties of word embeddings is the fact that they capture certain types of semantic and syntactic relationships. Recently, pre-trained language models such as BERT have achieved groundbreaking results across a wide range of Natural Language Processing tasks. However, it is unclear to what extent such models capture relational knowledge beyond what is already captured by standard word embeddings. To explore this question, we propose a methodology for distilling relational knowledge from a pre-trained language model. Starting from a few seed instances of a given relation, we first use a large text corpus to find sentences that are likely to express this relation. We then use a subset of these extracted sentences as templates. Finally, we fine-tune a language model to predict whether a given word pair is likely to be an instance of some relation, when given an instantiated template for that relation as input.
Zied Bouraoui, José Camacho-Collados, Steven Schockaert
AAAI2
2020 Go Simple and Pre-Train on Domain-Specific Corpora: On the Role of Training Data for Text Classification
abstract
Pre-trained language models provide the foundations for state-of-the-art performance across a wide range of natural language processing tasks, including text classification.However, most classification datasets assume a large amount labeled data, which is commonly not the case in practical settings.In particular, in this paper we compare the performance of a light-weight linear classifier based on word embeddings, i.e., fastText (Joulin et al., 2017), versus a pre-trained language model, i.e., BERT (Devlin et al., 2019), across a wide range of datasets and classification tasks.In general, results show the importance of domain-specific unlabeled data, both in the form of word embeddings or language models.As for the comparison, BERT outperforms all baselines in standard datasets with large training sets.However, in settings with small training datasets a simple method like fastText coupled with domain-specific word embeddings performs equally well or better than BERT, even when pre-trained on domain-specific data.
Aleksandra Edwards, José Camacho-Collados, Hélène de Ribaupierre, Alun D. Preece
COLING2
2020 Understanding the Source of Semantic Regularities in Word Embeddings
abstract
Semantic relations are core to how humans understand and express concepts in the real world using language.Recently, there has been a thread of research aimed at modeling these relations by learning vector representations from text corpora.Most of these approaches focus strictly on leveraging the co-occurrences of relationship word pairs within sentences.In this paper, we investigate the hypothesis that examples of a lexical relation in a corpus are fundamental to a neural word embedding's ability to complete analogies involving the relation.Our experiments, in which we remove all known examples of a relation from training corpora, show only marginal degradation in analogy completion performance involving the removed relation.This finding enhances our understanding of neural word embeddings, showing that co-occurrence information of a particular semantic relation is the not the main source of their structural regularity.
Hsiao-Yu Chiang, José Camacho-Collados, Zachary A. Pardos
CoNLL2
2020 Capturing Word Order in Averaging Based Sentence Embeddings
abstract
One of the most remarkable findings in the literature on sentence embeddings has been that simple word vector averaging can compete with state-of-the-art models in many tasks. While counter-intuitive, a convincing explanation has been provided by Arora et al., who showed that the bag-of-words representation of a sentence can be recovered from its word vector average with almost perfect accuracy. Beyond word vector averaging, however, most sentence embedding models are essentially black boxes: while there is abundant empirical evidence about their strengths and weaknesses, it is not clear why and how different embedding strategies are able to capture particular properties of sentences. In this paper, we focus in particular on how sentence embedding models are able to capture word order. For instance, it seems intuitively puzzling that simple LSTM autoencoders are able to learn sentence vectors from which the original sentence can be reconstructed almost perfectly. With the aim of elucidating this phenomenon, we show that to capture word order, it is in fact sufficient to supplement standard word vector averages with averages of bigram and trigram vectors. To this end, we first study the problem of reconstructing bags-of-bigrams, focusing in particular on how suitable bigram vectors should be encoded. We then show that LSTMs are capable, in principle, of learning our proposed sentence embeddings. Empirically, we find that our embeddings outperform those learned by LSTM autoencoders on the task of sentence reconstruction, while needing almost no training data.
Jae Hee Lee 0001, José Camacho-Collados, Luis Espinosa Anke, Steven Schockaert
ECAI2
2020 Don't Neglect the Obvious: On the Role of Unambiguous Words in Word Sense Disambiguation
abstract
State-of-the-art methods for Word Sense Disambiguation (WSD) combine two different features: the power of pre-trained language models and a propagation method to extend the coverage of such models.This propagation is needed as current sense-annotated corpora lack coverage of many instances in the underlying sense inventory (usually WordNet).At the same time, unambiguous words make for a large portion of all words in WordNet, while being poorly covered in existing senseannotated corpora.In this paper, we propose a simple method to provide annotations for most unambiguous words in a large corpus.We introduce the UWA (Unambiguous Word Annotations) dataset and show how a state-of-theart propagation-based model can use it to extend the coverage and quality of its word sense embeddings by a significant margin, improving on its original results on WSD.
Daniel Loureiro, José Camacho-Collados
EMNLP (1)2
2020 XL-WiC: A Multilingual Benchmark for Evaluating Semantic Contextualization
abstract
The ability to correctly model distinct meanings of a word is crucial for the effectiveness of semantic representation techniques.However, most existing evaluation benchmarks for assessing this criterion are tied to sense inventories (usually WordNet), restricting their usage to a small subset of knowledge-based representation techniques.The Word-in-Context dataset (WiC) addresses the dependence on sense inventories by reformulating the standard disambiguation task as a binary classification problem; but, it is limited to the English language.We put forward a large multilingual benchmark, XL-WiC, featuring gold standards in 12 new languages from varied language families and with different degrees of resource availability, opening room for evaluation scenarios such as zero-shot cross-lingual transfer.We perform a series of experiments to determine the reliability of the datasets and to set performance baselines for several recent contextualized multilingual models.Experimental results show that even when no tagged instances are available for a target language, models trained solely on the English data can attain competitive performance in the task of distinguishing different meanings of a word, even for distant languages.XL-WiC is available at https://pilehvar.github.io/xlwic/.
Alessandro Raganato, Tommaso Pasini, José Camacho-Collados, Mohammad Taher Pilehvar
EMNLP (1)3
2020 Learning Cross-Lingual Word Embeddings from Twitter via Distant Supervision
José Camacho-Collados, Yerai Doval, Eugenio Martínez-Cámara, Luis Espinosa Anke, Francesco Barbieri, Steven Schockaert
ICWSM1
2020 On the Robustness of Unsupervised and Semi-supervised Cross-lingual Word Embedding Learning
abstract
Cross-lingual word embeddings are vector representations of words in different languages where words with similar meaning are represented by similar vectors, regardless of the language. Recent developments which construct these embeddings by aligning monolingual spaces have shown that accurate alignments can be obtained with little or no supervision, which usually comes in the form of bilingual dictionaries. However, the focus has been on a particular controlled scenario for evaluation, and there is no strong evidence on how current state-of-the-art systems would fare with noisy text or for language pairs with major linguistic differences. In this paper we present an extensive evaluation over multiple cross-lingual embedding models, analyzing their strengths and limitations with respect to different variables such as target language, training corpora and amount of supervision. Our conclusions put in doubt the view that high-quality cross-lingual embeddings can always be learned without much supervision.
Yerai Doval, José Camacho-Collados, Luis Espinosa Anke, Steven Schockaert
LREC2
2020 A Short Survey on Sense-Annotated Corpora
abstract
Large sense-annotated datasets are increasingly necessary for training deep supervised systems in Word Sense Disambiguation. However, gathering high-quality sense-annotated data for as many instances as possible is a laborious and expensive task. This has led to the proliferation of automatic and semi-automatic methods for overcoming the so-called knowledge-acquisition bottleneck. In this short survey we present an overview of sense-annotated corpora, annotated either manually- or (semi)automatically, that are currently available for different languages and featuring distinct lexical resources as inventory of senses, i.e. WordNet, Wikipedia, BabelNet. Furthermore, we provide the reader with general statistics of each dataset and an analysis of their specific features.
Tommaso Pasini, José Camacho-Collados
LREC2
2019 Relational Word Embeddings
abstract
While word embeddings have been shown to implicitly encode various forms of attributional knowledge, the extent to which they capture relational information is far more limited.In previous work, this limitation has been addressed by incorporating relational knowledge from external knowledge bases when learning the word embedding.Such strategies may not be optimal, however, as they are limited by the coverage of available resources and conflate similarity with other forms of relatedness.As an alternative, in this paper we propose to encode relational knowledge in a separate word embedding, which is aimed to be complementary to a given standard word embedding.This relational word embedding is still learned from co-occurrence statistics, and can thus be used even when no external knowledge base is available.Our analysis shows that relational word vectors do indeed capture information that is complementary to what is encoded in standard word embeddings.
José Camacho-Collados, Luis Espinosa Anke, Steven Schockaert
ACL (1)1
2019 A Latent Variable Model for Learning Distributional Relation Vectors
abstract
Recently a number of unsupervised approaches have been proposed for learning vectors that capture the relationship between two words. Inspired by word embedding models, these approaches rely on co-occurrence statistics that are obtained from sentences in which the two target words appear. However, the number of such sentences is often quite small, and most of the words that occur in them are not relevant for characterizing the considered relationship. As a result, standard co-occurrence statistics typically lead to noisy relation vectors. To address this issue, we propose a latent variable model that aims to explicitly determine what words from the given sentences best characterize the relationship between the two target words. Relation vectors then correspond to the parameters of a simple unigram language model which is estimated from these words.
José Camacho-Collados, Luis Espinosa Anke, Shoaib Jameel, Steven Schockaert
IJCAI1
2019 Knowledge-enhanced document embeddings for text classification
abstract
Accurate semantic representation models are essential in text mining applications. For a successful application of the text mining process, the text representation adopted must keep the interesting patterns to be discovered. Although competitive results for automatic text classification may be achieved with traditional bag of words, such representation model cannot provide satisfactory classification performances on hard settings where richer text representations are required. In this paper, we present an approach to represent document collections based on embedded representations of words and word senses. We bring together the power of word sense disambiguation and the semantic richness of word- and word-sense embedded vectors to construct embedded representations of document collections. Our approach results in semantically enhanced and low-dimensional representations. We overcome the lack of interpretability of embedded vectors, which is a drawback of this kind of representation, with the use of word sense embedded vectors. Moreover, the experimental evaluation indicates that the use of the proposed representations provides stable classifiers with strong quantitative results, especially in semantically-complex classification scenarios
Roberta Akemi Sinoara, José Camacho-Collados, Rafael Geraldeli Rossi, Roberto Navigli, Solange Oliveira Rezende
Knowl. Based Syst.2
2018 Interpretable Emoji Prediction via Label-Wise Attention LSTMs
abstract
Human language has evolved towards newer forms of communication such as social media, where emojis (i.e., ideograms bearing a visual meaning) play a key role.While there is an increasing body of work aimed at the computational modeling of emoji semantics, there is currently little understanding about what makes a computational model represent or predict a given emoji in a certain way.In this paper we propose a label-wise attention mechanism with which we attempt to better understand the nuances underlying emoji prediction.In addition to advantages in terms of interpretability, we show that our proposed architecture improves over standard baselines in emoji prediction, and does particularly well when predicting infrequent emojis.
Francesco Barbieri, Luis Espinosa Anke, José Camacho-Collados, Steven Schockaert, Horacio Saggion
EMNLP3
2018 Improving Cross-Lingual Word Embeddings by Meeting in the Middle
abstract
Cross-lingual word embeddings are becoming increasingly important in multilingual NLP.Recently, it has been shown that these embeddings can be effectively learned by aligning two disjoint monolingual vector spaces through linear transformations, using no more than a small bilingual dictionary as supervision.In this work, we propose to apply an additional transformation after the initial alignment step, which moves cross-lingual synonyms towards a middle point between them.By applying this transformation our aim is to obtain a better cross-lingual integration of the vector spaces.In addition, and perhaps surprisingly, the monolingual spaces also improve by this transformation.This is in contrast to the original alignment, which is typically learned such that the structure of the monolingual spaces is preserved.Our experiments confirm that the resulting cross-lingual embeddings outperform state-of-the-art models in both monolingual and cross-lingual evaluation tasks.
Yerai Doval, José Camacho-Collados, Luis Espinosa Anke, Steven Schockaert
EMNLP2
2018 From Word To Sense Embeddings: A Survey on Vector Representations of Meaning
abstract
Over the past years, distributed semantic representations have proved to be effective and flexible keepers of prior knowledge to be integrated into downstream applications. This survey focuses on the representation of meaning. We start from the theoretical background behind word vector space models and highlight one of their major limitations: the meaning conflation deficiency, which arises from representing a word with all its possible meanings as a single vector. Then, we explain how this deficiency can be addressed through a transition from the word level to the more fine-grained level of word senses (in its broader acceptation) as a method for modelling unambiguous lexical meaning. We present a comprehensive overview of the wide range of techniques in the two main branches of sense representation, i.e., unsupervised and knowledge-based. Finally, this survey covers the main evaluation procedures and applications for this type of representation, and provides an analysis of four of its important aspects: interpretability, sense granularity, adaptability to different domains and compositionality.
José Camacho-Collados, Mohammad Taher Pilehvar
J. Artif. Intell. Res.1
2018 Applying automatic text-based detection of deceptive language to police reports: Extracting behavioral patterns from a multi-step classification model to understand how we lie to the police
Lara Quijano Sánchez, Federico Liberatore, José Camacho-Collados, Miguel Camacho-Collados
Knowl. Based Syst.3
2017 Towards a Seamless Integration of Word Senses into Downstream NLP Applications
abstract
Lexical ambiguity can impede NLP systems from accurate understanding of semantics.Despite its potential benefits, the integration of sense-level information into NLP systems has remained understudied.By incorporating a novel disambiguation algorithm into a state-of-the-art classification model, we create a pipeline to integrate sense-level information into downstream NLP applications.We show that a simple disambiguation of the input text can lead to consistent performance improvement on multiple topic categorization and polarity detection datasets, particularly when the fine granularity of the underlying sense inventory is reduced and the document is sufficiently large.Our results also point to the need for sense representation research to focus more on in vivo evaluations which target the performance in downstream NLP applications rather than artificial benchmarks.
Mohammad Taher Pilehvar, José Camacho-Collados, Roberto Navigli, Nigel Collier
ACL (1)2
2017 Embedding Words and Senses Together via Joint Knowledge-Enhanced Training
abstract
Word embeddings are widely used in Natural Language Processing, mainly due to their success in capturing semantic information from massive corpora.However, their creation process does not allow the different meanings of a word to be automatically separated, as it conflates them into a single vector.We address this issue by proposing a new model which learns word and sense embeddings jointly.Our model exploits large corpora and knowledge from semantic networks in order to produce a unified vector space of word and sense embeddings.We evaluate the main features of our approach both qualitatively and quantitatively in a variety of tasks, highlighting the advantages of the proposed method in comparison to stateof-the-art word-and sense-based models.
Massimiliano Mancini, José Camacho-Collados, Ignacio Iacobacci, Roberto Navigli
CoNLL2
2017 Word Sense Disambiguation: A Unified Evaluation Framework and Empirical Comparison
abstract
Word Sense Disambiguation is a longstanding task in Natural Language Processing, lying at the core of human language understanding.However, the evaluation of automatic systems has been problematic, mainly due to the lack of a reliable evaluation framework.In this paper we develop a unified evaluation framework and analyze the performance of various Word Sense Disambiguation systems in a fair setup.The results show that supervised systems clearly outperform knowledge-based models.Among the supervised systems, a linear classifier trained on conventional local features still proves to be a hard baseline to beat.Nonetheless, recent approaches exploiting neural networks on unlabeled corpora achieve promising results, surpassing this hard baseline in most test sets.
Alessandro Raganato, José Camacho-Collados, Roberto Navigli
EACL (1)2
2016 Extending WordNet with Fine-Grained Collocational Information via Supervised Distributional Learning
abstract
WordNet is probably the best known lexical resource in Natural Language Processing. While it is widely regarded as a high quality repository of concepts and semantic relations, updating and extending it manually is costly. One important type of relation which could potentially add enormous value to WordNet is the inclusion of collocational information, which is paramount in tasks such as Machine Translation, Natural Language Generation and Second Language Learning. In this paper, we present ColWordNet (CWN), an extended WordNet version with fine-grained collocational information, automatically introduced thanks to a method exploiting linear relations between analogous sense-level embeddings spaces. We perform both intrinsic and extrinsic evaluations, and release CWN for the use and scrutiny of the community.
Luis Espinosa Anke, José Camacho-Collados, Sara Rodríguez-Fernández, Horacio Saggion, Leo Wanner
COLING2
2016 Supervised Distributional Hypernym Discovery via Domain Adaptation
abstract
Comunicació presentada a la Conference on Empirical Methods in Natural Language Processing celebrada els dies 1 a 5 de novembre de 2016 a Austin, Texas.
Luis Espinosa Anke, José Camacho-Collados, Claudio Delli Bovi, Horacio Saggion
EMNLP2
2016 A Large-Scale Multilingual Disambiguation of Glosses
José Camacho-Collados, Claudio Delli Bovi, Alessandro Raganato, Roberto Navigli
LREC1
2016 Nasari: Integrating explicit knowledge and corpus statistics for a multilingual representation of concepts and entities
José Camacho-Collados, Mohammad Taher Pilehvar, Roberto Navigli
Artif. Intell.1
2015 A Unified Multilingual Semantic Representation of Concepts
abstract
José Camacho-Collados, Mohammad Taher Pilehvar, Roberto Navigli. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
José Camacho-Collados, Mohammad Taher Pilehvar, Roberto Navigli
ACL (1)1
2015 NASARI: a Novel Approach to a Semantically-Aware Representation of Items
abstract
José Camacho-Collados, Mohammad Taher Pilehvar, Roberto Navigli. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015.
José Camacho-Collados, Mohammad Taher Pilehvar, Roberto Navigli
HLT-NAACL1