EDBT 2026 Demo / reviewers in the wild / expert
Eneko Agirre
dblp:a/EnekoAgirre
· DBLP profile ↗
110ranked-venue papers
26as first author
20since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 99 · 23 first-author · 18 since 2021Databases, data management, data science and information retrieval · 10 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Emergent Abilities of Large Language Models under Continued Pre-training for Language AdaptationabstractContinued pretraining (CPT) is a popular approach to adapt existing large language models (LLMs) to new languages.When doing so, it is common practice to include a portion of English data in the mixture, but its role has not been carefully studied to date.In this work, we show that including English does not impact validation perplexity, yet it is critical for the emergence of downstream capabilities in the target language.We introduce a language-agnostic benchmark for in-context learning (ICL), which reveals catastrophic forgetting early on CPT when English is not included.This in turn damages the ability of the model to generalize to downstream prompts in the target language as measured by perplexity, even if it does not manifest in terms of accuracy until later in training, and can be tied to a big shift in the model parameters.Based on these insights, we introduce curriculum learning and exponential moving average (EMA) of weights as effective alternatives to mitigate the need for English.All in all, our work sheds light into the dynamics by which emergent abilities arise when doing CPT for language adaptation, and can serve as a foundation to design more effective methods in the future. Ahmed Elhady, Eneko Agirre, Mikel Artetxe |
ACL (1) | 2 |
| 2025 | Instructing Large Language Models for Low-Resource Languages: A Systematic Study for BasqueabstractOscar Sainz, Naiara Perez, Julen Etxaniz, Joseba Fernandez de Landa, Itziar Aldabe, Iker García-Ferrero, Aimar Zabala, Ekhi Azurmendi, German Rigau, Eneko Agirre, Mikel Artetxe, Aitor Soroa. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Oscar Sainz, Naiara Pérez, Julen Etxaniz, Joseba Fernandez de Landa, Itziar Aldabe, Iker García-Ferrero, Aimar Zabala, Ekhi Azurmendi, German Rigau, Eneko Agirre, Mikel Artetxe, Aitor Soroa |
EMNLP | 10 |
| 2024 | PixT3: Pixel-based Table-To-Text GenerationabstractTable -to-text generation involves generating appropriate textual descriptions given structured tabular data.It has attracted increasing attention in recent years thanks to the popularity of neural network models and the availability of large-scale datasets.A common feature across existing methods is their treatment of the input as a string, i.e., by employing linearization techniques that do not always preserve information in the table, are verbose, and lack space efficiency.We propose to rethink data-to-text generation as a visual recognition task, removing the need for rendering the input in a string format.We present PixT3, a multimodal tableto-text model that overcomes the challenges of linearization and input size limitations encountered by existing models.PixT3 is trained with a new self-supervised learning objective to reinforce table structure awareness and is applicable to open-ended and controlled generation settings.Experiments on the ToTTo (Parikh et al., 2020a) and Logic2Text (Chen et al., 2020c) benchmarks show that PixT3 is competitive and, in some settings, superior to generators that operate solely on text. 1 Iñigo Alonso 0001, Eneko Agirre, Mirella Lapata |
ACL (1) | 2 |
| 2024 | Latxa: An Open Language Model and Evaluation Suite for BasqueabstractJulen Etxaniz, Oscar Sainz, Naiara Perez, Itziar Aldabe, German Rigau, Eneko Agirre, Aitor Ormazabal, Mikel Artetxe, Aitor Soroa. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Julen Etxaniz, Oscar Sainz, Naiara Miguel, Itziar Aldabe, German Rigau, Eneko Agirre, Aitor Ormazabal, Mikel Artetxe, Aitor Soroa |
ACL (1) | 6 |
| 2024 | Event Extraction in Basque: Typologically Motivated Cross-Lingual Transfer-Learning AnalysisabstractCross-lingual transfer-learning is widely used in Event Extraction for low-resource languages and involves a Multilingual Language Model that is trained in a source language and applied to the target language. This paper studies whether the typological similarity between source and target languages impacts the performance of cross-lingual transfer, an under-explored topic. We first focus on Basque as the target language, which is an ideal target language because it is typologically different from surrounding languages. Our experiments on three Event Extraction tasks show that the shared linguistic characteristic between source and target languages does have an impact on transfer quality. Further analysis of 72 language pairs reveals that for tasks that involve token classification such as entity and event trigger identification, common writing script and morphological features produce higher quality cross-lingual transfer. In contrast, for tasks involving structural prediction like argument extraction, common word order is the most relevant feature. In addition, we show that when increasing the training size, not all the languages scale in the same way in the cross-lingual setting. To perform the experiments we introduce EusIE, an event extraction dataset for Basque, which follows the Multilingual Event Extraction dataset (MEE). The dataset and code are publicly available. Mikel Zubillaga, Oscar Sainz, Ainara Estarrona, Oier Lopez de Lacalle, Eneko Agirre |
LREC/COLING | 5 |
| 2024 | GoLLIE: Annotation Guidelines improve Zero-Shot Information-ExtractionabstractLarge Language Models (LLMs) combined with instruction tuning have made significant progress when generalizing to unseen tasks. However, they have been less successful in Information Extraction (IE), lagging behind task-specific models. Typically, IE tasks are characterized by complex annotation guidelines which describe the task and give examples to humans. Previous attempts to leverage such information have failed, even with the largest models, as they are not able to follow the guidelines out-of-the-box. In this paper we propose GoLLIE (Guideline-following Large Language Model for IE), a model able to improve zero-shot results on unseen IE tasks by virtue of being fine-tuned to comply with annotation guidelines. Comprehensive evaluation empirically demonstrates that GoLLIE is able to generalize to and follow unseen guidelines, outperforming previous attempts at zero-shot information extraction. The ablation study shows that detailed guidelines is key for good results. Code, data and models will be made publicly available. Oscar Sainz, Iker García-Ferrero, Rodrigo Agerri, Oier Lopez de Lacalle, German Rigau, Eneko Agirre |
ICLR | 6 |
| 2024 | BiVLC: Extending Vision-Language Compositionality Evaluation with Text-to-Image RetrievalabstractExisting Vision-Language Compositionality (VLC) benchmarks like SugarCrepe are formulated as image-to-text retrieval problems, where, given an image, the models need to select between the correct textual description and a synthetic hard negative text. In this work, we present the Bidirectional Vision-Language Compositionality (BiVLC) dataset. The novelty of BiVLC is to add a synthetic hard negative image generated from the synthetic text, resulting in two image-to-text retrieval examples (one for each image) and, more importantly, two text-to-image retrieval examples (one for each text). Human annotators filter out ill-formed examples ensuring the validity of the benchmark. The experiments on BiVLC uncover a weakness of current multimodal models, as they perform poorly in the text-to-image direction. In fact, when considering both retrieval directions, the conclusions obtained in previous works change significantly. In addition to the benchmark, weshow that a contrastive model trained using synthetic images and texts significantly improves over the base model in SugarCrepe and in BiVLC for both retrieval directions. The gap to human performance in BiVLC confirms that Vision-Language Compositionality is still a challenging problem. Imanol Miranda, Ander Salaberria, Eneko Agirre, Gorka Azkune |
NeurIPS | 3 |
| 2024 | Automatic Logical Forms improve fidelity in Table-to-Text generation
Iñigo Alonso 0001, Eneko Agirre |
Expert Syst. Appl. | 2 |
| 2024 | Grounding spatial relations in text-only language modelsabstractThis paper shows that text-only Language Models (LM) can learn to ground spatial relations like left of or below if they are provided with explicit location information of objects and they are properly trained to leverage those locations. We perform experiments on a verbalized version of the Visual Spatial Reasoning (VSR) dataset, where images are coupled with textual statements which contain real or fake spatial relations between two objects of the image. We verbalize the images using an off-the-shelf object detector, adding location tokens to every object label to represent their bounding boxes in textual form. Given the small size of VSR, we do not observe any improvement when using locations, but pretraining the LM over a synthetic dataset automatically derived by us improves results significantly when using location tokens. We thus show that locations allow LMs to ground spatial relations, with our text-only LMs outperforming Vision-and-Language Models and setting the new state-of-the-art for the VSR dataset. Our analysis show that our text-only LMs can generalize beyond the relations seen in the synthetic dataset to some extent, learning also more useful information than that encoded in the spatial rules we used to create the synthetic dataset itself. Gorka Azkune, Ander Salaberria, Eneko Agirre |
Neural Networks | 3 |
| 2023 | CombLM: Adapting Black-Box Language Models through Small Fine-Tuned ModelsabstractMethods for adapting language models (LMs) to new tasks and domains have traditionally assumed white-box access to the model, and work by modifying its parameters.However, this is incompatible with a recent trend in the field, where the highest quality models are only available as black-boxes through inference APIs.Even when the model weights are available, the computational cost of fine-tuning large LMs can be prohibitive for most practitioners.In this work, we present a lightweight method for adapting large LMs to new domains and tasks, assuming no access to their weights or intermediate activations.Our approach fine-tunes a small white-box LM and combines it with the large black-box LM at the probability level through a small network, learned on a small validation set.We validate our approach by adapting a large LM (OPT-30B) to several domains and a downstream task (machine translation), observing improved performance in all cases, of up to 9%, while using a domain expert 23x smaller. Aitor Ormazabal, Mikel Artetxe, Eneko Agirre |
EMNLP | 3 |
| 2023 | What do Language Models know about word senses? Zero-Shot WSD with Language Models and Domain InventoriesabstractLanguage Models are the core for almost any Natural Language Processing system nowadays.One of their particularities is their contextualized representations, a game changer feature when a disambiguation between word senses is necessary.In this paper we aim to explore to what extent language models are capable of discerning among senses at inference time.We performed this analysis by prompting commonly used Languages Models such as BERT or RoBERTa to perform the task of Word Sense Disambiguation (WSD).We leverage the relation between word senses and domains, and cast WSD as a textual entailment problem, where the different hypothesis refer to the domains of the word senses.Our results show that this approach is indeed effective, close to supervised systems. Oscar Sainz, Oier Lopez de Lacalle, Eneko Agirre, German Rigau |
GWC | 3 |
| 2023 | Image captioning for effective use of language models in knowledge-based visual question answeringabstractIntegrating outside knowledge for reasoning in visio-linguistic tasks such as visual question answering (VQA) is an open problem. Given that pretrained language models have been shown to include world knowledge, we propose to use a unimodal (text-only) train and inference procedure based on automatic off-the-shelf captioning of images and pretrained language models. More specifically, we verbalize the image contents and allow language models to better leverage their implicit knowledge to solve knowledge-intensive tasks. Focusing on a visual question answering task which requires external knowledge (OK-VQA), our contributions are: (i) a text-only model that outperforms pretrained multimodal (image-text) models of comparable number of parameters; (ii) confirmation that our text-only method is specially effective for tasks requiring external knowledge, as it is less effective in standard a VQA task (VQA 2.0); and (iii) our method attains results in the state-of-the-art when increasing the size of the language model. We also significantly outperform current multimodal systems, even though augmented with external knowledge. Our qualitative analysis on OK-VQA reveals that automatic captions often fail to capture relevant information in the images, which seems to be balanced by the better inference ability of the text-only language models. Our work opens up possibilities to further improve inference in visio-linguistic tasks. Ander Salaberria, Gorka Azkune, Oier Lopez de Lacalle, Aitor Soroa, Eneko Agirre |
Expert Syst. Appl. | 5 |
| 2022 | Principled Paraphrase Generation with Parallel CorporaabstractRound-trip Machine Translation (MT) is a popular choice for paraphrase generation, which leverages readily available parallel corpora for supervision.In this paper, we formalize the implicit similarity function induced by this approach, and show that it is susceptible to nonparaphrase pairs sharing a single ambiguous translation.Based on these insights, we design an alternative similarity metric that mitigates this issue by requiring the entire translation distribution to match, and implement a relaxation of it through the Information Bottleneck method.Our approach incorporates an adversarial term into MT training in order to learn representations that encode as much information about the reference translation as possible, while keeping as little information about the input as possible.Paraphrases can be generated by decoding back to the source from this representation, without having to generate pivot translations.In addition to being more principled and efficient than round-trip MT, our approach offers an adjustable parameter to control the fidelity-diversity trade-off, and obtains better results in our experiments. Aitor Ormazabal, Mikel Artetxe, Aitor Soroa, Gorka Labaka, Eneko Agirre |
ACL (1) | 5 |
| 2022 | Few-shot Information Extraction is Here: Pre-train, Prompt and EntailabstractDeep Learning has made tremendous progress in Natural Language Processing (NLP), where large pre-trained language models (PLM) fine-tuned on the target task have become the predominant tool. More recently, in a process called prompting, NLP tasks are rephrased as natural language text, allowing us to better exploit linguistic knowledge learned by PLMs and resulting in significant improvements. Still, PLMs have limited inference ability. In the Textual Entailment task, systems need to output whether the truth of a certain textual hypothesis follows from the given premise text. Manually annotated entailment datasets covering multiple inference phenomena have been used to infuse inference capabilities to PLMs. Eneko Agirre |
SIGIR | 1 |
| 2022 | Information retrieval and question answering: A case study on COVID-19 scientific literature
Arantxa Otegi, Iñaki San Vicente, Xabier Saralegi, Anselmo Peñas, Borja Lozano, Eneko Agirre |
Knowl. Based Syst. | 6 |
| 2021 | Beyond Offline Mapping: Learning Cross-lingual Word Embeddings through Context AnchoringabstractRecent research on cross-lingual word embeddings has been dominated by unsupervised mapping approaches that align monolingual embeddings. Such methods critically rely on those embeddings having a similar structure, but it was recently shown that the separate training in different languages causes departures from this assumption. In this paper, we propose an alternative approach that does not have this limitation, while requiring a weak seed dictionary (e.g., a list of identical words) as the only form of supervision. Rather than aligning two fixed embedding spaces, our method works by fixing the target language embeddings, and learning a new set of embeddings for the source language that are aligned with them. To that end, we use an extension of skip-gram that leverages translated context words as anchor points, and incorporates self-learning and iterative restarts to reduce the dependency on the initial dictionary. Our approach outperforms conventional mapping methods on bilingual lexicon induction, and obtains competitive results in the downstream XNLI task. Aitor Ormazabal, Mikel Artetxe, Aitor Soroa, Gorka Labaka, Eneko Agirre |
ACL/IJCNLP (1) | 5 |
| 2021 | Label Verbalization and Entailment for Effective Zero and Few-Shot Relation ExtractionabstractRelation extraction systems require large amounts of labeled examples which are costly to annotate.In this work we reformulate relation extraction as an entailment task, with simple, hand-made, verbalizations of relations produced in less than 15 minutes per relation.The system relies on a pretrained textual entailment engine which is run as-is (no training examples, zero-shot) or further fine-tuned on labeled examples (few-shot or fully trained).In our experiments on TACRED we attain 63% F1 zero-shot, 69% with 16 examples per relation (17% points better than the best supervised system on the same conditions), and only 4 points short of the state-of-the-art (which uses 20 times more training data).We also show that the performance can be improved significantly with larger entailment models, up to 12 points in zero-shot, giving the best results to date on TACRED when fully trained.The analysis shows that our few-shot systems are especially effective when discriminating between relations, and that the performance difference in low data regimes comes mainly from identifying no-relation cases. Oscar Sainz, Oier Lopez de Lacalle, Gorka Labaka, Ander Barrena, Eneko Agirre |
EMNLP (1) | 5 |
| 2021 | Towards zero-shot cross-lingual named entity disambiguationabstractIn cross-Lingual Named Entity Disambiguation (XNED) the task is to link Named Entity mentions in text in some native language to English entities in a knowledge graph. XNED systems usually require training data for each native language, limiting their application for low resource languages with small amounts of training data. Prior work have proposed so-called zero-shot transfer systems which are only trained in English training data, but required native prior probabilities of entities with respect to mentions, which had to be estimated from native training examples, limiting their practical interest. In this work we present a zero-shot XNED architecture where, instead of a single disambiguation model, we have a model for each possible mention string, thus eliminating the need for native prior probabilities. Our system improves over prior work in XNED datasets in Spanish and Chinese by 32 and 27 points, and matches the systems which do require native prior information. We experiment with different multilingual transfer strategies, showing that better results are obtained with a purpose-built multilingual pre-training method compared to state-of-the-art generic multilingual models such as XLM-R. We also discovered, surprisingly, that English is not necessarily the most effective zero-shot training language for XNED into English. For instance, Spanish is more effective when training a zero-shot XNED system that disambiguates Basque mentions with respect to an English knowledge graph. Ander Barrena, Aitor Soroa, Eneko Agirre |
Expert Syst. Appl. | 3 |
| 2021 | A large reproducible benchmark of ontology-based methods and word embeddings for word similarity
Juan J. Lastra-Díaz, Josu Goikoetxea, Mohamed Ali Hadj Taieb, Ana García-Serrano, Mohamed Benaouicha, Eneko Agirre, David Sánchez 0001 |
Inf. Syst. | 6 |
| 2021 | Inferring spatial relations from textual descriptions of images
Aitzol Elu, Gorka Azkune, Oier Lopez de Lacalle, Ignacio Arganda-Carreras, Aitor Soroa, Eneko Agirre |
Pattern Recognit. | 6 |
| 2020 | A Call for More Rigor in Unsupervised Cross-lingual LearningabstractWe review motivations, definition, approaches, and methodology for unsupervised crosslingual learning and call for a more rigorous position in each of them.An existing rationale for such research is based on the lack of parallel data for many of the world's languages.However, we argue that a scenario without any parallel data and abundant monolingual data is unrealistic in practice.We also discuss different training signals that have been used in previous work, which depart from the pure unsupervised setting.We then describe common methodological issues in tuning and evaluation of unsupervised cross-lingual models and present best practices.Finally, we provide a unified outlook for different types of research in this area (i.e., cross-lingual word embeddings, deep multilingual pretraining, and unsupervised machine translation) and argue for comparable evaluation of these models. Mikel Artetxe, Sebastian Ruder, Dani Yogatama, Gorka Labaka, Eneko Agirre |
ACL | 5 |
| 2020 | DoQA - Accessing Domain-Specific FAQs via Conversational QAabstractThe goal of this work is to build conversational Question Answering (QA) interfaces for the large body of domain-specific information available in FAQ sites.We present DoQA, a dataset with 2,437 dialogues and 10,917 QA pairs.The dialogues are collected from three Stack Exchange sites using the Wizard of Oz method with crowdsourcing.Compared to previous work, DoQA comprises well-defined information needs, leading to more coherent and natural conversations with less factoid questions and is multi-domain.In addition, we introduce a more realistic information retrieval (IR) scenario where the system needs to find the answer in any of the FAQ documents.The results of an existing, strong, system show that, thanks to transfer learning from a Wikipedia QA dataset and fine tuning on a single FAQ domain, it is possible to build high quality conversational QA systems for FAQs without indomain training data.The good results carry over into the more challenging IR scenario.In both cases, there is still ample room for improvement, as indicated by the higher human upperbound. Jon Ander Campos, Arantxa Otegi, Aitor Soroa, Jan Deriu, Mark Cieliebak, Eneko Agirre |
ACL | 6 |
| 2020 | A Methodology for Creating Question Answering Corpora Using Inverse Data AnnotationabstractJan Deriu, Katsiaryna Mlynchyk, Philippe Schläpfer, Alvaro Rodrigo, Dirk von Grünigen, Nicolas Kaiser, Kurt Stockinger, Eneko Agirre, Mark Cieliebak. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020. Jan Deriu, Katsiaryna Mlynchyk, Philippe Schläpfer, Álvaro Rodrigo, Dirk Von Gruenigen, Nicolas Kaiser, Kurt Stockinger, Eneko Agirre, Mark Cieliebak |
ACL | 8 |
| 2020 | Improving Conversational Question Answering Systems after Deployment using Feedback-Weighted LearningabstractThe interaction of conversational systems with users poses an exciting opportunity for improving them after deployment, but little evidence has been provided of its feasibility.In most applications, users are not able to provide the correct answer to the system, but they are able to provide binary (correct, incorrect) feedback.In this paper we propose feedback-weighted learning based on importance sampling to improve upon an initial supervised system using binary user feedback.We perform simulated experiments on document classification (for development) and Conversational Question Answering datasets like QuAC and DoQA, where binary user feedback is derived from gold annotations.The results show that our method is able to improve over the initial supervised system, getting close to a fully-supervised system that has access to the same labeled examples in in-domain experiments (QuAC), and even matching in out-of-domain experiments (DoQA).Our work opens the prospect to exploit interactions with real users and improve conversational systems after deployment. Jon Ander Campos, Kyunghyun Cho, Arantxa Otegi, Aitor Soroa, Eneko Agirre, Gorka Azkune |
COLING | 5 |
| 2020 | Evaluating Multimodal Representations on Visual Semantic Textual SimilarityabstractThe combination of visual and textual representations has produced excellent results in tasks such as image captioning and visual question answering, but the inference capabilities of multimodal representations are largely untested. In the case of textual representations, inference tasks such as Textual Entailment and Semantic Textual Similarity have been often used to benchmark the quality of textual representations. The long term goal of our research is to devise multimodal representation techniques that improve current inference capabilities. We thus present a novel task, Visual Semantic Textual Similarity (vSTS), where such inference ability can be tested directly. Given two items comprised each by an image and its accompanying caption, vSTS systems need to assess the degree to which the captions in context are semantically equivalent to each other. Our experiments using simple multimodal representations show that the addition of image representations produces better inference, compared to text-only representations. The improvement is observed both when directly computing the similarity between the representations of the two items, and when learning a siamese network based on vSTS training data. Our work shows, for the first time, the successful contribution of visual information to textual inference, with ample room for benchmarking more complex multimodal representation options. Oier Lopez de Lacalle, Ander Salaberria, Aitor Soroa, Gorka Azkune, Eneko Agirre |
ECAI | 5 |
| 2020 | Translation Artifacts in Cross-lingual Transfer LearningabstractBoth human and machine translation play a central role in cross-lingual transfer learning: many multilingual datasets have been created through professional translation services, and using machine translation to translate either the test set or the training set is a widely used transfer technique.In this paper, we show that such translation process can introduce subtle artifacts that have a notable impact in existing cross-lingual models.For instance, in natural language inference, translating the premise and the hypothesis independently can reduce the lexical overlap between them, which current models are highly sensitive to.We show that some previous findings in cross-lingual transfer learning need to be reconsidered in the light of this phenomenon.Based on the gained insights, we also improve the state-of-the-art in XNLI for the translate-test and zero-shot approaches by 4.3 and 2.8 points, respectively.3 We use FastAlign (Dyer et al., 2013) for word alignment, and discard the few questions for which the mapping method fails (when none of the tokens in the answer span are aligned).4 We use the same procedure as for the training set except that (i) given the small size of the test set, we combine it with WikiMatrix (Schwenk et al., 2019) to aid word alignment, (ii) we use Jieba for Chinese segmentation instead of the Moses tokenizer, and (iii) for the few unaligned spans, we return the English answer. Mikel Artetxe, Gorka Labaka, Eneko Agirre |
EMNLP (1) | 3 |
| 2020 | Spot The Bot: A Robust and Efficient Framework for the Evaluation of Conversational Dialogue SystemsabstractJan Deriu, Don Tuggener, Pius von Däniken, Jon Ander Campos, Alvaro Rodrigo, Thiziri Belkacem, Aitor Soroa, Eneko Agirre, Mark Cieliebak. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020. Jan Deriu, Don Tuggener, Pius von Däniken, Jon Ander Campos, Álvaro Rodrigo, Thiziri Belkacem, Aitor Soroa, Eneko Agirre, Mark Cieliebak |
EMNLP (1) | 8 |
| 2020 | Give your Text Representation Models some Love: the Case for BasqueabstractWord embeddings and pre-trained language models allow to build rich representations of text and have enabled improvements across most NLP tasks. Unfortunately they are very expensive to train, and many small companies and research groups tend to use models that have been pre-trained and made available by third parties, rather than building their own. This is suboptimal as, for many languages, the models have been trained on smaller (or lower quality) corpora. In addition, monolingual pre-trained models for non-English languages are not always available. At best, models for those languages are included in multilingual versions, where each language shares the quota of substrings and parameters with the rest of the languages. This is particularly true for smaller languages such as Basque. In this paper we show that a number of monolingual models (FastText word embeddings, FLAIR and BERT language models) trained with larger Basque corpora produce much better results than publicly available versions in downstream NLP tasks, including topic classification, sentiment classification, PoS tagging and NER. This work sets a new state-of-the-art in those tasks for Basque. All benchmarks and models used in this work are publicly available. Rodrigo Agerri, Iñaki San Vicente, Jon Ander Campos, Ander Barrena, Xabier Saralegi, Aitor Soroa, Eneko Agirre |
LREC | 7 |
| 2020 | Conversational Question Answering in Low Resource Scenarios: A Dataset and Case Study for BasqueabstractConversational Question Answering (CQA) systems meet user information needs by having conversations with them, where answers to the questions are retrieved from text. There exist a variety of datasets for English, with tens of thousands of training examples, and pre-trained language models have allowed to obtain impressive results. The goal of our research is to test the performance of CQA systems under low-resource conditions which are common for most non-English languages: small amounts of native annotations and other limitations linked to low resource languages, like lack of crowdworkers or smaller wikipedias. We focus on the Basque language, and present the first non-English CQA dataset and results. Our experiments show that it is possible to obtain good results with low amounts of native data thanks to cross-lingual transfer, with quality comparable to those obtained for English. We also discovered that dialogue history models are not directly transferable to another language, calling for further research. The dataset is publicly available. Arantxa Otegi, Aitor Gonzalez-Agirre, Jon Ander Campos, Aitor Soroa, Eneko Agirre |
LREC | 5 |
| 2020 | Cross-Lingual Word EmbeddingsabstractThe representation of words across languages is of interest since the early days of interlingual machine translation, as it allows us to connect the meaning of words in different languages and to generalize lexical semantic properties and relations across languages (Hutchins 2000). Structured representations such as multilingual lexical knowledge bases represent polysemy, as well as language internal and cross-lingual relations, but they require costly manual construction and maintenance (Vossen 1998). Alternatively, corpus-based methods have been used to automatically induce monolingual word representations like word embeddings with great success (Mikolov et al. 2013). Word embeddings represent the words in the vocabulary of a language as vectors in n-dimensional space, where words that are similar being located close to each other. Cross-lingual word embeddings (CLWE for short) extend the idea, and represent translation-equivalent words from two (or more) languages close to each other in a common, cross-lingual space.The interest in cross-lingual word embeddings has grown in recent years. This is partly because of their success in cross-lingual transfer, where NLP tools trained in a resource-rich language such as English are transferred to another language with smaller or no annotated data. For instance, given training data for a text-classification task in English, a model using CLWE can classify foreign language documents. Beyond language pairs, CLWE allows us to represent words of several languages in a common space, and thus pave the way to build multilingual NLP tools that use the same model to process text in different languages.This comprehensive and, at the same time, dense book has been written by Anders Søgaard, Ivan Vulić, Sebastian Ruder, and Manaal Faruqui. It covers all key issues as well as the most relevant work in CLWE, including the most recent research (up to May 2019) in this vibrant research area. It does a great job of organizing different approaches in a typology, according to the kind of bilingual resources needed, and differentiating word-level, sentence-level, and document-level models. The book also covers extensions to CLWE that are able to represent multiple languages in the same space, as well as unsupervised learning, where the systems only use monolingual resources to build the cross-lingual space. The book is structured in 12 chapters.Chapter 1 is a brief introduction, which includes an explanation of the notation used in the book. The book tries to establish a formal relation between several word-level alignment models, and it does a thorough job of describing the methods using a consistent mathematical formalization. This consistency allows the authors to describe methods in a more compact way and helps to better see the common patterns across seemingly different approaches. In this sense, introducing the notation in the first chapter makes life easier for later. The formulas are quite dense and demanding, but although a mathematically naive reader might have a hard time following them, casual readers can also get value from a higher level read of the book.Chapter 2 makes a brief introduction of the main monolingual word embedding models, focusing on the formalization of the loss function being optimized by each of the models.Chapter 3 introduces a typology of supervised CLWE models, that is, methods that use some kind of bilingual signal. The typology leaves aside multilingual and unsupervised methods, which are covered in later chapters. The typology is based on data requirements along two dimensions: the type of bilingual signal (at the level of words, sentences, or documents), and whether the method requires parallel resources, or comparable resources suffice.Chapter 4 introduces work on cross-lingual word representations that pre-dates the introduction of word embeddings. The chapter covers work on cross-lingual clusters, delexicalization strategies for cross-lingual transfer, earlier use of seed dictionaries, together with distributional vector space models, cross-lingual word alignment in machine translation, and latent cross-lingual concepts. This chapter draws connections with earlier work, and is a must-read for anyone wanting to take a step back from current techniques, to look at the big picture and draw inspiration in the larger picture of (non-embbeding-related) NLP methods.The book then follows with three chapters organized according to typology. Chapter 5 covers the most popular family of models, those based on word-level information. The models that require parallel data in the form of bilingual dictionaries (or word alignments induced from parallel corpora) are further classified into those that learn separate monolingual spaces for each language to then learn a mapping, or those that learn the cross-lingual space for the two languages jointly, as well as mixed approaches. The authors make an effort to show that some methods coming from mapping, joint, and mixed approaches are very similar. In addition, the chapter also covers methods that ground words into images or image features. Most of the space in the chapter is taken by mapping methods, as they take the bulk of recent publications.Chapter 6 introduces sentence-level information, usually in the form of sentence-aligned parallel translations. The additional supervision is used to learn either shared sentence representations, bilingual encoders, or a bilingual version of the monolingual skip-gram loss.Chapter 7 introduces methods that use information from comparable documents only, which offers less supervision compared with the methods in the previous two chapters. These methods typically use Wikipedia articles from different languages as comparable documents.Chapter 8 introduces multilingual CLWE, where more than two languages are involved. Apart from the practical interest, some works show that the use of multiple languages improves the quality of word embeddings. Most works use bilingual CLWE learning methods taking a pivot language (e.g., English).Chapter 9 is devoted to unsupervised methods, that is, those that learn the CLWE space without any bilingual information. At the core, unsupervised methods apply one of the supervised methods (e.g., a word-level mapping method from Chapter 5). An initial small or low-quality seed lexicon is produced using some method, and iteratively, better dictionaries are obtained and used as seed lexicons.Chapter 10 gathers a wide array of applications and intrinsic evaluation tasks that have been used across the literature. Arguably, bilingual dictionary induction is the reference evaluation task for word-level mapping algorithms, but other methods have chosen to evaluate on other tasks, making comparison across methods in the same family difficult, and comparison across types of methods unfeasible. Note that the book does not provide information about the experimental performance of the methods, which is understandable, given the lack of an agreed-upon evaluation task or data set that covers all methods in the typology. That said, some experimental evaluation information, although limited, would have been of interest to the reader, as it would allow one to have a grasp of the relative standing of some relevant methods introduced in the book.On the practical side, Chapter 11 introduces a comprehensive list of monolingual corpora and embeddings, bilingual dictionaries, parallel corpora, and CLWE open source models. It also lists some relevant evaluation data sets and applications.The book finishes in Chapter 12 with general challenges and future direction, where the authors outline some of the current challenges in this field, alongside some specific open problems.In summary, this book provides a comprehensive and in-depth overview to cross-lingual word embeddings, covering the breadth of techniques and resources used. This book is recommended not only to researchers, students, and practitioners who work on the area of cross-lingual and multilingual word embeddings, but also to a wider range of readers who have interest in cross-lingual and multilingual processing.Note that the book is very similar to a contemporaneous journal survey (Ruder, Vulić, and Søgaard 2019), written by three of the authors. The book is more detailed and contains a separate section on unsupervised methods and another section with useful data and software. On the other hand, the journal survey contains a handful of newer references, up through ACL 2019. The field is moving forward fast, and seeing the latest developments (at the time of writing this review), it seems the authors chose a very good time to publish the book. The irruption of contextual word embedding models, where the representation of a word depends on the context of occurrence, has put on the table a new family of alternative methods that learn cross-lingual word representations. It is good to see that one of the authors has already started to cover some of the newer methods in an excellent blog.1 Eneko Agirre |
Comput. Linguistics | 1 |
| 2020 | Cross-environment activity recognition using word embeddings for sensor and activity representation
Gorka Azkune, Aitor Almeida, Eneko Agirre |
Neurocomputing | 3 |
| 2020 | Do all roads lead to Rome? Understanding the role of initialization in iterative back-translation
Mikel Artetxe, Gorka Labaka, Noe Casas, Eneko Agirre |
Knowl. Based Syst. | 4 |
| 2019 | An Effective Approach to Unsupervised Machine TranslationabstractWhile machine translation has traditionally relied on large amounts of parallel corpora, a recent research line has managed to train both Neural Machine Translation (NMT) and Statistical Machine Translation (SMT) systems using monolingual corpora only. In this paper, we identify and address several deficiencies of existing unsupervised SMT approaches by exploiting subword information, developing a theoretically well founded unsupervised tuning method, and incorporating a joint refinement procedure. Moreover, we use our improved SMT system to initialize a dual NMT model, which is further fine-tuned through on-the-fly back-translation. Together, we obtain large improvements over the previous state-of-the-art in unsupervised machine translation. For instance, we get 22.5 BLEU points in English-to-German WMT 2014, 5.5 points more than the previous best unsupervised system, and 0.5 points more than the (supervised) shared task winner back in 2014. Mikel Artetxe, Gorka Labaka, Eneko Agirre |
ACL (1) | 3 |
| 2019 | Bilingual Lexicon Induction through Unsupervised Machine TranslationabstractA recent research line has obtained strong results on bilingual lexicon induction by aligning independently trained word embeddings in two languages and using the resulting crosslingual embeddings to induce word translation pairs through nearest neighbor or related retrieval methods.In this paper, we propose an alternative approach to this problem that builds on the recent work on unsupervised machine translation.This way, instead of directly inducing a bilingual lexicon from cross-lingual embeddings, we use them to build a phrasetable, combine it with a language model, and use the resulting machine translation system to generate a synthetic parallel corpus, from which we extract the bilingual lexicon using statistical word alignment techniques.As such, our method can work with any word embedding and cross-lingual mapping technique, and it does not require any additional resource besides the monolingual corpus used to train the embeddings.When evaluated on the exact same cross-lingual embeddings, our proposed method obtains an average improvement of 6 accuracy points over nearest neighbor and 4 points over CSLS retrieval, establishing a new state-of-the-art in the standard MUSE dataset. Mikel Artetxe, Gorka Labaka, Eneko Agirre |
ACL (1) | 3 |
| 2019 | Analyzing the Limitations of Cross-lingual Word Embedding MappingsabstractRecent research in cross-lingual word embeddings has almost exclusively focused on offline methods, which independently train word embeddings in different languages and map them to a shared space through linear transformations.While several authors have questioned the underlying isomorphism assumption, which states that word embeddings in different languages have approximately the same structure, it is not clear whether this is an inherent limitation of mapping approaches or a more general issue when learning crosslingual embeddings.So as to answer this question, we experiment with parallel corpora, which allows us to compare offline mapping to an extension of skip-gram that jointly learns both embedding spaces.We observe that, under these ideal conditions, joint learning yields to more isomorphic embeddings, is less sensitive to hubness, and obtains stronger results in bilingual lexicon induction.We thus conclude that current mapping methods do have strong limitations, calling for further research to jointly learn cross-lingual embeddings with a weaker cross-lingual signal. Aitor Ormazabal, Mikel Artetxe, Gorka Labaka, Aitor Soroa, Eneko Agirre |
ACL (1) | 5 |
| 2019 | Probing for Semantic Classes: Diagnosing the Meaning Content of Word EmbeddingsabstractWord embeddings typically represent different meanings of a word in a single conflated vector.Empirical analysis of embeddings of ambiguous words is currently limited by the small size of manually annotated resources and by the fact that word senses are treated as unrelated individual concepts.We present a large dataset based on manual Wikipedia annotations and word senses, where word senses from different words are related by semantic classes.This is the basis for novel diagnostic tests for an embedding's content: we probe word embeddings for semantic classes and analyze the embedding space by classifying embeddings into semantic classes.Our main findings are: (i) Information about a sense is generally represented well in a single-vector embedding -if the sense is frequent.(ii) A classifier can accurately predict whether a word is single-sense or multi-sense, based only on its embedding.(iii) Although rare senses are not well represented in single-vector embeddings, this does not have negative impact on an NLP application whose performance depends on frequent senses. Yadollah Yaghoobzadeh, Katharina Kann, Timothy J. Hazen, Eneko Agirre, Hinrich Schütze |
ACL (1) | 4 |
| 2019 | A reproducible survey on word embeddings and ontology-based methods for word similarity: Linear combinations outperform the state of the artabstractHuman similarity and relatedness judgements between concepts underlie most of cognitive capabilities, such as categorisation, memory, decision-making and reasoning. For this reason, the proposal of methods for the estimation of the degree of similarity and relatedness between words and concepts has been a very active line of research in the fields of artificial intelligence, information retrieval and natural language processing among others. Main approaches proposed in the literature can be categorised in two large families as follows: (1) Ontology-based semantic similarity Measures (OM) and (2) distributional measures whose most recent and successful methods are based on Word Embedding (WE) models. However, the lack of a deep analysis of both families of methods slows down the advance of this line of research and its applications. This work introduces the largest, reproducible and detailed experimental survey of OM measures and WE models reported in the literature which is based on the evaluation of both families of methods on a same software platform, with the aim of elucidating what is the state of the problem. We show that WE models which combine distributional and ontology-based information get the best results, and in addition, we show for the first time that a simple average of two best performing WE models with other ontology-based measures or WE models is able to improve the state of the art by a large margin. In addition, we provide a very detailed reproducibility protocol together with a collection of software tools and datasets as supplementary material to allow the exact replication of our results. Juan J. Lastra-Díaz, Josu Goikoetxea, Mohamed Ali Hadj Taieb, Ana García-Serrano, Mohamed Benaouicha, Eneko Agirre |
Eng. Appl. Artif. Intell. | 6 |
| 2019 | Word n-gram attention models for sentence similarity and inference
Iñigo Lopez-Gazpio, Montse Maritxalar, Mirella Lapata, Eneko Agirre |
Expert Syst. Appl. | 4 |
| 2018 | Generalizing and Improving Bilingual Word Embedding Mappings with a Multi-Step Framework of Linear TransformationsabstractUsing a dictionary to map independently trained word embeddings to a shared space has shown to be an effective approach to learn bilingual word embeddings. In this work, we propose a multi-step framework of linear transformations that generalizes a substantial body of previous work. The core step of the framework is an orthogonal transformation, and existing methods can be explained in terms of the additional normalization, whitening, re-weighting, de-whitening and dimensionality reduction steps. This allows us to gain new insights into the behavior of existing methods, including the effectiveness of inverse regression, and design a novel variant that obtains the best published results in zero-shot bilingual lexicon extraction. The corresponding software is released as an open source project. Mikel Artetxe, Gorka Labaka, Eneko Agirre |
AAAI | 3 |
| 2018 | A robust self-learning method for fully unsupervised cross-lingual mappings of word embeddingsabstractRecent work has managed to learn crosslingual word embeddings without parallel data by mapping monolingual embeddings to a shared space through adversarial training.However, their evaluation has focused on favorable conditions, using comparable corpora or closely-related languages, and we show that they often fail in more realistic scenarios.This work proposes an alternative approach based on a fully unsupervised initialization that explicitly exploits the structural similarity of the embeddings, and a robust self-learning algorithm that iteratively improves this solution.Our method succeeds in all tested scenarios and obtains the best published results in standard datasets, even surpassing previous supervised systems.Our implementation is released as an open source project at https://github. com/artetxem/vecmap. Mikel Artetxe, Gorka Labaka, Eneko Agirre |
ACL (1) | 3 |
| 2018 | Uncovering Divergent Linguistic Information in Word Embeddings with Lessons for Intrinsic and Extrinsic EvaluationabstractFollowing the recent success of word embeddings, it has been argued that there is no such thing as an ideal representation for words, as different models tend to capture divergent and often mutually incompatible aspects like semantics/syntax and similarity/relatedness.In this paper, we show that each embedding model captures more information than directly apparent.A linear transformation that adjusts the similarity order of the model without any external resource can tailor it to achieve better results in those aspects, providing a new perspective on how embeddings encode divergent linguistic information.In addition, we explore the relation between intrinsic and extrinsic evaluation, as the effect of our transformations in downstream tasks is higher for unsupervised systems than for supervised ones. Mikel Artetxe, Gorka Labaka, Iñigo Lopez-Gazpio, Eneko Agirre |
CoNLL | 4 |
| 2018 | Learning Text Representations for 500K Classification Tasks on Named Entity DisambiguationabstractNamed Entity Disambiguation algorithms typically learn a single model for all target entities.In this paper we present a word expert model and train separate deep learning models for each target entity string, yielding 500K classification tasks.This gives us the opportunity to benchmark popular text representation alternatives on this massive dataset.In order to face scarce training data we propose a simple data-augmentation technique and transfer-learning.We show that bagof-word-embeddings are better than LSTMs for tasks with scarce training data, while the situation is reversed when having larger amounts.Transferring an LSTM which is learned on all datasets is the most effective context representation option for the word experts in all frequency bands.The experiments show that our system trained on out-ofdomain Wikipedia data surpasses comparable NED systems which have been trained on indomain training data. Ander Barrena, Aitor Soroa, Eneko Agirre |
CoNLL | 3 |
| 2018 | Unsupervised Statistical Machine TranslationabstractWhile modern machine translation has relied on large parallel corpora, a recent line of work has managed to train Neural Machine Translation (NMT) systems from monolingual corpora only (Artetxe et al., 2018c;Lample et al., 2018).Despite the potential of this approach for low-resource settings, existing systems are far behind their supervised counterparts, limiting their practical interest.In this paper, we propose an alternative approach based on phrase-based Statistical Machine Translation (SMT) that significantly closes the gap with supervised systems.Our method profits from the modular architecture of SMT: we first induce a phrase table from monolingual corpora through cross-lingual embedding mappings, combine it with an n-gram language model, and fine-tune hyperparameters through an unsupervised MERT variant.In addition, iterative backtranslation improves results further, yielding, for instance, 14.08 and 26.22 BLEU points in WMT 2014 English-German and English-French, respectively, an improvement of more than 7-10 BLEU points over previous unsupervised systems, and closing the gap with supervised SMT (Moses trained on Europarl) down to 2-5 BLEU points.Our implementation is available at https:// github.com/artetxem/monoses. Mikel Artetxe, Gorka Labaka, Eneko Agirre |
EMNLP | 3 |
| 2018 | Unsupervised Neural Machine Translation
Mikel Artetxe, Gorka Labaka, Eneko Agirre, Kyunghyun Cho |
ICLR (Poster) | 3 |
| 2018 | Bilingual embeddings with random walks over multilingual wordnets
Josu Goikoetxea, Aitor Soroa, Eneko Agirre |
Knowl. Based Syst. | 3 |
| 2017 | Learning bilingual word embeddings with (almost) no bilingual dataabstractMost methods to learn bilingual word embeddings rely on large parallel corpora, which is difficult to obtain for most language pairs.This has motivated an active research line to relax this requirement, with methods that use document-aligned corpora or bilingual dictionaries of a few thousand words instead.In this work, we further reduce the need of bilingual resources using a very simple self-learning approach that can be combined with any dictionary-based mapping technique.Our method exploits the structural similarity of embedding spaces, and works with as little bilingual evidence as a 25 word dictionary or even an automatically generated list of numerals, obtaining results comparable to those of systems that use richer resources. Mikel Artetxe, Gorka Labaka, Eneko Agirre |
ACL (1) | 3 |
| 2017 | Interpretable semantic textual similarity: Finding and explaining differences between sentences
Iñigo Lopez-Gazpio, Montse Maritxalar, Aitor Gonzalez-Agirre, German Rigau, Larraitz Uria, Eneko Agirre |
Knowl. Based Syst. | 6 |
| 2016 | Single or Multiple? Combining Word Representations Independently Learned from Text and WordNetabstractText and Knowledge Bases are complementary sources of information. Given the success of distributed word representations learned from text, several techniques to infuse additional information from sources like WordNet into word representations have been proposed. In this paper, we follow an alternative route. We learn word representations from text and WordNet independently, and then explore simple and sophisticated methods to combine them. The combined representations are applied to an extensive set of datasets on word similarity and relatedness. Simple combination methods happen to perform better that more complex methods like CCA or retrofitting, showing that, in the case of WordNet, learning word representations separately is preferable to learning one single representation space or adding WordNet information directly. A key factor, which we illustrate with examples, is that the WordNet-based representations captures similarity relations encoded in WordNet better than retrofitting. In addition, we show that the average of the similarities from six word representations yields results beyond the state-of-the-art in several datasets, reinforcing the opportunities to explore further combination techniques. Josu Goikoetxea, Eneko Agirre, Aitor Soroa |
AAAI | 2 |
| 2016 | Alleviating Poor Context with Background Knowledge for Named Entity DisambiguationabstractNamed Entity Disambiguation (NED) algorithms disambiguate mentions of named entities with respect to a knowledge-base, but sometimes the context might be poor or misleading.In this paper we introduce the acquisition of two kinds of background information to alleviate that problem: entity similarity and selectional preferences for syntactic positions.We show, using a generative Näive Bayes model for NED, that the additional sources of context are complementary, and improve results in the CoNLL 2003 and TAC KBP DEL 2014 datasets, yielding the third best and the best results, respectively.We provide examples and analysis which show the value of the acquired background information. Ander Barrena, Aitor Soroa, Eneko Agirre |
ACL (1) | 3 |
| 2016 | Improving Translation Selection with SupersensesabstractSelecting appropriate translations for source words with multiple meanings still remains a challenge for statistical machine translation (SMT). One reason for this is that most SMT systems are not good at detecting the proper sense for a polysemic word when it appears in different contexts. In this paper, we adopt a supersense tagging method to annotate source words with coarse-grained ontological concepts. In order to enable the system to choose an appropriate translation for a word or phrase according to the annotated supersense of the word or phrase, we propose two translation models with supersense knowledge: a maximum entropy based model and a supersense embedding model. The effectiveness of our proposed models is validated on a large-scale English-to-Spanish translation task. Results indicate that our method can significantly improve translation quality via correctly conveying the meaning of the source language to the target language. Haiqing Tang, Deyi Xiong, Oier Lopez de Lacalle, Eneko Agirre |
COLING | 4 |
| 2016 | Learning principled bilingual mappings of word embeddings while preserving monolingual invarianceabstractMapping word embeddings of different languages into a single space has multiple applications.In order to map from a source space into a target space, a common approach is to learn a linear mapping that minimizes the distances between equivalences listed in a bilingual dictionary.In this paper, we propose a framework that generalizes previous work, provides an efficient exact method to learn the optimal linear transformation and yields the best bilingual results in translation induction while preserving monolingual performance in an analogy task. Mikel Artetxe, Gorka Labaka, Eneko Agirre |
EMNLP | 3 |
| 2016 | A comparison of Named-Entity Disambiguation and Word Sense Disambiguation
Angel X. Chang, Valentin I. Spitkovsky, Christopher D. Manning, Eneko Agirre |
LREC | 4 |
| 2016 | Word Sense-Aware Machine Translation: Including Senses as Contextual Features for Improved Translation Models
Steven Neale, Luís Gomes 0002, Eneko Agirre, Oier Lopez de Lacalle, António Branco |
LREC | 3 |
| 2016 | QTLeap WSD/NED Corpora: Semantic Annotation of Parallel Corpora in Six Languages
Arantxa Otegi, Nora Aranberri, António Branco, Jan Hajic 0001, Martin Popel, Kiril Ivanov Simov, Eneko Agirre, Petya Osenova, Rita Valadas Pereira, João Silva 0004, Steven Neale |
LREC | 7 |
| 2016 | Addressing the MFS Bias in WSD systems
Marten Postma, Rubén Izquierdo, Eneko Agirre, German Rigau, Piek Vossen |
LREC | 3 |
| 2016 | Evaluating Translation Quality and CLIR Performance of Query Sessions
Xabier Saralegi, Eneko Agirre, Iñaki Alegria |
LREC | 2 |
| 2016 | Why are these similar? Investigating item similarity types in a large digital libraryabstractWe introduce a new problem, identifying the type of relation that holds between a pair of similar items in a digital library. Being able to provide a reason why items are similar has applications in recommendation, personalization, and search. We investigate the problem within the context of Europeana, a large digital library containing items related to cultural heritage. A range of types of similarity in this collection were identified. A set of 1,500 pairs of items from the collection were annotated using crowdsourcing. A high intertagger agreement (average 71.5 Pearson correlation) was obtained and demonstrates that the task is well defined. We also present several approaches to automatically identifying the type of similarity. The best system applies linear regression and achieves a mean Pearson correlation of 71.3, close to human performance. The problem formulation and data set described here were used in a public evaluation exercise, the *SEM shared task on Semantic Textual Similarity. The task attracted the participation of 6 teams, who submitted 14 system runs. All annotations, evaluation scripts, and system runs are freely available. Aitor Gonzalez-Agirre, German Rigau, Eneko Agirre, Nikolaos Aletras, Mark Stevenson 0001 |
J. Assoc. Inf. Sci. Technol. | 3 |
| 2015 | Random Walks and Neural Network Language Models on Knowledge BasesabstractRandom walks over large knowledge bases like WordNet have been successfully used in word similarity, relatedness and disambiguation tasks. Unfortunately, those algorithms are relatively slow for large repositories, with significant memory footprints. In this paper we present a novel algorithm which encodes the structure of a knowledge base in a continuous vector space, combining random walks and neural net language models in order to produce novel word representations. Evaluation in word relatedness and similar- ity datasets yields equal or better results than those of a random walk algorithm, using a dense representation (300 dimensions instead of 117K). Furthermore, the word representations are complementary to those of the random walk algorithm and to corpus-based continuous representations, improving the state- of-the-art in the similarity dataset. Our technique opens up exciting opportunities to combine distributional and knowledge-based word representations. Josu Goikoetxea, Aitor Soroa, Eneko Agirre |
HLT-NAACL | 3 |
| 2015 | Diamonds in the Rough: Event Extraction from Imperfect Microblog DataabstractAnder Intxaurrondo, Eneko Agirre, Oier Lopez de Lacalle, Mihai Surdeanu. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015. Ander Intxaurrondo, Eneko Agirre, Oier Lopez de Lacalle, Mihai Surdeanu |
HLT-NAACL | 2 |
| 2015 | Using knowledge-based relatedness for information retrieval
Arantxa Otegi, Xabier Arregi, Olatz Ansa, Eneko Agirre |
Knowl. Inf. Syst. | 4 |
| 2014 | "One Entity per Discourse" and "One Entity per Collocation" Improve Named-Entity Disambiguation
Ander Barrena, Eneko Agirre, Bernardo Cabaleiro, Anselmo Peñas, Aitor Soroa |
COLING | 2 |
| 2014 | Random Walks for Knowledge-Based Word Sense DisambiguationabstractWord Sense Disambiguation (WSD) systems automatically choose the intended meaning of a word in context. In this article we present a WSD algorithm based on random walks over large Lexical Knowledge Bases (LKB). We show that our algorithm performs better than other graph-based methods when run on a graph built from WordNet and eXtended WordNet. Our algorithm and LKB combination compares favorably to other knowledge-based approaches in the literature that use similar knowledge on a variety of English data sets and a data set on Spanish. We include a detailed analysis of the factors that affect the algorithm. The algorithm and the LKBs used are publicly available, and the results easily reproducible. Eneko Agirre, Oier Lopez de Lacalle, Aitor Soroa |
Comput. Linguistics | 1 |
| 2014 | Evaluating hierarchical organisation structures for exploring digital libraries
Mark M. Hall, Samuel Fernando, Paul D. Clough, Aitor Soroa, Eneko Agirre, Mark Stevenson 0001 |
Inf. Retr. | 5 |
| 2014 | Improving search over Electronic Health Records using UMLS-based query expansion through random walks
David Martínez 0001, Arantxa Otegi, Aitor Soroa, Eneko Agirre |
J. Biomed. Informatics | 4 |
| 2013 | PATHSenrich: A Web Service Prototype for Automatic Cultural Heritage Item Enrichment
Eneko Agirre, Ander Barrena, Kike Fernández, Esther Miranda, Arantxa Otegi, Aitor Soroa |
TPDL | 1 |
| 2013 | Information seeking in digital cultural heritage with PATHSabstractCurrent Information Retrieval systems for digital cultural heritage support only the actual search aspect of the information seeking process. This demonstration presents the second PATHS system which provides the exploration, analysis, and sense-making features to support the full information seeking process. Mark M. Hall, Paul D. Clough, Samuel Fernando, Paula Goodale, Mark Stevenson 0001, Eneko Agirre, Arantxa Otegi, Aitor Soroa, Kate Fernie, Jillian Griffiths, Runar Bergheim |
SIGIR | 6 |
| 2013 | Selectional Preferences for Semantic Role ClassificationabstractThis paper focuses on a well-known open issue in Semantic Role Classification (SRC) research: the limited influence and sparseness of lexical features. We mitigate this problem using models that integrate automatically learned selectional preferences (SP). We explore a range of models based on WordNet and distributional-similarity SPs. Furthermore, we demonstrate that the SRC task is better modeled by SP models centered on both verbs and prepositions, rather than verbs alone. Our experiments with SP-based models in isolation indicate that they outperform a lexical baseline with 20 F1 points in domain and almost 40 F1 points out of domain. Furthermore, we show that a state-of-the-art SRC system extended with features based on selectional preferences performs significantly better, both in domain (17% error reduction) and out of domain (13% error reduction). Finally, we show that in an end-to-end semantic role labeling system we obtain small but statistically significant improvements, even though our modified SRC model affects only approximately 4% of the argument candidates. Our post hoc error analysis indicates that the SP-based features help mostly in situations where syntactic information is either incorrect or insufficient to disambiguate the correct role. Beñat Zapirain, Eneko Agirre, Lluís Màrquez, Mihai Surdeanu |
Comput. Linguistics | 2 |
| 2012 | Contribution of Complex Lexical Information to Solve Syntactic Ambiguity in Basque
Aitziber Atutxa, Eneko Agirre, Kepa Sarasola |
COLING | 2 |
| 2012 | Comparing Taxonomies for Organising Collections of Documents
Samuel Fernando, Mark M. Hall, Eneko Agirre, Aitor Soroa, Paul D. Clough, Mark Stevenson 0001 |
COLING | 3 |
| 2012 | PATHS - Exploring Digital Cultural Heritage Spaces
Mark M. Hall, Eneko Agirre, Nikolaos Aletras, Runar Bergheim, Konstantinos Chandrinos, Paul D. Clough, Samuel Fernando, Kate Fernie, Paula Goodale, Jillian Griffiths, Oier Lopez de Lacalle, Andrea de Polo, Aitor Soroa, Mark Stevenson 0001 |
TPDL | 2 |
| 2012 | Matching Cultural Heritage items to Wikipedia
Eneko Agirre, Ander Barrena, Oier Lopez de Lacalle, Aitor Soroa, Samuel Fernando, Mark Stevenson 0001 |
LREC | 1 |
| 2012 | Exploiting domain information for Word Sense Disambiguation of medical documentsabstractOBJECTIVE: Current techniques for knowledge-based Word Sense Disambiguation (WSD) of ambiguous biomedical terms rely on relations in the Unified Medical Language System Metathesaurus but do not take into account the domain of the target documents. The authors' goal is to improve these methods by using information about the topic of the document in which the ambiguous term appears. DESIGN: The authors proposed and implemented several methods to extract lists of key terms associated with Medical Subject Heading terms. These key terms are used to represent the document topic in a knowledge-based WSD system. They are applied both alone and in combination with local context. MEASUREMENTS: A standard measure of accuracy was calculated over the set of target words in the widely used National Library of Medicine WSD dataset. RESULTS AND DISCUSSION: The authors report a significant improvement when combining those key terms with local context, showing that domain information improves the results of a WSD system based on the Unified Medical Language System Metathesaurus alone. The best results were obtained using key terms obtained by relevance feedback and weighted by inverse document frequency. Mark Stevenson 0001, Eneko Agirre, Aitor Soroa |
J. Am. Medical Informatics Assoc. | 2 |
| 2011 | Two birds with one stone: learning semantic models for text categorization and word sense disambiguationabstractIn this paper we present a novel approach to learning semantic models for multiple domains, which we use to categorize Wikipedia pages and to perform domain Word Sense Disambiguation (WSD). In order to learn a semantic model for each domain we first extract relevant terms from the texts in the domain and then use these terms to initialize a random walk over the WordNet graph. Given an input text, we check the semantic models, choose the appropriate domain for that text and use the best-matching model to perform WSD. Our results show considerable improvements on text categorization and domain WSD tasks. Roberto Navigli, Stefano Faralli 0001, Aitor Soroa, Oier Lopez de Lacalle, Eneko Agirre |
CIKM | 5 |
| 2011 | Query Expansion for IR using Knowledge-Based Relatedness
Arantxa Otegi, Xabier Arregi, Eneko Agirre |
IJCNLP | 3 |
| 2010 | Plagiarism Detection across Distant Language Pairs
Alberto Barrón-Cedeño, Paolo Rosso, Eneko Agirre, Gorka Labaka |
COLING | 3 |
| 2010 | Exploring Knowledge Bases for Similarity
Eneko Agirre, Montse Cuadros, German Rigau, Aitor Soroa |
LREC | 1 |
| 2010 | Improving Semantic Role Classification with Selectional Preferences
Beñat Zapirain, Eneko Agirre, Lluís Màrquez, Mihai Surdeanu |
HLT-NAACL | 2 |
| 2010 | Graph-based Word Sense Disambiguation of biomedical documentsabstractMOTIVATION: Word Sense Disambiguation (WSD), automatically identifying the meaning of ambiguous words in context, is an important stage of text processing. This article presents a graph-based approach to WSD in the biomedical domain. The method is unsupervised and does not require any labeled training data. It makes use of knowledge from the Unified Medical Language System (UMLS) Metathesaurus which is represented as a graph. A state-of-the-art algorithm, Personalized PageRank, is used to perform WSD. RESULTS: When evaluated on the NLM-WSD dataset, the algorithm outperforms other methods that rely on the UMLS Metathesaurus alone. AVAILABILITY: The WSD system is open source licensed and available from http://ixa2.si.ehu.es/ukb/. The UMLS, MetaMap program and NLM-WSD corpus are available from the National Library of Medicine https://www.nlm.nih.gov/research/umls/, http://mmtx.nlm.nih.gov and http://wsd.nlm.nih.gov. Software to convert the NLM-WSD corpus into a format that can be used by our WSD system is available from http://www.dcs.shef.ac.uk/∼marks/biomedical_wsd under open source license. Eneko Agirre, Aitor Soroa, Mark Stevenson 0001 |
Bioinform. | 1 |
| 2009 | Supervised Domain Adaption for WSD
Eneko Agirre, Oier Lopez de Lacalle |
EACL | 1 |
| 2009 | Personalizing PageRank for Word Sense Disambiguation
Eneko Agirre, Aitor Soroa |
EACL | 1 |
| 2009 | Use of Rich Linguistic Information to Translate Prepositions and Grammar Cases to Basque
Eneko Agirre, Aitziber Atutxa, Gorka Labaka, Mikel Lersundi, Aingeru Mayor, Kepa Sarasola |
EAMT | 1 |
| 2009 | Knowledge-Based WSD and Specific Domains: Performing Better than Generic Supervised WSD
Eneko Agirre, Oier Lopez de Lacalle, Aitor Soroa |
IJCAI | 1 |
| 2009 | A Study on Similarity and Relatedness Using Distributional and WordNet-based Approaches
Eneko Agirre, Enrique Alfonseca, Keith B. Hall, Jana Kravalova, Marius Pasca, Aitor Soroa |
HLT-NAACL | 1 |
| 2008 | Improving Parsing and PP Attachment Performance with Sense Information
Eneko Agirre, Timothy Baldwin, David Martínez 0001 |
ACL | 1 |
| 2008 | Robustness and Generalization of Role Sets: PropBank vs. VerbNet
Beñat Zapirain, Eneko Agirre, Lluís Màrquez |
ACL | 2 |
| 2008 | A Preliminary Study on the Robustness and Generalization of Role Sets for Semantic Role Labeling
Beñat Zapirain, Eneko Agirre, Lluís Màrquez |
CICLing | 2 |
| 2008 | On Robustness and Domain Adaptation using SVD for Word Sense Disambiguation
Eneko Agirre, Oier Lopez de Lacalle |
COLING | 1 |
| 2008 | Using the Multilingual Central Repository for Graph-Based Word Sense Disambiguation
Eneko Agirre, Aitor Soroa |
LREC | 1 |
| 2008 | WNTERM: Enriching the MCR with a Terminological Dictionary
Eli Pociello, Antton Gurrutxaga, Eneko Agirre, Izaskun Aldezabal, German Rigau |
LREC | 3 |
| 2008 | KYOTO: a System for Mining, Structuring and Distributing Knowledge across Languages and Cultures
Piek Vossen, Eneko Agirre, Nicoletta Calzolari, Christiane Fellbaum, Shu-Kai Hsieh, Chu-Ren Huang, Hitoshi Isahara, Kyoko Kanzaki, Andrea Marchetti, Monica Monachini, Federico Neri, Remo Raffaelli, German Rigau, Maurizio Tesconi, Joop VanGent |
LREC | 2 |
| 2008 | On the Use of Automatically Acquired Examples for All-Nouns Word Sense Disambiguation abstractThis article focuses on Word Sense Disambiguation (WSD), which is a Natural Language Processing task that is thought to be important for many Language Technology applications, such as Information Retrieval, Information Extraction, or Machine Translation. One of the main issues preventing the deployment of WSD technology is the lack of training examples for Machine Learning systems, also known as the Knowledge Acquisition Bottleneck. A method which has been shown to work for small samples of words is the automatic acquisition of examples. We have previously shown that one of the most promising example acquisition methods scales up and produces a freely available database of 150 million examples from Web snippets for all polysemous nouns in WordNet. This paper focuses on the issues that arise when using those examples, all alone or in addition to manually tagged examples, to train a supervised WSD system for all nouns. The extensive evaluation on both lexical-sample and all-words Senseval benchmarks shows that we are able to improve over commonly used baselines and to achieve top-rank performance. The good use of the prior distributions from the senses proved to be a crucial factor. David Martínez 0001, Oier Lopez de Lacalle, Eneko Agirre |
J. Artif. Intell. Res. | 3 |
| 2006 | Two graph-based algorithms for state-of-the-art WSD
Eneko Agirre, David Martínez 0001, Oier Lopez de Lacalle, Aitor Soroa |
EMNLP | 1 |
| 2006 | A methodology for the joint development of the Basque WordNet and Semcor
Eneko Agirre, Izaskun Aldezabal, Jone Etxeberria, Eli Izagirre, Karmele Mendizabal, Eli Pociello, Mikel Quintian |
LREC | 1 |
| 2006 | A Preliminary Study for Building the Basque PropBank
Eneko Agirre, Izaskun Aldezabal, Jone Etxeberria, Eli Pociello |
LREC | 1 |
| 2004 | Unsupervised WSD based on Automatically Retrieved Examples: The Importance of Bias
Eneko Agirre, David Martínez 0001 |
EMNLP | 1 |
| 2004 | Exploring Portability of Syntactic Information from English to Basque
Eneko Agirre, Aitziber Atutxa, Koldo Gojenola, Kepa Sarasola |
LREC | 1 |
| 2004 | Publicly Available Topic Signatures for all WordNet Nominal Senses
Eneko Agirre, Oier Lopez de Lacalle |
LREC | 1 |
| 2004 | Cross-Language Acquisition of Semantic Models for Verbal Predicates
Jordi Atserias Batalla, Bernardo Magnini, Octavian Popescu, Eneko Agirre, Aitziber Atutxa, German Rigau, John Carroll 0001, Rob Koeling |
LREC | 4 |
| 2004 | The Effect of Bias on an Automatically-built Word Sense Corpus
David Martínez 0001, Eneko Agirre |
LREC | 2 |
| 2002 | Syntactic Features for High Precision Word Sense Disambiguation
David Martínez 0001, Eneko Agirre, Lluís Màrquez |
COLING | 2 |
| 2000 | A word-grammar based morphological analyzer for agglutinative languages
Itziar Aduriz, Eneko Agirre, Izaskun Aldezabal, Iñaki Alegria, Xabier Arregi, Jose Maria Arriola, Xabier Artola, Koldo Gojenola, Montse Maritxalar, Kepa Sarasola, Miriam Urkia |
COLING | 2 |
| 2000 | One Sense per Collocation and Genre/Topic VariationsabstractThis paper revisits the one sense per collocation hypothesis using fine-grained sense distinctions and two different corpora. We show that the hypothesis is weaker for fine-grained sense distinctions (70% vs. 99% reported earlier on 2-way ambiguities). We also show that one sense per collocation does hold across corpora, but that collocations vary from one corpus to the other, following genre and topic variations. This explains the low results when performing word sense disambiguation across corpora. In fact, we demonstrate that when two independent corpora share a related genre/topic, the word sense disambiguation results would be better. Future work on word sense disambiguation will have to take into account genre and topic as important parameters on their models. David Martínez 0001, Eneko Agirre |
EMNLP | 2 |
| 2000 | A Word-level Morphosyntactic Analyzer for Basque
Itziar Aduriz, Eneko Agirre, Izaskun Aldezabal, Xabier Arregi, Jose Maria Arriola, Xabier Artola, Koldo Gojenola, A. Maritxalar, Kepa Sarasola, Miriam Urkia |
LREC | 2 |
| 2000 | A Methodology for Building Translator-oriented Dictionary Systems
Eneko Agirre, Xabier Arregi, Xabier Artola, Arantza Díaz de Ilarraza, Kepa Sarasola, Aitor Soroa |
Mach. Transl. | 1 |
| 1999 | MLDS: A translator-oriented MultiLingual dictionary systemabstractThis paper focuses on the design methodology of the MultiLingual Dictionary-System (MLDS), which is a human-oriented tool for assisting in the task of translating lexical units, oriented to translators and conceived from studies carried out with translators. We describe the model adopted for the representation of multilingual dictionary-knowledge. Such a model allows an enriched exploitation of the lexical-semantic relations extracted from dictionaries. In addition, MLDS is supplied with knowledge about the use of the dictionaries in the process of lexical translation, which was elicitated by means of empirical methods and specified in a formal language. The dictionary-knowledge along with the task-oriented knowledge are used to offer the translator active, anticipative and intelligent assistance. Eneko Agirre, Xabier Arregi, Xabier Artola, Arantza Díaz de Ilarraza, Kepa Sarasola, Aitor Soroa |
Nat. Lang. Eng. | 1 |
| 1997 | Combining Unsupervised Lexical Knowledge Methods for Word Sense DisambiguationabstractThis paper presents a method to combine a set of unsupervised algorithms that can accurately disambiguate word senses in a large, completely untagged corpus. Although most of the techniques for word sense resolution have been presented as stand-alone, it is our belief that full-fledged lexical ambiguity resolution should combine several information sources and techniques. The set of techniques have been applied in a combined way to disambiguate the genus terms of two machine-readable dictionaries (MRD), enabling us to construct complete taxonomies for Spanish and French. Texted accuracy is above 80% overall and 95% for two-way ambiguous genus terms, showing that texonomy building is not limited to structured dictionaries such as LDOCE. German Rigau, Jordi Atserias Batalla, Eneko Agirre |
ACL | 3 |
| 1996 | Word Sense Disambiguation using Conceptual Density
Eneko Agirre, German Rigau |
COLING | 1 |
| 1996 | Constructing an intelligent dictionary help systemabstractThis paper discusses different issues in the construction and knowledge representation of an intelligent dictionary help system. The Intelligent Dictionary Help System (IDHS) is conceived as a monolingual (explanatory) dictionary system for human use (Artola and Evrard, 1992). The fact that it is intended for people instead of automatic processing distinguishes it from other systems dealing with the acquisition of semantic knowledge from conventional dictionaries. The system provides various access possibilities to the data, allowing to deduce implicit knowledge from the explicit dictionary information. IDHS deals with reasoning mechanisms analogous to those used by humans when they consult a dictionary. User level functionality of the system has been specified and a prototype has been implemented (Agirre et al., 1994a). A methodology for the extraction of semantic knowledge from a conventional dictionary is described. The method followed in the construction of the phrasal pattern hierarchies required by the parser (Alshawi, 1989) is based on an empirical study carried out on the structure of definition sentences. The results of its application to a real dictionary has shown that the parsing method is particularly suited to the analysis of short definition sentences, as it was the case of the source dictionary. As a result of this process, the characterization of the different lexical-semantic relations between senses is established by means of semantic rules (attached to the patterns); these rules are used for the initial construction of the Dictionary Knowledge Base (DKB). The representation schema proposed for the DKB (Agirre et al., 1994b) is basically a semantic network of frames representing word senses. After construction of the initial DKB, several enrichment processes are performed on the DKB to add new facts to it; these processes are based on the exploitation of the properties of lexical-semantic relations, and also on specially conceived deduction mechanisms. The result of the enrichment processes show the suitability of the representation schema chosen to deduce implicit knowledge. Erroneous deductions are mainly due to incorrect word sense disambiguation. Eneko Agirre, Xabier Arregi, Xabier Artola, Arantza Díaz de Ilarraza, Kepa Sarasola, Aitor Soroa |
Nat. Lang. Eng. | 1 |
| 1994 | Lexical, Knowledge Representation In An Intelligent Dictionary Help System
Eneko Agirre, Xabier Arregi, Xabier Artola, Arantza Díaz de Ilarraza, Kepa Sarasola |
COLING | 1 |
| 1993 | A Morphological Analysis Based Method for Spelling Correction
Itziar Aduriz, Eneko Agirre, Iñaki Alegria, Xabier Arregi, Jose Maria Arriola, Xabier Artola, Arantza Díaz de Ilarraza, Nerea Ezeiza, Montse Maritxalar, Kepa Sarasola, Miriam Urkia |
EACL | 2 |