Aitor Soroa

dblp:03/6734 · DBLP profile ↗
← Back
64ranked-venue papers
0as first author
18since 2021 · last 2026
0000-0001-8573-2654ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 55 · 18 since 2021Databases, data management, data science and information retrieval · 7Graphics, computer vision, multimedia, augmented reality and games · 3Applied, interdisciplinary, general and emerging computing · 3
YearPublicationVenuePosition
2026 CRITICS: Critical Science Without Borders by Translation of Scientific Knowledge
abstract
The CRITICS project addresses science accessibility and literacy through the convergence of advanced Machine Translation (MT) based on Large Language Models (LLMs) and educational technology. By leveraging MT systems specifically optimized for scientific content, educational institutions can provide accurate, culturally relevant translations of scientific materials in higher-education students’ native languages, ensuring that complex scientific concepts are comprehensible while maintaining technical accuracy. Novel research on MT specifically tailored for scientific documents aims to break down language barriers in accessing cutting-edge research and educational materials currently only available in high-resourced languages such as English, thereby facilitating the democratization of scientific knowledge across linguistic boundaries.
Rodrigo Agerri, Itziar Aldabe, Elena Cabrio, Mark Cieliebak, Jan Deriu, Mariana Flores, Jurgita Kapociute-Dzikiene, Dovile Kuiziniene, Arantza Rico, Aritz Ruiz-González, Aitor Soroa, Mantas Vaskevicius, Serena Villata
EAMT (2)11
2026 Conditioning LLMs to Generate Code-Switched Text
Maite Heredia, Gorka Labaka, Jeremy Barnes 0001, Aitor Soroa
LREC4
2026 Assessing Logical Coherence of LLMs via Fine-Grained NLI
Jon F. Apaolaza, Begoña Altuna, Aitor Soroa, Iñigo Lopez-Gazpio
LREC3
2026 Benchmarking Mathematical Reasoning in a Low-Resource Language: Structured Prompting and Evaluation in Basque
Inigo Martinez-Criado, Aitor Soroa, Jeremy Barnes 0001
LREC2
2026 Judging Instruction Responses in a Low-Resource Language: A Case Study on Basque
David Ponce, Harritxu Gete, Thierry Etchegoyhen, Irune Zubiaga, Aitor Soroa
LREC5
2026 MeteoGalEus: An Iberian Multilingual Weather Dataset in Galician, Euskera, and Spanish
Ainhoa Vivel-Couso, Nella Zabrina Pramata, David Robredo, Aitor Soroa, Jose Maria Alonso-Moral
LREC4
2025 Instructing Large Language Models for Low-Resource Languages: A Systematic Study for Basque
abstract
Oscar Sainz, Naiara Perez, Julen Etxaniz, Joseba Fernandez de Landa, Itziar Aldabe, Iker García-Ferrero, Aimar Zabala, Ekhi Azurmendi, German Rigau, Eneko Agirre, Mikel Artetxe, Aitor Soroa. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Oscar Sainz, Naiara Pérez, Julen Etxaniz, Joseba Fernandez de Landa, Itziar Aldabe, Iker García-Ferrero, Aimar Zabala, Ekhi Azurmendi, German Rigau, Eneko Agirre, Mikel Artetxe, Aitor Soroa
EMNLP12
2024 Latxa: An Open Language Model and Evaluation Suite for Basque
abstract
Julen Etxaniz, Oscar Sainz, Naiara Perez, Itziar Aldabe, German Rigau, Eneko Agirre, Aitor Ormazabal, Mikel Artetxe, Aitor Soroa. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Julen Etxaniz, Oscar Sainz, Naiara Miguel, Itziar Aldabe, German Rigau, Eneko Agirre, Aitor Ormazabal, Mikel Artetxe, Aitor Soroa
ACL (1)9
2024 XNLIeu: a dataset for cross-lingual NLI in Basque
abstract
Maite Heredia, Julen Etxaniz, Muitze Zulaika, Xabier Saralegi, Jeremy Barnes, Aitor Soroa. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Maite Heredia, Julen Etxaniz, Muitze Zulaika, Xabier Saralegi, Jeremy Barnes 0001, Aitor Soroa
NAACL-HLT6
2024 BertaQA: How Much Do Language Models Know About Local Culture?
abstract
Large Language Models (LLMs) exhibit extensive knowledge about the world, but most evaluations have been limited to global or anglocentric subjects. This raises the question of how well these models perform on topics relevant to other cultures, whose presence on the web is not that prominent. To address this gap, we introduce BertaQA, a multiple-choice trivia dataset that is parallel in English and Basque. The dataset consists of a local subset with questions pertinent to the Basque culture, and a global subset with questions of broader interest. We find that state-of-the-art LLMs struggle with local cultural knowledge, even as they excel on global topics. However, we show that continued pre-training in Basque significantly improves the models' performance on Basque culture, even when queried in English. To our knowledge, this is the first solid evidence of knowledge transfer from a low-resource to a high-resource language. Our analysis sheds light on the complex interplay between language and knowledge, and reveals that some prior findings do not fully hold when reassessed on local topics. Our dataset and evaluation code are available under open licenses at https://github.com/juletx/BertaQA.
Julen Etxaniz, Gorka Azkune, Aitor Soroa, Oier Lopez de Lacalle, Mikel Artetxe
NeurIPS3
2023 Image captioning for effective use of language models in knowledge-based visual question answering
abstract
Integrating outside knowledge for reasoning in visio-linguistic tasks such as visual question answering (VQA) is an open problem. Given that pretrained language models have been shown to include world knowledge, we propose to use a unimodal (text-only) train and inference procedure based on automatic off-the-shelf captioning of images and pretrained language models. More specifically, we verbalize the image contents and allow language models to better leverage their implicit knowledge to solve knowledge-intensive tasks. Focusing on a visual question answering task which requires external knowledge (OK-VQA), our contributions are: (i) a text-only model that outperforms pretrained multimodal (image-text) models of comparable number of parameters; (ii) confirmation that our text-only method is specially effective for tasks requiring external knowledge, as it is less effective in standard a VQA task (VQA 2.0); and (iii) our method attains results in the state-of-the-art when increasing the size of the language model. We also significantly outperform current multimodal systems, even though augmented with external knowledge. Our qualitative analysis on OK-VQA reveals that automatic captions often fail to capture relevant information in the images, which seems to be balanced by the better inference ability of the text-only language models. Our work opens up possibilities to further improve inference in visio-linguistic tasks.
Ander Salaberria, Gorka Azkune, Oier Lopez de Lacalle, Aitor Soroa, Eneko Agirre
Expert Syst. Appl.4
2022 Principled Paraphrase Generation with Parallel Corpora
abstract
Round-trip Machine Translation (MT) is a popular choice for paraphrase generation, which leverages readily available parallel corpora for supervision.In this paper, we formalize the implicit similarity function induced by this approach, and show that it is susceptible to nonparaphrase pairs sharing a single ambiguous translation.Based on these insights, we design an alternative similarity metric that mitigates this issue by requiring the entire translation distribution to match, and implement a relaxation of it through the Information Bottleneck method.Our approach incorporates an adversarial term into MT training in order to learn representations that encode as much information about the reference translation as possible, while keeping as little information about the input as possible.Paraphrases can be generated by decoding back to the source from this representation, without having to generate pivot translations.In addition to being more principled and efficient than round-trip MT, our approach offers an adjustable parameter to control the fidelity-diversity trade-off, and obtains better results in our experiments.
Aitor Ormazabal, Mikel Artetxe, Aitor Soroa, Gorka Labaka, Eneko Agirre
ACL (1)3
2022 Does Corpus Quality Really Matter for Low-Resource Languages?
abstract
The vast majority of non-English corpora are derived from automatically filtered versions of CommonCrawl.While prior work has identified major issues on the quality of these datasets (Kreutzer et al., 2021), it is not clear how this impacts downstream performance.Taking representation learning in Basque as a case study, we explore tailored crawling-manually identifying and scraping websites with high-quality content-as an alternative to filtering Common-Crawl.Our new corpus, called EusCrawl, is similar in size to the Basque portion of popular multilingual corpora like CC100 and mC4, yet it has a much higher quality according to native annotators.For instance, 66% of documents are rated as high-quality for EusCrawl, in contrast with < 33% for both mC4 and CC100.Nevertheless, we obtain similar results on downstream NLU tasks regardless of the corpus used for pre-training.Our work suggests that NLU performance in low-resource languages is not primarily constrained by the quality of the data, and other factors like corpus size and domain coverage can play a more important role.
Mikel Artetxe, Itziar Aldabe, Rodrigo Agerri, Olatz Perez-de-Viñaspre, Aitor Soroa
EMNLP5
2022 BasqueGLUE: A Natural Language Understanding Benchmark for Basque
abstract
Natural Language Understanding (NLU) technology has improved significantly over the last few years and multitask benchmarks such as GLUE are key to evaluate this improvement in a robust and general way. These benchmarks take into account a wide and diverse set of NLU tasks that require some form of language understanding, beyond the detection of superficial, textual clues. However, they are costly to develop and language-dependent, and therefore they are only available for a small number of languages. In this paper, we present BasqueGLUE, the first NLU benchmark for Basque, a less-resourced language, which has been elaborated from previously existing datasets and following similar criteria to those used for the construction of GLUE and SuperGLUE. We also report the evaluation of two state-of-the-art language models for Basque on BasqueGLUE, thus providing a strong baseline to compare upon. BasqueGLUE is freely available under an open license.
Gorka Urbizu, Iñaki San Vicente, Xabier Saralegi, Rodrigo Agerri, Aitor Soroa
LREC5
2022 The BigScience ROOTS Corpus: A 1.6TB Composite Multilingual Dataset
abstract
As language models grow ever larger, the need for large-scale high-quality text datasets has never been more pressing, especially in multilingual settings. The BigScience workshop, a 1-year international and multidisciplinary initiative, was formed with the goal of researching and training large language models as a values-driven undertaking, putting issues of ethics, harm, and governance in the foreground. This paper documents the data creation and curation efforts undertaken by BigScience to assemble the Responsible Open-science Open-collaboration Text Sources (ROOTS) corpus, a 1.6TB dataset spanning 59 languages that was used to train the 176-billion-parameter BigScience Large Open-science Open-access Multilingual (BLOOM) language model. We further release a large initial subset of the corpus and analyses thereof, and hope to empower large-scale monolingual and multilingual modeling projects with both the data and the processing tools, as well as stimulate research around this large multilingual corpus.
Hugo Laurençon, Lucile Saulnier, Thomas Wang, Christopher Akiki, Albert Villanova del Moral, Teven Le Scao, Leandro von Werra, Chenghao Mou, Eduardo G. Ponferrada, Huu Nguyen, Jörg Frohberg, Mario Sasko, Quentin Lhoest, Angelina McMillan-Major, Gérard Dupont, Stella Biderman, Anna Rogers, Loubna Ben Allal, Francesco De Toni, Giada Pistilli, Olivier Nguyen, Somaieh Nikpoor, Maraim Masoud, Pierre Colombo, Javier de la Rosa 0001, Paulo Villegas, Tristan Thrush, Shayne Longpre, Sebastian Nagel 0005, Leon Weber-Genzel, Manuel Muñoz, Daniel van Strien, Zaid Alyafeai, Khalid Almubarak, Minh Chien Vu, Itziar Gonzalez-Dios, Aitor Soroa, Kyle Lo, Manan Dey, Pedro Ortiz Suarez, Aaron Gokaslan, Shamik Bose, David Ifeoluwa Adelani, Long Phan, Hieu Tran, Ian Yu, Suhas Pai, Jenny Chim, Violette Lepercq, Suzana Ilic, Margaret Mitchell, Sasha Luccioni, Yacine Jernite
NeurIPS38
2021 Beyond Offline Mapping: Learning Cross-lingual Word Embeddings through Context Anchoring
abstract
Recent research on cross-lingual word embeddings has been dominated by unsupervised mapping approaches that align monolingual embeddings. Such methods critically rely on those embeddings having a similar structure, but it was recently shown that the separate training in different languages causes departures from this assumption. In this paper, we propose an alternative approach that does not have this limitation, while requiring a weak seed dictionary (e.g., a list of identical words) as the only form of supervision. Rather than aligning two fixed embedding spaces, our method works by fixing the target language embeddings, and learning a new set of embeddings for the source language that are aligned with them. To that end, we use an extension of skip-gram that leverages translated context words as anchor points, and incorporates self-learning and iterative restarts to reduce the dependency on the initial dictionary. Our approach outperforms conventional mapping methods on bilingual lexicon induction, and obtains competitive results in the downstream XNLI task.
Aitor Ormazabal, Mikel Artetxe, Aitor Soroa, Gorka Labaka, Eneko Agirre
ACL/IJCNLP (1)3
2021 Towards zero-shot cross-lingual named entity disambiguation
abstract
In cross-Lingual Named Entity Disambiguation (XNED) the task is to link Named Entity mentions in text in some native language to English entities in a knowledge graph. XNED systems usually require training data for each native language, limiting their application for low resource languages with small amounts of training data. Prior work have proposed so-called zero-shot transfer systems which are only trained in English training data, but required native prior probabilities of entities with respect to mentions, which had to be estimated from native training examples, limiting their practical interest. In this work we present a zero-shot XNED architecture where, instead of a single disambiguation model, we have a model for each possible mention string, thus eliminating the need for native prior probabilities. Our system improves over prior work in XNED datasets in Spanish and Chinese by 32 and 27 points, and matches the systems which do require native prior information. We experiment with different multilingual transfer strategies, showing that better results are obtained with a purpose-built multilingual pre-training method compared to state-of-the-art generic multilingual models such as XLM-R. We also discovered, surprisingly, that English is not necessarily the most effective zero-shot training language for XNED into English. For instance, Spanish is more effective when training a zero-shot XNED system that disambiguates Basque mentions with respect to an English knowledge graph.
Ander Barrena, Aitor Soroa, Eneko Agirre
Expert Syst. Appl.2
2021 Inferring spatial relations from textual descriptions of images
Aitzol Elu, Gorka Azkune, Oier Lopez de Lacalle, Ignacio Arganda-Carreras, Aitor Soroa, Eneko Agirre
Pattern Recognit.5
2020 DoQA - Accessing Domain-Specific FAQs via Conversational QA
abstract
The goal of this work is to build conversational Question Answering (QA) interfaces for the large body of domain-specific information available in FAQ sites.We present DoQA, a dataset with 2,437 dialogues and 10,917 QA pairs.The dialogues are collected from three Stack Exchange sites using the Wizard of Oz method with crowdsourcing.Compared to previous work, DoQA comprises well-defined information needs, leading to more coherent and natural conversations with less factoid questions and is multi-domain.In addition, we introduce a more realistic information retrieval (IR) scenario where the system needs to find the answer in any of the FAQ documents.The results of an existing, strong, system show that, thanks to transfer learning from a Wikipedia QA dataset and fine tuning on a single FAQ domain, it is possible to build high quality conversational QA systems for FAQs without indomain training data.The good results carry over into the more challenging IR scenario.In both cases, there is still ample room for improvement, as indicated by the higher human upperbound.
Jon Ander Campos, Arantxa Otegi, Aitor Soroa, Jan Deriu, Mark Cieliebak, Eneko Agirre
ACL3
2020 Improving Conversational Question Answering Systems after Deployment using Feedback-Weighted Learning
abstract
The interaction of conversational systems with users poses an exciting opportunity for improving them after deployment, but little evidence has been provided of its feasibility.In most applications, users are not able to provide the correct answer to the system, but they are able to provide binary (correct, incorrect) feedback.In this paper we propose feedback-weighted learning based on importance sampling to improve upon an initial supervised system using binary user feedback.We perform simulated experiments on document classification (for development) and Conversational Question Answering datasets like QuAC and DoQA, where binary user feedback is derived from gold annotations.The results show that our method is able to improve over the initial supervised system, getting close to a fully-supervised system that has access to the same labeled examples in in-domain experiments (QuAC), and even matching in out-of-domain experiments (DoQA).Our work opens the prospect to exploit interactions with real users and improve conversational systems after deployment.
Jon Ander Campos, Kyunghyun Cho, Arantxa Otegi, Aitor Soroa, Eneko Agirre, Gorka Azkune
COLING4
2020 Evaluating Multimodal Representations on Visual Semantic Textual Similarity
abstract
The combination of visual and textual representations has produced excellent results in tasks such as image captioning and visual question answering, but the inference capabilities of multimodal representations are largely untested. In the case of textual representations, inference tasks such as Textual Entailment and Semantic Textual Similarity have been often used to benchmark the quality of textual representations. The long term goal of our research is to devise multimodal representation techniques that improve current inference capabilities. We thus present a novel task, Visual Semantic Textual Similarity (vSTS), where such inference ability can be tested directly. Given two items comprised each by an image and its accompanying caption, vSTS systems need to assess the degree to which the captions in context are semantically equivalent to each other. Our experiments using simple multimodal representations show that the addition of image representations produces better inference, compared to text-only representations. The improvement is observed both when directly computing the similarity between the representations of the two items, and when learning a siamese network based on vSTS training data. Our work shows, for the first time, the successful contribution of visual information to textual inference, with ample room for benchmarking more complex multimodal representation options.
Oier Lopez de Lacalle, Ander Salaberria, Aitor Soroa, Gorka Azkune, Eneko Agirre
ECAI3
2020 Spot The Bot: A Robust and Efficient Framework for the Evaluation of Conversational Dialogue Systems
abstract
Jan Deriu, Don Tuggener, Pius von Däniken, Jon Ander Campos, Alvaro Rodrigo, Thiziri Belkacem, Aitor Soroa, Eneko Agirre, Mark Cieliebak. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020.
Jan Deriu, Don Tuggener, Pius von Däniken, Jon Ander Campos, Álvaro Rodrigo, Thiziri Belkacem, Aitor Soroa, Eneko Agirre, Mark Cieliebak
EMNLP (1)7
2020 Give your Text Representation Models some Love: the Case for Basque
abstract
Word embeddings and pre-trained language models allow to build rich representations of text and have enabled improvements across most NLP tasks. Unfortunately they are very expensive to train, and many small companies and research groups tend to use models that have been pre-trained and made available by third parties, rather than building their own. This is suboptimal as, for many languages, the models have been trained on smaller (or lower quality) corpora. In addition, monolingual pre-trained models for non-English languages are not always available. At best, models for those languages are included in multilingual versions, where each language shares the quota of substrings and parameters with the rest of the languages. This is particularly true for smaller languages such as Basque. In this paper we show that a number of monolingual models (FastText word embeddings, FLAIR and BERT language models) trained with larger Basque corpora produce much better results than publicly available versions in downstream NLP tasks, including topic classification, sentiment classification, PoS tagging and NER. This work sets a new state-of-the-art in those tasks for Basque. All benchmarks and models used in this work are publicly available.
Rodrigo Agerri, Iñaki San Vicente, Jon Ander Campos, Ander Barrena, Xabier Saralegi, Aitor Soroa, Eneko Agirre
LREC6
2020 Conversational Question Answering in Low Resource Scenarios: A Dataset and Case Study for Basque
abstract
Conversational Question Answering (CQA) systems meet user information needs by having conversations with them, where answers to the questions are retrieved from text. There exist a variety of datasets for English, with tens of thousands of training examples, and pre-trained language models have allowed to obtain impressive results. The goal of our research is to test the performance of CQA systems under low-resource conditions which are common for most non-English languages: small amounts of native annotations and other limitations linked to low resource languages, like lack of crowdworkers or smaller wikipedias. We focus on the Basque language, and present the first non-English CQA dataset and results. Our experiments show that it is possible to obtain good results with low amounts of native data thanks to cross-lingual transfer, with quality comparable to those obtained for English. We also discovered that dialogue history models are not directly transferable to another language, calling for further research. The dataset is publicly available.
Arantxa Otegi, Aitor Gonzalez-Agirre, Jon Ander Campos, Aitor Soroa, Eneko Agirre
LREC4
2019 Analyzing the Limitations of Cross-lingual Word Embedding Mappings
abstract
Recent research in cross-lingual word embeddings has almost exclusively focused on offline methods, which independently train word embeddings in different languages and map them to a shared space through linear transformations.While several authors have questioned the underlying isomorphism assumption, which states that word embeddings in different languages have approximately the same structure, it is not clear whether this is an inherent limitation of mapping approaches or a more general issue when learning crosslingual embeddings.So as to answer this question, we experiment with parallel corpora, which allows us to compare offline mapping to an extension of skip-gram that jointly learns both embedding spaces.We observe that, under these ideal conditions, joint learning yields to more isomorphic embeddings, is less sensitive to hubness, and obtains stronger results in bilingual lexicon induction.We thus conclude that current mapping methods do have strong limitations, calling for further research to jointly learn cross-lingual embeddings with a weaker cross-lingual signal.
Aitor Ormazabal, Mikel Artetxe, Gorka Labaka, Aitor Soroa, Eneko Agirre
ACL (1)4
2018 Learning Text Representations for 500K Classification Tasks on Named Entity Disambiguation
abstract
Named Entity Disambiguation algorithms typically learn a single model for all target entities.In this paper we present a word expert model and train separate deep learning models for each target entity string, yielding 500K classification tasks.This gives us the opportunity to benchmark popular text representation alternatives on this massive dataset.In order to face scarce training data we propose a simple data-augmentation technique and transfer-learning.We show that bagof-word-embeddings are better than LSTMs for tasks with scarce training data, while the situation is reversed when having larger amounts.Transferring an LSTM which is learned on all datasets is the most effective context representation option for the word experts in all frequency bands.The experiments show that our system trained on out-ofdomain Wikipedia data surpasses comparable NED systems which have been trained on indomain training data.
Ander Barrena, Aitor Soroa, Eneko Agirre
CoNLL2
2018 Bilingual embeddings with random walks over multilingual wordnets
Josu Goikoetxea, Aitor Soroa, Eneko Agirre
Knowl. Based Syst.2
2017 A scalable architecture for data-intensive natural language processing
abstract
Abstract Computational power needs have greatly increased during the last years, and this is also the case in the Natural Language Processing (NLP) area, where thousands of documents must be processed, i.e., linguistically analyzed, in a reasonable time frame. These computing needs have implied a radical change in the computing architectures and big-scale text processing techniques used in NLP. In this paper, we present a scalable architecture for distributed language processing. The architecture uses Storm to combine diverse NLP modules into a processing chain, which carries out the linguistic analysis of documents. Scalability requires designing solutions that are able to run distributed programs in parallel and across large machine clusters. Using the architecture presented here, it is possible to integrate a set of third-party NLP modules into a unique processing chain which can be deployed onto a distributed environment, i.e., a cluster of machines, so allowing the language-processing modules run in parallel. No restrictions are placed a priori on the NLP modules apart of being able to consume and produce linguistic annotations following a given format. We show the feasibility of our approach by integrating two linguistic processing chains for English and Spanish. Moreover, we provide several scripts that allow building from scratch a whole distributed architecture that can be then easily installed and deployed onto a cluster of machines. The scripts and the NLP modules used in the paper are publicly available and distributed under free licenses. In the paper, we also describe a series of experiments carried out in the context of the NewsReader project with the goal of testing how the system behaves in different scenarios.
Zuhaitz Beloki, Xabier Artola, Aitor Soroa
Nat. Lang. Eng.3
2016 Single or Multiple? Combining Word Representations Independently Learned from Text and WordNet
abstract
Text and Knowledge Bases are complementary sources of information. Given the success of distributed word representations learned from text, several techniques to infuse additional information from sources like WordNet into word representations have been proposed. In this paper, we follow an alternative route. We learn word representations from text and WordNet independently, and then explore simple and sophisticated methods to combine them. The combined representations are applied to an extensive set of datasets on word similarity and relatedness. Simple combination methods happen to perform better that more complex methods like CCA or retrofitting, showing that, in the case of WordNet, learning word representations separately is preferable to learning one single representation space or adding WordNet information directly. A key factor, which we illustrate with examples, is that the WordNet-based representations captures similarity relations encoded in WordNet better than retrofitting. In addition, we show that the average of the similarities from six word representations yields results beyond the state-of-the-art in several datasets, reinforcing the opportunities to explore further combination techniques.
Josu Goikoetxea, Eneko Agirre, Aitor Soroa
AAAI3
2016 Alleviating Poor Context with Background Knowledge for Named Entity Disambiguation
abstract
Named Entity Disambiguation (NED) algorithms disambiguate mentions of named entities with respect to a knowledge-base, but sometimes the context might be poor or misleading.In this paper we introduce the acquisition of two kinds of background information to alleviate that problem: entity similarity and selectional preferences for syntactic positions.We show, using a generative Näive Bayes model for NED, that the additional sources of context are complementary, and improve results in the CoNLL 2003 and TAC KBP DEL 2014 datasets, yielding the third best and the best results, respectively.We provide examples and analysis which show the value of the acquired background information.
Ander Barrena, Aitor Soroa, Eneko Agirre
ACL (1)2
2016 Exploring Comparative Evaluation of Semantic Enrichment Tools for Cultural Heritage Metadata
Hugo Manguinhas, Nuno Freire 0001, Antoine Isaac, Juliane Stiller, Valentine Charles, Aitor Soroa, Rainer Simon, Vladimir Alexiev
TPDL6
2016 Interoperability of Annotation Schemes: Using the Pepper Framework to Display AWA Documents in the ANNIS Interface
Talvany Carlotto, Zuhaitz Beloki, Xabier Artola, Aitor Soroa
LREC4
2016 Two Architectures for Parallel Processing of Huge Amounts of Text
Mathijs Kattenberg, Zuhaitz Beloki, Aitor Soroa, Xabier Artola, Antske Fokkens, Paul Huygen, Kees Verstoep
LREC3
2016 Building event-centric knowledge graphs from news
Marco Rospocher, Marieke van Erp, Piek Vossen, Antske Fokkens, Itziar Aldabe, German Rigau, Aitor Soroa, Thomas Ploeger, Tessel Bogaard
J. Web Semant.7
2015 Random Walks and Neural Network Language Models on Knowledge Bases
abstract
Random walks over large knowledge bases like WordNet have been successfully used in word similarity, relatedness and disambiguation tasks. Unfortunately, those algorithms are relatively slow for large repositories, with significant memory footprints. In this paper we present a novel algorithm which encodes the structure of a knowledge base in a continuous vector space, combining random walks and neural net language models in order to produce novel word representations. Evaluation in word relatedness and similar- ity datasets yields equal or better results than those of a random walk algorithm, using a dense representation (300 dimensions instead of 117K). Furthermore, the word representations are complementary to those of the random walk algorithm and to corpus-based continuous representations, improving the state- of-the-art in the similarity dataset. Our technique opens up exciting opportunities to combine distributional and knowledge-based word representations.
Josu Goikoetxea, Aitor Soroa, Eneko Agirre
HLT-NAACL2
2015 Big data for Natural Language Processing: A streaming approach
Rodrigo Agerri, Xabier Artola, Zuhaitz Beloki, German Rigau, Aitor Soroa
Knowl. Based Syst.5
2014 "One Entity per Discourse" and "One Entity per Collocation" Improve Named-Entity Disambiguation
Ander Barrena, Eneko Agirre, Bernardo Cabaleiro, Anselmo Peñas, Aitor Soroa
COLING5
2014 A stream computing approach towards scalable NLP
Xabier Artola, Zuhaitz Beloki, Aitor Soroa
LREC3
2014 Random Walks for Knowledge-Based Word Sense Disambiguation
abstract
Word Sense Disambiguation (WSD) systems automatically choose the intended meaning of a word in context. In this article we present a WSD algorithm based on random walks over large Lexical Knowledge Bases (LKB). We show that our algorithm performs better than other graph-based methods when run on a graph built from WordNet and eXtended WordNet. Our algorithm and LKB combination compares favorably to other knowledge-based approaches in the literature that use similar knowledge on a variety of English data sets and a data set on Spanish. We include a detailed analysis of the factors that affect the algorithm. The algorithm and the LKBs used are publicly available, and the results easily reproducible.
Eneko Agirre, Oier Lopez de Lacalle, Aitor Soroa
Comput. Linguistics3
2014 Evaluating hierarchical organisation structures for exploring digital libraries
Mark M. Hall, Samuel Fernando, Paul D. Clough, Aitor Soroa, Eneko Agirre, Mark Stevenson 0001
Inf. Retr.4
2014 Improving search over Electronic Health Records using UMLS-based query expansion through random walks
David Martínez 0001, Arantxa Otegi, Aitor Soroa, Eneko Agirre
J. Biomed. Informatics3
2013 PATHSenrich: A Web Service Prototype for Automatic Cultural Heritage Item Enrichment
Eneko Agirre, Ander Barrena, Kike Fernández, Esther Miranda, Arantxa Otegi, Aitor Soroa
TPDL6
2013 Information seeking in digital cultural heritage with PATHS
abstract
Current Information Retrieval systems for digital cultural heritage support only the actual search aspect of the information seeking process. This demonstration presents the second PATHS system which provides the exploration, analysis, and sense-making features to support the full information seeking process.
Mark M. Hall, Paul D. Clough, Samuel Fernando, Paula Goodale, Mark Stevenson 0001, Eneko Agirre, Arantxa Otegi, Aitor Soroa, Kate Fernie, Jillian Griffiths, Runar Bergheim
SIGIR8
2012 Comparing Taxonomies for Organising Collections of Documents
Samuel Fernando, Mark M. Hall, Eneko Agirre, Aitor Soroa, Paul D. Clough, Mark Stevenson 0001
COLING4
2012 PATHS - Exploring Digital Cultural Heritage Spaces
Mark M. Hall, Eneko Agirre, Nikolaos Aletras, Runar Bergheim, Konstantinos Chandrinos, Paul D. Clough, Samuel Fernando, Kate Fernie, Paula Goodale, Jillian Griffiths, Oier Lopez de Lacalle, Andrea de Polo, Aitor Soroa, Mark Stevenson 0001
TPDL13
2012 Matching Cultural Heritage items to Wikipedia
Eneko Agirre, Ander Barrena, Oier Lopez de Lacalle, Aitor Soroa, Samuel Fernando, Mark Stevenson 0001
LREC4
2012 Exploiting domain information for Word Sense Disambiguation of medical documents
abstract
OBJECTIVE: Current techniques for knowledge-based Word Sense Disambiguation (WSD) of ambiguous biomedical terms rely on relations in the Unified Medical Language System Metathesaurus but do not take into account the domain of the target documents. The authors' goal is to improve these methods by using information about the topic of the document in which the ambiguous term appears. DESIGN: The authors proposed and implemented several methods to extract lists of key terms associated with Medical Subject Heading terms. These key terms are used to represent the document topic in a knowledge-based WSD system. They are applied both alone and in combination with local context. MEASUREMENTS: A standard measure of accuracy was calculated over the set of target words in the widely used National Library of Medicine WSD dataset. RESULTS AND DISCUSSION: The authors report a significant improvement when combining those key terms with local context, showing that domain information improves the results of a WSD system based on the Unified Medical Language System Metathesaurus alone. The best results were obtained using key terms obtained by relevance feedback and weighted by inverse document frequency.
Mark Stevenson 0001, Eneko Agirre, Aitor Soroa
J. Am. Medical Informatics Assoc.3
2011 Two birds with one stone: learning semantic models for text categorization and word sense disambiguation
abstract
In this paper we present a novel approach to learning semantic models for multiple domains, which we use to categorize Wikipedia pages and to perform domain Word Sense Disambiguation (WSD). In order to learn a semantic model for each domain we first extract relevant terms from the texts in the domain and then use these terms to initialize a random walk over the WordNet graph. Given an input text, we check the semantic models, choose the appropriate domain for that text and use the best-matching model to perform WSD. Our results show considerable improvements on text categorization and domain WSD tasks.
Roberto Navigli, Stefano Faralli 0001, Aitor Soroa, Oier Lopez de Lacalle, Eneko Agirre
CIKM3
2010 Exploring Knowledge Bases for Similarity
Eneko Agirre, Montse Cuadros, German Rigau, Aitor Soroa
LREC4
2010 Graph-based Word Sense Disambiguation of biomedical documents
abstract
MOTIVATION: Word Sense Disambiguation (WSD), automatically identifying the meaning of ambiguous words in context, is an important stage of text processing. This article presents a graph-based approach to WSD in the biomedical domain. The method is unsupervised and does not require any labeled training data. It makes use of knowledge from the Unified Medical Language System (UMLS) Metathesaurus which is represented as a graph. A state-of-the-art algorithm, Personalized PageRank, is used to perform WSD. RESULTS: When evaluated on the NLM-WSD dataset, the algorithm outperforms other methods that rely on the UMLS Metathesaurus alone. AVAILABILITY: The WSD system is open source licensed and available from http://ixa2.si.ehu.es/ukb/. The UMLS, MetaMap program and NLM-WSD corpus are available from the National Library of Medicine https://www.nlm.nih.gov/research/umls/, http://mmtx.nlm.nih.gov and http://wsd.nlm.nih.gov. Software to convert the NLM-WSD corpus into a format that can be used by our WSD system is available from http://www.dcs.shef.ac.uk/∼marks/biomedical_wsd under open source license.
Eneko Agirre, Aitor Soroa, Mark Stevenson 0001
Bioinform.2
2009 Personalizing PageRank for Word Sense Disambiguation
Eneko Agirre, Aitor Soroa
EACL2
2009 Knowledge-Based WSD and Specific Domains: Performing Better than Generic Supervised WSD
Eneko Agirre, Oier Lopez de Lacalle, Aitor Soroa
IJCAI3
2009 A Study on Similarity and Relatedness Using Distributional and WordNet-based Approaches
Eneko Agirre, Enrique Alfonseca, Keith B. Hall, Jana Kravalova, Marius Pasca, Aitor Soroa
HLT-NAACL6
2009 Dealing With Complex Linguistic Annotations Within a Language Processing Framework
abstract
In this paper we present AWA, a general purpose Annotation Web Architecture for representing, storing, and accessing the information produced by different linguistic processors. The objective of AWA is to establish a coherent and flexible representation scheme that will be the basis for the exchange and use of linguistic information. In morphologically-rich languages as Basque it is necessary to represent and provide easy access to complex phenomena such as intraword structure, declension, derivation and composition features, constituent discontinousness (in multiword expressions) and so on. AWA provides a well-suited schema to deal with these phenomena. The annotation model relies on XML technologies for data representation, storage and retrieval. Typed feature structures are used as a representation schema for linguistic analyses. A consistent underlying data model, which captures the structure and relations contained in the information to be manipulated, has been identified and implemented. AWA is integrated into LPAF, a multilayered Language Processing and Annotation Framework, whose goal is the management and integration of diverse NLP components and resources. Moreover, we introduce EULIA, an annotation tool which exploits and manipulates the data created by the linguistic processors. Two real corpora have been processed and annotated within this framework.
Xabier Artola, Arantza Díaz de Ilarraza, Aitor Soroa, Aitor Sologaistoa
IEEE Trans. Speech Audio Process.3
2008 Using the Multilingual Central Repository for Graph-Based Word Sense Disambiguation
Eneko Agirre, Aitor Soroa
LREC2
2008 Spelling Correction: from Two-Level Morphology to Open Source
Iñaki Alegria, Klara Ceberio, Nerea Ezeiza, Aitor Soroa, Gregorio Hernández
LREC4
2008 ELHISA: An architecture for the integration of heterogeneous lexical information
abstract
Abstract The design and construction of lexical resources is a critical issue in Natural Language Processing (NLP). Real-world NLP systems need large-scale lexica, which provide rich information about words and word senses at all levels: morphologic, syntactic, lexical semantics, etc., but the construction of lexical resources is a difficult and costly task. The last decade has been highly influenced by the notion of reusability, that is, the use of the information of existing lexical resources in constructing new ones. It is unrealistic, however, to expect that the great variety of available lexical information resources could be converted into a single and standard representation schema in the near future. The purpose of this article is to present the ELHISA system, a software architecture for the integration of heterogeneous lexical information. We address, from the point of view of the information integration area, the problem of querying very different existing lexical information sources using a unique and common query language. The integration in ELHISA is performed in a logical way, so that the lexical resources do not suffer any modification when integrating them into the system. ELHISA is primarily defined as a consultation system for accessing structured lexical information, and therefore it does not have the capability to modify or update the underlying information. For this purpose, a General Conceptual Model (GCM) for describing diverse lexical data has been conceived. The GCM establishes a fixed vocabulary describing objects in the lexical information domain, their attributes, and the relationships among them. To integrate the lexical resources into the federation, a Source Conceptual Model (SCM) is built on the top of each one, which represents the lexical objects concurring in each particular source. To answer the user queries, ELHISA must access the integrated resources, and, hence, it must translate the query expressed in GCM terms into queries formulated in terms of the SCM of each source. The relation between the GCM and the SCMs is explicitly described by means of mapping rules called Content Description Rules. Data integration at the extensional level is achieved by means of the data cleansing process, needed if we want to compare the data arriving from different sources. In this process, the object identification step is carried out. Based on this architecture, a prototype named ELHISA has been built, and five resources covering a broad scope have been integrated into it so far for testing purposes. The fact that such heterogeneous resources have been integrated with ease into the system shows, in the opinion of the authors, the suitability of the approach taken.
Xabier Artola, Aitor Soroa
Nat. Lang. Eng.2
2006 Two graph-based algorithms for state-of-the-art WSD
Eneko Agirre, David Martínez 0001, Oier Lopez de Lacalle, Aitor Soroa
EMNLP4
2006 Structure, Annotation and Tools in the Basque ZT Corpus
Nerea Areta, Antton Gurrutxaga, Igor Leturia, Ziortza Polin, Rafa Saiz, Iñaki Alegria, Xabier Artola, Arantza Díaz de Ilarraza, Nerea Ezeiza, Aitor Sologaistoa, Aitor Soroa, Andoni Valverde
LREC11
2002 A Class Library for the Integration of NLP Tools: Definition and implementation of an Abstract Data Type Collection for the manipulation of SGML documents in a context of stand-off linguistic annotation
Xabier Artola, Arantza Díaz de Ilarraza, Nerea Ezeiza, Koldo Gojenola, Gregorio Hernández, Aitor Soroa
LREC6
2000 A Proposal for the Integration of NLP Tools using SGML-Tagged Documents
Xabier Artola, Arantza Díaz de Ilarraza, Nerea Ezeiza, Koldo Gojenola, A. Maritxalar, Aitor Soroa
LREC6
2000 A Methodology for Building Translator-oriented Dictionary Systems
Eneko Agirre, Xabier Arregi, Xabier Artola, Arantza Díaz de Ilarraza, Kepa Sarasola, Aitor Soroa
Mach. Transl.6
1999 MLDS: A translator-oriented MultiLingual dictionary system
abstract
This paper focuses on the design methodology of the MultiLingual Dictionary-System (MLDS), which is a human-oriented tool for assisting in the task of translating lexical units, oriented to translators and conceived from studies carried out with translators. We describe the model adopted for the representation of multilingual dictionary-knowledge. Such a model allows an enriched exploitation of the lexical-semantic relations extracted from dictionaries. In addition, MLDS is supplied with knowledge about the use of the dictionaries in the process of lexical translation, which was elicitated by means of empirical methods and specified in a formal language. The dictionary-knowledge along with the task-oriented knowledge are used to offer the translator active, anticipative and intelligent assistance.
Eneko Agirre, Xabier Arregi, Xabier Artola, Arantza Díaz de Ilarraza, Kepa Sarasola, Aitor Soroa
Nat. Lang. Eng.6
1996 Constructing an intelligent dictionary help system
abstract
This paper discusses different issues in the construction and knowledge representation of an intelligent dictionary help system. The Intelligent Dictionary Help System (IDHS) is conceived as a monolingual (explanatory) dictionary system for human use (Artola and Evrard, 1992). The fact that it is intended for people instead of automatic processing distinguishes it from other systems dealing with the acquisition of semantic knowledge from conventional dictionaries. The system provides various access possibilities to the data, allowing to deduce implicit knowledge from the explicit dictionary information. IDHS deals with reasoning mechanisms analogous to those used by humans when they consult a dictionary. User level functionality of the system has been specified and a prototype has been implemented (Agirre et al., 1994a). A methodology for the extraction of semantic knowledge from a conventional dictionary is described. The method followed in the construction of the phrasal pattern hierarchies required by the parser (Alshawi, 1989) is based on an empirical study carried out on the structure of definition sentences. The results of its application to a real dictionary has shown that the parsing method is particularly suited to the analysis of short definition sentences, as it was the case of the source dictionary. As a result of this process, the characterization of the different lexical-semantic relations between senses is established by means of semantic rules (attached to the patterns); these rules are used for the initial construction of the Dictionary Knowledge Base (DKB). The representation schema proposed for the DKB (Agirre et al., 1994b) is basically a semantic network of frames representing word senses. After construction of the initial DKB, several enrichment processes are performed on the DKB to add new facts to it; these processes are based on the exploitation of the properties of lexical-semantic relations, and also on specially conceived deduction mechanisms. The result of the enrichment processes show the suitability of the representation schema chosen to deduce implicit knowledge. Erroneous deductions are mainly due to incorrect word sense disambiguation.
Eneko Agirre, Xabier Arregi, Xabier Artola, Arantza Díaz de Ilarraza, Kepa Sarasola, Aitor Soroa
Nat. Lang. Eng.6