German Rigau

dblp:66/1456 · DBLP profile ↗
← Back
88ranked-venue papers
1as first author
17since 2021 · last 2026
0000-0003-1119-0930ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 82 · 1 first-author · 17 since 2021Databases, data management, data science and information retrieval · 16 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4Theory of computation · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Merge and Conquer: Instructing Multilingual Models by Adding Target Language Weights
Eneko Valero, Maria Ribalta i Albado, Oscar Sainz, Naiara Pérez, German Rigau
LREC5
2026 SemBench: A Universal Semantic Framework for LLM Evaluation
Mikel Zubillaga, Naiara Pérez, Oscar Sainz, German Rigau
LREC4
2025 IberoBench: A Benchmark for LLM Evaluation in Iberian Languages
abstract
The current best practice to measure the performance of base Large Language Models is to establish a multi-task benchmark that covers a range of capabilities of interest. Currently, however, such benchmarks are only available in a few high-resource languages. To address this situation, we present IberoBench, a multilingual, multi-task benchmark for Iberian languages (i.e., Basque, Catalan, Galician, European Spanish and European Portuguese) built on the LM Evaluation Harness framework. The benchmark consists of 62 tasks divided into 179 subtasks. We evaluate 33 existing LLMs on IberoBench on 0- and 5-shot settings. We also explore the issues we encounter when working with the Harness and our approach to solving them to ensure high-quality evaluation.
Irene Baucells de la Peña, Javier Aula-Blasco, Iria de-Dios-Flores, Silvia Paniagua Suárez, Naiara Pérez, Anna Salles, Susana Sotelo Docío, Júlia Falcão, José Javier Saiz, Robiert Sepúlveda-Torres, Jeremy Barnes 0001, Pablo Gamallo 0001, Aitor Gonzalez-Agirre, German Rigau, Marta Villegas
COLING14
2025 Instructing Large Language Models for Low-Resource Languages: A Systematic Study for Basque
abstract
Oscar Sainz, Naiara Perez, Julen Etxaniz, Joseba Fernandez de Landa, Itziar Aldabe, Iker García-Ferrero, Aimar Zabala, Ekhi Azurmendi, German Rigau, Eneko Agirre, Mikel Artetxe, Aitor Soroa. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Oscar Sainz, Naiara Pérez, Julen Etxaniz, Joseba Fernandez de Landa, Itziar Aldabe, Iker García-Ferrero, Aimar Zabala, Ekhi Azurmendi, German Rigau, Eneko Agirre, Mikel Artetxe, Aitor Soroa
EMNLP9
2024 Latxa: An Open Language Model and Evaluation Suite for Basque
abstract
Julen Etxaniz, Oscar Sainz, Naiara Perez, Itziar Aldabe, German Rigau, Eneko Agirre, Aitor Ormazabal, Mikel Artetxe, Aitor Soroa. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Julen Etxaniz, Oscar Sainz, Naiara Miguel, Itziar Aldabe, German Rigau, Eneko Agirre, Aitor Ormazabal, Mikel Artetxe, Aitor Soroa
ACL (1)5
2024 MedMT5: An Open-Source Multilingual Text-to-Text LLM for the Medical Domain
abstract
Research on language technology for the development of medical applications is currently a hot topic in Natural Language Understanding and Generation. Thus, a number of large language models (LLMs) have recently been adapted to the medical domain, so that they can be used as a tool for mediating in human-AI interaction. While these LLMs display competitive performance on automated medical texts benchmarks, they have been pre-trained and evaluated with a focus on a single language (English mostly). This is particularly true of text-to-text models, which typically require large amounts of domain-specific pre-training data, often not easily accessible for many languages. In this paper, we address these shortcomings by compiling, to the best of our knowledge, the largest multilingual corpus for the medical domain in four languages, namely English, French, Italian and Spanish. This new corpus has been used to train Medical mT5, the first open-source text-to-text multilingual model for the medical domain. Additionally, we present two new evaluation benchmarks for all four languages with the aim of facilitating multilingual research in this domain. A comprehensive evaluation shows that Medical mT5 outperforms both encoders and similarly sized text-to-text models for the Spanish, French, and Italian benchmarks, while being competitive with current state-of-the-art LLMs in English.
Iker García-Ferrero, Rodrigo Agerri, Aitziber Atutxa, Elena Cabrio, Iker de la Iglesia, Alberto Lavelli, Bernardo Magnini, Benjamin Molinet, Johanna Ramirez-Romero, German Rigau, Jose Maria Villa-Gonzalez, Serena Villata, Andrea Zaninello
LREC/COLING10
2024 GoLLIE: Annotation Guidelines improve Zero-Shot Information-Extraction
abstract
Large Language Models (LLMs) combined with instruction tuning have made significant progress when generalizing to unseen tasks. However, they have been less successful in Information Extraction (IE), lagging behind task-specific models. Typically, IE tasks are characterized by complex annotation guidelines which describe the task and give examples to humans. Previous attempts to leverage such information have failed, even with the largest models, as they are not able to follow the guidelines out-of-the-box. In this paper we propose GoLLIE (Guideline-following Large Language Model for IE), a model able to improve zero-shot results on unseen IE tasks by virtue of being fine-tuned to comply with annotation guidelines. Comprehensive evaluation empirically demonstrates that GoLLIE is able to generalize to and follow unseen guidelines, outperforming previous attempts at zero-shot information extraction. The ablation study shows that detailed guidelines is key for good results. Code, data and models will be made publicly available.
Oscar Sainz, Iker García-Ferrero, Rodrigo Agerri, Oier Lopez de Lacalle, German Rigau, Eneko Agirre
ICLR5
2023 This is not a Dataset: A Large Negation Benchmark to Challenge Large Language Models
abstract
Although large language models (LLMs) have apparently acquired a certain level of grammatical knowledge and the ability to make generalizations, they fail to interpret negation, a crucial step in Natural Language Processing.We try to clarify the reasons for the sub-optimal performance of LLMs understanding negation.We introduce a large semi-automatically generated dataset of circa 400,000 descriptive sentences about commonsense knowledge that can be true or false in which negation is present in about 2/3 of the corpus in different forms.We have used our dataset with the largest available open LLMs in a zero-shot approach to grasp their generalization and inference capability and we have also fine-tuned some of the models to assess whether the understanding of negation can be trained.Our findings show that, while LLMs are proficient at classifying affirmative sentences, they struggle with negative sentences and lack a deep understanding of negation, often relying on superficial cues.Although finetuning the models on negative sentences improves their performance, the lack of generalization in handling negation is persistent, highlighting the ongoing challenges of LLMs regarding negation understanding and generalization.The dataset and code are publicly available: https://github.com/hitz-zentroa/ This-is-not-a-Dataset
Iker García-Ferrero, Begoña Altuna, Javier Álvez, Itziar Gonzalez-Dios, German Rigau
EMNLP5
2023 Towards Effective Correction Methods Using WordNet Meronymy Relations
abstract
In this paper, we analyse and compare several correction methods of knowledge resources with the purpose of improving the abilities of systems that require commonsense reasoning with the least possible human-effort.To this end, we cross-check the WordNet meronymy relation member against the knowledge encoded in a SUMO-based first-order logic ontology on the basis of the mapping between WordNet and SUMO.In particular, we focus on the knowledge in WordNet regarding the taxonomy of animals and plants.Despite being created manually, these knowledge resources -WordNet, SUMO and their mapping-are not free of errors and discrepancies.Thus, we propose three correction methods by semi-automatically improving the alignment between WordNet and SUMO, by performing some few corrections in SUMO and by combining the above two strategies.The evaluation of each method includes the required human-effort and the achieved improvement on unseen data from the WebChild project, that is tested using first-order logic automated theorem provers.
Javier Álvez, Itziar Gonzalez-Dios, German Rigau
GWC3
2023 What do Language Models know about word senses? Zero-Shot WSD with Language Models and Domain Inventories
abstract
Language Models are the core for almost any Natural Language Processing system nowadays.One of their particularities is their contextualized representations, a game changer feature when a disambiguation between word senses is necessary.In this paper we aim to explore to what extent language models are capable of discerning among senses at inference time.We performed this analysis by prompting commonly used Languages Models such as BERT or RoBERTa to perform the task of Word Sense Disambiguation (WSD).We leverage the relation between word senses and domains, and cast WSD as a textual entailment problem, where the different hypothesis refer to the domains of the word senses.Our results show that this approach is indeed effective, close to supervised systems.
Oscar Sainz, Oier Lopez de Lacalle, Eneko Agirre, German Rigau
GWC4
2023 Towards the integration of WordNet into ClinIDMap
abstract
This paper presents the integration of Word-Net knowledge resource into ClinIDMap tool, which aims to map identifiers between clinical ontologies and lexical resources.ClinIDMap interlinks identifiers from UMLS, SMOMED-CT, ICD-10 and the corresponding Wikidata and Wikipedia articles for concepts from the UMLS Metathesaurus.The main goal of the tool is to provide semantic interoperability across the clinical concepts from various knowledge bases.As a side effect, the mapping enriches already annotated medical corpora in multiple languages with new labels.In this new release, we add WordNet 3.0 and 3.1 synsets using the available mappings through Wikidata.Thanks to cross-lingual links in MCR we also include the corresponding synsets in other languages and also, extend further ClinIDMap with different domain information.Finally, the final resource helps in the task of enriching of already annotated clinical corpora with additional semantic annotations.
Elena Zotova, Montse Cuadros, German Rigau
GWC3
2023 Negation and speculation processing: A study on cue-scope labelling and assertion classification in Spanish clinical text
abstract
Natural Language Processing (NLP) based on new deep learning technology is contributing to the emergence of powerful solutions that help healthcare providers and researchers discover valuable patterns within insurmountable volumes of health records and scientific literature. Fundamental to the success of such solutions is the processing of negation and speculation. The article addresses this problem with state-of-the-art deep learning approaches from two perspectives: cue and scope labelling, and assertion classification. In light of the real struggle to access clinical annotated data, the study (a) proposes a methodology to automatically convert cue-scope annotations to assertion annotations; and (b) includes a range of scenarios with varying amounts of training data and adversarial test examples. The results expose the clear advantage of Transformer-based models in this regard, managing to overpass a series of baselines and the related work in the public corpus NUBes of clinical Spanish text.
Naiara Pérez, Montse Cuadros, German Rigau
Artif. Intell. Medicine3
2023 A modular approach for multilingual timex detection and normalization using deep learning and grammar-based methods
abstract
Detecting and normalizing temporal expressions is an essential step for many NLP tasks. While a variety of methods have been proposed for detection, best normalization approaches rely on hand-crafted rules. Furthermore, most of them have been designed only for English. In this paper we present a modular multilingual temporal processing system combining a fine-tuned Masked Language Model for detection, and a grammar-based normalizer. We experiment in Spanish and English and compare with HeidelTime, the state-of-the-art in multilingual temporal processing. We obtain best results in gold timex normalization, timex detection and type recognition, and competitive performance in the combined TempEval-3 relaxed value metric. A detailed error analysis shows that detecting only those timexes for which it is feasible to provide a normalization is highly beneficial in this last metric. This raises the question of which is the best strategy for timex processing, namely, leaving undetected those timexes for which is not easy to provide normalization rules or aiming for high coverage.
Nayla Escribano, German Rigau, Rodrigo Agerri
Knowl. Based Syst.2
2022 Overview of the ELE Project
abstract
This paper provides an overview of the ongoing European Language Equality(ELE) project, an 18-month action funded by the European Commission which involves 52 partners. The primary goal of ELE is to prepare the European Language Equality Programme, in the form of a strategic research, innovation and implementation agenda and a roadmap for achieving full digital language equality (DLE) in Europe by 2030.
Itziar Aldabe, Jane Dunne, Aritz Farwell, Owen Gallagher, Federico Gaspari, Maria Giagkou, Jan Hajic 0001, Jens Peter Kückens, Teresa Lynn, Georg Rehm, German Rigau, Katrin Marheinecke, Stelios Piperidis, Natália Resende, Tereza Vojtechová, Andy Way
EAMT11
2022 ClinIDMap: Towards a Clinical IDs Mapping for Data Interoperability
abstract
This paper presents ClinIDMap, a tool for mapping identifiers between clinical ontologies and lexical resources. ClinIDMap interlinks identifiers from UMLS, SMOMED-CT, ICD-10 and the corresponding Wikipedia articles for concepts from the UMLS Metathesaurus. Our main goal is to provide semantic interoperability across the clinical concepts from various knowledge bases. As a side effect, the mapping enriches already annotated corpora in multiple languages with new labels. For instance, spans manually annotated with IDs from UMLS can be annotated with Semantic Types and Groups, and its corresponding SNOMED CT and ICD-10 IDs. We also experiment with sequence labelling models for detecting Diagnosis and Procedures concepts and for detecting UMLS Semantic Groups trained on Spanish, English, and bilingual corpora obtained with the new mapping procedure. The ClinIDMap tool is publicly available.
Elena Zotova, Montse Cuadros, German Rigau
LREC3
2021 Ask2Transformers: Zero-Shot Domain labelling with Pretrained Language Models
abstract
In this paper we present a system that exploits different pre-trained Language Models for assigning domain labels to WordNet synsets without any kind of supervision.Furthermore, the system is not restricted to use a particular set of domain labels.We exploit the knowledge encoded within different off-theshelf pre-trained Language Models and task formulations to infer the domain label of a particular WordNet definition.The proposed zero-shot system achieves a new state-of-theart on the English dataset used in the evaluation.
Oscar Sainz, German Rigau
GWC2
2021 Semi-automatic generation of multilingual datasets for stance detection in Twitter
Elena Zotova, Rodrigo Agerri, German Rigau
Expert Syst. Appl.3
2020 Applying the Closed World Assumption to SUMO-Based FOL Ontologies for Effective Commonsense Reasoning
abstract
Most commonly, the Open World Assumption is adopted as a standard strategy for the design, construction and use of ontologies. This strategy limits the inferencing capabilities of any system because non-asserted statements (missing knowledge) could be assumed to be alternatively true or false. As we will demonstrate, this is especially the case of first-order logic (FOL) ontologies where non-asserted statements is nowadays one of the main obstacles to its practical application in automated commonsense reasoning tasks. In this paper, we investigate the application of the Closed World Assumption (CWA) to enable a better exploitation of FOL ontologies by using state-of-the-art automated theorem provers. To that end, we explore different CWA formulations for the structural knowledge encoded in a FOL translation of the SUMO ontology, discovering that almost 30 % of the structural knowledge is missing. We evaluate these formulations on a practical experimentation using a very large commonsense benchmark obtained from WordNet through its mapping to SUMO. The results show that the competency of the ontology improves more than 50 % when reasoning under the CWA. Thus, applying the CWA automatically to FOL ontologies reduces their ambiguity and more commonsense questions can be answered
Javier Álvez, Itziar Gonzalez-Dios, German Rigau
ECAI3
2020 Language Independent Sequence Labelling for Opinion Target Extraction (Extended Abstract)
abstract
In this paper we present a language independent system to model Opinion Target Extraction (OTE) as a sequence labelling task. The system consists of a combination of clustering features implemented on top of a simple set of shallow local features. Experiments on the well known Aspect Based Sentiment Analysis (ABSA) benchmarks show that our approach is very competitive across languages, obtaining, at the time of writing, best results for six languages in seven different datasets. Furthermore, the results provide further insights into the behaviour of clustering features for sequence labeling tasks. Finally, we also show that these results can be outperformed by recent advances in contextual word embeddings and the transformer architecture. The system and models generated in this work are available for public use and to facilitate reproducibility of results.
Rodrigo Agerri, German Rigau
IJCAI2
2020 NUBes: A Corpus of Negation and Uncertainty in Spanish Clinical Texts
abstract
This paper introduces the first version of the NUBes corpus (Negation and Uncertainty annotations in Biomedical texts in Spanish). The corpus is part of an on-going research and currently consists of 29,682 sentences obtained from anonymised health records annotated with negation and uncertainty. The article includes an exhaustive comparison with similar corpora in Spanish, and presents the main annotation and design decisions. Additionally, we perform preliminary experiments using deep learning algorithms to validate the annotated dataset. As far as we know, NUBes is the largest available corpora for negation in Spanish and the first that also incorporates the annotation of speculation cues, scopes, and events.
Salvador Lima, Naiara Pérez, Montse Cuadros, German Rigau
LREC4
2020 Multilingual Stance Detection in Tweets: The Catalonia Independence Corpus
abstract
Stance detection aims to determine the attitude of a given text with respect to a specific topic or claim. While stance detection has been fairly well researched in the last years, most the work has been focused on English. This is mainly due to the relative lack of annotated data in other languages. The TW-10 referendum Dataset released at IberEval 2018 is a previous effort to provide multilingual stance-annotated data in Catalan and Spanish. Unfortunately, the TW-10 Catalan subset is extremely imbalanced. This paper addresses these issues by presenting a new multilingual dataset for stance detection in Twitter for the Catalan and Spanish languages, with the aim of facilitating research on stance detection in multilingual and cross-lingual settings. The dataset is annotated with stance towards one topic, namely, the ndependence of Catalonia. We also provide a semi-automatic method to annotate the dataset based on a categorization of Twitter users. We experiment on the new corpus with a number of supervised approaches, including linear classifiers and deep learning methods. Comparison of our new corpus with the with the TW-1O dataset shows both the benefits and potential of a well balanced corpus for multilingual and cross-lingual research on stance detection. Finally, we establish new state-of-the-art results on the TW-10 dataset, both for Catalan and Spanish.
Elena Zotova, Rodrigo Agerri, Manuel Núñez 0005, German Rigau
LREC4
2020 Cross-lingual semantic annotation of biomedical literature: experiments in Spanish and English
abstract
MOTIVATION: Biomedical literature is one of the most relevant sources of information for knowledge mining in the field of Bioinformatics. In spite of English being the most widely addressed language in the field; in recent years, there has been a growing interest from the natural language processing community in dealing with languages other than English. However, the availability of language resources and tools for appropriate treatment of non-English texts is lacking behind. Our research is concerned with the semantic annotation of biomedical texts in the Spanish language, which can be considered an under-resourced language where biomedical text processing is concerned. RESULTS: We have carried out experiments to assess the effectiveness of several methods for the automatic annotation of biomedical texts in Spanish. One approach is based on the linguistic analysis of Spanish texts and their annotation using an information retrieval and concept disambiguation approach. A second method takes advantage of a Spanish-English machine translation process to annotate English documents and transfer annotations back to Spanish. A third method takes advantage of the combination of both procedures. Our evaluation shows that a combined system has competitive advantages over the two individual procedures. AVAILABILITY AND IMPLEMENTATION: UMLSMapper (https://snlt.vicomtech.org/umlsmapper) and the annotation transfer tool (http://scientmin.taln.upf.edu/anntransfer/) are freely available for research purposes as web services and/or demos. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Naiara Pérez, Pablo Accuosto, Àlex Bravo, Montse Cuadros, Eva Martínez Garcia, Horacio Saggion, German Rigau
Bioinform.7
2019 Exploiting Metonymy from Available Knowledge Resources
Itziar Gonzalez-Dios, Javier Álvez, German Rigau
CICLing (1)3
2019 Commonsense Reasoning Using WordNet and SUMO: a Detailed Analysis
abstract
We describe a detailed analysis of a sample of large benchmark of commonsense reasoning problems that has been automatically obtained from WordNet, SUMO and their mapping.The objective is to provide a better assessment of the quality of both the benchmark and the involved knowledge resources for advanced commonsense reasoning tasks.By means of this analysis, we are able to detect some knowledge misalignments, mapping errors and lack of knowledge and resources.Our final objective is the extraction of some guidelines towards a better exploitation of this commonsense knowledge framework by the improvement of the included resources.
Javier Álvez, Itziar Gonzalez-Dios, German Rigau
GWC3
2019 Language independent sequence labelling for Opinion Target Extraction
Rodrigo Agerri, German Rigau
Artif. Intell.2
2019 Automatic white-box testing of first-order logic ontologies
abstract
Formal ontologies are axiomatizations in a logic-based formalism. The development of formal ontologies is generating considerable research on the use of automated reasoning techniques and tools that help in ontology engineering. One of the main aims is to refine and to improve axiomatizations for enabling automated reasoning tools to efficiently infer reliable information. Defects in the axiomatization cannot only cause wrong inferences, but can also hinder the inference of expected information, either by increasing the computational cost of or even preventing the inference. In this paper, we introduce a novel, fully automatic white-box testing framework for first-order logic (FOL) ontologies. Our methodology is based on the detection of inference-based redundancies in the given axiomatization. The application of the proposed testing method is fully automatic since (i) the automated generation of tests is guided only by the syntax of axioms and (ii) the evaluation of tests is performed by automated theorem provers (ATPs). Our proposal enables the detection of defects and serves to certify the grade of suitability—for reasoning purposes—of every axiom. We formally define the set of tests that are (automatically) generated from any axiom and prove that every test is logically related to redundancies in the axiom from which the test has been generated. We have implemented our method and used this implementation to automatically detect several non-trivial defects that were hidden in various FOL ontologies. Throughout the paper we provide illustrative examples of these defects, explain how they were found and how each proof—given by an ATP—provides useful hints on the nature of each defect. Additionally, by correcting all the detected defects, we have obtained an improved version of one of the tested ontologies: Adimen-SUMO.
Javier Álvez, Montserrat Hermo, Paqui Lucio, German Rigau
J. Log. Comput.4
2018 Building Named Entity Recognition Taggers via Parallel Corpora
Rodrigo Agerri, Yiling Chung, Itziar Aldabe, Nora Aranberri, Gorka Labaka, German Rigau
LREC6
2018 Developing New Linguistic Resources and Tools for the Galician Language
Rodrigo Agerri, Xavier Gómez Guinovart, German Rigau, Miguel Anxo Solla Portela
LREC3
2018 Cross-checking WordNet and SUMO Using Meronymy
Javier Álvez, Itziar Gonzalez-Dios, German Rigau
LREC3
2018 Biomedical term normalization of EHRs with UMLS
Naiara Pérez, Montse Cuadros, German Rigau
LREC3
2018 Towards Cross-checking WordNet and SUMO Using Meronymy
abstract
We describe the practical application of a black-box testing methodology for the validation of the knowledge encoded in WordNet, SUMO and their mapping by using automated theorem provers.In this paper, we concentrate on the part-whole information provided by WordNet and create a large set of tests on the basis of few question patterns.From our preliminary evaluation results, we report on some of the detected inconsistencies.
Javier Álvez, German Rigau
GWC2
2018 W2VLDA: Almost unsupervised system for Aspect Based Sentiment Analysis
Aitor García-Pablos, Montse Cuadros, German Rigau
Expert Syst. Appl.3
2017 Robust Multilingual Named Entity Recognition with Shallow Semi-supervised Features (Extended Abstract)
abstract
We present a multilingual Named Entity Recognition approach based on a robust and general set of features across languages and datasets. Our system combines shallow local information with clustering semi-supervised features induced on large amounts of unlabeled text. Understanding via empiricalexperimentation how to effectively combine various types of clustering features allows us to seamlessly export our system to other datasets and languages. The result is a simple but highly competitive system which obtains state of the art results across five languages and twelve datasets. The results are reported on standard shared task evaluation data such as CoNLL for English, Spanish and Dutch. Furthermore, and despite the lack of linguistically motivated features, we also report best results for languages such as Basque and German. In addition, we demonstrate that our method also obtains very competitive results even when the amount of supervised data is cut by half, alleviating the dependency on manually annotated data. Finally, the results show that our emphasis on clustering features is crucial to develop robust out-of-domain models. The system and models are freely available to facilitate its use and guarantee the reproducibility of results.
Rodrigo Agerri, German Rigau
IJCAI2
2017 Multi-lingual and Cross-lingual timeline extraction
Egoitz Laparra, Rodrigo Agerri, Itziar Aldabe, German Rigau
Knowl. Based Syst.4
2017 Interpretable semantic textual similarity: Finding and explaining differences between sentences
Iñigo Lopez-Gazpio, Montse Maritxalar, Aitor Gonzalez-Agirre, German Rigau, Larraitz Uria, Eneko Agirre
Knowl. Based Syst.4
2016 A Multilingual Predicate Matrix
Maddalen Lopez de Lacalle, Egoitz Laparra, Itziar Aldabe, German Rigau
LREC4
2016 A Comparison of Domain-based Word Polarity Estimation using different Word Embeddings
Aitor García-Pablos, Montse Cuadros, German Rigau
LREC3
2016 Addressing the MFS Bias in WSD systems
Marten Postma, Rubén Izquierdo, Eneko Agirre, German Rigau, Piek Vossen
LREC4
2016 The Event and Implied Situation Ontology (ESO): Application and Evaluation
Roxane Segers, Marco Rospocher, Piek Vossen, Egoitz Laparra, German Rigau, Anne-Lyse Minard
LREC5
2016 The Predicate Matrix and the Event and Implied Situation Ontology: Making More of Events
abstract
This paper presents the Event and Implied Situation Ontology (ESO), a resource which formalizes the pre and post situations of events and the roles of the entities affected by an event.The ontology reuses and maps across existing resources such as WordNet, SUMO, VerbNet, Prop-Bank and FrameNet.We describe how ESO is injected into a new version of the Predicate Matrix and illustrate how these resources are used to detect information in large document collections that otherwise would have remained implicit.The model targets interpretations of situations rather than the semantics of verbs per se.The event is interpreted as a situation using RDF taking all event components into account.Hence, the ontology and the linked resources need to be considered from the perspective of this interpretation model.
Roxane Segers, Egoitz Laparra, Marco Rospocher, Piek Vossen, German Rigau, Filip Ilievski
GWC5
2016 Robust multilingual Named Entity Recognition with shallow semi-supervised features
Rodrigo Agerri, German Rigau
Artif. Intell.2
2016 Why are these similar? Investigating item similarity types in a large digital library
abstract
We introduce a new problem, identifying the type of relation that holds between a pair of similar items in a digital library. Being able to provide a reason why items are similar has applications in recommendation, personalization, and search. We investigate the problem within the context of Europeana, a large digital library containing items related to cultural heritage. A range of types of similarity in this collection were identified. A set of 1,500 pairs of items from the collection were annotated using crowdsourcing. A high intertagger agreement (average 71.5 Pearson correlation) was obtained and demonstrates that the task is well defined. We also present several approaches to automatically identifying the type of similarity. The best system applies linear regression and achieves a mean Pearson correlation of 71.3, close to human performance. The problem formulation and data set described here were used in a public evaluation exercise, the *SEM shared task on Semantic Textual Similarity. The task attracted the participation of 6 teams, who submitted 14 system runs. All annotations, evaluation scripts, and system runs are freely available.
Aitor Gonzalez-Agirre, German Rigau, Eneko Agirre, Nikolaos Aletras, Mark Stevenson 0001
J. Assoc. Inf. Sci. Technol.2
2016 NewsReader: Using knowledge resources in a cross-lingual reading machine to generate more knowledge from massive streams of news
abstract
In this article, we describe a system that reads news articles in four different languages and detects what happened, who is involved, where and when. This event-centric information is represented as episodic situational knowledge on individuals in an interoperable RDF format that allows for reasoning on the implications of the events. Our system covers the complete path from unstructured text to structured knowledge, for which we defined a formal model that links interpreted textual mentions of things to their representation as instances. The model forms the skeleton for interoperable interpretation across different sources and languages. The real content, however, is defined using multilingual and cross-lingual knowledge resources, both semantic and episodic. We explain how these knowledge resources are used for the processing of text and ultimately define the actual content of the episodic situational knowledge that is reported in the news. The knowledge and model in our system can be seen as an example how the Semantic Web helps NLP. However, our systems also generate massive episodic knowledge of the same type as the Semantic Web is built on. We thus envision a cycle of knowledge acquisition and NLP improvement on a massive scale. This article reports on the details of the system but also on the performance of various high-level components. We demonstrate that our system performs at state-of-the-art level for various subtasks in the four languages of the project, but that we also consider the full integration of these tasks in an overall system with the purpose of reading text. We applied our system to millions of news articles, generating billions of triples expressing formal semantic properties. This shows the capacity of the system to perform at an unprecedented scale.
Piek Vossen, Rodrigo Agerri, Itziar Aldabe, Agata Cybulska, Marieke van Erp, Antske Fokkens, Egoitz Laparra, Anne-Lyse Minard, Alessio Palmero Aprosio, German Rigau, Marco Rospocher, Roxane Segers
Knowl. Based Syst.10
2016 Building event-centric knowledge graphs from news
Marco Rospocher, Marieke van Erp, Piek Vossen, Antske Fokkens, Itziar Aldabe, German Rigau, Aitor Soroa, Thomas Ploeger, Tessel Bogaard
J. Web Semant.6
2015 Improving the Competency of First-Order Ontologies
abstract
We introduce a new framework to evaluate and improve first-order (FO) ontologies using automated theorem provers (ATPs) on the basis of competency questions (CQs). Our framework includes both the adaptation of a methodology for evaluating ontologies to the framework of first-order logic and a new set of non-trivial CQs designed to evaluate FO versions of SUMO, which significantly extends the very small set of CQs proposed in the literature. Most of these new CQs have been automatically generated from a small set of patterns and the mapping of WordNet to SUMO. Applying our framework, we demonstrate that Adimen-SUMO v2.2 outperforms TPTP-SUMO. In addition, using the feedback provided by ATPs we have set an improved version of Adimen-SUMO (v2.4). This new version outperforms the previous ones in terms of competency. For instance, "Humans can reason" is automatically inferred from Adimen-SUMO v2.4, while it is neither deducible from TPTP-SUMO nor Adimen-SUMO v2.2.
Javier Álvez, Paqui Lucio, German Rigau
K-CAP3
2015 Evaluating the Competency of a First-Order Ontology
abstract
We report on the results of evaluating the competency of a first-order ontology for its use with automated theorem provers (ATPs). The evaluation follows the adaptation of the methodology based on competency questions (CQs) [4] to the framework of first-order logic, which is presented in [2], and is applied to Adimen-SUMO [1]. The set of CQs used for this evaluation has been automatically generated from a small set of semantic patterns and the mapping of WordNet to SUMO. Analysing the results, we can conclude that it is feasible to use ATPs for working with Adimen-SUMO v2.4, enabling the resolution of goals by means of performing non-trivial inferences.
Javier Álvez, Paqui Lucio, German Rigau
K-CAP3
2015 Word vs. Class-Based Word Sense Disambiguation
abstract
As empirically demonstrated by the Word Sense Disambiguation (WSD) tasks of the last SensEval/SemEval exercises, assigning the appropriate meaning to words in context has resisted all attempts to be successfully addressed. Many authors argue that one possible reason could be the use of inappropriate sets of word meanings. In particular, WordNet has been used as a de-facto standard repository of word meanings in most of these tasks. Thus, instead of using the word senses defined in WordNet, some approaches have derived semantic classes representing groups of word senses. However, the meanings represented by WordNet have been only used for WSD at a very fine-grained sense level or at a very coarse-grained semantic class level (also called SuperSenses). We suspect that an appropriate level of abstraction could be on between both levels. The contributions of this paper are manifold. First, we propose a simple method to automatically derive semantic classes at intermediate levels of abstraction covering all nominal and verbal WordNet meanings. Second, we empirically demonstrate that our automatically derived semantic classes outperform classical approaches based on word senses and more coarse-grained sense groupings. Third, we also demonstrate that our supervised WSD system benefits from using these new semantic classes as additional semantic features while reducing the amount of training examples. Finally, we also demonstrate the robustness of our supervised semantic class-based WSD system when tested on out of domain corpus.
Rubén Izquierdo, Armando Suárez, German Rigau
J. Artif. Intell. Res.3
2015 Big data for Natural Language Processing: A streaming approach
Rodrigo Agerri, Xabier Artola, Zuhaitz Beloki, German Rigau, Aitor Soroa
Knowl. Based Syst.4
2014 Multilingual, Efficient and Easy NLP Processing with IXA Pipeline
abstract
IXA pipeline is a modular set of Natural Language Processing tools (or pipes) which provide easy access to NLP technology.It aims at lowering the barriers of using NLP technology both for research purposes and for small industrial developers and SMEs by offering robust and efficient linguistic annotation to both researchers and non-NLP experts.IXA pipeline can be used "as is" or exploit its modularity to pick and change different components.This paper describes the general data-centric architecture of IXA pipeline and presents competitive results in several NLP annotations for English and Spanish.
Rodrigo Agerri, Josu Bermudez, German Rigau
EACL3
2014 Simple, Robust and (almost) Unsupervised Generation of Polarity Lexicons for Multiple Languages
abstract
This paper presents a simple, robust and (almost) unsupervised dictionary-based method, qwn-ppv (Q-WordNet as Personalized PageRanking Vector) to automatically generate polarity lexicons.We show that qwn-ppv outperforms other automatically generated lexicons for the four extrinsic evaluations presented here.It also shows very competitive and robust results with respect to manually annotated ones.Results suggest that no single lexicon is best for every task and dataset and that the intrinsic evaluation of polarity lexicons is not a good performance indicator on a Sentiment Analysis task.The qwn-ppv method allows to easily create quality polarity lexicons whenever no domain-based annotated corpora are available for a given language.
Iñaki San Vicente, Rodrigo Agerri, German Rigau
EACL3
2014 IXA pipeline: Efficient and Ready to Use Multilingual NLP tools
Rodrigo Agerri, Josu Bermudez, German Rigau
LREC3
2014 Predicate Matrix: extending SemLink through WordNet mappings
Maddalen Lopez de Lacalle, Egoitz Laparra, German Rigau
LREC3
2014 NewsReader: recording history from daily news streams
Piek Vossen, German Rigau, Luciano Serafini, Pim Stouten, Francis Irving, Willem Robert van Hage
LREC2
2014 First steps towards a Predicate Matrix
abstract
This paper presents the first steps towards building the Predicate Matrix, a new lexical resource resulting from the integration of multiple sources of predicate information including FrameNet (Baker et al., 1997), VerbNet (Kipper, 2005), PropBank (Palmer et al., 2005) and WordNet (Fellbaum, 1998).By using the Predicate Matrix, we expect to provide a more robust interoperable lexicon by discovering and solving inherent inconsistencies among the resources.Moreover, we plan to extend the coverage of current predicate resources (by including from WordNet morphologically related nominal and verbal concepts), to enrich WordNet with predicate information, and possibly to extend predicate information to languages other than English (by exploiting the local wordnets aligned to the English WordNet).
Maddalen Lopez de Lacalle, Egoitz Laparra, German Rigau
GWC3
2013 ImpAr: A Deterministic Algorithm for Implicit Semantic Role Labelling
Egoitz Laparra, German Rigau
ACL (1)2
2012 A Graph-Based Method to Improve WordNet Domains
Aitor Gonzalez-Agirre, German Rigau, Mauro Castillo
CICLing (1)2
2012 Highlighting relevant concepts from Topic Signatures
Montse Cuadros, Lluís Padró 0001, German Rigau
LREC3
2012 A proposal for improving WordNet Domains
Aitor Gonzalez-Agirre, Mauro Castillo, German Rigau
LREC3
2012 Multilingual Central Repository version 3.0
Aitor Gonzalez-Agirre, Egoitz Laparra, German Rigau
LREC3
2012 Mapping WordNet to the Kyoto ontology
Egoitz Laparra, German Rigau, Piek Vossen
LREC2
2012 Adimen-SUMO: Reengineering an Ontology for First-Order Reasoning
abstract
In this paper, the authors present Adimen-SUMO, an operational ontology to be used by first-order theorem provers in intelligent systems that require sophisticated reasoning capabilities (e.g. Natural Language Processing, Knowledge Engineering, Semantic Web infrastructure, etc.). Adimen-SUMO has been obtained by automatically translating around 88% of the original axioms of SUMO (Suggested Upper Merged Ontology). Their main interest is to present in a practical way the advantages of using first-order theorem provers during the design and development of first-order ontologies. First-order theorem provers are applied as inference engines for reengineering a large and complex ontology in order to allow for formal reasoning. In particular, the authors’ study focuses on providing first-order reasoning support to SUMO. During the process, they detect, explain and repair several important design flaws and problems of the SUMO axiomatization. As a by-product, they also provide general design decisions and good practices for creating operational first-order ontologies of any kind.
Javier Álvez, Paqui Lucio, German Rigau
Int. J. Semantic Web Inf. Syst.3
2011 Using Semantic Classes as Document Keywords
Rubén Izquierdo, Armando Suárez, German Rigau
NLDB3
2010 Exploring Knowledge Bases for Similarity
Eneko Agirre, Montse Cuadros, German Rigau, Aitor Soroa
LREC3
2010 Integrating a Large Domain Ontology of Species into WordNet
Montse Cuadros, Egoitz Laparra, German Rigau, Piek Vossen, Wauter Bosma
LREC3
2010 eXtended WordFrameNet
Egoitz Laparra, German Rigau
LREC2
2010 Wikicorpus: A Word-Sense Disambiguated Multilingual Wikipedia Corpus
Samuel Reese, Gemma Boleda, Montse Cuadros, Lluís Padró 0001, German Rigau
LREC5
2009 An Empirical Study on Class-Based Word Sense Disambiguation
Rubén Izquierdo, Armando Suárez, German Rigau
EACL3
2008 KnowNet: Building a Large Net of Knowledge from the Web
Montse Cuadros, German Rigau
COLING2
2008 Complete and Consistent Annotation of WordNet using the Top Concept Ontology
Javier Álvez, Jordi Atserias Batalla, Jordi Carrera, Salvador Climent 0001, Egoitz Laparra, Antoni Oliver 0001, German Rigau
LREC7
2008 WNTERM: Enriching the MCR with a Terminological Dictionary
Eli Pociello, Antton Gurrutxaga, Eneko Agirre, Izaskun Aldezabal, German Rigau
LREC5
2008 KYOTO: a System for Mining, Structuring and Distributing Knowledge across Languages and Cultures
Piek Vossen, Eneko Agirre, Nicoletta Calzolari, Christiane Fellbaum, Shu-Kai Hsieh, Chu-Ren Huang, Hitoshi Isahara, Kyoko Kanzaki, Andrea Marchetti, Monica Monachini, Federico Neri, Remo Raffaelli, German Rigau, Maurizio Tesconi, Joop VanGent
LREC13
2006 Quality Assessment of Large Scale Knowledge Resources
Montse Cuadros, German Rigau
EMNLP2
2005 Combining Knowledge- and Corpus-based Word-Sense-Disambiguation Methods
abstract
In this paper we concentrate on the resolution of the lexical ambiguity that arises when a given word has several different meanings. This specific task is commonly referred to as word sense disambiguation (WSD). The task of WSD consists of assigning the correct sense to words using an electronic dictionary as the source of word definitions. We present two WSD methods based on two main methodological approaches in this research area: a knowledge-based method and a corpus-based method. Our hypothesis is that word-sense disambiguation requires several knowledge sources in order to solve the semantic ambiguity of the words. These sources can be of different kinds--- for example, syntagmatic, paradigmatic or statistical information. Our approach combines various sources of knowledge, through combinations of the two WSD methods mentioned above. Mainly, the paper concentrates on how to combine these methods and sources of information in order to achieve good results in the disambiguation. Finally, this paper presents a comprehensive study and experimental work on evaluation of the methods and their combinations.
Andrés Montoyo, Armando Suárez, German Rigau, Manuel Palomar
J. Artif. Intell. Res.3
2004 Towards the Meaning Top Ontology: Sources of Ontological Meaning
Jordi Atserias Batalla, Salvador Climent 0001, German Rigau
LREC3
2004 Cross-Language Acquisition of Semantic Models for Verbal Predicates
Jordi Atserias Batalla, Bernardo Magnini, Octavian Popescu, Eneko Agirre, Aitziber Atutxa, German Rigau, John Carroll 0001, Rob Koeling
LREC6
2004 Spanish WordNet 1.6: Porting the Spanish Wordnet Across Princeton Versions
Jordi Atserias Batalla, Luis Villarejo, German Rigau
LREC3
2004 Automatic Acquisition of Sense Examples Using ExRetriever
Mauro Castillo, German Rigau, Jordi Atserias Batalla, Jordi Turmo
LREC3
2003 Offline Compilation of Chains for Head-Driven Generation with Constraint-Based Grammars
Antoni Tuells, German Rigau, Horacio Rodríguez
CICLing2
2001 Interface for WordNet Enrichment with Classification Systems
Andrés Montoyo, Manuel Palomar, German Rigau
DEXA3
2000 Mapping WordNets Using Structural Information
abstract
We present a robust approach for linking already existing lexical/semantic hierarchies. We used a constraint satisfaction algorithm (relaxation labeling) to select - among a set of candidates- the node in a target taxonomy that bests matches each node in a source taxonomy. In particular, we use it to map the nominal part of WordNet 1.5 onto WordNet 1.6, with a very high precision and a very low remaining ambiguity.
Jordi Daudé, Lluís Padró 0001, German Rigau
ACL3
2000 Naive Bayes and Exemplar-based Approaches to Word Sense Disambiguation Revisited
Gerard Escudero, Lluís Màrquez, German Rigau
ECAI3
2000 Boosting Applied toe Word Sense Disambiguation
Gerard Escudero, Lluís Màrquez, German Rigau
ECML3
2000 An Empirical Study of the Domain Dependence of Supervised Word Disambiguation Systems
abstract
This paper describes a set of experiments carried out to explore the domain dependence of alternative supervised Word Sense Disambiguation algorithms. The aim of the work is threefold: studying the performance of these algorithms when tested on a different corpus from that they were trained on; exploring their ability to tune to new domains, and demonstrating empirically that the Lazy-Boosting algorithm outperforms state-of-the-art supervised WSD algorithms in both previous situations.
Gerard Escudero, Lluís Màrquez, German Rigau
EMNLP3
1999 Mapping Multilingual Hierarchies Using Relaxation Labeling
Jordi Daudé, Lluís Padró 0001, German Rigau
EMNLP3
1997 Combining Unsupervised Lexical Knowledge Methods for Word Sense Disambiguation
abstract
This paper presents a method to combine a set of unsupervised algorithms that can accurately disambiguate word senses in a large, completely untagged corpus. Although most of the techniques for word sense resolution have been presented as stand-alone, it is our belief that full-fledged lexical ambiguity resolution should combine several information sources and techniques. The set of techniques have been applied in a combined way to disambiguate the genus terms of two machine-readable dictionaries (MRD), enabling us to construct complete taxonomies for Spanish and French. Texted accuracy is above 80% overall and 95% for two-way ambiguous genus terms, showing that texonomy building is not limited to structured dictionaries such as LDOCE.
German Rigau, Jordi Atserias Batalla, Eneko Agirre
ACL1
1996 Word Sense Disambiguation using Conceptual Density
Eneko Agirre, German Rigau
COLING2
1994 TGE: Tlinks Generation Environment
Alicia Ageno, Francesc Ribas, German Rigau, Horacio Rodríguez, Anna Samiotou
COLING3
1994 Acquisition of lexical translation relations from MRDS
Ann Gopestake, Ted Briscoe, Piek Vossen, Alicia Ageno, Irene Castellón, Francesc Ribas, German Rigau, Horacio Rodríguez, Anna Samiotou
Mach. Transl.7