VLDB 2026 Research / reviewers in the wild / expert
Rodrigo Agerri
dblp:57/5047
· DBLP profile ↗
44ranked-venue papers
14as first author
24since 2021 · last 2026
0000-0002-7303-7598ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 43 · 14 first-author · 23 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CRITICS: Critical Science Without Borders by Translation of Scientific KnowledgeabstractThe CRITICS project addresses science accessibility and literacy through the convergence of advanced Machine Translation (MT) based on Large Language Models (LLMs) and educational technology. By leveraging MT systems specifically optimized for scientific content, educational institutions can provide accurate, culturally relevant translations of scientific materials in higher-education students’ native languages, ensuring that complex scientific concepts are comprehensible while maintaining technical accuracy. Novel research on MT specifically tailored for scientific documents aims to break down language barriers in accessing cutting-edge research and educational materials currently only available in high-resourced languages such as English, thereby facilitating the democratization of scientific knowledge across linguistic boundaries. Rodrigo Agerri, Itziar Aldabe, Elena Cabrio, Mark Cieliebak, Jan Deriu, Mariana Flores, Jurgita Kapociute-Dzikiene, Dovile Kuiziniene, Arantza Rico, Aritz Ruiz-González, Aitor Soroa, Mantas Vaskevicius, Serena Villata |
EAMT (2) | 1 |
| 2026 | Physical Commonsense Reasoning for Lower-Resourced Languages and Dialects: A Study on Basque
Jaione Bengoetxea, Itziar Gonzalez-Dios, Rodrigo Agerri |
LREC | 3 |
| 2026 | FactOReS: Fact-checking with an Evidence-based Open Resource in Spanish
Nagore Bravo, Jaione Bengoetxea, Iker García-Ferrero, Alba Bonet-Jover, Estela Saquete Boró, Rodrigo Agerri |
LREC | 6 |
| 2025 | Truth Knows No Language: Evaluating Truthfulness Beyond EnglishabstractBlanca Calvo Figueras, Eneko Sagarzazu, Julen Etxaniz, Jeremy Barnes, Pablo Gamallo, Iria de-Dios-Flores, Rodrigo Agerri. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Blanca Calvo Figueras, Eneko Sagarzazu, Julen Etxaniz, Jeremy Barnes 0001, Pablo Gamallo 0001, Iria de-Dios-Flores, Rodrigo Agerri |
ACL (1) | 7 |
| 2025 | La Leaderboard: A Large Language Model Leaderboard for Spanish Varieties and Languages of Spain and Latin AmericaabstractMaría Grandury, Javier Aula-Blasco, Júlia Falcão, Clémentine Fourrier, Miguel González Saiz, Gonzalo Martínez, Gonzalo Santamaria Gomez, Rodrigo Agerri, Nuria Aldama García, Luis Chiruzzo, Javier Conde, Helena Gomez Adorno, Marta Guerrero Nieto, Guido Ivetta, Natàlia López Fuertes, Flor Miriam Plaza-del-Arco, María-Teresa Martín-Valdivia, Helena Montoro Zamorano, Carmen Muñoz Sanz, Pedro Reviriego, Leire Rosado Plaza, Alejandro Vaca Serrano, Estrella Vallecillo-Rodríguez, Jorge Vallego, Irune Zubiaga. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. María Grandury, Javier Aula-Blasco, Júlia Falcão, Clémentine Fourrier, Miguel González Saiz, Gonzalo Martínez 0001, Gonzalo Santamaría Gómez, Rodrigo Agerri, Nuria Aldama-García, Luis Chiruzzo, Javier Conde, Helena Gómez-Adorno, Marta Guerrero Nieto, Guido Ivetta, Natàlia Fuertes, Flor Miriam Plaza del Arco, María Teresa Martín Valdivia, Helena Montoro Zamorano, Carmen Muñoz Sanz, Pedro Reviriego, Leire Rosado Plaza, Alejandro Vaca Serrano, María Estrella Vallecillo Rodríguez, Jorge Vallego, Irune Zubiaga |
ACL (1) | 8 |
| 2025 | Optimal strategies to perform multilingual analysis of social content for a novel dataset in the tourism domainabstractThe rising influence of social media platforms in various domains, including tourism, has highlighted the growing need for efficient and automated Natural Language Processing (NLP) strategies to take advantage of this valuable resource. However, the transformation of multilingual, unstructured, and informal texts into structured knowledge still poses significant challenges, most notably the never-ending requirement for manually annotated data to train deep learning classifiers. In this work, we study different NLP techniques to establish the best ones to obtain competitive performances while keeping the need for training annotated data to a minimum. To do so, we built the first publicly available multilingual dataset (French, English, and Spanish) for the tourism domain, composed of tourism-related tweets. The dataset includes multilayered, manually revised annotations for Named Entity Recognition (NER) for Locations and Fine-grained Thematic Concepts Extraction mapped to the Thesaurus of Tourism and Leisure Activities of the World Tourism Organization, as well as for Sentiment Analysis at the tweet level. Extensive experimentation comparing various few-shot and fine-tuning techniques with modern language models demonstrate that modern few-shot techniques allow us to obtain competitive results for all three tasks with very little annotation data: 5 tweets per label (15 in total) for Sentiment Analysis, 30 tweets for Named Entity Recognition of Locations and 1K tweets annotated with fine-grained thematic concepts, a highly fine-grained sequence labeling task based on an inventory of 315 classes. We believe that our results, grounded in a novel dataset, pave the way for applying NLP to new domain-specific applications, reducing the need for manual annotations and circumventing the complexities of rule-based, ad-hoc solutions. Maxime Masson, Rodrigo Agerri, Christian Sallaberry, Marie-Noëlle Bessagnet, Annig Lacayrelle, Philippe Roose |
Knowl. Based Syst. | 2 |
| 2024 | Argument Mining in Data Scarce Settings: Cross-lingual Transfer and Few-shot TechniquesabstractRecent research on sequence labelling has been exploring different strategies to mitigate the lack of manually annotated data for the large majority of the world languages.Among others, the most successful approaches have been based on (i) the cross-lingual transfer capabilities of multilingual pre-trained language models (model-transfer), (ii) data translation and label projection (data-transfer) and (iii), promptbased learning by reusing the mask objective to exploit the few-shot capabilities of pre-trained language models (few-shot).Previous work seems to conclude that model-transfer outperforms data-transfer methods and that few-shot techniques based on prompting are superior to updating the model's weights via fine-tuning.In this paper, we empirically demonstrate that, for Argument Mining, a sequence labelling task which requires the detection of long and complex discourse structures, previous insights on cross-lingual transfer or few-shot learning do not apply.Contrary to previous work, we show that for Argument Mining data transfer obtains better results than model-transfer and that finetuning outperforms few-shot methods.Regarding the former, the domain of the dataset used for data-transfer seems to be a deciding factor, while, for few-shot, the type of task (length and complexity of the sequence spans) and sampling method prove to be crucial. Anar Yeginbergen, Maite Oronoz, Rodrigo Agerri |
ACL (1) | 3 |
| 2024 | Basque and Spanish Counter Narrative Generation: Data Creation and EvaluationabstractCounter Narratives (CNs) are non-negative textual responses to Hate Speech (HS) aiming at defusing online hatred and mitigating its spreading across media. Despite the recent increase in HS content posted online, research on automatic CN generation has been relatively scarce and predominantly focused on English. In this paper, we present CONAN-EUS, a new Basque and Spanish dataset for CN generation developed by means of Machine Translation (MT) and professional post-edition. Being a parallel corpus, also with respect to the original English CONAN, it allows to perform novel research on multilingual and crosslingual automatic generation of CNs. Our experiments on CN generation with mT5, a multilingual encoder-decoder model, shows that generation greatly benefits from training on post-edited data, as opposed to relying on silver MT data only. These results are confirmed by their correlation with a qualitative manual evaluation, demonstrating that manually revised training data remains crucial for the quality of the generated CNs. Furthermore, multilingual data augmentation improves results over monolingual settings for structurally similar languages such as English and Spanish, while being detrimental for Basque, a language isolate. Similar findings occur in zero-shot crosslingual evaluations, where model transfer (fine-tuning in English and generating in a different target language) outperforms fine-tuning mT5 on machine translated data for Spanish but not for Basque. This provides an interesting insight into the asymmetry in the multilinguality of generative models, a challenging topic which is still open to research. Data and code will be made publicly available upon publication. Jaione Bengoetxea, Yi-Ling Chung, Marco Guerini, Rodrigo Agerri |
LREC/COLING | 4 |
| 2024 | MedMT5: An Open-Source Multilingual Text-to-Text LLM for the Medical DomainabstractResearch on language technology for the development of medical applications is currently a hot topic in Natural Language Understanding and Generation. Thus, a number of large language models (LLMs) have recently been adapted to the medical domain, so that they can be used as a tool for mediating in human-AI interaction. While these LLMs display competitive performance on automated medical texts benchmarks, they have been pre-trained and evaluated with a focus on a single language (English mostly). This is particularly true of text-to-text models, which typically require large amounts of domain-specific pre-training data, often not easily accessible for many languages. In this paper, we address these shortcomings by compiling, to the best of our knowledge, the largest multilingual corpus for the medical domain in four languages, namely English, French, Italian and Spanish. This new corpus has been used to train Medical mT5, the first open-source text-to-text multilingual model for the medical domain. Additionally, we present two new evaluation benchmarks for all four languages with the aim of facilitating multilingual research in this domain. A comprehensive evaluation shows that Medical mT5 outperforms both encoders and similarly sized text-to-text models for the Spanish, French, and Italian benchmarks, while being competitive with current state-of-the-art LLMs in English. Iker García-Ferrero, Rodrigo Agerri, Aitziber Atutxa, Elena Cabrio, Iker de la Iglesia, Alberto Lavelli, Bernardo Magnini, Benjamin Molinet, Johanna Ramirez-Romero, German Rigau, Jose Maria Villa-Gonzalez, Serena Villata, Andrea Zaninello |
LREC/COLING | 2 |
| 2024 | Evaluating Shortest Edit Script Methods for Contextual LemmatizationabstractModern contextual lemmatizers often rely on automatically induced Shortest Edit Scripts (SES), namely, the number of edit operations to transform a word form into its lemma. In fact, different methods of computing SES have been proposed as an integral component in the architecture of several state-of-the-art contextual lemmatizers currently available. However, previous work has not investigated the direct impact of SES in the final lemmatization performance. In this paper we address this issue by focusing on lemmatization as a token classification task where the only input that the model receives is the word-label pairs in context, where the labels correspond to previously induced SES. Thus, by modifying in our lemmatization system only the SES labels that the model needs to learn, we may then objectively conclude which SES representation produces the best lemmatization results. We experiment with seven languages of different morphological complexity, namely, English, Spanish, Basque, Russian, Czech, Turkish and Polish, using multilingual and language-specific pre-trained masked language encoder-only models as a backbone to build our lemmatizers. Comprehensive experimental results, both in- and out-of-domain, indicate that computing the casing and edit operations separately is beneficial overall, but much more clearly for languages with high-inflected morphology. Notably, multilingual pre-trained language models consistently outperform their language-specific counterparts in every evaluation setting. Olia Toporkov, Rodrigo Agerri |
LREC/COLING | 2 |
| 2024 | ANTIDOTE: ArgumeNtaTIon-Driven explainable artificial intelligence fOr digiTal mEdicineabstractThe need for transparent AI systems in sensitive domains like medicine has become key. In this paper we present ANTIDOTE, a software suite proposing different tools for argumentation-driven explainable Artificial Intelligence for digital medicine. Our system offers the following functionalities: multilingual argumentative analysis for the medical domain, explanation extraction and generation of clinical diagnoses, multilingual large language models for the medical domain, and the first multilingual benchmark for medical question-answering. Experimental results demonstrate the efficacy of ANTIDOTE across different tasks, highlighting its potential as an asset in medical research and practice and fostering transparency, which is crucial for informed decision-making in healthcare. Cristian Cardellino, Theo Alkibiades Collias, Benjamin Molinet, Erwan Hain, Rodrigo Agerri, Serena Villata, Elena Cabrio |
ECAI | 6 |
| 2024 | CasiMedicos-Arg: A Medical Question Answering Dataset Annotated with Explanatory Argumentative StructuresabstractExplaining Artificial Intelligence (AI) decisions is a major challenge nowadays in AI, in particular when applied to sensitive scenarios like medicine and law.However, the need to explain the rationale behind decisions is a main issue also for human-based deliberation as it is important to justify why a certain decision has been taken.Resident medical doctors for instance are required not only to provide a (possibly correct) diagnosis, but also to explain how they reached a certain conclusion.Developing new tools to aid residents to train their explanation skills is therefore a central objective of AI in education.In this paper, we follow this direction, and we present, to the best of our knowledge, the first multilingual dataset for Medical Question Answering where correct and incorrect diagnoses for a clinical case are enriched with a natural language explanation written by doctors.These explanations have been manually annotated with argument components (i.e., premise, claim) and argument relations (i.e., attack, support).The Multilingual CasiMedicosarg dataset consists of 558 clinical cases in four languages (English, Spanish, French, Italian) with explanations, where we annotated 5021 claims, 2313 premises, 2431 support relations, and 1106 attack relations.We conclude by showing how competitive baselines perform over this challenging dataset for the argument mining task. Ekaterina Sviridova, Anar Yeginbergen, Ainara Estarrona, Elena Cabrio, Serena Villata, Rodrigo Agerri |
EMNLP | 6 |
| 2024 | GoLLIE: Annotation Guidelines improve Zero-Shot Information-ExtractionabstractLarge Language Models (LLMs) combined with instruction tuning have made significant progress when generalizing to unseen tasks. However, they have been less successful in Information Extraction (IE), lagging behind task-specific models. Typically, IE tasks are characterized by complex annotation guidelines which describe the task and give examples to humans. Previous attempts to leverage such information have failed, even with the largest models, as they are not able to follow the guidelines out-of-the-box. In this paper we propose GoLLIE (Guideline-following Large Language Model for IE), a model able to improve zero-shot results on unseen IE tasks by virtue of being fine-tuned to comply with annotation guidelines. Comprehensive evaluation empirically demonstrates that GoLLIE is able to generalize to and follow unseen guidelines, outperforming previous attempts at zero-shot information extraction. The ablation study shows that detailed guidelines is key for good results. Code, data and models will be made publicly available. Oscar Sainz, Iker García-Ferrero, Rodrigo Agerri, Oier Lopez de Lacalle, German Rigau, Eneko Agirre |
ICLR | 3 |
| 2024 | MedExpQA: Multilingual benchmarking of Large Language Models for Medical Question AnsweringabstractLarge Language Models (LLMs) have the potential of facilitating the development of Artificial Intelligence technology to assist medical experts for interactive decision support. This potential has been illustrated by the state-of-the-art performance obtained by LLMs in Medical Question Answering, with striking results such as passing marks in licensing medical exams. However, while impressive, the required quality bar for medical applications remains far from being achieved. Currently, LLMs remain challenged by outdated knowledge and by their tendency to generate hallucinated content. Furthermore, most benchmarks to assess medical knowledge lack reference gold explanations which means that it is not possible to evaluate the reasoning of LLMs predictions. Finally, the situation is particularly grim if we consider benchmarking LLMs for languages other than English which remains, as far as we know, a totally neglected topic. In order to address these shortcomings, in this paper we present MedExpQA, the first multilingual benchmark based on medical exams to evaluate LLMs in Medical Question Answering. To the best of our knowledge, MedExpQA includes for the first time reference gold explanations, written by medical doctors, of the correct and incorrect options in the exams. Comprehensive multilingual experimentation using both the gold reference explanations and Retrieval Augmented Generation (RAG) approaches show that performance of LLMs, with best results around 75 accuracy for English, still has large room for improvement, especially for languages other than English, for which accuracy drops 10 points. Therefore, despite using state-of-the-art RAG methods, our results also demonstrate the difficulty of obtaining and integrating readily available medical knowledge that may positively impact results on downstream evaluations for Medical Question Answering. Data, code, and fine-tuned models will be made publicly available.1 Iñigo Alonso 0001, Maite Oronoz, Rodrigo Agerri |
Artif. Intell. Medicine | 3 |
| 2024 | Explanatory argument extraction of correct answers in resident medical examsabstractDeveloping technology to assist medical experts in their everyday decision-making is currently a hot topic in the field of Artificial Intelligence (AI). This is specially true within the framework of Evidence-Based Medicine (EBM), where the aim is to facilitate the extraction of relevant information using natural language as a tool for mediating in human-AI interaction. In this context, AI techniques can be beneficial in finding arguments for past decisions in evolution notes or patient journeys, especially when different doctors are involved in a patient's care. In those documents the decision-making process towards treating the patient is reported. Thus, applying Natural Language Processing (NLP) techniques has the potential to assist doctors in extracting arguments for a more comprehensive understanding of the decisions made. This work focuses on the explanatory argument identification step by setting up the task in a Question Answering (QA) scenario in which clinicians ask questions to the AI model to assist them in identifying those arguments. In order to explore the capabilities of current AI-based language models, we present a new dataset which, unlike previous work: (i) includes not only explanatory arguments for the correct hypothesis, but also arguments to reason on the incorrectness of other hypotheses; (ii) the explanations are written originally in Spanish by doctors to reason over cases from the Spanish Residency Medical Exams. Furthermore, this new benchmark allows us to set up a novel extractive task by identifying the explanation written by medical doctors that supports the correct answer within an argumentative text. An additional benefit of our approach lies in its ability to evaluate the extractive performance of language models using automatic metrics, which in the Antidote CasiMedicos dataset corresponds to a 74.47 F1 score. Comprehensive experimentation shows that our novel dataset and approach is an effective technique to help practitioners in identifying relevant evidence-based explanations for medical questions. Iakes Goenaga, Aitziber Atutxa, Koldo Gojenola, Maite Oronoz, Rodrigo Agerri |
Artif. Intell. Medicine | 5 |
| 2024 | On the Role of Morphological Information for Contextual LemmatizationabstractAbstract Lemmatization is a natural language processing (NLP) task that consists of producing, from a given inflected word, its canonical form or lemma. Lemmatization is one of the basic tasks that facilitate downstream NLP applications, and is of particular importance for high-inflected languages. Given that the process to obtain a lemma from an inflected word can be explained by looking at its morphosyntactic category, including fine-grained morphosyntactic information to train contextual lemmatizers has become common practice, without considering whether that is the optimum in terms of downstream performance. In order to address this issue, in this article we empirically investigate the role of morphological information to develop contextual lemmatizers in six languages within a varied spectrum of morphological complexity: Basque, Turkish, Russian, Czech, Spanish, and English. Furthermore, and unlike the vast majority of previous work, we also evaluate lemmatizers in out-of-domain settings, which constitutes, after all, their most common application use. The results of our study are rather surprising. It turns out that providing lemmatizers with fine-grained morphological features during training is not that beneficial, not even for agglutinative languages. In fact, modern contextual word representations seem to implicitly encode enough morphological information to obtain competitive contextual lemmatizers without seeing any explicit morphological signal. Moreover, our experiments suggest that the best lemmatizers out-of-domain are those using simple UPOS tags or those trained without morphology and, lastly, that current evaluation practices for lemmatization are not adequate to clearly discriminate between models. Olia Toporkov, Rodrigo Agerri |
Comput. Linguistics | 2 |
| 2023 | APs: A Proxemic Framework for Social Media Interactions Modeling and Analysis
Maxime Masson, Philippe Roose, Christian Sallaberry, Rodrigo Agerri, Marie-Noëlle Bessagnet, Annig Lacayrelle |
IDA | 4 |
| 2023 | A modular approach for multilingual timex detection and normalization using deep learning and grammar-based methodsabstractDetecting and normalizing temporal expressions is an essential step for many NLP tasks. While a variety of methods have been proposed for detection, best normalization approaches rely on hand-crafted rules. Furthermore, most of them have been designed only for English. In this paper we present a modular multilingual temporal processing system combining a fine-tuned Masked Language Model for detection, and a grammar-based normalizer. We experiment in Spanish and English and compare with HeidelTime, the state-of-the-art in multilingual temporal processing. We obtain best results in gold timex normalization, timex detection and type recognition, and competitive performance in the combined TempEval-3 relaxed value metric. A detailed error analysis shows that detecting only those timexes for which it is feasible to provide a normalization is highly beneficial in this last metric. This raises the question of which is the best strategy for timex processing, namely, leaving undetected those timexes for which is not easy to provide normalization rules or aiming for high coverage. Nayla Escribano, German Rigau, Rodrigo Agerri |
Knowl. Based Syst. | 3 |
| 2022 | Leveraging a New Spanish Corpus for Multilingual and Cross-lingual Metaphor DetectionabstractThe lack of wide coverage datasets annotated with everyday metaphorical expressions for languages other than English is striking.This means that most research on supervised metaphor detection has been published only for that language.In order to address this issue, this work presents the first corpus annotated with naturally occurring metaphors in Spanish large enough to develop systems to perform metaphor detection.The presented dataset, CoMeta, includes texts from various domains, namely, news, political discourse, Wikipedia and reviews.In order to label CoMeta, we apply the MIPVU method, the guidelines most commonly used to systematically annotate metaphor on real data.We use our newly created dataset to provide competitive baselines by fine-tuning several multilingual and monolingual state-of-the-art large language models.Furthermore, by leveraging the existing VUAM English data in addition to CoMeta, we present the, to the best of our knowledge, first cross-lingual experiments on supervised metaphor detection.Finally, we perform a detailed error analysis that explores the seemingly high transfer of everyday metaphor across these two languages and datasets. Elisa Sanchez-Bayona, Rodrigo Agerri |
CoNLL | 2 |
| 2022 | Does Corpus Quality Really Matter for Low-Resource Languages?abstractThe vast majority of non-English corpora are derived from automatically filtered versions of CommonCrawl.While prior work has identified major issues on the quality of these datasets (Kreutzer et al., 2021), it is not clear how this impacts downstream performance.Taking representation learning in Basque as a case study, we explore tailored crawling-manually identifying and scraping websites with high-quality content-as an alternative to filtering Common-Crawl.Our new corpus, called EusCrawl, is similar in size to the Basque portion of popular multilingual corpora like CC100 and mC4, yet it has a much higher quality according to native annotators.For instance, 66% of documents are rated as high-quality for EusCrawl, in contrast with < 33% for both mC4 and CC100.Nevertheless, we obtain similar results on downstream NLU tasks regardless of the corpus used for pre-training.Our work suggests that NLU performance in low-resource languages is not primarily constrained by the quality of the data, and other factors like corpus size and domain coverage can play a more important role. Mikel Artetxe, Itziar Aldabe, Rodrigo Agerri, Olatz Perez-de-Viñaspre, Aitor Soroa |
EMNLP | 3 |
| 2022 | BasqueParl: A Bilingual Corpus of Basque Parliamentary TranscriptionsabstractParliamentary transcripts provide a valuable resource to understand the reality and know about the most important facts that occur over time in our societies. Furthermore, the political debates captured in these transcripts facilitate research on political discourse from a computational social science perspective. In this paper we release the first version of a newly compiled corpus from Basque parliamentary transcripts. The corpus is characterized by heavy Basque-Spanish code-switching, and represents an interesting resource to study political discourse in contrasting languages such as Basque and Spanish. We enrich the corpus with metadata related to relevant attributes of the speakers and speeches (language, gender, party...) and process the text to obtain named entities and lemmas. The obtained metadata is then used to perform a detailed corpus analysis which provides interesting insights about the language use of the Basque political representatives across time, parties and gender. Nayla Escribano, Jon Ander González, Julen Orbegozo-Terradillos, Ainara Larrondo-Ureta, Simón Peña-Fernández, Olatz Perez-de-Viñaspre, Rodrigo Agerri |
LREC | 7 |
| 2022 | BasqueGLUE: A Natural Language Understanding Benchmark for BasqueabstractNatural Language Understanding (NLU) technology has improved significantly over the last few years and multitask benchmarks such as GLUE are key to evaluate this improvement in a robust and general way. These benchmarks take into account a wide and diverse set of NLU tasks that require some form of language understanding, beyond the detection of superficial, textual clues. However, they are costly to develop and language-dependent, and therefore they are only available for a small number of languages. In this paper, we present BasqueGLUE, the first NLU benchmark for Basque, a less-resourced language, which has been elaborated from previously existing datasets and following similar criteria to those used for the construction of GLUE and SuperGLUE. We also report the evaluation of two state-of-the-art language models for Basque on BasqueGLUE, thus providing a strong baseline to compare upon. BasqueGLUE is freely available under an open license. Gorka Urbizu, Iñaki San Vicente, Xabier Saralegi, Rodrigo Agerri, Aitor Soroa |
LREC | 4 |
| 2022 | A Domain-Independent Method for Thematic Dataset Building from Social Media: The Case of Tourism on Twitter
Maxime Masson, Christian Sallaberry, Rodrigo Agerri, Marie-Noëlle Bessagnet, Philippe Roose, Annig Lacayrelle |
WISE | 3 |
| 2021 | Semi-automatic generation of multilingual datasets for stance detection in Twitter
Elena Zotova, Rodrigo Agerri, German Rigau |
Expert Syst. Appl. | 2 |
| 2020 | Language Independent Sequence Labelling for Opinion Target Extraction (Extended Abstract)abstractIn this paper we present a language independent system to model Opinion Target Extraction (OTE) as a sequence labelling task. The system consists of a combination of clustering features implemented on top of a simple set of shallow local features. Experiments on the well known Aspect Based Sentiment Analysis (ABSA) benchmarks show that our approach is very competitive across languages, obtaining, at the time of writing, best results for six languages in seven different datasets. Furthermore, the results provide further insights into the behaviour of clustering features for sequence labeling tasks. Finally, we also show that these results can be outperformed by recent advances in contextual word embeddings and the transformer architecture. The system and models generated in this work are available for public use and to facilitate reproducibility of results. Rodrigo Agerri, German Rigau |
IJCAI | 1 |
| 2020 | Give your Text Representation Models some Love: the Case for BasqueabstractWord embeddings and pre-trained language models allow to build rich representations of text and have enabled improvements across most NLP tasks. Unfortunately they are very expensive to train, and many small companies and research groups tend to use models that have been pre-trained and made available by third parties, rather than building their own. This is suboptimal as, for many languages, the models have been trained on smaller (or lower quality) corpora. In addition, monolingual pre-trained models for non-English languages are not always available. At best, models for those languages are included in multilingual versions, where each language shares the quota of substrings and parameters with the rest of the languages. This is particularly true for smaller languages such as Basque. In this paper we show that a number of monolingual models (FastText word embeddings, FLAIR and BERT language models) trained with larger Basque corpora produce much better results than publicly available versions in downstream NLP tasks, including topic classification, sentiment classification, PoS tagging and NER. This work sets a new state-of-the-art in those tasks for Basque. All benchmarks and models used in this work are publicly available. Rodrigo Agerri, Iñaki San Vicente, Jon Ander Campos, Ander Barrena, Xabier Saralegi, Aitor Soroa, Eneko Agirre |
LREC | 1 |
| 2020 | Multilingual Stance Detection in Tweets: The Catalonia Independence CorpusabstractStance detection aims to determine the attitude of a given text with respect to a specific topic or claim. While stance detection has been fairly well researched in the last years, most the work has been focused on English. This is mainly due to the relative lack of annotated data in other languages. The TW-10 referendum Dataset released at IberEval 2018 is a previous effort to provide multilingual stance-annotated data in Catalan and Spanish. Unfortunately, the TW-10 Catalan subset is extremely imbalanced. This paper addresses these issues by presenting a new multilingual dataset for stance detection in Twitter for the Catalan and Spanish languages, with the aim of facilitating research on stance detection in multilingual and cross-lingual settings. The dataset is annotated with stance towards one topic, namely, the ndependence of Catalonia. We also provide a semi-automatic method to annotate the dataset based on a categorization of Twitter users. We experiment on the new corpus with a number of supervised approaches, including linear classifiers and deep learning methods. Comparison of our new corpus with the with the TW-1O dataset shows both the benefits and potential of a well balanced corpus for multilingual and cross-lingual research on stance detection. Finally, we establish new state-of-the-art results on the TW-10 dataset, both for Catalan and Spanish. Elena Zotova, Rodrigo Agerri, Manuel Núñez 0005, German Rigau |
LREC | 2 |
| 2019 | Language independent sequence labelling for Opinion Target Extraction
Rodrigo Agerri, German Rigau |
Artif. Intell. | 1 |
| 2018 | Building Named Entity Recognition Taggers via Parallel Corpora
Rodrigo Agerri, Yiling Chung, Itziar Aldabe, Nora Aranberri, Gorka Labaka, German Rigau |
LREC | 1 |
| 2018 | Developing New Linguistic Resources and Tools for the Galician Language
Rodrigo Agerri, Xavier Gómez Guinovart, German Rigau, Miguel Anxo Solla Portela |
LREC | 1 |
| 2018 | Annotating Abstract Meaning Representations for Spanish
Noelia Migueles-Abraira, Rodrigo Agerri, Arantza Díaz de Ilarraza |
LREC | 2 |
| 2017 | Robust Multilingual Named Entity Recognition with Shallow Semi-supervised Features (Extended Abstract)abstractWe present a multilingual Named Entity Recognition approach based on a robust and general set of features across languages and datasets. Our system combines shallow local information with clustering semi-supervised features induced on large amounts of unlabeled text. Understanding via empiricalexperimentation how to effectively combine various types of clustering features allows us to seamlessly export our system to other datasets and languages. The result is a simple but highly competitive system which obtains state of the art results across five languages and twelve datasets. The results are reported on standard shared task evaluation data such as CoNLL for English, Spanish and Dutch. Furthermore, and despite the lack of linguistically motivated features, we also report best results for languages such as Basque and German. In addition, we demonstrate that our method also obtains very competitive results even when the amount of supervised data is cut by half, alleviating the dependency on manually annotated data. Finally, the results show that our emphasis on clustering features is crucial to develop robust out-of-domain models. The system and models are freely available to facilitate its use and guarantee the reproducibility of results. Rodrigo Agerri, German Rigau |
IJCAI | 1 |
| 2017 | Multi-lingual and Cross-lingual timeline extraction
Egoitz Laparra, Rodrigo Agerri, Itziar Aldabe, German Rigau |
Knowl. Based Syst. | 2 |
| 2016 | Robust multilingual Named Entity Recognition with shallow semi-supervised features
Rodrigo Agerri, German Rigau |
Artif. Intell. | 1 |
| 2016 | NewsReader: Using knowledge resources in a cross-lingual reading machine to generate more knowledge from massive streams of newsabstractIn this article, we describe a system that reads news articles in four different languages and detects what happened, who is involved, where and when. This event-centric information is represented as episodic situational knowledge on individuals in an interoperable RDF format that allows for reasoning on the implications of the events. Our system covers the complete path from unstructured text to structured knowledge, for which we defined a formal model that links interpreted textual mentions of things to their representation as instances. The model forms the skeleton for interoperable interpretation across different sources and languages. The real content, however, is defined using multilingual and cross-lingual knowledge resources, both semantic and episodic. We explain how these knowledge resources are used for the processing of text and ultimately define the actual content of the episodic situational knowledge that is reported in the news. The knowledge and model in our system can be seen as an example how the Semantic Web helps NLP. However, our systems also generate massive episodic knowledge of the same type as the Semantic Web is built on. We thus envision a cycle of knowledge acquisition and NLP improvement on a massive scale. This article reports on the details of the system but also on the performance of various high-level components. We demonstrate that our system performs at state-of-the-art level for various subtasks in the four languages of the project, but that we also consider the full integration of these tasks in an overall system with the purpose of reading text. We applied our system to millions of news articles, generating billions of triples expressing formal semantic properties. This shows the capacity of the system to perform at an unprecedented scale. Piek Vossen, Rodrigo Agerri, Itziar Aldabe, Agata Cybulska, Marieke van Erp, Antske Fokkens, Egoitz Laparra, Anne-Lyse Minard, Alessio Palmero Aprosio, German Rigau, Marco Rospocher, Roxane Segers |
Knowl. Based Syst. | 2 |
| 2015 | Big data for Natural Language Processing: A streaming approach
Rodrigo Agerri, Xabier Artola, Zuhaitz Beloki, German Rigau, Aitor Soroa |
Knowl. Based Syst. | 1 |
| 2014 | Multilingual, Efficient and Easy NLP Processing with IXA PipelineabstractIXA pipeline is a modular set of Natural Language Processing tools (or pipes) which provide easy access to NLP technology.It aims at lowering the barriers of using NLP technology both for research purposes and for small industrial developers and SMEs by offering robust and efficient linguistic annotation to both researchers and non-NLP experts.IXA pipeline can be used "as is" or exploit its modularity to pick and change different components.This paper describes the general data-centric architecture of IXA pipeline and presents competitive results in several NLP annotations for English and Spanish. Rodrigo Agerri, Josu Bermudez, German Rigau |
EACL | 1 |
| 2014 | Simple, Robust and (almost) Unsupervised Generation of Polarity Lexicons for Multiple LanguagesabstractThis paper presents a simple, robust and (almost) unsupervised dictionary-based method, qwn-ppv (Q-WordNet as Personalized PageRanking Vector) to automatically generate polarity lexicons.We show that qwn-ppv outperforms other automatically generated lexicons for the four extrinsic evaluations presented here.It also shows very competitive and robust results with respect to manually annotated ones.Results suggest that no single lexicon is best for every task and dataset and that the intrinsic evaluation of polarity lexicons is not a good performance indicator on a Sentiment Analysis task.The qwn-ppv method allows to easily create quality polarity lexicons whenever no domain-based annotated corpora are available for a given language. Iñaki San Vicente, Rodrigo Agerri, German Rigau |
EACL | 2 |
| 2014 | IXA pipeline: Efficient and Ready to Use Multilingual NLP tools
Rodrigo Agerri, Josu Bermudez, German Rigau |
LREC | 1 |
| 2014 | Generating Polarity Lexicons with WordNet propagation in 5 languages
Isa Maks, Rubén Izquierdo, Francesca Frontini, Rodrigo Agerri, Piek Vossen, Andoni Azpeitia |
LREC | 4 |
| 2012 | SUMAT: Data Collection and Parallel Corpus Compilation for Machine Translation of Subtitles
Volha Petukhova, Rodrigo Agerri, Mark Fishel, Sergio Penkale, Arantza del Pozo, Mirjam Sepesy Maucec, Andy Way, Panayota Georgakopoulou, Martin Volk 0001 |
LREC | 2 |
| 2010 | On the Automatic Generation of Intermediate Logic Forms for WordNet Glosses
Rodrigo Agerri, Anselmo Peñas |
CICLing | 1 |
| 2010 | Q-WordNet: Extracting Polarity from WordNet Senses
Rodrigo Agerri, Ana García-Serrano |
LREC | 1 |
| 2007 | On the formalization of Invariant Mappings for Metaphor Interpretation
Rodrigo Agerri, John A. Barnden, Mark Lee 0001, Alan M. Wallington |
ACL | 1 |