Aitor Gonzalez-Agirre

dblp:126/6335 · DBLP profile ↗
← Back
17ranked-venue papers
5as first author
11since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 4 first-author · 11 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 ACAData: Parallel Dataset of Academic Data for Machine Translation
abstract
We present ACADATA, a high-quality parallel dataset for academic translation, that consists of two subsets: ACAD-TRAIN, which contains approximately 1.5 million author-generated paragraph pairs across 96 language directions and ACAD-BENCH, a curated evaluation set of almost 6,000 translations covering 12 directions. To validate its utility, we fine-tune two Large Language Models (LLMs) on ACAD-TRAIN and benchmark them on ACAD-BENCH against specialized machine-translation systems, general-purpose, open-weight LLMs, and several large-scale proprietary models. Experimental results demonstrate that fine-tuning on ACAD-TRAIN leads to improvements in academic translation quality by +6.1 and +12.4 d-BLEU points on average for 7B and 2B models respectively, while also improving long-context translation in a general domain by up to 24.9% when translating out of English. The fine-tuned top-performing model surpasses the best propietary and open-weight models on academic translation domain. By releasing ACAD-TRAIN, ACAD-BENCH and the fine-tuned models, we provide the community with a valuable resource to advance research in academic domain and long-context translation.
Iñaki Lacunza, Javier García Gilabert, Francesca de Luca Fornaciari, Javier Aula-Blasco, Aitor Gonzalez-Agirre, Maite Melero, Marta Villegas
LREC5
2026 EsBBQ and CaBBQ: The Spanish and Catalan Bias Benchmarks for Question Answering
abstract
Previous literature has largely shown that Large Language Models (LLMs) perpetuate social biases learnt from their pre-training data. Given the notable lack of resources for social bias evaluation in languages other than English, and for social contexts outside of the United States, this paper introduces the Spanish and the Catalan Bias Benchmarks for Question Answering (EsBBQ and CaBBQ). Based on the original BBQ, these two parallel datasets are designed to assess social bias across 10 categories using a multiple-choice QA setting, now adapted to the Spanish and Catalan languages and to the social context of Spain. We report evaluation results on different LLMs, factoring in model family, size and variant. Our results show that models tend to fail to choose the correct answer in ambiguous scenarios, and that high QA accuracy often correlates with greater reliance on social biases.
Valle Ruíz-Fernández, Mario Mina, Júlia Falcão, Luis Vasquez-Reina, Anna Salles, Aitor Gonzalez-Agirre, Olatz Perez-de-Viñaspre
LREC6
2025 XDoGE: Multilingual Data Reweighting to Enhance Language Inclusivity in LLMs
Iñaki Lacunza, José Javier Saiz, Alexander Shvets, Aitor Gonzalez-Agirre, Marta Villegas
IEEE Big Data4
2025 VeritasQA: A Truthfulness Benchmark Aimed at Multilingual Transferability
abstract
As Large Language Models (LLMs) become available in a wider range of domains and applications, evaluating the truthfulness of multilingual LLMs is an issue of increasing relevance. TruthfulQA (Lin et al., 2022) is one of few benchmarks designed to evaluate how models imitate widespread falsehoods. However, it is strongly English-centric and starting to become outdated. We present VeritasQA, a context- and time-independent truthfulness benchmark built with multilingual transferability in mind, and available in Spanish, Catalan, Galician and English. VeritasQA comprises a set of 353 questions and answers inspired by common misconceptions and falsehoods that are not tied to any particular country or recent event. We release VeritasQA under an open license and present the evaluation results of 15 models of various architectures and sizes.
Javier Aula-Blasco, Júlia Falcão, Susana Sotelo, Silvia Paniagua Suárez, Aitor Gonzalez-Agirre, Marta Villegas
COLING5
2025 IberoBench: A Benchmark for LLM Evaluation in Iberian Languages
abstract
The current best practice to measure the performance of base Large Language Models is to establish a multi-task benchmark that covers a range of capabilities of interest. Currently, however, such benchmarks are only available in a few high-resource languages. To address this situation, we present IberoBench, a multilingual, multi-task benchmark for Iberian languages (i.e., Basque, Catalan, Galician, European Spanish and European Portuguese) built on the LM Evaluation Harness framework. The benchmark consists of 62 tasks divided into 179 subtasks. We evaluate 33 existing LLMs on IberoBench on 0- and 5-shot settings. We also explore the issues we encounter when working with the Harness and our approach to solving them to ensure high-quality evaluation.
Irene Baucells de la Peña, Javier Aula-Blasco, Iria de-Dios-Flores, Silvia Paniagua Suárez, Naiara Pérez, Anna Salles, Susana Sotelo Docío, Júlia Falcão, José Javier Saiz, Robiert Sepúlveda-Torres, Jeremy Barnes 0001, Pablo Gamallo 0001, Aitor Gonzalez-Agirre, German Rigau, Marta Villegas
COLING13
2025 Cognitive Biases, Task Complexity, and Result Intepretability in Large Language Models
Mario Mina, Valle Ruíz-Fernández, Júlia Falcão, Luis Vasquez-Reina, Aitor Gonzalez-Agirre
COLING5
2025 Multi-LMentry: Can Multilingual LLMs Solve Elementary Tasks Across Languages?
abstract
Luca Moroni, Javier Aula-Blasco, Simone Conia, Irene Baucells, Naiara Perez, Silvia Paniagua Suárez, Anna Sallés, Malte Ostendorff, Júlia Falcão, Guijin Son, Aitor Gonzalez-Agirre, Roberto Navigli, Marta Villegas. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Luca Moroni, Javier Aula-Blasco, Simone Conia, Irene Baucells de la Peña, Naiara Pérez, Silvia Paniagua Suárez, Anna Salles, Malte Ostendorff, Júlia Falcão, Guijin Son, Aitor Gonzalez-Agirre, Roberto Navigli, Marta Villegas
EMNLP11
2024 FLOR: On the Effectiveness of Language Adaptation
abstract
Large language models have amply proven their great capabilities, both in downstream tasks and real-life settings. However, low- and mid-resource languages do not have access to the necessary means to train such models from scratch, and often have to rely on multilingual models despite being underrepresented in the training data. For the particular case of the Catalan language, we prove that continued pre-training with vocabulary adaptation is a better alternative to take the most out of already pre-trained models, even if these have not seen any Catalan data during their pre-training phase. We curate a 26B tokens corpus and use it to further pre-train BLOOM, giving rise to the FLOR models. We perform an extensive evaluation to assess the effectiveness of our method, obtaining consistent gains across Catalan and Spanish tasks. The models, training data, and evaluation framework are made freely available under permissive licenses.
Severino Da Dalt, Joan Llop-Palao, Irene Baucells de la Peña, Marc Pàmies, Yishi Xu, Aitor Gonzalez-Agirre, Marta Villegas
LREC/COLING6
2024 Building a Data Infrastructure for a Mid-Resource Language: The Case of Catalan
abstract
Current LLM-based applications are becoming steadily available for everyone with a reliable access to technology and the internet. These applications offer benefits to their users that leave those without access to them at a serious disadvantage. Given the vastly large amount of data needed to train LLMs, the gap between languages with access to such quantity of data and those without it is currently larger than ever. Aimed at saving this gap, the Aina Project was created to provide Catalan with the necessary resources to keep being relevant in the context of AI/NLP applications based on LLMs. We thus present a set of strategies to consider when improving technology support for a mid- or low-resource language, specially addressing sustainability of high-quality data acquisition and the challenges involved in the process. We also introduce a large amount of new annotated data for Catalan. Our hope is that those interested in replicating this work for another language can learn from what worked for us, the challenges that we faced, and the sometimes disheartening truth of working with mid- and low-resource languages.
Aitor Gonzalez-Agirre, Montserrat Marimon, Carlos Rodríguez Penagos, Javier Aula-Blasco, Irene Baucells de la Peña, Carme Armentano-Oller, Jorge Palomar-Giner, Baybars Külebi, Marta Villegas
LREC/COLING1
2024 A CURATEd CATalog: Rethinking the Extraction of Pretraining Corpora for Mid-Resourced Languages
abstract
We present and describe two language resources in this paper: CATalog 1.0, the largest text corpus in Catalan to date, and CURATE (Corpus Utility for RAting TExt), a modular, parallelizable pipeline used for processing and scoring documents based on text quality that we have optimised to run in High Performance Cluster (HPC) environments. In the coming sections we describe our data preprocessing pipeline at length; traditional pipelines usually implement a set of binary filters such that a given document is either in or out. In our experience with Catalan, in lower-resource settings it is more practical to instead assign a document a soft score to allow for more flexible decision-making. We describe how the document score is calculated and highlight its interpretability by showing that it is significantly correlated with human judgements as obtained from a comparative judgement experiment. We additionally describe the different subcorpora that make up CATalog 1.0.
Jorge Palomar-Giner, José Javier Saiz, Ferran Espuña, Mario Mina, Severino Da Dalt, Joan Llop-Palao, Malte Ostendorff, Pedro Ortiz Suarez, Georg Rehm, Aitor Gonzalez-Agirre, Marta Villegas
LREC/COLING10
2023 A weakly supervised textual entailment approach to zero-shot text classification
abstract
Marc Pàmies, Joan Llop, Francesco Multari, Nicolau Duran-Silva, César Parra-Rojas, Aitor Gonzalez-Agirre, Francesco Alessandro Massucci, Marta Villegas. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023.
Marc Pàmies, Joan Llop-Palao, Francesco Multari, Nicolau Duran-Silva, César Parra-Rojas, Aitor Gonzalez-Agirre, Francesco Alessandro Massucci, Marta Villegas
EACL6
2020 Conversational Question Answering in Low Resource Scenarios: A Dataset and Case Study for Basque
abstract
Conversational Question Answering (CQA) systems meet user information needs by having conversations with them, where answers to the questions are retrieved from text. There exist a variety of datasets for English, with tens of thousands of training examples, and pre-trained language models have allowed to obtain impressive results. The goal of our research is to test the performance of CQA systems under low-resource conditions which are common for most non-English languages: small amounts of native annotations and other limitations linked to low resource languages, like lack of crowdworkers or smaller wikipedias. We focus on the Basque language, and present the first non-English CQA dataset and results. Our experiments show that it is possible to obtain good results with low amounts of native data thanks to cross-lingual transfer, with quality comparable to those obtained for English. We also discovered that dialogue history models are not directly transferable to another language, calling for further research. The dataset is publicly available.
Arantxa Otegi, Aitor Gonzalez-Agirre, Jon Ander Campos, Aitor Soroa, Eneko Agirre
LREC2
2017 Interpretable semantic textual similarity: Finding and explaining differences between sentences
Iñigo Lopez-Gazpio, Montse Maritxalar, Aitor Gonzalez-Agirre, German Rigau, Larraitz Uria, Eneko Agirre
Knowl. Based Syst.3
2016 Why are these similar? Investigating item similarity types in a large digital library
abstract
We introduce a new problem, identifying the type of relation that holds between a pair of similar items in a digital library. Being able to provide a reason why items are similar has applications in recommendation, personalization, and search. We investigate the problem within the context of Europeana, a large digital library containing items related to cultural heritage. A range of types of similarity in this collection were identified. A set of 1,500 pairs of items from the collection were annotated using crowdsourcing. A high intertagger agreement (average 71.5 Pearson correlation) was obtained and demonstrates that the task is well defined. We also present several approaches to automatically identifying the type of similarity. The best system applies linear regression and achieves a mean Pearson correlation of 71.3, close to human performance. The problem formulation and data set described here were used in a public evaluation exercise, the *SEM shared task on Semantic Textual Similarity. The task attracted the participation of 6 teams, who submitted 14 system runs. All annotations, evaluation scripts, and system runs are freely available.
Aitor Gonzalez-Agirre, German Rigau, Eneko Agirre, Nikolaos Aletras, Mark Stevenson 0001
J. Assoc. Inf. Sci. Technol.1
2012 A Graph-Based Method to Improve WordNet Domains
Aitor Gonzalez-Agirre, German Rigau, Mauro Castillo
CICLing (1)1
2012 A proposal for improving WordNet Domains
Aitor Gonzalez-Agirre, Mauro Castillo, German Rigau
LREC1
2012 Multilingual Central Repository version 3.0
Aitor Gonzalez-Agirre, Egoitz Laparra, German Rigau
LREC1