Marta Villegas

dblp:52/4795 · DBLP profile ↗
← Back
25ranked-venue papers
8as first author
11since 2021 · last 2026
0000-0003-0711-0029ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 23 · 7 first-author · 11 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author
YearPublicationVenuePosition
2026 ACAData: Parallel Dataset of Academic Data for Machine Translation
abstract
We present ACADATA, a high-quality parallel dataset for academic translation, that consists of two subsets: ACAD-TRAIN, which contains approximately 1.5 million author-generated paragraph pairs across 96 language directions and ACAD-BENCH, a curated evaluation set of almost 6,000 translations covering 12 directions. To validate its utility, we fine-tune two Large Language Models (LLMs) on ACAD-TRAIN and benchmark them on ACAD-BENCH against specialized machine-translation systems, general-purpose, open-weight LLMs, and several large-scale proprietary models. Experimental results demonstrate that fine-tuning on ACAD-TRAIN leads to improvements in academic translation quality by +6.1 and +12.4 d-BLEU points on average for 7B and 2B models respectively, while also improving long-context translation in a general domain by up to 24.9% when translating out of English. The fine-tuned top-performing model surpasses the best propietary and open-weight models on academic translation domain. By releasing ACAD-TRAIN, ACAD-BENCH and the fine-tuned models, we provide the community with a valuable resource to advance research in academic domain and long-context translation.
Iñaki Lacunza, Javier García Gilabert, Francesca de Luca Fornaciari, Javier Aula-Blasco, Aitor Gonzalez-Agirre, Maite Melero, Marta Villegas
LREC7
2025 XDoGE: Multilingual Data Reweighting to Enhance Language Inclusivity in LLMs
Iñaki Lacunza, José Javier Saiz, Alexander Shvets, Aitor Gonzalez-Agirre, Marta Villegas
IEEE Big Data5
2025 VeritasQA: A Truthfulness Benchmark Aimed at Multilingual Transferability
abstract
As Large Language Models (LLMs) become available in a wider range of domains and applications, evaluating the truthfulness of multilingual LLMs is an issue of increasing relevance. TruthfulQA (Lin et al., 2022) is one of few benchmarks designed to evaluate how models imitate widespread falsehoods. However, it is strongly English-centric and starting to become outdated. We present VeritasQA, a context- and time-independent truthfulness benchmark built with multilingual transferability in mind, and available in Spanish, Catalan, Galician and English. VeritasQA comprises a set of 353 questions and answers inspired by common misconceptions and falsehoods that are not tied to any particular country or recent event. We release VeritasQA under an open license and present the evaluation results of 15 models of various architectures and sizes.
Javier Aula-Blasco, Júlia Falcão, Susana Sotelo, Silvia Paniagua Suárez, Aitor Gonzalez-Agirre, Marta Villegas
COLING6
2025 IberoBench: A Benchmark for LLM Evaluation in Iberian Languages
abstract
The current best practice to measure the performance of base Large Language Models is to establish a multi-task benchmark that covers a range of capabilities of interest. Currently, however, such benchmarks are only available in a few high-resource languages. To address this situation, we present IberoBench, a multilingual, multi-task benchmark for Iberian languages (i.e., Basque, Catalan, Galician, European Spanish and European Portuguese) built on the LM Evaluation Harness framework. The benchmark consists of 62 tasks divided into 179 subtasks. We evaluate 33 existing LLMs on IberoBench on 0- and 5-shot settings. We also explore the issues we encounter when working with the Harness and our approach to solving them to ensure high-quality evaluation.
Irene Baucells de la Peña, Javier Aula-Blasco, Iria de-Dios-Flores, Silvia Paniagua Suárez, Naiara Pérez, Anna Salles, Susana Sotelo Docío, Júlia Falcão, José Javier Saiz, Robiert Sepúlveda-Torres, Jeremy Barnes 0001, Pablo Gamallo 0001, Aitor Gonzalez-Agirre, German Rigau, Marta Villegas
COLING15
2025 Multi-LMentry: Can Multilingual LLMs Solve Elementary Tasks Across Languages?
abstract
Luca Moroni, Javier Aula-Blasco, Simone Conia, Irene Baucells, Naiara Perez, Silvia Paniagua Suárez, Anna Sallés, Malte Ostendorff, Júlia Falcão, Guijin Son, Aitor Gonzalez-Agirre, Roberto Navigli, Marta Villegas. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Luca Moroni, Javier Aula-Blasco, Simone Conia, Irene Baucells de la Peña, Naiara Pérez, Silvia Paniagua Suárez, Anna Salles, Malte Ostendorff, Júlia Falcão, Guijin Son, Aitor Gonzalez-Agirre, Roberto Navigli, Marta Villegas
EMNLP13
2024 Becoming a High-Resource Language in Speech: The Catalan Case in the Common Voice Corpus
abstract
Collecting voice resources for speech recognition systems is a multifaceted challenge, involving legal, technical, and diversity considerations. However, it is crucial to ensure fair access to voice-driven technology across diverse linguistic backgrounds. We describe an ongoing effort to create an extensive, high-quality, publicly available voice dataset for future development of speech technologies in Catalan through the Mozilla Common Voice crowd-sourcing platform. We detail the specific approaches used to address the challenges faced in recruiting contributors and managing the collection, validation, and recording of sentences. This detailed overview can serve as a source of guidance for similar initiatives across other projects and linguistic contexts. The success of this project is evident in the latest corpus release, version 16.1, where Catalan ranks as the most prominent language in the corpus, both in terms of recorded hours and when considering validated hours. This establishes Catalan as a language with significant speech resources for language technology development and significantly raises its international visibility.
Carme Armentano-Oller, Montserrat Marimon, Marta Villegas
LREC/COLING3
2024 FLOR: On the Effectiveness of Language Adaptation
abstract
Large language models have amply proven their great capabilities, both in downstream tasks and real-life settings. However, low- and mid-resource languages do not have access to the necessary means to train such models from scratch, and often have to rely on multilingual models despite being underrepresented in the training data. For the particular case of the Catalan language, we prove that continued pre-training with vocabulary adaptation is a better alternative to take the most out of already pre-trained models, even if these have not seen any Catalan data during their pre-training phase. We curate a 26B tokens corpus and use it to further pre-train BLOOM, giving rise to the FLOR models. We perform an extensive evaluation to assess the effectiveness of our method, obtaining consistent gains across Catalan and Spanish tasks. The models, training data, and evaluation framework are made freely available under permissive licenses.
Severino Da Dalt, Joan Llop-Palao, Irene Baucells de la Peña, Marc Pàmies, Yishi Xu, Aitor Gonzalez-Agirre, Marta Villegas
LREC/COLING7
2024 Building a Data Infrastructure for a Mid-Resource Language: The Case of Catalan
abstract
Current LLM-based applications are becoming steadily available for everyone with a reliable access to technology and the internet. These applications offer benefits to their users that leave those without access to them at a serious disadvantage. Given the vastly large amount of data needed to train LLMs, the gap between languages with access to such quantity of data and those without it is currently larger than ever. Aimed at saving this gap, the Aina Project was created to provide Catalan with the necessary resources to keep being relevant in the context of AI/NLP applications based on LLMs. We thus present a set of strategies to consider when improving technology support for a mid- or low-resource language, specially addressing sustainability of high-quality data acquisition and the challenges involved in the process. We also introduce a large amount of new annotated data for Catalan. Our hope is that those interested in replicating this work for another language can learn from what worked for us, the challenges that we faced, and the sometimes disheartening truth of working with mid- and low-resource languages.
Aitor Gonzalez-Agirre, Montserrat Marimon, Carlos Rodríguez Penagos, Javier Aula-Blasco, Irene Baucells de la Peña, Carme Armentano-Oller, Jorge Palomar-Giner, Baybars Külebi, Marta Villegas
LREC/COLING9
2024 A CURATEd CATalog: Rethinking the Extraction of Pretraining Corpora for Mid-Resourced Languages
abstract
We present and describe two language resources in this paper: CATalog 1.0, the largest text corpus in Catalan to date, and CURATE (Corpus Utility for RAting TExt), a modular, parallelizable pipeline used for processing and scoring documents based on text quality that we have optimised to run in High Performance Cluster (HPC) environments. In the coming sections we describe our data preprocessing pipeline at length; traditional pipelines usually implement a set of binary filters such that a given document is either in or out. In our experience with Catalan, in lower-resource settings it is more practical to instead assign a document a soft score to allow for more flexible decision-making. We describe how the document score is calculated and highlight its interpretability by showing that it is significantly correlated with human judgements as obtained from a comparative judgement experiment. We additionally describe the different subcorpora that make up CATalog 1.0.
Jorge Palomar-Giner, José Javier Saiz, Ferran Espuña, Mario Mina, Severino Da Dalt, Joan Llop-Palao, Malte Ostendorff, Pedro Ortiz Suarez, Georg Rehm, Aitor Gonzalez-Agirre, Marta Villegas
LREC/COLING11
2023 A weakly supervised textual entailment approach to zero-shot text classification
abstract
Marc Pàmies, Joan Llop, Francesco Multari, Nicolau Duran-Silva, César Parra-Rojas, Aitor Gonzalez-Agirre, Francesco Alessandro Massucci, Marta Villegas. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023.
Marc Pàmies, Joan Llop-Palao, Francesco Multari, Nicolau Duran-Silva, César Parra-Rojas, Aitor Gonzalez-Agirre, Francesco Alessandro Massucci, Marta Villegas
EACL8
2023 MarIA: Spanish Language Models
Marta Villegas
ICAART1
2020 BioASQ at CLEF2020: Large-Scale Biomedical Semantic Indexing and Question Answering
Martin Krallinger, Anastasia Krithara, Anastasios Nentidis, Georgios Paliouras, Marta Villegas
ECIR (2)5
2016 Leveraging RDF Graphs for Crossing Multiple Bilingual Dictionaries
Marta Villegas, Maite Melero, Núria Bel, Jorge Gracia
LREC1
2014 Metadata as Linked Open Data: mapping disparate XML metadata registries into one RDF/OWL registry
Marta Villegas, Maite Melero, Núria Bel
LREC1
2012 The IULA Treebank
Montserrat Marimon, Beatríz Fisas, Núria Bel, Jorge Vivaldi, Sergi Torner, Mercè Lorente, Silvia Vázquez, Marta Villegas
LREC8
2012 Using Language Resources in Humanities research
Marta Villegas, Núria Bel, Carlos Gonzalo, Amparo Moreno, Nuria Simelio
LREC1
2010 A Case Study on Interoperability for Language Resources and Applications
Marta Villegas, Núria Bel, Santiago Bel, Víctor Rodríguez-Doncel
LREC1
2009 Integrating Full-Text Search and Linguistic Analyses on Disperse Data for Humanities and Social Sciences Research Projects
abstract
The research reported in this paper is part of the activities carried out within the CLARIN (Common Language Resources and Technology Infrastructure) project, a large-scale pan-European project to create, coordinate and make Language Resources and Technologies (LRT) available and readily useable. CLARIN is devoted to the creation of a persistent and stable infrastructure serving the needs of the European Humanities and Social Sciences (HSS) research community. HSS researchers will be able to efficiently access distributed resources and apply analysis and exploitation tools relevant for their research. Hereby we present a real use case addressed as a CLARIN scenario and the implementation of a demonstrator that enables us to foresee the potential problems and contributes to the planning of the implementation phase. It deals with how to support researchers interested in harvesting and analyzing data from historical press archives. Therefore, we address the integration and interoperability of distributed and heterogeneous research data and analysis tools.
Marta Villegas, Carla Parra
eScience1
2008 COLDIC, a Lexicographic Platform for LMF compliant lexica
Núria Bel, Sergio Espeja, Montserrat Marimon, Marta Villegas
LREC4
2004 Cost-effective Cross-lingual Document Classification
Núria Bel, Cornelis H. A. Koster, Marta Villegas
LREC3
2002 From Resources to Applications. Designing the Multilingual ISLE Lexical Entry
Sue Atkins, Núria Bel, Francesca Bertagna, Pierrette Bouillon, Nicoletta Calzolari, Christiane Fellbaum, Ralph Grishman, Alessandro Lenci, Catherine Macleod, Martha Palmer, Gregor Thurmair, Marta Villegas, Antonio Zampolli
LREC12
2002 From DTD to relational dB. An automatic generation of a lexicographical station out off ISLE guidelines
Marta Villegas, Núria Bel
LREC1
2001 The ISLE in the ocean. Transatlantic standards for multilingual lexicons (with an eye to machine translation)
abstract
The ISLE project is a continuation of the long standing EAGLES initiative, carried out under the Human Language Technology (HLT) programme in collaboration between American and European groups in the framework of the EU-US International Research Co-operation, supported by NSF and EC. In this paper we concentrate on the current position of the ISLE Computational Lexicon Working Group (CLWG), whose activities aim at defining a general schema for a multilingual lexical entry (MILE), as the basis for a standard framework for multilingual computational lexicons. The needs and features of existing Machine Translation systems provide the main reference points for the process of consensual definition of the MILE. The overall structure of the MILE will be illustrated with particular attention to some of the issues raised for multilingual lexicons by the need of expressing complex transfer conditions among translation equivalents
Nicoletta Calzolari, Alessandro Lenci, Antonio Zampolli, Núria Bel, Marta Villegas, Gregor Thurmair
MTSummit5
2000 SIMPLE: A General Framework for the Development of Multilingual Lexicons
Núria Bel, Federica Busa, Nicoletta Calzolari, Elisabetta Gola, Alessandro Lenci, Monica Monachini, Antoine Ogonowski, Ivonne Peters, Wim Peters, Nilda Ruimy, Marta Villegas, Antonio Zampolli
LREC11
2000 Multilingual Linguistic Resources: From Monolingual Lexicons to Bilingual Interrelated Lexicons
Marta Villegas, Núria Bel, Alessandro Lenci, Nicoletta Calzolari, Nilda Ruimy, Antonio Zampolli, Teresa Sadurní, Joan Soler
LREC1