VLDB 2026 Research / reviewers in the wild / expert
Xabier Saralegi
dblp:34/7625
· DBLP profile ↗
21ranked-venue papers
10as first author
10since 2021 · last 2026
0000-0003-1959-4737ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 8 first-author · 9 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Are Social Biases in LLMs Consistent across Generative Tasks? A Case Study for Basque
Muitze Zulaika, Xabier Saralegi, Julia Shershneva, Lia Gonzalez, Arkaitz Fullaondo |
LREC | 2 |
| 2025 | BasqBBQ: A QA Benchmark for Assessing Social Biases in LLMs for Basque, a Low-Resource LanguageabstractThe rise of pre-trained language models has revolutionized natural language processing (NLP) tasks, but concerns about the propagation of social biases in these models remain, particularly in under-resourced languages like Basque. This paper introduces BasqBBQ, the first benchmark designed to assess social biases in Basque across eight domains, using a multiple-choice question-answering (QA) task. We evaluate various autoregressive large language models (LLMs), including multilingual and those adapted for Basque, to analyze both their accuracy and bias transmission. Our results show that while larger models generally achieve better accuracy, ambiguous cases remain challenging. In terms of bias, larger models exhibit lower negative bias. However, high negative bias persists in specific categories such as Disability Status, Age and Physical Appearance, especially in ambiguous contexts. Conversely, categories such as Sexual Orientation, Gender Identity, and Race/Ethnicity show the least bias in ambiguous contexts. The continual pre-training based adaptation process for Basque has a limited impact on bias when compared with English. This work represents a key step toward creating more ethical LLMs for low-resource languages. Xabier Saralegi, Muitze Zulaika |
COLING | 1 |
| 2025 | Pipeline Analysis for Developing Instruct LLMs in Low-Resource Languages: A Case Study on BasqueabstractAnder Corral, Ixak Sarasua Antero, Xabier Saralegi. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Ander Corral, Ixak Sarasua, Xabier Saralegi |
NAACL (Long Papers) | 3 |
| 2024 | How Well Can BERT Learn the Grammar of an Agglutinative and Flexible-Order Language? The Case of BasqueabstractThis work investigates the acquisition of formal linguistic competence by neural language models, hypothesizing that languages with complex grammar, such as Basque, present substantial challenges during the pre-training phase. Basque is distinguished by its complex morphology and flexible word order, potentially complicating grammar extraction. In our analysis, we evaluated the grammatical knowledge of BERT models trained under various pre-training configurations, considering factors such as corpus size, model size, number of epochs, and the use of lemmatization. To assess this grammatical knowledge, we constructed the BL2MP (Basque L2 student-based Minimal Pairs) test set. This test set consists of minimal pairs, each containing both a grammatically correct and an incorrect sentence, sourced from essays authored by students at different proficiency levels in the Basque language. Additionally, our analysis explores the difficulties in learning various grammatical phenomena, the challenges posed by flexible word order, and the influence of the student’s proficiency level on the difficulty of correcting grammar errors. Gorka Urbizu, Muitze Zulaika, Xabier Saralegi, Ander Corral |
LREC/COLING | 3 |
| 2024 | MULTILINGTOOL, Development of an Automatic Multilingual Subtitling and Dubbing SystemabstractIn this paper, we present the MULTILINGTOOL project, led by the Elhuyar Foundation and funded by the European Commission under the CREA-MEDIA2022-INNOVBUSMOD call. The aim of the project is to develop an advanced platform for automatic multilingual subtitling and dubbing. It will provide support for Spanish, English, and French, as well as the co-official languages of Spain, namely Basque, Catalan, and Galician. Xabier Saralegi, Ander Corral, Igor Leturia, Xabier Sarasola, Josu Murua, Iker Manterola, Itziar Cortés Etxabe |
EAMT (2) | 1 |
| 2024 | XNLIeu: a dataset for cross-lingual NLI in BasqueabstractMaite Heredia, Julen Etxaniz, Muitze Zulaika, Xabier Saralegi, Jeremy Barnes, Aitor Soroa. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Maite Heredia, Julen Etxaniz, Muitze Zulaika, Xabier Saralegi, Jeremy Barnes 0001, Aitor Soroa |
NAACL-HLT | 4 |
| 2022 | TANDO: A Corpus for Document-level Machine TranslationabstractDocument-level Neural Machine Translation aims to increase the quality of neural translation models by taking into account contextual information. Properly modelling information beyond the sentence level can result in improved machine translation output in terms of coherence, cohesion and consistency. Suitable corpora for context-level modelling are necessary to both train and evaluate context-aware systems, but are still relatively scarce. In this work we describe TANDO, a document-level corpus for the under-resourced Basque-Spanish language pair, which we share with the scientific community. The corpus is composed of parallel data from three different domains and has been prepared with context-level information. Additionally, the corpus includes contrastive test sets for fine-grained evaluations of gender and register contextual phenomena on both source and target language sides. To establish the usefulness of the corpus, we trained and evaluated baseline Transformer models and context-aware variants based on context concatenation. Our results indicate that the corpus is suitable for fine-grained evaluation of document-level machine translation systems. Harritxu Gete, Thierry Etchegoyhen, David Ponce, Gorka Labaka, Nora Aranberri, Ander Corral, Xabier Saralegi, Igor Ellakuria, Maite Martín |
LREC | 7 |
| 2022 | BasqueGLUE: A Natural Language Understanding Benchmark for BasqueabstractNatural Language Understanding (NLU) technology has improved significantly over the last few years and multitask benchmarks such as GLUE are key to evaluate this improvement in a robust and general way. These benchmarks take into account a wide and diverse set of NLU tasks that require some form of language understanding, beyond the detection of superficial, textual clues. However, they are costly to develop and language-dependent, and therefore they are only available for a small number of languages. In this paper, we present BasqueGLUE, the first NLU benchmark for Basque, a less-resourced language, which has been elaborated from previously existing datasets and following similar criteria to those used for the construction of GLUE and SuperGLUE. We also report the evaluation of two state-of-the-art language models for Basque on BasqueGLUE, thus providing a strong baseline to compare upon. BasqueGLUE is freely available under an open license. Gorka Urbizu, Iñaki San Vicente, Xabier Saralegi, Rodrigo Agerri, Aitor Soroa |
LREC | 3 |
| 2022 | Information retrieval and question answering: A case study on COVID-19 scientific literature
Arantxa Otegi, Iñaki San Vicente, Xabier Saralegi, Anselmo Peñas, Borja Lozano, Eneko Agirre |
Knowl. Based Syst. | 3 |
| 2021 | Fine-Tuning BERT for COVID-19 Domain Ad-Hoc IR by Using Pseudo-qrels
Xabier Saralegi, Iñaki San Vicente |
ECIR (2) | 1 |
| 2020 | Give your Text Representation Models some Love: the Case for BasqueabstractWord embeddings and pre-trained language models allow to build rich representations of text and have enabled improvements across most NLP tasks. Unfortunately they are very expensive to train, and many small companies and research groups tend to use models that have been pre-trained and made available by third parties, rather than building their own. This is suboptimal as, for many languages, the models have been trained on smaller (or lower quality) corpora. In addition, monolingual pre-trained models for non-English languages are not always available. At best, models for those languages are included in multilingual versions, where each language shares the quota of substrings and parameters with the rest of the languages. This is particularly true for smaller languages such as Basque. In this paper we show that a number of monolingual models (FastText word embeddings, FLAIR and BERT language models) trained with larger Basque corpora produce much better results than publicly available versions in downstream NLP tasks, including topic classification, sentiment classification, PoS tagging and NER. This work sets a new state-of-the-art in those tasks for Basque. All benchmarks and models used in this work are publicly available. Rodrigo Agerri, Iñaki San Vicente, Jon Ander Campos, Ander Barrena, Xabier Saralegi, Aitor Soroa, Eneko Agirre |
LREC | 5 |
| 2020 | Building a Task-oriented Dialog System for Languages with no Training Data: the Case for BasqueabstractThis paper presents an approach for developing a task-oriented dialog system for less-resourced languages in scenarios where training data is not available. Both intent classification and slot filling are tackled. We project the existing annotations in rich-resource languages by means of Neural Machine Translation (NMT) and posterior word alignments. We then compare training on the projected monolingual data with direct model transfer alternatives. Intent Classifiers and slot filling sequence taggers are implemented using a BiLSTM architecture or by fine-tuning BERT transformer models. Models learnt exclusively from Basque projected data provide better accuracies for slot filling. Combining Basque projected train data with rich-resource languages data outperforms consistently models trained solely on projected data for intent classification. At any rate, we achieve competitive performance in both tasks, with accuracies of 81% for intent classification and 77% for slot filling. Maddalen Lopez de Lacalle, Xabier Saralegi, Iñaki San Vicente |
LREC | 2 |
| 2016 | Evaluating Translation Quality and CLIR Performance of Query Sessions
Xabier Saralegi, Eneko Agirre, Iñaki Alegria |
LREC | 1 |
| 2016 | Polarity Lexicon Building: to what Extent Is the Manual Effort Worth?
Iñaki San Vicente, Xabier Saralegi |
LREC | 2 |
| 2013 | Analyzing the Sense Distribution of Concordances Obtained by Web as Corpus Approach
Xabier Saralegi, Pablo Gamallo 0001 |
CICLing (1) | 1 |
| 2013 | Cross-Lingual Projections vs. Corpora Extracted Subjectivity Lexicons for Less-Resourced Languages
Xabier Saralegi, Iñaki San Vicente, Irati Ugarteburu |
CICLing (2) | 1 |
| 2012 | Building a Basque-Chinese Dictionary by Using English as Pivot
Xabier Saralegi, Iker Manterola, Iñaki San Vicente |
LREC | 1 |
| 2011 | Analyzing Methods for Improving Precision of Pivot Based Bilingual Dictionaries
Xabier Saralegi, Iker Manterola, Iñaki San Vicente |
EMNLP | 1 |
| 2010 | Estimating Translation Probabilities from the Web for Structured Queries on CLIR
Xabier Saralegi, Maddalen Lopez de Lacalle |
ECIR | 1 |
| 2010 | Dictionary and Monolingual Corpus-based Query Translation for Basque-English CLIR
Xabier Saralegi, Maddalen Lopez de Lacalle |
LREC | 1 |
| 2004 | A XML-Based Term Extraction Tool for Basque
Iñaki Alegria, Antton Gurrutxaga, P. Lizaso, Xabier Saralegi, Sahats Ugartetxea, Ruben Urizar |
LREC | 4 |