VLDB 2026 Research / reviewers in the wild / expert
Luis Chiruzzo
dblp:162/1557
· DBLP profile ↗
20ranked-venue papers
6as first author
12since 2021 · last 2026
0000-0002-1697-4614ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 6 first-author · 11 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Generating Sign Language Poses from HamNoSys and Natural Language Descriptions
Santiago Máximo, Luis Chiruzzo |
LREC | 2 |
| 2025 | Why AI Is WEIRD and Shouldn't Be This Way: Towards AI for Everyone, with Everyone, by EveryoneabstractThis paper presents a vision for creating AI systems that are inclusive at every stage of development, from data collection to model design and evaluation. We address key limitations in the current AI pipeline and its WEIRD* representation, such as lack of data diversity, biases in model performance, and narrow evaluation metrics. We also focus on the need for diverse representation among the developers of these systems, as well as incentives that are not skewed toward certain groups. We highlight opportunities to develop AI systems that are for everyone (with diverse stakeholders in mind), with everyone (inclusive of diverse data and annotators), and by everyone (designed and developed by a globally diverse workforce). *WEIRD = an acronym coined by Joseph Henrich to highlight the coverage limitations of many psychological studies, referring to populations that are Western, Educated, Industrialized, Rich, and Democratic; while we do not fully adopt this term for AI, as its current scope does not perfectly align with the WEIRD dimensions, we believe that today's AI has a similarly "weird" coverage, particularly in terms of who is involved in its development and who benefits from it. Rada Mihalcea, Oana Ignat, Longju Bai, Angana Borah, Luis Chiruzzo, Zhijing Jin 0001, Claude Kwizera, Joan Nwatu, Soujanya Poria, Thamar Solorio |
AAAI | 5 |
| 2025 | La Leaderboard: A Large Language Model Leaderboard for Spanish Varieties and Languages of Spain and Latin AmericaabstractMaría Grandury, Javier Aula-Blasco, Júlia Falcão, Clémentine Fourrier, Miguel González Saiz, Gonzalo Martínez, Gonzalo Santamaria Gomez, Rodrigo Agerri, Nuria Aldama García, Luis Chiruzzo, Javier Conde, Helena Gomez Adorno, Marta Guerrero Nieto, Guido Ivetta, Natàlia López Fuertes, Flor Miriam Plaza-del-Arco, María-Teresa Martín-Valdivia, Helena Montoro Zamorano, Carmen Muñoz Sanz, Pedro Reviriego, Leire Rosado Plaza, Alejandro Vaca Serrano, Estrella Vallecillo-Rodríguez, Jorge Vallego, Irune Zubiaga. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. María Grandury, Javier Aula-Blasco, Júlia Falcão, Clémentine Fourrier, Miguel González Saiz, Gonzalo Martínez 0001, Gonzalo Santamaría Gómez, Rodrigo Agerri, Nuria Aldama-García, Luis Chiruzzo, Javier Conde, Helena Gómez-Adorno, Marta Guerrero Nieto, Guido Ivetta, Natàlia Fuertes, Flor Miriam Plaza del Arco, María Teresa Martín Valdivia, Helena Montoro Zamorano, Carmen Muñoz Sanz, Pedro Reviriego, Leire Rosado Plaza, Alejandro Vaca Serrano, María Estrella Vallecillo Rodríguez, Jorge Vallego, Irune Zubiaga |
ACL (1) | 10 |
| 2024 | Null Subjects in Spanish as a Machine Translation ProblemabstractIn this study we approach the detection of null subjects and impersonal constructions in Spanish using a machine translation methodology. We repurpose the Spanish AnCora corpus, converting it to a parallel set that transforms Spanish sentences into a format that allows us to detect and classify verbs, and train LSTM-based neural machine translation systems to perform this task. Various models differing on output format and hyperparameters were evaluated. Experimental results proved this approach to be highly resource-effective, obtaining results comparable to or surpassing the state of the art found in existing literature, while employing modest computational resources. Additionally, an improved dataset for training and evaluating Spanish null-subject detection tools was elaborated for this project, that could aid in the creation and serve as a benchmark for further developments in the area. Jose Diego Suarez, Luis Chiruzzo |
LREC/COLING | 2 |
| 2024 | Bootstrapping Pre-trained Word Embedding Models for Sign Language Gloss TranslationabstractThis paper explores a novel method to modify existing pre-trained word embedding models of spoken languages for Sign Language glosses. These newly-generated embeddings are described, visualised, and then used in the encoder and/or decoder of models for the Text2Gloss and Gloss2Text task of machine translation. In two translation settings (one including data augmentation-based pre-training and a baseline), we find that bootstrapped word embeddings for glosses improve translation across four Signed/spoken language pairs. Many improvements are statistically significant, including those where the bootstrapped gloss embedding models are used.Languages included: American Sign Language, Finnish Sign Language, Spanish Sign Language, Sign Language of The Netherlands. Euan McGill, Luis Chiruzzo, Horacio Saggion |
EAMT (1) | 2 |
| 2024 | PAYADOR: A Minimalist Approach to Grounding Language Models on Structured Data for Interactive Storytelling and Role-playing Games
Santiago Góngora, Luis Chiruzzo, Gonzalo Méndez 0001, Pablo Gervás |
ICCC | 2 |
| 2024 | Grammar-based Data Augmentation for Low-Resource Languages: The Case of Guarani-Spanish Neural Machine TranslationabstractAgustín Lucas, Alexis Baladón, Victoria Pardiñas, Marvin Agüero-Torales, Santiago Góngora, Luis Chiruzzo. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Agustín Lucas, Alexis Baladón, Victoria Pardiñas, Marvin M. Agüero-Torales, Santiago Góngora, Luis Chiruzzo |
NAACL-HLT | 6 |
| 2023 | Meeting the Needs of Low-Resource Languages: The Value of Automatic Alignments via Pretrained ModelsabstractAbteen Ebrahimi, Arya D. McCarthy, Arturo Oncevay, John E. Ortega, Luis Chiruzzo, Gustavo Giménez-Lugo, Rolando Coto-Solano, Katharina Kann. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023. Abteen Ebrahimi, Arya McCarthy, Arturo Oncevay, John E. Ortega, Luis Chiruzzo, Gustavo Giménez Lugo, Rolando Coto-Solano, Katharina Kann |
EACL | 5 |
| 2023 | Initial Experiments for Building a Guarani WordNetabstractThis paper presents a work in progress about creating a Guarani version of the WordNet database.Guarani is an indigenous South American language and is a low-resource language from the NLP perspective.Following the expand approach, we aim to find Guarani lemmas that correspond to the concepts defined in WordNet.We do this through three strategies that try to select the correct lemmas from Guarani-Spanish datasets.We ran them through three different bilingual dictionaries and had native speakers assess the results.This procedure found Guarani lemmas for about 6.5 thousand synsets, including 27% of the base WordNet concepts.However, more work on the quality of the selected words will be needed in order to create a final version of the dataset. Luis Chiruzzo, Marvin M. Agüero-Torales, Aldo Alvarez 0001, Yliana Rodríguez |
GWC | 1 |
| 2022 | AmericasNLI: Evaluating Zero-shot Natural Language Understanding of Pretrained Multilingual Models in Truly Low-resource LanguagesabstractAbteen Ebrahimi, Manuel Mager, Arturo Oncevay, Vishrav Chaudhary, Luis Chiruzzo, Angela Fan, John Ortega, Ricardo Ramos, Annette Rios, Ivan Vladimir Meza Ruiz, Gustavo Giménez-Lugo, Elisabeth Mager, Graham Neubig, Alexis Palmer, Rolando Coto-Solano, Thang Vu, Katharina Kann. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Abteen Ebrahimi, Manuel Mager, Arturo Oncevay, Vishrav Chaudhary, Luis Chiruzzo, Angela Fan, John E. Ortega, Ricardo Ramos, Annette Rios, Iván V. Meza, Gustavo Giménez Lugo, Elisabeth Mager, Graham Neubig, Alexis Palmer, Rolando Coto-Solano, Ngoc Thang Vu, Katharina Kann |
ACL (1) | 5 |
| 2022 | Jojajovai: A Parallel Guarani-Spanish Corpus for MT BenchmarkingabstractThis work presents a parallel corpus of Guarani-Spanish text aligned at sentence level. The corpus contains about 30,000 sentence pairs, and is structured as a collection of subsets from different sources, further split into training, development and test sets. A sample of sentences from the test set was manually annotated by native speakers in order to incorporate meta-linguistic annotations about the Guarani dialects present in the corpus and also the correctness of the alignment and translation. We also present some baseline MT experiments and analyze the results in terms of the subsets. We hope this corpus can be used as a benchmark for testing Guarani-Spanish MT systems, and aim to expand and improve the quality of the corpus in future iterations. Luis Chiruzzo, Santiago Góngora, Aldo Alvarez 0001, Gustavo Giménez Lugo, Marvin M. Agüero-Torales, Yliana Rodríguez |
LREC | 1 |
| 2022 | Linguistically Enhanced Text to Sign Gloss Machine Translation
Santiago Egea Gómez, Luis Chiruzzo, Euan McGill, Horacio Saggion |
NLDB | 2 |
| 2020 | A Multi-level Annotated Corpus of Scientific Papers for Scientific Document Summarization and Cross-document Relation DiscoveryabstractRelated work sections or literature reviews are an essential part of every scientific article being crucial for paper reviewing and assessment. The automatic generation of related work sections can be considered an instance of the multi-document summarization problem. In order to allow the study of this specific problem, we have developed a manually annotated, machine readable data-set of related work sections, cited papers (e.g. references) and sentences, together with an additional layer of papers citing the references. We additionally present experiments on the identification of cited sentences, using as input citation contexts. The corpus alongside the gold standard are made available for use by the scientific community. Ahmed AbuRa'ed, Horacio Saggion, Luis Chiruzzo |
LREC | 3 |
| 2020 | Development of a Guarani - Spanish Parallel CorpusabstractThis paper presents the development of a Guarani - Spanish parallel corpus with sentence-level alignment. The Guarani sentences of the corpus use the Jopara Guarani dialect, the dialect of Guarani spoken in Paraguay, which is based on Guarani grammar and may include several Spanish loanwords or neologisms. The corpus has around 14,500 sentence pairs aligned using a semi-automatic process, containing 228,000 Guarani tokens and 336,000 Spanish tokens extracted from web sources. Luis Chiruzzo, Pedro J. Amarilla, Adolfo A. Rios, Gustavo Giménez Lugo |
LREC | 1 |
| 2020 | HAHA 2019 Dataset: A Corpus for Humor Analysis in SpanishabstractThis paper presents the development of a corpus of 30,000 Spanish tweets that were crowd-annotated with humor value and funniness score. The corpus contains approximately 38.6% of humorous tweets with an average score of 2.04 in a scale from 1 to 5 for the humorous tweets. The corpus has been used in an automatic humor recognition and analysis competition, obtaining encouraging results from the participants. Luis Chiruzzo, Santiago Castro, Aiala Rosá |
LREC | 1 |
| 2019 | Building a supertagger for Spanish HPSG
Luis Chiruzzo, Dina Wonsever |
Comput. Speech Lang. | 1 |
| 2018 | Spanish HPSG Treebank based on the AnCora Corpus
Luis Chiruzzo, Dina Wonsever |
LREC | 1 |
| 2018 | Using Context to Improve the Spanish WordNet TranslationabstractWe present some strategies for improving the Spanish version of WordNet, part of the MCR, selecting new lemmas for the Spanish synsets by translating the lemmas of the corresponding English synsets.We used four simple selectors that resulted in a considerable improvement of the Spanish WordNet coverage, but with relatively lower precision, then we defined two context based selectors that improved the precision of the translations. Alfonso Methol, Guillermo López, Juan Álvarez, Luis Chiruzzo, Dina Wonsever |
GWC | 4 |
| 2016 | Automatic definition extraction and crosswords generation from news textabstractThis paper describes the design and implementation of a system that takes Spanish text and generates crosswords (board and definitions) in a fully automatic way using the definitions extracted from those texts. Our solution divides the problem in two parts: a definition extraction module that applies pattern matching implemented in Python, and a crossword generation module that uses a greedy strategy implemented in Prolog. The system achieves 73% precision and builds crosswords similar to those built by humans. Jennifer Esteche, Romina Romero, Luis Chiruzzo, Aiala Rosá |
CLEI | 3 |
| 2016 | Some strategies for the improvement of a Spanish WordNetabstractAlthough there are currently several versions of Princeton WordNet for different languages, the lack of development of some of these versions does not make it possible to use them in different Natural Language Processing applications.So is the case of the Spanish Wordnet contained in the Multilingual Central Repository (MCR), which we tried unsuccessfully to incorporate into an anaphora resolution application and also in search terms expansion.In this situation, different strategies to improve MCR Spanish Word-Net coverage were put forward and tested, obtaining encouraging results.A specific process was conducted to increase the number of adverbs, and a few simple processes were applied which made it possible to increase, at a very low cost, the number of terms in the Spanish WordNet.Finally, a more complex method based on distributional semantics was proposed, using the relations between English Wordnet synsets, also returning positive results. Matías Herrera, Luis Chiruzzo, Dina Wonsever |
GWC | 3 |