VLDB 2026 Research / reviewers in the wild / expert
Álvaro Rodrigo
dblp:57/1853 · also Álvaro Rodrigo-Yuste
· DBLP profile ↗
21ranked-venue papers
5as first author
7since 2021 · last 2026
0000-0002-6331-4117ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 3 first-author · 4 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Theory of computation · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TalentCLEF at CLEF2026: Skill and Job Title Intelligence for Human Capital Management
Luis Gascó, Hermenegildo Fabregat, Laura García-Sardiña, Paula Estrella, Casimiro Pio Carrino, Daniel Deniz, Álvaro Rodrigo, Rabih Zbib |
ECIR (4) | 7 |
| 2025 | TalentCLEF at CLEF2025: Skill and Job Title Intelligence for Human Capital Management
Luis Gascó, Hermenegildo Fabregat, Laura García-Sardiña, Daniel Deniz, Álvaro Rodrigo, Paula Estrella, Rabih Zbib |
ECIR (5) | 5 |
| 2025 | ViClaim: A Multilingual Multilabel Dataset for Automatic Claim Detection in VideosabstractThe growing influence of video content as a medium for communication and misinformation underscores the urgent need for effective tools to analyze claims in multilingual and multi-topic settings.Existing efforts in misinformation detection largely focus on written text, leaving a significant gap in addressing the complexity of spoken text in video transcripts.We introduce ViClaim, a dataset of 1,798 annotated video transcripts across three languages (English, German, Spanish) and six topics.Each sentence in the transcripts is labeled with three claim-related categories: factcheck-worthy, fact-non-check-worthy, or opinion.We developed a custom annotation tool to facilitate the highly complex annotation process.Experiments with state-of-the-art multilingual language models demonstrate strong performance in cross-validation (macro F1 up to 0.896) but reveal challenges in generalization to unseen topics, particularly for distinct domains.Our findings highlight the complexity of claim detection in video transcripts.ViClaim offers a robust foundation for advancing misinformation detection in video-based communication, addressing a critical gap in multimodal analysis. Patrick Giedemann, Pius von Däniken, Jan Deriu, Álvaro Rodrigo, Anselmo Peñas, Mark Cieliebak |
EMNLP | 4 |
| 2025 | Simulating Misinformation Diffusion on Social Media Through CoNVaI: A Textual- and Agent-Based Diffusion ModelabstractMisinformation has experienced increased online diffusion, leveraging strategies, such as emotional manipulation, to influence users' opinions. Efforts are underway to develop tools to mitigate its effects, such as misinformation propagation models used to simulate the diffusion of information. There are different approaches within these models, although, they show a significant limitation by disregarding the content of the information shared, crucial to the diffusion. We consider it the central aspect of modeling information dissemination. To this end, we focus on Agent-Based Modeling due to its suitability to simulate the complex interactions and heterogeneous behaviors observed on social media. We base our approach on a state-of-the-art Agent-Based Model that we modify and extend to account for the texts of the messages shared, focusing on two aspects that influence agents' decisions: i) the novelty of the content and; ii) its diffusion and behavior over time. To determine whether this content proves informative, we conduct an empirical evaluation using social media data from Twitter. Based on our experimental results, we observe that our textual-based approach reflects information diffusion more realistically than the state of the art, reducing the error regarding real diffusion. Raquel Rodríguez-García, Roberto Centeno, Álvaro Rodrigo |
IJCAI | 3 |
| 2025 | None of the above: comparing scenarios for answerability detection in question answering systemsabstractAbstract Question Answering (QA) is often used to assess the reasoning capabilities of NLP systems. For a QA system, it is crucial to have the capability to determine answerability– whether the question can be answered with the information at hand. Previous works have studied answerability by including a fixed proportion of unanswerable questions in a collection without explaining the reasons for such proportion or the impact on systems’ results. Furthermore, they do not answer the question of whether systems learn to determine answerability. This work aims to answer that question, providing a systematic analysis of how unanswerable question ratios in training data impact QA systems. To that end, we create a series of versions of the well-known Multiple-Choice QA dataset RACE by modifying different amounts of questions to make them unanswerable, and then train and evaluate several Large Language Models on them. We show that LLMs tend to overfit the distribution of unanswerable questions encountered during training, while the ability to decide on answerability always comes at the expense of finding the answer when it exists. Our experiments also show that a proportion of unanswerable questions around 30%– as found in existing datasets– produces the most discriminating systems. We hope these findings offer useful guidelines for future dataset designers looking to address the problem of answerability. Julio Reyes-Montesinos, Álvaro Rodrigo, Anselmo Peñas |
Appl. Intell. | 2 |
| 2024 | Improving Quantification with Minimal In-Domain Annotations: Beyond Classify and CountabstractQuantification is the task of estimating the class distribution in a given collection. With the growing availability of classification models, the use of classifiers for quantification has become increasingly popular, carrying the promise of eliminating the need for manual annotation. However, the naive classify and count approach presents clear limitations, especially evident in the face of domain discrepancies. In this work, we introduce two novel quantification methods, called CPCC and BCC, which can adapt to new target datasets with a small number of annotated in-domain samples (N = 100). To explore their real-world applicability, we apply our methods to a range of quantification tasks in the realm of hateful and offensive language, where they perform markedly better than classify and count and other existing methods. Pius von Däniken, Jan Deriu, Álvaro Rodrigo, Mark Cieliebak |
ICWSM | 3 |
| 2022 | Study of a lifelong learning scenario for question answeringabstractQuestion Answering (QA) systems have witnessed a significant advance in the last years due to the development of neural architectures employing pre-trained large models like BERT. However, once the QA model is fine-tuned for a task (e.g., a particular type of questions over a particular domain), system performance drops when new tasks are added along time, (e.g., new types of questions or new domains). Therefore, the system requires a retraining but, since the data distribution has shifted away from the previous learning, performance over previous tasks drops significantly. Hence, we need strategies to make our systems resistant to the passage of time. Lifelong Learning (LL) aims to study how systems can take advantage of the previous learning and the knowledge acquired to maintain or improve performance over time. In this article, we explore a scenario where the same LL based QA system suffers along time several shifts in the data distribution, represented as the addition of new different QA datasets. In this setup, the following research questions arise: (i) How LL based QA systems can benefit from previously learned tasks? (ii) Is there any strategy general enough to maintain or improve the performance over time when new tasks are added? and finally, (iii) How to detect a lack of knowledge that impedes the answering of questions and must trigger a new learning process? To answer these questions, we systematically try all possible training sequences over three well known QA datasets. Our results show how the learning of a new dataset is sensitive to previous training sequences and that we can find a strategy general enough to avoid the combinatorial explosion of testing all possible training sequences. Thus, when a new dataset is added to the system, the best way to retrain the system without dropping performance over the previous datasets is to randomly merge the new training material with the previous one. Guillermo Echegoyen, Álvaro Rodrigo, Anselmo Peñas |
Expert Syst. Appl. | 2 |
| 2020 | A Methodology for Creating Question Answering Corpora Using Inverse Data AnnotationabstractJan Deriu, Katsiaryna Mlynchyk, Philippe Schläpfer, Alvaro Rodrigo, Dirk von Grünigen, Nicolas Kaiser, Kurt Stockinger, Eneko Agirre, Mark Cieliebak. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020. Jan Deriu, Katsiaryna Mlynchyk, Philippe Schläpfer, Álvaro Rodrigo, Dirk Von Gruenigen, Nicolas Kaiser, Kurt Stockinger, Eneko Agirre, Mark Cieliebak |
ACL | 4 |
| 2020 | Spot The Bot: A Robust and Efficient Framework for the Evaluation of Conversational Dialogue SystemsabstractJan Deriu, Don Tuggener, Pius von Däniken, Jon Ander Campos, Alvaro Rodrigo, Thiziri Belkacem, Aitor Soroa, Eneko Agirre, Mark Cieliebak. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020. Jan Deriu, Don Tuggener, Pius von Däniken, Jon Ander Campos, Álvaro Rodrigo, Thiziri Belkacem, Aitor Soroa, Eneko Agirre, Mark Cieliebak |
EMNLP (1) | 5 |
| 2019 | The effect of answer validation on the performance of Question-Answering systems
Álvaro Rodrigo, Jesús Herrera, Anselmo Peñas |
Expert Syst. Appl. | 1 |
| 2018 | Do systems pass university entrance exams?
Álvaro Rodrigo, Anselmo Peñas, Yusuke Miyao, Noriko Kando |
Inf. Process. Manag. | 1 |
| 2017 | A study about the future evaluation of Question-Answering systems
Álvaro Rodrigo, Anselmo Peñas |
Knowl. Based Syst. | 1 |
| 2015 | On Evaluating the Contribution of Validation for Question AnsweringabstractValidation is arising as a crucial component of new architectures aimed at improving Question Answering technologies. Hence, there is a strong need of appropriate measures for evaluating Validation and, even more, its impact on question answering results. However, common Validation measures do not allow a clear study of this impact, and they might even lead researchers to obtain wrong conclusions. We propose a new approach for evaluating Validation technologies, which offers clear and useful information about the impact on QA performance. We compare our proposal with classic evaluation measures, showing the benefits of our scheme. Álvaro Rodrigo, Anselmo Peñas |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2013 | Answering questions about European legislation
Álvaro Rodrigo, Joaquín Pérez-Iglesias, Anselmo Peñas, Guillermo Garrido, Lourdes Araujo |
Expert Syst. Appl. | 1 |
| 2012 | Temporally Anchored Relation Extraction
Guillermo Garrido, Anselmo Peñas, Bernardo Cabaleiro, Álvaro Rodrigo |
ACL (1) | 4 |
| 2012 | Evaluating Machine Reading Systems through Comprehension Tests
Anselmo Peñas, Eduard H. Hovy, Pamela Forner, Álvaro Rodrigo, Richard F. E. Sutcliffe, Corina Forascu, Caroline Sporleder |
LREC | 4 |
| 2011 | A Simple Measure to Assess Non-response
Anselmo Peñas, Álvaro Rodrigo |
ACL | 2 |
| 2010 | Evaluating Multilingual Question Answering Systems at CLEF
Pamela Forner, Danilo Giampiccolo, Bernardo Magnini, Anselmo Peñas, Álvaro Rodrigo, Richard F. E. Sutcliffe |
LREC | 5 |
| 2010 | GikiCLEF: Crosscultural Issues in Multilingual Information Access
Diana Santos, Luís Miguel Cabral, Corina Forascu, Pamela Forner, Fredric C. Gey, Katrin Lamm, Thomas Mandl 0001, Petya Osenova, Anselmo Peñas, Álvaro Rodrigo, Julia Maria Struß, Yvonne Skalban, Erik F. Tjong Kim Sang |
LREC | 10 |
| 2008 | Testing the Reasoning for Question Answering ValidationabstractQuestion answering (QA) is a task that deserves more collaboration between natural language processing (NLP) and knowledge representation (KR) communities, not only to introduce reasoning when looking for answers or making use of answer type taxonomies and encyclopaedic knowledge, but also, as discussed here, for answer validation (AV), that is to say, to decide whether the responses of a QA system are correct or not. This was one of the motivations for the first Answer Validation Exercise at CLEF 2006 (AVE 2006). The starting point for the AVE 2006 was the reformulation of the answer validation as a recognizing textual entailment (RTE) problem, under the assumption that a hypothesis can be automatically generated instantiating a hypothesis pattern with a QA system answer. The test collections that we developed in seven different languages at AVE 2006 are specially oriented to the development and evaluation of answer validation systems. We show in this article the methodology followed for developing these collections taking advantage of the human assessments already made in the evaluation of QA systems. We also propose an evaluation framework for AV linked to a QA evaluation track. We quantify and discuss the source of errors introduced by the reformulation of the answer validation problem in terms of textual entailment (around 2%, in the range of inter-annotator disagreement). We also show the evaluation results of the first answer validation exercise at CLEF 2006 where 11 groups have participated with 38 runs in seven different languages. The most extensively used techniques were Machine Learning and overlapping measures, but systems with broader knowledge resources and richer representation formalisms obtained the best results. Anselmo Peñas, Álvaro Rodrigo, Valentín Sama Rojo, M. Felisa Verdejo |
J. Log. Comput. | 2 |
| 2006 | SPARTE, a Test Suite for Recognising Textual Entailment in Spanish
Anselmo Peñas, Álvaro Rodrigo, M. Felisa Verdejo |
CICLing | 2 |