EDBT 2026 Demo / reviewers in the wild / expert
Pedro Rodríguez 0001
dblp:67/4105-1 · also Pedro Rodriguez 0001
· DBLP profile ↗
11ranked-venue papers
3as first author
7since 2021 · last 2025
0000-0001-8572-0725ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Theory of computation · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
8 papers |
Language models and text generation · 64% Question answering and dialogue systems · 22% Trustworthy machine learning · 14% | |
| Databases, data mining, and information retrieval
2 papers |
Information retrieval · 67% Web and social media mining · 33% | |
| Computer graphics and multimedia
1 paper |
Virtual and augmented reality · 77% Multimedia analysis and retrieval · 23% | |
| Software engineering, system software, and programming languages
1 paper |
Empirical software engineering · 100% |
Topics — the 19 heaviest of 23, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
instruction tuning |
1.5 | 2 | 2024 | RA-DIT: Retrieval-Augmented Dual Instruction Tuning · ICLR 2024 Instruction-tuned Language Models are Better Knowledge Learners · ACL (1) 2024 |
Natural language and speech › Language models and text generation › language modeling
byte-level language model |
0.9 | 1 | 2025 | Byte Latent Transformer: Patches Scale Better Than Tokens · ACL (1) 2025 |
Natural language and speech › Language models and text generation › language modeling
language model architecture |
0.9 | 1 | 2025 | Byte Latent Transformer: Patches Scale Better Than Tokens · ACL (1) 2025 |
Natural language and speech › Language models and text generation
tokenization |
0.9 | 1 | 2025 | Byte Latent Transformer: Patches Scale Better Than Tokens · ACL (1) 2025 |
Natural language and speech › Language models and text generation
knowledge learning |
0.8 | 1 | 2024 | Instruction-tuned Language Models are Better Knowledge Learners · ACL (1) 2024 |
Natural language and speech › Language models and text generation
retrieval-augmented language models |
0.8 | 1 | 2024 | RA-DIT: Retrieval-Augmented Dual Instruction Tuning · ICLR 2024 |
Natural language and speech › Question answering and dialogue systems › task-oriented dialogue
multimodal task-oriented dialogue |
0.7 | 1 | 2023 | SIMMC-VR: A Task-oriented Multimodal Dialog Dataset with Situated and Immersive VR Streams · ACL (1) 2023 |
Natural language and speech › Question answering and dialogue systems
question answering evaluation |
0.5 | 1 | 2021 | Evaluation Paradigms in Question Answering · EMNLP (1) 2021 |
Natural language and speech › Question answering and dialogue systems › interactive question answering
conversational question answering |
0.4 | 1 | 2020 | Information Seeking in the Spirit of Learning: A Dataset for Conversational Curiosity · EMNLP (1) 2020 |
Natural language and speech › Question answering and dialogue systems › task-oriented dialogue
information-seeking dialogue |
0.4 | 1 | 2020 | Information Seeking in the Spirit of Learning: A Dataset for Conversational Curiosity · EMNLP (1) 2020 |
Web and social media mining
user engagement |
0.4 | 1 | 2020 | Information Seeking in the Spirit of Learning: A Dataset for Conversational Curiosity · EMNLP (1) 2020 |
Machine learning › Trustworthy machine learning › robustness
adversarial examples |
0.3 | 1 | 2018 | Pathologies of Neural Models Make Interpretation Difficult · EMNLP 2018 |
Machine learning › Trustworthy machine learning › calibration
confidence calibration |
0.3 | 1 | 2018 | Pathologies of Neural Models Make Interpretation Difficult · EMNLP 2018 |
Machine learning › Trustworthy machine learning
interpretability |
0.3 | 1 | 2018 | Pathologies of Neural Models Make Interpretation Difficult · EMNLP 2018 |
Machine learning › Trustworthy machine learning
robustness |
0.3 | 1 | 2018 | Pathologies of Neural Models Make Interpretation Difficult · EMNLP 2018 |
Natural language and speech › Language models and text generation › language modeling
long-context language modeling |
0.2 | 1 | 2024 | Instruction-tuned Language Models are Better Knowledge Learners · ACL (1) 2024 |
Natural language and speech › Language models and text generation › large language model evaluation
NLP evaluation |
0.1 | 1 | 2021 | Evaluation Examples are not Equally Informative: How should that change NLP Leaderboards? · ACL/IJCNLP (1) 2021 |
Computational social science and digital humanities › science and technology studies
history of computing |
0.1 | 1 | 2021 | Evaluation Paradigms in Question Answering · EMNLP (1) 2021 |
Information retrieval
evaluation |
0.1 | 1 | 2020 | Information Seeking in the Spirit of Learning: A Dataset for Conversational Curiosity · EMNLP (1) 2020 |
Methods — techniques the papers use, named apart from their topics
instruction tuning · 2.3retrieval augmentation · 1.5large language model · 1.5dataset construction · 1.3position paper · 1.0latent transformer · 0.9wizard-of-oz · 0.9multi-task model · 0.9BERT · 0.9gradient-based attribution · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Byte Latent Transformer: Patches Scale Better Than TokensabstractArtidoro Pagnoni, Ramakanth Pasunuru, Pedro Rodriguez, John Nguyen, Benjamin Muller, Margaret Li, Chunting Zhou, Lili Yu, Jason E Weston, Luke Zettlemoyer, Gargi Ghosh, Mike Lewis, Ari Holtzman, Srini Iyer. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Artidoro Pagnoni, Ramakanth Pasunuru, Pedro Rodríguez 0001, John Nguyen, Benjamin Muller, Margaret Li, Chunting Zhou, Lili Yu, Jason Weston, Luke Zettlemoyer, Gargi Ghosh, Mike Lewis, Ari Holtzman, Srinivasan Iyer 0001 |
ACL (1) | 3 |
| 2024 | Instruction-tuned Language Models are Better Knowledge LearnersabstractZhengbao Jiang, Zhiqing Sun, Weijia Shi, Pedro Rodriguez, Chunting Zhou, Graham Neubig, Xi Lin, Wen-tau Yih, Srini Iyer. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Zhengbao Jiang, Zhiqing Sun, Pedro Rodríguez 0001, Chunting Zhou, Graham Neubig, Xi Victoria Lin, Scott Yih, Srinivasan Iyer 0001 |
ACL (1) | 4 |
| 2024 | RA-DIT: Retrieval-Augmented Dual Instruction TuningabstractRetrieval-augmented language models (RALMs) improve performance by accessing long-tail and up-to-date knowledge from external data stores, but are challenging to build. Existing approaches require either expensive retrieval-specific modifications to LM pre-training or use post-hoc integration of the data store that leads to suboptimal performance. We introduce Retrieval-Augmented Dual Instruction Tuning (RA-DIT), a lightweight fine-tuning methodology that provides a third option by retrofitting any LLM with retrieval capabilities. Our approach operates in two distinct fine-tuning steps: (1) one updates a pre-trained LM to better use retrieved information, while (2) the other updates the retriever to return more relevant results, as preferred by the LM. By fine-tuning over tasks that require both knowledge utilization and contextual awareness, we demonstrate that each stage yields significant performance improvements, and using both leads to additional gains. Our best model, RA-DIT 65B, achieves state-of-the-art performance across a range of knowledge-intensive zero- and few-shot learning benchmarks, significantly outperforming existing in-context RALM approaches by up to +8.9% in 0-shot setting and +1.4% in 5-shot setting on average. Xi Victoria Lin, Xilun Chen 0002, Mingda Chen, Maria Lomeli, Richard James 0001, Pedro Rodríguez 0001, Jacob Kahn, Gergely Szilvasy, Mike Lewis, Luke Zettlemoyer, Scott Yih |
ICLR | 7 |
| 2023 | SIMMC-VR: A Task-oriented Multimodal Dialog Dataset with Situated and Immersive VR StreamsabstractTe-Lin Wu, Satwik Kottur, Andrea Madotto, Mahmoud Azab, Pedro Rodriguez, Babak Damavandi, Nanyun Peng, Seungwhan Moon. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Te-Lin Wu, Satwik Kottur, Andrea Madotto, Mahmoud Azab, Pedro Rodríguez 0001, Babak Damavandi, Nanyun Peng 0001, Seungwhan Moon |
ACL (1) | 5 |
| 2023 | <tt>py-irt</tt>: A Scalable Item Response Theory Library for Pythonabstractpy-irt is a Python library for fitting Bayesian item response theory (IRT) models. At present, there is no Python package for fitting large-scale IRT models. py-irt estimates latent traits of subjects and items, making it appropriate for use in IRT tasks as well as in ideal point models. py-irt is built on top of the Pyro and PyTorch frameworks and uses GPU-accelerated training to scale to large data sets. It is the first Python package for large-scale IRT model fitting. py-irt is easy to use for practitioners and also allows for researchers to build and fit custom IRT models. py-irt is available as open-source software and can be installed from GitHub or the Python Package Index. History: Accepted by Ted Ralphs, Area Editor for software tools. Supplemental Material: The software that supports the findings of this study is available within the paper and its Supplementary Information [ https://pubsonline.informs.org/doi/suppl/10.1287/ijoc.2022.1250 ] or is available from the IJOC GitHub software repository ( https://github.com/INFORMSJoC ) at [ http://dx.doi.org/10.5281/zenodo.6818509 ]. John Lalor, Pedro Rodríguez 0001 |
INFORMS J. Comput. | 2 |
| 2021 | Evaluation Examples are not Equally Informative: How should that change NLP Leaderboards?abstractPedro Rodriguez, Joe Barrow, Alexander Miserlis Hoyle, John P. Lalor, Robin Jia, Jordan Boyd-Graber. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Pedro Rodríguez 0001, Joe Barrow, Alexander Miserlis Hoyle, John Lalor, Robin Jia, Jordan L. Boyd-Graber |
ACL/IJCNLP (1) | 1 |
| 2021 | Evaluation Paradigms in Question AnsweringabstractQuestion answering (QA) primarily descends from two branches of research: (1) Alan Turing's investigation of machine intelligence at Manchester University and (2) Cyril Cleverdon's comparison of library card catalog indices at Cranfield University.This position paper names and distinguishes these paradigms.Despite substantial overlap, subtle but significant distinctions exert an outsize influence on research.While one evaluation paradigm values creating more intelligent QA systems, the other paradigm values building QA systems that appeal to users.By better understanding the epistemic heritage of QA, researchers, academia, and industry can more effectively accelerate QA research. Pedro Rodríguez 0001, Jordan L. Boyd-Graber |
EMNLP (1) | 1 |
| 2020 | Information Seeking in the Spirit of Learning: A Dataset for Conversational CuriosityabstractOpen-ended human learning and information-seeking are increasingly mediated by digital assistants. However, such systems often ignore the user's pre-existing knowledge. Assuming a correlation between engagement and user responses such as "liking" messages or asking followup questions, we design a Wizard-of-Oz dialog task that tests the hypothesis that engagement increases when users are presented with facts related to what they know. Through crowd-sourcing of this experiment, we collect and release 14K dialogs (181K utterances) where users and assistants converse about geographic topics like geopolitical entities and locations. This dataset is annotated with pre-existing user knowledge, message-level dialog acts, grounding to Wikipedia, and user reactions to messages. Responses using a user's prior knowledge increase engagement. We incorporate this knowledge into a multi-task model that reproduces human assistant policies and improves over a BERT content model by 13 mean reciprocal rank points. Pedro Rodríguez 0001, Paul A. Crook, Seungwhan Moon, Zhiguang Wang |
EMNLP (1) | 1 |
| 2019 | Mitigating Noisy Inputs for Question AnsweringabstractNatural language processing systems are often downstream of unreliable inputs: machine translation, optical character recognition, or speech recognition. For instance, virtual assistants can only answer your questions after understanding your speech. We investigate and mitigate the effects of noise from Automatic Speech Recognition systems on two factoid Question Answering (QA) tasks. Integrating confidences into the model and forced decoding of unknown words are empirically shown to improve the accuracy of downstream neural QA systems. We create and train models on a synthetic corpus of over 500,000 noisy sentences and evaluate on two human corpora from Quizbowl and Jeopardy! competitions. Denis Peskov, Joe Barrow, Pedro Rodríguez 0001, Graham Neubig, Jordan L. Boyd-Graber |
INTERSPEECH | 3 |
| 2019 | Trick Me If You Can: Human-in-the-loop Generation of Adversarial Question Answering ExamplesabstractAdversarial evaluation stress-tests a model’s understanding of natural language. Because past approaches expose superficial patterns, the resulting adversarial examples are limited in complexity and diversity. We propose human- in-the-loop adversarial generation, where human authors are guided to break models. We aid the authors with interpretations of model predictions through an interactive user interface. We apply this generation framework to a question answering task called Quizbowl, where trivia enthusiasts craft adversarial questions. The resulting questions are validated via live human–computer matches: Although the questions appear ordinary to humans, they systematically stump neural and information retrieval models. The adversarial questions cover diverse phenomena from multi-hop reasoning to entity type distractors, exposing open challenges in robust question answering. Eric Wallace, Pedro Rodríguez 0001, Shi Feng 0005, Ikuya Yamada, Jordan L. Boyd-Graber |
Trans. Assoc. Comput. Linguistics | 2 |
| 2018 | Pathologies of Neural Models Make Interpretation DifficultabstractOne way to interpret neural model predictions is to highlight the most important input features-for example, a heatmap visualization over the words in an input sentence.In existing interpretation methods for NLP, a word's importance is determined by either input perturbation-measuring the decrease in model confidence when that word is removed-or by the gradient with respect to that word.To understand the limitations of these methods, we use input reduction, which iteratively removes the least important word from the input.This exposes pathological behaviors of neural models: the remaining words appear nonsensical to humans and are not the ones determined as important by interpretation methods.As we confirm with human experiments, the reduced examples lack information to support the prediction of any label, but models still make the same predictions with high confidence.To explain these counterintuitive results, we draw connections to adversarial examples and confidence calibration: pathological behaviors reveal difficulties in interpreting neural models trained with maximum likelihood.To mitigate their deficiencies, we fine-tune the models by encouraging high entropy outputs on reduced examples.Fine-tuned models become more interpretable under input reduction without accuracy loss on regular examples. Shi Feng 0005, Eric Wallace, Alvin Grissom II, Mohit Iyyer, Pedro Rodríguez 0001, Jordan L. Boyd-Graber |
EMNLP | 5 |