VLDB 2026 Research / reviewers in the wild / expert
Dheeru Dua
dblp:194/5251
· DBLP profile ↗
8ranked-venue papers
6as first author
5since 2021 · last 2023
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 6 first-author · 5 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
Question answering and dialogue systems · 58% Transfer learning and domain adaptation · 17% Trustworthy machine learning · 8% |
Topics — the 16 heaviest of 18, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Question answering and dialogue systems
machine reading comprehension |
1.4 | 3 | 2021 | Learning with Instance Bundles for Reading Comprehension · EMNLP (1) 2021 Dynamic Sampling Strategies for Multi-Task Reading Comprehension · ACL 2020 Benefits of Intermediate Annotations in Reading Comprehension · ACL 2020 |
Machine learning › Transfer learning and domain adaptation
domain adaptation |
0.7 | 1 | 2023 | To Adapt or to Annotate: Challenges and Interventions for Domain Adaptation in Open-Domain Question Answering · ACL (1) 2023 |
Natural language and speech › Question answering and dialogue systems
open-domain question answering |
0.7 | 1 | 2023 | To Adapt or to Annotate: Challenges and Interventions for Domain Adaptation in Open-Domain Question Answering · ACL (1) 2023 |
Natural language and speech › Question answering and dialogue systems › knowledge base question answering
complex question answering |
0.6 | 1 | 2022 | Successive Prompting for Decomposing Complex Questions · EMNLP 2022 |
Natural language and speech › Question answering and dialogue systems › question understanding
question decomposition |
0.6 | 1 | 2022 | Successive Prompting for Decomposing Complex Questions · EMNLP 2022 |
Natural language and speech › Question answering and dialogue systems › reasoning-based question answering
multi-hop question answering |
0.5 | 1 | 2021 | Generative Context Pair Selection for Multi-hop Question Answering · EMNLP (1) 2021 |
Natural language and speech › Question answering and dialogue systems › machine reading comprehension
multi-hop reading comprehension |
0.5 | 1 | 2021 | Learning with Instance Bundles for Reading Comprehension · EMNLP (1) 2021 |
Computer vision › Vision and language › multimodal reasoning
compositional reasoning |
0.4 | 1 | 2020 | Benefits of Intermediate Annotations in Reading Comprehension · ACL 2020 |
Natural language and speech › Question answering and dialogue systems › machine reading comprehension
multi-task machine reading comprehension |
0.4 | 1 | 2020 | Dynamic Sampling Strategies for Multi-Task Reading Comprehension · ACL 2020 |
Machine learning › Trustworthy machine learning › robustness
adversarial examples |
0.3 | 1 | 2018 | Generating Natural Adversarial Examples · ICLR (Poster) 2018 |
Machine learning › Trustworthy machine learning › robustness
adversarial robustness |
0.3 | 1 | 2018 | Generating Natural Adversarial Examples · ICLR (Poster) 2018 |
Natural language and speech › Language models and text generation › prompting
iterative prompting |
0.2 | 1 | 2022 | Successive Prompting for Decomposing Complex Questions · EMNLP 2022 |
Natural language and speech › Language models and text generation
prompting |
0.2 | 1 | 2022 | Successive Prompting for Decomposing Complex Questions · EMNLP 2022 |
Machine learning › Representation and self-supervised learning
contrastive estimation |
0.1 | 1 | 2021 | Learning with Instance Bundles for Reading Comprehension · EMNLP (1) 2021 |
Machine learning › Probabilistic and Bayesian machine learning › sampling
adaptive sampling |
0.1 | 1 | 2020 | Dynamic Sampling Strategies for Multi-Task Reading Comprehension · ACL 2020 |
Machine learning › Learning paradigms
multi-task learning |
0.1 | 1 | 2020 | Dynamic Sampling Strategies for Multi-Task Reading Comprehension · ACL 2020 |
Methods — techniques the papers use, named apart from their topics
data augmentation · 0.7synthetic data generation · 0.6in-context learning · 0.6few-shot prompting · 0.6instance bundles · 0.5generative model · 0.5cross-entropy loss · 0.5contrastive estimation · 0.5intermediate annotation · 0.4dynamic sampling · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | To Adapt or to Annotate: Challenges and Interventions for Domain Adaptation in Open-Domain Question AnsweringabstractRecent advances in open-domain question answering (ODQA) have demonstrated impressive accuracy on general-purpose domains like Wikipedia.While some work has been investigating how well ODQA models perform when tested for out-of-domain (OOD) generalization, these studies have been conducted only under conservative shifts in data distribution and typically focus on a single component (i.e., retriever or reader) rather than an end-to-end system.This work proposes a more realistic endto-end domain shift evaluation setting covering five diverse domains.We not only find that endto-end models fail to generalize but that high retrieval scores often still yield poor answer prediction accuracy.To address these failures, we investigate several interventions, in the form of data augmentations, for improving model adaption and use our evaluation set to elucidate the relationship between the efficacy of an intervention scheme and the particular type of dataset shifts we consider.We propose a generalizability test that estimates the type of shift in a target dataset without training a model in the target domain and that the type of shift is predictive of which data augmentation schemes will be effective for domain adaption.Overall, we find that these interventions increase end-to-end performance by up to ∼24 points.* *This work was done while authors were at Google.Average F1 over all target datasets Average F1 over target datasets with specific shifts Dheeru Dua, Emma Strubell, Sameer Singh 0001, Patrick Verga |
ACL (1) | 1 |
| 2022 | Successive Prompting for Decomposing Complex QuestionsabstractAnswering complex questions that require making latent decisions is a challenging task, especially when limited supervision is available.Recent works leverage the capabilities of large language models (LMs) to perform complex question answering in a few-shot setting by demonstrating how to output intermediate rationalizations while solving the complex question in a single pass.We introduce "Successive Prompting", where we iteratively break down a complex task into a simple task, solve it, and then repeat the process until we get the final solution.Successive prompting decouples the supervision for decomposing complex questions from the supervision for answering simple questions, allowing us to (1) have multiple opportunities to query in-context examples at each reasoning step (2) learn question decomposition separately from question answering, including using synthetic data, and (3) use bespoke (fine-tuned) components for reasoning steps where a large LM does not perform well.The intermediate supervision is typically manually written, which can be expensive to collect.We introduce a way to generate a synthetic dataset which can be used to bootstrap a model's ability to decompose and answer intermediate questions.Our best model (with successive prompting) achieves an improvement of ∼5% absolute F1 on a few-shot version of the DROP dataset when compared with a stateof-the-art model with the same supervision. Dheeru Dua, Shivanshu Gupta, Sameer Singh 0001, Matt Gardner 0001 |
EMNLP | 1 |
| 2022 | Tricks for Training Sparse Translation ModelsabstractDheeru Dua, Shruti Bhosale, Vedanuj Goswami, James Cross, Mike Lewis, Angela Fan. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Dheeru Dua, Shruti Bhosale, Vedanuj Goswami, James Cross 0003, Mike Lewis, Angela Fan |
NAACL-HLT | 1 |
| 2021 | Learning with Instance Bundles for Reading ComprehensionabstractWhen training most modern reading comprehension models, all the questions associated with a context are treated as being independent from each other.However, closely related questions and their corresponding answers are not independent, and leveraging these relationships could provide a strong supervision signal to a model.Drawing on ideas from contrastive estimation, we introduce several new supervision losses that compare question-answer scores across multiple related instances.Specifically, we normalize these scores across various neighborhoods of closely contrasting questions and/or answers, adding a cross entropy loss term in addition to traditional maximum likelihood estimation.Our techniques require bundles of related question-answer pairs, which we either mine from within existing data or create using automated heuristics.We empirically demonstrate the effectiveness of training with instance bundles on two datasets-HotpotQA and ROPES-showing up to 9% absolute gains in accuracy. Dheeru Dua, Pradeep Dasigi, Sameer Singh 0001, Matt Gardner 0001 |
EMNLP (1) | 1 |
| 2021 | Generative Context Pair Selection for Multi-hop Question AnsweringabstractDheeru Dua, Cicero Nogueira dos Santos, Patrick Ng, Ben Athiwaratkun, Bing Xiang, Matt Gardner, Sameer Singh. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021. Dheeru Dua, Cícero Nogueira dos Santos, Patrick Ng, Ben Athiwaratkun, Bing Xiang, Matt Gardner 0001, Sameer Singh 0001 |
EMNLP (1) | 1 |
| 2020 | Benefits of Intermediate Annotations in Reading ComprehensionabstractComplex, compositional reading comprehension datasets require performing latent sequential decisions that are learned via supervision from the final answer.A large combinatorial space of possible decision paths that result in the same answer, compounded by the lack of intermediate supervision to help choose the right path, makes the learning particularly hard for this task.In this work, we study the benefits of collecting intermediate reasoning supervision along with the answer during data collection.We find that these intermediate annotations can provide two-fold benefits.First, we observe that for any collection budget, spending a fraction of it on intermediate annotations results in improved model performance, for two complex compositional datasets: DROP and Quoref.Second, these annotations encourage the model to learn the correct latent reasoning steps, helping combat some of the biases introduced during the data collection process. Dheeru Dua, Sameer Singh 0001, Matt Gardner 0001 |
ACL | 1 |
| 2020 | Dynamic Sampling Strategies for Multi-Task Reading ComprehensionabstractBuilding general reading comprehension systems, capable of solving multiple datasets at the same time, is a recent aspirational goal in the research community.Prior work has focused on model architectures or generalization to held out datasets, and largely passed over the particulars of the multi-task learning set up.We show that a simple dynamic sampling strategy, selecting instances for training proportional to the multi-task model's current performance on a dataset relative to its singletask performance, gives substantive gains over prior multi-task sampling strategies, mitigating the catastrophic forgetting that is common in multi-task learning.We also demonstrate that allowing instances of different tasks to be interleaved as much as possible between each epoch and batch has a clear benefit in multitask performance over forcing task homogeneity at the epoch or batch level.Our final model shows greatly increased performance over the best model on ORB, a recently-released multitask reading comprehension benchmark. Ananth Gottumukkala, Dheeru Dua, Sameer Singh 0001, Matt Gardner 0001 |
ACL | 2 |
| 2018 | Generating Natural Adversarial Examples
Zhengli Zhao, Dheeru Dua, Sameer Singh 0001 |
ICLR (Poster) | 2 |