VLDB 2026 Research / reviewers in the wild / expert
Sainandan Ramakrishnan
dblp:192/2508
· DBLP profile ↗
1ranked-venue papers
1as first author
0since 2021 · last 2018
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Vision and language · 75% Trustworthy machine learning · 25% |
Topics — the 4 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning
interpretability |
0.3 | 1 | 2018 | Overcoming Language Priors in Visual Question Answering with Adversarial Regularization · NeurIPS 2018 |
Computer vision › Vision and language › visual question answering
language prior |
0.3 | 1 | 2018 | Overcoming Language Priors in Visual Question Answering with Adversarial Regularization · NeurIPS 2018 |
Computer vision › Vision and language
visual grounding |
0.3 | 1 | 2018 | Overcoming Language Priors in Visual Question Answering with Adversarial Regularization · NeurIPS 2018 |
Computer vision › Vision and language
visual question answering |
0.3 | 1 | 2018 | Overcoming Language Priors in Visual Question Answering with Adversarial Regularization · NeurIPS 2018 |
Methods — techniques the papers use, named apart from their topics
mutual information maximization · 0.3adversarial regularization · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2018 | Overcoming Language Priors in Visual Question Answering with Adversarial RegularizationabstractModern Visual Question Answering (VQA) models have been shown to rely heavily on superficial correlations between question and answer words learned during training -- \eg overwhelmingly reporting the type of room as kitchen or the sport being played as tennis, irrespective of the image. Most alarmingly, this shortcoming is often not well reflected during evaluation because the same strong priors exist in test distributions; however, a VQA system that fails to ground questions in image content would likely perform poorly in real-world settings. In this work, we present a novel regularization scheme for VQA that reduces this effect. We introduce a question-only model that takes as input the question encoding from the VQA model and must leverage language biases in order to succeed. We then pose training as an adversarial game between the VQA model and this question-only adversary -- discouraging the VQA model from capturing language biases in its question encoding.Further, we leverage this question-only model to estimate the mutual information between the image and answer given the question, which we maximize explicitly to encourage visual grounding. Our approach is a model agnostic training procedure and simple to implement. We show empirically that it can improve performance significantly on a bias-sensitive split of the VQA dataset for multiple base models -- achieving state-of-the-art on this task. Further, on standard VQA tasks, our approach shows significantly less drop in accuracy compared to existing bias-reducing VQA models. Sainandan Ramakrishnan, Aishwarya Agrawal, Stefan Lee |
NeurIPS | 1 |