VLDB 2026 Research / reviewers in the wild / expert
Abhishek Das 0002
dblp:40/5262-2
· DBLP profile ↗
18ranked-venue papers
6as first author
5since 2021 · last 2025
0000-0002-4718-4316ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 6 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 3 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
16 papers |
Question answering and dialogue systems · 22% Reinforcement learning · 20% Robot navigation and mapping · 16% | |
| Interdisciplinary, comprehensive, and emerging computing
3 papers |
Computational science and engineering · 92% Bioinformatics and computational biology · 8% |
Topics — the 30 heaviest of 36, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Question answering and dialogue systems
visual dialog |
1.8 | 5 | 2020 | Large-Scale Pretraining for Visual Dialog: A Simple State-of-the-Art Baseline · ECCV (18) 2020 Visual Dialog · IEEE Trans. Pattern Anal. Mach. Intell. 2019 Audio Visual Scene-Aware Dialog · CVPR 2019 |
Robotics › Robot navigation and mapping
object goal navigation |
1.2 | 2 | 2023 | PIRLNav: Pretraining with Imitation and RL Finetuning for OBJECTNAV · CVPR 2023 Auxiliary Tasks and Exploration Enable ObjectGoal Navigation · ICCV 2021 |
Machine learning › Graph learning
graph neural network |
1.1 | 2 | 2022 | Spherical Channels for Modeling Atomic Interactions · NeurIPS 2022 Towards Training Billion Parameter Graph Neural Networks for Atomic Simulations · ICLR 2022 |
Machine learning › Reinforcement learning
exploration |
0.9 | 2 | 2021 | Auxiliary Tasks and Exploration Enable ObjectGoal Navigation · ICCV 2021 IR-VIC: Unsupervised Discovery of Sub-goals for Transfer in RL · IJCAI 2020 |
Computational science and engineering › model simulation
atomistic simulation |
0.9 | 1 | 2025 | UMA: A Family of Universal Models for Atoms · NeurIPS 2025 |
Computational science and engineering › model simulation › atomistic simulation
machine learning interatomic potential |
0.9 | 1 | 2025 | UMA: A Family of Universal Models for Atoms · NeurIPS 2025 |
Machine learning › Trustworthy machine learning
interpretability |
0.7 | 2 | 2020 | Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization · Int. J. Comput. Vis. 2020 Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization · ICCV 2017 |
Machine learning › Trustworthy machine learning › interpretability
visual explanation |
0.7 | 2 | 2020 | Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization · Int. J. Comput. Vis. 2020 Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization · ICCV 2017 |
Natural language and speech › Question answering and dialogue systems › multimodal question answering
embodied question answering |
0.7 | 2 | 2019 | Embodied Question Answering in Photorealistic Environments With Point Cloud Perception · CVPR 2019 Embodied Question Answering · CVPR 2018 |
Machine learning › Reinforcement learning › reinforcement learning for NLP
reinforcement fine-tuning |
0.7 | 1 | 2023 | PIRLNav: Pretraining with Imitation and RL Finetuning for OBJECTNAV · CVPR 2023 |
Computer vision › Vision and language
visual question answering |
0.6 | 2 | 2019 | Visual Dialog · IEEE Trans. Pattern Anal. Mach. Intell. 2019 Human Attention in Visual Question Answering: Do Humans and Deep Networks look at the same regions? · EMNLP 2016 |
Machine learning › Graph learning › graph neural network › geometric graph neural network
equivariant message passing |
0.6 | 1 | 2022 | Spherical Channels for Modeling Atomic Interactions · NeurIPS 2022 |
Computational science and engineering
computational chemistry |
0.6 | 1 | 2022 | Spherical Channels for Modeling Atomic Interactions · NeurIPS 2022 |
Computational science and engineering › computational chemistry
molecular simulation |
0.6 | 1 | 2022 | Towards Training Billion Parameter Graph Neural Networks for Atomic Simulations · ICLR 2022 |
Robotics › Robot navigation and mapping
embodied navigation |
0.5 | 2 | 2021 | Embodied Question Answering in Photorealistic Environments With Point Cloud Perception · CVPR 2019 Auxiliary Tasks and Exploration Enable ObjectGoal Navigation · ICCV 2021 |
Machine learning › Reinforcement learning
hierarchical reinforcement learning |
0.4 | 1 | 2020 | IR-VIC: Unsupervised Discovery of Sub-goals for Transfer in RL · IJCAI 2020 |
Machine learning › Reinforcement learning › hierarchical reinforcement learning › option discovery
subgoal discovery |
0.4 | 1 | 2020 | IR-VIC: Unsupervised Discovery of Sub-goals for Transfer in RL · IJCAI 2020 |
Computer vision › Vision and language
visual grounding |
0.4 | 2 | 2018 | Visual Dialog · CVPR 2017 Embodied Question Answering · CVPR 2018 |
Natural language and speech › Question answering and dialogue systems › multimodal dialogue system
audio-visual scene-aware dialog |
0.4 | 1 | 2019 | Audio Visual Scene-Aware Dialog · CVPR 2019 |
Computer vision › Video understanding and tracking
video question answering |
0.4 | 1 | 2019 | Audio Visual Scene-Aware Dialog · CVPR 2019 |
Robotics › Robot navigation and mapping › visual navigation
image-goal navigation |
0.3 | 1 | 2018 | Embodied Question Answering · CVPR 2018 |
Machine learning › Trustworthy machine learning › interpretability › visual explanation
class activation map |
0.3 | 1 | 2017 | Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization · ICCV 2017 |
Machine learning › Reinforcement learning › multi-agent reinforcement learning › multi-agent communication
cooperative multi-agent communication |
0.3 | 1 | 2017 | Learning Cooperative Visual Dialog Agents with Deep Reinforcement Learning · ICCV 2017 |
Natural language and speech › Question answering and dialogue systems › visual dialog
goal-oriented visual dialogue |
0.3 | 1 | 2017 | Learning Cooperative Visual Dialog Agents with Deep Reinforcement Learning · ICCV 2017 |
Computer vision › Image recognition and object detection › object localization
weakly supervised object localization |
0.3 | 1 | 2017 | Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization · ICCV 2017 |
Machine learning › Deep learning architectures and training
mixture of experts |
0.3 | 1 | 2025 | UMA: A Family of Universal Models for Atoms · NeurIPS 2025 |
Bioinformatics and computational biology
molecular property prediction |
0.3 | 1 | 2025 | UMA: A Family of Universal Models for Atoms · NeurIPS 2025 |
Usability and user experience research › cognitive modeling
attention modeling |
0.2 | 1 | 2016 | Human Attention in Visual Question Answering: Do Humans and Deep Networks look at the same regions? · EMNLP 2016 |
Robotics › Robot navigation and mapping › learning-based navigation
imitation learning for navigation |
0.2 | 1 | 2023 | PIRLNav: Pretraining with Imitation and RL Finetuning for OBJECTNAV · CVPR 2023 |
Machine learning › Efficient and distributed learning › distributed training
large-scale training |
0.2 | 1 | 2022 | Towards Training Billion Parameter Graph Neural Networks for Atomic Simulations · ICLR 2022 |
Methods — techniques the papers use, named apart from their topics
scaling laws · 1.7spherical harmonics · 1.1graph neural network · 1.1equivariance · 1.1behavior cloning · 1.0reinforcement learning · 1.0memory network · 1.0gradient-based localization · 0.7exploration reward · 0.5auxiliary tasks · 0.5multimodal fusion · 0.4dialog history modeling · 0.4late fusion · 0.3hierarchical recurrent encoder · 0.3encoder-decoder · 0.3attention supervision · 0.2attention annotation · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | UMA: A Family of Universal Models for AtomsabstractThe ability to quickly and accurately compute properties from atomic simulations is critical for advancing a large number of applications in chemistry and materials science including drug discovery, energy storage, and semiconductor manufacturing. To address this need, we present a family of Universal Models for Atoms (UMA), designed to push the frontier of speed, accuracy, and generalization. UMA models are trained on half a billion unique 3D atomic structures (the largest training runs to date) by compiling data across multiple chemical domains, e.g. molecules, materials, and catalysts. We develop empirical scaling laws to help understand how to increase model capacity alongside dataset size to achieve the best accuracy. The UMA small and medium models utilize a novel architectural design we refer to as mixture of linear experts that enables increasing model capacity without sacrificing speed. For example, UMA-medium has 1.4B parameters but only $\sim$50M active parameters per atomic structure. We evaluate UMA models on a diverse set of applications across multiple domains and find that, remarkably, a single model without any fine-tuning can perform similarly or better than specialized models. We are releasing the UMA code, weights, and associated data to accelerate computational workflows and enable the community to build increasingly capable AI models. Brandon M. Wood, Misko Dzamba, Xiang Fu 0005, Muhammed Shuaibi, Luis Barroso-Luque, Kareem Abdelmaqsoud, Vahe Gharakhanyan, John R. Kitchin, Daniel S. Levine 0003, Kyle Michel, Anuroop Sriram, Taco Cohen, Abhishek Das 0002, Sushree Jagriti Sahoo, Ammar Rizvi, Zachary W. Ulissi, C. Lawrence Zitnick |
NeurIPS | 14 |
| 2023 | PIRLNav: Pretraining with Imitation and RL Finetuning for OBJECTNAVabstractWe study ObjectGoal Navigation - where a virtual robot situated in a new environment is asked to navigate to an object. Prior work [1] has shown that imitation learning (IL) using behavior cloning (BC) on a dataset of human demonstrations achieves promising results. However, this has limitations - 1) BC policies generalize poorly to new states, since the training mimics actions not their consequences, and 2) collecting demonstrations is expensive. On the other hand, reinforcement learning (RL) is trivially scalable, but requires careful reward engineering to achieve desirable behavior. We present PIRLNav, a two-stage learning scheme for BC pretraining on human demonstrations followed by RL-finetuning. This leads to a policy that achieves a success rate of 65.0% on OBJECTNAV (+5.0% absolute over previous state-of-the-art). Using this BC→RL training recipe, we present a rigorous empirical analysis of design choices. First, we investigate whether human demonstrations can be replaced with ‘free’ (automatically generated) sources of demonstrations, e.g. shortest paths (SP) or task-agnostic frontier exploration (FE) trajectories. We find that BC→RL on human demonstrations outperforms BC→RL on SP and FE trajectories, even when controlled for the same BC-pretraining success on TRAIN, and even on a subset of VAL episodes where BC-pretraining success favors the SP or FE policies. Next, we study how RL-finetuning performance scales with the size of the BC pretraining dataset. We find that as we increase the size of the BC-pretraining dataset and get to high BC accuracies, the improvements from RL-finetuning are smaller, and that 90% of the performance of our best BC→RL policy can be achieved with less than half the number of BC demonstrations. Finally, we analyze failure modes of our OBJECTNAV policies, and present guidelines for further improving them. Project page: ram81.github.io/projects/pirlnav. Ram Ramrakhya, Dhruv Batra, Erik Wijmans, Abhishek Das 0002 |
CVPR | 4 |
| 2022 | Towards Training Billion Parameter Graph Neural Networks for Atomic Simulations
Anuroop Sriram, Abhishek Das 0002, Brandon M. Wood, Siddharth Goyal, C. Lawrence Zitnick |
ICLR | 2 |
| 2022 | Spherical Channels for Modeling Atomic InteractionsabstractModeling the energy and forces of atomic systems is a fundamental problem in computational chemistry with the potential to help address many of the world’s most pressing problems, including those related to energy scarcity and climate change. These calculations are traditionally performed using Density Functional Theory, which is computationally very expensive. Machine learning has the potential to dramatically improve the efficiency of these calculations from days or hours to seconds.We propose the Spherical Channel Network (SCN) to model atomic energies and forces. The SCN is a graph neural network where nodes represent atoms and edges their neighboring atoms. The atom embeddings are a set of spherical functions, called spherical channels, represented using spherical harmonics. We demonstrate, that by rotating the embeddings based on the 3D edge orientation, more information may be utilized while maintaining the rotational equivariance of the messages. While equivariance is a desirable property, we find that by relaxing this constraint in both message passing and aggregation, improved accuracy may be achieved. We demonstrate state-of-the-art results on the large-scale Open Catalyst 2020 dataset in both energy and force prediction for numerous tasks and metrics. C. Lawrence Zitnick, Abhishek Das 0002, Adeesh Kolluru, Janice Lan, Muhammed Shuaibi, Anuroop Sriram, Zachary W. Ulissi, Brandon M. Wood |
NeurIPS | 2 |
| 2021 | Auxiliary Tasks and Exploration Enable ObjectGoal NavigationabstractObjectGoal Navigation (ObjectNav) is an embodied task wherein agents are to navigate to an object instance in an unseen environment. Prior works have shown that end-to-end ObjectNav agents that use vanilla visual and recurrent modules, e.g. a CNN+RNN, perform poorly due to overfitting and sample inefficiency. This has motivated current state-of-the-art methods to mix analytic and learned components and operate on explicit spatial maps of the environment. We instead re-enable a generic learned agent by adding auxiliary learning tasks and an exploration reward. Our agents achieve 24.5% success and 8.1% SPL, a 37% and 8% relative improvement over prior state-of-the-art, respectively, on the Habitat ObjectNav Challenge [35]. From our analysis, we propose that agents will act to simplify their visual inputs so as to smooth their RNN dynamics, and that auxiliary tasks reduce overfitting by minimizing effective RNN dimensionality; i.e. a performant ObjectNav agent that must maintain coherent plans over long horizons does so by learning smooth, low-dimensional recurrent dynamics. Site: joel99.github.io/objectnav/ Joel Ye, Dhruv Batra, Abhishek Das 0002, Erik Wijmans |
ICCV | 3 |
| 2020 | Large-Scale Pretraining for Visual Dialog: A Simple State-of-the-Art Baseline
Vishvak Murahari, Dhruv Batra, Devi Parikh, Abhishek Das 0002 |
ECCV (18) | 4 |
| 2020 | IR-VIC: Unsupervised Discovery of Sub-goals for Transfer in RLabstractWe propose a novel framework to identify sub-goals useful for exploration in sequential decision making tasks under partial observability. We utilize the variational intrinsic control framework (Gregor et.al., 2016) which maximizes empowerment -- the ability to reliably reach a diverse set of states and show how to identify sub-goals as states with high necessary option information through an information theoretic regularizer. Despite being discovered without explicit goal supervision, our sub-goals provide better exploration and sample complexity on challenging grid-world navigation tasks compared to supervised counterparts in prior work. Nirbhay Modhe, Prithvijit Chattopadhyay, Abhishek Das 0002, Devi Parikh, Dhruv Batra, Ramakrishna Vedantam |
IJCAI | 4 |
| 2020 | Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization
Ramprasaath R. Selvaraju, Michael Cogswell, Abhishek Das 0002, Ramakrishna Vedantam, Devi Parikh, Dhruv Batra |
Int. J. Comput. Vis. | 3 |
| 2019 | Audio Visual Scene-Aware DialogabstractWe introduce the task of scene-aware dialog. Our goal is to generate a complete and natural response to a question about a scene, given video and audio of the scene and the history of previous turns in the dialog. To answer successfully, agents must ground concepts from the question in the video while leveraging contextual cues from the dialog history. To benchmark this task, we introduce the Audio Visual Scene-Aware Dialog (AVSD) Dataset. For each of more than 11,000 videos of human actions from the Charades dataset, our dataset contains a dialog about the video, plus a final summary of the video by one of the dialog participants. We train several baseline systems for this task and evaluate the performance of the trained models using both qualitative and quantitative metrics. Our results indicate that models must utilize all the available inputs (video, audio, question, and dialog history) to perform best on this dataset. Huda AlAmri, Vincent Cartillier, Abhishek Das 0002, Jue Wang 0010, Anoop Cherian, Irfan A. Essa, Dhruv Batra, Tim K. Marks, Chiori Hori, Stefan Lee, Devi Parikh |
CVPR | 3 |
| 2019 | Embodied Question Answering in Photorealistic Environments With Point Cloud PerceptionabstractTo help bridge the gap between internet vision-style problems and the goal of vision for embodied perception we instantiate a large-scale navigation task -- Embodied Question Answering [1] in photo-realistic environments (Matterport 3D). We thoroughly study navigation policies that utilize 3D point clouds, RGB images, or their combination. Our analysis of these models reveals several key findings. We find that two seemingly naive navigation baselines, forward-only and random, are strong navigators and challenging to outperform, due to the specific choice of the evaluation setting presented by [1]. We find a novel loss-weighting scheme we call Inflection Weighting to be important when training recurrent models for navigation with behavior cloning and are able to out perform the baselines with this technique. We find that point clouds provide a richer signal than RGB images for learning obstacle avoidance, motivating the use (and continued study) of 3D deep learning models for embodied navigation. Erik Wijmans, Samyak Datta, Oleksandr Maksymets, Abhishek Das 0002, Georgia Gkioxari, Stefan Lee, Irfan A. Essa, Devi Parikh, Dhruv Batra |
CVPR | 4 |
| 2019 | End-to-end Audio Visual Scene-aware Dialog Using Multimodal Attention-based Video FeaturesabstractIn order for machines interacting with the real world to have conversations with users about the objects and events around them, they need to understand dynamic audiovisual scenes. The recent revolution of neural network models allows us to combine various modules into a single end-to-end differentiable network. As a result, Audio Visual Scene-Aware Dialog (AVSD) systems for real-world applications can be developed by integrating state-of-the-art technologies from multiple research areas, including end-to-end dialog technologies, visual question answering (VQA) technologies, and video description technologies. In this paper, we introduce a new data set of dialogs about videos of human behaviors, as well as an end-to-end Audio Visual Scene-Aware Dialog (AVSD) model, trained using this new data set, that generates responses in a dialog about a video. By using features that were developed for multimodal attention-based video description, our system improves the quality of generated dialog about dynamic video scenes. Chiori Hori, Huda AlAmri, Jue Wang 0010, Gordon Wichern, Takaaki Hori, Anoop Cherian, Tim K. Marks, Vincent Cartillier, Raphael Gontijo Lopes, Abhishek Das 0002, Irfan A. Essa, Dhruv Batra, Devi Parikh |
ICASSP | 10 |
| 2019 | Visual DialogabstractWe introduce the task of Visual Dialog, which requires an AI agent to hold a meaningful dialog with humans in natural, conversational language about visual content. Specifically, given an image, a dialog history, and a question about the image, the agent has to ground the question in image, infer context from history, and answer the question accurately. Visual Dialog is disentangled enough from a specific downstream task so as to serve as a general test of machine intelligence, while being sufficiently grounded in vision to allow objective evaluation of individual responses and benchmark progress. We develop a novel two-person real-time chat data-collection protocol to curate a large-scale Visual Dialog dataset (VisDial). VisDial v0.9 has been released and consists of$\sim$1.2M dialog question-answer pairs from 10-round, human-human dialogs grounded in$\sim$120k images from the COCO dataset. We introduce a family of neural encoder-decoder models for Visual Dialog with 3 encoders—Late Fusion, Hierarchical Recurrent Encoder and Memory Network (optionally with attention over image features)—and 2 decoders (generative and discriminative), which outperform a number of sophisticated baselines. We propose a retrieval-based evaluation protocol for Visual Dialog where the AI agent is asked to sort a set of candidate answers and evaluated on metrics such as mean-reciprocal-rank and recall$@k$of human response. We quantify the gap between machine and human performance on the Visual Dialog task via human studies. Putting it all together, we demonstrate the first ‘visual chatbot’! Our dataset, code, pretrained models and visual chatbot are available onhttps://visualdialog.org. Abhishek Das 0002, Satwik Kottur, Khushi Gupta, Avi Singh, Deshraj Yadav, Stefan Lee, José M. F. Moura, Devi Parikh, Dhruv Batra |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2018 | Embodied Question AnsweringabstractWe present a new AI task - Embodied Question Answering(EmbodiedQA) - where an agent is spawned at a random location in a 3D environment and asked a question ('What color is the car?'). In order to answer, the agent must first intelligently navigate to explore the environment, gather necessary visual information through first-person (egocentric) vision, and then answer the question ('orange'). EmbodiedQA requires a range of AI skills - language understanding, visual recognition, active perception, goal-driven navigation, commonsense reasoning, long-term memory, and grounding language into actions. In this work, we develop a dataset of questions and answers in House3D environments [1], evaluation metrics, and a hierarchical model trained with imitation and reinforcement learning. Abhishek Das 0002, Samyak Datta, Georgia Gkioxari, Stefan Lee, Devi Parikh, Dhruv Batra |
CVPR | 1 |
| 2017 | Visual DialogabstractWe introduce the task of Visual Dialog, which requires an AI agent to hold a meaningful dialog with humans in natural, conversational language about visual content. Specifically, given an image, a dialog history, and a question about the image, the agent has to ground the question in image, infer context from history, and answer the question accurately. Visual Dialog is disentangled enough from a specific downstream task so as to serve as a general test of machine intelligence, while being grounded in vision enough to allow objective evaluation of individual responses and benchmark progress. We develop a novel two-person chat data-collection protocol to curate a large-scale Visual Dialog dataset (VisDial). VisDial contains 1 dialog (10 question-answer pairs) on ~140k images from the COCO dataset, with a total of ~1.4M dialog question-answer pairs. We introduce a family of neural encoder-decoder models for Visual Dialog with 3 encoders (Late Fusion, Hierarchical Recurrent Encoder and Memory Network) and 2 decoders (generative and discriminative), which outperform a number of sophisticated baselines. We propose a retrieval-based evaluation protocol for Visual Dialog where the AI agent is asked to sort a set of candidate answers and evaluated on metrics such as mean-reciprocal-rank of human response. We quantify gap between machine and human performance on the Visual Dialog task via human studies. Our dataset, code, and trained models will be released publicly at https://visualdialog.org. Putting it all together, we demonstrate the first visual chatbot!. Abhishek Das 0002, Satwik Kottur, Khushi Gupta, Avi Singh, Deshraj Yadav, José M. F. Moura, Devi Parikh, Dhruv Batra |
CVPR | 1 |
| 2017 | Learning Cooperative Visual Dialog Agents with Deep Reinforcement LearningabstractWe introduce the first goal-driven training for visual question answering and dialog agents. Specifically, we pose a cooperative `image guessing' game between two agents - Q-BOT and A-BOT- who communicate in natural language dialog so that Q-BOT can select an unseen image from a lineup of images. We use deep reinforcement learning (RL) to learn the policies of these agents end-to-end - from pixels to multi-agent multi-round dialog to game reward.,,We demonstrate two experimental results.,,First, as a `sanity check' demonstration of pure RL (from scratch), we show results on a synthetic world, where the agents communicate in ungrounded vocabularies, i.e., symbols with no pre-specified meanings (X, Y, Z). We find that two bots invent their own communication protocol and start using certain symbols to ask/answer about certain visual attributes (shape/color/style). Thus, we demonstrate the emergence of grounded language and communication among `visual' dialog agents with no human supervision.,,Second, we conduct large-scale real-image experiments on the VisDial dataset [5], where we pretrain on dialog data with supervised learning (SL) and show that the RL finetuned agents significantly outperform supervised pretraining. Interestingly, the RL Q-BOT learns to ask questions that A-BOT is good at, ultimately resulting in more informative dialog and a better team. Abhishek Das 0002, Satwik Kottur, José M. F. Moura, Stefan Lee, Dhruv Batra |
ICCV | 1 |
| 2017 | Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based LocalizationabstractWe propose a technique for producing `visual explanations' for decisions from a large class of Convolutional Neural Network (CNN)-based models, making them more transparent. Our approach - Gradient-weighted Class Activation Mapping (Grad-CAM), uses the gradients of any target concept (say logits for `dog' or even a caption), flowing into the final convolutional layer to produce a coarse localization map highlighting the important regions in the image for predicting the concept. Unlike previous approaches, Grad- CAM is applicable to a wide variety of CNN model-families: (1) CNNs with fully-connected layers (e.g. VGG), (2) CNNs used for structured outputs (e.g. captioning), (3) CNNs used in tasks with multi-modal inputs (e.g. visual question answering) or reinforcement learning, without architectural changes or re-training. We combine Grad-CAM with existing fine-grained visualizations to create a high-resolution class-discriminative visualization, Guided Grad-CAM, and apply it to image classification, image captioning, and visual question answering (VQA) models, including ResNet-based architectures. In the context of image classification models, our visualizations (a) lend insights into failure modes of these models (showing that seemingly unreasonable predictions have reasonable explanations), (b) outperform previous methods on the ILSVRC-15 weakly-supervised localization task, (c) are more faithful to the underlying model, and (d) help achieve model generalization by identifying dataset bias. For image captioning and VQA, our visualizations show even non-attention based models can localize inputs. Finally, we design and conduct human studies to measure if Grad-CAM explanations help users establish appropriate trust in predictions from deep networks and show that Grad-CAM helps untrained users successfully discern a `stronger' deep network from a `weaker' one even when both make identical predictions. Our code is available at https: //github.com/ramprs/grad-cam/ along with a demo on CloudCV [2] and video at youtu.be/COjUB9Izk6E. Ramprasaath R. Selvaraju, Michael Cogswell, Abhishek Das 0002, Ramakrishna Vedantam, Devi Parikh, Dhruv Batra |
ICCV | 3 |
| 2017 | Human Attention in Visual Question Answering: Do Humans and Deep Networks Look at the Same Regions?
Abhishek Das 0002, Harsh Agrawal, C. Lawrence Zitnick, Devi Parikh, Dhruv Batra |
Comput. Vis. Image Underst. | 1 |
| 2016 | Human Attention in Visual Question Answering: Do Humans and Deep Networks look at the same regions?abstractWe conduct large-scale studies on ‘human attention’ in Visual Question Answering (VQA) to understand where humans choose to look to answer questions about images. We design and test multiple game-inspired novel attention-annotation interfaces that require the subject to sharpen regions of a blurred image to answer a question. Thus, we introduce the VQA-HAT (Human ATtention) dataset. We evaluate attention maps generated by state-of-the-art VQA models against human attention both qualitatively (via visualizations) and quantitatively (via rank-order correlation). Our experiments show that current attention models in VQA do not seem to be looking at the same regions as humans. Finally, we train VQA models with explicit attention supervision, and find that it improves VQA performance. Abhishek Das 0002, Harsh Agrawal, C. Lawrence Zitnick, Devi Parikh, Dhruv Batra |
EMNLP | 1 |