Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Abhishek Das 0002

dblp:40/5262-2 · DBLP profile ↗
← Back
18ranked-venue papers
6as first author
5since 2021 · last 2025
0000-0002-4718-4316ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 6 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 3 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
16 papers
Question answering and dialogue systems · 22% Reinforcement learning · 20% Robot navigation and mapping · 16%
Interdisciplinary, comprehensive, and emerging computing
3 papers
Computational science and engineering · 92% Bioinformatics and computational biology · 8%

Topics — the 30 heaviest of 36, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Question answering and dialogue systems
visual dialog
1.852020
Large-Scale Pretraining for Visual Dialog: A Simple State-of-the-Art Baseline · ECCV (18) 2020
Visual Dialog · IEEE Trans. Pattern Anal. Mach. Intell. 2019
Audio Visual Scene-Aware Dialog · CVPR 2019
Robotics › Robot navigation and mapping
object goal navigation
1.222023
PIRLNav: Pretraining with Imitation and RL Finetuning for OBJECTNAV · CVPR 2023
Auxiliary Tasks and Exploration Enable ObjectGoal Navigation · ICCV 2021
Machine learning › Graph learning
graph neural network
1.122022
Spherical Channels for Modeling Atomic Interactions · NeurIPS 2022
Towards Training Billion Parameter Graph Neural Networks for Atomic Simulations · ICLR 2022
Machine learning › Reinforcement learning
exploration
0.922021
Auxiliary Tasks and Exploration Enable ObjectGoal Navigation · ICCV 2021
IR-VIC: Unsupervised Discovery of Sub-goals for Transfer in RL · IJCAI 2020
Computational science and engineering › model simulation
atomistic simulation
0.912025
UMA: A Family of Universal Models for Atoms · NeurIPS 2025
Computational science and engineering › model simulation › atomistic simulation
machine learning interatomic potential
0.912025
UMA: A Family of Universal Models for Atoms · NeurIPS 2025
Machine learning › Trustworthy machine learning
interpretability
0.722020
Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization · Int. J. Comput. Vis. 2020
Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization · ICCV 2017
Machine learning › Trustworthy machine learning › interpretability
visual explanation
0.722020
Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization · Int. J. Comput. Vis. 2020
Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization · ICCV 2017
Natural language and speech › Question answering and dialogue systems › multimodal question answering
embodied question answering
0.722019
Embodied Question Answering in Photorealistic Environments With Point Cloud Perception · CVPR 2019
Embodied Question Answering · CVPR 2018
Machine learning › Reinforcement learning › reinforcement learning for NLP
reinforcement fine-tuning
0.712023
PIRLNav: Pretraining with Imitation and RL Finetuning for OBJECTNAV · CVPR 2023
Computer vision › Vision and language
visual question answering
0.622019
Visual Dialog · IEEE Trans. Pattern Anal. Mach. Intell. 2019
Human Attention in Visual Question Answering: Do Humans and Deep Networks look at the same regions? · EMNLP 2016
Machine learning › Graph learning › graph neural network › geometric graph neural network
equivariant message passing
0.612022
Spherical Channels for Modeling Atomic Interactions · NeurIPS 2022
Computational science and engineering
computational chemistry
0.612022
Spherical Channels for Modeling Atomic Interactions · NeurIPS 2022
Computational science and engineering › computational chemistry
molecular simulation
0.612022
Towards Training Billion Parameter Graph Neural Networks for Atomic Simulations · ICLR 2022
Robotics › Robot navigation and mapping
embodied navigation
0.522021
Embodied Question Answering in Photorealistic Environments With Point Cloud Perception · CVPR 2019
Auxiliary Tasks and Exploration Enable ObjectGoal Navigation · ICCV 2021
Machine learning › Reinforcement learning
hierarchical reinforcement learning
0.412020
IR-VIC: Unsupervised Discovery of Sub-goals for Transfer in RL · IJCAI 2020
Machine learning › Reinforcement learning › hierarchical reinforcement learning › option discovery
subgoal discovery
0.412020
IR-VIC: Unsupervised Discovery of Sub-goals for Transfer in RL · IJCAI 2020
Computer vision › Vision and language
visual grounding
0.422018
Visual Dialog · CVPR 2017
Embodied Question Answering · CVPR 2018
Natural language and speech › Question answering and dialogue systems › multimodal dialogue system
audio-visual scene-aware dialog
0.412019
Audio Visual Scene-Aware Dialog · CVPR 2019
Computer vision › Video understanding and tracking
video question answering
0.412019
Audio Visual Scene-Aware Dialog · CVPR 2019
Robotics › Robot navigation and mapping › visual navigation
image-goal navigation
0.312018
Embodied Question Answering · CVPR 2018
Machine learning › Trustworthy machine learning › interpretability › visual explanation
class activation map
0.312017
Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization · ICCV 2017
Machine learning › Reinforcement learning › multi-agent reinforcement learning › multi-agent communication
cooperative multi-agent communication
0.312017
Learning Cooperative Visual Dialog Agents with Deep Reinforcement Learning · ICCV 2017
Natural language and speech › Question answering and dialogue systems › visual dialog
goal-oriented visual dialogue
0.312017
Learning Cooperative Visual Dialog Agents with Deep Reinforcement Learning · ICCV 2017
Computer vision › Image recognition and object detection › object localization
weakly supervised object localization
0.312017
Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization · ICCV 2017
Machine learning › Deep learning architectures and training
mixture of experts
0.312025
UMA: A Family of Universal Models for Atoms · NeurIPS 2025
Bioinformatics and computational biology
molecular property prediction
0.312025
UMA: A Family of Universal Models for Atoms · NeurIPS 2025
Usability and user experience research › cognitive modeling
attention modeling
0.212016
Human Attention in Visual Question Answering: Do Humans and Deep Networks look at the same regions? · EMNLP 2016
Robotics › Robot navigation and mapping › learning-based navigation
imitation learning for navigation
0.212023
PIRLNav: Pretraining with Imitation and RL Finetuning for OBJECTNAV · CVPR 2023
Machine learning › Efficient and distributed learning › distributed training
large-scale training
0.212022
Towards Training Billion Parameter Graph Neural Networks for Atomic Simulations · ICLR 2022

Methods — techniques the papers use, named apart from their topics

scaling laws · 1.7spherical harmonics · 1.1graph neural network · 1.1equivariance · 1.1behavior cloning · 1.0reinforcement learning · 1.0memory network · 1.0gradient-based localization · 0.7exploration reward · 0.5auxiliary tasks · 0.5multimodal fusion · 0.4dialog history modeling · 0.4late fusion · 0.3hierarchical recurrent encoder · 0.3encoder-decoder · 0.3attention supervision · 0.2attention annotation · 0.2
YearPublicationVenuePosition
2025 UMA: A Family of Universal Models for Atoms
abstract
The ability to quickly and accurately compute properties from atomic simulations is critical for advancing a large number of applications in chemistry and materials science including drug discovery, energy storage, and semiconductor manufacturing. To address this need, we present a family of Universal Models for Atoms (UMA), designed to push the frontier of speed, accuracy, and generalization. UMA models are trained on half a billion unique 3D atomic structures (the largest training runs to date) by compiling data across multiple chemical domains, e.g. molecules, materials, and catalysts. We develop empirical scaling laws to help understand how to increase model capacity alongside dataset size to achieve the best accuracy. The UMA small and medium models utilize a novel architectural design we refer to as mixture of linear experts that enables increasing model capacity without sacrificing speed. For example, UMA-medium has 1.4B parameters but only $\sim$50M active parameters per atomic structure. We evaluate UMA models on a diverse set of applications across multiple domains and find that, remarkably, a single model without any fine-tuning can perform similarly or better than specialized models. We are releasing the UMA code, weights, and associated data to accelerate computational workflows and enable the community to build increasingly capable AI models.
Brandon M. Wood, Misko Dzamba, Xiang Fu 0005, Muhammed Shuaibi, Luis Barroso-Luque, Kareem Abdelmaqsoud, Vahe Gharakhanyan, John R. Kitchin, Daniel S. Levine 0003, Kyle Michel, Anuroop Sriram, Taco Cohen, Abhishek Das 0002, Sushree Jagriti Sahoo, Ammar Rizvi, Zachary W. Ulissi, C. Lawrence Zitnick
NeurIPS14
2023 PIRLNav: Pretraining with Imitation and RL Finetuning for OBJECTNAV
abstract
We study ObjectGoal Navigation - where a virtual robot situated in a new environment is asked to navigate to an object. Prior work [1] has shown that imitation learning (IL) using behavior cloning (BC) on a dataset of human demonstrations achieves promising results. However, this has limitations - 1) BC policies generalize poorly to new states, since the training mimics actions not their consequences, and 2) collecting demonstrations is expensive. On the other hand, reinforcement learning (RL) is trivially scalable, but requires careful reward engineering to achieve desirable behavior. We present PIRLNav, a two-stage learning scheme for BC pretraining on human demonstrations followed by RL-finetuning. This leads to a policy that achieves a success rate of 65.0% on OBJECTNAV (+5.0% absolute over previous state-of-the-art). Using this BC→RL training recipe, we present a rigorous empirical analysis of design choices. First, we investigate whether human demonstrations can be replaced with ‘free’ (automatically generated) sources of demonstrations, e.g. shortest paths (SP) or task-agnostic frontier exploration (FE) trajectories. We find that BC→RL on human demonstrations outperforms BC→RL on SP and FE trajectories, even when controlled for the same BC-pretraining success on TRAIN, and even on a subset of VAL episodes where BC-pretraining success favors the SP or FE policies. Next, we study how RL-finetuning performance scales with the size of the BC pretraining dataset. We find that as we increase the size of the BC-pretraining dataset and get to high BC accuracies, the improvements from RL-finetuning are smaller, and that 90% of the performance of our best BC→RL policy can be achieved with less than half the number of BC demonstrations. Finally, we analyze failure modes of our OBJECTNAV policies, and present guidelines for further improving them. Project page: ram81.github.io/projects/pirlnav.
Ram Ramrakhya, Dhruv Batra, Erik Wijmans, Abhishek Das 0002
CVPR4
2022 Towards Training Billion Parameter Graph Neural Networks for Atomic Simulations
Anuroop Sriram, Abhishek Das 0002, Brandon M. Wood, Siddharth Goyal, C. Lawrence Zitnick
ICLR2
2022 Spherical Channels for Modeling Atomic Interactions
abstract
Modeling the energy and forces of atomic systems is a fundamental problem in computational chemistry with the potential to help address many of the world’s most pressing problems, including those related to energy scarcity and climate change. These calculations are traditionally performed using Density Functional Theory, which is computationally very expensive. Machine learning has the potential to dramatically improve the efficiency of these calculations from days or hours to seconds.We propose the Spherical Channel Network (SCN) to model atomic energies and forces. The SCN is a graph neural network where nodes represent atoms and edges their neighboring atoms. The atom embeddings are a set of spherical functions, called spherical channels, represented using spherical harmonics. We demonstrate, that by rotating the embeddings based on the 3D edge orientation, more information may be utilized while maintaining the rotational equivariance of the messages. While equivariance is a desirable property, we find that by relaxing this constraint in both message passing and aggregation, improved accuracy may be achieved. We demonstrate state-of-the-art results on the large-scale Open Catalyst 2020 dataset in both energy and force prediction for numerous tasks and metrics.
C. Lawrence Zitnick, Abhishek Das 0002, Adeesh Kolluru, Janice Lan, Muhammed Shuaibi, Anuroop Sriram, Zachary W. Ulissi, Brandon M. Wood
NeurIPS2
2021 Auxiliary Tasks and Exploration Enable ObjectGoal Navigation
abstract
ObjectGoal Navigation (ObjectNav) is an embodied task wherein agents are to navigate to an object instance in an unseen environment. Prior works have shown that end-to-end ObjectNav agents that use vanilla visual and recurrent modules, e.g. a CNN+RNN, perform poorly due to overfitting and sample inefficiency. This has motivated current state-of-the-art methods to mix analytic and learned components and operate on explicit spatial maps of the environment. We instead re-enable a generic learned agent by adding auxiliary learning tasks and an exploration reward. Our agents achieve 24.5% success and 8.1% SPL, a 37% and 8% relative improvement over prior state-of-the-art, respectively, on the Habitat ObjectNav Challenge [35]. From our analysis, we propose that agents will act to simplify their visual inputs so as to smooth their RNN dynamics, and that auxiliary tasks reduce overfitting by minimizing effective RNN dimensionality; i.e. a performant ObjectNav agent that must maintain coherent plans over long horizons does so by learning smooth, low-dimensional recurrent dynamics. Site: joel99.github.io/objectnav/
Joel Ye, Dhruv Batra, Abhishek Das 0002, Erik Wijmans
ICCV3
2020 Large-Scale Pretraining for Visual Dialog: A Simple State-of-the-Art Baseline
Vishvak Murahari, Dhruv Batra, Devi Parikh, Abhishek Das 0002
ECCV (18)4
2020 IR-VIC: Unsupervised Discovery of Sub-goals for Transfer in RL
abstract
We propose a novel framework to identify sub-goals useful for exploration in sequential decision making tasks under partial observability. We utilize the variational intrinsic control framework (Gregor et.al., 2016) which maximizes empowerment -- the ability to reliably reach a diverse set of states and show how to identify sub-goals as states with high necessary option information through an information theoretic regularizer. Despite being discovered without explicit goal supervision, our sub-goals provide better exploration and sample complexity on challenging grid-world navigation tasks compared to supervised counterparts in prior work.
Nirbhay Modhe, Prithvijit Chattopadhyay, Abhishek Das 0002, Devi Parikh, Dhruv Batra, Ramakrishna Vedantam
IJCAI4
2020 Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization
Ramprasaath R. Selvaraju, Michael Cogswell, Abhishek Das 0002, Ramakrishna Vedantam, Devi Parikh, Dhruv Batra
Int. J. Comput. Vis.3
2019 Audio Visual Scene-Aware Dialog
abstract
We introduce the task of scene-aware dialog. Our goal is to generate a complete and natural response to a question about a scene, given video and audio of the scene and the history of previous turns in the dialog. To answer successfully, agents must ground concepts from the question in the video while leveraging contextual cues from the dialog history. To benchmark this task, we introduce the Audio Visual Scene-Aware Dialog (AVSD) Dataset. For each of more than 11,000 videos of human actions from the Charades dataset, our dataset contains a dialog about the video, plus a final summary of the video by one of the dialog participants. We train several baseline systems for this task and evaluate the performance of the trained models using both qualitative and quantitative metrics. Our results indicate that models must utilize all the available inputs (video, audio, question, and dialog history) to perform best on this dataset.
Huda AlAmri, Vincent Cartillier, Abhishek Das 0002, Jue Wang 0010, Anoop Cherian, Irfan A. Essa, Dhruv Batra, Tim K. Marks, Chiori Hori, Stefan Lee, Devi Parikh
CVPR3
2019 Embodied Question Answering in Photorealistic Environments With Point Cloud Perception
abstract
To help bridge the gap between internet vision-style problems and the goal of vision for embodied perception we instantiate a large-scale navigation task -- Embodied Question Answering [1] in photo-realistic environments (Matterport 3D). We thoroughly study navigation policies that utilize 3D point clouds, RGB images, or their combination. Our analysis of these models reveals several key findings. We find that two seemingly naive navigation baselines, forward-only and random, are strong navigators and challenging to outperform, due to the specific choice of the evaluation setting presented by [1]. We find a novel loss-weighting scheme we call Inflection Weighting to be important when training recurrent models for navigation with behavior cloning and are able to out perform the baselines with this technique. We find that point clouds provide a richer signal than RGB images for learning obstacle avoidance, motivating the use (and continued study) of 3D deep learning models for embodied navigation.
Erik Wijmans, Samyak Datta, Oleksandr Maksymets, Abhishek Das 0002, Georgia Gkioxari, Stefan Lee, Irfan A. Essa, Devi Parikh, Dhruv Batra
CVPR4
2019 End-to-end Audio Visual Scene-aware Dialog Using Multimodal Attention-based Video Features
abstract
In order for machines interacting with the real world to have conversations with users about the objects and events around them, they need to understand dynamic audiovisual scenes. The recent revolution of neural network models allows us to combine various modules into a single end-to-end differentiable network. As a result, Audio Visual Scene-Aware Dialog (AVSD) systems for real-world applications can be developed by integrating state-of-the-art technologies from multiple research areas, including end-to-end dialog technologies, visual question answering (VQA) technologies, and video description technologies. In this paper, we introduce a new data set of dialogs about videos of human behaviors, as well as an end-to-end Audio Visual Scene-Aware Dialog (AVSD) model, trained using this new data set, that generates responses in a dialog about a video. By using features that were developed for multimodal attention-based video description, our system improves the quality of generated dialog about dynamic video scenes.
Chiori Hori, Huda AlAmri, Jue Wang 0010, Gordon Wichern, Takaaki Hori, Anoop Cherian, Tim K. Marks, Vincent Cartillier, Raphael Gontijo Lopes, Abhishek Das 0002, Irfan A. Essa, Dhruv Batra, Devi Parikh
ICASSP10
2019 Visual Dialog
abstract
We introduce the task of Visual Dialog, which requires an AI agent to hold a meaningful dialog with humans in natural, conversational language about visual content. Specifically, given an image, a dialog history, and a question about the image, the agent has to ground the question in image, infer context from history, and answer the question accurately. Visual Dialog is disentangled enough from a specific downstream task so as to serve as a general test of machine intelligence, while being sufficiently grounded in vision to allow objective evaluation of individual responses and benchmark progress. We develop a novel two-person real-time chat data-collection protocol to curate a large-scale Visual Dialog dataset (VisDial). VisDial v0.9 has been released and consists of$\sim$1.2M dialog question-answer pairs from 10-round, human-human dialogs grounded in$\sim$120k images from the COCO dataset. We introduce a family of neural encoder-decoder models for Visual Dialog with 3 encoders—Late Fusion, Hierarchical Recurrent Encoder and Memory Network (optionally with attention over image features)—and 2 decoders (generative and discriminative), which outperform a number of sophisticated baselines. We propose a retrieval-based evaluation protocol for Visual Dialog where the AI agent is asked to sort a set of candidate answers and evaluated on metrics such as mean-reciprocal-rank and recall$@k$of human response. We quantify the gap between machine and human performance on the Visual Dialog task via human studies. Putting it all together, we demonstrate the first ‘visual chatbot’! Our dataset, code, pretrained models and visual chatbot are available onhttps://visualdialog.org.
Abhishek Das 0002, Satwik Kottur, Khushi Gupta, Avi Singh, Deshraj Yadav, Stefan Lee, José M. F. Moura, Devi Parikh, Dhruv Batra
IEEE Trans. Pattern Anal. Mach. Intell.1
2018 Embodied Question Answering
abstract
We present a new AI task - Embodied Question Answering(EmbodiedQA) - where an agent is spawned at a random location in a 3D environment and asked a question ('What color is the car?'). In order to answer, the agent must first intelligently navigate to explore the environment, gather necessary visual information through first-person (egocentric) vision, and then answer the question ('orange'). EmbodiedQA requires a range of AI skills - language understanding, visual recognition, active perception, goal-driven navigation, commonsense reasoning, long-term memory, and grounding language into actions. In this work, we develop a dataset of questions and answers in House3D environments [1], evaluation metrics, and a hierarchical model trained with imitation and reinforcement learning.
Abhishek Das 0002, Samyak Datta, Georgia Gkioxari, Stefan Lee, Devi Parikh, Dhruv Batra
CVPR1
2017 Visual Dialog
abstract
We introduce the task of Visual Dialog, which requires an AI agent to hold a meaningful dialog with humans in natural, conversational language about visual content. Specifically, given an image, a dialog history, and a question about the image, the agent has to ground the question in image, infer context from history, and answer the question accurately. Visual Dialog is disentangled enough from a specific downstream task so as to serve as a general test of machine intelligence, while being grounded in vision enough to allow objective evaluation of individual responses and benchmark progress. We develop a novel two-person chat data-collection protocol to curate a large-scale Visual Dialog dataset (VisDial). VisDial contains 1 dialog (10 question-answer pairs) on ~140k images from the COCO dataset, with a total of ~1.4M dialog question-answer pairs. We introduce a family of neural encoder-decoder models for Visual Dialog with 3 encoders (Late Fusion, Hierarchical Recurrent Encoder and Memory Network) and 2 decoders (generative and discriminative), which outperform a number of sophisticated baselines. We propose a retrieval-based evaluation protocol for Visual Dialog where the AI agent is asked to sort a set of candidate answers and evaluated on metrics such as mean-reciprocal-rank of human response. We quantify gap between machine and human performance on the Visual Dialog task via human studies. Our dataset, code, and trained models will be released publicly at https://visualdialog.org. Putting it all together, we demonstrate the first visual chatbot!.
Abhishek Das 0002, Satwik Kottur, Khushi Gupta, Avi Singh, Deshraj Yadav, José M. F. Moura, Devi Parikh, Dhruv Batra
CVPR1
2017 Learning Cooperative Visual Dialog Agents with Deep Reinforcement Learning
abstract
We introduce the first goal-driven training for visual question answering and dialog agents. Specifically, we pose a cooperative `image guessing' game between two agents - Q-BOT and A-BOT- who communicate in natural language dialog so that Q-BOT can select an unseen image from a lineup of images. We use deep reinforcement learning (RL) to learn the policies of these agents end-to-end - from pixels to multi-agent multi-round dialog to game reward.,,We demonstrate two experimental results.,,First, as a `sanity check' demonstration of pure RL (from scratch), we show results on a synthetic world, where the agents communicate in ungrounded vocabularies, i.e., symbols with no pre-specified meanings (X, Y, Z). We find that two bots invent their own communication protocol and start using certain symbols to ask/answer about certain visual attributes (shape/color/style). Thus, we demonstrate the emergence of grounded language and communication among `visual' dialog agents with no human supervision.,,Second, we conduct large-scale real-image experiments on the VisDial dataset [5], where we pretrain on dialog data with supervised learning (SL) and show that the RL finetuned agents significantly outperform supervised pretraining. Interestingly, the RL Q-BOT learns to ask questions that A-BOT is good at, ultimately resulting in more informative dialog and a better team.
Abhishek Das 0002, Satwik Kottur, José M. F. Moura, Stefan Lee, Dhruv Batra
ICCV1
2017 Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization
abstract
We propose a technique for producing `visual explanations' for decisions from a large class of Convolutional Neural Network (CNN)-based models, making them more transparent. Our approach - Gradient-weighted Class Activation Mapping (Grad-CAM), uses the gradients of any target concept (say logits for `dog' or even a caption), flowing into the final convolutional layer to produce a coarse localization map highlighting the important regions in the image for predicting the concept. Unlike previous approaches, Grad- CAM is applicable to a wide variety of CNN model-families: (1) CNNs with fully-connected layers (e.g. VGG), (2) CNNs used for structured outputs (e.g. captioning), (3) CNNs used in tasks with multi-modal inputs (e.g. visual question answering) or reinforcement learning, without architectural changes or re-training. We combine Grad-CAM with existing fine-grained visualizations to create a high-resolution class-discriminative visualization, Guided Grad-CAM, and apply it to image classification, image captioning, and visual question answering (VQA) models, including ResNet-based architectures. In the context of image classification models, our visualizations (a) lend insights into failure modes of these models (showing that seemingly unreasonable predictions have reasonable explanations), (b) outperform previous methods on the ILSVRC-15 weakly-supervised localization task, (c) are more faithful to the underlying model, and (d) help achieve model generalization by identifying dataset bias. For image captioning and VQA, our visualizations show even non-attention based models can localize inputs. Finally, we design and conduct human studies to measure if Grad-CAM explanations help users establish appropriate trust in predictions from deep networks and show that Grad-CAM helps untrained users successfully discern a `stronger' deep network from a `weaker' one even when both make identical predictions. Our code is available at https: //github.com/ramprs/grad-cam/ along with a demo on CloudCV [2] and video at youtu.be/COjUB9Izk6E.
Ramprasaath R. Selvaraju, Michael Cogswell, Abhishek Das 0002, Ramakrishna Vedantam, Devi Parikh, Dhruv Batra
ICCV3
2017 Human Attention in Visual Question Answering: Do Humans and Deep Networks Look at the Same Regions?
Abhishek Das 0002, Harsh Agrawal, C. Lawrence Zitnick, Devi Parikh, Dhruv Batra
Comput. Vis. Image Underst.1
2016 Human Attention in Visual Question Answering: Do Humans and Deep Networks look at the same regions?
abstract
We conduct large-scale studies on ‘human attention’ in Visual Question Answering (VQA) to understand where humans choose to look to answer questions about images. We design and test multiple game-inspired novel attention-annotation interfaces that require the subject to sharpen regions of a blurred image to answer a question. Thus, we introduce the VQA-HAT (Human ATtention) dataset. We evaluate attention maps generated by state-of-the-art VQA models against human attention both qualitatively (via visualizations) and quantitatively (via rank-order correlation). Our experiments show that current attention models in VQA do not seem to be looking at the same regions as humans. Finally, we train VQA models with explicit attention supervision, and find that it improves VQA performance.
Abhishek Das 0002, Harsh Agrawal, C. Lawrence Zitnick, Devi Parikh, Dhruv Batra
EMNLP1