Prithvijit Chattopadhyay

dblp:179/2452 · DBLP profile ↗
← Back
14ranked-venue papers
6as first author
7since 2021 · last 2024
0000-0001-7555-3644ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 5 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
13 papers
Trustworthy machine learning · 27% Transfer learning and domain adaptation · 19% Image recognition and object detection · 9%
Human-computer interaction and pervasive computing
1 paper
Human-AI interaction · 100%

Topics — the 30 heaviest of 37, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning
robustness
1.832023
LANCE: Stress-testing Visual Models by Generating Language-guided Counterfactual Images · NeurIPS 2023
Benchmarking Low-Shot Robustness to Natural Distribution Shifts · ICCV 2023
RobustNav: Towards Benchmarking Robustness in Embodied Navigation · ICCV 2021
Machine learning › Transfer learning and domain adaptation
sim-to-real transfer
1.422024
AUGCAL: Improving Sim2Real Adaptation by Uncertainty Calibration on Augmented Synthetic Images · ICLR 2024
Pasta: Proportional Amplitude Spectrum Training Augmentation for Syn-to-Real Domain Generalization · ICCV 2023
Computer vision › 3D vision
aerial image analysis
0.812024
SKYSCENES: A Synthetic Dataset for Aerial Scene Understanding · ECCV (79) 2024
Computer vision › 3D vision › remote sensing
aerial imagery
0.812024
SKYSCENES: A Synthetic Dataset for Aerial Scene Understanding · ECCV (79) 2024
Machine learning › Trustworthy machine learning
calibration
0.812024
AUGCAL: Improving Sim2Real Adaptation by Uncertainty Calibration on Augmented Synthetic Images · ICLR 2024
Computer vision › Segmentation and scene understanding
scene understanding
0.812024
SKYSCENES: A Synthetic Dataset for Aerial Scene Understanding · ECCV (79) 2024
Machine learning › Transfer learning and domain adaptation › domain adaptation
unsupervised domain adaptation
0.812024
AUGCAL: Improving Sim2Real Adaptation by Uncertainty Calibration on Augmented Synthetic Images · ICLR 2024
Machine learning › Generative modeling › conditional generative model
counterfactual image generation
0.712023
LANCE: Stress-testing Visual Models by Generating Language-guided Counterfactual Images · NeurIPS 2023
Machine learning › Transfer learning and domain adaptation
few-shot learning
0.712023
Benchmarking Low-Shot Robustness to Natural Distribution Shifts · ICCV 2023
Machine learning › Deep learning architectures and training › data augmentation
image augmentation
0.712023
Pasta: Proportional Amplitude Spectrum Training Augmentation for Syn-to-Real Domain Generalization · ICCV 2023
Computer vision › Vision and language › language-guided learning › language-guided vision
language-guided image editing
0.712023
LANCE: Stress-testing Visual Models by Generating Language-guided Counterfactual Images · NeurIPS 2023
Computer vision › Image recognition and object detection
object detection
0.712023
Pasta: Proportional Amplitude Spectrum Training Augmentation for Syn-to-Real Domain Generalization · ICCV 2023
Machine learning › Trustworthy machine learning › robustness › distribution shift
robustness to distribution shift
0.712023
Benchmarking Low-Shot Robustness to Natural Distribution Shifts · ICCV 2023
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning
self-supervised representation learning
0.712023
Battle of the Backbones: A Large-Scale Comparison of Pretrained Models across Computer Vision Tasks · NeurIPS 2023
Machine learning › Trustworthy machine learning › robustness › model robustness evaluation
stress testing
0.712023
LANCE: Stress-testing Visual Models by Generating Language-guided Counterfactual Images · NeurIPS 2023
Machine learning › Transfer learning and domain adaptation › domain generalization
synthetic-to-real generalization
0.712023
Pasta: Proportional Amplitude Spectrum Training Augmentation for Syn-to-Real Domain Generalization · ICCV 2023
Robotics › Robot navigation and mapping
embodied navigation
0.512021
RobustNav: Towards Benchmarking Robustness in Embodied Navigation · ICCV 2021
Machine learning › Trustworthy machine learning › robustness evaluation
robustness benchmark
0.512021
RobustNav: Towards Benchmarking Robustness in Embodied Navigation · ICCV 2021
Machine learning › Transfer learning and domain adaptation
domain generalization
0.412020
Learning to Balance Specificity and Invariance for In and Out of Domain Generalization · ECCV (9) 2020
Machine learning › Reinforcement learning
exploration
0.412020
IR-VIC: Unsupervised Discovery of Sub-goals for Transfer in RL · IJCAI 2020
Machine learning › Reinforcement learning
hierarchical reinforcement learning
0.412020
IR-VIC: Unsupervised Discovery of Sub-goals for Transfer in RL · IJCAI 2020
Machine learning › Trustworthy machine learning
invariance
0.412020
Learning to Balance Specificity and Invariance for In and Out of Domain Generalization · ECCV (9) 2020
Machine learning › Reinforcement learning › hierarchical reinforcement learning › option discovery
subgoal discovery
0.412020
IR-VIC: Unsupervised Discovery of Sub-goals for Transfer in RL · IJCAI 2020
Computer vision › Vision and language
visual question answering
0.422018
Do explanations make VQA models more predictable to a human? · EMNLP 2018
Counting Everyday Objects in Everyday Scenes · CVPR 2017
Natural language and speech › Question answering and dialogue systems
dialogue generation
0.412019
Improving Generative Visual Dialog by Answering Diverse Questions · EMNLP/IJCNLP (1) 2019
Natural language and speech › Question answering and dialogue systems › question generation
diverse question generation
0.412019
Improving Generative Visual Dialog by Answering Diverse Questions · EMNLP/IJCNLP (1) 2019
Natural language and speech › Question answering and dialogue systems
visual dialog
0.412019
Improving Generative Visual Dialog by Answering Diverse Questions · EMNLP/IJCNLP (1) 2019
Knowledge, reasoning and agents › Knowledge representation and reasoning › knowledge engineering › knowledge integration
domain knowledge integration
0.312018
Choose Your Neuron: Incorporating Domain Knowledge Through Neuron-Importance · ECCV (13) 2018
Computer vision › Vision and language › visual question answering
explainable visual question answering
0.312018
Do explanations make VQA models more predictable to a human? · EMNLP 2018
Machine learning › Trustworthy machine learning
interpretability
0.312018
Choose Your Neuron: Incorporating Domain Knowledge Through Neuron-Importance · ECCV (13) 2018

Methods — techniques the papers use, named apart from their topics

data augmentation · 1.3uncertainty calibration · 0.8synthetic dataset generation · 0.8text-based image editing · 0.7pre-training · 0.7large language model · 0.7fourier domain augmentation · 0.7fine-tuning · 0.7benchmarking · 0.7amplitude spectrum perturbation · 0.7human-in-the-loop evaluation · 0.3
YearPublicationVenuePosition
2024 SKYSCENES: A Synthetic Dataset for Aerial Scene Understanding
Sahil Khose, Anisha Pal, Aayushi Agarwal 0001, Deepanshi 0002, Judy Hoffman, Prithvijit Chattopadhyay
ECCV (79)6
2024 AUGCAL: Improving Sim2Real Adaptation by Uncertainty Calibration on Augmented Synthetic Images
abstract
Synthetic data (Sim) drawn from simulators have emerged as a popular alternativefor training models where acquiring annotated real-world images is difficult. However, transferring models trained on synthetic images to real-world applicationscan be challenging due to appearance disparities. A commonly employed solution to counter this Sim2Real gap is unsupervised domain adaptation, where models are trained using labeled Sim data and unlabeled Real data. Mispredictions made by such Sim2Real adapted models are often associated with miscalibration – stemming from overconfident predictions on real data. In this paper, we introduce AUGCAL, a simple training-time patch for unsupervised adaptation that improves Sim2Real adapted models by – (1) reducing overall miscalibration, (2) reducing overconfidence in incorrect predictions and (3) improving confidence score reliability by better guiding misclassification detection – all while retaining or improving Sim2Real performance. Given a base Sim2Real adaptation algorithm, at training time, AUGCAL involves replacing vanilla Sim images with strongly augmented views (AUG intervention) and additionally optimizing for a training time calibration loss on augmented Sim predictions (CAL intervention). We motivate AUGCAL using a brief analytical justification of how to reduce miscalibration on unlabeled REAL data. Through our experiments, we empirically show the efficacy of AUGCAL across multiple adaptation methods, backbones, tasks and shifts.
Prithvijit Chattopadhyay, Bharat Goyal, Boglarka Ecsedi, Viraj Prabhu, Judy Hoffman
ICLR1
2023 Pasta: Proportional Amplitude Spectrum Training Augmentation for Syn-to-Real Domain Generalization
abstract
Synthetic data offers the promise of cheap and bountiful training data for settings where labeled real-world data is scarce. However, models trained on synthetic data significantly underperform when evaluated on real-world data. In this paper, we propose Proportional Amplitude Spectrum Training Augmentation (Pasta), a simple and effective augmentation strategy to improve out-of-the-box synthetic-to-real (syn-to-real) generalization performance. Pasta perturbs the amplitude spectra of synthetic images in the Fourier domain to generate augmented views. Specifically, with Pasta we propose a structured perturbation strategy where high-frequency components are perturbed relatively more than the low-frequency ones. For the tasks of semantic segmentation (GTAV→Real), object detection (Sim10K→Real), and object recognition (VisDA-C Syn→Real), across a total of 5 syn-to-real shifts, we find that Pasta outperforms more complex state-of-the-art generalization methods while being complementary to the same.
Prithvijit Chattopadhyay, Kartik Sarangmath, Vivek Vijaykumar, Judy Hoffman
ICCV1
2023 Benchmarking Low-Shot Robustness to Natural Distribution Shifts
abstract
Robustness to natural distribution shifts has seen remarkable progress thanks to recent pre-training strategies combined with better fine-tuning methods. However, such fine-tuning assumes access to large amounts of labelled data, and the extent to which the observations hold when the amount of training data is not as high remains unknown. We address this gap by performing the first in-depth study of robustness to various natural distribution shifts in different low-shot regimes: spanning datasets, architectures, pre-trained initializations, and state-of-the-art robustness interventions. Most importantly, we find that there is no single model of choice that is often more robust than others, and existing interventions can fail to improve robustness on some datasets even if they do so in the full-shot regime. We hope that our work will motivate the community to focus on this problem of practical importance. Our code and low-shot subsets are publicly available at this url.
Aaditya Singh, Kartik Sarangmath, Prithvijit Chattopadhyay, Judy Hoffman
ICCV3
2023 Battle of the Backbones: A Large-Scale Comparison of Pretrained Models across Computer Vision Tasks
abstract
Neural network based computer vision systems are typically built on a backbone, a pretrained or randomly initialized feature extractor. Several years ago, the default option was an ImageNet-trained convolutional neural network. However, the recent past has seen the emergence of countless backbones pretrained using various algorithms and datasets. While this abundance of choice has led to performance increases for a range of systems, it is difficult for practitioners to make informed decisions about which backbone to choose. Battle of the Backbones (BoB) makes this choice easier by benchmarking a diverse suite of pretrained models, including vision-language models, those trained via self-supervised learning, and the Stable Diffusion backbone, across a diverse set of computer vision tasks ranging from classification to object detection to OOD generalization and more. Furthermore, BoB sheds light on promising directions for the research community to advance computer vision by illuminating strengths and weakness of existing approaches through a comprehensive analysis conducted on more than 1500 training runs. While vision transformers (ViTs) and self-supervised learning (SSL) are increasingly popular, we find that convolutional neural networks pretrained in a supervised fashion on large training sets still perform best on most tasks among the models we consider. Moreover, in apples-to-apples comparisons on the same architectures and similarly sized pretraining datasets, we find that SSL backbones are highly competitive, indicating that future works should perform SSL pretraining with advanced architectures and larger pretraining datasets. We release the raw results of our experiments along with code that allows researchers to put their own backbones through the gauntlet here: https://github.com/hsouri/Battle-of-the-Backbones.
Micah Goldblum, Hossein Souri, Renkun Ni, Manli Shu, Viraj Prabhu, Gowthami Somepalli, Prithvijit Chattopadhyay, Mark Ibrahim, Adrien Bardes, Judy Hoffman, Rama Chellappa, Andrew Gordon Wilson, Tom Goldstein
NeurIPS7
2023 LANCE: Stress-testing Visual Models by Generating Language-guided Counterfactual Images
abstract
We propose an automated algorithm to stress-test a trained visual model by generating language-guided counterfactual test images (LANCE). Our method leverages recent progress in large language modeling and text-based image editing to augment an IID test set with a suite of diverse, realistic, and challenging test images without altering model weights. We benchmark the performance of a diverse set of pre-trained models on our generated data and observe significant and consistent performance drops. We further analyze model sensitivity across different types of edits, and demonstrate its applicability at surfacing previously unknown class-level model biases in ImageNet. Code is available at https://github.com/virajprabhu/lance.
Viraj Prabhu, Sriram Yenamandra, Prithvijit Chattopadhyay, Judy Hoffman
NeurIPS3
2021 RobustNav: Towards Benchmarking Robustness in Embodied Navigation
abstract
As an attempt towards assessing the robustness of embodied navigation agents, we propose RobustNav, a framework to quantify the performance of embodied navigation agents when exposed to a wide variety of visual – affecting RGB inputs – and dynamics – affecting transition dynamics – corruptions. Most recent efforts in visual navigation have typically focused on generalizing to novel target environments with similar appearance and dynamics characteristics. With RobustNav, we find that some standard embodied navigation agents significantly underperform (or fail) in the presence of visual or dynamics corruptions. We systematically analyze the kind of idiosyncrasies that emerge in the behavior of such agents when operating under corruptions. Finally, for visual corruptions in RobustNav, we show that while standard techniques to improve robustness such as data-augmentation and self-supervised adaptation offer some zero-shot resistance and improvements in navigation performance, there is still a long way to go in terms of recovering lost performance relative to clean "non-corrupt" settings, warranting more research in this direction. Our code is available at https://github.com/allenai/robustnav.
Prithvijit Chattopadhyay, Judy Hoffman, Roozbeh Mottaghi, Aniruddha Kembhavi
ICCV1
2020 Learning to Balance Specificity and Invariance for In and Out of Domain Generalization
Prithvijit Chattopadhyay, Yogesh Balaji, Judy Hoffman
ECCV (9)1
2020 IR-VIC: Unsupervised Discovery of Sub-goals for Transfer in RL
abstract
We propose a novel framework to identify sub-goals useful for exploration in sequential decision making tasks under partial observability. We utilize the variational intrinsic control framework (Gregor et.al., 2016) which maximizes empowerment -- the ability to reliably reach a diverse set of states and show how to identify sub-goals as states with high necessary option information through an information theoretic regularizer. Despite being discovered without explicit goal supervision, our sub-goals provide better exploration and sample complexity on challenging grid-world navigation tasks compared to supervised counterparts in prior work.
Nirbhay Modhe, Prithvijit Chattopadhyay, Abhishek Das 0002, Devi Parikh, Dhruv Batra, Ramakrishna Vedantam
IJCAI2
2019 Improving Generative Visual Dialog by Answering Diverse Questions
abstract
Vishvak Murahari, Prithvijit Chattopadhyay, Dhruv Batra, Devi Parikh, Abhishek Das. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Vishvak Murahari, Prithvijit Chattopadhyay, Dhruv Batra, Devi Parikh
EMNLP/IJCNLP (1)2
2018 Choose Your Neuron: Incorporating Domain Knowledge Through Neuron-Importance
Ramprasaath R. Selvaraju, Prithvijit Chattopadhyay, Tilak Sharma, Dhruv Batra, Devi Parikh, Stefan Lee
ECCV (13)2
2018 Do explanations make VQA models more predictable to a human?
abstract
A rich line of research attempts to make deep neural networks more transparent by generating human-interpretable 'explanations' of their decision process, especially for interactive tasks like Visual Question Answering (VQA).In this work, we analyze if existing explanations indeed make a VQA model -its responses as well as failures -more predictable to a human.Surprisingly, we find that they do not.On the other hand, we find that humanin-the-loop approaches that treat the model as a black-box do.
Arjun Chandrasekaran, Viraj Prabhu, Deshraj Yadav, Prithvijit Chattopadhyay, Devi Parikh
EMNLP4
2017 Counting Everyday Objects in Everyday Scenes
abstract
We are interested in counting the number of instances of object classes in natural, everyday images. Previous counting approaches tackle the problem in restricted domains such as counting pedestrians in surveillance videos. Counts can also be estimated from outputs of other vision tasks like object detection. In this work, we build dedicated models for counting designed to tackle the large variance in counts, appearances, and scales of objects found in natural scenes. Our approach is inspired by the phenomenon of subitizing – the ability of humans to make quick assessments of counts given a perceptual signal, for small count values. Given a natural scene, we employ a divide and conquer strategy while incorporating context across the scene to adapt the subitizing idea to counting. Our approach offers consistent improvements over numerous baseline approaches for counting on the PASCAL VOC 2007 and COCO datasets. Subsequently, we study how counting can be used to improve object detection. We then show a proof of concept application of our counting methods to the task of Visual Question Answering, by studying the how many? questions in the VQA and COCO-QA datasets.
Prithvijit Chattopadhyay, Ramakrishna Vedantam, Ramprasaath R. Selvaraju, Dhruv Batra, Devi Parikh
CVPR1
2017 Evaluating Visual Conversational Agents via Cooperative Human-AI Games
abstract
As AI continues to advance, human-AI teams are inevitable. However, progress in AI is routinely measured in isolation, without a human in the loop. It is crucial to benchmark progress in AI, not just in isolation, but also in terms of how it translates to helping humans perform certain tasks, i.e., the performance of human-AI teams. In this work, we design a cooperative game — GuessWhich — to measure human-AI team performance in the specific context of the AI being a visual conversational agent. GuessWhich involves live interaction between the human and the AI. The AI, which we call ALICE, is provided an image which is unseen by the human. Following a brief description of the image, the human questions ALICE about this secret image to identify it from a fixed pool of images. We measure performance of the human-ALICE team by the number of guesses it takes the human to correctly identify the secret image after a fixed number of dialog rounds with ALICE. We compare performance of the human-ALICE teams for two versions of ALICE. Our human studies suggest a counterintuitive trend – that while AI literature shows that one version outperforms the other when paired with an AI questioner bot, we find that this improvement in AI-AI performance does not translate to improved human-AI performance. This suggests a mismatch between benchmarking of AI in isolation and in the context of human-AI teams.
Prithvijit Chattopadhyay, Deshraj Yadav, Viraj Prabhu, Arjun Chandrasekaran, Stefan Lee, Dhruv Batra, Devi Parikh
HCOMP1