Chaithanya Kumar Mummadi

dblp:208/6386 · also Mummadi Chaithanya Kumar · DBLP profile ↗
← Back
11ranked-venue papers
2as first author
7since 2021 · last 2026
0000-0002-1173-2720ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
10 papers
Trustworthy machine learning · 50% Transfer learning and domain adaptation · 11% Segmentation and scene understanding · 11%
Network and information security
1 paper
Security and privacy of machine learning · 100%

Topics — the 30 heaviest of 31, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning
robustness
1.322024
PerceptionCLIP: Visual Classification by Inferring and Conditioning on Contexts · ICLR 2024
DiagViB-6: A Diagnostic Benchmark Suite for Vision Models in the Presence of Shortcut and Generalization Opportunities · ICCV 2021
Machine learning › Trustworthy machine learning › robustness
adversarial robustness
1.232022
Give Me Your Attention: Dot-Product Attention Considered Harmful for Adversarial Patch Robustness · CVPR 2022
Defending Against Universal Perturbations With Shared Adversarial Training · ICCV 2019
Universal Adversarial Perturbations Against Semantic Image Segmentation · ICCV 2017
Machine learning › Trustworthy machine learning › robustness
shortcut learning
1.122022
Overcoming Shortcut Learning in a Target Domain by Generalizing Basic Visual Factors from a Source Domain · ECCV (25) 2022
DiagViB-6: A Diagnostic Benchmark Suite for Vision Models in the Presence of Shortcut and Generalization Opportunities · ICCV 2021
Machine learning › Efficient and distributed learning
federated learning
0.812024
Federated Text-driven Prompt Generation for Vision-Language Models · ICLR 2024
Computer vision › Vision and language › vision-language model
prompt learning
0.812024
Federated Text-driven Prompt Generation for Vision-Language Models · ICLR 2024
Machine learning › Trustworthy machine learning › debiasing
spurious feature mitigation
0.812024
PerceptionCLIP: Visual Classification by Inferring and Conditioning on Contexts · ICLR 2024
Machine learning › Transfer learning and domain adaptation › zero-shot learning
zero-shot classification
0.812024
PerceptionCLIP: Visual Classification by Inferring and Conditioning on Contexts · ICLR 2024
Machine learning › Transfer learning and domain adaptation
domain generalization
0.612022
Overcoming Shortcut Learning in a Target Domain by Generalizing Basic Visual Factors from a Source Domain · ECCV (25) 2022
Security and privacy of machine learning › adversarial attack › physical adversarial attack
adversarial patch
0.612022
Give Me Your Attention: Dot-Product Attention Considered Harmful for Adversarial Patch Robustness · CVPR 2022
Machine learning › Learning theory
generalization
0.512021
DiagViB-6: A Diagnostic Benchmark Suite for Vision Models in the Presence of Shortcut and Generalization Opportunities · ICCV 2021
Machine learning › Trustworthy machine learning
out-of-distribution generalization
0.512021
DiagViB-6: A Diagnostic Benchmark Suite for Vision Models in the Presence of Shortcut and Generalization Opportunities · ICCV 2021
Machine learning › Trustworthy machine learning › robustness
robustness to corruption
0.512021
Does enhanced shape bias improve neural network robustness to common corruptions? · ICLR 2021
Computer vision › Image recognition and object detection
shape bias
0.512021
Does enhanced shape bias improve neural network robustness to common corruptions? · ICLR 2021
Machine learning › Trustworthy machine learning › robustness › learning with noisy labels
label filtering
0.412020
SELF: Learning to Filter Noisy Labels with Self-Ensembling · ICLR 2020
Machine learning › Trustworthy machine learning › robustness
learning with noisy labels
0.412020
SELF: Learning to Filter Noisy Labels with Self-Ensembling · ICLR 2020
Machine learning › Representation and self-supervised learning
self-ensembling
0.412020
SELF: Learning to Filter Noisy Labels with Self-Ensembling · ICLR 2020
Computer vision › Segmentation and scene understanding
semantic segmentation
0.422019
Universal Adversarial Perturbations Against Semantic Image Segmentation · ICCV 2017
Defending Against Universal Perturbations With Shared Adversarial Training · ICCV 2019
Machine learning › Trustworthy machine learning › robustness › adversarial robustness
adversarial training
0.412019
Defending Against Universal Perturbations With Shared Adversarial Training · ICCV 2019
Computer vision › Segmentation and scene understanding › pseudo-label learning
pseudo-label refinement
0.412019
DeepUSPS: Deep Robust Unsupervised Saliency Prediction via Self-supervision · NeurIPS 2019
Computer vision › Segmentation and scene understanding › saliency detection
salient object detection
0.412019
DeepUSPS: Deep Robust Unsupervised Saliency Prediction via Self-supervision · NeurIPS 2019
Machine learning › Trustworthy machine learning › adversarial machine learning › adversarial defense
universal perturbation defense
0.412019
Defending Against Universal Perturbations With Shared Adversarial Training · ICCV 2019
Computer vision › Segmentation and scene understanding › saliency detection
unsupervised saliency detection
0.412019
DeepUSPS: Deep Robust Unsupervised Saliency Prediction via Self-supervision · NeurIPS 2019
Machine learning › Trustworthy machine learning › robustness › adversarial examples
universal adversarial perturbation
0.312017
Universal Adversarial Perturbations Against Semantic Image Segmentation · ICCV 2017
Computer vision › Image recognition and object detection
image classification
0.322021
DiagViB-6: A Diagnostic Benchmark Suite for Vision Models in the Presence of Shortcut and Generalization Opportunities · ICCV 2021
Defending Against Universal Perturbations With Shared Adversarial Training · ICCV 2019
Computer vision › Vision and language › vision-language model
CLIP
0.212024
PerceptionCLIP: Visual Classification by Inferring and Conditioning on Contexts · ICLR 2024
Machine learning › Transfer learning and domain adaptation
generalization to unseen classes
0.212024
Federated Text-driven Prompt Generation for Vision-Language Models · ICLR 2024
Computer vision › Vision and language
vision-language model
0.212024
PerceptionCLIP: Visual Classification by Inferring and Conditioning on Contexts · ICLR 2024
Machine learning › Deep learning architectures and training › transformer
vision transformer
0.212022
Give Me Your Attention: Dot-Product Attention Considered Harmful for Adversarial Patch Robustness · CVPR 2022
Machine learning › Representation and self-supervised learning › representation learning
visual representation learning
0.212022
Overcoming Shortcut Learning in a Target Domain by Generalizing Basic Visual Factors from a Source Domain · ECCV (25) 2022
Machine learning › Deep learning architectures and training
convolutional neural network
0.112021
Does enhanced shape bias improve neural network robustness to common corruptions? · ICLR 2021

Methods — techniques the papers use, named apart from their topics

dot-product attention analysis · 1.1adversarial patch optimization · 1.1vision-language model · 0.8training-free prompting · 0.8prompt generation · 0.8contextual attribute inference · 0.8shape bias analysis · 0.5metrics · 0.5diagnostic benchmark · 0.5data augmentation · 0.5
YearPublicationVenuePosition
2026 RadarSim: Simulating Single-Chip Radar Via Multimodal Neural Fields
abstract
Radars are an ideal complement to cameras: both are inexpensive, solid-state sensors, with cameras offering fine angular resolution, while radars provide metric depth and robustness under adverse weather. However, radar data is more difficult to interpret than camera images and varies significantly between sensors, necessitating increased reliance on simulation for prototyping sensors and processing pipelines. Recent work treating radar reconstruction as a novel view synthesis problem has shown great promise in reconstructing radar-relevant geometry and simulating lowlevel radar data. However, such methods are constrained by the low spatial resolution of the underlying radar. To address this, we propose a unified differentiable renderer, RadarSim, which leverages the high angular resolution of RGB cameras to generate Doppler radar range images from a camera-initialized neural field. Using a novel data set of calibrated radar camera recordings from a custom handheld rig, we demonstrate that RadarSim produces sharper geometry and Doppler range frames than radar-only reconstructions.
Chuhan Chen, Tianshu Huang, Akarsh Prabhakara, Chaithanya Kumar Mummadi, Zhongxiao Cong, Anthony Rowe 0001, Matthew O'Toole, Deva Ramanan
3DV4
2024 PerceptionCLIP: Visual Classification by Inferring and Conditioning on Contexts
abstract
Vision-language models like CLIP are widely used in zero-shot image classification due to their ability to understand various visual concepts and natural language descriptions. However, how to fully leverage CLIP's unprecedented human-like understanding capabilities to achieve better performance is still an open question. This paper draws inspiration from the human visual perception process: when classifying an object, humans first infer contextual attributes (e.g., background and orientation) which help separate the foreground object from the background, and then classify the object based on this information. Inspired by it, we observe that providing CLIP with contextual attributes improves zero-shot image classification and mitigates reliance on spurious features. We also observe that CLIP itself can reasonably infer the attributes from an image. With these observations, we propose a training-free, two-step zero-shot classification method PerceptionCLIP. Given an image, it first infers contextual attributes (e.g., background) and then performs object classification conditioning on them. Our experiments show that PerceptionCLIP achieves better generalization, group robustness, and interpretability.
Bang An 0001, Sicheng Zhu, Michael-Andrei Panaitescu-Liess, Chaithanya Kumar Mummadi, Furong Huang
ICLR4
2024 Federated Text-driven Prompt Generation for Vision-Language Models
abstract
Prompt learning for vision-language models, e.g., CoOp, has shown great success in adapting CLIP to different downstream tasks, making it a promising solution for federated learning due to computational reasons. Existing prompt learning techniques replace hand-crafted text prompts with learned vectors that offer improvements on seen classes, but struggle to generalize to unseen classes. Our work addresses this challenge by proposing Federated Text-driven Prompt Generation (FedTPG), which learns a unified prompt generation network across multiple remote clients in a scalable manner. The prompt generation network is conditioned on task-related text input, thus is context-aware, making it suitable to generalize for both seen and unseen classes. Our comprehensive empirical evaluations on nine diverse image classification datasets show that our method is superior to existing federated prompt learning methods, achieving better overall generalization on both seen and unseen classes, as well as datasets.
Chaithanya Kumar Mummadi, Madan Ravi Ganesh, Lu Peng 0001, Wan-Yi Lin
ICLR3
2022 Give Me Your Attention: Dot-Product Attention Considered Harmful for Adversarial Patch Robustness
abstract
Neural architectures based on attention such as vision transformers are revolutionizing image recognition. Their main benefit is that attention allows reasoning about all parts of a scene jointly. In this paper, we show how the global reasoning of (scaled) dot-product attention can be the source of a major vulnerability when confronted with adversarial patch attacks. We provide a theoretical understanding of this vulnerability and relate it to an adversary's ability to misdirect the attention of all queries to a single key token under the control of the adversarial patch. We propose novel adversarial objectives for crafting adversarial patches which target this vulnerability explicitly. We show the effectiveness of the proposed patch attacks on popular image classification (ViTs and DeiTs) and object detection models (DETR). We find that adversarial patches occupying 0.5% of the input can lead to robust accuracies as low as 0% for ViT on ImageNet, and reduce the mAP of DETR on MS COCO to less than 3%.
Giulio Lovisotto, Nicole Finnie, Mauricio Munoz, Chaithanya Kumar Mummadi, Jan Hendrik Metzen
CVPR4
2022 Overcoming Shortcut Learning in a Target Domain by Generalizing Basic Visual Factors from a Source Domain
Piyapat Saranrittichai, Chaithanya Kumar Mummadi, Claudia Blaiotta, Mauricio Munoz, Volker Fischer 0003
ECCV (25)2
2021 DiagViB-6: A Diagnostic Benchmark Suite for Vision Models in the Presence of Shortcut and Generalization Opportunities
abstract
Common deep neural networks (DNNs) for image classification have been shown to rely on shortcut opportunities (SO) in the form of predictive and easy-to-represent visual factors. This is known as shortcut learning and leads to impaired generalization. In this work, we show that common DNNs also suffer from shortcut learning when predicting only basic visual object factors of variation (FoV) such as shape, color, or texture. We argue that besides shortcut opportunities, generalization opportunities (GO) are also an inherent part of real-world vision data and arise from partial independence between predicted classes and FoVs. We also argue that it is necessary for DNNs to exploit GO to overcome shortcut learning. Our core contribution is to introduce the Diagnostic Vision Benchmark suite DiagViB-6, which includes datasets and metrics to study a network’s shortcut vulnerability and generalization capability for six independent FoV. In particular, DiagViB-6 allows controlling the type and degree of SO and GO in a dataset. We benchmark a wide range of popular vision architectures and show that they can exploit GO only to a limited extent.
Elias Eulig, Piyapat Saranrittichai, Chaithanya Kumar Mummadi, Kilian Rambach, William Beluch, Xiahan Shi, Volker Fischer 0003
ICCV3
2021 Does enhanced shape bias improve neural network robustness to common corruptions?
Chaithanya Kumar Mummadi, Ranjitha Subramaniam, Robin Hutmacher, Julien Vitay, Volker Fischer 0003, Jan Hendrik Metzen
ICLR1
2020 SELF: Learning to Filter Noisy Labels with Self-Ensembling
Duc Tam Nguyen, Chaithanya Kumar Mummadi, Thi-Phuong-Nhung Ngo, Thi Hoai Phuong Nguyen, Laura Beggel, Thomas Brox
ICLR2
2019 Defending Against Universal Perturbations With Shared Adversarial Training
abstract
Classifiers such as deep neural networks have been shown to be vulnerable against adversarial perturbations on problems with high-dimensional input space. While adversarial training improves the robustness of image classifiers against such adversarial perturbations, it leaves them sensitive to perturbations on a non-negligible fraction of the inputs. In this work, we show that adversarial training is more effective in preventing universal perturbations, where the same perturbation needs to fool a classifier on many inputs. Moreover, we investigate the trade-off between robustness against universal perturbations and performance on unperturbed data and propose an extension of adversarial training that handles this trade-off more gracefully. We present results for image classification and semantic segmentation to showcase that universal perturbations that fool a model hardened with adversarial training become clearly perceptible and show patterns of the target scene.
Chaithanya Kumar Mummadi, Thomas Brox, Jan Hendrik Metzen
ICCV1
2019 DeepUSPS: Deep Robust Unsupervised Saliency Prediction via Self-supervision
abstract
Deep neural network (DNN) based salient object detection in images based on high-quality labels is expensive. Alternative unsupervised approaches rely on careful selection of multiple handcrafted saliency methods to generate noisy pseudo-ground-truth labels. In this work, we propose a two-stage mechanism for robust unsupervised object saliency prediction, where the first stage involves refinement of the noisy pseudo labels generated from different handcrafted methods. Each handcrafted method is substituted by a deep network that learns to generate the pseudo labels. These labels are refined incrementally in multiple iterations via our proposed self-supervision technique. In the second stage, the refined labels produced from multiple networks representing multiple saliency methods are used to train the actual saliency detection network. We show that this self-learning procedure outperforms all the existing unsupervised methods over different datasets. Results are even comparable to those of fully-supervised state-of-the-art approaches.
Duc Tam Nguyen, Maximilian Dax, Chaithanya Kumar Mummadi, Thi-Phuong-Nhung Ngo, Thi Hoai Phuong Nguyen, Zhongyu Lou, Thomas Brox
NeurIPS3
2017 Universal Adversarial Perturbations Against Semantic Image Segmentation
abstract
While deep learning is remarkably successful on perceptual tasks, it was also shown to be vulnerable to adversarial perturbations of the input. These perturbations denote noise added to the input that was generated specifically to fool the system while being quasi-imperceptible for humans. More severely, there even exist universal perturbations that are input-agnostic but fool the network on the majority of inputs. While recent work has focused on image classification, this work proposes attacks against semantic image segmentation: we present an approach for generating (universal) adversarial perturbations that make the network yield a desired target segmentation as output. We show empirically that there exist barely perceptible universal noise patterns which result in nearly the same predicted segmentation for arbitrary inputs. Furthermore, we also show the existence of universal noise which removes a target class (e.g., all pedestrians) from the segmentation while leaving the segmentation mostly unchanged otherwise.
Jan Hendrik Metzen, Chaithanya Kumar Mummadi, Thomas Brox, Volker Fischer 0003
ICCV2