Rohit Gandikota

dblp:230/4602 · DBLP profile ↗
← Back
13ranked-venue papers
11as first author
11since 2021 · last 2026
0000-0003-0581-6810ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 5 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Finding the Unicorn within a Million Patches: A Material-Driven Journey into the Hidden Layers of a Diffusion Model
abstract
Today's AI models learn rich internal representations, such as the visual features inside diffusion models that produce generated images, and offer a new kind of material for co-creation. However, interfaces for creating with generative AI typically operate on the level of inputs and outputs, obscuring the material formation that unfolds in between. To address this gap, this pictorial documents a material-driven journey of designing an interface for interacting with the hidden layers of diffusion models.
Imke Grabe, Jaden Fiotto-Kaufman, Rohit Gandikota, David Bau, Tom Jenkins
Creativity & Cognition3
2026 Distilling Diversity and Control in Diffusion Models
abstract
Distilled diffusion models generate images in far fewer timesteps but suffer from reduced sample diversity when generating multiple outputs from the same prompt. To understand this phenomenon, we first investigate whether distillation damages concept representations by examining if the required diversity is properly learned. Surprisingly, distilled models retain the base model’s representational structure: control mechanisms like Concept Sliders and LoRAs transfer seamlessly without retraining, and Slider-Space analysis reveals distilled models possess variational directions needed for diversity yet fail to activate them. This redirects our investigation to understanding how the generation dynamics differ between base and distilled models. Using ${{\hat{\mathbf x}}_0}$ trajectory visualization, we discover distilled models commit to their final image structure almost immediately at the first timestep, while base models distribute structural decisions across many steps. To test whether this first-step commitment causes the diversity loss, we introduce diversity distillation, a hybrid approach using the base model for only the first critical timestep before switching to the distilled model. This single intervention restores sample diversity while maintaining computational efficiency. We provide both causal validation and theoretical support showing why the very first timestep concentrates the diversity bottleneck in distilled models. Our code and data are available at distillation.baulab.info
Rohit Gandikota, David Bau
WACV1
2025 SliderSpace: Decomposing the Visual Capabilities of Diffusion Models
abstract
We present SliderSpace, a framework for automatically decomposing the visual capabilities of diffusion models into controllable and human-understandable directions. Unlike existing control methods that require a user to specify attributes for each edit direction individually, SliderSpace discovers multiple interpretable and diverse directions simultaneously from a single text prompt. Each direction is trained as a low-rank adaptor, enabling compositional control and the discovery of surprising possibilities in the model's latent space. Through extensive experiments on state-of-the-art diffusion models, we demonstrate SliderSpace's effectiveness across three applications: concept decomposition, artistic style exploration, and diversity enhancement. Our quantitative evaluation shows that SliderSpace-discovered directions decompose the visual structure of model's knowledge effectively, offering insights into the latent capabilities encoded within diffusion models. User studies further validate that our method produces more diverse and useful variations compared to baselines. Our code, data and trained weights are available at https://sliderspace.baulab.info
Rohit Gandikota, Zongze Wu 0002, Richard Zhang 0001, David Bau, Eli Shechtman, Nicholas I. Kolkin
ICCV1
2025 Erasing Conceptual Knowledge from Language Models
abstract
In this work, we introduce Erasure of Language Memory (ELM), a principled approach to concept-level unlearning that operates by matching distributions defined by the model's own introspective classification capabilities. Our key insight is that effective unlearning should leverage the model's ability to evaluate its own knowledge, using the language model itself as a classifier to identify and reduce the likelihood of generating content related to undesired concepts. ELM applies this framework to create targeted low-rank updates that reduce generation probabilities for concept-specific content while preserving the model's broader capabilities. We demonstrate ELM's efficacy on biosecurity, cybersecurity, and literary domain erasure tasks. Comparative evaluation reveals that ELM-modified models achieve near-random performance on assessments targeting erased concepts, while simultaneously preserving generation coherence, maintaining benchmark performance on unrelated tasks, and exhibiting strong robustness to adversarial attacks.
Rohit Gandikota, Sheridan Feucht, Samuel Marks, David Bau
NeurIPS1
2025 When Are Concepts Erased From Diffusion Models?
abstract
In concept erasure, a model is modified to selectively prevent it from generating a target concept. Despite the rapid development of new methods, it remains unclear how thoroughly these approaches remove the target concept from the model. We begin by proposing two conceptual models for the erasure mechanism in diffusion models: (i) interfering with the model’s internal guidance processes, and (ii) reducing the unconditional likelihood of generating the target concept, potentially removing it entirely. To assess whether a concept has been truly erased from the model, we introduce a comprehensive suite of independent probing techniques: supplying visual context, modifying the diffusion trajectory, applying classifier guidance, and analyzing the model's alternative generations that emerge in place of the erased concept. Our results shed light on the value of exploring concept erasure robustness outside of adversarial text inputs, and emphasize the importance of comprehensive evaluations for erasure in diffusion models.
Nicky Kriplani, Rohit Gandikota, Minh Pham 0005, David Bau, Chinmay Hegde, Niv Cohen
NeurIPS3
2024 Concept Sliders: LoRA Adaptors for Precise Control in Diffusion Models
Rohit Gandikota, Joanna Materzynska, Tingrui Zhou, Antonio Torralba 0001, David Bau
ECCV (40)1
2024 Unified Concept Editing in Diffusion Models
abstract
Text-to-image models suffer from various safety issues that may limit their suitability for deployment. Previous methods have separately addressed individual issues of bias, copyright, and offensive content in text-to-image models. However, in the real world, all of these issues appear simultaneously in the same model. We present a method that tackles all issues with a single approach. Our method, Unified Concept Editing (UCE), edits the model without training using a closed-form solution, and scales seamlessly to concurrent edits on text-conditional diffusion models.We present scalable simultaneous debiasing, style erasure, and content moderation by editing text-to-image projections, and perform extensive experiments demonstrating improved efficacy and scalability over prior work. Our code is available at unified.baulab.info.
Rohit Gandikota, Hadas Orgad, Yonatan Belinkov, Joanna Materzynska, David Bau
WACV1
2023 Erasing Concepts from Diffusion Models
abstract
Motivated by concerns that large-scale diffusion models can produce undesirable output such as sexually explicit content or copyrighted artistic styles, we study erasure of specific concepts from diffusion model weights. We propose a fine-tuning method that can erase a visual concept from a pre-trained diffusion model, given only the name of the style and using negative guidance as a teacher. We benchmark our method against previous approaches that remove sexually explicit content and demonstrate its effectiveness, performing on par with Safe Latent Diffusion and censored training. To evaluate artistic style removal, we conduct experiments erasing five modern artists from the network and conduct a user study to assess the human perception of the removed styles. Unlike previous methods, our approach can remove concepts from a diffusion model permanently rather than modifying the output at the inference time, so it cannot be circumvented even if a user has access to model weights. Our code, data, and results are available at erasing.baulab.info.
Rohit Gandikota, Joanna Materzynska, Jaden Fiotto-Kaufman, David Bau
ICCV1
2022 Pro-DDPM: Progressive Growing of Variable Denoising Diffusion Probabilistic Models for Faster Convergence
Rohit Gandikota, Nicholas Brown
BMVC1
2022 Dismon-Gan: 24×7 All-Weather Optical Domain Surveillance Using Progressively Growing Adversarial Networks with Patch Discriminator
abstract
Cyclones and floods are significant disasters seen by south- Asian countries. With cloud occlusions on the affected regions, monitoring the disaster becomes tedious. Synthetic Aperture Radars (SAR) can penetrate through the clouds (due to their microwave frequencies) and image the area under- Being active sensors, they also can image around the clock. These properties enable the domain experts to use SAR images for disaster monitoring; however, even professionals find it challenging to interpret the data since the human eye is unfamiliar with the impact of distance-dependent imaging, signal intensities observed in the radar spectrum, and image features associated to speckle or postprocessing procedures. This manuscript exploits the valuable imaging properties of SAR images to propose a Generative Adversarial Network (GAN) to synthesize realistic and semantic optical images by conditioning them over the microwave satellite images. It en- the model to synthesize optical images of an affected area in all weather and around the clock conditions, making it a critical disaster monitoring assistance tool.
Rohit Gandikota, Deepak Mishra 0002
IGARSS1
2022 Pixel Noise Localization Algorithm for Indian Satellite Data Quality Control: A Novel Approach
abstract
The noise correction and image enhancements are done at the Data Processing Generation System (DPGS) at the National Remote Sensing Center (NRSC) satellite data production chain. Even after the noise correction algorithms, pixel noise is observed at the checkout by the Product Quality Control (PQC) system. These are detected manually at PQC and an alert is raised at the DPGS to recorrect the image. This process is time consuming due to the manual intervention and the application of noise correction algorithm on the entire image. This manuscript proposes a novel 2-stage, sliding window-based algorithm to automatically detect pixel noise in satellite images and localise the noise pixels so that the correction algorithms can be applied in a localised fashion. This local noise correction also restores the overall SNR of the image compared to global correction. This algorithm is realized and tested on data obtained from the Indian remote sensing (IRS) satellites like Cartosat-2S, Resourcesat-2/2A. The dataset is not open-sourced, and hence very minimal information is provided regarding the IRS data. However, we use Landsat-8 data to conduct a few analyses on the algorithm's performance. The role of patch size considered during the detection of pixel noise in satellite data is also analyzed.
Rohit Gandikota, ManjuSarma M
IGARSS1
2020 RTC-GAN: REAL-TIME CLASSIFICATION OF SATELLITE IMAGERY USING DEEP GENERATIVE ADVERSARIAL NETWORKS WITH INFUSED SPECTRAL INFORMATION
abstract
This paper implements a deep learning-based Convolutional Neural Network (CNN) with adversarial training and infused pixel information to classify multi-spectral data into 4 LULC classes and cloud. The network is capable of classifying the image on a real-time basis at acquisition time in pixel level by considering the various spectral band values at the pixel and a spatial region around the pixel to collect the spatial features. This way, both spatial information, and spectral information are considered to classify the image. This novel GAN architecture named RTC-GAN is generalized over all the satellites that have their sensors in and around standard NIR, R and G spectral bands while being able to classify the images in realtime. This network is realized and tested on data obtained from satellites Landsat 8 Sentinel2 and Indian Remote Sensing (IRS) satellites like Cartosat-2S, Resourcesat-2/2A. The dataset is not open-sourced and hence very minimal information is provided regarding the IRS data.
Rohit Gandikota, Radha Krishna K, Anupama Sharma, ManjuSarma M, Vinod M. Bothale
IGARSS1
2019 How You See Me: Understanding Convolutional Neural Networks
abstract
Convolutional Neural networks(CNN) are one of the most powerful tools in the present era of science. There has been a lot of research done to improve their performance and robustness while their internal working was left unexplored to much extent. They are often defined as black boxes that can map non-linear data effectively. This paper answers the question, “How does a CNN look at an image?”. Visual results are also provided to strongly support the proposed method. The proposed algorithm exploits the basic math behind CNN to backtrack the important pixels. This is a generic approach which can be applied to any architecture of a neural network. This doesn't require any additional training or architectural changes. In literature, few attempts have been made to explain how learning happens in CNN internally, by exploiting the convolution filter maps. This is a simple algorithm as it does not involve any cost functions, filter exploitation, gradient calculations or probability scores. Further, we demonstrate that the proposed scheme can be used in some important computer vision tasks such as object detection, salient region proposal, etc.
Rohit Gandikota, Deepak Mishra 0002
TENCON1