Amil Dravid

dblp:272/9123 · DBLP profile ↗
← Back
7ranked-venue papers
4as first author
7since 2021 · last 2025
0000-0001-6007-0690ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-author · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Generative modeling · 39% Trustworthy machine learning · 17% Deep learning architectures and training · 14%
Computer graphics and multimedia
1 paper
Visual content generation and editing · 100%

Topics — the 17 heaviest of 21, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning
interpretability
1.422025
Vision Transformers Don't Need Trained Registers · NeurIPS 2025
Visual Explanations for Convolutional Neural Networks via Latent Traversal of Generative Adversarial Networks (Student Abstract) · AAAI 2022
Machine learning › Deep learning architectures and training › transformer
register token
0.912025
Vision Transformers Don't Need Trained Registers · NeurIPS 2025
Machine learning › Deep learning architectures and training › transformer
vision transformer
0.912025
Vision Transformers Don't Need Trained Registers · NeurIPS 2025
Machine learning › Generative modeling
diffusion model
0.812024
Interpreting the Weight Space of Customized Diffusion Models · NeurIPS 2024
Machine learning › Representation and self-supervised learning
latent space
0.812024
Interpreting the Weight Space of Customized Diffusion Models · NeurIPS 2024
Machine learning › Generative modeling
latent space interpretation
0.812024
Interpreting the Weight Space of Customized Diffusion Models · NeurIPS 2024
Machine learning › Generative modeling › diffusion model › diffusion distillation
one-step generation
0.812024
Idempotent Generative Network · ICLR 2024
Machine learning › Generative modeling › diffusion model › diffusion model adaptation
personalized diffusion model
0.812024
Interpreting the Weight Space of Customized Diffusion Models · NeurIPS 2024
Computer vision › Face, body and person analysis › human pose estimation
3d pose estimation
0.712023
BKinD-3D: Self-Supervised 3D Keypoint Discovery from Multi-View Videos · CVPR 2023
Computer vision › 3D vision
pose estimation
0.712023
BKinD-3D: Self-Supervised 3D Keypoint Discovery from Multi-View Videos · CVPR 2023
Machine learning › Generative modeling
generative adversarial network
0.612022
Visual Explanations for Convolutional Neural Networks via Latent Traversal of Generative Adversarial Networks (Student Abstract) · AAAI 2022
Machine learning › Generative modeling
latent space exploration
0.612022
Visual Explanations for Convolutional Neural Networks via Latent Traversal of Generative Adversarial Networks (Student Abstract) · AAAI 2022
Machine learning › Trustworthy machine learning › interpretability
visual explanation
0.612022
Visual Explanations for Convolutional Neural Networks via Latent Traversal of Generative Adversarial Networks (Student Abstract) · AAAI 2022
Computer vision › Vision and language
vision-language model
0.312025
Vision Transformers Don't Need Trained Registers · NeurIPS 2025
Visual content generation and editing
image editing
0.212024
Interpreting the Weight Space of Customized Diffusion Models · NeurIPS 2024
Machine learning › Representation and self-supervised learning › representation matching › feature alignment
cross-model representation alignment
0.212023
Rosetta Neurons: Mining the Common Units in a Model Zoo · ICCV 2023
Machine learning › Trustworthy machine learning › interpretability › neural network interpretation
interpretability of generative models
0.212023
Rosetta Neurons: Mining the Common Units in a Model Zoo · ICCV 2023

Methods — techniques the papers use, named apart from their topics

inversion · 1.5fine-tuning · 1.5neuron analysis · 0.9activation shifting · 0.9idempotent operator training · 0.8distribution matching · 0.8joint length constraints · 0.7encoder-decoder architecture · 0.7dictionary learning · 0.73d volumetric heatmap · 0.7latent interpolation · 0.6Grad-CAM · 0.6GAN · 0.6
YearPublicationVenuePosition
2025 Vision Transformers Don't Need Trained Registers
abstract
We investigate the mechanism underlying a previously identified phenomenon in Vision Transformers -- the emergence of high-norm tokens that lead to noisy attention maps (Darcet et al., 2024). We observe that in multiple models (e.g., CLIP, DINOv2), a sparse set of neurons is responsible for concentrating high-norm activations on outlier tokens, leading to irregular attention patterns and degrading downstream visual processing. While the existing solution for removing these outliers involves retraining models from scratch with additional learned $\textit{register tokens}$, we use our findings to create a training-free approach to mitigate these artifacts. By shifting the high-norm activations from our discovered $\textit{register neurons}$ into an additional untrained token, we can mimic the effect of register tokens on a model already trained without registers. We demonstrate that our method produces cleaner attention and feature maps, enhances performance over base models across multiple downstream visual tasks, and achieves results comparable to models explicitly trained with register tokens. We then extend test-time registers to off-the-shelf vision-language models, yielding cleaner attention-based, text-to-image attribution. Finally, we outline a simple mathematical model that reflects the observed behavior of register neurons and high norm tokens. Our results suggest that test-time registers effectively take on the role of register tokens at test-time, offering a training-free solution for any pre-trained model released without them.
Nick Jiang, Amil Dravid, Alexei A. Efros, Yossi Gandelsman
NeurIPS2
2024 Idempotent Generative Network
abstract
We propose a new approach for generative modeling based on training a neural network to be idempotent. An idempotent operator is one that can be applied sequentially without changing the result beyond the initial application, namely $f(f(z))=f(z)$. The proposed model $f$ is trained to map a source distribution (e.g, Gaussian noise) to a target distribution (e.g. realistic images) using the following objectives: (1) Instances from the target distribution should map to themselves, namely $f(x)=x$. We define the target manifold as the set of all instances that $f$ maps to themselves. (2) Instances that form the source distribution should map onto the defined target manifold. This is achieved by optimizing the idempotence term, $f(f(z))=f(z)$ which encourages the range of $f(z)$ to be on the target manifold. Under ideal assumptions such a process provably converges to the target distribution. This strategy results in a model capable of generating an output in one step, maintaining a consistent latent space, while also allowing sequential applications for refinement. Additionally, we find that by processing inputs from both target and source distributions, the model adeptly projects corrupted or modified data back to the target manifold. This work is a first step towards a ``global projector'' that enables projecting any input into a target data distribution.
Assaf Shocher, Amil Dravid, Yossi Gandelsman, Inbar Mosseri, Michael Rubinstein, Alexei A. Efros
ICLR2
2024 Interpreting the Weight Space of Customized Diffusion Models
abstract
We investigate the space of weights spanned by a large collection of customized diffusion models. We populate this space by creating a dataset of over 60,000 models, each of which is a base model fine-tuned to insert a different person's visual identity. We model the underlying manifold of these weights as a subspace, which we term $\textit{weights2weights}$. We demonstrate three immediate applications of this space that result in new diffusion models -- sampling, editing, and inversion. First, sampling a set of weights from this space results in a new model encoding a novel identity. Next, we find linear directions in this space corresponding to semantic edits of the identity (e.g., adding a beard), resulting in a new model with the original identity edited. Finally, we show that inverting a single image into this space encodes a realistic identity into a model, even if the input image is out of distribution (e.g., a painting). We further find that these linear properties of the diffusion model weight space extend to other visual concepts. Our results indicate that the weight space of fine-tuned diffusion models can behave as an interpretable $\textit{meta}$-latent space producing new models.
Amil Dravid, Yossi Gandelsman, Kuan-Chieh Wang, Rameen Abdal, Gordon Wetzstein, Alexei A. Efros, Kfir Aberman
NeurIPS1
2023 BKinD-3D: Self-Supervised 3D Keypoint Discovery from Multi-View Videos
abstract
Quantifying motion in 3D is important for studying the behavior of humans and other animals, but manual pose annotations are expensive and time-consuming to obtain. Self-supervised keypoint discovery is a promising strategy for estimating 3D poses without annotations. However, current keypoint discovery approaches commonly process single 2D views and do not operate in the 3D space. We propose a new method to perform self-supervised keypoint discovery in 3D from multi-view videos of behaving agents, without any keypoint or bounding box supervision in 2D or 3D. Our method, BKinD-3D, uses an encoder-decoder architecture with a 3D volumetric heatmap, trained to reconstruct spatiotemporal differences across multiple views, in addition to joint length constraints on a learned 3D skeleton of the subject. In this way, we discover keypoints without requiring manual supervision in videos of humans and rats, demonstrating the potential of 3D keypoint discovery for studying behavior.
Jennifer J. Sun, Lili Karashchuk, Amil Dravid, Serim Ryou, Sonia Fereidooni, John C. Tuthill, Aggelos K. Katsaggelos, Bingni W. Brunton, Georgia Gkioxari, Ann Kennedy, Yisong Yue, Pietro Perona
CVPR3
2023 Rosetta Neurons: Mining the Common Units in a Model Zoo
abstract
Do different neural networks, trained for various vision tasks, share some common representations? In this paper, we demonstrate the existence of common features we call "Rosetta Neurons" across a range of models with different architectures, different tasks (generative and discriminative), and different types of supervision (class-supervised, text-supervised, self-supervised). We present an algorithm for mining a dictionary of Rosetta Neurons across several popular vision models: Class Supervised-ResNet50, DINO-ResNet50, DINO-ViT, MAE, CLIP-ResNet50, Big-GAN, StyleGAN-2, StyleGAN-XL. Our findings suggest that certain visual concepts and structures are inherently embedded in the natural world and can be learned by different models regardless of the specific task or architecture, and without the use of semantic labels. We can visualize shared concepts directly due to generative models included in our analysis. The Rosetta Neurons facilitate model-to-model translation enabling various inversion-based manipulations, including cross-class alignments, shifting, zooming, and more, without the need for specialized training.
Amil Dravid, Yossi Gandelsman, Alexei A. Efros, Assaf Shocher
ICCV1
2022 Visual Explanations for Convolutional Neural Networks via Latent Traversal of Generative Adversarial Networks (Student Abstract)
abstract
Lack of explainability in artificial intelligence, specifically deep neural networks, remains a bottleneck for implementing models in practice. Popular techniques such as Gradient-weighted Class Activation Mapping (Grad-CAM) provide a coarse map of salient features in an image, which rarely tells the whole story of what a convolutional neural network(CNN) learned. Using COVID-19 chest X-rays, we present a method for interpreting what a CNN has learned by utilizing Generative Adversarial Networks (GANs). Our GAN framework disentangles lung structure from COVID-19 features. Using this GAN, we can visualize the transition of a pair of COVID negative lungs in a chest radiograph to a COVID positive pair by interpolating in the latent space of the GAN, which provides fine-grained visualization of how the CNN responds to varying features within the lungs.
Amil Dravid, Aggelos K. Katsaggelos
AAAI1
2022 Investigating the Potential of Auxiliary-Classifier Gans for Image Classification in Low Data Regimes
abstract
Generative Adversarial Networks (GANs) have shown promise in augmenting datasets and boosting convolutional neural network (CNN) performance on image classification tasks. But they introduce more hyperparameters to tune as well as the need for additional time and computational power to train, supplementary to the CNN. In this work, we examine the potential for Auxiliary-Classifier GANs (AC-GANs) as a ’one-stop-shop’ architecture for image classification, particularly in low data regimes. Additionally, we explore modifications to the typical AC-GAN framework, changing the generator’s latent space sampling scheme and employing a Wasserstein loss with gradient penalty to stabilize the simultaneous training of image synthesis and classification. Through experiments on images of varying resolutions and complexity, we demonstrate that AC-GANs show promise in image classification, achieving competitive performance with standard CNNs. These methods can be employed as an ’all-in-one’ framework with particular utility in the absence of large amounts of training data.
Amil Dravid, Florian Schiffers, Yunan Wu, Oliver Cossairt, Aggelos K. Katsaggelos
ICASSP1