EDBT 2026 Demo / reviewers in the wild / expert
Assaf Shocher
dblp:211/8006
· DBLP profile ↗
13ranked-venue papers
4as first author
8since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 4 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 4 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
11 papers |
Generative modeling · 46% Trustworthy machine learning · 24% Representation and self-supervised learning · 12% | |
| Computer graphics and multimedia
7 papers |
Image and video processing · 57% Visual content generation and editing · 24% Image and video coding · 19% |
Topics — the 30 heaviest of 37, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling
generative adversarial network |
1.4 | 4 | 2022 | Semantic Pyramid for Image Generation · CVPR 2020 Blind Super-Resolution Kernel Estimation using an Internal-GAN · NeurIPS 2019 InGAN: Capturing and Retargeting the "DNA" of a Natural Image · ICCV 2019 |
Machine learning › Trustworthy machine learning › uncertainty estimation
confidence estimation |
0.9 | 1 | 2025 | IT3: Idempotent Test-Time Training · ICML 2025 |
Machine learning › Transfer learning and domain adaptation
test-time adaptation |
0.9 | 1 | 2025 | IT3: Idempotent Test-Time Training · ICML 2025 |
Machine learning › Trustworthy machine learning
uncertainty estimation |
0.9 | 1 | 2025 | IT3: Idempotent Test-Time Training · ICML 2025 |
Image and video coding
video compression |
0.9 | 1 | 2025 | RL-RC-DoT: A Block-level RL agent for Task-Aware Video Compression · CVPR 2025 |
Machine learning › Trustworthy machine learning › interpretability
concept-based explanation |
0.8 | 1 | 2024 | The Hidden Language of Diffusion Models · ICLR 2024 |
Machine learning › Generative modeling
diffusion model |
0.8 | 1 | 2024 | The Hidden Language of Diffusion Models · ICLR 2024 |
Machine learning › Trustworthy machine learning
interpretability |
0.8 | 1 | 2024 | The Hidden Language of Diffusion Models · ICLR 2024 |
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning › masked modeling
masked image modeling |
0.8 | 1 | 2024 | Stochastic positional embeddings improve masked image modeling · ICML 2024 |
Machine learning › Generative modeling › diffusion model › diffusion distillation
one-step generation |
0.8 | 1 | 2024 | Idempotent Generative Network · ICLR 2024 |
Machine learning › Deep learning architectures and training
positional encoding |
0.8 | 1 | 2024 | Stochastic positional embeddings improve masked image modeling · ICML 2024 |
Machine learning › Generative modeling › diffusion model › text-to-image generation
text-to-image diffusion model |
0.8 | 1 | 2024 | The Hidden Language of Diffusion Models · ICLR 2024 |
Machine learning › Generative modeling
video generation |
0.6 | 1 | 2022 | Diverse Generation from a Single Video Made Possible · ECCV (17) 2022 |
Machine learning › Generative modeling › video generation › video frame synthesis
video synthesis |
0.6 | 1 | 2022 | Diverse Generation from a Single Video Made Possible · ECCV (17) 2022 |
Visual content generation and editing › image generation
single-image generation |
0.6 | 1 | 2022 | Drop the GAN: In Defense of Patches Nearest Neighbors as Single Image Generative Models · CVPR 2022 |
Machine learning › Generative modeling
image generation |
0.4 | 1 | 2020 | Semantic Pyramid for Image Generation · CVPR 2020 |
Machine learning › Generative modeling › image generation › conditional image synthesis
semantic image synthesis |
0.4 | 1 | 2020 | Semantic Pyramid for Image Generation · CVPR 2020 |
Machine learning › Deep learning architectures and training › deep generative model
deep image prior |
0.4 | 1 | 2019 | "Double-DIP": Unsupervised Image Decomposition via Coupled Deep-Image-Priors · CVPR 2019 |
Image and video processing › super-resolution › image super-resolution
blind super-resolution |
0.4 | 1 | 2019 | Blind Super-Resolution Kernel Estimation using an Internal-GAN · NeurIPS 2019 |
Image and video processing
image decomposition |
0.4 | 1 | 2019 | "Double-DIP": Unsupervised Image Decomposition via Coupled Deep-Image-Priors · CVPR 2019 |
Visual content generation and editing
image retargeting |
0.4 | 1 | 2019 | InGAN: Capturing and Retargeting the "DNA" of a Natural Image · ICCV 2019 |
Image and video processing › super-resolution
image super-resolution |
0.4 | 1 | 2019 | Blind Super-Resolution Kernel Estimation using an Internal-GAN · NeurIPS 2019 |
Image and video processing › image decomposition › image separation
layer separation |
0.4 | 1 | 2019 | "Double-DIP": Unsupervised Image Decomposition via Coupled Deep-Image-Priors · CVPR 2019 |
Image and video processing
super-resolution |
0.3 | 1 | 2018 | "Zero-Shot" Super-Resolution Using Deep Internal Learning · CVPR 2018 |
Image and video processing › super-resolution
unsupervised image super-resolution |
0.3 | 1 | 2018 | "Zero-Shot" Super-Resolution Using Deep Internal Learning · CVPR 2018 |
Image and video processing › super-resolution › image super-resolution
zero-shot super-resolution |
0.3 | 1 | 2018 | "Zero-Shot" Super-Resolution Using Deep Internal Learning · CVPR 2018 |
Computer vision › Image recognition and object detection
image classification |
0.3 | 1 | 2025 | IT3: Idempotent Test-Time Training · ICML 2025 |
Machine learning › Generative modeling › diffusion model
text-to-image generation |
0.2 | 1 | 2024 | The Hidden Language of Diffusion Models · ICLR 2024 |
Machine learning › Deep learning architectures and training › transformer
vision transformer |
0.2 | 1 | 2024 | Stochastic positional embeddings improve masked image modeling · ICML 2024 |
Machine learning › Representation and self-supervised learning › representation matching › feature alignment
cross-model representation alignment |
0.2 | 1 | 2023 | Rosetta Neurons: Mining the Common Units in a Model Zoo · ICCV 2023 |
Methods — techniques the papers use, named apart from their topics
test-time training · 0.9reinforcement learning · 0.9idempotence · 0.9stochastic positional embedding · 0.8idempotent operator training · 0.8gaussian distribution · 0.8distribution matching · 0.8diffusion model interpretation · 0.8concept decomposition · 0.8deep internal learning · 0.7model-to-model translation · 0.7dictionary learning · 0.7patch nearest neighbor · 0.6generative adversarial network · 0.6pre-trained classification features · 0.4deep feature pyramid · 0.4internal patch distribution learning · 0.4deep image prior · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | RL-RC-DoT: A Block-level RL agent for Task-Aware Video CompressionabstractVideo encoders optimize compression for human perception by minimizing reconstruction error under bit-rate constraints. In many modern applications such as autonomous driving, an overwhelming majority of videos serve as input for AI systems performing tasks like object recognition or segmentation, rather than being watched by humans. It is therefore useful to optimize the encoder for a downstream task instead of for perceptual image quality. However, a major challenge is how to combine such downstream optimization with existing standard video encoders, which are highly efficient and popular. Here, we address this challenge by controlling the Quantization Parameters (QPs) at the macro-block level to optimize the downstream task. This granular control allows us to prioritize encoding for taskrelevant regions within each frame. We formulate this optimization problem as a Reinforcement Learning (RL) task, where the agent learns to balance long-term implications of choosing QPs on both task performance and bit-rate constraints. Notably, our policy does not require the downstream task as an input during inference, making it suitable for streaming applications and edge devices such as vehicles. We demonstrate significant improvements in two tasks, car detection, and ROI (saliency) encoding. Our approach improves task performance for a given bit rate compared to traditional task agnostic encoding methods, paving the way for more efficient task-aware video compression. Uri Gadot, Assaf Shocher, Shie Mannor, Gal Chechik, Assaf Hallak |
CVPR | 2 |
| 2025 | IT3: Idempotent Test-Time TrainingabstractDeep learning models often struggle when deployed in real-world settings due to distribution shifts between training and test data. While existing approaches like domain adaptation and test-time training (TTT) offer partial solutions, they typically require additional data or domain-specific auxiliary tasks. We present Idempotent Test-Time Training (IT3), a novel approach that enables on-the-fly adaptation to distribution shifts using only the current test instance, without any auxiliary task design. Our key insight is that enforcing idempotence---where repeated applications of a function yield the same result---can effectively replace domain-specific auxiliary tasks used in previous TTT methods. We theoretically connect idempotence to prediction confidence and demonstrate that minimizing the distance between successive applications of our model during inference leads to improved out-of-distribution performance. Extensive experiments across diverse domains (including image classification, aerodynamics prediction, and aerial segmentation) and architectures (MLPs, CNNs, GNNs) show that IT3 consistently outperforms existing approaches while being simpler and more widely applicable. Our results suggest that idempotence provides a universal principle for test-time adaptation that generalizes across domains and architectures. Nikita Durasov, Assaf Shocher, Doruk Öner, Gal Chechik, Alexei A. Efros, Pascal Fua |
ICML | 2 |
| 2024 | The Hidden Language of Diffusion ModelsabstractText-to-image diffusion models have demonstrated an unparalleled ability to generate high-quality, diverse images from a textual prompt. However, the internal representations learned by these models remain an enigma. In this work, we present Conceptor, a novel method to interpret the internal representation of a textual concept by a diffusion model. This interpretation is obtained by decomposing the concept into a small set of human-interpretable textual elements. Applied over the state-of-the-art Stable Diffusion model, Conceptor reveals non-trivial structures in the representations of concepts. For example, we find surprising visual connections between concepts, that transcend their textual semantics. We additionally discover concepts that rely on mixtures of exemplars, biases, renowned artistic styles, or a simultaneous fusion of multiple meanings of the concept.
Through a large battery of experiments, we demonstrate Conceptor's ability to provide meaningful, robust, and faithful decompositions for a wide variety of abstract, concrete, and complex textual concepts, while allowing to naturally connect each decomposition element to its corresponding visual impact on the generated images. Hila Chefer, Oran Lang, Mor Geva, Volodymyr Polosukhin, Assaf Shocher, Michal Irani, Inbar Mosseri, Lior Wolf |
ICLR | 5 |
| 2024 | Idempotent Generative NetworkabstractWe propose a new approach for generative modeling based on training a neural network to be idempotent. An idempotent operator is one that can be applied sequentially without changing the result beyond the initial application, namely $f(f(z))=f(z)$. The proposed model $f$ is trained to map a source distribution (e.g, Gaussian noise) to a target distribution (e.g. realistic images) using the following objectives:
(1) Instances from the target distribution should map to themselves, namely $f(x)=x$. We define the target manifold as the set of all instances that $f$ maps to themselves.
(2) Instances that form the source distribution should map onto the defined target manifold. This is achieved by optimizing the idempotence term, $f(f(z))=f(z)$ which encourages the range of $f(z)$ to be on the target manifold. Under ideal assumptions such a process provably converges to the target distribution. This strategy results in a model capable of generating an output in one step, maintaining a consistent latent space, while also allowing sequential applications for refinement. Additionally, we find that by processing inputs from both target and source distributions, the model adeptly projects corrupted or modified data back to the target manifold. This work is a first step towards a ``global projector'' that enables projecting any input into a target data distribution. Assaf Shocher, Amil Dravid, Yossi Gandelsman, Inbar Mosseri, Michael Rubinstein, Alexei A. Efros |
ICLR | 1 |
| 2024 | Stochastic positional embeddings improve masked image modelingabstractMasked Image Modeling (MIM) is a promising self-supervised learning approach that enables learning from unlabeled images. Despite its recent success, learning good representations through MIM remains challenging because it requires predicting the right semantic content in accurate locations. For example, given an incomplete picture of a dog, we can guess that there is a tail, but we cannot determine its exact location. In this work, we propose to incorporate location uncertainty to MIM by using stochastic positional embeddings (StoP). Specifically, we condition the model on stochastic masked token positions drawn from a gaussian distribution. We show that using StoP reduces overfitting to location features and guides the model toward learning features that are more robust to location uncertainties. Quantitatively, using StoP improves downstream MIM performance on a variety of downstream tasks. For example, linear probing on ImageNet using ViT-B is improved by $+1.7\%$, and by $2.5\%$ for ViT-H using 1% of the data. Amir Bar, Florian Bordes, Assaf Shocher, Mido Assran, Pascal Vincent, Nicolas Ballas, Trevor Darrell, Amir Globerson, Yann LeCun |
ICML | 3 |
| 2023 | Rosetta Neurons: Mining the Common Units in a Model ZooabstractDo different neural networks, trained for various vision tasks, share some common representations? In this paper, we demonstrate the existence of common features we call "Rosetta Neurons" across a range of models with different architectures, different tasks (generative and discriminative), and different types of supervision (class-supervised, text-supervised, self-supervised). We present an algorithm for mining a dictionary of Rosetta Neurons across several popular vision models: Class Supervised-ResNet50, DINO-ResNet50, DINO-ViT, MAE, CLIP-ResNet50, Big-GAN, StyleGAN-2, StyleGAN-XL. Our findings suggest that certain visual concepts and structures are inherently embedded in the natural world and can be learned by different models regardless of the specific task or architecture, and without the use of semantic labels. We can visualize shared concepts directly due to generative models included in our analysis. The Rosetta Neurons facilitate model-to-model translation enabling various inversion-based manipulations, including cross-class alignments, shifting, zooming, and more, without the need for specialized training. Amil Dravid, Yossi Gandelsman, Alexei A. Efros, Assaf Shocher |
ICCV | 4 |
| 2022 | Drop the GAN: In Defense of Patches Nearest Neighbors as Single Image Generative ModelsabstractImage manipulation dates back long before the deep learning era. The classical prevailing approaches were based on maximizing patch similarity between the input and generated output. Recently, single-image GANs were introduced as a superior and more sophisticated solution to image manipulation tasks. Moreover, they offered the opportunity not only to manipulate a given image, but also to generate a large and diverse set of different outputs from a single natural image. This gave rise to new tasks, which are considered “GAN-only”. However, despite their impressiveness, single-image GANs require long training time (usually hours) for each image and each task and often suffer from visual artifacts. In this paper we revisit the classical patch-based methods, and show that - unlike previously believed - classical methods can be adapted to tackle these novel “GAN-only” tasks. Moreover, they do so better and faster than single-image GAN-based methods. More specifically, we show that: (i) by introducing slight modifications, classical patch-based methods are able to unconditionally generate diverse images based on a single natural image; (ii) the generated output visual quality exceeds that of single-image GANs by a large margin (confirmed both quantitatively and qualitatively); (iii) they are orders of magnitude faster (runtime reduced from hours to seconds).22This project received funding from the European Research Council (ERC) under the European Union's Horizon 2020 research and innovation programme (grant agreement No 788535), and the Carolito Stiftung. Dr Bagon is a Robin Chemers Neustein AI Fellow. Niv Granot, Ben Feinstein, Assaf Shocher, Shai Bagon, Michal Irani |
CVPR | 3 |
| 2022 | Diverse Generation from a Single Video Made Possible
Niv Haim, Ben Feinstein, Niv Granot, Assaf Shocher, Shai Bagon, Tali Dekel, Michal Irani |
ECCV (17) | 4 |
| 2020 | Semantic Pyramid for Image GenerationabstractWe present a novel GAN-based model that utilizes the space of deep features learned by a pre-trained classification model. Inspired by classical image pyramid representations, we construct our model as a Semantic Generation Pyramid -- a hierarchical framework which leverages the continuum of semantic information encapsulated in such deep features; this ranges from low level information contained in fine features to high level, semantic information contained in deeper features. More specifically, given a set of features extracted from a reference image, our model generates diverse image samples, each with matching features at each semantic level of the classification model. We demonstrate that our model results in a versatile and flexible framework that can be used in various classic and novel image generation tasks. These include: generating images with a controllable extent of semantic similarity to a reference image, and different manipulation tasks such as semantically-controlled inpainting and compositing; all achieved with the same model, with no further training. Assaf Shocher, Yossi Gandelsman, Inbar Mosseri, Michal Yarom, Michal Irani, William T. Freeman, Tali Dekel |
CVPR | 1 |
| 2019 | "Double-DIP": Unsupervised Image Decomposition via Coupled Deep-Image-PriorsabstractMany seemingly unrelated computer vision tasks can be viewed as a special case of image decomposition into separate layers. For example, image segmentation (separation into foreground and background layers); transparent layer separation (into reflection and transmission layers); Image dehazing (separation into a clear image and a haze map), and more. In this paper we propose a unified framework for unsupervised layer decomposition of a single image, based on coupled "Deep-image-Prior" (DIP) networks. It was shown [Ulyanov et al] that the structure of a single DIP generator network is sufficient to capture the low-level statistics of a single image. We show that coupling multiple such DIPs provides a powerful tool for decomposing images into their basic components, for a wide variety of applications. This capability stems from the fact that the internal statistics of a mixture of layers is more complex than the statistics of each of its individual components. We show the power of this approach for Image-Dehazing, Fg/Bg Segmentation, Watermark-Removal, Transparency Separation in images and video, and more. These capabilities are achieved in a totally unsupervised way, with no training examples other than the input image/video itself. Yossi Gandelsman, Assaf Shocher, Michal Irani |
CVPR | 2 |
| 2019 | InGAN: Capturing and Retargeting the "DNA" of a Natural ImageabstractGenerative Adversarial Networks (GANs) typically learn a distribution of images in a large image dataset, and are then able to generate new images from this distribution. However, each natural image has its own internal statistics, captured by its unique distribution of patches. In this paper we propose an "Internal GAN'' (InGAN) - an image-specific GAN - which trains on a single input image and learns its internal distribution of patches. It is then able to synthesize a plethora of new natural images of significantly different sizes, shapes and aspect-ratios - all with the same internal patch-distribution (same "DNA'') as the input image. In particular, despite large changes in global size/shape of the image, all elements inside the image maintain their local size/shape. InGAN is fully unsupervised, requiring no additional data other than the input image itself. Once trained on the input image, it can remap the input to any size or shape in a single feedforward pass, while preserving the same internal patch distribution. InGAN provides a unified framework for a variety of tasks, bridging the gap between textures and natural images. Assaf Shocher, Shai Bagon, Phillip Isola, Michal Irani |
ICCV | 1 |
| 2019 | Blind Super-Resolution Kernel Estimation using an Internal-GANabstractSuper resolution (SR) methods typically assume that the low-resolution (LR) image was downscaled from the unknown high-resolution (HR) image by a fixed `ideal’ downscaling kernel (e.g. Bicubic downscaling). However, this is rarely the case in real LR images, in contrast to synthetically generated SR datasets. When the assumed downscaling kernel deviates from the true one, the performance of SR methods significantly deteriorates. This gave rise to Blind-SR - namely, SR when the downscaling kernel (SR-kernel’’) is unknown. It was further shown that the true SR-kernel is the one that maximizes the recurrence of patches across scales of the LR image. In this paper we show how this powerful cross-scale recurrence property can be realized using Deep Internal Learning. We introduceKernelGAN’’, an image-specific Internal-GAN, which trains solely on the LR test image at test time, and learns its internal distribution of patches. Its Generator is trained to produce a downscaled version of the LR test image, such that its Discriminator cannot distinguish between the patch distribution of the downscaled image, and the patch distribution of the original LR image. The Generator, once trained, constitutes the downscaling operation with the correct image-specific SR-kernel. KernelGAN is fully unsupervised, requires no training data other than the input image itself, and leads to state-of-the-art results in Blind-SR when plugged into existing SR algorithms. Sefi Bell-Kligler, Assaf Shocher, Michal Irani |
NeurIPS | 2 |
| 2018 | "Zero-Shot" Super-Resolution Using Deep Internal LearningabstractDeep Learning has led to a dramatic leap in SuperResolution (SR) performance in the past few years. However, being supervised, these SR methods are restricted to specific training data, where the acquisition of the low-resolution (LR) images from their high-resolution (HR) counterparts is predetermined (e.g., bicubic downscaling), without any distracting artifacts (e.g., sensor noise, image compression, non-ideal PSF, etc). Real LR images, however, rarely obey these restrictions, resulting in poor SR results by SotA (State of the Art) methods. In this paper we introduce "Zero-Shot" SR, which exploits the power of Deep Learning, but does not rely on prior training. We exploit the internal recurrence of information inside a single image, and train a small image-specific CNN at test time, on examples extracted solely from the input image itself. As such, it can adapt itself to different settings per image. This allows to perform SR of real old photos, noisy images, biological data, and other images where the acquisition process is unknown or non-ideal. On such images, our method outperforms SotA CNN-based SR methods, as well as previous unsupervised SR methods. To the best of our knowledge, this is the first unsupervised CNN-based SR method. Assaf Shocher, Michal Irani |
CVPR | 1 |