Dvir Samuel

dblp:262/3701 · DBLP profile ↗
← Back
11ranked-venue papers
7as first author
11since 2021 · last 2026
0000-0003-3573-2220ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 5 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 4 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Generative modeling · 44% Representation and self-supervised learning · 15% Trustworthy machine learning · 8%
Computer graphics and multimedia
4 papers
Visual content generation and editing · 83% Multimedia analysis and retrieval · 13% Image and video processing · 4%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%

Topics — the 25 heaviest of 26, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
diffusion model
3.042025
Add-it: Training-Free Object Insertion in Images With Pretrained Diffusion Models · ICLR 2025
Where's Waldo: Diffusion Features For Personalized Segmentation and Retrieval · NeurIPS 2024
Generating Images of Rare Concepts Using Pre-trained Diffusion Models · AAAI 2024
Visual content generation and editing
image editing
1.722025
Add-it: Training-Free Object Insertion in Images With Pretrained Diffusion Models · ICLR 2025
Lightning-Fast Image Inversion and Editing for Text-to-Image Diffusion Models · ICLR 2025
Machine learning › Generative modeling › diffusion model
text-to-image generation
1.422024
Generating Images of Rare Concepts Using Pre-trained Diffusion Models · AAAI 2024
Norm-guided latent space exploration for text-to-image generation · NeurIPS 2023
Machine learning › Representation and self-supervised learning › foundation model representation
foundation model features
0.912025
EffoVPR: Effective Foundation Model Utilization for Visual Place Recognition · ICLR 2025
Computer vision › Image recognition and object detection
object detection
0.912025
Find your Needle: Small Object Image Retrieval via Multi-Object Attention Optimization · NeurIPS 2025
Machine learning › Generative modeling › diffusion model › image editing
text-guided image editing
0.912025
Add-it: Training-Free Object Insertion in Images With Pretrained Diffusion Models · ICLR 2025
Robotics › Robot navigation and mapping › place recognition
visual place recognition
0.912025
EffoVPR: Effective Foundation Model Utilization for Visual Place Recognition · ICLR 2025
Multimedia analysis and retrieval
image retrieval
0.912025
Find your Needle: Small Object Image Retrieval via Multi-Object Attention Optimization · NeurIPS 2025
Visual content generation and editing › image editing › image compositing
object insertion
0.912025
Add-it: Training-Free Object Insertion in Images With Pretrained Diffusion Models · ICLR 2025
Visual content generation and editing › image generation
text-to-image generation
0.912025
Lightning-Fast Image Inversion and Editing for Text-to-Image Diffusion Models · ICLR 2025
Visual content generation and editing › video editing
video layer decomposition
0.912025
OmnimatteZero: Fast Training-free Omnimatte with Pre-trained Video Diffusion Models · SIGGRAPH Asia 2025
Visual content generation and editing › video editing
video object removal
0.912025
OmnimatteZero: Fast Training-free Omnimatte with Pre-trained Video Diffusion Models · SIGGRAPH Asia 2025
Computer vision › Segmentation and scene understanding
personalized segmentation
0.812024
Where's Waldo: Diffusion Features For Personalized Segmentation and Retrieval · NeurIPS 2024
Information retrieval › image retrieval
instance retrieval
0.812024
Where's Waldo: Diffusion Features For Personalized Segmentation and Retrieval · NeurIPS 2024
Information retrieval
personalized search
0.812024
Where's Waldo: Diffusion Features For Personalized Segmentation and Retrieval · NeurIPS 2024
Machine learning › Learning paradigms › class imbalance
long-tailed learning
0.722023
Distributional Robustness Loss for Long-tail Learning · ICCV 2021
Norm-guided latent space exploration for text-to-image generation · NeurIPS 2023
Machine learning › Representation and self-supervised learning › latent space
latent space manipulation
0.712023
Norm-guided latent space exploration for text-to-image generation · NeurIPS 2023
Machine learning › Trustworthy machine learning › robustness
distributional robustness
0.512021
Distributional Robustness Loss for Long-tail Learning · ICCV 2021
Machine learning › Trustworthy machine learning
robustness
0.512021
Distributional Robustness Loss for Long-tail Learning · ICCV 2021
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning › self-supervised visual representation learning
self-supervised vision transformer
0.312025
EffoVPR: Effective Foundation Model Utilization for Visual Place Recognition · ICLR 2025
Image and video processing › video frame interpolation › interpolation
image interpolation
0.312025
Lightning-Fast Image Inversion and Editing for Text-to-Image Diffusion Models · ICLR 2025
Visual content generation and editing › video generation
video diffusion
0.312025
OmnimatteZero: Fast Training-free Omnimatte with Pre-trained Video Diffusion Models · SIGGRAPH Asia 2025
Machine learning › Deep learning architectures and training
data augmentation
0.212024
Generating Images of Rare Concepts Using Pre-trained Diffusion Models · AAAI 2024
Machine learning › Deep learning architectures and training › data augmentation
semantic augmentation
0.212024
Generating Images of Rare Concepts Using Pre-trained Diffusion Models · AAAI 2024
Machine learning › Transfer learning and domain adaptation
few-shot learning
0.212023
Norm-guided latent space exploration for text-to-image generation · NeurIPS 2023

Methods — techniques the papers use, named apart from their topics

diffusion model · 2.5object masks · 1.7multi-object pre-training · 1.7extended attention · 1.7attention-based feature extraction · 1.7vit pooling · 0.9temporal attention guidance · 0.9spatial attention guidance · 0.9self-attention features · 0.9reranking · 0.9numerical analysis · 0.9newton-raphson · 0.9latent arithmetic · 0.9DINOv2 · 0.9self-supervised foundation models · 0.8seed selection · 0.8
YearPublicationVenuePosition
2026 Story2Board: A Training-Free Approach for Expressive Visual Storytelling
abstract
Abstract We present Story2Board, a training‐free framework for expressive storyboard generation from natural language. Existing methods narrowly focus on subject identity, overlooking key aspects of visual storytelling such as spatial composition, background evolution, and narrative pacing. To address this, we introduce a lightweight consistency framework composed of two components: Latent Panel Anchoring, which preserves a shared character reference across panels, and Reciprocal Attention Value Mixing, which softly blends visual features between token pairs with strong reciprocal attention. Together, these mechanisms enhance coherence without architectural changes or fine‐tuning, enabling state‐of‐the‐art diffusion models to generate visually diverse yet consistent storyboards. To structure generation, we use an off‐the‐shelf language model to convert free‐form stories into grounded panel‐level prompts. To evaluate, we propose the Rich Storyboard Benchmark , a suite of open‐domain narratives designed to assess layout diversity and background‐grounded storytelling, in addition to consistency. We also introduce a new Scene Diversity metric that quantifies spatial and pose variation across storyboards. Our qualitative and quantitative results, as well as a user study, show that Story2Board produces more dynamic, coherent, and narratively engaging storyboards than existing baselines. Project page : https://daviddinkevich.github.io/Story2Board/
David Dinkevich, Matan Levy, Omri Avrahami, Dvir Samuel, Dani Lischinski
Comput. Graph. Forum4
2025 Lightning-Fast Image Inversion and Editing for Text-to-Image Diffusion Models
abstract
Diffusion inversion is the problem of taking an image and a text prompt that describes it and finding a noise latent that would generate the exact same image. Most current deterministic inversion techniques operate by approximately solving an implicit equation and may converge slowly or yield poor reconstructed images. We formulate the problem by finding the roots of an implicit equation and devlop a method to solve it efficiently. Our solution is based on Newton-Raphson (NR), a well-known technique in numerical analysis. We show that a vanilla application of NR is computationally infeasible while naively transforming it to a computationally tractable alternative tends to converge to out-of-distribution solutions, resulting in poor reconstruction and editing. We therefore derive an efficient guided formulation that fastly converges and provides high-quality reconstructions and editing. We showcase our method on real image editing with three popular open-sourced diffusion models: Stable Diffusion, SDXL-Turbo, and Flux with different deterministic schedulers. Our solution, **Guided Newton-Raphson Inversion**, inverts an image within 0.4 sec (on an A100 GPU) for few-step models (SDXL-Turbo and Flux.1), opening the door for interactive image editing. We further show improved results in image interpolation and generation of rare objects.
Dvir Samuel, Barak Meiri, Haggai Maron, Yoad Tewel, Nir Darshan, Shai Avidan, Gal Chechik, Rami Ben-Ari
ICLR1
2025 Add-it: Training-Free Object Insertion in Images With Pretrained Diffusion Models
abstract
Adding Object into images based on text instructions is a challenging task in semantic image editing, requiring a balance between preserving the original scene and seamlessly integrating the new object in a fitting location. Despite extensive efforts, existing models often struggle with this balance, particularly with finding a natural location for adding an object in complex scenes. We introduce Add-it, a training-free approach that extends diffusion models' attention mechanisms to incorporate information from three key sources: the scene image, the text prompt, and the generated image itself. Our weighted extended-attention mechanism maintains structural consistency and fine details while ensuring natural object placement. Without task-specific fine-tuning, Add-it achieves state-of-the-art results on both real and generated image insertion benchmarks, including our newly constructed "Additing Affordance Benchmark" for evaluating object placement plausibility, outperforming supervised methods. Human evaluations show that Add-it is preferred in over 80% of cases, and it also demonstrates improvements in various automated metrics.
Yoad Tewel, Rinon Gal, Dvir Samuel, Yuval Atzmon, Lior Wolf, Gal Chechik
ICLR3
2025 EffoVPR: Effective Foundation Model Utilization for Visual Place Recognition
abstract
The task of Visual Place Recognition (VPR) is to predict the location of a query image from a database of geo-tagged images. Recent studies in VPR have highlighted the significant advantage of employing pre-trained foundation models like DINOv2 for the VPR task. However, these models are often deemed inadequate for VPR without further fine-tuning on VPR-specific data. In this paper, we present an effective approach to harness the potential of a foundation model for VPR. We show that features extracted from self-attention layers can act as a powerful re-ranker for VPR, even in a zero-shot setting. Our method not only outperforms previous zero-shot approaches but also introduces results competitive with several supervised methods. We then show that a single-stage approach utilizing internal ViT layers for pooling can produce global features that achieve state-of-the-art performance, with impressive feature compactness down to 128D. Moreover, integrating our local foundation features for re-ranking further widens this performance gap. Our method also demonstrates exceptional robustness and generalization, setting new state-of-the-art performance, while handling challenging conditions such as occlusion, day-night transitions, and seasonal variations.
Issar Tzachor, Boaz Lerner, Matan Levy, Tal Berkovitz Shalev, Gavriel Habib, Dvir Samuel, Noam Korngut Zailer, Or Shimshi, Nir Darshan, Rami Ben-Ari
ICLR7
2025 Find your Needle: Small Object Image Retrieval via Multi-Object Attention Optimization
abstract
We address the challenge of Small Object Image Retrieval (SoIR), where the goal is to retrieve images containing a specific small object, in a cluttered scene. The key challenge in this setting is constructing a single image descriptor, for scalable and efficient search, that effectively represents all objects in the image. In this paper, we first analyze the limitations of existing methods on this challenging task and then introduce new benchmarks to support SoIR evaluation. Next, we introduce \ours (\oursMI), a novel retrieval framework which incorporates a dedicated multi-object pre-training phase. This is followed by a refinement process that leverages attention-based feature extraction with object masks, integrating them into a single unified image descriptor. Our \oursMI approach significantly outperforms existing retrieval methods and strong baselines, achieving notable improvements in both zero-shot and lightweight multi-object fine-tuning. We hope this work will pave the way and inspire further research to enhance retrieval performance for this highly practical task.
Matan Levy, Issar Tzachor, Dvir Samuel, Nir Darshan, Rami Ben-Ari
NeurIPS4
2025 OmnimatteZero: Fast Training-free Omnimatte with Pre-trained Video Diffusion Models
abstract
In Omnimatte, one aims to decompose a given video into semantically meaningful layers, including the background and individual objects along with their associated effects, such as shadows and reflections. Existing methods often require extensive training or costly self-supervised optimization. In this paper, we present OmnimatteZero, a training-free approach that leverages off-the-shelf pre-trained video diffusion models for omnimatte. It can remove objects from videos, extract individual object layers along with their effects, and composite those objects onto new videos. These are accomplished by adapting zero-shot image inpainting techniques for video object removal, a task they fail to handle effectively out-of-the-box. To overcome this, we introduce temporal and spatial attention guidance modules that steer the diffusion process for accurate object removal and temporally consistent background reconstruction. We further show that self-attention maps capture information about the object and its footprints and use them to inpaint the object’s effects, leaving a clean background. Additionally, through simple latent arithmetic, object layers can be isolated and recombined seamlessly with new video layers to produce new videos. Evaluations show that OmnimatteZero not only achieves superior performance in terms of background reconstruction but also sets a new record for the fastest Omnimatte approach, achieving real-time performance with minimal frame runtime. Project Page.
Dvir Samuel, Matan Levy, Nir Darshan, Gal Chechik, Rami Ben-Ari
SIGGRAPH Asia1
2024 Generating Images of Rare Concepts Using Pre-trained Diffusion Models
abstract
Text-to-image diffusion models can synthesize high quality images, but they have various limitations. Here we highlight a common failure mode of these models, namely, generating uncommon concepts and structured concepts like hand palms. We show that their limitation is partly due to the long-tail nature of their training data: web-crawled data sets are strongly unbalanced, causing models to under-represent concepts from the tail of the distribution. We characterize the effect of unbalanced training data on text-to-image models and offer a remedy. We show that rare concepts can be correctly generated by carefully selecting suitable generation seeds in the noise space, using a small reference set of images, a technique that we call SeedSelect. SeedSelect does not require retraining or finetuning the diffusion model. We assess the faithfulness, quality and diversity of SeedSelect in creating rare objects and generating complex formations like hand images, and find it consistently achieves superior performance. We further show the advantage of SeedSelect in semantic data augmentation. Generating semantically appropriate images can successfully improve performance in few-shot recognition benchmarks, for classes from the head and from the tail of the training data of diffusion models.
Dvir Samuel, Rami Ben-Ari, Simon Raviv, Nir Darshan, Gal Chechik
AAAI1
2024 Where's Waldo: Diffusion Features For Personalized Segmentation and Retrieval
abstract
Personalized retrieval and segmentation aim to locate specific instances within a dataset based on an input image and a short description of the reference instance. While supervised methods are effective, they require extensive labeled data for training. Recently, self-supervised foundation models have been introduced to these tasks showing comparable results to supervised methods. However, a significant flaw in these models is evident: they struggle to locate a desired instance when other instances within the same class are presented. In this paper, we explore text-to-image diffusion models for these tasks. Specifically, we propose a novel approach called PDM for Personalized Diffusion Features Matching, that leverages intermediate features of pre-trained text-to-image models for personalization tasks without any additional training. PDM demonstrates superior performance on popular retrieval and segmentation benchmarks, outperforming even supervised methods. We also highlight notable shortcomings in current instance and segmentation datasets and propose new benchmarks for these tasks.
Dvir Samuel, Rami Ben-Ari, Matan Levy, Nir Darshan, Gal Chechik
NeurIPS1
2023 Norm-guided latent space exploration for text-to-image generation
abstract
Text-to-image diffusion models show great potential in synthesizing a large variety of concepts in new compositions and scenarios. However, the latent space of initial seeds is still not well understood and its structure was shown to impact the generation of various concepts. Specifically, simple operations like interpolation and finding the centroid of a set of seeds perform poorly when using standard Euclidean or spherical metrics in the latent space. This paper makes the observation that, in current training procedures, diffusion models observed inputs with a narrow range of norm values. This has strong implications for methods that rely on seed manipulation for image generation, with applications to few-shot and long-tail learning tasks. To address this issue, we propose a novel method for interpolating between two seeds and demonstrate that it defines a new non-Euclidean metric that takes into account a norm-based prior on seeds. We describe a simple yet efficient algorithm for approximating this interpolation procedure and use it to further define centroids in the latent seed space. We show that our new interpolation and centroid techniques significantly enhance the generation of rare concept images. This further leads to state-of-the-art performance on few-shot and long-tail benchmarks, improving prior approaches in terms of generation speed, image quality, and semantic content.
Dvir Samuel, Rami Ben-Ari, Nir Darshan, Haggai Maron, Gal Chechik
NeurIPS1
2021 Distributional Robustness Loss for Long-tail Learning
abstract
Real-world data is often unbalanced and long-tailed, but deep models struggle to recognize rare classes in the presence of frequent classes. To address unbalanced data, most studies try balancing the data, the loss, or the classifier to reduce classification bias towards head classes. Far less attention has been given to the latent representations learned with unbalanced data. We show that the feature extractor part of deep networks suffers greatly from this bias. We propose a new loss based on robustness theory, which encourages the model to learn high-quality representations for both head and tail classes. While the general form of the robustness loss may be hard to compute, we further derive an easy-to-compute upper bound that can be minimized efficiently. This procedure reduces representation bias towards head classes in the feature space and achieves new SOTA results on CIFAR100-LT, ImageNet-LT, and iNaturalist long-tail benchmarks. We find that training with robustness increases recognition accuracy of tail classes while largely maintaining the accuracy of head classes. The new robustness loss can be combined with various classifier balancing techniques and can be applied to representations at several layers of the deep model.
Dvir Samuel, Gal Chechik
ICCV1
2021 From generalized zero-shot learning to long-tail with class descriptors
abstract
Real-world data is predominantly unbalanced and long-tailed, but deep models struggle to recognize rare classes in the presence of frequent classes. Often, classes can be accompanied by side information like textual descriptions, but it is not fully clear how to use them for learning with unbalanced long-tail data. Such descriptions have been mostly used in (Generalized) Zero-shot learning (ZSL), suggesting that ZSL with class descriptions may also be useful for long- tail distributions.We describe Dragon, a late-fusion architecture for long-tail learning with class descriptors. It learns to (1) correct the bias towards head classes on a sample- by-sample basis; and (2) fuse information from class- descriptions to improve the tail-class accuracy. We also introduce new benchmarks CUB-LT, SUN-LT, AWA-LT for long-tail learning with class-descriptions, building on existing learning-with-attributes datasets and a version of Imagenet-LT with class descriptors. Dragon outperforms state-of-the-art models on the new benchmark. It is also a new SoTA on existing benchmarks for GFSL with class descriptors (GFSL-d) and standard (vision-only) long-tailed learning ImageNet-LT, CIFAR-10, 100, and Places365-LT.
Dvir Samuel, Yuval Atzmon, Gal Chechik
WACV1