Nir Darshan

dblp:342/7307 · DBLP profile ↗
← Back
11ranked-venue papers
0as first author
11since 2021 · last 2025
0009-0004-0652-9519ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
Generative modeling · 34% Representation and self-supervised learning · 17% Face, body and person analysis · 16%
Databases, data mining, and information retrieval
3 papers
Information retrieval · 100%
Computer graphics and multimedia
3 papers
Visual content generation and editing · 77% Multimedia analysis and retrieval · 18% Image and video processing · 5%

Topics — the 28 heaviest of 30, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
diffusion model
2.232024
Where's Waldo: Diffusion Features For Personalized Segmentation and Retrieval · NeurIPS 2024
Generating Images of Rare Concepts Using Pre-trained Diffusion Models · AAAI 2024
Norm-guided latent space exploration for text-to-image generation · NeurIPS 2023
Machine learning › Generative modeling › diffusion model
text-to-image generation
1.422024
Generating Images of Rare Concepts Using Pre-trained Diffusion Models · AAAI 2024
Norm-guided latent space exploration for text-to-image generation · NeurIPS 2023
Machine learning › Representation and self-supervised learning › foundation model representation
foundation model features
0.912025
EffoVPR: Effective Foundation Model Utilization for Visual Place Recognition · ICLR 2025
Computer vision › Face, body and person analysis › gait analysis
gait recognition
0.912025
CarGait: Cross-Attention Based Re-ranking for Gait Recognition · ICCV 2025
Computer vision › Image recognition and object detection
object detection
0.912025
Find your Needle: Small Object Image Retrieval via Multi-Object Attention Optimization · NeurIPS 2025
Robotics › Robot navigation and mapping › place recognition
visual place recognition
0.912025
EffoVPR: Effective Foundation Model Utilization for Visual Place Recognition · ICLR 2025
Visual content generation and editing
image editing
0.912025
Lightning-Fast Image Inversion and Editing for Text-to-Image Diffusion Models · ICLR 2025
Multimedia analysis and retrieval
image retrieval
0.912025
Find your Needle: Small Object Image Retrieval via Multi-Object Attention Optimization · NeurIPS 2025
Visual content generation and editing › image generation
text-to-image generation
0.912025
Lightning-Fast Image Inversion and Editing for Text-to-Image Diffusion Models · ICLR 2025
Visual content generation and editing › video editing
video layer decomposition
0.912025
OmnimatteZero: Fast Training-free Omnimatte with Pre-trained Video Diffusion Models · SIGGRAPH Asia 2025
Visual content generation and editing › video editing
video object removal
0.912025
OmnimatteZero: Fast Training-free Omnimatte with Pre-trained Video Diffusion Models · SIGGRAPH Asia 2025
Computer vision › Segmentation and scene understanding
personalized segmentation
0.812024
Where's Waldo: Diffusion Features For Personalized Segmentation and Retrieval · NeurIPS 2024
Information retrieval › image retrieval
composed image retrieval
0.812024
Data Roaming and Quality Assessment for Composed Image Retrieval · AAAI 2024
Information retrieval › image retrieval
instance retrieval
0.812024
Where's Waldo: Diffusion Features For Personalized Segmentation and Retrieval · NeurIPS 2024
Information retrieval
multimedia analysis and retrieval
0.812024
Data Roaming and Quality Assessment for Composed Image Retrieval · AAAI 2024
Information retrieval
personalized search
0.812024
Where's Waldo: Diffusion Features For Personalized Segmentation and Retrieval · NeurIPS 2024
Machine learning › Representation and self-supervised learning › latent space
latent space manipulation
0.712023
Norm-guided latent space exploration for text-to-image generation · NeurIPS 2023
Information retrieval › interactive information retrieval › conversational information seeking › conversational search
conversational image retrieval
0.712023
Chatting Makes Perfect: Chat-based Image Retrieval · NeurIPS 2023
Information retrieval
image retrieval
0.712023
Chatting Makes Perfect: Chat-based Image Retrieval · NeurIPS 2023
Information retrieval
interactive information retrieval
0.712023
Chatting Makes Perfect: Chat-based Image Retrieval · NeurIPS 2023
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning › self-supervised visual representation learning
self-supervised vision transformer
0.312025
EffoVPR: Effective Foundation Model Utilization for Visual Place Recognition · ICLR 2025
Image and video processing › video frame interpolation › interpolation
image interpolation
0.312025
Lightning-Fast Image Inversion and Editing for Text-to-Image Diffusion Models · ICLR 2025
Visual content generation and editing › video generation
video diffusion
0.312025
OmnimatteZero: Fast Training-free Omnimatte with Pre-trained Video Diffusion Models · SIGGRAPH Asia 2025
Computer vision › Vision and language
cross-modal retrieval
0.212024
Data Roaming and Quality Assessment for Composed Image Retrieval · AAAI 2024
Machine learning › Deep learning architectures and training
data augmentation
0.212024
Generating Images of Rare Concepts Using Pre-trained Diffusion Models · AAAI 2024
Machine learning › Deep learning architectures and training › data augmentation
semantic augmentation
0.212024
Generating Images of Rare Concepts Using Pre-trained Diffusion Models · AAAI 2024
Machine learning › Transfer learning and domain adaptation
few-shot learning
0.212023
Norm-guided latent space exploration for text-to-image generation · NeurIPS 2023
Machine learning › Learning paradigms › class imbalance
long-tailed learning
0.212023
Norm-guided latent space exploration for text-to-image generation · NeurIPS 2023

Methods — techniques the papers use, named apart from their topics

cross-attention · 2.4object masks · 1.7multi-object pre-training · 1.7attention-based feature extraction · 1.7auxiliary task training · 1.5zero-shot image inpainting · 0.9vit pooling · 0.9video diffusion model · 0.9temporal attention guidance · 0.9spatial attention guidance · 0.9self-attention features · 0.9reranking · 0.9numerical analysis · 0.9newton-raphson · 0.9latent arithmetic · 0.9DINOv2 · 0.9self-supervised foundation models · 0.8diffusion model · 0.8
YearPublicationVenuePosition
2025 CarGait: Cross-Attention Based Re-ranking for Gait Recognition
abstract
Gait recognition is a computer vision task that identifies individuals based on their walking patterns. Gait recognition performance is commonly evaluated by ranking a gallery of candidates and measuring the accuracy at the top Rank-$K$. Existing models are typically single-staged, i.e. searching for the probe's nearest neighbors in a gallery using a single global feature representation. Although these models typically excel at retrieving the correct identity within the top-$K$ predictions, they struggle when hard negatives appear in the top short-list, leading to relatively low performance at the highest ranks (e.g., Rank-1). In this paper, we introduce CarGait, a Cross-Attention Re-ranking method for gait recognition, that involves re-ordering the top-$K$ list leveraging the fine-grained correlations between pairs of gait sequences through cross-attention between gait strips. This re-ranking scheme can be adapted to existing single-stage models to enhance their final results. We demonstrate the capabilities of CarGait by extensive experiments on three common gait datasets, Gait3D, GREW, and OU-MVLP, and seven different gait models, showing consistent improvements in Rank-1,5 accuracy, superior results over existing re-ranking methods, and strong baselines.
Gavriel Habib, Noa Barzilay, Or Shimshi, Rami Ben-Ari, Nir Darshan
ICCV5
2025 Lightning-Fast Image Inversion and Editing for Text-to-Image Diffusion Models
abstract
Diffusion inversion is the problem of taking an image and a text prompt that describes it and finding a noise latent that would generate the exact same image. Most current deterministic inversion techniques operate by approximately solving an implicit equation and may converge slowly or yield poor reconstructed images. We formulate the problem by finding the roots of an implicit equation and devlop a method to solve it efficiently. Our solution is based on Newton-Raphson (NR), a well-known technique in numerical analysis. We show that a vanilla application of NR is computationally infeasible while naively transforming it to a computationally tractable alternative tends to converge to out-of-distribution solutions, resulting in poor reconstruction and editing. We therefore derive an efficient guided formulation that fastly converges and provides high-quality reconstructions and editing. We showcase our method on real image editing with three popular open-sourced diffusion models: Stable Diffusion, SDXL-Turbo, and Flux with different deterministic schedulers. Our solution, **Guided Newton-Raphson Inversion**, inverts an image within 0.4 sec (on an A100 GPU) for few-step models (SDXL-Turbo and Flux.1), opening the door for interactive image editing. We further show improved results in image interpolation and generation of rare objects.
Dvir Samuel, Barak Meiri, Haggai Maron, Yoad Tewel, Nir Darshan, Shai Avidan, Gal Chechik, Rami Ben-Ari
ICLR5
2025 EffoVPR: Effective Foundation Model Utilization for Visual Place Recognition
abstract
The task of Visual Place Recognition (VPR) is to predict the location of a query image from a database of geo-tagged images. Recent studies in VPR have highlighted the significant advantage of employing pre-trained foundation models like DINOv2 for the VPR task. However, these models are often deemed inadequate for VPR without further fine-tuning on VPR-specific data. In this paper, we present an effective approach to harness the potential of a foundation model for VPR. We show that features extracted from self-attention layers can act as a powerful re-ranker for VPR, even in a zero-shot setting. Our method not only outperforms previous zero-shot approaches but also introduces results competitive with several supervised methods. We then show that a single-stage approach utilizing internal ViT layers for pooling can produce global features that achieve state-of-the-art performance, with impressive feature compactness down to 128D. Moreover, integrating our local foundation features for re-ranking further widens this performance gap. Our method also demonstrates exceptional robustness and generalization, setting new state-of-the-art performance, while handling challenging conditions such as occlusion, day-night transitions, and seasonal variations.
Issar Tzachor, Boaz Lerner, Matan Levy, Tal Berkovitz Shalev, Gavriel Habib, Dvir Samuel, Noam Korngut Zailer, Or Shimshi, Nir Darshan, Rami Ben-Ari
ICLR10
2025 Find your Needle: Small Object Image Retrieval via Multi-Object Attention Optimization
abstract
We address the challenge of Small Object Image Retrieval (SoIR), where the goal is to retrieve images containing a specific small object, in a cluttered scene. The key challenge in this setting is constructing a single image descriptor, for scalable and efficient search, that effectively represents all objects in the image. In this paper, we first analyze the limitations of existing methods on this challenging task and then introduce new benchmarks to support SoIR evaluation. Next, we introduce \ours (\oursMI), a novel retrieval framework which incorporates a dedicated multi-object pre-training phase. This is followed by a refinement process that leverages attention-based feature extraction with object masks, integrating them into a single unified image descriptor. Our \oursMI approach significantly outperforms existing retrieval methods and strong baselines, achieving notable improvements in both zero-shot and lightweight multi-object fine-tuning. We hope this work will pave the way and inspire further research to enhance retrieval performance for this highly practical task.
Matan Levy, Issar Tzachor, Dvir Samuel, Nir Darshan, Rami Ben-Ari
NeurIPS5
2025 OmnimatteZero: Fast Training-free Omnimatte with Pre-trained Video Diffusion Models
abstract
In Omnimatte, one aims to decompose a given video into semantically meaningful layers, including the background and individual objects along with their associated effects, such as shadows and reflections. Existing methods often require extensive training or costly self-supervised optimization. In this paper, we present OmnimatteZero, a training-free approach that leverages off-the-shelf pre-trained video diffusion models for omnimatte. It can remove objects from videos, extract individual object layers along with their effects, and composite those objects onto new videos. These are accomplished by adapting zero-shot image inpainting techniques for video object removal, a task they fail to handle effectively out-of-the-box. To overcome this, we introduce temporal and spatial attention guidance modules that steer the diffusion process for accurate object removal and temporally consistent background reconstruction. We further show that self-attention maps capture information about the object and its footprints and use them to inpaint the object’s effects, leaving a clean background. Additionally, through simple latent arithmetic, object layers can be isolated and recombined seamlessly with new video layers to produce new videos. Evaluations show that OmnimatteZero not only achieves superior performance in terms of background reconstruction but also sets a new record for the fastest Omnimatte approach, achieving real-time performance with minimal frame runtime. Project Page.
Dvir Samuel, Matan Levy, Nir Darshan, Gal Chechik, Rami Ben-Ari
SIGGRAPH Asia3
2024 Data Roaming and Quality Assessment for Composed Image Retrieval
abstract
The task of Composed Image Retrieval (CoIR) involves queries that combine image and text modalities, allowing users to express their intent more effectively. However, current CoIR datasets are orders of magnitude smaller compared to other vision and language (V&L) datasets. Additionally, some of these datasets have noticeable issues, such as queries containing redundant modalities. To address these shortcomings, we introduce the Large Scale Composed Image Retrieval (LaSCo) dataset, a new CoIR dataset which is ten times larger than existing ones. Pre-training on our LaSCo, shows a noteworthy improvement in performance, even in zero-shot. Furthermore, we propose a new approach for analyzing CoIR datasets and methods, which detects modality redundancy or necessity, in queries. We also introduce a new CoIR baseline, the Cross-Attention driven Shift Encoder (CASE). This baseline allows for early fusion of modalities using a cross-attention module and employs an additional auxiliary task during training. Our experiments demonstrate that this new baseline outperforms the current state-of-the-art methods on established benchmarks like FashionIQ and CIRR.
Matan Levy, Rami Ben-Ari, Nir Darshan, Dani Lischinski
AAAI3
2024 Generating Images of Rare Concepts Using Pre-trained Diffusion Models
abstract
Text-to-image diffusion models can synthesize high quality images, but they have various limitations. Here we highlight a common failure mode of these models, namely, generating uncommon concepts and structured concepts like hand palms. We show that their limitation is partly due to the long-tail nature of their training data: web-crawled data sets are strongly unbalanced, causing models to under-represent concepts from the tail of the distribution. We characterize the effect of unbalanced training data on text-to-image models and offer a remedy. We show that rare concepts can be correctly generated by carefully selecting suitable generation seeds in the noise space, using a small reference set of images, a technique that we call SeedSelect. SeedSelect does not require retraining or finetuning the diffusion model. We assess the faithfulness, quality and diversity of SeedSelect in creating rare objects and generating complex formations like hand images, and find it consistently achieves superior performance. We further show the advantage of SeedSelect in semantic data augmentation. Generating semantically appropriate images can successfully improve performance in few-shot recognition benchmarks, for classes from the head and from the tail of the training data of diffusion models.
Dvir Samuel, Rami Ben-Ari, Simon Raviv, Nir Darshan, Gal Chechik
AAAI4
2024 Where's Waldo: Diffusion Features For Personalized Segmentation and Retrieval
abstract
Personalized retrieval and segmentation aim to locate specific instances within a dataset based on an input image and a short description of the reference instance. While supervised methods are effective, they require extensive labeled data for training. Recently, self-supervised foundation models have been introduced to these tasks showing comparable results to supervised methods. However, a significant flaw in these models is evident: they struggle to locate a desired instance when other instances within the same class are presented. In this paper, we explore text-to-image diffusion models for these tasks. Specifically, we propose a novel approach called PDM for Personalized Diffusion Features Matching, that leverages intermediate features of pre-trained text-to-image models for personalization tasks without any additional training. PDM demonstrates superior performance on popular retrieval and segmentation benchmarks, outperforming even supervised methods. We also highlight notable shortcomings in current instance and segmentation datasets and propose new benchmarks for these tasks.
Dvir Samuel, Rami Ben-Ari, Matan Levy, Nir Darshan, Gal Chechik
NeurIPS4
2024 Watch Where You Head: A View-biased Domain Gap in Gait Recognition and Unsupervised Adaptation
abstract
Gait Recognition is a computer vision task aiming to identify people by their walking patterns. Although existing methods often show high performance on specific datasets, they lack the ability to generalize to unseen scenarios. Unsupervised Domain Adaptation (UDA) tries to adapt a model, pre-trained in a supervised manner on a source domain, to an unlabelled target domain. There are only a few works on UDA for gait recognition proposing solutions to limited scenarios. In this paper, we reveal a fundamental phenomenon in adaptation of gait recognition models, caused by the bias in the target domain to viewing angle or walking direction. We then suggest a remedy to reduce this bias with a novel triplet selection strategy combined with curriculum learning. To this end, we present Gait Orientation-based method for Unsupervised Domain Adaptation (GOUDA). We provide extensive experiments on four widely-used gait datasets, CASIA-B, OU-MVLP, GREW, and Gait3D, and on three backbones, GaitSet, Gait-Part, and GaitGL, justifying the view bias and showing the superiority of our proposed method over prior UDA works.
Gavriel Habib, Noa Barzilay, Or Shimshi, Rami Ben-Ari, Nir Darshan
WACV5
2023 Chatting Makes Perfect: Chat-based Image Retrieval
abstract
Chats emerge as an effective user-friendly approach for information retrieval, and are successfully employed in many domains, such as customer service, healthcare, and finance. However, existing image retrieval approaches typically address the case of a single query-to-image round, and the use of chats for image retrieval has been mostly overlooked. In this work, we introduce ChatIR: a chat-based image retrieval system that engages in a conversation with the user to elicit information, in addition to an initial query, in order to clarify the user's search intent. Motivated by the capabilities of today's foundation models, we leverage Large Language Models to generate follow-up questions to an initial image description. These questions form a dialog with the user in order to retrieve the desired image from a large corpus. In this study, we explore the capabilities of such a system tested on a large dataset and reveal that engaging in a dialog yields significant gains in image retrieval. We start by building an evaluation pipeline from an existing manually generated dataset and explore different modules and training strategies for ChatIR. Our comparison includes strong baselines derived from related applications trained with Reinforcement Learning. Our system is capable of retrieving the target image from a pool of 50K images with over 78% success rate after 5 dialogue rounds, compared to 75% when questions are asked by humans, and 64% for a single shot text-to-image retrieval. Extensive evaluations reveal the strong capabilities and examine the limitations of CharIR under different settings. Project repository is available at https://github.com/levymsn/ChatIR.
Matan Levy, Rami Ben-Ari, Nir Darshan, Dani Lischinski
NeurIPS3
2023 Norm-guided latent space exploration for text-to-image generation
abstract
Text-to-image diffusion models show great potential in synthesizing a large variety of concepts in new compositions and scenarios. However, the latent space of initial seeds is still not well understood and its structure was shown to impact the generation of various concepts. Specifically, simple operations like interpolation and finding the centroid of a set of seeds perform poorly when using standard Euclidean or spherical metrics in the latent space. This paper makes the observation that, in current training procedures, diffusion models observed inputs with a narrow range of norm values. This has strong implications for methods that rely on seed manipulation for image generation, with applications to few-shot and long-tail learning tasks. To address this issue, we propose a novel method for interpolating between two seeds and demonstrate that it defines a new non-Euclidean metric that takes into account a norm-based prior on seeds. We describe a simple yet efficient algorithm for approximating this interpolation procedure and use it to further define centroids in the latent seed space. We show that our new interpolation and centroid techniques significantly enhance the generation of rare concept images. This further leads to state-of-the-art performance on few-shot and long-tail benchmarks, improving prior approaches in terms of generation speed, image quality, and semantic content.
Dvir Samuel, Rami Ben-Ari, Nir Darshan, Haggai Maron, Gal Chechik
NeurIPS3