Tal Shaharabany

dblp:255/5390 · DBLP profile ↗
← Back
9ranked-venue papers
5as first author
8since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 7 since 2021Artificial intelligence and machine learning · 5 · 3 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2025 Classifier-Guided Captioning Across Modalities
abstract
Most current captioning systems use language models trained on data from specific settings, such as image-based captioning via Amazon Mechanical Turk, limiting their ability to generalize to other modality distributions and contexts. This limitation hinders performance in tasks like audio or video captioning, where different semantic cues are needed. Addressing this challenge is crucial for creating more adaptable and versatile captioning frameworks applicable across diverse real-world contexts. In this work, we introduce a method to adapt captioning networks to the semantics of alternative settings, such as capturing audibility in audio captioning, where it is crucial to describe sounds and their sources. Our framework consists of two main components: (i) a frozen captioning system incorporating a language model (LM), and (ii) a text classifier that guides the captioning system. The classifier is trained on a dataset automatically generated by GPT-4, using tailored prompts specifically designed to enhance key aspects of the generated captions. Importantly, the framework operates solely during inference, eliminating the need for further training of the underlying captioning model. We evaluated the framework on various models and modalities, with a focus on audio captioning, and report promising results. Notably, when combined with an existing zero-shot audio captioning system, our framework improves its quality and sets state-of-the-art performance in zero-shot audio captioning.
Ariel Shaulov, Tal Shaharabany, Eitan Shaar, Gal Chechik, Lior Wolf
ICASSP2
2023 AutoSAM: Adapting SAM to Medical Images by Overloading the Prompt Encoder
Tal Shaharabany, Aviad Dahan, Raja Giryes, Lior Wolf
BMVC1
2023 Similarity Maps for Self-Training Weakly-Supervised Phrase Grounding
abstract
A phrase grounding model receives an input image and a text phrase and outputs a suitable localization map. We present an effective way to refine a phrase ground model by considering self-similarity maps extracted from the latent representation of the model's image encoder. Our main insights are that these maps resemble localization maps and that by combining such maps, one can obtain useful pseudo-labels for performing self-training. Our results surpass, by a large margin, the state of the art in weakly supervised phrase grounding. A similar gap in performance is obtained for a recently proposed downstream task called WWbL, in which only the image is input, without any text. Our code is available at https://github.com/talshaharabany/Similarity-Maps-for-Self-Training-Weakly-Supervised-Phrase-Grounding
Tal Shaharabany, Lior Wolf
CVPR1
2023 Learning a Weight Map for Weakly-Supervised Localization
abstract
In the weakly supervised localization setting, supervision is given as an image-level label. We propose employing an image classifier f and training a generative network g that outputs, given the input image, a per-pixel weight map that indicates the location of the object within the image. Network g is trained by minimizing the discrepancy between the output of the classifier f on the original image and its output given the same image weighted by the output of g. Our results indicate that the method outperforms existing localization methods on the challenging fine-grained classification datasets.
Tal Shaharabany, Lior Wolf
ICASSP1
2023 Box-based Refinement for Weakly Supervised and Unsupervised Localization Tasks
abstract
It has been established that training a box-based detector network can enhance the localization performance of weakly supervised and unsupervised methods. Moreover, we extend this understanding by demonstrating that these detectors can be utilized to improve the original network, paving the way for further advancements. To accomplish this, we train the detectors on top of the network output instead of the image data and apply suitable loss backpropagation. Our findings reveal a significant improvement in phrase grounding for the "what is where by looking" task, as well as various methods of unsupervised object discovery. Our code is available at https://github.com/eyalgomel/box-based-refinement.
Eyal Gomel, Tal Shaharabany, Lior Wolf
ICCV2
2023 Annotator Consensus Prediction for Medical Image Segmentation with Diffusion Models
Tomer Amit, Shmuel Shichrur, Tal Shaharabany, Lior Wolf
MICCAI (4)3
2022 End-to-End Segmentation of Medical Images via Patch-Wise Polygons Prediction
Tal Shaharabany, Lior Wolf
MICCAI (5)1
2022 What is Where by Looking: Weakly-Supervised Open-World Phrase-Grounding without Text Inputs
abstract
Given an input image, and nothing else, our method returns the bounding boxes of objects in the image and phrases that describe the objects. This is achieved within an open world paradigm, in which the objects in the input image may not have been encountered during the training of the localization mechanism. Moreover, training takes place in a weakly supervised setting, where no bounding boxes are provided. To achieve this, our method combines two pre-trained networks: the CLIP image-to-text matching score and the BLIP image captioning tool. Training takes place on COCO images and their captions and is based on CLIP. Then, during inference, BLIP is used to generate a hypothesis regarding various regions of the current image. Our work generalizes weakly supervised segmentation and phrase grounding and is shown empirically to outperform the state of the art in both domains. It also shows very convincing results in the novel task of weakly-supervised open-world purely visual phrase-grounding presented in our work.For example, on the datasets used for benchmarking phrase-grounding, our method results in a very modest degradation in comparison to methods that employ human captions as an additional input.
Tal Shaharabany, Yoad Tewel, Lior Wolf
NeurIPS1
2020 End to End Trainable Active Contours via Differentiable Rendering
Shir Gur, Tal Shaharabany, Lior Wolf
ICLR2