VLDB 2026 Research / reviewers in the wild / expert
Yoad Tewel
dblp:307/5258
· DBLP profile ↗
8ranked-venue papers
4as first author
8since 2021 · last 2025
0009-0006-8042-0428ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 3 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-author · 4 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Generative modeling · 50% Vision and language · 29% 3D vision · 9% | |
| Computer graphics and multimedia
3 papers |
Visual content generation and editing · 95% Image and video processing · 5% |
Topics — the 17 heaviest of 18, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling
diffusion model |
2.5 | 3 | 2025 | Add-it: Training-Free Object Insertion in Images With Pretrained Diffusion Models · ICLR 2025 Make It Count: Text-to-Image Generation with an Accurate Number of Objects · CVPR 2025 Training-Free Consistent Text-to-Image Generation · ACM Trans. Graph. 2024 |
Visual content generation and editing
image editing |
1.7 | 2 | 2025 | Add-it: Training-Free Object Insertion in Images With Pretrained Diffusion Models · ICLR 2025 Lightning-Fast Image Inversion and Editing for Text-to-Image Diffusion Models · ICLR 2025 |
Machine learning › Generative modeling › diffusion model
text-to-image generation |
1.6 | 2 | 2025 | Make It Count: Text-to-Image Generation with an Accurate Number of Objects · CVPR 2025 Training-Free Consistent Text-to-Image Generation · ACM Trans. Graph. 2024 |
Visual content generation and editing › image generation
text-to-image generation |
1.6 | 2 | 2025 | Lightning-Fast Image Inversion and Editing for Text-to-Image Diffusion Models · ICLR 2025 Training-Free Consistent Text-to-Image Generation · ACM Trans. Graph. 2024 |
Computer vision › 3D vision › 3d scene understanding
room layout estimation |
0.9 | 1 | 2025 | Make It Count: Text-to-Image Generation with an Accurate Number of Objects · CVPR 2025 |
Machine learning › Generative modeling › diffusion model › image editing
text-guided image editing |
0.9 | 1 | 2025 | Add-it: Training-Free Object Insertion in Images With Pretrained Diffusion Models · ICLR 2025 |
Visual content generation and editing › image editing › image compositing
object insertion |
0.9 | 1 | 2025 | Add-it: Training-Free Object Insertion in Images With Pretrained Diffusion Models · ICLR 2025 |
Visual content generation and editing › image generation › text-to-image generation
consistent subject generation |
0.8 | 1 | 2024 | Training-Free Consistent Text-to-Image Generation · ACM Trans. Graph. 2024 |
Computer vision › Vision and language
image captioning |
0.6 | 1 | 2022 | ZeroCap: Zero-Shot Image-to-Text Generation for Visual-Semantic Arithmetic · CVPR 2022 |
Machine learning › Trustworthy machine learning
open-world recognition |
0.6 | 1 | 2022 | What is Where by Looking: Weakly-Supervised Open-World Phrase-Grounding without Text Inputs · NeurIPS 2022 |
Computer vision › Vision and language › visual grounding
phrase grounding |
0.6 | 1 | 2022 | What is Where by Looking: Weakly-Supervised Open-World Phrase-Grounding without Text Inputs · NeurIPS 2022 |
Computer vision › Vision and language
vision-language model |
0.6 | 1 | 2022 | ZeroCap: Zero-Shot Image-to-Text Generation for Visual-Semantic Arithmetic · CVPR 2022 |
Computer vision › Image recognition and object detection › object localization
weakly supervised object localization |
0.6 | 1 | 2022 | What is Where by Looking: Weakly-Supervised Open-World Phrase-Grounding without Text Inputs · NeurIPS 2022 |
Computer vision › Vision and language › visual grounding › phrase grounding
weakly supervised phrase grounding |
0.6 | 1 | 2022 | What is Where by Looking: Weakly-Supervised Open-World Phrase-Grounding without Text Inputs · NeurIPS 2022 |
Computer vision › Vision and language › image captioning › low-shot image captioning
zero-shot image captioning |
0.6 | 1 | 2022 | ZeroCap: Zero-Shot Image-to-Text Generation for Visual-Semantic Arithmetic · CVPR 2022 |
Image and video processing › video frame interpolation › interpolation
image interpolation |
0.3 | 1 | 2025 | Lightning-Fast Image Inversion and Editing for Text-to-Image Diffusion Models · ICLR 2025 |
Visual content generation and editing › image generation
image customization |
0.2 | 1 | 2024 | Training-Free Consistent Text-to-Image Generation · ACM Trans. Graph. 2024 |
Methods — techniques the papers use, named apart from their topics
extended attention · 1.7diffusion model · 1.7shared attention · 1.5correspondence-based feature injection · 1.5numerical analysis · 0.9newton-raphson · 0.9instance identity features · 0.9denoising guidance · 0.9visual-semantic models · 0.6large language model · 0.6contrastive learning · 0.6BLIP · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Make It Count: Text-to-Image Generation with an Accurate Number of ObjectsabstractDespite the unprecedented success of text-to-image diffusion models, controlling the number of depicted objects using text is surprisingly hard. This is important for various applications from technical documents, to children’s books to illustrating cooking recipes. Generating object-correct counts is fundamentally challenging because the generative model needs to keep a sense of separate identity for every instance of the object, even if several objects look identical or overlap, and then carry out a global computation implicitly during generation. It is still unknown if such representations exist. To address count-correct generation, we first identify features within the diffusion model that can carry the object identity information. We then use them to separate and count instances of objects during the denoising process and detect over-generation and under-generation. We fix the latter by training a model that predicts both the shape and location of a missing object, based on the layout of existing ones, and show how it can be used to guide denoising with correct object count. Our approach, CountGen, does not depend on external source to determine object layout, but rather uses the prior from the diffusion model itself, creating prompt-dependent and seed-dependent layouts. Evaluated on two benchmark datasets, we find that CountGen strongly outperforms the count-accuracy of existing baselines. Project page: https://make-it-count-paper.github.io/ Lital Binyamin, Yoad Tewel, Hilit Segev, Eran Hirsch, Royi Rassin, Gal Chechik |
CVPR | 2 |
| 2025 | Lightning-Fast Image Inversion and Editing for Text-to-Image Diffusion ModelsabstractDiffusion inversion is the problem of taking an image and a text prompt that describes it and finding a noise latent that would generate the exact same image.
Most current deterministic inversion techniques operate by approximately solving an implicit equation and may converge slowly or yield poor reconstructed images. We formulate the problem by finding the roots of an implicit equation and devlop a method to solve it efficiently. Our solution is based on Newton-Raphson (NR), a well-known technique in numerical analysis. We show that a vanilla application of NR is computationally infeasible while naively transforming it to a computationally tractable alternative tends to converge to out-of-distribution solutions, resulting in poor reconstruction and editing. We therefore derive an efficient guided formulation that fastly converges and provides high-quality reconstructions and editing. We showcase our method on real image editing with three popular open-sourced diffusion models: Stable Diffusion, SDXL-Turbo, and Flux with different deterministic schedulers. Our solution, **Guided Newton-Raphson Inversion**, inverts an image within 0.4 sec (on an A100 GPU) for few-step models (SDXL-Turbo and Flux.1),
opening the door for interactive image editing. We further show improved results in image interpolation and generation of rare objects. Dvir Samuel, Barak Meiri, Haggai Maron, Yoad Tewel, Nir Darshan, Shai Avidan, Gal Chechik, Rami Ben-Ari |
ICLR | 4 |
| 2025 | Add-it: Training-Free Object Insertion in Images With Pretrained Diffusion ModelsabstractAdding Object into images based on text instructions is a challenging task in semantic image editing, requiring a balance between preserving the original scene and seamlessly integrating the new object in a fitting location. Despite extensive efforts, existing models often struggle with this balance, particularly with finding a natural location for adding an object in complex scenes. We introduce Add-it, a training-free approach that extends diffusion models' attention mechanisms to incorporate information from three key sources: the scene image, the text prompt, and the generated image itself. Our weighted extended-attention mechanism maintains structural consistency and fine details while ensuring natural object placement. Without task-specific fine-tuning, Add-it achieves state-of-the-art results on both real and generated image insertion benchmarks, including our newly constructed "Additing Affordance Benchmark" for evaluating object placement plausibility, outperforming supervised methods. Human evaluations show that Add-it is preferred in over 80% of cases, and it also demonstrates improvements in various automated metrics. Yoad Tewel, Rinon Gal, Dvir Samuel, Yuval Atzmon, Lior Wolf, Gal Chechik |
ICLR | 1 |
| 2025 | Padding Tone: A Mechanistic Analysis of Padding Tokens in T2I ModelsabstractMichael Toker, Ido Galil, Hadas Orgad, Rinon Gal, Yoad Tewel, Gal Chechik, Yonatan Belinkov. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Michael Toker, Ido Galil, Hadas Orgad, Rinon Gal, Yoad Tewel, Gal Chechik, Yonatan Belinkov |
NAACL (Long Papers) | 5 |
| 2024 | Training-Free Consistent Text-to-Image GenerationabstractText-to-image models offer a new level of creative flexibility by allowing users to guide the image generation process through natural language. However, using these models to consistently portraythe samesubject across diverse prompts remains challenging. Existing approaches fine-tune the model to teach it new words that describe specific user-provided subjects or add image conditioning to the model. These methods require lengthy persubject optimization or large-scale pre-training. Moreover, they struggle to align generated images with text prompts and face difficulties in portraying multiple subjects. Here, we presentConsiStory, atraining-freeapproach that enables consistent subject generation by sharing the internal activations of the pretrained model. We introduce a subject-driven shared attention block and correspondence-based feature injection to promote subject consistency between images. Additionally, we develop strategies to encourage layout diversity while maintaining subject consistency. We compareConsiStoryto a range of baselines, and demonstrate state-of-the-art performance on subject consistency and text alignment, without requiring a single optimization step. Finally,ConsiStorycan naturally extend to multi-subject scenarios, and even enable training-freepersonalizationfor common objects. Yoad Tewel, Omri Kaduri, Rinon Gal, Yoni Kasten, Lior Wolf, Gal Chechik, Yuval Atzmon |
ACM Trans. Graph. | 1 |
| 2023 | Zero-Shot Video Captioning by Evolving Pseudo-tokens
Yoad Tewel, Yoav Shalev, Roy Nadler, Idan Schwartz, Lior Wolf |
BMVC | 1 |
| 2022 | ZeroCap: Zero-Shot Image-to-Text Generation for Visual-Semantic ArithmeticabstractRecent text-to-image matching models apply contrastive learning to large corpora of uncurated pairs of images and sentences. While such models can provide a powerful score for matching and subsequent zero-shot tasks, they are not capable of generating caption given an image. In this work, we repurpose such models to generate a descriptive text given an image at inference time, without any further training or tuning step. This is done by combining the visual-semantic model with a large language model, benefiting from the knowledge in both web-scale models. The resulting captions are much less restrictive than those obtained by supervised captioning methods. Moreover, as a zero-shot learning method, it is extremely flexible and we demonstrate its ability to perform image arithmetic in which the inputs can be either images or text and the output is a sentence. This enables novel high-level vision capabilities such as comparing two images or solving visual analogy tests. Our code is available at: https://github.com/YoadTew/zero-shot-image-to-text. Yoad Tewel, Yoav Shalev, Idan Schwartz, Lior Wolf |
CVPR | 1 |
| 2022 | What is Where by Looking: Weakly-Supervised Open-World Phrase-Grounding without Text InputsabstractGiven an input image, and nothing else, our method returns the bounding boxes of objects in the image and phrases that describe the objects. This is achieved within an open world paradigm, in which the objects in the input image may not have been encountered during the training of the localization mechanism. Moreover, training takes place in a weakly supervised setting, where no bounding boxes are provided. To achieve this, our method combines two pre-trained networks: the CLIP image-to-text matching score and the BLIP image captioning tool. Training takes place on COCO images and their captions and is based on CLIP. Then, during inference, BLIP is used to generate a hypothesis regarding various regions of the current image. Our work generalizes weakly supervised segmentation and phrase grounding and is shown empirically to outperform the state of the art in both domains. It also shows very convincing results in the novel task of weakly-supervised open-world purely visual phrase-grounding presented in our work.For example, on the datasets used for benchmarking phrase-grounding, our method results in a very modest degradation in comparison to methods that employ human captions as an additional input. Tal Shaharabany, Yoad Tewel, Lior Wolf |
NeurIPS | 2 |