VLDB 2026 Research / reviewers in the wild / expert
Noam Rotstein
dblp:304/2369
· DBLP profile ↗
6ranked-venue papers
3as first author
6since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CaricatureGS: Exaggerating 3D Gaussian Splatting Faces with Gaussian CurvatureabstractA photorealistic and controllable 3D caricaturization framework for faces is introduced. We start with an intrinsic Gaussian curvature-based surface exaggeration technique, which, when coupled with texture, tends to produce oversmoothed renders. To address this, we resort to 3D Gaussian Splatting (3DGS), which has recently been shown to produce realistic free-viewpoint avatars. Given a multiview sequence, we extract a FLAME mesh, solve a curvatureweighted Poisson equation, and obtain its exaggerated form. However, directly deforming the Gaussians yields poor results, necessitating the synthesis of pseudo–groundtruth caricature images by warping each frame to its exaggerated 2D representation using local affine transformations. We then devise a training scheme that alternates real and synthesized supervision, enabling a single Gaussian collection to represent both natural and exaggerated avatars. This scheme improves fidelity, supports local edits, and allows continuous control over the intensity of the caricature. In order to achieve real-time deformations, an efficient interpolation between the original and exaggerated surfaces is introduced. We further analyze and show that it has a bounded deviation from closed-form solutions. In both quantitative and qualitative evaluations, our results outperform prior work, delivering photorealistic, geometry-controlled caricature avatars. Project page: https://c4ricaturegs.github.io Eldad Matmon, Amit Bracha, Noam Rotstein, Ron Kimmel |
3DV | 3 |
| 2025 | Pathways on the Image Manifold: Image Editing via Video GenerationabstractRecent advances in image editing, driven by image diffusion models, have shown remarkable progress. However, significant challenges remain, as these models often struggle to follow complex edit instructions accurately and frequently compromise fidelity by altering key elements of the original image. Simultaneously, video generation has made remarkable strides, with models that effectively function as consistent and continuous world simulators. In this paper, we propose merging these two fields by utilizing image-to-video models for image editing. We reformulate image editing as a temporal process, using pretrained video models to create smooth transitions from the original image to the desired edit. This approach traverses the image manifold continuously, ensuring consistent edits while preserving the original image’s key aspects. Our approach achieves stateof-the-art results on text-based image editing, demonstrating significant improvements in both edit accuracy and image preservation. Visit our project page. Noam Rotstein, Gal Yona, Daniel Silver, Roy Velich, David Bensaïd, Ron Kimmel |
CVPR | 1 |
| 2025 | Paint by Inpaint: Learning to Add Image Objects by Removing Them FirstabstractImage editing has advanced significantly with the introduction of text-conditioned diffusion models. Despite this progress, seamlessly adding objects to images based on textual instructions without requiring user-provided input masks remains a challenge. We address this by leveraging the insight that removing objects (Inpaint) is significantly simpler than its inverse process of adding them (Paint), attributed to inpainting models that benefit from segmentation mask guidance. Capitalizing on this realization, by implementing an automated and extensive pipeline, we curate a filtered large-scale image dataset containing pairs of images and their corresponding object-removed versions. Using these pairs, we train a diffusion model to inverse the inpainting process, effectively adding objects into images. Unlike other editing datasets, ours features natural target images instead of synthetic ones while ensuring source-target consistency by construction. Additionally, we utilize a large Vision-Language Model to provide detailed descriptions of the removed objects and a Large Language Model to convert these descriptions into diverse, natural-language instructions. Our quantitative and qualitative results show that the trained model surpasses existing models in both object addition and general editing tasks. Visit our project page for the released dataset and trained models. Navve Wasserman, Noam Rotstein, Roy Ganz, Ron Kimmel |
CVPR | 2 |
| 2024 | FuseCap: Leveraging Large Language Models for Enriched Fused Image CaptionsabstractThe advent of vision-language pre-training techniques enhanced substantial progress in the development of models for image captioning. However, these models frequently produce generic captions and may omit semantically important image details. This limitation can be traced back to the image-text datasets; while their captions typically offer a general description of image content, they frequently omit salient details. Considering the magnitude of these datasets, manual reannotation is impractical, emphasizing the need for an automated approach. To address this challenge, we leverage existing captions and explore augmenting them with visual details using "frozen" vision experts including an object detector, an attribute recognizer, and an Optical Character Recognizer (OCR). Our proposed method, FuseCap, fuses the outputs of such vision experts with the original captions using a large language model (LLM), yielding comprehensive image descriptions. We automatically curate a training set of 12M image-enriched caption pairs. These pairs undergo extensive evaluation through both quantitative and qualitative analyses. Subsequently, this data is utilized to train a captioning generation BLIP-based model. This model outperforms current state-of-the-art approaches, producing more precise and detailed descriptions, demonstrating the effectiveness of the proposed data-centric approach. We release this large-scale dataset of enriched image-caption pairs for the community. Noam Rotstein, David Bensaïd, Shaked Brody, Roy Ganz, Ron Kimmel |
WACV | 1 |
| 2023 | Partial Matching of Nonrigid Shapes by Learning Piecewise Smooth FunctionsabstractAbstract Learning functions defined on non‐flat domains, such as outer surfaces of non‐rigid shapes, is a central task in computer vision and geometry processing. Recent studies have explored the use of neural fields to represent functions like light reflections in volumetric domains and textures on curved surfaces by operating in the embedding space. Here, we choose a different line of thought and introduce a novel formulation of partial shape matching by learning a piecewise smooth function on a surface. Our method begins with pairing sparse landmarks defined on a full shape and its part, using feature similarity. Next, a neural representation is optimized to fit these landmarks, efficiently interpolating between the matched features that act as anchors. This process results in a function that accurately captures the partiality. Unlike previous methods, the proposed neural model of functions is intrinsically defined on the given curved surface, rather than the classical embedding Euclidean space. This representation is shown to be particularly well‐suited for representing piecewise smooth functions. We further extend the proposed framework to the more challenging part‐to‐part setting, where both shapes exhibit missing parts. Comprehensive experiments highlight that the proposed method effectively addresses partiality in shape matching and significantly outperforms leading state‐of‐the‐art methods in challenging benchmarks. Code is available at https://github.com/davidgip74/Learning-Partiality-with-Implicit-Intrinsic-Functions David Bensaïd, Noam Rotstein, Nelson Goldenstein, Ron Kimmel |
Comput. Graph. Forum | 2 |
| 2022 | Multimodal Colored Point Cloud to Image AlignmentabstractReconstruction of geometric structures from images using supervised learning suffers from limited available amount of accurate data. One type of such data is accurate real-world RGB-D images. A major challenge in acquiring such ground truth data is the accurate alignment between RGB images and the point cloud measured by a depth scanner. To overcome this difficulty, we consider a differential optimization method that aligns a colored point cloud with a given color image through iterative geometric and color matching. In the proposed framework, the optimization minimizes the photometric difference between the colors of the point cloud and the corresponding colors of the image pixels. Unlike other methods that try to reduce this photometric error, we analyze the computation of the gradient on the image plane and propose a different direct scheme. We assume that the colors produced by the geometric scanner camera and the color camera sensor are different and therefore characterized by different chromatic acquisition properties. Under these multimodal conditions, we find the transformation between the camera image and the point cloud colors. We alternately optimize for aligning the position of the point cloud and matching the different color spaces. The alignments produced by the proposed method are demonstrated on both synthetic data with quantitative evaluation and real scenes with qualitative results. Noam Rotstein, Amit Bracha, Ron Kimmel |
CVPR | 1 |